EconBase
← Back to paper

Counterfactual Inference in Duration Models with Random Censoring

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

56,045 characters · 9 sections · 76 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Counterfactual Inference in Duration Models with Random Censoring

abstractWe propose a counterfactual Kaplan-Meier estimator that incorporates exogenous covariates and unobserved heterogeneity of unrestricted dimensionality in duration models with random censoring. Under some regularity conditions, we establish the joint weak convergence of the proposed counterfactual estimator and the unconditional Kaplan-Meier (KaplanMeier1958) estimator. Applying the functional delta method, we make inference on the cumulative hazard policy effect, that is, the change of duration dependence in response to a counterfactual policy. We also evaluate the finite sample performance of the proposed counterfactual estimation method in a Monte Carlo study. Keywords: Counterfactual policy effect, Random censoring, Duration analysis, Kaplan-Meier estimator JEL Classification: C14, C41, C24

\fontsize{12}{18pt}\selectfont

Introduction

Policy evaluation is one of the important areas in social science. Counterfactual analysis is an approach that provides policy recommendations when a policy is not implemented yet or when a similar quasi-experiment is infeasible. Recent studies on counterfactual analysis, for example Rothe2010 and ChernozhukovFernandez-ValEtAl2013a, emphasize the unconditional distributional effect of an exogenous manipulation of covariates on an outcome variable of interest. Methods in these studies are usually based on data that are completely observed; however, sampling schemes may generate incomplete data and thus restrict their applicability. For example, duration data on unemployment spells, collected by the Current Population Survey, are commonly believed to be subject to right censoring, as explained by Kiefer1988.

The main objectives of this paper are to estimate the unconditional distribution of a duration variable affected by a counterfactual policy that exogenously manipulates covariates, and to evaluate associated policy effects by the comparison between the counterfactual and original unconditional distribution of the duration variable. Specifically, we consider a nonseparable model

align[align omitted — 58 chars of source]

where $T$ is a nonnegative duration variable of interest, $X$ is a $d$-dimensional vector of time-invariant covariates, $\varepsilon$ is individual unobserved heterogeneity in an arbitrary measurable space of unrestricted dimensionality, and $\varphi$ is a structural function that is unknown to researchers. In addition, the right censoring may make $T$ unobserved; instead, the observable data are the vector $X$ of covariates,

align[align omitted — 95 chars of source]

where $C$ is a censoring random variable, which is only observed for censored observations, and $\mathbbm{1}_{[\cdot]}$ is an indicator function. Policy makers consider the counterfactual scenario that exogenously changes $X$ to $X^{*}$ and leads to the counterfactual duration variable

align[align omitted — 68 chars of source]

and attempt to evaluate the policy effect

align*[align* omitted — 58 chars of source]

where $F_{T}$ and $F_{T^{*}}$ are the cumulative distribution functions (CDFs) of $T$ and $T^{*}$, respectively, and $\nu$ is some functional defined on the collection of all CDFs. Such an effect $\nu(F_{T^{*}})-\nu(F_{T})$ may be important in policy evaluation. For example, although unemployment insurance benefits may smooth the income fluctuation of the unemployed, it may discourage the unemployed from searching for jobs. Therefore, policy makers would be interested in the effect of reducing wage replacement ratio on the cumulative hazard rate of unemployment spells.

commentFor example, although unemployment insurance benefits may improve workers' ability to avoid income fluctuation, it reduces the willingness of the unemployed to search for employment as suggested by search theory. Therefore, a policy maker would be interested in evaluating the effects on the unemployment spell of different ways that decrease the unemployment insurance benefits.

In this case, $T$ is the unemployment duration, $X$ is the wage replacement ratio, and the functional $\nu$ is a map from a CDF to its cumulative hazard function, that is, $\nu:F\mapsto \int_{[0,\cdot]} \frac{1}{1-F^{-}}\;\mathrm{d} F$.\footnote{ For any c\`{a}dl\`{a}g function $F$, we write $F^{-}$ for its left-continuous version, that is, $F^{-}(t)\equiv \lim_{s\uparrow t}F(s)$. }

We propose a nonparametric estimation method of the unconditional CDF $F_{T^{*}}$ arising from an exogenous manipulation of covariates $X$ on the duration variable $T$. This proposed nonparametric estimation method allows researchers to conduct a counterfactual policy analysis, rather than just a descriptive analysis, of duration data.\footnote{ Lancaster1992 and CameronTrivedi2005 indicate that the nonparametric Kaplan-Meier estimation is traditionally viewed as a descriptive analysis. } Specifically, we evaluate $\nu(F_{T^{*}})-\nu(F_{T})$ by replacing $F_{T^{*}}$ and $F_{T}$ with their nonparametric estimator, respectively. On the one hand, we construct the unconditional Kaplan-Meier (KaplanMeier1958) estimator $\hat{F}_{T}$. On the other hand, under regularity conditions, the unconditional CDF of $T^{*}$ is recovered by $F_{T^{*}}(t)=\mathbb{E}(F_{T|X}(t|X^{*}))$ where $F_{T|X}$ is the conditional CDF of $T$ given $X$; thus, we follow the analogy principle to propose a two-stage fully nonparametric estimator of the counterfactual CDF of $T^{*}$. In particular, we first construct a variant of Beran1981's (Beran1981) conditional Kaplan-Meier estimator $\hat{F}_{T|X}$ and then take average of $\hat{F}_{T|X}$ with respect to the empirical distribution of $X^{*}$ to obtain a counterfactual estimator $\hat{F}_{T^{*}}$. The first-stage nonparametric estimation can avoid the misspecification of the conditional CDF, which is emphasized in RotheWied2013. Moreover, the proposed first-stage estimator, instead of the kernel CDF estimator in Rothe2010, is essential to avoid the estimation bias in the presence of censoring. Indeed, our simulation experiments show that the proposed estimator $\hat{F}_{T^{*}}$, compared with Rothe's counterfactual CDF estimator, has smaller mean integrated absolute error (MIAE) and root mean integrated squared error (RMISE) when duration data are subject to censoring.

To establish the validity of the proposed approach, we show that under some regularity conditions, the vector $(\hat{F}_{T^{*}}-F_{T^{*}},\hat{F}_{T}-F_{T})^{\top}$ converges weakly to a two dimensional centered Gaussian process over a specific compact subset of $\mathbb{R}^{2}_{+}$ at the rate $\sqrt{n}$. This convergence rate avoids the curse of dimensionality even though the first-stage estimator $\hat{F}_{T|X}$ converges at a rate less than $\sqrt{n}$. Applying the functional delta method, we can obtain the asymptotic distribution of the counterfactual policy effect $\sqrt{n}(\nu(\hat{F}_{T^{*}})-\nu(\hat{F}_{T}))$, which allows us to evaluate the effect of counterfactual policy intervention on the cumulative hazard function.

The proposed method also complements the literature on decomposition methods. Decomposition methods are usually used to explain the difference in unconditional distributional features of an outcome variable across two different demographic groups or time periods. The between-group difference is usually decomposed into a structure effect and a composition effect.\footnote{ A structure effect arises because structural functions are different between two groups, and a composition effect reflects the differences in covariates between two groups. The early development in decomposition methods is well surveyed by FortinLemieuxEtAl2011. Recently, Rothe2015 further investigates a detailed decomposition, which attributes the composition effect to each covariate. } In the presence of complete data, Rothe2010 proposes a two-stage fully nonparametric estimation of the composition effect, whereas ChernozhukovFernandez-ValEtAl2013a develop a two-stage semiparametric estimation of this effect by either distribution regression or quantile regression. Taking random censoring into account, Garcia-Suaza2016 studies the effect based on the proportional hazard specification. In contrast, the method in this paper is fully nonparametric; additionally, as explained by Rothe2010, we can regard $X^{*}$ as observable covariates of a different group, and $\nu(F_{T^{*}})-\nu(F_{T})$ as the composition effect in the setup ((ref))-((ref)) of random censoring.

Throughout this paper, all random variables are defined on the same probability space $(\Omega,\mathcal{A},\mathbb{P})$. We denote $\mathbb{D}[c_{1},c_{2}]$ and $l^{\infty}[c_{1},c_{2}]$ by the set of c\`{a}dl\`{a}g and bounded functions defined on the interval $[c_{1},c_{2}]$, respectively. We write $\Rightarrow$ for weak convergence in a function space equipped with the uniform norm, and $a\wedge b$ for the minimum of $a$ and $b$. We also denote the density of $X$ by $m$, and the density of $X^{*}$ by $m^{*}$. For a generic random variable $U$, we write $F_{U}$ for the CDF of $U$, $f_{U}$ for the derivative of $F_{U}$, $F_{U|X}$ for the conditional CDF of $U$ given $X$, and $f_{U|X}(u|x)$ for the derivative with respect to $u$ of $F_{U|X}(u|x)$; additionally, let $F^{\delta}_{U}(u)=\mathbb{P}(U\leq u, \delta=1)$ and $F^{\delta}_{U|X}(u|x)=\mathbb{P}(U\leq u, \delta=1|X=x)$. We assume that $T,C,X$, and $X^{*}$ are absolutely continuous random variables. The absolute continuity of the duration variable $T$ is reasonable because $T$ is expected to be generated by a transition process, which is usually modeled in continuous time. (See CameronTrivedi2005 and FlorensFougereEtAl2008 for example.) Furthermore, we only consider absolutely continuous covariates for ease of exposition because the proposed estimation method can be revised to include discrete covariates.

commentWhen the covariates are mixed, say $x=(x^{c},x^{d})$ with continuous $x^{c}$ and discrete $x^{d}$, we consider the weights \begin{align*} B_{n_{j}}(x;h)=\frac{K\left(\frac{x^{c}-X^{c}_{j}}{h}\right)\mathbbm{1}_{[x^{d}=X^{d}_{j}]}} {\sum_{k=1}^{n}K\left(\frac{x^{c}-X^{c}_{k}}{h}\right)\mathbbm{1}_{[x^{d}=X^{d}_{k}]}} \quad\quad j=1,2,\cdots,n. \end{align*}

Alternatively, in the case of a binary policy variable, SantAnna2016 extends the method of Kaplan-Meier integrals and studies various treatment effects when the outcome may be right censored.

The remainder of this paper is organized as follows. Section (ref) discusses the setup of duration analysis, the objects of interest, and the counterfactual Kaplan-Meier estimator. Section (ref) shows the asymptotic theory of the proposed estimator and statistical inference on the associated policy effects. Section (ref) presents the results of Monte Carlo simulation. Section (ref) concludes. Technical proofs are deferred to Appendix.

Model and estimation

Setup and objects of interest

The flexible duration model in ((ref)) can avoid several types of model misspecification. First, LuWhite2014 point out that the nonseparability of $\varepsilon$ enables treatment effect and marginal effect could depend on unobservable heterogeneity. In addition, the unrestricted dimensionality of $\varepsilon$ can avoid incorrect inference due to the inclusion of limited heterogeneity, as argued by BrowningCarro2007. HoderleinMammen2009 also indicate that $\varepsilon$ can be viewed as an element of an infinitely dimensional function space; for example, it could be individual preference for leisure in the analysis of unemployment spells. Finally, both the marginal distribution of $\varepsilon$ and the conditional distribution of $T$ given $\varepsilon$ are unspecified to avoid inappropriate inference caused by parametric assumptions.\footnote{ Lancaster1992 documents many alternatives of parametric assumptions about the hazard function in mixture models. Parametric specification of the heterogeneity distribution and duration dependence is, however, a well-known issue in econometrics. See the discussion in HausmanWoutersen2014. }

The random censoring feature in ((ref)) is prevalent in duration analysis and usually arises because of sampling schemes, for example, a random failure to follow up an individual during the study period. We refer readers to Moore2016 for more underlying reasons of random censoring. In this paper, we consider the simple counterfactual scenario that policy intervention does not affect the structural function $\varphi$ in ((ref)); however, we allow a change in the censoring variable $C$ after policy intervention.\footnote{ As suggested in FortinLemieuxEtAl2011, policy intervention may result in an alternative structural function $\varphi^{*}$ in general equilibrium. }

A counterfactual policy that changes the duration variable from $T$ to $T^{*}$ yields the distribution policy effect

align*[align* omitted — 59 chars of source]

For instance, since the shape of the distribution of unemployment spells may affect the government expenditures on unemployment insurance, the distribution policy effect matters for policy makers concerning fiscal deficits. In addition, the literature on duration models especially emphasizes the duration dependence, that is, the shape of the hazard function.\footnote{ The hazard function of a nonnegative duration variable $T$ is defined as

align*[align* omitted — 90 chars of source]

} The change of duration dependence in response to a counterfactual policy can be answered by the cumulative hazard policy effect

align*[align* omitted — 76 chars of source]

where $\Lambda_{T^{*}}(t)=\int_{0}^{t}\frac{F_{T^{*}}(\!\!\;\mathrm{d} u)}{1-F_{T^{*}}^{-}(u)}$ and $\Lambda_{T}(t)=\int_{0}^{t}\frac{F_{T}(\!\!\;\mathrm{d} u)}{1-F_{T}^{-}(u)}$ are the cumulative hazard functions of $T^{*}$ and $T$, respectively. In the case of unemployment spells, policy makers would be interested in the cumulative hazard policy effect because it would evaluate whether a counterfactual policy is beneficial for a target group, for example the long-term unemployed, to escape the unemployment trap. Other counterfactual policy effects, such as quantile policy effect and Lorenz curve policy effect, may also be of interest. See Bhattacharya2007 and Rothe2010 for treatment of these and further examples. When the objects of interest are the aforementioned policy effects, the identification of $\varphi$ is not necessary, as indicated by Rothe2010; thus, we maintain the flexible specification of the structural function in ((ref)).\footnote{ Matzkin2003 provides conditions such that in the absence of censoring, the structural function $\varphi$ can be identified if $\varphi(x,\varepsilon)$ is strictly increasing in unobserved scalar heterogeneity $\varepsilon$ for each $x$. }

Nonparametric identification and estimation

Since a counterfactual policy effect can be generally written as $\nu(F_{T^{*}})-\nu(F_{T})$ for some specific functional $\nu$, we start by identifying the CDFs $F_{T}$ and $F_{T^{*}}$. We first introduce the following assumptions.\\[0.2cm] Assumption D (Data)

enumerate[label=D\arabic*] • Both $\{(Y_{i},T_{i},C_{i},\delta_{i},X_{i})\}_{i=1}^{n}$ and $\{X_{j}^{*}\}_{j=1}^{n^{*}}$ are independent and identically distributed across individuals. • (i) $\{(Y_{i},\delta_{i},X_{i})\}_{i=1}^{n}$ are observable; (ii) $\{X_{j}^{*}\}_{j=1}^{n^{*}}$ are observable and $n^{*}=n$.

Assumption (ref) is common in models of cross-sectional data. Assumption (ref)(i) is also common in duration models where researchers know whether the observed duration variable is censored. Assumption (ref)(ii) is innocuous when we consider the counterfactual policy that shifts $X$ to $X^{*}=\pi(X)$ for some measurable function $\pi$, whereas this assumption is imposed for convenience when we treat $X^{*}$ as observable covariates of a different group in the analysis of the composition effect.\\

Assumption I (Identification)

enumerate[label=I\arabic*] • $T$ and $C$ are independent. • $T$ and $C$ are conditionally independent given $X$. • $\varepsilon$ is independent of both $X$ and $X^{*}$. • The support of $X^{*}$ is a subset of the support of $X$.

Assumptions (ref) and (ref) are commonly imposed in survival analysis, for example Lancaster1992 and KalbfleischPrentice2002 for Assumption (ref) and Dabrowska1989, Iglesias-PerezGonzalez-Manteiga1999, and Gneyou2014 for Assumption (ref). Assumption (ref) ensures that the random censoring is non-informative; that is, $C$ does not provide any information about $T$, and vice versa.

comment\textcolor{red}{Page 99 of KleinMoeschberger2003 “Both the Nelson-Aalen and Product-Limit estimator are based on an assumption of non-informative censoring which means that knowledge of a censoring time for an individual provides no further information about this person’s likelihood of survival at a future time had the individual continued on the study.”}

Assumption (ref) holds if random censoring is non-informative when covariates are controlled. Note that neither Assumption (ref) nor Assumption (ref) is stronger.\footnote{ See the examples on page 65 of Stoyanov2014.} If $C$ is independent of $(X,T)$, then Assumptions (ref) and (ref) are satisfied; however, these two assumptions are not sufficient for the independence between $C$ and $X$.\footnote{ The independence between $C$ and $X$ holds if Assumptions (ref) and (ref) hold and the family of distributions of $T$ given $X$ is boundedly complete; that is, for a bounded function $g$, $\operatorname*{\mathbb{E}}[g(T)|X]=0$ almost surely implies $g(T)=0$ almost surely. See the discussion in Dawid1998. } Assumptions (ref) and (ref) are imposed for counterfactual analysis as in Rothe2010. Assumption (ref) requires that all covariates are exogenous and thus may be strong in some empirical studies. If we are interested in the effects arising from the manipulation of some policy variables, this assumption can be weaken by conditional exogeneity of $\varepsilon$ given observable covariates. To be precise, let $X=(X_{\text{p}},X_{\text{c}})^{\top}$ where $X_{\text{p}}$ and $X_{\text{c}}$ are the vector of policy variables and vector of covariates, respectively. Policy intervention changes $X_{\text{p}}$ to $X^{*}_{\text{p}}$ but keeps $X_{\text{c}}$ unchanged. Proposition (ref) below is still valid if Assumption (ref) is replaced with the assumption that $\varepsilon$ is independent of $(X_{\text{p}},X^{*}_{\text{p}})$ conditional on $X_{\text{c}}$.\footnote{ Similarly, Assumptions (ref) can be relaxed by the control function approach proposed by BlundellPowell2003 and ImbensNewey2009. The application of the control function approach is however beyond the scope this paper. See Lee2015 for the analysis of counterfactual effects by the control function approach in the absence of random censoring. } Assumptions (ref) and (ref) imply that censoring occurs exogenously provided that $\varphi(x,\cdot)$ is invertible for all $x$.\footnote{ Suppose that invertibility of $\varphi(x,\cdot)$ holds for all $x$. Assumption (ref) implies that $\varepsilon$ and $C$ are conditionally independent given $X$ by Lemmas 4.1 and 4.2 of Dawid1979. If moreover Assumption (ref) holds, then $\varepsilon$ is independent of $(C,X)$ by Lemma 4.2 of Dawid1979. } Since nonparametric analyses of counterfactual policy effects resulting from an extrapolation of covariates may be invalid, we impose the overlap condition in Assumption (ref).

Assumption I guarantees the identification of $(F_{T^{*}},F_{T})^{\top}$ over a subset of $\mathbb{R}_{+}^{2}$. StuteWang1993 show that under Assumption (ref), $F_{T}(t)$ is identified for each $t<\tau\equiv\inf\{t:F_{Y}(t)=1\}$. Under Assumptions (ref) and (ref), we can express $F_{T^{*}}$ as the population average, taken with respect to the distribution of $X^{*}$, of the conditional CDF of $T$ given $X$; to be precise, $F_{T^{*}}(t)=\mathbb{E}(F_{T|X}(t|X^{*}))$. The identification of $F_{T^{*}}(t)$ thus follows that of $F_{T|X}(t|x)$, and the latter is achieved under Assumption (ref) for $(t,x)\in[0,\tau)\times\mathbb{R}^{d}$. Alternatively, the joint CDF of $(T,X)$ can be identified on $[0,\tau)\times\mathbb{R}^{d}$ by replacing Assumption (ref) with the assumption that $\delta$ and $X$ are conditionally independent given $T$; that is, given the duration, covariates provide no further information on whether censoring occurs.\footnote{ Suppose that Assumption (ref) holds. If $\delta$ and $X$ are conditionally independent given $T$, then the joint distribution of $(T,X)$ on $[0,\tau)\times \mathbb{R}^{d}$ can be recovered by

align*[align* omitted — 223 chars of source]

where $H_{Y}^{0}(y)=\mathbb{P}(Y\leq y, \delta=0)$ and $H_{YX}^{1}(y,x)=\mathbb{P}(Y\leq y, X\leq x,\delta=1)$. The availability of data on $\{(Y_{i},\delta_{i},X_{i})\}_{i=1}^{n}$ implies that $F_{T|X}(t|x)$ is identified for $(t,x)\in [0,\tau)\times \mathbb{R}^{d}$. More general results are shown in Equation (1.2) of Stute1996. } This assumption is imposed in recent studies on duration analysis, for example SantAnna2016 (SantAnna2016, SantAnna2017), Garcia-Suaza2016, and references cited therein. We summarize the discussion in the following proposition.

ProSuppose that Assumptions (ref) and (ref) hold. Under Assumption (ref), $F_{T^{*}}(t)$ is identified for $t \in[0,\tau)$. If in addition Assumption (ref) holds, then $F_{T|X}(t|x)$ is identified for $(t,x)\in [0,\tau)\times\mathbb{R}^{d}$. Moreover, if Assumptions (ref) and (ref) are also satisfied, we have $F_{T^{*}}(t)=\mathbb{E}(F_{T|X}(t|X^{*}))$ for $t\in [0,\tau)$.\qed

Proposition (ref) suggests that we follow the analogy principle to construct an estimator of $F_{T^{*}}(t)$ by

equation[equation omitted — 116 chars of source]

where $\hat{F}_{T|X}$ is the variant of Beran1981's (Beran1981) conditional Kaplan-Meier estimator, that is

equation[equation omitted — 220 chars of source]

$\{B_{n_{\ell}}(x;h)\}_{\ell=1}^{n}$ are appropriate weights and $h$ is a tuning parameter.\footnote{ In fact, the estimator $\hat{F}_{T|X}$ is the exponential transformation of the Nalson-Aalen estimator of the cumulative hazard function of $F_{T|X}$. Additionally, Beran1981's (Beran1981) conditional Kaplan-Meier estimator

align*[align* omitted — 216 chars of source]

can be viewed as the first-order Taylor series approximation of $\hat{F}_{T|X}$. } Different choices of weights are documented in the literature on the conditional Kaplan-Meier estimator.\footnote{ These choices include Gasser-Muller weights and Nadaraya-Watson weights. See for example Gonzalez-ManteigaCadarso-Suarez1994, Dabrowska1989, and Iglesias-PerezGonzalez-Manteiga1999. } In this paper, we construct the counterfactual Kaplan-Meier estimator in ((ref)) based on the Nadaraya-Watson weights

displaymathB_{n_{\ell}}(x;h_{n})=\frac{K(\frac{x-X_{\ell}}{h_{n}})}{\sum_{i=1}^{n}K(\frac{x-X_{i}}{h_{n}})},\quad\quad \ell=1,2,\cdots,n ,

for some kernel function $K$ and bandwidth $h_{n}$. For ease of notation, we suppress the dependence on $h_{n}$ for $\hat{F}_{T^{*}}$ and $\hat{F}_{T|X}$ hereafter. To estimate $F_{T}(t)$, we adopt the unconditional Kaplan-Meier (KaplanMeier1958) estimator

equation[equation omitted — 136 chars of source]

where $\{(Y_{(j)},\delta_{(j)})\}_{j=1}^{n}$ are the $n$ pairs of observations ordered on the order statistics of $\{Y_{i}\}_{i=1}^{n}$.

Asymptotic theory

Representations

Asymptotic properties of the unconditional Kaplan-Meier estimator $\hat{F}_{T}$ in ((ref)) have been studied extensively in survival analysis. One attractive feature is that $\hat{F}_{T}(t)-F_{T}(t)$ can be approximated by an average of independent and identically distributed random variables with mean zero.\footnote{ Another appealing feature is the strong approximation for $\sqrt{n}(\hat{F}_{T}-F_{T})$ by a sequence of Gaussian processes. See for example BurkeCsoergoEtAl1988 and MajorRejto1988. } We state this representation in the following proposition for completeness.

ProUnder Assumptions (ref)-(ref) and (ref), for any $\zeta<\tau=\inf\{t:F_{Y}(t)=1\}$ and $t\in[0,\zeta]$, \begin{align*} \sqrt{n}\left(\hat{F}_{T}(t)-F_{T}(t)\right) =\frac{1}{\sqrt{n}}\sum_{i=1}^{n}\xi(Y_{i},\delta_{i};t)+R_{n}(t), \end{align*} where \begin{align*} \xi(y,\delta;t)=\left[1-F_{T}(t)\right]\left[\frac{\mathbbm{1}_{[y\leq t, \delta=1]}}{1-F_{Y}(y)}-\int_{0}^{y\wedge t}\frac{F^{\delta}_{Y}(\!\!\;\mathrm{d} u)}{(1-F_{Y}(u))^{2}}\right], \end{align*} and \begin{align*} \sup_{t\in [0,\zeta]}|R_{n}(t)|=\operatorname{o}_{p}\left(1\right). \end{align*} \qed

The influence function of $\hat{F}_{T}$ in Proposition (ref) is centered at zero. Moreover, the precise rate of approximation error $\sup_{t\in [0,\zeta]}|R_{n}(t)|$ differs if different assumptions about the data generating process are imposed (cf. LoSingh1986, Cai1998, ChenLo1997).

In addition to the representation of $\hat{F}_{T}$, a similar representation of $\hat{F}_{T|X}$ in ((ref)) has been established in the case of a univariate covariate by Iglesias-PerezGonzalez-Manteiga1999, and further extended to the case of multivariate covariates and dependent data by LiangUna-AlvarezEtAl2012. Since the counterfactual Kaplan-Meier estimator $\hat{F}_{T^{*}}$ in ((ref)) is constructed by taking average of $\hat{F}_{T|X}$ with respect to the empirical distribution of $X^{*}$, applying the representation of $\hat{F}_{T|X}$ allows us to approximate $\hat{F}_{T^{*}}$ by an average of independent and identically distributed random variables with mean zero. To obtain the approximation of $\hat{F}_{T^{*}}$, we need the following assumptions about the kernel and bandwidth.\\[0.2cm] Assumption K (Kernel)\\ The kernel function $K:\mathbb{R}^{d}\rightarrow \mathbb{R}$ satisfies the following conditions.

enumerate[label=K\arabic*] • $K$ is of bounded variation, vanishes outside $[-1,1]^{d}$, and satisfies $\int K(u)\;\mathrm{d} u=1$. • There is a positive integer $r\geq 2$ such that \begin{align*} \int \left(\prod_{\ell=1}^{d}u_{\ell}^{\lambda_{\ell}}\right)K(u)\;\mathrm{d} u=0 \end{align*} for any $d$-dimensional vector $\lambda=(\lambda_{1},\ldots,\lambda_{d})^{\top}$ of nonnegative integers with $\sum_{\ell=1}^{d}\lambda_{\ell}\leq r-1$. • For $u\in[-1,1]^{d}$, $K(u)$ is $r$-times differentiable with respect to $u$ and the derivatives are uniformly continuous and bounded. • For $u\in[-1,1]^{d}$, $K(u)=K(|u|)$.

Assumption B (Bandwidth)\\ The sequence $\{h_{n}\}_{n=1}^{\infty}$ of bandwidths satisfies the following conditions.

enumerate[label=B\arabic*] • $h_{n}\rightarrow 0$. • $n^{1/2}\left(\frac{\log n}{nh_{n}^{d}}\right)^{3/4}\rightarrow 0$. • $n^{1/2}h_{n}^{r}\rightarrow 0$.

These assumptions about the kernel and bandwidth are mild. Assumption (ref) restricts the choice of kernels so that the estimator $\hat{F}_{T|X}(t|x)$ is $r$-times differentiable with respect to $x$ and these derivatives are uniformly continuous and bounded. Assumptions (ref) implies the remainder term in the representation of $\hat{F}_{T|X}$ in Proposition (ref) below is of order less than $n^{-1/2}$. Assumptions (ref)-(ref) and (ref) are imposed to ensure the bias terms of $\hat{F}_{T^{*}}$ in Proposition (ref) are also of order less than $n^{-1/2}$. Note that a necessary condition to make Assumptions (ref) and (ref) valid simultaneously is $3d<2r$. Thus, we use a higher order kernel to construct $\hat{F}_{T^{*}}$ under Assumption (ref) if multivariate policy variables are of interest, that is, $d\geq 2$. Assumption (ref) is valid if $K$ is a product kernel function $K(u)=\prod_{\ell=1}^{d}k_{\ell}(u_{\ell})$ and each $k_{\ell}$ is a univariate kernel function that is symmetric around zero.

Moreover, we need conditions about the support and smoothness of densities and distributions as follows.\\[0.2cm] Assumption SP (Support)

enumerate[label=SP\arabic*] • The support of $X^{*}$ is the compact subset $J^{*}\equiv\prod_{\ell=1}^{d}[\underline{X_{\ell}^{*}},\overline{X_{\ell}^{*}}]$ of the interior of the support of $X$, say $J\equiv\prod_{\ell=1}^{d}[\underline{X_{\ell}},\overline{X_{\ell}}]$. • There exist positive numbers $u_{0}$ and $v_{0}$ such that $\inf\left\{m(x):x\in J_{v_{0}}^{*}\right\}\geq u_{0}$ where $J_{v_{0}}^{*}=\prod_{\ell=1}^{d}[\underline{X_{\ell}^{*}}-v_{0},\overline{X_{\ell}^{*}}+v_{0}]$. • There exist positive numbers $\zeta^{*}$ and $v^{*}$ such that $\inf\{1-F_{Y|X}(\zeta^{*}|x):x\in J\}\geq v^{*}$.

Assumption (ref), stronger than Assumption (ref), requires the support of $X^{*}$ to be a proper subset of the support of $X$. When $J^{*}$ is close to $J$, Assumption (ref) is valid provided the density of $X$ on the boundary of $J^{*}$ is still bounded away from zero. Assumption (ref) requires the conditional survival function of $Y$ given $X$ is uniformly bounded away from zero on $[0,\zeta^{*}]\times J$; in addition, it implies $\zeta^{*}<\tau=\inf\{t:F_{Y}(t)=1\}$.\\[0.2cm] Assumption SM (Smoothness)

enumerate[label=SM\arabic*] • The function $m(x)$ is $r$-times differentiable with respect to $x$ on the interior of $J$, and its derivatives are bounded and uniformly continuous. • For all $(t,x)\in \mathbb{R}\times J_{v_{0}}^{*}$, the first $r$ partial derivative with respect to $x$ of $F_{T|X}(t|x)$, $F_{C|X}(t|x)$, $f_{T|X}(t|x)$ and $f_{C|X}(t|x)$ are bounded. • For all $(t,x)\in \mathbb{R}\times J_{v_{0}}^{*}$, the first derivative with respect to $t$ of $f_{T|X}(t|x)$ and $f_{C|X}(t|x)$ are bounded. • The function $m^{*}(x)$ is $r$-times differentiable with respect to $x$ on the interior of $J$, and its derivatives are bounded and uniformly continuous.\footnote{Let $m^{*}(x)=0$ if $x\notin J^{*}$.} • Both $\int (\sup_{s\in[0,\zeta^{*}]}f_{T|X}(s|x))^{2}m(x)\;\mathrm{d} x$ and $\int [m^{*}(x)]^{2}/m(x)\;\mathrm{d} x$ are finite.

Assumptions (ref)-(ref) are imposed to obtain the representation of $\hat{F}_{T|X}$ in Proposition (ref). Similar conditions are used in Iglesias-PerezGonzalez-Manteiga1999 and LiangUna-AlvarezEtAl2012. As in Rothe2010, we impose Assumption (ref) to establish the representation of $\hat{F}_{T^{*}}$ in Proposition (ref) by standard kernel smoothing techniques. Assumption (ref) is technical and valid if the second moments of $\sup_{s\in[0,\zeta^{*}]}f_{T|X}(s|X)$ and $m^{*}(X)/m(X)$ are finite.

Pro(i) Under Assumptions (ref)-(ref), (ref)-(ref), (ref)-(ref), (ref)-(ref), (ref)-(ref), and (ref)-(ref), for $(t,x)\in [0,\zeta^{*}]\times J^{*}$, \begin{displaymath} \hat{F}_{T|X}(t|x)-F_{T|X}(t|x)=\sum_{i=1}^{n}\xi^{*}(Y_{i},\delta_{i};t,x)B_{n_{i}}(x)+r_{n}(t,x), \end{displaymath} where \begin{displaymath} \xi^{*}(y,\delta;t,x)=\left[1-F_{T|X}(t|x)\right]\left[\frac{\mathbbm{1}_{[y\leq t, \delta=1]}}{1-F_{Y|X}(y|x)}-\int_{0}^{y\wedge t}\frac{F^{\delta}_{Y|X}(\!\!\;\mathrm{d} u|x)}{(1-F_{Y|X}(u|x))^{2}}\right], \end{displaymath} and \begin{displaymath} \sup_{(t,x)\in [0,\zeta^{*}]\times J^{*}}|r_{n}(t,x)|=\operatorname{O}_{as}\left(\left(\frac{\log n}{nh^{d}}\right)^{3/4}\right). \end{displaymath} (ii) If in addition Assumptions (ref)-(ref), (ref), and (ref)-(ref) hold, then for $t\in [0,\zeta^{*}]$, \begin{align*} & \sqrt{n}\left(\hat{F}_{T^{*}}(t)-F_{T^{*}}(t)\right)\\ = & \frac{1}{\sqrt{n}}\sum_{i=1}^{n}\left(F_{T|X}(t|X_{i}^{*})-F_{T^{*}}(t)\right) +\frac{1}{\sqrt{n}}\sum_{i=1}^{n}\xi^{*}(Y_{i},\delta_{i};t,X_{i})\frac{m^{*}(X_{i})}{m(X_{i})}+R^{*}_{n}(t), \end{align*} where \begin{displaymath} \xi^{*}(y,\delta;t,x)=\left[1-F_{T|X}(t|x)\right]\left[\frac{\mathbbm{1}_{[y\leq t, \delta=1]}}{1-F_{Y|X}(y|x)}-\int_{0}^{y\wedge t}\frac{F^{\delta}_{Y|X}(\!\!\;\mathrm{d} u|x)}{(1-F_{Y|X}(u|x))^{2}}\right]. \end{displaymath} and \begin{displaymath} \sup_{(t,x)\in [0,\zeta^{*}]}|R^{*}_{n}(t)|=\operatorname{o}_{p}\left(1\right). \end{displaymath} \qed

The influence function of $\hat{F}_{T^{*}}$ in Proposition (ref) has mean zero and can be decomposed into two components: the first part arises from the sample variation in $X^{*}$, and the second part results from the estimate of $F_{T|X}(t|x)$. The influence function of $\hat{F}_{T^{*}}$ is different from that of the counterfactual CDF estimator in Rothe2010 because we use the estimator $\hat{F}_{T|X}(t|x)$ in ((ref)) to recover the conditional CDF $F_{T|X}(t|x)$ in the presence of random censoring.

Asymptotic properties

Propositions (ref) and (ref) show that the estimators $\hat{F}_{T}(t)$ and $\hat{F}_{T^{*}}(t)$ can be represented by the average of functions of independent and identically distributed random variables $(Y, \delta, X, X^{*})^{\top}$ plus asymptotic negligible terms. The counterfactual estimator $\hat{F}_{T^{*}}$ is uniformly consistent for $F_{T^{*}}$ on $[0,\zeta^{*}]$ because the two classes $\{x\mapsto F_{T|X}(t|x):t\in[0,\zeta^{*}]\}$ and $\{(y,\delta,x)\mapsto\xi^{*}(y,\delta;t,x):t\in [0,\zeta^{*}]\}$ are both Euclidean under the assumptions imposed. Moreover, for each $t\in[0,\zeta^{*}]$, the proposed estimator $\hat{F}_{T^{*}}(t)$ can avoid the curse of dimensionality, namely convergence at the usual parametric rate $\sqrt{n}$, even if the first-stage estimator $\hat{F}_{T|X}(t|x)$ converges to $F_{T|X}(t|x)$ at a rate slower than $\sqrt{n}$ for each $x\in J^{*}$.\footnote{ Details can be found in Dabrowska1989, Iglesias-PerezGonzalez-Manteiga1999, Iglesias-Perez2003, and LiangUna-AlvarezEtAl2012. }

The representations of $\hat{F}_{T}$ and $\hat{F}_{T^{*}}$ further allow us to apply techniques in the literature on empirical processes to show that the random map

equation[equation omitted — 94 chars of source]

converges weakly to a two dimensional centered Gaussian process, where $\hat{\mathbf{F}}\equiv(\hat{F}_{T^{*}},\hat{F}_{T})^{\top}$, $\mathbf{F}\equiv(F_{T^{*}},F_{T})^{\top}$, and $t=(t_{1},t_{2})^{\top}$. Let $Z=(Y,\delta,X,X^{*})^{\top}$ and

align[align omitted — 253 chars of source]

where $\xi$ and $\xi^{*}$ are defined in Propositions (ref) and (ref), respectively. We establish the weak convergence of $\hat{\mathbf{F}}$ as follows.

ThmIf Assumptions D, I, K, B, SP, and SM hold, then in $\mathbb{D}[0,\zeta^{*}]\times\mathbb{D}[0,\zeta^{*}]$, \begin{align*} \sqrt{n}\left(\hat{\mathbf{F}}(\cdot)-\mathbf{F}(\cdot)\right)\Rightarrow \mathbb{F}(\cdot), \end{align*} where $\mathbb{F}$ is a two dimensional centered Gaussian process with covariance function $\Sigma(s,t)=\mathbb{E}\left(\Psi(s,Z)\Psi(t,Z)^{\top}\right)$ for every $s,t\in[0,\zeta^{*}]\times[0,\zeta^{*}]$.\qed

Theorem (ref) demonstrates that when the sample size is large, we can approximate the random map in ((ref)) by the two dimensional centered Gaussian process $\mathbb{F}$ with the covariance function

align*[align* omitted — 207 chars of source]

where

align*[align* omitted — 597 chars of source]

and

align*[align* omitted — 833 chars of source]

The second term in the last line is zero if $X^{*}$ is independent of $(Y,\delta,X)$, which is expected to be valid in the analysis of the composition effect. In contrast, if we consider a counterfactual manipulation with $X^{*}=\pi(X)$, then the second term should not be omitted in general.

We can make inference on the distribution policy effect $\triangle_{F}(t)$ by $\hat{\triangle}_{F}(t)\equiv \hat{F}_{T^{*}}(t)-\hat{F}_{T}(t)$ because $\sqrt{n}\left(\hat{\triangle}_{F}(\cdot)-\triangle_{F}(\cdot)\right)$ converges weakly to $(1,-1)\mathbb{F}(\cdot)$ in $\mathbb{D}[0,\zeta^{*}]\times\mathbb{D}[0,\zeta^{*}]$ by Theorem (ref) and the continuous mapping theorem. The covariance function $\Sigma(s,t)$ can be estimated by replacing the unknown functions with associated consistent estimators. For example, a plug-in estimator of $\Sigma_{11}(u,u')$ is

align*[align* omitted — 477 chars of source]

where $\hat{F}_{T^{*}}$ is defined in ((ref)), $\hat{F}_{T|X}$ is defined in ((ref)),

align*[align* omitted — 563 chars of source]

To analyze other counterfactual policy effects, we apply the functional delta method as follows.

ThmLet $\nu$ be a functional mapping from a subset of $\mathbb{D}[0,\zeta^{*}]\times\mathbb{D}[0,\zeta^{*}]$ to some normed space $\mathcal{V}$. Suppose that $\nu$ is Hadamard differentiable at $\mathbf{F}$ with derivative $\nu_{\mathbf{F}}'$. Let $Z=(Y,\delta,X,X^{*})^{\top}$ and $\Psi^{\nu}(t,Z)=\nu_{\mathbf{F}}'(\Psi)(t,Z)$, where $\Psi=(\psi_{1},\psi_{2})^{\top}$ is defined in ((ref)). Under the assumptions of Theorem (ref), we have that in $\mathcal{V}$, \begin{displaymath} \sqrt{n}\left(\nu(\hat{\mathbf{F}})(\cdot)-\nu(\mathbf{F})(\cdot)\right) \Rightarrow \nu_{\mathbf{F}}'(\mathbb{F})(\cdot)\equiv\mathbb{G}(\cdot), \end{displaymath} where $\mathbb{G}$ is a two dimensional centered Gaussian process with covariance function $\Sigma^{\nu}(s,t)=\mathbb{E}\left(\Psi^{\nu}(s,Z)\Psi^{\nu}(t,Z)^{\top}\right)$.\qed

Let $\hat{\Lambda}_{T}= \nu(\hat{F}_{T})$ and $\hat{\Lambda}_{T^{*}}= \nu(\hat{F}_{T^{*}})$ for the functional $\nu:F\mapsto\int_{[0,\cdot]} \frac{1}{1-F^{-}}\;\mathrm{d} F$. In addition, let $\hat{\mathbf{\Lambda}}\equiv(\hat{\Lambda}_{T^{*}},\hat{\Lambda}_{T})^{\top}$ and $\mathbf{\Lambda}\equiv(\Lambda_{T^{*}},\Lambda_{T})^{\top}$. Theorem (ref) immediately implies the following corollary. We can make inference on the cumulative hazard policy effect $\triangle_{\Lambda}(t)$ by $\hat{\triangle}_{\Lambda}(t)\equiv\hat{\Lambda}_{T^{*}}(t)-\hat{\Lambda}_{T}(t)$ because $\sqrt{n}\left(\hat{\triangle}_{\Lambda}(\cdot)-\triangle_{\Lambda}(\cdot)\right)$ converges weakly in $\mathbb{D}[0,\zeta^{*}]\times\mathbb{D}[0,\zeta^{*}]$ by the continuous mapping theorem.

comment\begin{align*} \sqrt{n}\left(\hat{\triangle}_{\Lambda}(\cdot)-\triangle_{\Lambda}(\cdot)\right)\Rightarrow (1,-1)\mathbb{A}(\cdot) \end{align*}
CorUnder the assumptions of Theorem (ref), \begin{align*} \sqrt{n}\left(\hat{\mathbf{\Lambda}}(\cdot)-\mathbf{\Lambda}(\cdot)\right) \Rightarrow & \left[ \begin{array}{c} \int_{0}^{\cdot}\frac{1}{1-F^{-}_{T^{*}}(u)}\mathbb{F}_{1}(\!\!\;\mathrm{d} u) +\int_{0}^{\cdot}\frac{\mathbb{F}^{-}_{1}(u)}{\left(1-F^{-}_{T^{*}}(u)\right)^{2}}F_{T^{*}}(\!\!\;\mathrm{d} u) \\ \int_{0}^{\cdot}\frac{1}{1-F^{-}_{T}(u)}\mathbb{F}_{2}(\!\!\;\mathrm{d} u) +\int_{0}^{\cdot}\frac{\mathbb{F}^{-}_{2}(u)}{\left(1-F^{-}_{T}(u)\right)^{2}}F_{T}(\!\!\;\mathrm{d} u) \end{array} \right] \equiv\mathbb{A}(\cdot) \end{align*} in $\mathbb{D}[0,\zeta^{*}]\times\mathbb{D}[0,\zeta^{*}]$. The two dimensional process $\mathbb{A}$ is centered Gaussian with covariance function $\Sigma^{\Lambda}(s,t)=\mathbb{E}\left(\Psi^{\Lambda}(s,Z)\Psi^{\Lambda}(t,Z)^{\top}\right)$, where $\Psi^{\Lambda}(t,Z)=\left(\frac{\psi_{1}(t_{1},Z)}{1-F_{T^{*}}(t_{1})}, \frac{\psi_{2}(t_{2},Z)}{1-F_{T}(t_{2})}\right)^{\top}$, $Z=(Y,\delta,X,X^{*})^{\top}$, and $(\psi_{1},\psi_{2})^{\top}$ is defined in ((ref)).\qed
commentWe first establish the weak convergence of the quantile process and the cumulative hazard process in the following two propositions. \begin{Pro} Let $[q_{1}^{*},q_{2}^{*}]$ and $[q_{1},q_{2}]$ be subintervals of $(0,F_{T^{*}}(\zeta^{*}))$ and $(0,F_{T}(\zeta))$ respectively. Suppose that $F_{T^{*}}$ and $F_{T}$ are continuously differentiable on $[F_{T^{*}}^{-1}(q_{1}^{*})-\varepsilon, F_{T^{*}}^{-1}(q_{2}^{*})+\varepsilon]$ and $[F_{T}^{-1}(q_{1})-\varepsilon, F_{T}^{-1}(q_{2})+\varepsilon]$ for some $\varepsilon>0$, with strictly positive derivative $f_{T^{*}}$ and $f_{T}$ respectively. Then we have \begin{displaymath} \sqrt{n}\left(\hat{\mathbf{Q}}(\cdot)-\mathbf{Q}(\cdot)\right) \Rightarrow -\left(\frac{\mathbb{F}_{1}}{f_{T^{*}}},\frac{\mathbb{F}_{2}}{f_{T}}\right)^{\top}\mathbf{Q}(\cdot) \equiv\mathbb{Q}(\cdot) \end{displaymath} and given the data $(Y_{i}, \delta_{i}, X_{i})_{i=1}^{n}$ and $(X_{i}^{*})_{i=1}^{n^{*}}$, \begin{displaymath} \sqrt{n}\left(\hat{\mathbf{Q}}^{b}(\cdot)-\hat{\mathbf{Q}}(\cdot)\right) \Rightarrow \mathbb{Q}(\cdot) \end{displaymath} in probability, where $\mathbb{Q}$ is a two dimensional Gaussian process with mean zero and covariance function $\Sigma^{Q}(\tau,\tau')=\mathbb{E}\left(\Psi^{Q}(\tau,Z)\Psi^{Q}(\tau',Z)^{\top}\right)$ with $\Psi^{Q}(\tau,Z)=$ $\left(\frac{\psi_{1}^{F}(Q_{T^{*}}(\tau_{1}),Z)}{f_{T^{*}}(Q_{T^{*}}(\tau_{1}))}, \frac{\psi_{2}^{F}(Q_{T}(\tau_{2}),Z)}{f_{T}(Q_{T}(\tau_{2}))}\right)^{\top}$ and the convergence is in $\ell^{\infty}(0,F_{T^{*}}({\zeta^{*}}))\times\ell^{\infty}(0,F_{T}({\zeta}))$. \end{Pro} \begin{proof} Consider the functional $\nu(\phi_{1},\phi_{2})=(\phi_{1}^{-1},\phi_{2}^{-1})^{\top}$ where $\phi_{i}^{-1}(\tau)=\inf\{s\in \mathbb{R}:\phi_{i}(s)\geq\tau\}$. By Lemma 21.4 in Vaart1998, the functional $\nu$ is Hadamard differentiable at $\mathbf{F}$ tangentially to $C[F_{T^{*}}^{-1}(q_{1}^{*})-\varepsilon, F_{T^{*}}^{-1}(q_{2}^{*})+\varepsilon]\times C[F_{T}^{-1}(q_{1})-\varepsilon, F_{T}^{-1}(q_{2})+\varepsilon]$ with derivative \begin{displaymath} \nu_{\mathbf{F}}^{'}(\phi_{1},\phi_{2})= -\left(\frac{\phi_{1}}{f_{T^{*}}},\frac{\phi_{2}}{f_{T}}\right)^{\top}\mathbf{F}^{-1}. \end{displaymath} Theorem (ref) therefore implies $\sqrt{n}(\hat{\mathbf{Q}}(\cdot)-\mathbf{Q}(\cdot))$ converges weakly to $\nu_{\mathbf{F}}^{'}(\mathbb{F})(\cdot)=\mathbb{Q}(\cdot)$ in $\ell^{\infty}(0,F_{T^{*}}({\zeta^{*}}))\times\ell^{\infty}(0,F_{T}({\zeta}))$. Moreover, the weak convergence of the empirical bootstrap quantile process $\hat{\mathbf{Q}}^{b}(\cdot)$ is the consequence of Theorem (ref). \end{proof} By Proposition (ref), together with continuous mapping theorem, we have \begin{displaymath} \sqrt{n}\left(\hat{\triangle}_{Q}(\cdot)-\triangle_{Q}(\cdot)\right)\Rightarrow (1,-1)\mathbb{Q}(\cdot). \end{displaymath} In addition, the $(1-\alpha)$ uniform confidence bands for can also be constructed by the empirical bootstrap method similar to the procedures in previous section.

Monte Carlo simulation

In this section, we evaluate the small-sample performance of the proposed estimator $\hat{F}_{T^{*}}$ and its associated counterfactual policy effects by Monte Carlo simulation. We consider the following data generating process (DGP):

align*[align* omitted — 100 chars of source]

where the covariates $(X_{1}$, $X_{2})$ follow the Beta distribution with shape parameters $(2,2)$, the unobserved heterogeneity $\varepsilon$ is exponentially distributed with mean 2, and $C$ is log-normally distributed with parameters $(2.5,1)$; additionally, $(X_{1}, X_{2}, \varepsilon, C)$ are mutually independent. The censoring rate in this design is approximate $23.45\%$. We study the policy intervention

align*[align* omitted — 93 chars of source]

and this intervention does not affect $(C,\varepsilon)$. The DGP and counterfactual policy are not meant to mimic any data set in empirical studies; instead, they are only used to illustrate the proposed method.

We consider the sample sizes $n=100, 200, 400$, and $800$. The number of simulation replications is $S=1000$. The criteria of evaluation include the mean integrated absolute error (MIAE) and the root mean integrated squared error (RMISE).\footnote{ For an estimator $\hat{f}$ of a generic real-valued function $f$, the mean integrated absolute error of $\hat{f}$ is

align*[align* omitted — 111 chars of source]

and the root mean integrated squared error of $\hat{f}$ is

align*[align* omitted — 124 chars of source]

} The CDF and cumulative hazard estimators in this Monte Carlo study are calculated over the eqidistant grids $\{4.25,4.30,4.35,\ldots,8.10,8.15\}$, and the numerical integration in MIAE and RMISE is taken over $[4.25,8.15]$, where $4.25$ and $8.15$ are the $10\%$ and $90\%$ quantile of $T$, respectively.

We evaluate the estimation of the CDFs $(F_{T^{*}},F_{T})$ and the estimation of the cumulative hazard functions $(\Lambda_{T^{*}},\Lambda_{T})$. We estimate $F_{T}$ by the unconditional Kaplan-Meier estimator $\hat{F}_{T}$ in ((ref)). To estimate $F_{T^{*}}$, we use the proposed estimator $\hat{F}_{T^{*}}$ in ((ref)) with the fourth order product kernel function $K(u_{1},u_{2})=k(u_{1})k(u_{2})$ where $k(u)=(15/32)(3-10u^{2}+7u^{4})\mathbbm{1}_{[|u|<1]}$ and the bandwidth $h_{n}=3n^{-1/7}$.\footnote{ Assumptions (ref)-(ref) and (ref)-(ref) are satisfied under this choice of kernel function and bandwidth} Table (ref) shows that the MIAE and RMISE of $(\hat{F}_{T^{*}},\hat{F}_{T})$ shrink as the sample size increases.

table[table omitted — 998 chars of source]

The MIAE and RMISE of $\hat{F}_{T^{*}}$ halves as the sample size quadruples; namely, this estimator converges at the rate $\sqrt{n}$. This confirms the theoretical analysis that the proposed estimator $\hat{F}_{T^{*}}$ does not suffer from the curse of dimensionality. We also consider an oracle estimator $\tilde{F}_{T^{*}}$, which is the unconditional Kaplan-Meier estimator of $F_{T^{*}}$ if $Y^{*}=\min\{T^{*},C\}$ and $\delta^{*}=\mathbbm{1}_{[T^{*}\leq C]}$ are observed. Surprisingly, this oracle estimator $\tilde{F}_{T^{*}}$ does not outweigh considerably the proposed estimator $\hat{F}_{T^{*}}$ in terms of MIAE and RMISE; however, $\tilde{F}_{T^{*}}$ is infeasible because $Y^{*}$ and $\delta^{*}$ are unobserved in practice. Moreover, we consider Rothe's (Rothe2010) counterfactual estimator $F^{\dagger}_{T^{*}}$, which is constructed under the assumption that data are not censored.\footnote{ Since the support of $(X^{*}_{1},X^{*}_{2})$ is a proper subset of the support of $(X_{1},X_{2})$, we construct Rothe's estimator based on the aforementioned fourth order kernel function and bandwidth $h_{n}=3n^{-1/7}$. } As shown in Table (ref), the neglect of censoring results in larger MIAE and RMISE of $F^{\dagger}_{T^{*}}$, compared with those of $\hat{F}_{T^{*}}$ and $\tilde{F}_{T^{*}}$. Table (ref) reports the MIAE and RMISE of the estimated cumulative hazard functions

align*[align* omitted — 242 chars of source]

where $\nu$ is the functional that maps $F$ to $-\log{(1-F)}$.

table[table omitted — 1,071 chars of source]

Similarly, the simulation results provide evidence that the proposed cumulative hazard function $\hat{\Lambda}_{T^{*}}$ converges at the rate $\sqrt{n}$. Moreover, $\hat{\Lambda}_{T^{*}}$ performs as well as the oracle estimator $\tilde{\Lambda}_{T^{*}}$. The neglect of censoring, however, causes relatively large bias of $\Lambda^{\dagger}_{T^{*}}$.

Conclusion

We have proposed a two-stage fully nonparametric estimator of the counterfactual CDF for a duration variable, which is subject to the random censoring. Since the nonseparable heterogeneity is of unrestricted dimensionality and its marginal distribution is unspecified, the duration analysis in this paper would avoid several types of model misspecification in empirical studies. The incorporation of covariates also enables researchers to evaluate the change of duration dependence in response to a counterfactual policy that changes exogenous covariates.

There are some directions of extension to this research. First, we may relax the assumption of exogenous covariates by the control function approach. Second, it would be important to establish the validity of a bootstrap method to construct a confidence band for the counterfactual policy effect. Finally, the inclusion of time-varying covariates might be relevant in some empirical studies.