Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
55,272 characters · 13 sections · 59 citation commands
Tests of exogeneity in duration models with censored data
We consider the setting in which a researcher is interested in identifying the causal effect of a treatment variable $Z$ on a duration of interest $T$. Note that $Z$ could be a policy intervention, a new medical treatment or other forms of exposure. A central challenge in this context arises from the problem of endogeneity, which refers to $Z$ not being independent of the error term in the structural model for $T$. Endogeneity can result from a variety of sources such as unobserved confounders, sample selection, measurement error and noncompliance rubin1974estimating,HeckmanJamesJ.1979SSBa,angrist1996identification. When $Z$ is endogenous, standard estimation methods can be severely biased. Because much applied work relies on observational data, addressing endogeneity should be a necessary step in any empirical analysis.
When the treatment $Z$ is exogenous, meaning independent of the error term in the structural model for the outcome $T$, estimating its causal effect is a well-posed problem for which standard estimation approaches are well-suited. By contrast, when $Z$ is endogenous, estimation of the treatment effect becomes more challenging. A common approach to deal with endogeneity is the use of instrumental variable (IV) methods. This approach relies on external sources of variation (instruments) that influence $Z$, but are independent of the unobserved heterogeneity affecting both $Z$ and $T$. However, with nonparametric IV methods, the estimation problem often takes the form of an ill-posed inverse problem newey2003instrumental,chernozhukov2005iv,darolles2011nonparametric,blundell2013nonparametric. This means that nonparametric estimation with instruments is more convoluted and unstable compared to standard estimation methods such as ordinary least squares or a general quantile estimator. As a result, researchers often focus on simpler estimands, such as averages, rather than attempting to recover full functional relationships. These issues highlight the importance of having a statistical test for the hypothesis that $Z$ is exogenous.
We consider the following nonparametric nonseparable model
where the map $u \mapsto \varphi(X,Z,u)$ is strictly increasing for each $(X,Z)$ and $U \sim \mathcal{U}[0,1]$. The variable $T$ is a duration of interest, $Z$ is a possibly endogenous treatment variable and $X$ represents a vector of baseline covariates. Note that if $\varphi(X,Z,\cdot)$ is strictly increasing for each $(X,Z)$, the rank invariance assumption as described by chernozhukov2005iv is implied dong2018testing. The rank invariance assumption is a condition on the potential outcomes and implies that individuals have the same unobserved rank when their treatment changes. Further, we will suppose that there exists an instrument $W$ that is independent of $U$ given $X$. We also allow for the duration time $T$ to be censored by a right censoring time $C$. Therefore, we only observe the follow-up time $Y=\min\{T,C\}$ and the censoring indicator $\Delta=\mathbbm{1}(Y=T)$ alongside $X,W$ and $Z$. We will assume that $X,W$ and $Z$ are categorical variables with finite support to facilitate the construction of the test statistics. Moreover, it is assumed that $C$ is independent of $T$ and $W$ jointly given $X$ and $Z$. The goal of this paper is to develop nonparametric tests for the hypothesis that $Z$ is exogenous, meaning that $Z$ is independent of $U$ conditional on $X$, without having to estimate the function $\varphi$.
There are two main approaches to test for exogeneity in a nonparametric nonseparable model. The first is based on estimating the function $\varphi$ under the identification assumptions of the nonparametric IV model, and comparing the estimate to general regression or quantile estimates. The second approach is to estimate the residuals under the exogeneity assumption and verify if the IV condition is satisfied. We will follow the second approach to avoid the difficulties associated with nonparametric IV estimation. This puts us in a similar framework to feve2018estimation, who developed a nonparametric test for exogeneity based on an estimator of the distribution of the conditional rank $F_{T\mid Z}(T\mid Z)$. Contrary to feve2018estimation, we allow for $T$ to be subject to right censoring. Moreover, we also allow for the inclusion of baseline covariates in model (ref). These generalizations introduce significant challenges during the construction of the test statistics. Note that if $T$ is right censored by $C$, this implies that $V_T=F_{T\mid X,Z}(T\mid X,Z)$ is right censored by $V_C=F_{T\mid X,Z}(C\mid X,Z)$. Therefore, conditional on $F_{T\mid X,Z}$, we only observe $V = \min\{V_T,V_C\}$ and $\Delta=\mathbbm{1}(Y=T)$. Because of this, also the estimated $\widehat{V}_T=\widehat{F}_{T\mid X,Z}(T\mid X,Z)$ is subject to right censoring, which complicates the estimation of the distribution of $V_T$ significantly. Whereas feve2018estimation permits $W$ and $Z$ to be continuous, we restrict $X,W$ and $Z$ to be categorical. This restriction has the advantage that we do not need to smooth over the covariates, treatment and duration time when estimating $F_{T\mid X,Z}$. Under suitable identification conditions on $\varphi$, we show that the independence between $V_T$ and $(X,W)$ is equivalent to $Z$ being independent of $U$ given $X$. Other tests of exogeneity for censored duration outcomes have been proposed in the literature by smith1986exogeneity and rivers1988limited among others. However, these tests assume fixed and fully observed censoring times in Tobit models. To the best of our knowledge, no other test for exogeneity in a nonparametric nonseparable model that allows for random right censoring is available in the literature.
The nonparametric nonseparable model is a common framework for analyzing the effects of endogenous treatments on censored duration outcomes. Several nonparametric contributions that focus on estimating local average treatment effects under a monotonicity assumption include frandsen2015treatment, sant2016program and blanco2020bounds, while chernozhukov2015quantile and beyhum2022nonparametric address identification and estimation of population-level treatment effects under a type of rank invariance assumption. See wuthrich2020comparison for a comparison of rank invariance to monotonicity for the identification of quantile treatment effects. Another strand of the literature achieves identification under separability assumptions abbring2003nonparametric,abbring2005social,bijwaard2005correcting. Some semiparametric approaches include tchetgen2015instrumental, li2015instrumental, chan2016reader kianian2021causal, beyhum2024instrumental and tedesco2025instrumental among others.
The remainder of the paper is organized as follows. Section (ref) discusses under which conditions exogeneity of $Z$ is equivalent to the independence between $V_T$ and $(X,W)$, how the test statistics are constructed and explains the estimation procedure. The test statistics' limiting distributions and two bootstrap approaches to approximate the critical values are given in Section (ref). The finite-sample performance is investigated in Section (ref) through Monte Carlo simulations and Section (ref) provides an empirical application to the National Job Training Partnership Act (JTPA) Study. Section (ref) discusses possible extensions and practical considerations. The technical proofs can be found in the Appendix.
In this paper, we develop nonparametric tests for the null hypothesis that $$H_0 : V_T \perp \!\!\! \perp (X,W),$$ where $ V_T = F_{T \mid X,Z}(T \mid X,Z)$ with $F_{T \mid X,Z}(t \mid x,z) = \mathbbm{P}(T \leq t\mid X=x,Z=z)$, $X$ is a vector of baseline covariates and $W$ is an instrumental variable such that $W \perp \!\!\! \perp U\mid X$ with $\perp \!\!\! \perp$ denoting statistical independence. Because $T$ is right censored by $C$, $V_T$ is right censored by $V_C = F_{T \mid X,Z}(C \mid X,Z)$. Therefore, conditionally on $F_{T \mid X,Z}$, we only observe $V=F_{T \mid X,Z}(Y \mid X,Z)=\min\{V_T,V_C\}$ and $\Delta = \mathbbm{1}(Y=T)$. Note that $T$ and $V_T$ both being subject to right censoring introduces significant challenges in constructing the test statistics. We now introduce the following model condition:
The model being identified means that if $$ T=\varphi_1(X,Z,U_1) = \varphi_2(X,Z,U_2), $$ and $U_1$ and $U_2$ are both independent of $W$ conditionally on $X$ and uniform on $[0,1]$, then $$ U_1=U_2 \text{ a.s.\ and } \varphi_1=\varphi_2. $$ Note that $U\sim \mathcal{U}[0,1]$ can be assumed without loss of generality as long as $U$ is continuously distributed and has positive density on its support. The equivalence between $V_T \perp \!\!\! \perp (X,W)$ and $Z$ being exogenous, meaning that $Z \perp \!\!\! \perp U \mid X$, is established by the following proposition.
Proposition (ref) shows that while our independence test between $V_T$ and $(X,W)$ does not rely on Assumption (ref), the independence property can be interpreted as an exogeneity property under this assumption. The identification of model (ref), which is nonparametric and nonseparable, using instrumental variables is a complex issue that has been studied by chesher2003identification, chernozhukov2005iv, chernozhukov2007instrumental and chen2014local among others. A sufficient condition for the identification of model (ref) can be found in Appendix A of feve2018estimation, which is a completeness type condition that has been proven to be nontestable canay2013testability.
Under $H_0$, it is clear that
is equal to zero for all $ (v,x,w)$. Equivalently, we have that
is equal to zero for all $v$ and all $(x,w)$ for which $\mathbbm{P}(X=x,W=w) > 0$. While it is possible to construct our test statistics based on an estimator of (ref), $V_T$ being subject to right censoring by $V_C$ substantially complicates estimation of $\mathbbm{P}(V_T \leq v, X = x, W =w)$. Nonetheless, possible estimators for this joint distribution have been proposed by stute1993consistent and akritas1994nearest. The estimator proposed by stute1993consistent is simple to implement, but it would require the additional assumption that $\mathbbm{P}(V_T \leq V_C \mid X,W,V_T)=\mathbbm{P}(V_T \leq V_C \mid V_T)$. On the other hand, the approach proposed by akritas1994nearest requires the estimation of $\mathbbm{P}(V_T \leq v\mid X = x,W = w)$ as a first step to estimate $\mathbbm{P}(V_T \leq v, X = x, W =w)$. Therefore, we will construct our test statistics based on a nonparametric estimator of $\eqref{condind}$. To facilitate the construction of the test statistics, we will make the following key assumptions:
Even though assumption (ref) might be restrictive in some settings, it will allow us to develop a tractable, fully nonparametric approach to test $H_0$. Moreover, it has the advantage that we do not need to smooth over the covariates, treatment and duration time when estimating $F_{T\mid X,Z}$. Note that under Assumption (ref), a necessary condition for Assumption (ref) to hold is that $d_w \geq d_z$. Possible extensions to allow for continuous covariates, instruments and treatments are discussed in Section (ref). Further, it follows that Assumption (ref) is equivalent to assuming that $(i)$ $C \perp \!\!\! \perp W \mid X,Z$ and $(ii)$ $T \perp \!\!\! \perp C \mid X,W,Z$. Clearly, $(i)$ implies that after conditioning on $(X,Z)$, the instrument should not affect the censoring time $C$. In particular, if $W$ indicates being randomized to a treatment or control group, this means that, conditional on the baseline covariates $X$ and treatment participation $Z$, being assigned to the treatment or control group should not influence the censoring time. Practitioners can check for violations of $(i)$ by comparing the censoring rates for different levels of $W$ conditional on $(X,Z)$. On the other hand, $(ii)$ implies that the duration and censoring times are independent given the baseline covariates, treatment assignment and treatment participation. Even though this assumption is widely adopted in the literature on censored duration outcomes, it is, in general, nontestable TsiatisA.1975ANAo and its violation can lead to biased estimates moeschberger_consequences_1984. Note that when $C$ is observed in addition to $Y$ and $\Delta$, frandsen2019testing proposed a test for censoring point independence. In Section (ref), we outline a possible extension to our approach that allows for some forms of dependence between $T$ and $C$ after conditioning on $(X,W,Z)$.
For all $i=1,\dots,n$, define $$ \widehat{V}_i = \widehat{F}_{T \mid X,Z}(Y_i \mid X_i,Z_i), $$ where $\widehat{F}_{T \mid X,Z}$ is a conditional Kaplan-Meier estimator, that is, $$ \widehat{F}_{T \mid X,Z}(t \mid x,z) =1-\prod_{i : Y_{(i)} \leq t} \left(1-\frac{d_{(i)}(x,z)}{r_{(i)}(x,z)}\right), $$ with $Y_{(1)},\dots,Y_{(m)}$ the $m$ ordered distinct follow-up times, the number of events $d_{(i)}(x,z) = \sum_{j=1}^n\mathbbm{1}(Y_j = Y_{(i)},\Delta_j = 1,X_j=x, Z_j=z)$ and the risk set $r_{(i)}(x,z)= \sum_{j=1}^n\mathbbm{1}(Y_j \geq Y_{(i)},X_j=x, Z_j=z).$ Because $\widehat{F}_{T \mid X,Z}( \cdot \mid x,z)$ is a monotonically increasing function, we have that $$\widehat{V}_i = \widehat{F}_{T \mid X,Z}(\min\left\{T_i,C_i\right\} \mid X_i,Z_i) = \min\left\{\widehat{F}_{T \mid X,Z}(T_i \mid X_i,Z_i),\widehat{F}_{T \mid X,Z}(C_i \mid X_i,Z_i)\right\},$$ such that $\widehat{V}_{T_i} = \widehat{F}_{T \mid X,Z}(T_i \mid X_i,Z_i)$ is censored by $\widehat{V}_{C_i} = \widehat{F}_{T \mid X,Z}(C_i \mid X_i,Z_i)$ with the same censoring indicator $\Delta_i=\mathbbm{1}(Y_i=T_i)$. It is important to note that, since $V_T \perp \!\!\! \perp X,Z$ by construction, Assumption (ref) implies that $V_T \perp \!\!\! \perp V_C$. Therefore, we can estimate the distribution function of $V_T=F_{T \mid X,Z}(T \mid X,Z)$ by a Kaplan-Meier estimator, plugging-in $\{\widehat{V}_i\}_{i=1,\dots,n}$ as the observed follow-up times. Specifically, let $$ \widehat{F}_{\widehat{V}_T}(v)= 1-\prod_{i:\widehat{V}_{(i)} \leq v} \left(1-\frac{\widehat{d}_{(i)}}{\widehat{r}_{(i)}}\right), $$ with $\widehat{V}_{(1)},\dots,\widehat{V}_{(k)}$ the $k$ ordered distinct estimated conditional ranks, $\widehat{d}_{(i)}=\sum_{j=1}^n\mathbbm{1}(\widehat{V}_j = \widehat{V}_{(i)},\Delta_j = 1)$ and $\widehat{r}_{(i)}=\sum_{j=1}^n\mathbbm{1}(\widehat{V}_j \geq \widehat{V}_{(i)})$. Lastly, it is important to note that we cannot estimate $F_{V_T\mid X,W}(v\mid x, w)=\mathbbm{P}(V_T \leq v\mid X = x,W = w)$ by a conditional Kaplan-Meier estimator, since our assumptions do not imply $V_T \perp \!\!\! \perp V_C \mid X,W$. However, Assumption (ref) does imply that $V_T \perp \!\!\! \perp V_C \mid X,Z,W$. Therefore, we can estimate $F_{V_T\mid X,W,Z}(v\mid x, w,z)=\mathbbm{P}(V_T \leq v\mid X=x,W=w,Z=z)$ by a conditional Kaplan-Meier estimator and marginalize out $Z$, that is, $$ \widehat{F}_{\widehat{V}_T\mid X,W}(v\mid x, w) =\sum_{z \in\mathcal{R}_Z}\widehat{F}_{\widehat{V}_T\mid X,W,Z}(v\mid x,w,z)\widehat{p}_{Z\mid X,W}(z\mid x,w),$$ where $$ \widehat{p}_{Z\mid X,W}(z\mid x,w) = \frac{ \sum_{i=1}^n\mathbbm{1}(X_i=x,W_i=w,Z_i=z)}{\sum_{i=1}^n\mathbbm{1}(X_i=x,W_i=w)}, $$ is an estimator of $p_{Z\mid X,W}(z\mid x,w)=\mathbbm{P}(Z=z\mid X=x,W=w)$ and $$ \widehat{F}_{\widehat{V}_T\mid X,W,Z}(v\mid x, w,z) =1-\prod_{i:\widehat{V}_{(i)} \leq v} \left(1-\frac{\widehat{d}_{(i)}(x,w,z)}{\widehat{r}_{(i)}(x,w,z)}\right), $$ with $\widehat{d}_{(i)}(x,w,z)= \sum_{j=1}^n\mathbbm{1}(\widehat{V}_j = \widehat{V}_{(i)},\Delta_j = 1,X_j = x, W_j = w, Z_j = z)$ and $\widehat{r}_{(i)}(x,w,z) = \sum_{j=1}^n\mathbbm{1}(\widehat{V}_j \geq \widehat{V}_{(i)}, X_j = x, W_j = w, Z_j = z).$ Finally, let $$\widehat{D}(v,x,w)=\widehat{F}_{\widehat{V}_T\mid X,W}(v\mid x,w) - \widehat{F}_{\widehat{V}_T}(v),$$ be an estimator for (ref). To check the null hypothesis, we can compute the following Kolmogorov-Smirnov statistic:
or the following weighted Cramér-von Mises statistic:
where $\widehat{\pi}(x,w)$ is a specified weight function, $\mathcal{R}_{X,W} = \{(x,w) \in \mathcal{R}_X \times \mathcal{R}_W : \mathbbm{P}(X=x,W=w) > 0 \}$ and $\mathcal{I} \subseteq [0,1-\gamma]$ for some small $\gamma > 0$ that will be defined by Assumption (ref). For example, one may take the weight function $\widehat{\pi}(x,w)$ of the Cramér-von Mises statistic to be constant or to equal the empirical proportion $n^{-1}\sum_{i=1}^n\mathbbm{1}(X_i=x,W_i=w)$. Note that $\gamma$ being strictly positive is only necessary for the asymptotic theory. In Section (ref), we will set $\gamma = 0$ and show that the test still performs well using Monte Carlo simulations.
Before stating the main theorems regarding the test statistics' limiting distributions, we need to introduce some more notation and assumptions. Firstly, let
and similarly for $S_{V,1\mid X,Z}$ and $S_{V\mid X,Z}$. Moreover, let $$\zeta_i(v,x,w,z) = \int^v_0\frac{N_{i,x,w,z}(u)dS_{V,1\mid X,W,Z}(u\mid x,w,z)}{S_{V\mid X,W,Z}(u\mid x,w,z)^2}-\int^v_0\frac{dN^1_{i,x,w,z}(u)}{S_{V\mid X,W,Z}(u\mid x,w,z)},$$ and $$\xi_i(u,x,z) = \int^u_0\frac{N_{i,x,z}(s)dS_{V,1\mid X,Z}(s\mid x,z)}{S_{V\mid X,Z}(s\mid x,z)^2}-\int^u_0\frac{dN^1_{i,x,z}(s)}{S_{V\mid X,Z}(s\mid x,z)},$$ with $N_{i,x,w,z}(v)= \mathbbm{1}(V_i\geq v,X_i=x,W_i = w,Z_i = z)$, $N^1_{i,x,w,z}(v)=\mathbbm{1}(V_i\geq v, \Delta_i = 1,X_i=x,W_i=w,Z_i = z)$ and similarly for $N_{i,x,z}(v)$ and $N^1_{i,x,z}(v)$. Lastly, we will need the following regularity assumptions:
We now have the following result.
Note that $F_{V_T\mid X,W}(v\mid x,w)=v$ for all $(v,x,w) \in \mathcal{I}\times \mathcal{R}_{X,W}$ under $H_0$. During the proof of this theorem, we also show in Appendix (ref) that $$\sup_{v \in \mathcal{I}}\left\lvert\widehat{F}_{\widehat{V}_T}(v)-v\right\rvert =o_p(n^{-1/2}),$$ meaning that $\widehat{F}_{\widehat{V}_T}(v)$ converges to $v$ at a rate faster than the usual parametric $n^{-1/2}$-rate. While this was already shown by feve2018estimation for uncensored conditional ranks, it is not obvious that the result would still hold in the presence of right censoring. Further, let $\ell^{\infty}(\mathcal{I}\times \mathcal{R}_{X,W})$ be the set of all uniformly bounded real-valued functions equipped with the supremum norm. We are now ready to give the main result regarding the limiting distributions of the proposed test statistics.
An explicit expression for the covariance function is not provided, as it would be excessively long and almost infeasible to estimate in practice. Instead, the limiting distributions are approximated using two possible bootstrap approaches that are detailed in Section (ref). Further, consider the local alternative $$H_a :F_{V_T\mid X,W}(v\mid x,w)=v+n^{-1/2}H(v,x,w) \text{ for all } (v,x,w) \in \mathcal{I} \times \mathcal{R}_{X,W},$$ where for all $(x,w) \in \mathcal{R}_{X,W}$ the function $H$ is such that $F_{V_T\mid X,W}$ remains a valid conditional distribution function under $H_a$. It follows that, under $H_a$, the test statistics converge in distribution to the same limiting distribution as under $H_0$, except for the additive bias $H(v,x,w)$, that is $$ T_n^{KS} \xrightarrow{\textit{d}} \sup_{v \in \mathcal{I},(x,w) \in \mathcal{R}_{X,W}}\left\lvert \mathbb{G}(v,x,w) +H(v,x,w) \right\rvert, $$ and $$ T^{CM}_{n} \xrightarrow{\textit{d}} \sum_{(x,w) \in \mathcal{R}_{X,W}} \pi(x,w) \int_{\mathcal{I} } \left[\mathbb{G}(v,x,w)+H(v,x,w)\right]^2 \mathop{}\!d F_{V_T}(v). $$ Therefore, the proposed tests can detect local alternatives that converge to $H_0$ at the parametric $n^{-1/2}$-rate.
Due to the complicated covariance structure of the process $\mathbb{G}(v,x,w)$, we use bootstrap approximations for the limiting distributions of $T_n^{KS}$ and $T_n^{CM}$. For the bootstrap procedure to be valid, we follow the approach of feve2018estimation and impose the slightly stronger null hypothesis $$H^*_0 : V_T \perp \!\!\! \perp (X,W,Z),$$ instead of $H_0:V_T \perp \!\!\! \perp (X,W)$. Note that $H^*_0$ is only slightly stronger than $H_0$ since $V_T \perp \!\!\! \perp (X,Z)$ by construction. We propose two bootstrap procedures that only differ in the construction of the bootstrap censoring times. Procedure A is the standard bootstrap for right censored data efron1981censored, in which both the duration and censoring times are generated from (conditional) Kaplan-Meier fits. Procedure B is based on the conditional random censoring algorithm robinson1983bootstrap. In this approach, the known censoring times from the original sample are kept for the censored observations. For the uncensored observations, censoring times are generated from a conditional Kaplan-Meier fit, conditionally on the information that the unknown censoring time is greater than the observed duration time. In Section (ref), we investigate whether there are scenarios under which one of the bootstrapping approaches performs better. The two procedures (Type A and B) are implemented as follows.
For each $b \in \{1,\dots,B\}:$
Finally, compute the $p-$value as $B^{-1}\sum^B_{b=1} \mathbbm{1}\left\{T^*_{n,b} > T_{n}\right\}$.
In this section, we investigate the finite-sample performance of the proposed tests described in Section (ref) and both of the bootstrap approximations described in Section (ref) through Monte Carlo simulations. To estimate the Cramér-von Mises test statistic, we set $\widehat{\pi}(x,w) = 1$ for all simulation settings. Moreover, we implement the warp-speed method of warpspeed to obtain the critical values. The procedure calculates a single bootstrap test statistic for each Monte Carlo sample and aggregates these statistics across simulations to approximate the critical value, which significantly reduces computation time. We will begin by describing the data-generating process, followed by a discussion of the proposed tests' performance under different degrees of endogeneity, censoring and instrument strength.
For $i=1,\dots,n$, let $$U_{T,i} \sim \mathcal{U}[0,1], \text{ } U_{C,i} \sim \mathcal{U}[0,1]\text{ and } X_i\sim \text{Bin}(0.45), $$ such that $$\pi_{W,i}=\frac{\exp(0.9 - 0.3 X_i)}{1+\exp(0.9 - 0.3 X_i)},$$ and $W_i \sim \text{Bin}(\pi_{W,i})$. Moreover, let $$\pi_{Z,i}=\frac{\exp(-2+0.2X_i+\eta W_i+\alpha \bar{U}_{T,i})}{1+\exp(-2+0.2X_i+\eta W_i+\alpha \bar{U}_{T,i})},$$ with $\bar{U}_{T,i} = U_{T,i}-0.5$ and $Z_i \sim \text{Bin}(\pi_{Z,i})$. The reason for centering $U_{T,i}$ is to keep the censoring rate somewhat stable when we vary $\alpha$, which determines the endogeneity of $Z$. This will allow us to compare the power of the test statistics under different degrees of endogeneity. Moreover, $Z$ is exogenous when $\alpha=0$ and $\eta$ controls the instrument strength. Further, let $$ T_i = \exp\left\{4-0.5X_i-Z_i +\Phi^{-1}(U_{T,i})\right\}, $$ with $\Phi^{-1}$ the inverse of the standard normal distribution. Furthermore, let $$C_i=-\frac{\log(1-U_{C,i})}{\exp\left\{\lambda+0.9X_i+0.8Z_i\right\}}.$$ Note that $\lambda$ controls the censoring rate. Lastly, we generate $Y_i=\min(T_i,C_i)$ and $\Delta_i=\mathbbm{1}(Y_i=T_i)$ to get the simulated data $\left\{Y_i,\Delta_i,X_i,W_i,Z_i\right\}_{i=1,\dots,n}$. Even though we only have one covariate $X$, the parameter values were chosen such that there are very few data points for which $W_i=0$ and $Z_i=1$. On average, 1.9% of the simulated data have $(X_i,W_i,Z_i)=(0,0,1)$ and 2.3% have $(X_i,W_i,Z_i)=(1,0,1)$. This is a similar situation to the empirical application described in Section (ref). For each of the following simulation settings, we performed 1000 Monte Carlo replications.
To generate the data used in this part of the simulation study, we let $(\eta,\lambda)=(2.4,-5.7)$. By choosing these parameter values, we have an instrument strength of around $\tau_{WZ}=0.45$ (measured by Kendall's Tau between $W$ and $Z$) and approximately 25% censoring. We use Kendall's tau as a proxy for instrument strength because it is simple to interpret and report. Note that these values were chosen such that the instrument strength and censoring rate are comparable to the empirical strata examined in Section (ref).
Looking at Figure (ref), we see that both the Kolmogorov-Smirnov (KS) and Cramér-von Mises (CM) test statistics are able to discriminate the null from the alternative hypotheses. It is also clear that the CM test statistic performs better, where for a sample size of $n=1000$, the distributions for $\alpha = 0$ and $\alpha =5$ are almost completely separated. Figure (ref) shows the power of the two test statistics as a function of the degree of endogeneity $\tau_{ZU}$ (measured by Kendall's Tau between $Z$ and $U_T$) for different sample sizes and both of the proposed bootstrap approximations. Again, we see that the CM test statistic outperforms the KS test statistic. As the endogeneity and/or the sample size increase, so does the power of the test statistics. Under the null $(\tau_{ZU} = 0)$, both test statistics reject at the nominal level of 5%, independent of the sample size. Comparing the two bootstrap approximations, there seems to be no noticeable difference. Looking at Figure (ref), we see that the $p$-values obtained from Monte Carlo are well approximated by both bootstrapping procedures. However, approach B seems to be slightly more conservative. Overall, we find that our test statistics have good power, even when there are almost no observations with $W=0$ and $Z=1$. Moreover, the CM test statistic consistently outperforms the KS test statistic regardless of sample size or the degree of endogeneity.
In this subsection, we will separately vary $\eta$ and $\lambda$ alongside $\alpha$ to assess the influence of the instrument strength and censoring rate, respectively, on the power of our proposed test statistics. All simulations use a sample size of $n=1000$ and the other parameters have the same value as described in Section (ref). We start by varying the instrument strength $\eta \in \{0.4,0.9,1.4,1.9,2.4,2.9,3.4\}$ and fixing $\lambda =-5.7$, such that there is around 25% censoring. Figure (ref) shows us that the power increases when the instrument strength increases, as would be expected. Moreover, when there is weak endogeneity ($\alpha =2$), the increase in power becomes less as the instrument strength increases. The CM test statistic again consistently outperforms the KS test statistic. Concerning the bootstrap approximations, it can be seen that approach B performs slightly worse than approach A for the KM test statistic when there is weak endogeneity. Under the null, the rejection rate remains at the nominal 5% level for both approximations, independent of the instrument strength.
To examine the impact of censoring, we vary the parameter $\lambda \in \{-5.7,-5.1,-4.6,-4.2,-3.8\}$ and fix $\eta=2.4$ such that there is an instrument strength of around $\tau_{WZ}=0.45$. As we would expect, Figure (ref) indicates that the power of both test statistics decreases when there is more censoring. However, until approximately 50% censoring, the power of the CM test statistic remains relatively stable. When there is weak endogeneity, we see that bootstrap approach B performs slightly better when the censoring rate is above 50%. Interestingly, the CM test statistic performs a little worse than the KM test statistic when there is weak endogeneity and high censoring. The rejection rate under the null remains somewhat stable as the censoring rate increases, but falls slightly below 5% when the censoring rate is around 75%. Overall, the test statistics still perform well despite reduced instrument strength and high censoring rates.
The data come from the National Job Training Partnership Act (JTPA) Study, a large randomized evaluation of more than 600 federally funded programs designed to increase the employability of eligible adults and out-of-school youths. The study was conducted between 1987 and 1989, enrolled over 20000 applicants and collected follow-up information on employment, earnings, and program participation. Random assignment placed unemployed individuals into either a treatment group (eligible for JTPA services) or a control group (ineligible for 18 months). Nevertheless, roughly 3% of control-group members received JTPA services despite their ineligibility. Because individuals may self-select into treatment in a nonrandom way, noncompliance can induce endogeneity of the treatment angrist1996identification. All participants were surveyed between 12 and 36 months after randomization (average of 21 months). A second follow-up survey was administered to a subsample of 5,468 respondents, focusing on the interval between the two interviews, and took place between 23 and 48 months after the initial randomization. The data can be downloaded at \url{https://www.upjohn.org/data-tools/employment-research-data-center/national-jtpa-study}.
Because the study combines a clear policy intervention with extensive follow-up on earnings and labor-market outcomes, the data have become a standard for evaluating causal estimands in both cross-sectional and duration contexts. In particular, bloom1997benefits, abadie2002instrumental and wuthrich2020comparison investigate the impact of JTPA services on the sum of earnings after treatment. The effect of JTPA services on unemployment duration has been investigated by frandsen2015treatment and beyhum2024instrumental under the (conditional) independent censoring assumption, while crommen2024instrumental and crommen2025estimation allow for dependent censoring. To make the independent censoring assumption more plausible, frandsen2015treatment and beyhum2024instrumental rely exclusively on data from the first follow-up survey. Because almost all participants were surveyed, it is reasonable to assume that the censoring is administrative. By contrast, crommen2024instrumental and crommen2025estimation incorporate data from the second follow-up survey. They argue that participants who were invited to a second survey but did not participate could introduce endogenous censoring. Because the validity of our test statistics relies on the conditional independent censoring assumption, we restrict our empirical analysis to data from the first follow-up interview. Our goal is to test whether instrumental variable methods are required to identify the causal effect of JTPA services on unemployment duration, i.e., if JTPA services are an endogenous treatment.
Due to the size of the data set, we will focus our attention on two strata. The first stratum consists of 1127 single white men without children who are 30 years or younger and reported having no job at the time of treatment assignment. The other stratum is similar to the first one, except that the 1017 men are non-white. We will refer to these strata as white and non-white men, respectively. The covariate HSGED equals 1 if an individual held a high school diploma or GED at the time of treatment assignment and 0 otherwise. Approximately 46% of white men and 40% of non-white men had a high school diploma or obtained a GED at the time of treatment assignment. Moreover, treatment assignment will be our instrument $W$ ($W=1$ if assigned to the treatment group, $0$ for the control group). We deem treatment assignment to be a valid instrument as it is randomly assigned, correlated with JTPA participation (Kendall's Tau of 0.457 for white men and 0.475 for non-white men) and influences time to employment only through treatment participation. Note that $Z=1$ if they participated in JTPA services, and $0$ otherwise. Around 68% of white men and 71% of non-white men were assigned to the treatment group. In both strata, only 47% actually participated in JTPA services.
Figure (ref) displays histograms of the observed follow-up times for each stratum, where a darker shading indicates a higher censoring rate. Consistent with frandsen2015treatment, we find that most observations are censored at approximately 600 days, around which time most of the follow-up interviews took place. Beyond 600 days, nearly all observations are censored. Moreover, Table (ref) shows the censoring rates for each level of HSGED, $W$ and $Z$. keeping HSGED and $Z$ fixed, we find that the censoring rate remains stable for different levels of $W$. Therefore, we do not find any major violations of Assumption (ref). Table (ref) also shows the number of observations for each level of HSGED, $W$ and $Z$. As expected, there are very few observations from the control group who participated in JTPA services, i.e., $W=0$ and $Z=1$. However, as shown in Section (ref), the test still performs well when there are only a few observations with $W=0$ and $Z=1$.
We compute the $p$-values of both test statistics using the two bootstrap approximations described in Section (ref), based on 1000 resamples. Similarly to Section (ref), we set $\widehat{\pi}(x,w)=1$ for the Cramér–von Mises test statistic. The results are reported in Table (ref) and Figure (ref) displays the Kaplan–Meier curves, conditional on high school diploma/GED status and treatment participation, for both the white and non-white men. For the stratum of white men, we fail to reject the null hypothesis that $Z$ is exogenous. Therefore, we can simply compare the estimated Kaplan-Meier curves for the treated and untreated. A log-rank test indicates no statistically significant difference between the survival curves ($p$-value of 0.330). Even after conditioning on having a high school diploma or GED, no significant difference is found ($p$-value of 0.503 for both levels of HSGED). We therefore conclude that, within this stratum, there is no evidence of a significant effect of JTPA services on unemployment duration.
By contrast, for the stratum of non-white men, the null hypothesis that $Z$ is exogenous is rejected. This suggests the presence of unobserved heterogeneity influencing both treatment participation and unemployment duration. A log-rank test further reveals a significant difference between the survival functions for the treated and untreated ($p$-value = 0.0003). We find that, for non-white men without a high school diploma or GED, the treated have a significantly lower probability of being unemployed compared to the untreated ($p$-value of $8.87\times 10^{-5}$). However, for non-white men with a high school diploma or GED, there seems to be no significant difference ($p$-value of 0.403). It is important to emphasize that because $Z$ is endogenous in this stratum, we cannot interpret the observed survival differences as causal effects of JTPA participation. The effect of JTPA services among non-white men without a high school diploma or GED may, for example, reflect self-selection of more motivated or higher-ability individuals into treatment. To actually estimate the causal effect of JTPA participation under endogeneity, one would need to use instrumental variable methods that allow for right censoring. Possible approaches include the nonparametric local average treatment effect framework of frandsen2015treatment, as well as the population-level treatment effect methods developed by beyhum2022nonparametric,beyhum2024instrumental, which are nonparametric and semiparametric, respectively.
This paper develops nonparametric tests for exogeneity in nonparametric nonseparable duration models subject to right censoring. The tests are based on the independence between the conditional ranks, which are also right censored, and the instrumental variable. The proposed tests do not require estimation of the underlying structural model using nonparametric instrumental variable methods, which is even more challenging in the presence of right censoring. We derive the asymptotic properties of the proposed tests and illustrate their power under various scenarios using Monte Carlo simulations. The validity of the two bootstrap approaches to approximate the critical values is also analyzed via simulations. We find that the Cramér–von Mises statistic consistently outperforms the Kolmogorov–Smirnov statistic, and that both bootstrapping approaches perform equally well in almost all scenarios. An empirical application to the National Job Training Partnership Act (JTPA) Study illustrates the usefulness of the tests in applied settings. In particular, it provides practitioners with a test to decide if standard survival comparisons are appropriate or whether instrumental variable methods that allow for right censoring are needed to correctly identify treatment effects.
The tests can be extended in multiple directions. Firstly, extending the methodology to accommodate continuous covariates, instruments and treatments would substantially broaden its applicability. The most straightforward way to achieve this is by replacing the conditional Kaplan–Meier estimator with a smooth conditional estimator, such as beran1981nonparametric's estimator. Although conceptually straightforward, this extension introduces both theoretical and practical complications, such as bandwidth selection and the curse of dimensionality. A way of dealing with this could be to instead use semiparametric models, such as the CoxPHmodel proportional hazards model. Another important extension is to relax the conditional independent censoring assumption and allow for some forms of dependence between $T$ and $C$ after conditioning on $(X,W,Z)$. Copula-based methods, such as the copula-graphic estimator Rivest2001AMA,ZHENGMING1995Eoms, offer a natural approach to modeling such dependencies. However, other copula-based methods for dependent censoring could also be considered (see crommen2025recent for a recent review).
The computational resources and services used in this work were provided by the VSC (Flemish Supercomputer Center), funded by the Research Foundation Flanders (FWO) and the Flemish Government department EWI. G. Crommen is funded by a PhD fellowship from the Research Foundation - Flanders (grant number 11PKA24N). J.P. Florens acknowledges funding from the French National Research Agency (ANR) under the Investments for the Future (Investissement d'Avenir), grant ANR-17-EURE-0010. I. Van Keilegom acknowledges funding from the FWO and F.R.S. - FNRS (Excellence of Science programme, project ASTeRISK, grant no. 40007517), and from the FWO (senior research projects fundamental research, grant no. G047524N). Moreover, the authors would like to thank Jad Beyhum for helpful discussions.