Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
59,496 characters · 13 sections · 54 citation commands
Instrumental variable quantile regression under random right censoring
\corrfalse
Let us consider a setting where the researcher is interested in the causal effect of some regressors $Z$ on a duration outcome $T$. We wish to recover the structural quantile of $T$ if the treatment were set to a particular value. This task is often complicated by two issues. The first one is endogeneity, that is the variable $Z$ and the unobserved heterogeneity $U$ are dependent. In this case, the causal effect of the treatment is not characterized by the conditional distribution of $T$ given $Z$. The second problem is right censoring. It happens, for instance, when the period of observation of the subjects in the dataset is limited. If the length of the follow-up does not depend on the subject's characteristics, the censoring is called independent.
In this paper, we propose a new instrumental variable estimator which allows treating both issues. The endogeneity of the treatment is addressed through an instrumental variable $W$, independent of the error term $U$ of the model but sufficiently related to the treatment. We assume that censoring is random and independent. The structural quantile of $\log T$ is linear in the variables, which makes the model semiparametric. The causal effect at a given quantile $u$ is therefore characterized by a finite dimensional parameter vector $\beta_0(u)$. This model allows to handle several continuous or discrete variables and is standard in the literature on censored quantile regression (see below for related literature).
Specifically, we show that $\beta_0(u)$ is the solution to a continuum of integral equations. This result allows us to derive local and global identification results. Then, we propose a new minimum distance estimator solving an estimated version of the system of identification equations. {\ifcorr\color{red} \else \color{black} \fi We address censoring through a weighting scheme.} {\ifcorr\color{red} \else \color{black} \fi The estimator is asymptotically normal under some conditions}. {\ifcorr\color{red} \else \color{black} \fi We prove the validity of a bootstrap procedure for inference}. Finite {\ifcorr\color{red} \else \color{black} \fi sample} properties of the estimator are assessed through simulations. The procedure is illustrated {\ifcorr\color{red} \else \color{black} \fi through an} application to the national Job Training Partnership Act (JTPA) study. An appealing feature of our approach is that the instrument and the covariates can be either discrete or continuous. \\
\noindentRelated literature This paper is first related to the literature on censored quantile regression with exogenous regressors, which studies a model where the conditional quantile of $\log(T)$ is linear in the regressors without relying on instrumental variables. Earlier works, such as powell1984least,powell1986censored, khan2001two, fitzenberger199715, proposed estimators for this model in the special case where the censoring time is constant and observed. An alternative approach is buchinsky1998alternative where the censoring time is an unknown function of the regressors but does not need to be always observed. honore2002quantile later allowed for random right censoring independent of the outcome variable and the regressors. Then, chernozhukov2002three relaxed this condition by assuming only that the duration variable and the censoring time are independent conditional on the regressors. Both honore2002quantile and chernozhukov2002three required the censoring time to be observed. The approach developed in portnoy2003censored relies on the same independence assumption as chernozhukov2002three but allows the censoring time to be unobserved for uncensored observations. Alternative methods have been proposed in peng2008survival, which is based on martingales, or in yang2018new, which follows a data augmentation approach. The methods of portnoy2003censored, peng2008survival and yang2018new rely on the global assumption that all the conditional quantiles of $\log(T)$ are linear in the regressors. wang2009locally relaxed this global assumption by assuming linearity only for the quantile of interest (the approach of the present paper would also work under a local linearity assumption of this type). Other estimators which only impose similar local conditions are de2019adapted, which is based on adjusting the standard quantile loss function in order to accommodate randomly censored data, and de2020linear, who propose a minimum-distance estimator. For more details about each of these estimators, we refer the reader to the recent review by Peng2021.
There is also a large literature on instrumental variable approaches with randomly right censored duration outcomes, where it is not assumed that the quantile of $\log T$ is linear in the regressors. We only cite here some works for reasons of brevity. {\ifcorr\color{red} \else \color{black} \fi Some articles study different semiparametric models, see e.g. tchetgen2015instrumental for an additive hazard model or martinussen2019instrumental and wang2018learning for the Cox model}. On the other hand, frandsen2015treatment, sant2016program, richardson2017nonparametric, blanco2019bounds, sant2021nonparametric, beyhum2022nonparametric have developed nonparametric approaches with categorical treatment and instrument, whereas centorrinoflorens2021 proposed a nonparametric estimator when the variables are continuous and the model is additive.
Finally, this paper is most related to the literature on instrumental variable methods under right censoring assuming that the structural quantile of $\log T$ is linear in the regressors. First, there are papers, such as blundell2007censored, hong2003inference, chen2020semiparametric and wang2021moment, assuming that the censoring time $C$ is constant. In contrast, the present article allows $C$ to be random. Such a distinction is particularly relevant in studies where the duration of follow-up depends on the date of entry in the study. Next, although censoring is random in chernozhukov2015quantile, they assumed that $C$ is always observed, which we do not. Their approach is based on a control function requiring a separate specification for the relation between the regressors and the instrument. We do not need these additional structural assumptions. Also, they assume that the endogenous regressor is continuous, while our approach allows both discrete and continuous cases. {\ifcorr\color{red} \else \color{black} \fi khan2009inference present an alternative estimator}. Their method is designed to handle dependent censoring. It requires a restrictive support conditions {\ifcorr\color{red} \else \color{black} \fi (see condition IV2, page 110 in khan2009inference)}, which may not hold when censoring is independent. We do not need such a support assumption. Moreover, wei2021estimation studied the case where there are exogenous covariates, a single binary endogenous variable and a binary instrument. They imposed a monotonicity assumption (as in angrist1996identification), which means that the value of the treatment is increasing in the instrument. They identified and estimated the quantile treatment effects over the population of compliers, that is the subjects whose treatment status changes with the instrument. Instead, the present paper allows for nonbinary endogenous regressors and instruments. Additionally, thanks to a rank invariance assumption as in chernozhukov2005iv, we identify the quantile treatment effects on the whole population rather than only on the compliers.\\
\noindentOutline The paper is organised as follows. The model specification is given in Section 2. Identification results are presented in Section 3. Section 4 is devoted to estimation and inference. Sections 5 and 6 present the simulations and the empirical application, respectively. All proofs can be found in the supplementary material. The code for the simulations and the empirical application is in the replication package.
\setcounter{equation}{0} \setcounter{theorem}{0} \setcounter{definition}{0} Denote by $T$ the duration outcome variable with values in $\mathbb{R}_+$, by $Z$ a random vector of regressors with support $\mathcal{Z}\subset \mathbb{R}^K$. Let also $T(z)$ be the potential outcome of $T$ under treatment $z\in \mathcal{Z}$. By consistency, it holds that $T=T(Z)$. The random variable $U$ is the unobserved heterogeneity of the model. We normalize it to follow a uniform distribution on the interval $[0,1]$. We suppose that there exists quantile linear regression coefficients $\beta_0(\cdot):(0,1) \rightarrow\mathbb{R}^K$ such that the following relationship among the potential outcomes holds:
We also assume that the mapping $\beta_0(\cdot)$ is continuously differentiable everywhere with derivative denoted by $D\beta_0(\cdot)$ and for all $u\in (0,1)$ and $z \in \mathcal{Z}$, $z^\top D\beta_0(u)>0$. This ensures that $z^\top\beta_0(\cdot)$ is a well-behaved quantile function, in the sense that $u\in(0,1)\mapsto z^\top \beta_0(u)$ is strictly increasing. Under this last condition, for any value $u\in(0,1)$, the $u$-quantile of $\log T(z)$ is equal to $z^{\top}\beta_0(u)$, and $\beta_0(u)$ measures the causal effect of $Z$ on this structural $u$-quantile. Our goal is to identify and estimate $\beta_0(u)$, for some given $u\in(0,1)$. Note that we do not define $\beta_0(\cdot)$ on $\{0,1\}$. This choice avoids pathological behaviours happening at the boundaries and is innocuous since $P(U\in\{0,1\})=0$. Remark also that the approach proposed in this paper would still work if the quantile of $\log T(z)$ were linear only in a neighbourhood of the quantile $u$ of interest. We only assume linearity for all quantiles in $(0,1)$ in order to simplify the exposition.
The random vector $Z$ can contain both exogenous and endogenous variables. We possess an instrumental variable, denoted by $W$, with support $\mathcal{W}\subset \mathbb{R}^L$, that is independent of $U$. All the exogenous variables in $Z$ are included in $W$. The distributions of $Z$ and $W$ can be discrete or continuous. We also assume that the distribution of $U$ given $Z$ and $W$ is absolutely continuous. Note that, by the inverse function theorem and equation (ref), this continuity condition guarantees that the distribution of $T$ given $Z$ and $W$ is absolutely continuous too.
In addition, the outcome variable $T$ is considered to be right censored by a random variable $C$ with values in $\mathbb{R}_{+}$. Define $Y = \min(T,C)$ and $\delta = \mathds{1}\{ T\le C\}$, where $\mathds{1}\{\cdot\}$ denotes the indicator function. The observables are $(Y,\delta, Z,W)$.
Note that our model (ref) implies that
where $P(\epsilon(u)\le 0|W) = u$. Indeed, defining $\epsilon(u)= \log T(z) - z^\top \beta_0(u)$, we have $P(\epsilon(u)\le 0|W)= P(z^\top\beta_0(U)\le z^\top\beta_0(u)|W) = P(U\le u|W)=u$. Formulation (ref) corresponds to the way the quantile regression model is often written in the literature, see e.g. honore2002quantile,khan2009inference,wang2009locally.
We now summarize the conditions we imposed in this section, for the value $u\in(0,1)$ of interest.
\setcounter{equation}{0} \setcounter{theorem}{0} Now, we consider the identification of the parameter of interest $\beta_0(u)$. Using equation (ref) and the condition that $W \protect\mathpalette{\protect\independenT}{\perp} U$, it is possible to show that $\beta_0(u)$ is the solution to a continuum of identification equations, presented in the next subsection. This argument will lead to the aimed identification results, for both the case of censored and uncensored outcomes.
We use model (ref) to show that $\beta_0(u)$ is a solution of the following system of equations in $\beta \in\mathbb{R}^K$:
Indeed, using, the specification of $T$ in model (ref), Assumption (ref) (a), and the fact that $U|W\sim U[0,1]$, we obtain the following equalities:
{\ifcorr\color{red} \else \color{black} \fi In the next sections, we derive identification results for the parameter of interest based on the system of equations (ref)}.
In this section, to simplify the exposition, we derive identification results without taking into account the censoring mechanism ($C=\infty$). In this case, the identification properties are more readily obtained since the left-hand side of Equation (ref) is identified. This implies that, in this context, studying identification is equivalent to assessing the uniqueness of the solutions to (ref). Identification with censoring is discussed in Section (ref).
The present section contains two types of identification results (both without censoring as mentioned in the above paragraph). We first show identification under a general framework in Section (ref). The identification conditions in Section (ref) are technical and abstract, but standard in the literature on instrumental variable quantile regression models (see chernozhukov2005iv or feve2018estimation, among others). Next, we provide simpler and more interpretable identification results in the case of randomized experiments with noncompliance in Section (ref).
Let the parameter space $S$ be the set of continuously differentiable mappings $\beta(\cdot):(0,1)\xrightarrow{}\mathbb{R}^K$ with derivative denoted by $D\beta(\cdot)$ such that for all $u\in (0,1)$ and $z \in \mathcal{Z}$, $z^\top D\beta_0(u)>0$. The parameter $\beta_0(u)$ is identified (when there is no censoring) if, for $\beta(\cdot)\in S$, $$ E[\mathds{1}\{T\le\exp(Z^\top\beta(u))\}|W=w] = u\quad \text{for all}\quad w\in\mathcal{W}, $$ implies that $\beta(u)=\beta_0(u)$.
Rewrite the conditional distribution function $R(u|z,w)$ and density function $r(u|z,w)$ of $U$ given $Z=z,W=w$ as
where, $F_{T|Z,W}$ and $f_{T|Z,W}$ are, respectively, the cumulative distribution function and density of $T$ given $Z,W$.
Let $\Delta$ be a mapping from $\mathcal{Z}\times(0,1)$ to $\mathbb{R}$ differentiable in its second argument. For $z\in\mathcal{Z}$ and $u\in (0,1)$, we use the notation $\Delta_z(u)=\Delta(z,u)$. The derivative of $\Delta_z(\cdot)$ with respect to $u$ is denoted by $D\Delta_z(\cdot)$. For $\alpha\in[0,1]$, define the $\alpha$-perturbation of the quantities $R(u|z,w)$ and $r(u|z,w)$ in the direction $\Delta_z(u)\in \mathbb{R}$ as
We now assume that $Z$ is strongly complete by $W$ given $U=u$.
This type of strong completeness condition is studied in chernozhukov2005iv, both in the case where $Z$ and $W$ are continuous and the case where they are both categorical. It is an abstract condition requiring a certain degree of dependence between $Z$ and $W$. The next theorem contains the global identification result.
By identification here, we mean that $\beta_0(\cdot)$ is the unique solution in $S$ of the system of equations (ref). Note that we require that $E[ZZ^\top]$ has full rank. {\ifcorr\color{red} \else \color{black} \fi We need this assumption due to the linear relation between $\log(T)$ and the covariate $Z$ in model (ref).}
The conditions of Theorem (ref) are not minimal. Indeed, in some special cases of applied relevance, it is possible to obtain simpler and more interpretable identification conditions by following arguments similar to that of chernozhukov2005iv. We focus on the case where the treatment $Z$ and the instrument $W$ are binary. Such a setting corresponds to randomized experiments with noncompliance which naturally occur in empirical applications. An example is the JTPA experiment discussed in Section (ref). Note that, it would be possible to include an intercept or exogenous covariates in the analysis but, for simplicity, we decided to avoid it.
Let $\nu,\underline{f}>0$ be small constants. We define the set $\mathcal{L}$ by the closed rectangle of vectors $(t_0,t_1)\in\mathbb{R}^2$ satisfying the following conditions, where $t_{Z}=(1-Z)t_0+ Zt_1$:
We want to show that there exists a unique $t=(t_0,t_1)^\top \in \mathcal{L}$ such that $\Pi(t)=0$, where $$ \Pi(t) =
. $$ The Jacobian of $\Pi(t)$ with respect to $t$, denoted by $\Pi'(t)$, takes the following form:
where $f_{T,Z|W}(t,z|W)= f_{T|Z,W}(t|z,w)P(Z=z|W=w).$ We make the following Assumption.
Assumption (ref) (a) guarantees that $(0,\beta_0(u) )$ belongs to $\mathcal{L}$. It means that for every $z,w$ the density of $U$ given $Z=z,W=w$ is strictly positive. Assumption (ref) (b) is a full rank condition, which is standard in econometrics. Similarly as in chernozhukov2005iv, we can provide the following interpretation of Assumption (ref) (b). The matrix $\Pi'(t)$ has full rank for all $t\in\mathcal{L}$ if and only if
(or the same property with $<$ instead of $>$). This inequality can be interpreted as a monotone likelihood ratio condition: for all $t\in\mathcal{L}$, the instrument increases (or decreases) the likelihood ratio in (ref).
Let us also consider the special case of randomized experiments with one-sided noncompliance where $P(Z = 1 | W = 0) = 0$ (this equality approximately holds in the JTPA empirical application). Then, it can be seen that Assumption (ref) holds as long as $$P(Z = 1 | W = 1, U = u) > 0.$$ This condition means that subjects for which $U=u$ have a strictly positive probability to be treated when assigned to the treatment group.
We have the following theorem.
Notice that a similar result would hold in the case where there are additional exogenous covariates $X$, that is when $Z = (Z_1,X^\top)^\top$, where $Z_1\in\{0,1\}$ and $W = (W_1,X^\top)^\top$, where $W_1\in\{0,1\}$. The only difference would be that all the probabilities and densities in the definition of $\mathcal{L}$ and Assumption (ref) would be conditional on $X=x$ (for $x$ in the support of $X$), Assumption (ref) would need to hold for all $x$ in the support of $X$, and we would have to assume also that $E[XX^\top]$ has full rank.
In this section, we extend the identification results previously presented to the case where there is censoring. We assume that the censoring time $C$ satisfies the following assumption.
Assumption (ref) in particular implies that $T$ and $C$ are independent. {\ifcorr\color{red} \else \color{black} \fi We could replace it }by $C\protect\mathpalette{\protect\independenT}{\perp} U|Z,W$ but the latter assumption would complexify estimation and identification. Denote by $\bar c$ the upper bound of the support of $C$, and define $\bar u$ as
As we now explain, if $\bar u>0$, it is possible to identify the left-hand side of equation (ref) for any value of $u\in(0,\bar u)$. Let $\beta\in \mathbb{R}^K$ be such that $\exp(z^\top \beta)< \bar c$ for all $z\in\mathcal{Z}$. We have
where $G(s)=P(C\ge s)$ is the survival function of the censoring variable $C$. Note that, by definition of $\bar c$, we have $G(Y)>0$ on the event $\{Y\le \exp(Z^\top\beta)\}$. Hence, the left-hand side of (ref) is well-defined.
We can justify equation (ref), using that $\delta = 1$ corresponds to $Y = T$ and so {\allowdisplaybreaks
}
where in the last equality we use Assumption (ref), which implies that $E[\mathds{1}\{T\le C\} |T,Z,W] = G(T)$. {\ifcorr\color{red} \else \color{black} \fi We can give the following interpretation of equality (ref)}. The right-hand side of (ref) is an average of indicator functions $\mathds{1}\{T\le\exp(Z^\top\beta)\}$ which are only observed when $\delta=1$. We can replace this average of $\mathds{1}\{T\le\exp(Z^\top\beta)\}$ by the average of the observed indicators $ \mathds{1}\{Y\le\exp(Z^\top\beta)\}$, weighted by $\delta/G(Y)$ to make them representative of the full sample.
Notice that $G(t)$ is identified (from the distribution of $(Y,\delta)$) for all $t\in[0,\bar c]$ by standard arguments from the survival analysis literature. Hence, since $(Y,\delta, Z,W)$ are observed, equation (ref) implies that the left-hand side of equation (ref) is identified for all $\beta\in \mathbb{R}^K$ such that $\exp(z^\top \beta)< \bar c$ for all $z\in\mathcal{Z}$, and so, in particular, for $\beta_0(u)$ for all $u\in (0,\bar u)$. Therefore, all the identification results previously discussed can be readily adapted to obtain identification under censoring for all $u\in (0,\bar u)$.
\setcounter{equation}{0} \setcounter{theorem}{0} In this section, we provide an estimation procedure for the parameter vector $\beta_0(u)$. We possess an i.i.d. sample of size $n$ of the observables $(Y,\delta,Z, W)$, denoted by $\{(Y_i,\delta_i, Z_i, W_i)\}_{i=1}^n$.\\ Under Assumption (ref) and using equation (ref), equation (ref) is equivalent to $$ E\Bigg[\frac{\delta}{G(Y)} \mathds{1} \{Y\le \exp(Z^\top\beta(u))\}\bigg|W=w\Bigg] = u,\ \text{for all } w\in\mathcal{W}.$$ This continuum of conditional moment restrictions is equivalent to the following system of unconditional moment restrictions:
where by the event $\{W\le w\}$, we mean $\cap_{\ell=1}^L \{W_\ell\le w_\ell\}$. This new set of equations allows to design an estimation procedure which does not involve complex smoothing techniques. Let us define the following operator: $$ \mathcal{A}_u(\beta,w) = E\Bigg[\frac{\delta}{G(Y)} \mathds{1} \{Y\le \exp(Z^\top\beta), W\le w\}\Bigg] - uP(W\le w). $$ Consider now the following estimator $\hat{\mathcal{A}}_u(\beta,w)$ of $\mathcal{A}_u(\beta,w)$: $$ \hat{\mathcal{A}}_u(\beta,w) = \frac{1}{n}\sum_{i=1}^n \frac{\delta_i}{\hat G(Y_i)} \mathds{1} \{Y_i\le\exp(Z^\top_i\beta) , W_i\le w\} - \frac{u}{n}\sum_{i=1}^n\mathds{1}(W_i\le w), $$ where $\hat G(t)$ is the Kaplan-Meier estimator of $G$, that is
with $N(t) = \sum_{i=1}^n\mathds{1}\{Y_i\le t, \delta_i = 0\}$, $dN(s) = N(s)-\lim_{s'\xrightarrow{}s,s'<s}N(s')$ which denotes the jump of the process $N$ at the time $s$, and $Y(s) = \sum_{i=1}^n\mathds{1}(Y_i\ge s)$. Note that $\hat{\mathcal{A}}_u(\beta,w)$ is an average using only the uncensored observations which are made (approximately) representative of the full sample thanks to the weights $\{\delta_i/\hat{G}(Y_i)\}_{i=1}^n$. Our estimator $\hat\beta(u)$ of $\beta_0(u)$ is
where $\mathcal{B}$ is a compact parameter set in $\mathbb{R}^K$. This is a minimum distance estimator which attempts to solve estimated versions of the identification equation (ref) at the points $W_1,\dots,W_n$. Minimum distance estimators have been studied in Econometrics, see for instance brown2002weighted, poirier2017efficient, torgovitsky2017minimum. Our estimator has three distinctive features with respect to these papers. It concerns our specific semiparametric instrumental variable model, it deals with censoring and it solves the equation at the empirical distribution of $W$ rather than at a prespecified distribution. The last property avoids to let the econometrician choose the points (and the weights of these points) at which the equations are solved. {\ifcorr\color{red} \else \color{black} \fi We may be able to construct an alternative estimator attaining the semiparametric efficiency bound using the approach of poirier2017efficient}. This estimator would however be more complicated and rely on features that are chosen by the econometrician.
Define the map $L:\mathcal{B}\xrightarrow{}\mathbb{R}$, where $$ {L}(\beta) = E\left[ \mathcal{A}^2_u(\beta,W)\right], $$ and let $\bar t$ be the upper bound of the support of $T$. Consider also the following condition.
In Assumption (ref) (g), $\|\cdot\|$ corresponds to the Euclidean norm. Assumptions (ref) (a), (b) are standard in the literature on $M-$estimators (see newey1994large). Assumptions (ref) (c), (d) and (f) are mild conditions which depend simultaneously on the regularity of the map $\beta \xrightarrow{} \exp(z^\top \beta)$ and on the distribution of $(U,Z,W)$. Note that Assumption (ref) (g) holds when $\mathcal{Z}$ is compact but is more general. Assumption (ref) (e) ensures that $\mathcal{A}_u(\beta,w)$ is identified at all $\beta\in\mathcal{B}$ (another role of this assumption is to avoid running into consistency problems of the Kaplan-Meier estimator on the tails). We have the following theorem.
We present now a bootstrap procedure for inference regarding the proposed estimator. It avoids estimating its asymptotic variance matrix. Denote by $\{(Y_{bi},\delta_{bi}, Z_{bi}, W_{bi})\}_{i=1}^n$ a bootstrap sample drawn with replacement from the original sample $\{(Y_i,\delta_i, Z_i, W_i)\}_{i=1}^n$. In addition, denote by $\hat \beta_b(u)$ the value of the estimator computed on the bootstrap sample. The following asymptotic result holds.
\setcounter{equation}{0} \setcounter{theorem}{0} In this section, we present the results of numerical experiments to analyze the performance of the proposed estimator in finite samples. The data generating process (DGP) is as follows. A random variable $U\sim U[0,1]$ determines the structural quantile. Then, $Z = (1,Z_2,Z_3)^\top$ is a $3\times 1$ random vector, where the component $Z_2$ of $Z$ is endogenous and the component $Z_3$ is exogenous i.e. $Z_3\protect\mathpalette{\protect\independenT}{\perp} U$ but $Z_2\ \cancel{\protect\mathpalette{\protect\independenT}{\perp}}\ U$. Moreover, $W = (1,W_2,Z_3)^\top$ is the instrumental variable. We consider three different designs for the distributions of $W_2, Z_2, Z_3$. They are specified in Table (ref). In the first and third design, the endogenous component $Z_2$ has a discrete distribution, while in the second it is a continuous variable.
The parameter of interest is $\beta_0(u) = (u,u,u)^\top$, and the duration $T$ follows the model (ref), so $ T = \exp(Z^\top \beta_{0}(U))$. For each design, the random censoring time $C$ follows an exponential distribution with parameter $\lambda$ (i.e. its mean is $1/\lambda$). We consider $\lambda\in\{0.0068,0.176\}$ in the first design, $\lambda\in \{0.0173,0.065\}$ in the second design and $\lambda\in\{0.07, 0.175\}$ in the third design. These values of $\lambda$ ensure that there are 20% or 40% of censored observations. Therefore, the support of $C$ is $\mathbb{R}_{+}$ and $\bar u = 1$, where $\bar u$ is defined in (ref).
We estimate $\beta_0(u)$ for quantiles $u\in\{0.3,0.5,0.7\}$ and sample sizes $n\in\{500,1000\}$. We use the algorithm of nelder1965simplex for the minimization of the objective function in (ref). {\ifcorr\color{red} \else \color{black} \fi We search for the minimum} in the compact set $\mathcal{B}=[0, 1]^3$, and the algorithm starts from a random point taken in $\mathcal{B}$ (the initial value follows a uniform distribution on $\mathcal{B}$). {\ifcorr\color{red} \else \color{black} \fi The optimization algorithm starts at 100 random values}. Each of these starting values leads to a local minimum. The final estimate $\hat{\beta}(u)$ corresponds to the local minimum yielding the lowest value of the objective function in (ref). In Table (ref), we report, for each component $\beta_{0,k}$, $(k=1,2,3)$ of $\beta_0$, the bias of the estimator, its root-mean-squared error (RMSE), and the {\ifcorr\color{red} \else \color{black} \fi coverage of $95\%$ confidence intervals constructed by bootstrap percentiles. These results are based on 500 replications.} The RMSE is defined as the squared of the average over the simulations of the euclidean distance between the estimator and $\beta_0(u)$. {\ifcorr\color{red} \else \color{black} \fi We compute the coverage} using the method proposed in giacomini2013warp to speed up the simulations. Lastly, we display the bias and RMSE of an estimator for the standard censored quantile regression ($\hat\beta_{CQR}$) for which the method of wang2009locally is used. This CQR estimator ignores the endogeneity issue.
We see from the results that the bias of the proposed estimator is close to zero. Moreover, when the sample size increases, the bias and the coverage probability converge to their respective theoretical values of 0 and 0.95. The results improve when the proportion of censored observations is lower, and the RMSE tends to increase for higher quantiles. We also observe that the performance of the estimator is similar across the different designs. As expected, the standard censored quantile regression estimator of the endogenous regressor is biased.
In this section, we apply the proposed method to estimate the effect of publicly subsidized job training programs on unemployment durations. The data (jtpaDS) is collected from a large-scale randomized experiment known as the National JTPA Study, designed to evaluate programs funded by the Job Training Partnership Act of 1982. Different authors have analysed this dataset, for instance abadie2002instrumental and frandsen2015treatment.
The experiment started in 1987 and involved around 21,000 economically disadvantaged individuals, who were randomly assigned to a treatment group or a control group. The treated subjects were allowed to enrol in a JTPA-funded training program, while the individuals from the control group were not allowed to enrol for 18 months. Our application focuses on a subset of 802 individuals consisting of non-white single mothers unemployed at the time of randomization, who were surveyed in a follow-up interview taking place between 1 and 3 years following the random assignment.
About 75% of the subjects in the sample were assigned to the treatment group (524 subjects), and 35% to the control group (278 subjects). Among the subjects assigned to the treatment group, about 65% (339 out of 524) participated in a JTPA-funded program. Among the ones assigned to the control group, only 36 subjects (around 12%) enrolled in a JTPA-funded program (they enrolled more than 18 months after randomization, since they were forbidden from enrolling before).
The outcome variable $T$ is the duration (in days) between treatment assignment and finding employment. We are interested in the effect of two covariates on $T$. The regressors are $Z=(1,Z_2,Z_3)^\top$, where $Z_2$ is an indicator for participation in a JTPA-funded program ($Z_2=1$ if the subject participates, $Z_2=0$ otherwise) and $Z_3$ is the age of the subject. We call $Z_2$ the treatment and treat it as endogenous. The covariate age is treated as exogenous. The instrumental variable is $W=(1,W_2,Z_3)^\top$, where $W_2$ is an indicator for treatment group assignment. The validity of $W_2$ as an instrument for $Z_2$ can be justified by the facts that (i) the assignment is random, (ii) being assigned to treatment or control should have no impact on unemployment durations other than through participation in a JTPA-funded program, and (iii) being assigned to the treatment group and enrolling to the JTPA program are dependent.
The outcome variable is only observed for individuals who found a job before the follow-up survey, and it is censored for the other individuals. Hence, we observe $Y=\min(T,C)$ and $\delta= \mathds{1}\{T\le C\}$, where $C$ is the duration between the randomization and the follow-up interview. This censoring time varies between individuals and is therefore random. The follow-up surveys are initiated by the officers in charge of data collection. They contact the subjects participating in the experiment according to a set of rules depending only on the date of entry in the experiment. Hence, the censoring time is primarily determined by the date of randomization. As a result, as long as this date is independent of subjects characteristics, censoring should be independent. frandsen2015treatment provides additional reasons why censoring is likely to be independent.
The average (over uncensored observations) unemployment duration of treatment group members is 20 days lower than that of their control group counterparts. The proportion of censored observations is around 32%. The mode of the unemployment duration (when it is observed) is 26 days and the median is 361 days. A large number of observations of duration is around 600 days. This corresponds to the time around which many follow-up interviews were held. Beyond this point, almost all observations are censored.
To compute the estimator $\hat\beta(u)$ while avoiding local minima, we initialize the optimization algorithm at 1,000 random initialization points. We then select the estimate which minimizes the objective function among the resulting vectors. This procedure is applied for each $u \in \{0.1,0.2,..., 0.9\}$. {\ifcorr\color{red} \else \color{black} \fi We can use the results of the analysis to conduct an informal joint test of our assumptions. Indeed, recall that our identification conditions at $u$ include the fact that
(see Section (ref)). If our identification and estimation conditions were satisfied at $u$, then $\widehat{\beta}(u)$ would converge to $\beta(u)$ and therefore, by (ref), we should expect to have
with high probability. Since (ref) does not hold for the quantiles $u\in\{0.7, 0.8, 0.9\}$, we conclude that our conditions are not satisfied for these quantiles. In order to construct confidence intervals, we use the bootstrap approach discussed in Section (ref). }
{\ifcorr\color{red} \else \color{black} \fi We report the estimates and the 95% (percentile bootstrap) confidence intervals for $\beta_2$ and $\beta_3$ in Figure (ref).} The results indicate that the treatment has a significant and negative effect on $T$. In addition, there is some evidence that the treatment effect varies by quantile. We find a significant and negative effect of age on unemployment duration. In Section S3 of the supplementary material, we include a table reporting the exact values of the estimates and the bounds of the confidence intervals.
Financial support from the European Research Council (2016-2022, Horizon 2020 / ERC grant agreement No.\ 694409) is gratefully acknowledged.