EconBase
← Back to paper

Nonparametric instrumental regression with right censored duration outcomes

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

61,843 characters · 22 sections · 46 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Nonparametric instrumental regression with right censored duration outcomes

abstractThis paper analyzes the effect of a discrete treatment $Z$ on a duration $T$. The treatment is not randomly assigned. The confounding issue is treated using a discrete instrumental variable explaining the treatment and independent of the error term of the model. Our framework is nonparametric and allows for random right censoring. This specification generates a nonlinear inverse problem and the average treatment effect is derived from its solution. We provide local and global identification properties that rely on a nonlinear system of equations. We propose an estimation procedure to solve this system and derive rates of convergence and conditions under which the estimator is asymptotically normal. When censoring makes identification fail, we develop partial identification results. Our estimators exhibit good finite sample properties in simulations. We also apply our methodology to the Illinois Reemployment Bonus Experiment.

{{ Key Words:} Duration Models; Endogeneity; Instrumental variable; Nonseparability; Partial identification.} \\

\setcounter{footnote}{0} \setcounter{equation}{0}

Introduction

\setcounter{equation}{0}

The objective of this paper is to analyze treatment models when the outcome of the treatment is a duration possibly observed with censoring. Let $Z$ be the level of the treatment, the outcome will be a duration $T$, depending on $Z$ and on a random element $U$. If the treatment is randomly assigned, the model is formalized by the conditional distribution of $T$ given $Z$. However, in many cases the assignment mechanism of $Z$ is not independent of $U$ and then the conditional distribution will mix the effect of the treatment and of the assignment mechanism. In the econometric literature we are faced with a usual endogeneity problem, not specific to duration models. The distinguishing feature of this paper is that we introduce a random right censoring mechanism. The censoring duration is only observed for censored observations. We tackle the endogeneity issue using an instrumental variable $W$ independent of $U$ and sufficiently dependent of $Z$ given $U$. The model is nonparametric and nonseparable.

The problem is simplified by considering only the case where both $Z$ and $W$ are categorical and non-dynamic. This avoids the question of ill-posedness of the inversion and of regularity assumptions on the functional parameters. Under some usual conditions of completeness (see CH and FFV), we obtain point identification of the quantile regression function of $T$ on $Z$ in a subset of the parameter space, but outside of this set we can only identify regions of the parameter space. This type of non-unicity of the solution of an inverse problem due to a restriction of the range of the estimated operator seems fully new. Treatment effects on the quantiles of the duration, the survival probabilities and the hazard rates can be derived from the quantile regression function. Because we make a rank invariance assumption (see CH and wuthrich2020comparison), the identified treatment effects concern the whole population and not solely the subset of compliers. We propose a strategy to estimate the quantile regression function of $T$ on $Z$ by solving a nonlinear inverse problem (see D and C). Sufficient conditions under which the proposed estimators are asymptotically normal are derived. Inference results based on a bootstrap approach are given. Our estimation procedure exhibits good finite sample properties in simulations. We apply our methodology to the Illinois reemployment bonus experiment data.

This work borrows from the nonparametric instrumental regression literature (Darolles and CH). We extend the conventional framework to allow for censoring by using a particular conditional moment equation that can be estimated under random right censoring if the support of $T$ is included in the support of $C$ given $Z$, $W$. When the last condition fails, the regression function is only partially identified (manski1990nonparametric and manski2003partial). The identified set is characterized by a mix of conditional moment equalities and inequalities as in andrews2013inference. Note also that instrumental variables are not the only method developed in econometrics to address nonparametrically the endogenity issue, see for instance the control function approach (Newey) or the $g$-calculus (IR).

This paper contributes to the literature on program evaluation under censoring. Some papers introduce a `matching hypothesis' (conditional randomization) by assuming that the assignment of the treatment is conditionally independent of the outcome given a set of observed variables (see VdBBM, VdBBCD and sant2016program). In contrast, AVdB, AVdB2 specify a structural model involving separability assumptions and unobserved heterogeneity terms explaining the endogeneity of the treatment. Contrarily to the present article, their approach requires that the econometrician possesses two exogenous continuous regressors affecting both the duration and the outcome.

Similarly to ours, other works introduce an instrumental variable to solve the endogeneity issue. chernozhukov2015quantile consider a semiparametric framework under which, in contrast with the situation of this paper, the treatment is continuous and the censoring duration is observed. They propose an estimator and give conditions under which this estimator consistently estimates the quantile regression function of the whole population. Research in the field of biostatistics (tchetgen2015instrumental, li2015instrumental and chan2016reader) has studied an additive hazard model with instrumental variables. Such an additivity assumption is not needed in our framework.

Some papers focus on the case where both the treatment and the instrumental variable are binary. frandsen2015treatment, sant2016program, blanco2019bounds make a monotonicity assumption stating that there are no defiers (see angrist1996identification). Unlike ours, these articles are interested in local average treatment effects on the population of compliers which they are able to estimate thanks to this monotonicity assumption (see wuthrich2020comparison). Instead, by making a rank invariance assumption, we are able to estimate directly the quantile regression function of the whole population. Another difference between our paper and frandsen2015treatment and sant2016program is that the latter papers estimate the counterfactual cumulative distribution functions before inverting them to obtain estimates of the quantile treatment effects. This inversion process requires regularity assumptions that we do not make. Unlike in our paper, in frandsen2015treatment the censoring duration is observed. blanco2019bounds consider the same binary setting but allows for selection and endogenous censoring. Using principal stratification, it develops bounds on the quantile treatment effects. Finally, in the binary treatment and binary instrumental variable setting, BR study a semiparametric and separable model where there is full compliance in the control group.

This paper is organized as follows. The model specification is studied in Section 2. Identification results are derived in Section 3. Section 4 is devoted to the estimation theory. Section 5 describes our simulations and the empirical application. All the technical details are deferred to Appendices A, B and C. We draw conclusions in Section 6.

The model

Let us consider the following model:

eqnarray[eqnarray omitted — 60 chars of source]

where $T \in {\mathbb{R}}_+$, $Z$ is a categorical treatment with support $\{z_1,\dots, z_L\}$, $U$ has a unit exponential distribution and $\varphi$ belongs to $L_{\uparrow}^2(Z,U)$, the set of mappings from $\{z_1,\dots,z_L\}\times {\mathbb{R}}_+$ to ${\mathbb{R}}_+$ that are square integrable with respect to the distribution of $(Z,U)$, and strictly increasing and differentiable in their second argument. The variable $W$ is a categorical instrumental variable with support $\{w_1,\dots, w_K\}$, which is independent of $U$. We assume that the distribution ot $U$ given $(Z,W)$ is continuous. Note that our model is identical to the usual nonseparable IV model in econometrics, except that $U$ follows an exponential distribution instead of a uniform distribution, which is more natural in the context of duration models (the cumulative hazard of the duration follows such a distribution). Such a normalization of the distribution of $U$ is necessary to obtain identification. Note that knowledge of $\varphi$ is equivalent to knowledge of the counterfactual quantile function of $T$ and the two differ solely because the distribution of $U$ is not uniform on $[0,1]$. In particular, for $z\in\{z_1,\dots, z_L\}$ and $u\in{\mathbb{R}}_+$, $\varphi_z(u)$ is the $1-e^{-u}$-quantile of the potential outcome of the duration for treatment level $z$, that is of $\varphi_z(U)$. The variable $U$ represents both ex-ante unobserved heterogeneity, that is differences among individuals which are prior to the duration spell, and ex-post shocks occuring through the duration spell. The duration is right censored by a random variable $C$ with support in ${\mathbb{R}}_+$, and we observe $Y = \min(T,C)$, $Z$ , $W$ and $\delta = I(T \le C)$. We suppose that $C$ is independent of $T$ given $(Z,W)$.

{\color{black} The main object of interest is the regression function $\varphi$. Many quantities of interest can be derived from $\varphi$. First, the quantile treatment effect (QTE) of a change in treatment from $z^0$ to $z^1$ for an individual with $U=u\in \mathbb{R}^+$, $ \varphi(z^1,u)-\varphi(z^0,u).$ Note that this is the QTE at the $1-e^{-u}$-quantile. Then, there is the average treatment effect of a change in treatment from $z^0$ to $z^1$, $ \mathbb{E}[\varphi(z^1,U)-\varphi(z^0,U)].$ Consider also survival probabilities $\mathbb{P}(\varphi(z,U)\ge t)$ and hazard rates $\frac{\partial\mathbb{P}(\varphi(z,U)\le t)}{\partial t}/\mathbb{P}(\varphi(z,U)\ge t)$ at $t\in\mathbb{R}_+$ for $Z=z$.} Let us provide two different characterizations of $\varphi$ in order to further illustrate its relevance: one in terms of the conditional survival function of $T$ and one in terms of the conditional hazard rate of $T$. First, thanks to the exclusion restriction $U\protect\mathpalette{\protect\independenT}{\perp} W$ and to the fact that $U$ has an exponential distribution, $(\varphi_{z_\ell}(u))_{\ell=1}^L$ is a solution to the following system of equations in $\theta=(\theta_\ell)_{\ell=1}^L \in \mathbb{R}^L$:

equation[equation omitted — 124 chars of source]

where $S(t,z|w) = \mathbb{P}(T \ge t,Z=z|W=w)$. Indeed,

eqnarray*[eqnarray* omitted — 385 chars of source]

From now on, we assume that $S$ is differentiable in its first argument. Concerning the second characterization, in the next lemma, we re-express our model in terms of the conditional hazard function. The latter is more common in the context of duration models.

LemmaSuppose $T=\varphi(Z,U)$, $U$ and $W$ are independent, $U \sim \mbox{Exp}(1)$, and the density $f(\cdot|z,w)$ of T given $Z=z,W=w$ exists. Then, $$ \sum_{\ell=1}^L \int_0^{\varphi_{z_\ell}(u)} h(s|z_\ell,w) p(z_\ell|T\ge s,w) \, ds = u, $$ where $h(t|z,w) = f(t|z,w) / S(t|z,w)$ is the hazard function of $T$ given $Z=z,W=w$ (with $S$ being the corresponding survival function), and $p(z|T\ge t,w) = \mathbb{P}(Z=z|T \ge t,W=w)$.

The proof is given in Appendix (ref). To derive the identification results, we use the characterization (ref).

Identification

Exact versus partial identification

In this section, for the sake of simplicity, we assume that the (possibly infinite) upper bound of the support of the distribution of $C$ given $Z=z,W=w$ does not depend on $z\in\{z_1,\dots, z_L\}$ and $w\in\{w_1,\dots, w_K\}$ and we denote this upper bound by $c_0$. Such a case arises for instance when $C$ and $(T,Z,W)$ are independent. This happens, in our empirical application where all durations greater than $26$ are censored and the others are not. Because of censoring, $S$ is only identified on $[0,c_0]\times \{z_1,\dots, z_L\}\times \{w_1,\dots, w_K\}$. We introduce the following lemma that characterizes $\varphi$ in terms of observables.

LemmaFor $k=1,\dots,K$, let $R_{k,u}(\theta)=\sum_{\ell=1}^L S(\theta_\ell\wedge c_0,z_\ell|w_k)- e^{-u}$. The following hold: \begin{itemize} • If $(\varphi(z_\ell,u))_{\ell=1}^L\in [0,c_0)^{L}$, then $$ (\varphi(z_\ell,u))_{\ell=1}^L\in \Big\{\theta \in [0,c_0)^L \Big| R_{k,u}(\theta) = 0 \mbox { for all } k=1,\ldots,K \Big\};$$ • If $(\varphi(z_\ell,u))_{\ell=1}^L\notin [0,c_0)^{L}$ , then $$ (\varphi(z_\ell,u))_{\ell=1}^L\in\Big\{\theta \in \mathbb{R}_+^L \Big| \max_{\ell=1}^L \theta_\ell \ge c_0, \min_{k=1}^K R_{k,u}(\theta) \ge 0 \Big\}.$$ \end{itemize}

{\bf Proof.} Part (i) was proved in the previous section, thus we only prove (ii). We know that $\sum_{\ell=1}^L S(\varphi(z_\ell, u),z_\ell|w_k)= e^{-u}$. Note that $S(\cdot,z|w)$ is decreasing for any $z=z_\ell,\dots,z_L$ and $w=w_1,\dots, w_K$. Hence, $$\sum_{\ell=1}^L S(\varphi(z_\ell, u)\wedge c_0,z_\ell|w_k)\ge\sum_{\ell=1}^L S(\varphi(z_\ell, u),z_\ell|w_k)= e^{-u}.$$ $\Box$\\ We make use of Lemma (ref) to derive identification results in the two different cases mentioned in the lemma. Let $u_0 =\operatorname*{arg\,min}\limits_{\ell\in\{1,\dots, L\}}\varphi_{z_{\ell}}^{-1}(c_0)$ (if the support of $T$ is included in $[0,c_0)$, we set $u_0=\infty$). In Section (ref), we discuss exact identification of $\varphi$ on $[0,u_0)$. On the interval $[u_0,\infty)$ (the empty set if $u_0=\infty$), $\varphi$ is only partially identified. We show how to use Lemma (ref) (ii) to obtain an outer set to the identified set of $(\varphi_{z_{\ell}}(u))_{\ell=1}^L$ in Section (ref). We also discuss why the set in Lemma (ref) (ii) is not the identified set of $(\varphi_{z_{\ell}}(u))_{\ell=1}^L$ in general. Finally, Section (ref) obtains a smaller outer set of $(\varphi_{z_{\ell}}(u))_{\ell=1}^L$ than the one derived from Lemma (ref) (ii) in a special case which nests our empirical application.

Exact identification

In this subsection, we discuss identification of $\varphi$ on $[0,u_0)$. For notational purposes, let us define $\mathcal{F}_{\downarrow}^{Z,W}$, the set of mappings from $\mathbb{R}_+\times \{z_1,\dots,z_L\} \times \{w_1,\dots,w_K\}$ to ${\mathbb{R}}_+$ which are continuous and decreasing in their first argument. We know that $\varphi$ belongs to the set of solutions of the equations

equation[equation omitted — 43 chars of source]

where $A$ is the operator from $L_{\uparrow}^2(Z,U)\times \mathcal{F}^{Z,W}_\downarrow $ to the set of mappings from $[0,u_0)$ to $\mathbb{R}^K$ such that, for $\widetilde{\varphi} \in L_{\uparrow}^2(Z,U)$, $\widetilde{S}\in \mathcal{F}^{Z,W}_\downarrow $ and $u\in [0,u_0)$, $$A(\widetilde{\varphi},\widetilde{S})(u)= \Big(\sum_{\ell=1}^L \widetilde{S}(\widetilde{\varphi}_{z_\ell}(u),z_\ell|w_k)-e^{-u}\Big)_{k=1}^K.$$ If $\widetilde{S}\in \mathcal{F}^{Z,W}_\downarrow $ is differentiable in its first argument, then, for all $\widetilde{\varphi} \in L_{\uparrow}^2(Z,U)$, we define $\Gamma(\widetilde{\varphi},\widetilde{S})$, the Fr\'echet derivative of $A$ in its first argument at the point $(\widetilde{\varphi},\widetilde{S})$.

Local identification

Under weak assumptions, it is possible to show that the system (ref) has a unique solution in a neighborhood of $\varphi$. This is a local identification result. The Fr\'echet derivative of $A$ in its first argument at the point $(\varphi,S)$ is given by $$ \Big(\Gamma(\varphi,S) (u)\Big)_{k\ell}= - f \big(\varphi_{z_\ell}(u),z_\ell | w_k\big). $$ Also, let $g(z\vert u,w)=\mathbb{P}(Z=z\vert U=u,W=w)$ and let $G(u)$ be the $K\times L$ matrix such that $G_{ k \ell}(u)=g(z_\ell | u,w_k)$. We make the following assumption:

itemize• For all $u\in \mathbb{R}_+$, $\mathrm{rank}(G(u))\ge L$.

Note that this assumption is equivalent to the conditional completeness condition: $$ \mathbb{E}(g(Z,U)|W=w,U=u)=0 \, \mbox{ for all } w,u \, \Longrightarrow \, g=0 \mbox{ for all } g:\{z_1,\dots,z_L\}\times {\mathbb{R}}_+\mapsto {\mathbb{R}}. $$ The latter assumption (which implies that $K \ge L$), allows us to show the local identification of our model.

TheoremUnder assumption (L) and assuming that $(\varphi_{z_\ell}(u))_{\ell=1}^L\in [0,c_0)^{L}$, the model is locally identified, in the sense that if $ \Gamma(\varphi,S)(\tilde{\varphi}-\varphi) \equiv 0$ for $\tilde{\varphi} \in L_{\uparrow}^2(Z,U)$, then $\tilde\varphi \equiv \varphi$.

{\bf Proof.} Write

eqnarray[eqnarray omitted — 251 chars of source]

provided that $\varphi_{z_\ell}'(u) \neq 0$, where $g(u,z_\ell|w_k) = \varphi_{z_\ell}'(u) f \big(\varphi_{z_\ell}(u),z_\ell | w_k\big)$. As $g(u,z_\ell|w_k) = e^{-u} g(z_\ell|u,w_k)$, we have that (ref) equals zero if and only if $$ \sum_{\ell=1}^L \frac{\tilde\varphi_{z_\ell}(u)- \varphi_{z_\ell}(u)}{\varphi_{z_\ell}'(u)} g(z_\ell|u,w_k) = 0 \Leftrightarrow G(u) \Big(\frac{\tilde\varphi_{z_\ell}(u)- \varphi_{z_\ell}(u)}{\varphi_{z_\ell}'(u)}\Big)^L_{\ell=1}=0,$$ and hence it follows from assumption (L) that $\tilde\varphi \equiv \varphi$. $\Box$\\

It is useful to provide some intuition on assumption (L) in the case of our empirical application. If $Z$ is a binary treatment indicator and $W$ a binary instrument, the matrix $G(u)$ corresponds to $$G(u)=\left(

array[array omitted — 143 chars of source]

\right).$$ Therefore, assumption (L) is satisfied if and only if $$\mathbb{P}(Z=0\vert U=u,W=0)\ne \mathbb{P}(Z=0\vert U=u,W=1)$$ for all $u\in \mathbb{R}_+$, that is the instrument changes the treatment probability for any value of {\color{black} $U$}. This is similar to Example 1 in D.

Global identification

We continue with global identification. The aim is to find conditions under which the system (ref) has a unique solution. Our presentation is borrowed from Appendix A of FFV. Consider the density of $(U,Z)$ conditional on $W$. This density is perturbed in the direction of a function $\tilde\varphi$ and the amount of perturbation is characterized by the parameter $\mu>0$ : $$ g_{\mu,\tilde\varphi}(u,z|w) = \Big[\varphi_z'(u) + \mu \tilde\varphi_z'(u)\Big] f\big(\varphi(z,u)+\mu \tilde\varphi(z,u),z|w\big). $$ Here, $f(t,z|w) = -\frac{d}{dt} S(t,z|w)$. We make the following hypothesis:

itemize• If $\int_0^1 E_{g_{\mu,\tilde\varphi}}\big(\rho_\mu(Z,U) | U=u,W=w\big) d\mu=0$ for all $u,w$, then $\rho_\mu \equiv 0$, for any function $\rho_\mu:\{z_1,\dots, z_L\}\times{\mathbb{R}}_+\mapsto {\mathbb{R}} $.

Assumption (G) amounts to strong conditional completeness of $Z$ given $U$ and $W$ as introduced in CH and studied in FFV. Under this hypothesis we can now show the global identification of our model.

TheoremUnder assumption (G), $\varphi$ is globally identified on $[0,u_0)$, in the sense that if $ A(\tilde\varphi,S) \equiv A(\varphi,S)$ for $\tilde{\varphi} \in L_{\uparrow}^2(Z,U)$, then $\tilde\varphi \equiv \varphi$.

We refer to Appendix A of FFV for the proof of this result in the continuous case. The adaptation to the discrete case is immediate.

Partial identification

In this section, we consider pointwise identification at $u\in[u_0,\infty)$. In this case, by Lemma (ref), we know that

equation[equation omitted — 185 chars of source]

This is an outer set to the identified set and Lemma (ref) (ii) therefore constitutes a partial identification result. This outer set is not sharp. Section (ref) explains how to obtain a smaller outer set in the degenerate special case where $L=K=2$ and $S(t,1|0)=0$ for all $t\in{\mathbb{R}}_+$. The fact that (ref) is not sharp is not due to this specific situation where $S(\cdot|z,w)$ is zero for a given couple $(z,w)\in \{z_1,\dots,z_L\} \times \{w_1,\dots,w_K\}$. Indeed, in the general setting where $S(\cdot,z|w)$ is nonzero for all values of $z$ and $w$, we give in the supplementary material a counterexample where (ref) is not the identified set. This counterexample suggests that one reason why (ref) is not a sharp set is that it does not take into account the constraints of monotonicity and differentiability of $\varphi$. Let us now discuss how to compute this outer set as a finite union of product of intervals.

We consider first the case where $L=2$. In that case, the set can take the following four types of shapes depending on the shape of the set $\{\theta \in \mathbb{R}_+^L | \min\limits_{k=1,\dots, K} \sum_{\ell=1}^LS(\theta_\ell, z_\ell|w_k)= e^{-u} \}$:

itemize\begin{tabular}{ll} - & the whole positive quadrant $\mathbb{R}_+^2$ with the exception of $[0,c_0)\times[0,c_0)$ \\ - & $[0,\bar\theta_1] \times [c_0,\infty) \cup [c_0,\infty) \times [0,\bar\theta_2]$ \\ - & $[c_0,\infty) \times [0,\bar\theta_2]$ \\ - & $[0,\bar\theta_1] \times [c_0,\infty)$ \\[.2cm] \end{tabular}

for some $0 \le \bar\theta_1,\bar\theta_2 \le c_0$. See Figure (ref) for a graphical description.

figure[figure omitted — 2,163 chars of source]

Let us explain how to obtain Figure (ref). Take $\theta \in{\mathbb{R}}_+^2$ outside $[0,c_0)^2$. Because $S(\cdot \wedge c_0,z|w)$ is flat on $[c_0,\infty)$, $\theta$ belongs to the outer set if and only if $(\theta_1+ (c_0-\theta_1)I(\theta_1>c_0), \theta_2+ (c_0-\theta_2)I(\theta_2>c_0))$ is in the outer set. Therefore, we need to inspect the value of $\theta \rightarrow \min_{k=1}^K R_{k,u}(\theta)$ on $[0,c_0) \times \{c_0\}\cup\{c_0\}\times [0,c_0)$ to compute the outer set. Let us start with $[0,c_0)\times \{c_0\}$. Because $S(\cdot \wedge c_0,z|w)$ is decreasing and continuous, the set of pairs $(\theta_1,c_0)$ in $[0,c_0)\times \{c_0\}$ for which $\min_{k=1}^K R_{k,u}(\theta)\ge 0$ either form a line segment $[0,\bar{\theta}_1]\times{c_0}$ where $\bar{\theta}_1\in[0,c_0]$ or the empty set. First, we consider the case where the latter set is not empty. Because $S(\cdot \wedge c_0,z|w)$ is flat on $[c_0,\infty)$, $[0,\bar{\theta}_1]\times[c_0,\infty)$ belongs to the outer set. If instead the set is empty, $[0,c_0)\times[c_0,\infty)$ does not belong to the outer set. Then, apply the same procedure on $ \{c_0\}\times [0,c_0)$. Finally, if $\min_{k=1}^K R_{k,u}((c_0,c_0)^\top)\ge 0$, $[c_0,\infty)\times [c_0,\infty)$ also belongs to the outer set.

Next, we proceed by a recursion argument and explain how the identified set can be found in dimension $L$ if it is known how to find it in dimension $L-1$. Note that in $L$ dimensions, we can start by fixing the first coordinate $\theta_1$ to $c_0$ and specify the identified set for the remaining $L-1$ coordinates. If the so-selected set is not empty, it should be extrapolated by letting $\theta_1$ belong to $[c_0,\infty)$. This procedure should be repeated $L$ times by fixing each time one coordinate, and at the end the union of all obtained sets should be selected.

A special case

We conclude our discussion of identification with the special case where $L=K=2$ and ${\mathbb{P}}(Z=z_2|W=w_1)=0$. We have a triangular system of equations of the form

eqnarray[eqnarray omitted — 161 chars of source]

This particular setting corresponds to our empirical application. In this case, if $(\varphi_{z_\ell}(u))_{\ell=1}^L\notin [0,c_0)^{L}$ and the system (ref) has a unique solution (which for such a triangular system happens when $S$ is strictly decreasing and continuous), it is possible to obtain a smaller outer set to the identified set than $$ \Big\{\theta \in \mathbb{R}_+^L \Big| \max_{\ell=1}^L \theta_\ell \ge c_0, \min_{k=1}^K R_{k,u}(\theta) \ge 0 \Big\}.$$

We need to consider two cases. First, if $\varphi_{z_1}(u) < c_0$, then $\theta_1=\varphi_{z_1}(u)$ is the (unique) solution of the first equation of ((ref)). This value can then be inserted in the second equation, which will only contain $\theta_2$. If this equation has a solution, it should be $\varphi_{z_2}(u)$ and we are done. If it does not have a solution, the identified set contains all pairs $(\varphi_{z_1}(u),\theta_2)$ with $\theta_2 \ge c_0$.

In the second case $\varphi_{z_1}(u) \ge c_0$. Then, the first equation does not have a solution, and we need to solve the system of inequalities $$ \left\{

array[array omitted — 104 chars of source]

\right. $$ which leads to the identified set $[c_0,\infty) \times [0,\bar\theta_2]$, where $\bar\theta_2$ is the largest value of $\theta_2$ for which $S(\theta_2 \wedge c_0,z_2|w_2) \ge e^{-u}-S(c_0,z_1|w_2)$. See Figure~\ref{fig:red2} for a graphical illustration of the two aforementioned cases. Note that $\theta_2$ can be equal to infinity which implies that the height of the rectangle in the second part of Figure~\ref{fig:red2} can be greater than $c_0$.

figure[figure omitted — 907 chars of source]

Other special cases can be considered depending on which terms in the system of equations ((ref)) are absent. However, they require a case to case analysis, which we will not further develop here.

Estimation

Estimation procedure

We consider estimation with an i.i.d.\ sample of size $n$, $\{Y_i,Z_i,W_i,\delta_i\}_{i=1}^n$. We assume that we have an estimator $\widehat{S}$ of $S$ that satisfies properties to be specified later. Choices of $\widehat{S}$ are discussed in Section (ref). Let $\bar{U}<\infty$ be an upper bound up to which we wish to estimate $\varphi$. Let also $\bar{T}<\infty$ be an upper bound on $\max\limits_{\ell=1,\dots,L}\varphi(z_{\ell},\bar{U})$. Let $\mathcal{K}$ be the set of mappings from $[0,\bar{U}]$ to $\mathbb{R}^K$ and $\mathcal{F}_Z^{\bar{U},\bar{T}}= \{f: \{z_1,\dots,z_L\}\times [0,\bar{U}]\mapsto [0,\bar{T}]\}$. We will make assumptions guaranteeing that $\widehat{S}$ consistently estimates $S$ on $[0,\bar{T}]$ and that the system (ref) has a unique solution in $\mathcal{F}_Z^{\bar{U},\bar{T}}$. This implies that we are implicitly in the case where $(\varphi_{z_\ell}(u))_{\ell=1}^L$ is globally identified.

We introduce further notations. For $u\in [0,\bar{U}]$, let $V(u)$ be a positive definite $K\times K$ weighting matrix. For a $K\times K$ matrix $\bar V$ and a vector $v\in \mathbb{R}^K$, we define $\|v\|=\sqrt{v^{\top}v}$ and $\|v\|_{\bar V}=\sqrt{v^{\top}\bar Vv}$. Then, for $g\in \mathcal{K}$, we introduce $\|g\|^2=\int_{0}^{\bar{U}}\|g(u)\|^2du$, $\|g\|_{V}^2=\int_{0}^{\bar{U}}\|g(u)\|_{V(u)}^2du$. For $f\in \mathcal{F}_Z^{\bar U, \bar T}$, we use $\|f\|^2=\sup_{u\in[0,\bar{U}]}\|(f(z_{\ell},u))_{\ell=1}^L\|^2$. When we write that a random process $R$ is $O_P(a_n)$ or $o_P(a_n)$, we mean that $\|R\|=O_P(a_n)$ or $\|R\|=o_P(a_n)$, respectively. Let us now define the estimator

equation[equation omitted — 161 chars of source]

Consistency

We introduce the following assumptions:

itemize• For $u\in[0,\bar{U}]$, the eigenvalues of $V(u)$ are bounded from above and from below uniformly in $u$. The bounds are strictly positive constants. • \begin{itemize} • For all $\epsilon>0$, there exists $\nu>0$ such that $$\inf_{\theta\in\mathcal{F}_{Z}^{\bar{U},\bar{T}} : \|\theta-\varphi\|\ge \nu} \big\|A(\theta,S)\big\|_V\ge \epsilon ;$$ • There exist $\nu,c>0$ such that for any $\theta\in \mathcal{F}_Z^{\bar{U},\bar{T}}$ satisfying $\|\theta-\varphi \| \le \nu$, $$ \| \theta-\varphi \| \le c \| A(\theta,S) \|;$$ • There exists $(r_n)_n$, a real-valued sequence such that $r_n\to 0$ and $$\sup_{t\in[0,\bar{U}],\ z=z_1,\dots,z_L,\ w=w_1,\dots,w_K}\big|\widehat{S}(t,z|w)-S(t,z|w)\big|=O_P(r_n).$$ \end{itemize}

Assumptions C(i) and C(ii) are conditions related to the shape of the objective function $A$ and they ensure that there is a unique solution to the system of equations (ref) in $[0,\bar{U}]$ and, hence, that there is a unique minimum to the program (ref) if $\widehat{S}$ estimates $S$ well enough. One can develop primitive conditions for C(ii) involving the assumption that the two first derivatives of $S$ with respect to its first argument are uniformly bounded on $[0,\bar{T}]$ (see Lemma (ref) in Appendix (ref)). Assumption (C)(iii) implies that $\bar T$ is chosen lower than the minimum of the upper bound of the support of $C$ over all $z\in\{z_1,\dots, z_L\}$ and $w\in\{w_1,\dots, w_K\}$ such that $\mathbb{P}(Z=z|W=w)>0$. We prove the following theorem. The proof is in Appendix A.

TheoremUnder Assumptions (V) and (C), we have $\| \widehat{\varphi} -\varphi \|=O_P(r_n)$.

This theorem shows that the rate of convergence of $\widehat{\varphi}$ is the same as the one of $\widehat{S}$ in sup-norm. Hence, solving program (ref) does not deteriorate the convergence rate.

Asymptotic normality

We make the following assumption:

itemize\begin{itemize} • There exists a variance operator $\Omega$ such that $n^{1/2} A(\varphi,\widehat{S})$ converges weakly to a mean zero Gaussian process with variance operator $\Omega$; • $\widehat{S}$ is differentiable in its first argument for $u\in[0,\bar{U}]$, the mapping $\Gamma(\varphi,S)^\top V\Gamma(\varphi,S)$ is invertible on $[0,\bar{U}]$, $ \Gamma(\varphi,\widehat{S})\xrightarrow{P}\Sigma=\Gamma(\varphi,S)$ and $ \Gamma(\widehat{\varphi},\widehat{S})\xrightarrow{P}\Sigma$; • $A(\widehat{\varphi},\widehat{S})-A(\varphi,\widehat{S})-\Gamma(\varphi,\widehat{S})(\widehat{\varphi}-\varphi)=O_P(\| \widehat{\varphi}-\varphi \|^2);$$\underset{t\in[0,\bar{U}],\ z=z_1,\dots,z_L,\ w=w_1,\dots,w_K}{\sup}|\widehat{S}(t,z|w)-S(t,z|w)|=o_P(n^{-1/4})$. \end{itemize}

This assumption leads to the following theorem, of which the proof can be found in Appendix A.

TheoremUnder Assumptions (V), (C) and (N), $\sqrt{n}(\widehat{\varphi} -\varphi )$ converges to a mean zero Gaussian process with variance operator $[\Sigma^\top V\Sigma]^{-1} \Sigma^\top V \Omega V\Sigma[\Sigma^\top V\Sigma]^{-1} .$

Inference with bootstrap and optimality

We propose a bootstrap procedure for inference on the functional $f((\varphi_{z_\ell}(u))_{\ell=1}^L)$ of $(\varphi_{z_\ell}(u))_{\ell=1}^L$ for $u\in{\mathbb{R}}_+$ and $f:{\mathbb{R}}_+^L\mapsto {\mathbb{R}}$ differentiable. Examples of such functionals are mentioned in Section (ref).

itemize• Draw with replacement $B$ resamples of size $n$ from the original data $\{Y_i,Z_i,W_i,\delta_i\}_{i=1}^n$; • For each resample $b$, compute $\widehat{\varphi}_b$, the value of the estimator $\widehat{\varphi}$ in the resample; • Construct a $95\%$ confidence interval for $f((\varphi_{z_\ell}(u))_{\ell=1}^L)$ using for the lower bound the $2.5\%$ percentile of $\{f((\widehat{\varphi}_b(z_\ell,u))_{\ell=1}^L)\}_{b=1}^B$ and for the upper bound the $97.5\%$ percentile of $\{f((\widehat{\varphi}_b(z_\ell,u))_{\ell=1}^L)\}_{b=1}^B$.

{\color{black} Lemma (ref) in Appendix B proves that such a procedure works when $S$ is estimated with kernel smoothing as described in Section 4.5. The outlined procedure shows how to build pointwise confidence intervals. However, as demonstrated in Lemma (ref), a similar bootstrap approach also allows to obtain uniform confidence bands on $[0,\bar{U}]$.}

The optimal value of $V$ is $\Sigma^{-1}$ which can be estimated using $\widehat{\Sigma}=\Gamma(\widehat{\varphi},\widehat{S})$ where $\widehat{\Sigma}$ is obtained from a first-step estimator which sets $V(u)$ equal to the identity matrix for any $u\in [0,\bar{U}]$.

Choices of $\widehat{\boldsymbol{S}}$

In this subsection, we discuss choices for $\widehat{S}$ which are robust to random right censoring. Remark that $S(t,z\vert w)=S(t\vert z,w)p_{z,w}$ where $S(t\vert z,w)=\mathbb{P}(T\ge t\vert Z=z, W=w) $ and $p_{z,w}=\mathbb{P}(Z=z\vert W=w )$. We define the following stochastic processes:

align*[align* omitted — 203 chars of source]

We estimate $S(\cdot\vert z,w)$ using the Kaplan-Meier estimator

equation[equation omitted — 114 chars of source]

To provide an estimator $\widehat{S}$ satisfying assumption (N), one needs to smooth $\widehat{S}_{KM}$. Various techniques are available in the literature, including local polynomials and kernel smoothing. For instance, concerning the latter, if we use a kernel $K$ with a bandwidth $\epsilon$, we obtain

equation[equation omitted — 116 chars of source]

Our final estimator of $S$ is

equation[equation omitted — 101 chars of source]

where $\widehat{p}_{zw} =Y_{z,w}/Y_w$. In Appendix B, we state assumptions under which this choice of $\widehat{S}$ satisfies the conditions of Sections 4.2 and 4.3, and for which the bootstrap procedure of Section 4.4 works.

Practical implementation

Let us now discuss how to use the outlined estimation procedure in practice. Fist, one needs to choose $\bar T$. The latter has to be chosen such that Assumption (C)(iii) holds. In appendix B, the condition on $\bar T$ guaranteeing uniform convergence of $\widehat S$ when it is smoothed by kernel is that there exists $\xi>0$ for which $S(\bar T,z|w)/\mathbb{P}(Z=z|W=w)>\xi$ for all $z\in\{z_1,\dots, z_L\}$ and $w\in\{w_1,\dots, w_K\}$ such that $\mathbb{P}(Z=z|W=w)>0$ (condition (K)(v)). As in Section (ref), let us consider the case where the upper bound $c_0$ of the support of the distribution of $C$ given $Z=z,W=w$ does not depend on $z\in\{z_1,\dots, z_L\}$ and $w\in\{w_1,\dots, w_K\}$ and is finite. Then, if $[0,c_0)$ is strictly included in the support of $T$ (or $Y$) given $Z=z,W=w$ for all $z\in\{z_1,\dots, z_L\}$ and $w\in\{w_1,\dots, w_K\}$, the choice $\bar T=c_0$ ensures that condition (K)(v) holds. If the upper bound of the support of $C$ is $\infty$, then a simple valid choice for $\bar T$ is the minimum of the $\alpha$-quantile of $Y$ given $Z=z$ and $W=w$ over all $z\in\{z_1,\dots, z_L\}$ and $w\in\{w_1,\dots, w_K\}$ such that $\mathbb{P}(Z=z|W=w)>0$, where $\alpha\in (0,1)$. One may use $\alpha=0.95$ by convention.

In practice, the minimization program (ref) is not feasible. Instead we choose $u_1,\dots,u_M$, $M$ values at which we want to estimate $\varphi$ (for instance, $M$ i.i.d.\ replications of $U\sim \mathrm{Exp}(1)$ or a grid). For $m=1,\dots,M$, we estimate $\varphi(z,u_m)$ by $ \widehat{\varphi}(z,u_m)$ where

equation[equation omitted — 193 chars of source]

We obtain an estimate of $\varphi$ at the points $u_1,\dots, u_M$, which enables us to plot $\widehat{\varphi}(z,\cdot)$.

When $\varphi$ is not identified on all its support, the estimator $(\widehat \varphi(z_\ell, u_m))_{\ell=1}^L$ will not be consistent for some values of $u$. Under the conditions of Theorem (ref), by the continuous mapping theorem, $\big\|A(\widehat \varphi, \widehat{S})(u_m) \big\|_{V(u_m)}^2=o_P(1)$. Therefore, large values of $\big\|A(\widehat \varphi, \widehat{S})(u_m) \big\|_{V(u_m)}^2$ indicate that $(\widehat \varphi(z_\ell, u_m))_{\ell=1}^L$ is not identified. As a rule of thumb, we suggest to start using partial identification results from Sections (ref) and (ref), when $\big\|A(\widehat \varphi, \widehat{S})(u_m) \big\|_{V(u_m)}^2$ starts increasing significantly with $u$. This approach will be illustrated in Section (ref).

Numerical illustrations

We present two applications of our approach. One corresponds to a simulated example while the second one is an experiment concerning a bonus for finding a job proposed to job seekers in Illinois.

Simulations

We consider the following model. Let $W$ follow a Bernoulli distribution with parameter 0.7. To model the dependence between $Z$ and $(W,U)$ we set

equation[equation omitted — 71 chars of source]

where $\varepsilon\sim \mathcal{N}(0,1)$ and $U \sim Exp(1)$. This definition yields the following conditional treatment probabilities: $\mathbb{P}(Z=1|W=0)=0$ and $\mathbb{P}(Z=1|W=1)=0.76$ (average over 1,000,000 Monte Carlo replications). Note that under this design, $W=0$ implies that $Z=0$, which is chosen to mimic the setting of our empirical application. Next, let $T$ have an exponential distribution with hazard rate {\color{black} $1/10$} if $Z=0$, and {\color{black} $1/5$} if $Z= 1$, such that for $u\in\mathbb{R}_+$, $\varphi(0,u)=10u$ and $\varphi(1,u)=5u$. For $u\in{\mathbb{R}}_+$, the $1-e^{-u}$-quantile treatment effect is equal to $-5u$. Let $C=\max(15Exp(1), 10)$. Note that this implies that $C$ has a constant hazard rate of $1/15$ for durations lower than $10$. With such a censoring mechanism, around $22\%$ of the observations are censored (average of 1,000,000 replications). Finally, with $Y=\min(T,C)$ and $\delta=I(T \le C)$ we generate an i.i.d. sample $\{Y_i, Z_i,W_i, \delta_i\}_{i=1}^n$ of 10,000 observations having the same distribution as $(Y,Z,W,\delta)$. The sample size is chosen to be close to that of the empirical example. Note that under this data generating process, $S$ is identified on $\mathbb{R}_+$ and therefore we aim at point estimation of $\varphi$ and of the quantile treatment effects as in Section (ref).

$\bar{T}$ is equal to $10$, the upper bound of the support of $C$. We compute the estimators on a grid for $u$ between $0.01$ and $1.2$ with step size $0.01$. An Epanechnikov kernel is used to smooth the Kaplan-Meier estimator of the survival function. The bandwidth is chosen using the usual rule of thumb for normal densities. All results are averages over $1,000$ replications.

Figure (ref) reports the value of $\big\|A(\widehat \varphi, \widehat{S})(u) \big\|_{2}^2$ on these points. $\big\|A(\widehat \varphi, \widehat{S})(u) \big\|_{2}^2$ starts increasing slightly after $u$ reaches $0.9$. According to the rule of thumb of Section (ref), we use the estimation results of Section (ref) for $u$ below $0.9$. $\big\|A(\widehat \varphi, \widehat{S})(u) \big\|_{2}^2$ is large for $u$ close to zero because kernel estimators perform poorly near boundaries. Figure (ref) displays the estimated regression function $\widehat{\varphi}_0(\cdot)$. On the same graph, we also report the naive estimate of $\varphi_0(\cdot)$ obtained by inversion of the Nelson-Aalen estimator of the cumulative hazard of $T$ given $Z=0$. The true regression function is omitted because it is undistinguishable from our estimate on the plot (except for very low values of $u$). We also present the $95\%$ confidence intervals of $\widehat{\varphi}_0(u)$ computed using $200$ bootstrap draws with the approach described in Section 4.4. Figure (ref) contains the coverage of these confidence intervals. Figures (ref) and (ref) report the same information for $\widehat{\varphi}_1(\cdot)$. The estimated quantile treatment effects and corresponding confidence intervals are presented in Figures (ref) and (ref). The quantile level $1-e^{-u}$ ranges from 0 to 0.6.

It is clear that the naive estimator (in dashed blue) is biased. On the contrary, our estimator recovers precisely the shape of the true regression function. The confidence intervals have almost nominal coverage except for very low values of $u$ where $S$ is poorly estimated because of boundary properties of kernel estimators. In the supplementary material, we present the results of the simulations when a local polynomial of degree one is used to smooth the Kaplan-Meier estimator of the survival function. The coverages of the confidence intervals with local polynomial smoothing are much closer to $0.95$ for very low values of $u$. This is because local polynomials do not suffer from the bias of kernel estimators near boundaries.

figure[figure omitted — 189 chars of source]
figure[figure omitted — 476 chars of source]
figure[figure omitted — 476 chars of source]
figure[figure omitted — 482 chars of source]

Empirical application

Our empirical application revisits the so-called Illinois Reemployment Bonus Experiment. We refer to \url{https://www.upjohn.org/data-tools/employment-research-data-center/} \url{illinois-unemployment-incentive-experiments} for more details. Taking place between mid-1984 and mid-1985, this randomized control trial aimed at evaluating the effects of bonuses to employers or job seekers on unemployment duration. Specifically, eligible unemployment insurance beneficiaries were randomly allocated to one of the following groups:

itemize• The Job Search Incentive Experiment group (JSIE). Job seekers of this group were eligible to a cash bonus of 500\$ if they were to find a job of at least 30h/week within 11 weeks from the beginning of their unemployment spell and held that job for 4 months. In order to receive the bonus, job seekers had to read the experiment description and sign an agreement form at the beginning of their unemployment spell; • The Hiring Incentive Experiment group (HIE). Companies hiring job seekers within 11 weeks from the beginning of their unemployment spell for a job of more than 30h/week lasting at least 4 months were eligible to a cash bonus of 500\$. In order for their employer to receive the bonus, job seekers had to read the experiment description and sign an agreement form at the beginning of their unemployment spell; • The control group.

The following specifics are worth noting. First, to be part of one of the three aforementioned groups, claimants needed to be between 20 and 55 years old, and have a valid unemployment insurance (UI) claim. Second, the unemployment duration was recorded as the number of weeks during which participants received unemployment benefits, which implies that the data are discrete. In this particular case, the true unemployment duration is continuous but is only observed by interval. This setting creates another source of partial identification studied for instance in manski2002inference. Because an extension to such a setting is outside of the scope of this paper, we choose to neglect this issue and treat the data as continuous. Third, given that UI is granted for 26 weeks, we can only observe individuals up to the end of their UI claim, that is the data are right censored at 26 weeks. Around 21% of the observations are censored.

To clarify the relevance of our model to the experiment, let us assume that we want to evaluate the causal effect of the cash bonus of the JSIE (or HIE) experiment. Although the allocation to the groups is randomly chosen, the decision to agree to participate in the experiment is endogenous. Let us consider a framework with two treatments as described by the following variable $Z$: $$Z=\left\{

array[array omitted — 256 chars of source]

\right.$$ In order to deal with the selection bias, our strategy consists of using the group assignment as an instrument. Let us define $$W=\left\{

array[array omitted — 196 chars of source]

\right.$$\\ We report in Table \ref{tab2} the sample sizes for each combination of values of the couple $(W,Z)$. The refusal rates are respectively 16\% and 35\% in the JSIE and HIE experiment, which suggests large selection biases. In total, there are $12,101$ observations.

table[table omitted — 303 chars of source]

This dataset has been extensively studied in the literature. Comparing `naively' treatment groups to the control group, WS concluded that the JSIE bonus reduces (statistically significantly) the unemployment duration but found no significant evidence of an effect of the HIE bonus. Using a proportional hazards model with $W$ as explanatory variable, M88 reached the conclusion that the JSIE bonus increases the hazard rate. In BR, the hazard rate follows a mixed proportional hazards model which is semiparametric. The authors develop a two-stage instrumental variable estimator, and find a stronger positive effect of the JSIE experiment than M88.

Estimation under exact identification

We estimate quantile treatment effects of the two different bonuses using the method described in this paper. An Epanechnikov kernel is used to smooth the survival function. The bandwidth is equal to $2$, although the results are not sensitive to the choice of the bandwidth as long as it is sufficiently larger than $1$. $\bar{T}$ is equal to $26$, that is the censoring duration. We compute the estimators of $\varphi$ on a grid for $u$ between $0.01$ and $1$ with step size $0.01$. Figure (ref) reports the value of $\big\|A(\widehat \varphi, \widehat{S})(u) \big\|_{2}^2$ on the grid, when the analysis is restricted to the population for which $Z=0$ or $Z=1$ on the left and $Z=0$ or $Z=2$ on the right. It can be seen that $\big\|A(\widehat \varphi, \widehat{S})(u) \big\|_{2}^2$ starts increasing slightly after the value of $u$ reaches $0.7$. Applying the analysis of Section (ref), we choose to use the estimation results of Section (ref) for $u$ below $0.7$. The corresponding quantiles $1-e^{-u}$ are between 0 and around 0.5. Note that $\big\|A(\widehat \varphi, \widehat{S})(u) \big\|_{2}^2$ is large when $u$ is close to zero, this is because kernel estimators perform poorly near boundaries.

figure[figure omitted — 355 chars of source]

Figure (ref) reports estimated quantile treatment effects on this range. We estimate $95\%$ confidence intervals using $1000$ bootstrap draws. Consistent with previous studies, we conclude that the JSIE bonus has a stronger effect than the HIE bonus and both reduce the unemployment duration. For most values of $u$ in the selected range, the quantile treatment effect of the JSIE bonus is significant while the one of the HIE bonus is not. Estimation results using a local polynomial of degree one are similar and, hence, are not reported.

figure[figure omitted — 298 chars of source]

Partial identification

Let us now apply the partial identification results from Section (ref). This empirical application fits in the case of the triangular system of equations of Section (ref) if we restrict the data to individuals for which $W=0$ or $W=1$. Then, we can estimate the outer set of the identified set for values of $u$ for which the following system does not have an exact solution $(\theta_1,\theta_2)\in[0,c_0)\times [0,c_0)$, where $c_0=26$:

eqnarray[eqnarray omitted — 178 chars of source]

Let us define $u(c_0)=-\log(\widehat{S}(c_0 ,0|0) )$, which is equal to around 0.73. For $u\ge u(c_0)$ it is clear that the system (ref) has no solution in $[0,c_0)\times [0,c_0)$. We are, hence, in the second case described in Section (ref), and we can estimate an outer set to the identified set for $(\varphi_0(u(c_0)), (\varphi_1(u(c_0)))^\top$ by $[c_0,\infty)\times [0, \widehat{\theta}_2]$, where $\widehat{\theta}_2$ is defined by $\widehat{S}(\widehat{\theta}_2,1|1)= e^{-u}- \widehat{S}(c_0,0|1)$. For $u=u(c_0)$, we obtain $\widehat{\theta}_2=23.31$ and the estimated identified set is $[26,\infty)\times [0, 23.31]$. This implies that the estimated outer set of the quantile treatment effect of the JSIE bonus at $u(c_0)$ is $(-\infty,-2.69]$, which suggests that the true quantile treatment effect is lower than $-2.69$.

Conclusion and future research

In this paper, we studied the issue of identification and estimation of an endogenous categorical regressor on a duration, in the presence of a categorical instrumental variable. We developed partial, local and global identification results in the presence of random right censoring in a fully nonparametric framework. The causal effect is expressed as the solution of a nonlinear system of equations. We derived the asymptotic properties of the solution to an estimated version of this system. Our simulations exhibit excellent performance of our approach in finite samples. We showed how to revisit an empirical application using our methodology.

Several extensions of our model may be interesting. One could consider the case where $Z$ and/or $W$ are continuous variables, or even dynamic processes. For instance, $Z$ could represent a treatment that can be obtained throughout the unemployment spell of an individual, e.g. a job training. {\color{black} Specifically, if the variables $Z$ and $W$ are continuous (but not time dependent), equation (ref) is replaced by a nonlinear integral equation

equation[equation omitted — 71 chars of source]

where $\varphi$ is the functional parameter of interest and the function $S$ may be nonparametrically estimable. This equation generates an ill-posed problem and its resolution requires regularization (see e.g. kaltenbacher2008iterative or for econometric applications, C). The most convenient solution is to implement a sequential resolution algorithm stopped in a suitable way. The case with censoring and partial identification raises an original question in this framework. The extension of our approach to dynamic treatments (and possibly dynamic instruments) would be based on a generalisation of Lemma (ref). For example, $Z$ may be replaced by $Z_t = I(t \ge \epsilon)$, where $\epsilon$ is the random starting time of the treatment. Dynamic treatments also generate new censoring mechanisms and will be considered in future research. These problems are discussed in the unpublished working paper by florens2010endogeneity. }

Another valuable generalization would allow for additional covariates $X$. If the latter are discrete, they may be included in the variable $Z$ or the econometrician could use the approach of the present paper conditional on $X$. If they are continuous, the problem is more complicated. If $T=\varphi(Z,X,U)$, analogously to (ref), we would have the system $$ \int S(\varphi(z,x,u),z,x\vert w)dz=e^{-u}$$ or the equivalent with a sum if $Z$ is discrete. One would need to regularize the estimator as discussed in the previous paragraph.

As shown by the empirical application, an interesting extension would concern discrete outcomes. If a discrete outcome arises because $T$ is continuous but is only observed as an interval, then we have a second source of partial identification. Such a censoring mechanism has been studied in the partial identification literature, see for instance manski2002inference. If instead, the true outcome is discrete, then we need to allow $\varphi_Z(\cdot)$ to be a step function. In this case (ref) cannot be obtained because $\varphi_Z(\cdot)$ would not be strictly increasing.

Another possible source of partial identification is nonrandom censoring. Such a case was studied in blanco2019bounds using principal stratification. A similar approach may be tried in the context of this paper.

Finally, inference in the partially identified case due to censoring is an appealing topic. We conjecture that a bootstrap approach may be used. For each bootstrap simulation, the function $\min_k R_{k,u}(\theta)$ may be evaluated and then the corresponding outer set can be computed. Then, for any point outside of $[0,c_0)^L$, one may compute the bootstrap probability to belong to this outer set and confidence regions of the outer set may be derived. Another possibility would be to use results on models based on conditional moment inequalities such as in andrews2013inference.