Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
80,952 characters · 14 sections · 69 citation commands
Partial Identification and Inference in Duration Models with Endogenous Censoring
\setstretch{1.2}
Duration models are widely used in various empirical studies in economics and biomedical sciences, where outcomes of interest are durations up to the occurrence of some event. Durations of interest in economics include unemployment duration, strike duration, insurance claim duration, and the duration until the purchase of a durable good.\footnote{van_den_Berg_2001 surveys many applications of duration models.} In practice, duration data are often censored. For example, unemployment duration is likely to be censored due to some individuals dropping out of the survey, due to attrition say. Dealing with censoring has been a substantial challenge in duration analysis, and various methods have been proposed. A standard approach is to assume that censoring is independent of unobserved heterogeneity (conditional or unconditional on observed characteristics). Studies employing this approach include Cox_1972, Powell_1984, Ying_et_al_1995, Yang_1999, Honore_et_al_2002, Hong_Tamer_2003, and Khan_Tamer_2007, among others. However, in many cases, justifying this independence assumption is difficult. For example, in unemployment duration analysis, unemployed individuals with low motivation to find a job may tend to drop out of the survey at an early stage. Szydlowski_2019 presents a number of examples where censoring is correlated with unobserved heterogeneity (i.e., censoring is endogenous). In this paper, we study identification and inference in transformation models in the presence of endogenous censoring. The transformation model is expressed as
where $\Lambda$ is a non-degenerate monotone function; $Y^{\ast}$ is a dependent variable, which represents a duration outcome in this paper; $X$ is a $k$-dimensional vector of observed covariates, where $k\geq 2$; $\beta_{0}$ denotes a $k$-vector of regression parameters; and $U$ is unobserved heterogeneity that is independent of $X$. Many kinds of duration models, such as the accelerated failure time model, proportional hazard model, and mixed proportional hazard (MPH) model, can be viewed as transformation models.\footnote{Aside from duration models, a class of transformation models contains other important kinds of models, for example, the linear index model and Box-Cox transformation model.} In this paper, we consider a nonparametric transformation model in which neither the transformation function nor the distribution function of the unobserved heterogeneity is parametrically specified. One important model represented by the nonparametric transformation model is the nonparametric MPH model in which neither a baseline hazard function nor the distribution function of the unobserved heterogeneity is parametrically specified. Allowing for endogenous censoring, we develop bounds on the regression parameters $\beta_{0}$ and transformation function $\Lambda$ in model ((ref)). To the best of our knowledge, this is the first work to derive bounds on $\beta_{0}$ and $\Lambda$ in the nonparametric transformation model with endogenous censoring. The construction of the bounds is built on the rank properties of the nonparametric transformation model studied by Han_1987 and Chen_2002. Han_1987 shows that if there is no censoring, at least one element of $X$ has full-support on the real line, and $X$ is full-rank, the regression parameters are point-identified up to scale by looking at the rank correlation between the outcomes and covariates. In the presence of endogenous censoring, we develop bounds on $\beta_0$ by supposing that, in the rank property studied by Han_1987, each censored outcome takes an infinitely large value or is equal to the censoring time. This reflects the fact that concerning each censored outcome, all we know is that it may take any value larger than censoring time. When $\Lambda$ is strictly increasing, a lower bound on $\Lambda$ is also attained by incorporating endogenous censoring into the rank property studied by Chen_2002. Once the bounds on the regression parameters and the transformation function are obtained, we can also derive bounds on the distribution function of the unobserved heterogeneity $U$. The bounds on $\beta_0$ and $\Lambda$ are characterized by conditional moment inequalities whose sample moments are U-statistics. Based on these conditional moment inequalities, we construct inference methods for these parameters by extending the inference approach of Andrews_Shi_2013 for conditional moment inequality models to the case of U-statistics. The proposed inference approach can be applied not only to this work but also to other works involving conditional moment inequalities and U-statistics. In this sense, this paper also contributes to the literature on inference for conditional moment inequality models.\footnote{Various inference methods for conditional moment inequality models have been proposed, for example, by Andrews_Shi_2013, Andrews_Shi_2014, Andrews_Shi_2017, Chernozhukov_et_al_2013, Armstrong_2014, Armstrong_2015, Menzel_2014, Chernozhukov_et_al_2019, and so on. But none of them can be applied to sample moment functions of U-statistics.} The bounded sets of the parameters proposed in this paper are not necessarily sharp identified sets. On the other hand, using concepts from random set theory (e.g., Beresteanu_et_al_2011, Beresteanu_et_al_2012), we also characterize the sharp identified set of the regression parameters. However, constructing a feasible inference method based on it is difficult, whereas the proposed sets are tractable to construct feasible inference methods. In the paper, we also discuss conditions under which the proposed set of $\beta_0$ approaches the sharp identified set.
This paper is mostly related to works that study endogenous censoring. Khan_Tamer_2009, Khan_et_al_2011, Khan_et_al_2016, Li_Oka_2015, and Fan_Liu_2018 study identification and estimation of parameters in quantile regression models with endogenous censoring. For cross-sectional linear quantile regression models, Khan_Tamer_2009 provide a point identification result for the linear coefficients under a certain support condition, while Khan_et_al_2011 provide a partial identification result without this support condition. Under censoring characterized by a certain copula, Fan_Liu_2018 partially identify the linear coefficients of the same model. Li_Oka_2015 and Khan_et_al_2016 consider panel quantile regression models with endogenous censoring and provide partial identification results. In contrast to these works, the identification result in this paper does not rely on quantile modeling, copula characterization of censoring, or panel data. Aside from quantile models, Szydlowski_2019 considers the parametric MPH model and proposes a sharp identified set and inference method for its parameters. While Szydlowski_2019 considers the parametric MPH model, we consider the nonparametric one, which is robust to misspecification of the hazard function or the distribution function of unobserved heterogeneity.
For competing risks models, Honore_Lleras_Muney_2006 partially identify the parameters in the accelerated failure time model, and Kim_2018 derives computationally tractable bounds on distributions of latent durations by exploiting the discreteness of observed durations. In this paper, we allow for continuous observed durations and do not specify competing risks. In a sample selection model, Honore_Hu_2020 obtain the sharp identified set of linear regression coefficients by imposing a particular structure on the sample selection.
The remainder of this paper is structured as follows. Section (ref) describes the setup and assumptions and then provides the main results to develop bounds on the regression parameters and the transformation function. We also characterize the sharp identified set of the regression parameters and compare it to our proposed superset. Section (ref) provides an inference method for the regression parameters and derives its asymptotic properties. A joint inference method for the regression parameters and the transformation function is presented in Appendix (ref). Section (ref) presents numerical examples and Monte Carlo simulation results. The numerical examples show how the bounds on the regression parameters and transformation function vary depending on the degree of censoring and the support of covariates. The Monte Carlo simulation results show the finite sample properties of the proposed inference method for the regression parameters. Section (ref) presents an empirical illustration, where we apply our proposed inference methods to evaluate the effect of heart transplants on patients' survival duration using data from the Stanford Heart Transplant Study. We conclude this paper with some remarks in Section (ref). All proofs are presented in Appendices (ref) and (ref). Some additional numerical examples and Monte Carlo simulation results are presented in the supplementary material to this paper.
We first describe the setting of the paper and provide conditions to develop bounds on the regression parameters in Section (ref). Subsequently, in Section (ref), we present the main result to construct bounds on the regression parameters. In Section (ref), we characterize the sharp identified set of the regression parameters using concepts from random set theory, and compare this with our proposed superset. Section (ref) derives bounds on the transformation function and the distribution of the unobserved heterogeneity.
We consider the transformation model in the form of ((ref)). In the model, we do not specify the transformation function $\Lambda$ or the distribution function of the unobserved heterogeneity, which we denote by $F_{U}$. Because of this, we impose location and scale normalizations. For the location normalization, we suppose that the constant term is equal to zero (i.e., $X$ does not contain a constant term). For the scale normalization, we suppose that the absolute value of the first component of $\beta_{0}$ is equal to one (i.e., $\left|\beta_{0,1}\right|=1$), where $\beta_{0,j}$ denotes the $j$-th component of $\beta_{0}$. Later, in Section (ref), we impose an additional location normalization to fix $\Lambda$. Let $B\equiv \{-1,1\}\times \mathbb{R}^{k-1}$ denote the normalized regression parameter space. Our first main focus is on the identification of and inference on the normalized regression parameters $\beta_{0}$ in $B$. The transformation model contains many kinds of duration models as its special cases: the accelerated failure time model, Cox's proportional hazard model, and the MPH model.\footnote{If $\Lambda(Y^{\ast})=\exp( Y^{\ast})$, the transformation model corresponds to the accelerated failure time model; if $\Lambda(Y^{\ast})=\exp( \Delta(Y^{\ast}))$, where $\Delta(\cdot)$ is the integrated baseline hazard function and $U$ has the CDF $F(u)=1-\exp(-e^{u})$, the transformation model corresponds to Cox's proportional hazard model; if $\Lambda(Y^{\ast})=\exp(\Delta(Y^{\ast}))$ and $U=\epsilon + \nu$ where $\nu$ is unobserved heterogeneity and $\epsilon$ has the CDF $F(\epsilon)=1-\exp(-e^{\epsilon})$, the transformation model corresponds to the MPH model. For more details, see Horowitz_2009 (Horowitz_2009, Ch. 6).} In particular, the nonparametric MPH model is an important duration model represented by a nonparametric transformation model. The MPH model extends Cox's proportional hazard model by incorporating individual unobserved heterogeneity. Since introduced in Lancaster_1979, the MPH model has been widely used in various empirical studies in economics. In the nonparametric MPH model, the normalized regression parameters $\beta_{0}$ can be interpreted as the logs of the scale-normalized hazard ratios (see, e.g., Lancaster_1990). When data are subject to censoring, the duration outcome $Y^{\ast}$ cannot always be observed. Instead, for unit $i=1,\ldots,n$, we observe $W_{i}=\left(Y_{0i},D_{i},X_{i}\right)$ such that $Y_{0i}=\mbox{min}\left\{ Y_{i}^{\ast},C_{i}\right\} $ and $D_{i}=I\left[Y_{i}^{\ast}\leq C_{i}\right]$, where $C_{i}$ is a random censoring variable and $I\left[\cdot\right]$ denotes the indicator function. $D_{i}$ is a censoring indicator that takes the value zero if $Y_{i}^{\ast}$ is censored and the value one if $Y_{i}^{\ast}$ is observed. Note that we consider right censoring in the paper, but all the results presented below are easily extendable to left and interval censoring. Using $D_{i}$, $Y_{0i}$ can be expressed as $Y_{0i}=D_{i}Y_{i}^{\ast}+(1-D_{i})C_{i}$. Let $P$ denote the distribution function of a vector of random variables $\left(X,U,C\right)$ and $\mathcal{X} \subseteq \mathbb{R}^{k}$ denote the support of $X$. Throughout this paper, we suppose that the following assumptions hold.
Note that Assumption (ref) does not restrict the relationship between $U$ and $C$, allowing for endogenous censoring.\footnote{Chiappori_et_al_2015 study identification and estimation of the nonparametric transformation model when some covariates are endogenous and there is no censoring.}
This section constructs bounds on the regression parameters $\beta_{0}$. The construction is based on a rank property of the latent outcome $Y^{*}$ and the regression part $X^\prime \beta_{0}$ in the transformation model ((ref)). For explanatory purposes, we first introduce the point identification result of Han_1987 in the absence of censoring. Han_1987 supposes that there is no censoring (i.e., $Y^{\ast}$ is always observed). In this case, under Assumptions (ref)--(ref), a full-support condition on an element of $X$, a full-rank condition on $X$, and a continuous distribution of $U$, he shows that $\beta_{0}$ uniquely satisfies the following rank property,
for all $(x_{i},x_{j})\in \mathcal{X}^{2}$, where $P(\cdot \mid x_{i},x_{j})$ denotes the conditional probability given $(X_i,X_j)=(x_i,x_j)$.\footnote{Han_1987 actually considers a slightly different rank property. Theorem H.1 in the supplementary material to this paper shows that $\beta_0$ uniquely satisfies ((ref)) for all $(x_{i},x_{j})\in \mathcal{X}^{2}$ under the same assumptions as in his theorem.} This rank property means that, for any given pair of $(x_{i},x_{j})$, the probability that $Y_{i}^{\ast}$ is larger than or equal to $Y_{j}^{\ast}$ is greater than or equal to 1/2 if and only if $x_{i}^{\prime}\beta_{0}$ is larger than or equal to $x_{j}^{\prime}\beta_{0}$. Then $\beta_{0}$ is the unique value in $B$ that satisfies this rank relation for any pair of $(x_{i},x_{j}) \in \mathcal{X}^{2}$. In other words, for any $\beta\neq\beta_{0}$, there exists at least one pair $(x_{i},x_{j}) \in \mathcal{X}^{2}$ that violates the rank relation ((ref)). In this paper, we suppose that censoring exists and it may be endogenous. Hence, we cannot always observe $Y_{i}^{\ast}$ and do not have any information about the censoring mechanism. The censoring variable $C_{i}$ may be arbitrarily correlated with the observed covariates $X_{i}$ and unobserved heterogeneity $U_{i}$. In this situation, we can still construct bounds on the regression parameters $\beta_{0}$. Let $Y_{1i}\equiv D_{i}Y_{i}^{\ast}+(1-D_{i})(+\infty)$, which is an outcome variable that takes an arbitrary large value when the primary outcome is censored. Recall that $Y_{0i}=D_{i}Y_{i}+(1-D_{i})C_{i}$. Then because $P(Y_{1i}\geq Y_{0j}\mid x_{i},x_{j})\geq P(Y_{i}^{\ast}\geq Y_{j}^{\ast}\mid x_{i},x_{j})$ holds for all $(x_{i},x_{j}) \in \mathcal{X}^{2}$, the following rank property holds from ((ref)):
for all $(x_{i},x_{j}) \in \mathcal{X}^{2}$. Therefore, defining
$\beta_{0}$ is contained in $B_{I}$. This set is derived from a worst-case analysis where we suppose that censored outcomes may take extreme values, $C$ or $+\infty$, for any given value of $x$. This reflects the fact that concerning each censored outcome, all we know is that it may take any value at least larger than its censored time.
The following assumption ensures that $B_I$ is a proper subset of $B(=\{-1,1\}\times \mathbb{R}^{k-1})$.
The following theorem summarizes the main result in this section.
A proof of this theorem is given in Appendix (ref). We make several remarks about this theorem. First, in this theorem, we do not impose a full-support condition on an element of $X$, in contrast to many works in the semiparametric literature, as we no longer focus on point identification.\footnote{Magnac_Maurin_2008, Blevins_2011, and Komarova_2013 discuss the difficulties of justifying the full-support condition in a number of cases, and provide partial identification results for different semiparametric models in the absence of the full-support condition.} A full-rank condition on $X$ is necessary (but not sufficient) for $B_I$ to be bounded. Second, as we will see in the following subsection, $B_{I}$ is not a sharp identified set. However, this set is easy to compute and to construct a feasible inference method, as we will see in Section (ref). The following subsection shows how the sharp identified set can be characterized, and its computational difficulty.
In this section, we illustrate a way to characterize the sharp identified set using concepts from random set theory. Subsequently, we compare the sharp identified set with $B_{I}$. This comparison clarifies why $B_{I}$ is not a sharp set and in which situations $B_I$ approaches the sharp set. For the definitions and notations for random set theory used in this section, see, for example, Molchanov_2005 or Beresteanu_et_al_2012 (Beresteanu_et_al_2012, Appendix A). Throughout this section, for any variable $A$, we denote by $\tilde{A}$ an independent copy of $A$. Using concepts from random set theory, we can characterize the incomplete information for the latent outcome variable $Y^{\ast}$. For any random variable $A$, let $A_{x}$ denote a random variable that has the conditional distribution of $A$ given $X=x$. Then, for a given $x\in\mathcal{X}$, what we observe for the latent outcome variable in the presence of endogenous censoring can be expressed as the random set $\mathcal{Y}_{x}$ defined as
where $Y_{x}^{\ast}$ and $C_{x}$ are, respectively, a latent outcome and a censoring variable given $X=x$; $D_{x}=I\left[Y_{x}^{\ast}\leq C_{x}\right]$ is a censoring indicator given $X=x$. Hence, all the information for the latent outcome variable can be expressed by stating that $Y_{x}^{\ast}\in\mbox{Sel}(\mathcal{Y}_{x})$.\footnote{For any random set ${\cal Y}$, a random variable $Y$ is called a measurable selection of ${\cal Y}$ if $Y\in{\cal Y}$ a.s., and $\mbox{Sel}\left({\cal Y}\right)$ is defined to be the set of all measurable selections of ${\cal Y}$. See, for example, Molchanov_2005 (Molchanov_2005, Ch. 1) or Beresteanu_et_al_2012 (Beresteanu_et_al_2012, Appendix A).} Let $B_{0}$ denote the sharp identified set of $\beta_{0}$. Throughout this section, we suppose that the full-support condition on one element of $X$, the full-rank condition on $X$, and the continuity of $F_U$ hold to ensure the sharp identification result.\footnote{More specifically, we suppose that the following three conditions hold: (i) $\mathcal{X}$ is not contained in any proper linear subspace of $\mathbb{R}^k$; (ii) for almost every $x_{-1} = (x_2\ \ldots, x_k)$, the distribution of $X_{(1)}$ conditional on $(X_{(2)},\ldots,X_{(k)})=x_{-1}$ has an everywhere positive density, where $X_{(m)}$ denotes the $m$-th element of $X$; (iii) $F_U$ is a continuous distribution function. }\footnote{These conditions might not be needed to derive the sharp identified set; however, to our knowledge, there is no work that derives the sharp identified set for the nonparametric transformation model in the absence of these conditions.} Combining the above random set representation with Han_1987's (Han_1987) point identification result that $\beta_0$ uniquely satisfies the the rank property ((ref)) for all $(x_i,x_j) \in \mathcal{X}^2$, the sharp identified set $B_{0}$ is characterized as the set of $\beta$ such that there exists a family of pairs of selections $(Y_{x_{i}},\tilde{Y}_{x_{j}})\in\mbox{Sel}(\mathcal{Y}_{x_{i}})\times \mbox{Sel}(\widetilde{\mathcal{Y}}_{x_j})$ over $(x_{i},x_{j})\in {\cal{X}}^{2}$ that satisfy the following:
for all $(x_{i},x_{j})\in{\cal X}^{2}$, where $\widetilde{\mathcal{Y}}_{x}$ is the random set of $\tilde{Y}_{x}$. Therefore, $B_{0}$ is characterized as
We next look at how the proposed set $B_{I}$ can be characterized by the random set. Given $x \in \mathcal{X}$, by definition, $Y_{1x}$ and $Y_{0x}$ satisfy (i) $Y_{1x},Y_{0x}\in\mbox{Sel}(\mathcal{Y}_{x})$ a.s. and (ii) $Y_{0x} \leq Y_{x}\leq Y_{1x}$ a.s. for all $Y_{x}\in \mbox{Sel}(\mathcal{Y}_{x})$. Thus for any given pair $(x_{i},x_{j}) \in \mathcal{X}^2$, the parameter set
is equivalent to
which is the set of $\beta$ such that for the fixed $(x_{i},x_{j})$, there exists a pair of selections $\left(Y_{x_{i}},\tilde{Y}_{x_{j}}\right)\in\mbox{Sel}(\mathcal{Y}_{x_{i}})\times\mbox{Sel}(\widetilde{\mathcal{Y}}_{x_j})$ that satisfies the inequality ((ref)). Therefore, $B_{I}$ is characterized as the set of $\beta$ such that for any pair $(x_{i},x_{j})\in {\cal{X}}^{2}$, there exists a pair of selections $\left(Y_{x_{i}},\tilde{Y}_{x_{j}}\right)\in \mbox{Sel}(\mathcal{Y}_{x_{i}})\times\mbox{Sel}(\widetilde{\mathcal{Y}}_{x_{j}})$ that satisfy the inequality ((ref)). Formally, $B_{I}$ is characterized as
The difference between ((ref)) and ((ref)) (i.e., the different orders of “$\forall(x_{i},x_{j})\in{\cal X}^{2}$” and “$\exists\left(Y_{x_{i}},\tilde{Y}_{x_{j}}\right)\in \mbox{Sel}(\mathcal{Y}_{x_{i}})\times\mbox{Sel}(\widetilde{\mathcal{Y}}_{x_{j}})$” in ((ref)) and ((ref))) shows that $B_{0}$ is contained in $B_{I}$, but $B_0$ does not necessary contain $B_I$. Hence $B_{I}$ is not necessarily a sharp set. Some intuition for the non-sharpness of $B_{I}$ is as follows. Fix a triple $\left(x_{i},x_{j},x_{k}\right) \in \mathcal{X}^3$ and $\beta \in B$ such that $x_{i}^{\prime}\beta \leq x_{j}^{\prime}\beta \leq x_{k}^{\prime}\beta$. When we check whether $\beta \in B_{I}$ through the rank property ((ref)), in comparing $x_{j}$ and $x_{k}$, we suppose that the latent outcome variable $Y_{x_{j}}^{\ast}$ takes its smallest value, $C_{x_{j}}$; whereas, in comparing $x_{j}$ and $x_{i}$, we suppose that $Y_{x_{j}}^{\ast}$ takes its largest value, $+\infty$. However, when we check whether $\beta \in B_0$ through the characterization of ((ref)), we compare fixed selections $Y_{x} \in \mbox{Sel}(\mathcal{Y}_{x})$ over all $x\in \cal{X}$; that is, $Y_{x_{j}}$ does not change when comparing different $\tilde{Y}_{x_{i}}$ over $x_{i}\in \cal{X}$. This difference explains why $B_{I}$ is larger than $B_{0}$.
We next develop a lower bound on the transformation function in the presence of endogenous censoring. Knowing about the transformation function enables us to infer the type of duration model. We here suppose that $\Lambda$ is strictly monotonic.
Let $T\equiv \Lambda^{-1}$. We call $T$ the transformation function when no confusion arises. As an additional location normalization to fix $T$, we suppose $T(\tilde{y})=0$ for some specific outcome value $\tilde{y}<\infty$. We now construct a bound on the normalized $T(y)$ at a particular value of $y\in\mathbb{R}$. Construction of the bound is built on the rank property studied by Chen_2002. In the case of no censoring, provided that Assumptions (ref)--(ref) and $\ref{asm:strict monotonicity}$ hold and that the true regression parameters $\beta_{0}$ are given,\footnote{Under the supposed conditions with the full-support and full-rank conditions on $X$ and the continuity of $F_U$, $\beta_{0}$ can be point identified by, for example, applying Han_1987's Han_1987 maximum rank correlation approach.} Chen_2002 shows that $T\left(y\right)$ satisfies the following rank property:
for all $(x_{i},x_{j})\in{\cal X}^{2}$, where we recall that $\tilde{y}$ is such that $T\left(\tilde{y}\right)=0$ for the location normalization. Moreover, supposing further that the full-support and full-rank conditions on $X$ and the continuity of $F_U$ hold, Chen_2002 shows that $T\left(y\right)$ can be point identified as the minimum value of $t$ that satisfies the inequality ((ref)) with $t$ in place of $T(y)$ for all $(x_i,x_j)\in \mathcal{X}^2$. Chen_2002 also provides an inference method based on this identification result. In the presence of endogenous censoring, we can construct a lower bound on $T\left(y\right)$ using a similar idea to that presented in Section (ref). If $\beta_{0}$ were given, because $P(Y_{1i}\geq y\mid x_{i})\geq P(Y_{i}^{\ast}\geq y\mid x_{i})$ and $P(Y_{j}^{\ast}\geq y\mid x_{j})\geq P(Y_{0j}\geq y\mid x_{j})$ hold for all $(x_{i},x_{j}) \in \mathcal{X}^2$, it follows from ((ref)) that $T\left(y\right)$ is contained in the following set:
In the presence of endogenous censoring, we cannot point identify $\beta_{0}$; instead, we have $B_{I}$ which contains $\beta_{0}$. Thus, letting
and $T_{B_I}\left(y\right) \equiv \left\{ T_{I,\beta}\left(y\right)\mid\beta\in B_I\right\}$, we have that $T\left(y\right)\in T_{B_{I}}\left(y\right)$ a.s. Note that since any $t>T(y)$ satisfies all the inequality conditions in ((ref)), $T_{I,\beta}(y)$, as well as $T_{B_I}$, does not have a finite upper bound. Hence we can only obtain a lower bound on $T(y)$. Note also that, due to a similar reason as that discussed in Section (ref), $T_{I,\beta_{0}}\left(y\right)$ is not a sharp identified set of $T\left(y\right)$ even if $\beta_{0}$ is known. Hence, $\left\{ T_{I,\beta}\left(y\right)\mid\beta\in B_0\right\}$ is not a sharp identified set even if we had $B_{0}$.
The following theorem formalizes the main result in this subsection.
The proof is provided in Appendix (ref). For any fixed $y \in \mathbb{R}$, $T_{B_I}(y)$ has a finite lower bound if there exists at least one pair of $(x_i,x_j) \in \mathcal{X}^2$ such that $P\left(Y_{1i}\geq y\mid x_{i}\right)< P\left(Y_{0j}\geq\tilde{y}\mid x_{i}\right)$ holds and that $x_i^{\prime}\beta - x_j^{\prime}\beta>-\infty$ holds for all $\beta \in B_I$. The lower bound approaches the true value $T(y)$ as the likelihood of censoring diminishes.
This section provides a statistical inference approach for the regression parameters in model ((ref)) based on the result presented in Section (ref). We suggest a method to construct a confidence set that covers the true parameter value $\beta_{0}$ with a probability greater than or equal to $1-\alpha$ for $\alpha\in\left(0,1\right)$. Because $B_{I}$ is characterized by conditional moment inequalities involving U-statistics, we construct the inference method by extending the inference approach for conditional moment inequality models proposed by Andrews_Shi_2013 (hereafter AS) to the U-statistics case. The approach transforms conditional moment inequalities into an infinite number of unconditional ones, without information loss, to construct a test statistic. A confidence set is then constructed by inverting the test statistic and using critical values obtained via moment selection. The inference method proposed below is applicable to either continuous or discrete covariates. The inference method is for U-statistics of order two, but the approach is extendable to U-statistics of greater order with obvious modifications. A joint inference procedure for $\beta_{0}$ and $T(\cdot)$ is presented in Appendix (ref).
We first construct a test statistic and then describe the inference procedure. Let
where $W_{i} \equiv (Y_{1i},Y_{0i},X_{i})$ for $i=1,\ldots,n$. Then $B_{I}$ is a set of parameter values $\beta$ that satisfy the following conditional moment inequalities:
where $E_P[\cdot \mid x_i,x_j]$ denotes the conditional expectation under the distribution $P$ given $(X_i,X_j)=(x_i,x_j)$. To transform all of the conditional moment inequalities ((ref)) into unconditional ones without loss of information, we adopt AS's instrumental functions approach. We here suppose that the first $p$ ($p \leq k$) elements of $X$ are continuous variables or have infinite support and that the remaining elements of $X$ have finite support.\footnote{We can also allow all covariates to have finite or infinite support with modifications.} Let $\mathcal{X}_{1}$ and $\mathcal{X}_{2}$ denote the supports of $(X_{i,1},\ldots,X_{i,p})$ and $(X_{i,p+1},\ldots,X_{i,k})$, respectively, where $X_{i,j}$ denotes the $j$-th element of $X_i$. Without loss of generality, we also suppose that $(X_{i,1},\ldots,X_{i,p})$ is transformed via a one-to-one mapping so that each of its elements lies in $[0,1]$ (i.e., $\mathcal{X}_{1} \subseteq \left[0,1\right]^{p}$). \footnote{Let $X_{i}^{(1)}\equiv(X_{i,1},\ldots,X_{i,p})$. Following AS, the vector of transformed covariates in $X_{i}^{(1)}$ may be $\tilde{X}_{i}^{(1)}\equiv\Phi\left(\hat{\Sigma}_{X^{(1)},n}^{-1/2}\left(X_{i}^{(1)}-\bar{X}_{n}^{(1)}\right)\right)$ where $\bar{X}_{n}^{(1)}\equiv n^{-1}\sum_{i=1}^{n}X_{i}^{(1)}$, $\hat{\Sigma}_{X^{(1)},n}\equiv n^{-1}\sum_{i=1}^{n}\left(X_{i}^{(1)}-\bar{X}_{n}^{(1)}\right)\left(X_{i}^{(1)}-\bar{X}_{n}^{(1)}\right)^{\prime}$, and $\Phi\left(x^{(1)}\right)\equiv\left(\Phi\left(x_{1}\right),\ldots,\Phi\left(x_{p}\right)\right)^{\prime}$, where $\Phi\left(\cdot\right)$ denotes the standard normal cumulative distribution function and $x^{(1)}=\left(x_{1},\ldots,x_{p}\right)^{\prime}$.}
The set of instrumental functions that we consider is of the following form:
where
This set of instrumental functions transforms the conditional moment inequalities ((ref)) into infinitely many unconditional ones without loss of information. Accordingly, under Assumptions (ref)--(ref), $B_{I}$ is equivalent to
where $m(W_{i},W_{j},\beta,g)\equiv m\left(W_{i},W_{j},\beta\right)\cdot g(x_{i},x_{j})$ for $g\in{\cal G}$. We formalize this result as Lemma (ref) in Appendix (ref). Other kinds of instrumental functions introduced in AS are applicable with modifications. We define the sample moment function and sample variance function of $m(W_{i},W_{j},\beta,g)$, respectively, as
and
Note that $\bar{m}_{n}\left(\beta,g\right)$ and $\hat{\sigma}_{n}^{2}\left(\beta,g\right)$ are U-statistics of orders two and three, respectively. Because $\bar{m}_{n}\left(\beta,g\right)$ is a non-degenerate U-statistic of order two, the asymptotic variance of $\sqrt{n}\bar{m}_{n}\left(\beta,g\right)$ is $\mbox{Var}_{P}\left(E_{P}\left[m(W_{i},W_{j},\beta,g)\mid W_{i}\right]\right)$, which is equivalent to\footnote{For the variance of U-statistics, see, for example, van_der_Vaart_1998 (van_der_Vaart_1998, Ch. 12).}
Thus, $\hat{\sigma}_{n}^{2}\left(\beta,g\right)$ is a consistent estimator of the asymptotic variance of $\sqrt{n}\bar{m}_{n}\left(\beta,g\right)$. However, in practice, $\hat{\sigma}_{n}^{2}\left(\beta,g\right)$ could be zero for some $g \in \mathcal{G}$; as such we use the modification proposed by AS for $\hat{\sigma}_{n}^{2}\left(\beta,g\right)$. The modified version of $\hat{\sigma}_{n}^{2}\left(\beta,g\right)$ is
where $\hat{\sigma}_{n}^{2}=\hat{\sigma}_{n}^{2}\left(\beta,1\right)$, which is a consistent estimator of
and $\epsilon$ is a regularization parameter that takes some fixed positive value. In the simulation studies in Section (ref) and empirical application in Section (ref), we use $\epsilon=0.0001$ and $\epsilon =0.001$. Then, letting $g_{(a,b),(\tilde{a},\tilde{b}),r}(x_{i},x_{j})\equiv1\left[(x_{i},x_{j})\in J_{(a,b),(\tilde{a},\tilde{b}),r} \right]$, the test statistic at $\beta$ takes the form
where $\left[x\right]_{-}=-x$ if $x<0$ and $\left[x\right]_{-}=0$ if $x\geq0$. $\left((2r)^{p} \cdot (|\mathcal{X}|)^{k-p}\right)^{2}$ corresponds to the number of instrumental functions $g_{(a,b),(\tilde{a},\tilde{b}),r}$ for a fixed $r$. This test statistic is a version of AS's test statistic that is extended to U-statistics of order two. Here the inner summation is taken over two set of indices, $(a,b)$ and $(\tilde{a},\tilde{b})$. In the implementation, we instead use an approximate test statistic at $\beta$:
where $R$ is some truncation integer chosen by the researcher. To compute a critical value for $T_{n,R}(\beta)$, we propose using an asymptotic approximation version of the critical value. This is a simulated quantile of
where $\left(v_{n}\left(\beta,g\right)\right)_{g\in{\cal G}}$ is a zero mean Gaussian process with a covariance kernel evaluated by
In the expression of $T_{n,R}^{Asy}(\beta)$, $\left(v_{n}\left(\beta,g_{(a,b),(\tilde{a},\tilde{b}),r}\right)\right)_{(a,b),(\tilde{a},\tilde{b}),r}$ approximates the asymptotic distribution of
$\varphi_{n}\left(\beta,g_{(a,b),(\tilde{a},\tilde{b}),r}\right)$ is a generalized moment selection (GMS) function to select binding moment restrictions and is given by
where $B_{n}$ and $\kappa_{n}$ are two tuning parameters that should satisfy $B_{n}\rightarrow\infty$, $\kappa_{n}\rightarrow\infty$, and $\kappa_{n}/n^{1/2}\rightarrow 0$ as $n\rightarrow\infty$ a.s. In Sections (ref) and (ref), we use $B_{n}=\left(0.8\ln\left(n\right)/\ln\ln\left(n\right)\right)^{\frac{1}{2}}$ and $\kappa_{n}=\left(\left(1-\hat{p}_{1-D}^{1/3}\right)^{2/5}\times0.6\ln(n)\right)^{\frac{1}{2}}$, where $\hat{p}_{1-D}\equiv n^{-1}\sum_{i}^{n}\left(1-D_{i}\right)$ is the sample censoring rate.\footnote{The values for $\kappa_{n}$ and $B_{n}$ are different from the values recommend by AS, which are $\kappa_{n}=\left(0.3\ln(n)\right)^{\frac{1}{2}}$ and $B_{n}=\left(0.4\ln\left(n\right)/\ln\ln\left(n\right)\right)^{\frac{1}{2}}$, respectively.} This value of $\kappa_{n}$ decreases with the sample censoring rate. The following assumption summarizes the requirements for the tuning parameters in the GMS function.
For a significance level of $\alpha<1/2$, the critical value is set to be the $1-\alpha+\eta$ simulated quantile of $T_{n,R}^{Asy}(\beta)$, where $\eta$ is an arbitrarily small positive value (e.g., $10^{-6}$ following AS). Letting $\hat{c}_{n,\eta,1-\alpha}(\beta)$ be the $1-\alpha+\eta$ quantile of $T_{n,R}^{Asy}(\beta)$, a nominal level $1-\alpha$ confidence set is computed as
Note again that the inference approach presented above is for U-statistics of order two, but can be extended to U-statistics of a greater order with some modification. Namely, the class of instrumental functions, the moment function, and the sample variance function need to be modified for higher order; the inner double summations in the test statistics need to be replaced by more summations.
This subsection provides the uniform asymptotic size and power properties of the inference method. Let ${\cal Q}$ be the collection of all pairs of the regression parameters in $B$ and distributions, $(\beta,P)$, such that the conditional moment inequality condition ((ref)) holds, $\{W_i:i \geq 1\}$ are i.i.d. under $P$, and the support of $X$ is contained in $\mathcal{X}(=\mathcal{X}_{1}\times\mathcal{X}_{2})$. Define
which is the covariance kernel between $m(W_{i},W_{j},\beta,g )$ and $m(W_{i},W_{j},\beta,g^{\ast} )$ under the distribution $P$. Let ${\cal H}_{2}\equiv \{h_{2,P}(\beta,\cdot,\cdot):(\beta,P)\in \mathcal{Q}\}$ be the collection of all possible covariance kernel functions on ${\cal G}\times{\cal G}$. For any $\beta \in B$ and the true distribution $P$, let
The following theorem gives the uniform size and power properties of the proposed inference method.
A proof of this theorem is provided in Appendix (ref). Theorem (ref) (a) states that the proposed confidence set is asymptotically conservative, which corresponds to Theorem 2(b) of AS. The uniformity in the statement enables the asymptotic result to provide a good finite sample approximation, which is well discussed in AS. Theorem (ref) (b) states that the test is consistent against a fixed alternative.
This section presents numerical examples and Monte Carlo simulation results. The numerical examples show how $B_{I}$ and $T_{B_I}$ vary with the degree of censoring and the support of covariates. The Monte Carlo simulations show the finite sample performance of the proposed inference method for the regression parameters and demonstrate how it varies with the choice of the tuning parameters. Some additional numerical example and Monte Carlo simulation studies are presented in the supplementary material to this paper.
This subsection illustrates how $B_{I}$ and $T_{B_{I}}$ vary depending on the degree of censoring and the support of covariates. We consider three MPH models (Models 1--3) with endogenous censoring that have the following form:
In all the models, we set $\left(\beta_{1},\beta_{2}\right)=\left(0.5,1.5\right)$ and $\left(\gamma_{0},\gamma_{1},\gamma_{2}\right)=\left(-0.5,0.5,-1\right)$. In Models 1--3, we set $\alpha_{0}=+\infty$, $3$, and $1.6$, respectively. In Model 1, there is no censoring; in Models 2 and 3, there is censoring correlated with the covariates and unobserved heterogeneity. The outcome is more likely to be censored in Model 3 than in Model 2. In all the models, $U$, $V$, and $W$ have independent unit exponential distributions, and $X_{2}$ takes values in $\left\{ 0,1\right\} $. Regarding $X_{1}$, we consider three cases; $X_{1}$ takes values in (i) $\left\{ -2.5,-2.0,\ldots,2.5\right\} $, (ii) $\left\{ -5,-2.5,\ldots,5\right\} $, or (iii) $\left\{ -5,-4.8,\ldots,5\right\} $. In these data generating processes (DGPs), the censoring is endogenous because $Y^\ast$ is correlated with $C$ even conditional on $X_1$ and $X_2$, which occurs due to the presence of $U$ in both the equations for $\log Y^\ast$ and $\log C$. In the computation of $B_I$ and $T_{B_I}$, we set $|\beta_1|=1$ for the scale normalization and $T(\tilde{y})=0$ for the location normalization, where we set $\tilde{y}=0.77$, which is the median of $Y^\ast$ when $X_{1}$ is independently distributed as $N(0,2)$ and $X_{2}$ is independently and uniformly distributed over $\{0,1\}$. This joint distribution will be used in the Monte Carlo simulation in the following subsection. After the scale and location normalizations, the true normalized values of $\beta_2$ and $T(y)$ are $3$ and $T_{0}(y):=2(\log y - \log 0.77)$, respectively. Given the DGPs described above, $B_{I}$ is the set of regression parameter values that satisfy the conditional moment inequality ((ref)) for all pair of values of $\left(X_{1},X_{2}\right)$, and $T_{B_{I}}(y)$ for a fixed $y \in \mathbb{R}$ is the set of values $t$ that satisfy the inequality in ((ref)) for some $\beta \in B_{I}$ and all pair of values of $(X_{1},X_{2})$. We numerically obtain $B_{I}$ and $T_{B_I}(y)$ by simulating the distributions of $\log Y^\ast$ and $\log C$ for each pair of values of $\left(X_{1},X_{2}\right)$ using 20,000 random draws from each DGP. Table (ref) presents the numerical results for $B_{I}$. Each cell in the table shows the interval of $\beta_2$ obtained by the projection of the computed $B_I$ for each model and each support of $X_{1}$. As expected, the interval shrinks as the censoring rate decreases or the support becomes wider or finer. In each model, expanding the support from (i) to (ii) does not shrink the lower or upper bounds on $\beta_2$.
Figure (ref) illustrates the computed $T_{B_{I}}$ for each model and each support of $X_{1}$, along with the true normalized transformation function $T_0$. As mentioned in Section (ref), $T_{B_{I}}$ is not bounded from above. The result shows that as the censoring rate decreases or the support of $X_1$ becomes wider or finer, the lower bound on $T_{B_{I}}$ becomes closer to the true transformation function. The lower bound also becomes smoother as the support of $X_1$ becomes finer. The supplementary material to this paper gives two additional numerical examples for $B_{I}$. In the first example, we compute $B_{I}$ in models with three covariates, where the results in Table (ref) are extended to two-dimensional figures. In the second example, using slightly different models and supposing that the transformation model is parametrically specified (i.e., the functional forms of $\Lambda$ and $F_{U}$ are known), we compute the sharp identified set of $\beta_0$, based on the result in Szydlowski_2019, and compare it with $B_{I}$.
This subsection presents Monte Carlo simulation results to evaluate the finite sample size and power properties of the inference method for the regression parameters presented in Section (ref). We here use two DGPs (DGP1 and DGP2). In DGP1 and DGP2, the data are derived from Models 2 and 3, respectively, where $X_{1}$ is distributed as $N(0,2)$; $X_{2}$ is uniformly distributed over $\left\{ 0,1\right\}$; $U$, $V$, and $W$ have independent unit exponential distributions. The censoring rates in DGP1 and DGP2 are approximately $16\%$ and $30\%$, respectively. For the Monte Carlo experiments, 500 samples are randomly drawn with sample sizes $n=100$, $250$, and $500$. Based on the inference method, we conduct a test of $H_{0}:$ ((ref)) holds against $H_{1}:$ ((ref)) is violated at each value of $\left(\beta_{1},\beta_{2}\right)\in\left\{ 1\right\} \times\left\{ -1,-0.5,\ldots,9\right\} $. The true value of the normalized ${\beta}_{2}$ is 3. Critical values are simulated using 1,000 repetitions for the significance level $\alpha=0.05$. As a baseline case, we set the tuning parameters in the test statistics and GMS function to $R=5$, $\epsilon = 0.0001$, $B_{n}=B_{n}^{bc}\equiv\left(0.8\ln\left(n\right)/\ln\ln\left(n\right)\right)^{\frac{1}{2}}$, and $\kappa_{n}=\kappa_{n}^{bc}\equiv\left(\left(1-\hat{p}_{1-D}^{1/3}\right)^{2/5}\times0.6\ln(n)\right)^{\frac{1}{2}}$. We set $\eta=10^{-6}$ throughout this and the following sections. We do not assume that the researcher knows the exact distribution of $X_{1}$; hence, we transform $X_{1}$ into $\tilde{X}_{1}$ as described in Section (ref) and use $\tilde{X}_{1}$ instead of $X_{1}$. We assume that the researcher knows the support of $X_2$. Thus we apply the inference method presented in Section (ref) with $\mathcal{X}=\mathcal{X}_{1} \times \mathcal{X}_{2} = [0,1]\times \{0,1\}$. Figure (ref) shows the graphs of rejection frequencies for DGP1 and DGP2 given the baseline case of the tuning parameters with $\epsilon=0.0001$ or $\epsilon = 0.001$. All the rejection frequencies at the true value of normalized $\beta_2$ are close to or lower than the nominal size $\alpha=0.05$ (though the rejection frequency substantially exceeds 0.05 in DGP1 with $n=100$ and $\epsilon=0.0001$). The rejection frequencies are also close to or lower than 0.05 in the corresponding intervals computed in Table (ref) with Support (iii). When $\epsilon=0.001$, the test has less power but is more likely to be conservative even in small sample. The power of the test is higher at small values of $\beta_{2}$ than at large values. Table (ref) shows the rejection frequencies at the true parameter value and at $\left(\beta_{1},\beta_{2}\right)=\left(1,0\right)$ for several choices of the tuning parameters in DGP1 and DGP2. The point $\left(\beta_{1},\beta_{2}\right)=\left(1,0\right)$ is not contained in $B_{I}$, as seen from the results of the numerical examples. Table (ref) shows the degree of sensitivity of the test to variation in the sample size $n$, the choice of the truncation integer $R$ in the approximate test statistic, the value of $\epsilon$ for the modified variance estimator $\bar{\sigma}_{n}^{2}\left(\beta,g\right)$, and the choice of $B_n$ and $\kappa_n$ in the GMS function. In the case of the last row of Table (ref), $\kappa_{n}$ does not depend on the sample censoring rate $\hat{p}_{1-D}$. The results show that there is some sensitivity to the sample size, the choices of $\kappa_n$ and $R$, and the value of $\epsilon$. In particular, the sensitivity to the choice of $\kappa_{n}$ is high. A small value of $\epsilon$ leads to the test having high power, but can lead to the test having incorrect size in small samples. The supplementary material contains two additional Monte Carlo simulation results. One examines inference on the regression parameters in models with three covariates. The other evaluates the performance of a joint inference method for the regression parameters and transformation function introduced in Appendix (ref).
We apply the proposed inference method to evaluate the effect of heart transplants on patients' survival duration using the Stanford Heart Transplant Data taken from Kalbfleisch_Prentice_1980. The data set consists of survival times (in days) of 103 patients; an indicator of censoring, which takes the value one if the patient is dead (uncensored) or zero if the patient is censored; an indicator of receiving a heart transplant, which takes the value one if the patient receives a heart transplant or zero otherwise; and the age (in years) of patients at the time of acceptance into the program. Among the 103 patients, $27\%$ (28 patients) are censored due to attrition or administrative censoring. The censoring rates for the treated (receive a transplant) and untreated (do not receive a transplant) groups are $35\%$ and $22\%$, respectively. We consider the following censored transformation model,
where, for each patient $i$, $Y_{0i}$ is the observed survival time, $X_{i,age}$ is the age, $X_{i,treat}$ is the transplant indicator, $U_{i}$ is unobserved heterogeneity, $C_{i}$ is the censoring time, and $T$ is a strictly increasing function. Applying the proposed method, we allow the censoring to be arbitrarily correlated with the patient's age and unobserved heterogeneity. Furthermore, we do not specify the transformation function or the distribution function of the patient's unobserved heterogeneity. For scale normalization, we set $\left|\beta_{age}\right|=1$; for location normalization, we set $\beta_{cons}=0$ and $T(\tilde{y})=0$ with $\tilde{y}=90$ being the median of $Y_{0}$ in the sample. Our interest is in the normalized regression parameter $\beta_{treat}$ and the normalized transformation function $T$. We compare the proposed method with the partial rank estimator (PRE) proposed by Khan_Tamer_2007. This estimator is robust up to covariate-dependent censoring and consistently estimates the normalized regression parameters in the nonparametric transformation model. Table (ref) shows the inference results for $\beta_{treat}$. It presents the point estimate obtained from the PRE and 95% confidence intervals obtained from the PRE and the proposed method. The confidence interval obtained from the PRE is computed based on 1,000 bootstrap pseudo samples from the data. For the proposed method, we set the tuning parameters to $R=5$, $B_{n}=\left(0.8\ln\left(n\right)/\ln\ln\left(n\right)\right)^{\frac{1}{2}}$, $\kappa_{n}=\left(\left(1-\hat{p}_{1-D}^{1/3}\right)^{2/5}\times0.6\ln(n)\right)^{\frac{1}{2}}$, and $\eta=10^{-6}$ as in the baseline case in the Monte Carlo simulation in the previous section. For $\epsilon$, we use both $\epsilon=0.0001$ and $\epsilon=0.001$. Setting $\epsilon=0.001$ is more conservative in a small sample like the Stanford Heart Transplant data set according to the Monte Carlo simulation results in the previous section. The confidence intervals obtained from the proposed method do not have finite upper bounds. This would be because age does not have sufficiently large support to derive a finite upper bound on $\beta_{treat}$ in $B_I$. The estimate obtained from the PRE is positive and is significantly different from zero. The 95% confidence interval obtained from the proposed method is also entirely positive regardless of the choice of $\epsilon$. With the conservative choice of $\epsilon=0.001$, the $95\%$ confidence interval obtained from the proposed method covers the confidence interval obtained from the PRE. The inference results relating to the proposed method show that even if censoring is arbitrarily correlated with a patient's age or unobserved heterogeneity, a heart transplant has a positive effect on the patient's survival time.
We next apply the joint inference procedure for the regression parameters and transformation function presented in Appendix (ref). Figure (ref) shows the marginal $95\%$-confidence set of $T(y)$ at each $y \in [0,900]$, which is the projection of the three-dimensional confidence interval of $(\beta_{age},\beta_{treat},T(y))$ on the one-dimension. The same tuning parameters as those used in Table (ref) are used. Since $T_{I,\beta}(y)$ does not have a finite upper bound, neither do the confidence intervals. The estimated confidence intervals also do not have finite lower bounds for small values of $y$. The $95\%$-confidence lower bound on $T(y)$ increases rapidly with $y$ in the approximate interval $[100,250]$, but it changes slightly when $y$ is large.
In this paper, we propose a partial identification and inference approach for a nonparametric transformation model in the presence of endogenous censoring. We develop bounds on the regression parameters and the transformation function, each of which is characterized by conditional moment inequalities involving U-statistics. We also characterize the sharp identified set of the regression parameters, using concepts from random set theory, though this set is hard to compute. A comparison of the proposed set and the sharp identified set characterization makes it clear when the proposed set approaches the sharp set. Based on the identification result, we propose an inference method for the regression parameters by extending the inference approach for conditional moment inequality models, proposed by Andrews_Shi_2013, to the U-statistics case. We also derive the asymptotic properties of this approach. A joint inference procedure for the regression parameters and transformation function is presented in Appendix (ref). Numerical examples illustrate the characteristics of the proposed sets for the regression parameters and transformation function, and the results of Monte Carlo experiments demonstrate the size and power properties of the proposed inference methods. As an empirical application, we apply the inference methods to evaluate the effect of heart transplants on patients' survival duration using data from the Stanford Heart Transplant Study, for which we find that heart transplants have a positive effect on patients' survival duration regardless of the censoring mechanism.