EconBase
← Back to paper

Testing for Unobserved Heterogeneity in Censored Duration Models: EM Approach

The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.

60,097 characters

Testing for Unobserved Heterogeneity in Censored Duration Models: EM Approach



\title{\textbf{Testing for Unobserved Heterogeneity in Censored Duration Models: EM Approach}}
\author{Hiroyuki Kasahara \\
Vancouver School of Economics\\
University of British Columbia \\
[email removed]
\and
Hirokazu Matsuyama\\
DENTSU SOKEN INC.\\
[email removed]
\and
Katsumi Shimotsu\thanks{Corresponding author.}\\
Faculty of Economics \\
University of Tokyo\\
[email removed]
\and
Shota Takeishi\\
Department of Statistics and Data Science\\
Washington University in St.~Louis\\
[email removed]}
\date{September 29, 2026}

\maketitle

\begin{abstract}

Ignoring unobserved heterogeneity in duration models biases parameter estimates and invalidates inference, but testing for it is non-regular: the null hypothesis lies on the boundary of the parameter space and some parameters are unidentified under the null. These features render standard asymptotic theory inapplicable. This paper develops an EM test for unobserved heterogeneity in censored Weibull duration models, building on the EM approach of \citet*{lcm09bm}. The test statistic has an asymptotic null distribution equal to the square of $\max\{0, N(0,1)\}$, hence critical values require neither simulation nor bootstrap, and the test accommodates covariate-dependent censoring of arbitrary form. Monte Carlo simulations compare the EM test with the likelihood ratio test (LRT) of \citet{chowhite10joe}, information matrix tests, and Lagrange multiplier tests. The EM test has empirical size close to the nominal level for sample sizes of 500 or more, where the LRT remains markedly conservative and the other tests over-reject. Its size-adjusted power is comparable to that of the LRT and higher than that of the other tests. Because size adjustment requires knowledge of the data-generating process and is unavailable in practice, the EM test attains higher power than the LRT in most designs as the tests would actually be applied. In an application to the Stanford Heart Transplant data, the EM test rejects homogeneity in every covariate specification, whereas the LRT's conclusion depends on a user-chosen set of admissible parameter values and on the specification.
\end{abstract}

Keywords: censored models; duration models; EM test; unobserved heterogeneity; Weibull distribution.

JEL Classification Codes: C12, C24, C41

\section{Introduction}

Weibull duration models have been widely used in economics to analyze transitions, such as unemployment spells \citep{Lancaster1979}, plant survival \citep{Chen02ijio}, firm survival \citep{Kalnins13aej}, workers' compensation claims \citep{Butler01restat}, school-to-work transitions \citep{Pastore21labour}, and waiting times between stock transactions \citep{Engle98em}. \citet{Kiefer88jel}, \citet{Lancaster90}, and \citet{vandenberg01hdbk} review the theory and applications of duration models.

As demonstrated by \citet{heckmansinger84em}, parameter estimates in duration models can be substantially affected by assumptions concerning unobserved heterogeneity, and failing to account for it may lead to bias. Consequently, several tests for unobserved heterogeneity have been developed. \citet{Chesher1984} and \citet{Lancaster1984, Lancaster1985} develop information matrix (IM) tests based on the fact that neglected unobserved heterogeneity violates the information matrix equality of the homogeneous model. \citet{Kiefer1985}, \citet{Sharma1987}, and \citet{Prieger2003} propose Lagrange multiplier (LM) tests that exploit the Laguerre polynomial representations of the exponential and Weibull distributions. A parallel literature in biostatistics develops score tests for homogeneity across groups in survival data \citep{commengesandersen95lida, bolfarinevalenca05bj}.

\citet{chowhite10joe} develop a likelihood ratio test (LRT) to detect unobserved heterogeneity in exponential and Weibull duration models using a finite mixture framework. Their procedure tests the null hypothesis of no unobserved heterogeneity against a two-component finite mixture alternative. The LRT statistic has a nonstandard asymptotic distribution due to the presence of unidentified parameters and boundary issues under the null hypothesis. Their simulations show that the test has power against a broad range of heterogeneity distributions and better size and power properties than the IM and LM tests.

Many duration datasets contain censored observations due to survey-based sampling designs. For example, in unemployment duration studies, some data are obtained from individuals who are currently unemployed, rather than from tracking individuals over the full course of their unemployment spells. In such cases, only part of the spell is observed. Moreover, the censoring time may depend on individual characteristics such as age and education.

In censored duration models, existing tests for unobserved heterogeneity face challenges related to both size control and implementation. First, the LRT of \citet{chowhite10joe}, the IM test, and the LM tests suffer from finite-sample size distortions: the IM and LM tests are typically severely oversized, while the LRT tends to be undersized. Second, the LRT is difficult to apply under covariate-dependent censoring because its asymptotic null distribution is known only for the cases of fixed censoring and random (independent) censoring. Third, implementing the LRT is computationally intensive, as its asymptotic null distribution is a functional of a Gaussian process and is approximated by the weighted bootstrap of \citet{hansen96em}. Finally, the LRT requires the user to specify a set of admissible values for the mixture components, and its asymptotic null distribution depends on that set. In finite samples, the size and the power of the LRT, and hence its conclusions, are sensitive to that choice.

This paper develops a new test for unobserved heterogeneity in censored Weibull duration models that is straightforward to implement and exhibits good finite-sample performance. The proposed test, referred to as the EM test, adopts the finite mixture framework of \citet{chowhite10joe} and employs the EM approach \citep{lcm09bm, chenli09as, kasaharashimotsu15jasa} to test the null hypothesis of no unobserved heterogeneity against a two-component finite mixture alternative.

The proposed test offers several attractive features. First, the asymptotic null distribution of the test statistic is the square of $\max\{0, N(0,1)\}$. Critical values therefore require neither simulation nor the bootstrap. Second, the test accommodates covariate-dependent censoring of arbitrary form without requiring special adjustments, provided that the duration and the censoring time are conditionally independent given the covariates.
Third, the test demonstrates good finite-sample performance: its empirical size is close to the nominal level when the sample size is 500 or more, under both covariate-independent and covariate-dependent censoring. Moreover, it exhibits better size control than the LRT, IM, and LM tests. Fourth, for sample sizes of 500 or more, the proposed test achieves size-adjusted power higher than that of the IM and LM tests and comparable to that of the LRT. Size adjustment is unavailable in practice, since it requires knowledge of the data-generating process, including the censoring mechanism. Without size adjustment, the EM test attains higher power than the LRT in most designs because the LRT is markedly conservative. Furthermore, an application to the Stanford Heart Transplant data illustrates the practical difference: the EM test rejects homogeneity at the 1\% level in all five covariate specifications, whereas the LRT fails to reject at the 5\% level under its narrowest admissible set in every specification and rejects in most of them under wider admissible sets.

Testing and inference in finite mixture models have attracted considerable interest in econometrics. \citet{chenponomarevatamer14joe} study inference on the mixing probability in a two-component mixture and show that the profiled likelihood ratio statistic has asymptotic limits that vary discontinuously near the region where the mixture is not identified; they use a parametric bootstrap to restore uniform coverage. \citet{gukoenkervolgushev18et} develop an LRT based on a nonparametric maximum likelihood estimator of the mixing distribution. They derive its asymptotic distribution, but find that the asymptotic critical values are unlikely to control size and instead advocate a parametric bootstrap, whose consistency they establish.

A related literature relaxes the parametric and censoring assumptions of the Weibull model. \citet{KhanTamer07joe} develop a rank-based estimator of the regression coefficients in nonparametric censored duration models that allow for general forms of censoring. \citet{KhanTamer09joe} propose an instrumental variables estimator that allows for covariate-dependent censoring and endogenous censoring. \citet{Szydlowski19jae} and \citet{Sakaguchi24jae} analyze the set identification of parameters in duration models with endogenous censoring. \citet{Han1990} and \citet{Meyer1990} combine flexible hazard functions with parametric assumptions about unobserved heterogeneity.

The remainder of this paper is organized as follows. Section \ref{sec:Weibull_model} introduces the Weibull duration model with censoring. Section \ref{sec:tests} formulates unobserved heterogeneity as a two-point mixture, develops the EM test and derives its asymptotic null distribution, and reviews the LRT of \citet{chowhite10joe}. Section \ref{sec:MC} reports Monte Carlo simulation results under covariate-independent and covariate-dependent censoring. Section \ref{sec:realdata} applies the tests to data from the Stanford Heart Transplant Program. Section \ref{sec:conclusion} concludes. Appendices \ref{app:assumptions}--\ref{app:aux} collect the assumptions, proofs, and auxiliary results, respectively.

We use the following notation throughout. All limits below are taken as $n \rightarrow \infty$, unless stated otherwise. We use boldfaced letters to denote vectors and matrices. Let $\coloneq $ denote ``equals by definition.'' Let $\mathbb{R}^+ \coloneq (0,\infty)$. Let $\mathds{1}\{A\}$ denote an indicator function that takes the value of one when $A$ is true and zero otherwise. For a $k\times 1$ vector $\boldsymbol{a}$, let $\|\boldsymbol{a}\|$ and $\|\boldsymbol{a} \|_\infty$ denote the Euclidean norm and the uniform norm, respectively. For a $k\times 1$ vector $\boldsymbol{a}$ and a function $f(\boldsymbol{a})$, let $\nabla_{\boldsymbol{a}}f(\boldsymbol{a})$ denote the $k\times 1$ vector of the derivative $(\partial/ \partial {\boldsymbol{a}}) f(\boldsymbol{a})$, and let $\nabla_{\boldsymbol{a}\boldsymbol{a}^{\top}}f(\boldsymbol{a})$ denote the $k\times k$ matrix of the derivative $(\partial^2/ \partial{\boldsymbol{a}}\partial{\boldsymbol{a}}^{\top}) f(\boldsymbol{a})$.

\section{Weibull duration model with censoring}\label{sec:Weibull_model}

This section introduces the Weibull duration model with censoring. Our notation largely follows that of \citet{chowhite10joe}, except that we write $\varphi$ for their regression function $g$. Let $\{(Y_t, C_t, \boldsymbol{X}_t)\}$ be a strictly stationary geometric $\beta$-mixing stochastic process, where $Y_t$ and $C_t$ are scalar-valued and strictly positive, and $\boldsymbol{X}_t$ is $\mathbb{R}^k$-valued. The $\beta$-mixing condition allows $\boldsymbol{X}_t$ to depend on the entire history of $\{Y_s\}_{s<t}$, thus accommodating autoregressive conditional duration models. Here, $Y_t$ denotes the duration, and $C_t$ denotes the censoring variable. We allow for covariate-dependent censoring, meaning that $C_t$ may be arbitrarily correlated with the covariates $\boldsymbol{X}_t$. Under right censoring, the observed duration is $Y_t^{c}\coloneq \min\{Y_t, C_t\}$. Let $D_t \coloneq \mathds{1}\{Y_t^{c} = Y_t\}$, which equals 1 if the observed duration is not censored. In this setting, we observe $\{(Y_t^c, D_t,\boldsymbol{X}_t)\}_{t=1}^n$, although $C_t$ is not observed when $D_t = 1$.

Following \citet{KhanTamer07joe}, we assume throughout that $Y_t$ and $C_t$ are conditionally independent given $\boldsymbol{X}_t$.\footnote{For dependent observations, Assumption \ref{assn1}(ii) in the Appendix additionally requires the joint conditional law of $(Y_t,C_t)$ given $\boldsymbol{X}_t$ and the full past to depend only on $\boldsymbol{X}_t$.} This assumption rules out endogenous censoring, in which $C_t$ is related to the duration itself beyond what $\boldsymbol{X}_t$ captures.

We assume that the uncensored duration $Y_t$ conditional on $\boldsymbol{X}_t$ follows a Weibull distribution with the following parameterization
\begin{equation}\label{eq:wblpdf}
f(y| \boldsymbol{x}; \delta^*, \boldsymbol{\beta}^*, \gamma^*) = \delta^* \gamma^* \varphi(\boldsymbol{x}; \boldsymbol{\beta}^*)y^{\gamma^*-1}\exp(-\delta^* \varphi(\boldsymbol{x};\boldsymbol{\beta}^*) y^{\gamma^*}) ,
\end{equation}
where $(\delta^*,\boldsymbol{\beta}^*,\gamma^*) \in D \times B \times \Gamma \subset \mathbb{R}^+ \times \mathbb{R}^d \times \mathbb{R}^+$, and $\varphi(\cdot;\boldsymbol{\beta})$ is strictly positive for every $\boldsymbol{\beta} \in B$. We assume that $\boldsymbol{X}_t$ does not include a constant term, because an intercept in $\boldsymbol{\beta}^*$ would not be separately identified from $\delta^*$: in the widely used specification $\varphi(\boldsymbol{x}; \boldsymbol{\beta}^*) = \exp(\boldsymbol{x}^{\top}\boldsymbol{\beta}^*)$, we have $\delta^* \varphi(\boldsymbol{x}; \boldsymbol{\beta}^*) = \exp(\log \delta^* + \boldsymbol{x}^{\top}\boldsymbol{\beta}^*)$, so that $\log \delta^*$ serves as the intercept.

We model the joint conditional density of $(Y_t^{c}, D_t)$ given $\boldsymbol{X}_t$ as
\begin{equation}\label{eq:model1}
\begin{aligned}
f^{c}(y^{c},d | \boldsymbol{x};\delta^*,\boldsymbol{\beta}^*,\gamma^*) \coloneq & \left\{f(y^{c} | \boldsymbol{x};\delta^*,\boldsymbol{\beta}^*,\gamma^*)\right\}^{d}\left\{1-F(y^{c} | \boldsymbol{x};\delta^*,\boldsymbol{\beta}^*,\gamma^*)\right\}^{1-d} \\
& \times \left\{1-G(y^{c}|\boldsymbol{x})\right\}^{d}\left\{g(y^{c}|\boldsymbol{x})\right\}^{1-d},
\end{aligned}
\end{equation}
where $f(\cdot| \boldsymbol{x}; \delta^*, \boldsymbol{\beta}^*, \gamma^*)$ is defined by \eqref{eq:wblpdf}, $F(\cdot| \boldsymbol{x};\delta^*,\boldsymbol{\beta}^*,\gamma^*)$ is the distribution function of $Y_t$ given $\boldsymbol{X}_t$, and $G(\cdot|\boldsymbol{x})$ is the distribution function of $C_t$ given $\boldsymbol{X}_t$. Here $g(\cdot|\boldsymbol{x})$ denotes the density of $G(\cdot|\boldsymbol{x})$ with respect to a $\sigma$-finite measure $\nu_{\boldsymbol{x}}$ that dominates the conditional distribution of $C_t$ given $\boldsymbol{X}_t = \boldsymbol{x}$, so that $(Y_t, C_t)$ has a conditional density with respect to the product of Lebesgue measure and $\nu_{\boldsymbol{x}}$.\footnote{Correspondingly, \eqref{eq:model1} is a density of $(Y_t^c, D_t)$ with respect to the measure that equals Lebesgue measure on $\{d = 1\}$ and $\nu_{\boldsymbol{x}}$ on $\{d = 0\}$, because $Y_t^c = Y_t$ when $D_t = 1$ and $Y_t^c = C_t$ when $D_t = 0$. When $C_t$ is continuously distributed, $\nu_{\boldsymbol{x}}$ is Lebesgue measure and $g(\cdot|\boldsymbol{x})$ is the ordinary conditional density; when the conditional distribution of $C_t$ is discrete or degenerate, $\nu_{\boldsymbol{x}}$ is counting measure on its support.} This formulation imposes no restriction on the form of $G(\cdot|\boldsymbol{x})$ beyond the conditional independence of $Y_t$ and $C_t$ given $\boldsymbol{X}_t$. For maximum likelihood estimation in this model, we can drop the terms $\{1-G(y^{c}|\boldsymbol{x})\}^{d}$ and $\{g(y^{c}|\boldsymbol{x})\}^{1-d}$ from the likelihood function because they enter multiplicatively and do not depend on the model parameters.

Model \eqref{eq:model1} encompasses random censoring and fixed (Type I) censoring as special cases. In the case of random censoring, $C_t$ is independent of $(Y_t,\boldsymbol{X}_t)$. Consequently, model \eqref{eq:model1} simplifies to
\begin{align}\label{eq:model_random}
f^{cr}(y^{c},d | \boldsymbol{x};\delta^*,\boldsymbol{\beta}^*,\gamma^*) \coloneq & \left\{f(y^{c} | \boldsymbol{x};\delta^*,\boldsymbol{\beta}^*,\gamma^*)\right\}^{d}\left\{1-F(y^{c} | \boldsymbol{x};\delta^*,\boldsymbol{\beta}^*,\gamma^*)\right\}^{1-d}\notag \\
	& \times \left\{1-G(y^{c})\right\}^{d}\left\{g(y^{c})\right\}^{1-d}.
\end{align}
In the case of fixed censoring, $C_t=c<\infty$ with probability one. Then, model \eqref{eq:model1} reduces to
\begin{align}\label{eq:model_fixed}
f^{cf}(y^{c},d | \boldsymbol{x};\delta^*,\boldsymbol{\beta}^*,\gamma^*) \coloneq \left\{f(y^{c} | \boldsymbol{x};\delta^*,\boldsymbol{\beta}^*,\gamma^*)\right\}^{d}\left\{1-F(c| \boldsymbol{x};\delta^*,\boldsymbol{\beta}^*,\gamma^*)\right\}^{1-d}
\end{align}
on the support $\{(y^c,1): 0 < y^c < c\} \cup \{(c,0)\}$: uncensored observations satisfy $0 < y^c < c$, and censored ones form an atom at $(c,0)$ with probability $1-F(c|\boldsymbol{x};\delta^*,\boldsymbol{\beta}^*,\gamma^*)$.

\section{Tests for unobserved heterogeneity}\label{sec:tests}

This section develops a test for the presence of unobserved heterogeneity in $\delta^*$ based on a finite mixture model.

\subsection{Modeling unobserved heterogeneity using a mixture model}

Following \citet{heckmansinger84em} and \citet{chowhite10joe}, we model unobserved heterogeneity in $\delta^*$ using a two-point discrete distribution with point masses at $\delta_{1}^*$ and $\delta_{2}^*$. The discrete mixture specification has been widely adopted in the literature; see, for example, \citet{Gritz1993}, \citet{Ferrall1997}, and \citet{vandenberg98em}. As argued by \citet{chowhite10joe}, finite mixture models offer a flexible approximation to the unknown heterogeneity distribution and facilitate tests capable of detecting a broad range of unobserved heterogeneity. A commonly used alternative is the gamma mixture model, which assumes that $\delta^*$ follows a gamma distribution and offers an analytically convenient framework for incorporating unobserved heterogeneity into duration models \citep{Lancaster1979, Meyer1990, Han1990}. However, this approach has been criticized by \citet{heckmansinger84em} as being somewhat ad hoc. In particular, the gamma distribution may fail to capture key features of the true heterogeneity distribution, such as multimodality or skewness. Another alternative is to leave the heterogeneity distribution unspecified and estimate it by nonparametric maximum likelihood \citep{gukoenkervolgushev18et}.

With a two-point distribution of $\delta^*$, the joint conditional density of $(Y_t^{c}, D_t)$ given $\boldsymbol{X}_t$ is written as a mixture of Weibull densities
\begin{align}\label{eq:model_mixture}
f^{c}_{a}(y^{c},d | \boldsymbol{x};\pi^*,\delta_{1}^*, \delta_{2}^*,\boldsymbol{\beta}^*,\gamma^*) \coloneq & \pi^* f^{c}(y^{c},d | \boldsymbol{x};\delta_1^*,\boldsymbol{\beta}^*,\gamma^*) + (1-\pi^*)f^{c}(y^{c},d | \boldsymbol{x};\delta_2^*,\boldsymbol{\beta}^*,\gamma^*) ,
\end{align}
where $f^{c}(y^{c},d | \boldsymbol{x};\delta^*,\boldsymbol{\beta}^*,\gamma^*)$ is defined in \eqref{eq:model1} and $(\pi^*,\delta_{1}^*,\delta_{2}^*,\boldsymbol{\beta}^*,\gamma^*) \in [0,1] \times D\times D \times B \times \Gamma \subset [0,1] \times \mathbb{R}^{+} \times \mathbb{R}^{+}\times \mathbb{R}^{d} \times \mathbb{R}^{+}$. In this model, unobserved heterogeneity is absent when $\delta_{1}^* = \delta_{2}^*$ or $\pi^*(1-\pi^*) = 0$. For maximum likelihood estimation of the parameters in this model, we can drop the terms $\{1-G(y^{c}|\boldsymbol{x})\}^d$ and $\{g(y^{c}|\boldsymbol{x})\}^{1-d}$ in $f^{c}(y^{c},d | \boldsymbol{x};\delta,\boldsymbol{\beta},\gamma)$, because these terms enter multiplicatively in both $f^{c}(y^{c},d | \boldsymbol{x}; \delta_{1}, \boldsymbol{\beta}, \gamma)$ and $f^{c}(y^{c},d | \boldsymbol{x}; \delta_{2}, \boldsymbol{\beta}, \gamma)$ and $\pi + (1-\pi) = 1$.

Using model \eqref{eq:model_mixture}, we can express the null hypothesis of no unobserved heterogeneity as the parameter restrictions
\begin{equation} \label{eq:null_hypothesis}
H_0: \delta_{1}^* = \delta_{2}^*\ \text{ or }\ \pi^*(1-\pi^*) = 0\ \text{ against }\ H_1: \pi^* \in (0,1) \text{ and } \delta_{1}^* \neq \delta_{2}^*.
\end{equation}

\subsection{EM test for unobserved heterogeneity}

In this section, we develop an EM test for unobserved heterogeneity, building on the EM approach of \citet{lcm09bm}. The proposed test is primarily motivated by the limitations of the LRT when applied to testing $H_0$ against $H_1$ in equation \eqref{eq:null_hypothesis}. Under the null hypothesis, some parameters are not identified: in model \eqref{eq:model_mixture}, $\delta_2^*$ is not identified when $\pi^*= 1$, $\delta_1^*$ is not identified when $\pi^*= 0$, and $\pi^*$ is not identified when $\delta_1^* = \delta_2^*$. As documented by \citet{chowhite10joe}, this lack of identification under the null leads to several undesirable properties of the LRT in the context of censored duration models.

First, the asymptotic null distribution of the LRT statistic is unknown when censoring is covariate-dependent.\footnote{
\citet{chowhite10joe} derive the asymptotic distribution of the LRT statistic under fixed censoring and random censoring.} Second, even in the fixed and random censoring cases, the asymptotic null distribution of the LRT statistic is nonstandard and must be simulated because it is a functional of a Gaussian process whose covariance structure depends on the covariate and censoring mechanism in a complex way. Third, the LRT statistic and its asymptotic null distribution depend on the parameter space employed in maximizing the log-likelihood, so the choice of the parameter space affects the power of the LRT. Fourth, the LRT requires the lower bound of the parameter space of $\delta$ to exceed $0.5\delta^*$, a technical condition that ensures a finite variance of the likelihood ratio but has no economic justification.

Given these problems associated with the LRT, we focus on testing $\delta_{1}^* = \delta_{2}^*$ in $H_0$ in \eqref{eq:null_hypothesis}. Focusing on $\delta_{1}^* = \delta_{2}^*$ involves no loss of generality for the null data-generating process, because under either branch of $H_0$ the observed data are generated by a single Weibull density, which can always be represented with $\delta_{1}^* = \delta_{2}^*$. This restriction may reduce the test's power against alternatives with $\pi^*(1-\pi^*) \simeq 0$, such as discrete mixture 2 in Section \ref{sec:MC}. Our simulations indicate that this loss is small: under that design, the EM test and the LRT have similarly low power.

Define the log-likelihood function under the null hypothesis as
\begin{equation} \label{loglik0}
L_{0n}(\delta, \boldsymbol{\beta}, \gamma) \coloneq \sum_{t=1}^n \log f^{c}(Y_t^{c},D_t | \boldsymbol{X}_t;\delta,\boldsymbol{\beta},\gamma),
\end{equation}
where $f^{c}(y^{c},d | \boldsymbol{x};\delta,\boldsymbol{\beta},\gamma)$ is defined in \eqref{eq:model1}. Define the maximum likelihood estimator (MLE) of the parameters under the null hypothesis as
\[
(\widehat \delta, \widehat{\boldsymbol{\beta}}, \widehat \gamma) \coloneq \mathop{\arg\max}_{(\delta,\boldsymbol{\beta},\gamma) \in D \times B \times \Gamma} L_{0n}(\delta, \boldsymbol{\beta}, \gamma).
\]
Define the log-likelihood function under the two-point discrete mixture model as
\begin{equation} \label{loglik}
L_{n}(\pi, \delta_{1}, \delta_{2}, \boldsymbol{\beta}, \gamma) \coloneq \sum_{t=1}^n \log f_{a}^{c}(Y_t^{c},D_t | \boldsymbol{X}_t; \pi, \delta_{1}, \delta_{2}, \boldsymbol{\beta}, \gamma),
\end{equation}
with $f^{c}_{a}(y^{c},d | \boldsymbol{x};\pi,\delta_{1}, \delta_{2},\boldsymbol{\beta},\gamma)$ defined in \eqref{eq:model_mixture}. Define the penalized log-likelihood function under the two-point discrete mixture model as
\begin{equation} \label{pen_loglik}
PL_{n}(\pi, \delta_{1}, \delta_{2}, \boldsymbol{\beta}, \gamma) \coloneq L_{n}(\pi, \delta_{1}, \delta_{2}, \boldsymbol{\beta}, \gamma) + p(\pi),
\end{equation}
where $p(\pi)$ is a penalty function on $\pi$. We specify $p(\pi)$ in Section \ref{sec:MC}. Its purpose is to penalize values of $\pi$ that are too close to zero or one and to ensure that the model stays away from regions of the parameter space where identification is weak or lost.

Let $\boldsymbol{\theta} \coloneq (\delta_1, \delta_2, \boldsymbol{\beta}^{\top}, \gamma)^{\top} \in \Theta\coloneq D^2 \times B \times \Gamma$ and $\boldsymbol{\theta}^* \coloneq (\delta^*, \delta^*, (\boldsymbol{\beta}^*)^{\top}, \gamma^*)^{\top}$. Then, the parameter of model \eqref{eq:model_mixture} is given by $(\pi,\boldsymbol{\theta})$, and the penalized log-likelihood function is written as $PL_{n}(\pi,\boldsymbol{\theta})$. Let $\Pi$ be a finite subset of $(0,0.5]$. For each $\pi_0 \in \Pi$, define $\pi^{(1)}(\pi_0)=\pi_0$, and define $\boldsymbol{\theta}^{(1)}(\pi_0)$ as
\[
\boldsymbol{\theta}^{(1)}(\pi_0) \coloneq \mathop{\arg\max}_{\boldsymbol{\theta} \in \Theta} PL_{n}(\pi^{(1)}(\pi_0), \boldsymbol{\theta}).
\]
Starting from $(\pi^{(1)}(\pi_0),\boldsymbol{\theta}^{(1)}(\pi_0))$, we iteratively update $\pi$ and $\boldsymbol{\theta}$ using EM steps as in \citet{lcm09bm}. Suppose that $\pi^{(k)}(\pi_0)$ and $\boldsymbol{\theta}^{(k)}(\pi_0)$ have been calculated. For $t=1,\ldots,n$, define the weights for the E-step as
\begin{align*}
w_{1t}^{(k)}(\pi_0) &\coloneq \frac{\pi^{(k)}f^c(Y_t^{c}, D_t| \boldsymbol{X}_t; \delta_1^{(k)}(\pi_0), \boldsymbol{\beta}^{(k)}(\pi_0), \gamma^{(k)}(\pi_0) )}{f_{a}^c(Y_t^{c},D_t| \boldsymbol{X}_t; \pi^{(k)}(\pi_0), \boldsymbol{\theta}^{(k)}(\pi_0))}, \quad w_{2t}^{(k)}(\pi_0) \coloneq 1- w_{1t}^{(k)}(\pi_0).
\end{align*}
In the M-step, update $\pi$ and $\boldsymbol{\theta}$ as follows
\begin{align*}
\pi^{(k+1)}(\pi_0) &\coloneq \mathop{\arg \max}_{\pi \in [0,1]} \left\{ \sum_{t=1}^n w_{1t}^{(k)} \log(\pi) + \sum_{t=1}^n w_{2t}^{(k)} \log (1-\pi) + p(\pi) \right\},
\end{align*}
where we suppress the dependence of $w_{1t}^{(k)}(\pi_0)$ on $\pi_0$ for notational simplicity. The update for $\boldsymbol{\theta}$ is given by
\begin{align*}
\boldsymbol{\theta}^{(k+1)}(\pi_0) &\coloneq \mathop{\arg \max}_{\boldsymbol{\theta}\in \Theta} \left\{ \sum_{t=1}^n \sum_{j=1}^{2} w_{jt}^{(k)} \log f^c(Y_t^{c},D_t| \boldsymbol{X}_t; \delta_j, \boldsymbol{\beta}, \gamma) \right\}.
\end{align*}
Note that an EM step never decreases the penalized log-likelihood. For a pre-specified integer $K \geq 1$, define the test statistic for $\pi_0 \in \Pi$ after $K-1$ EM steps as
\[
M_n^{K}(\pi_0) \coloneq 2\left\{PL_{n}(\pi^{(K)}(\pi_0),\boldsymbol{\theta}^{(K)}(\pi_0)) - L_{0n}(\widehat \delta, \widehat{\boldsymbol{\beta}}, \widehat \gamma) \right\},
\]
where $L_{0n}(\widehat{\delta}, \widehat{\boldsymbol{\beta}}, \widehat{\gamma})$ is the maximized log-likelihood under the null.

The \textit{EM test statistic} is defined as the maximum of $M_n^{K}(\pi_0)$ over $\pi_0\in \Pi$,
\[
\text{EM}_n^{K} \coloneq \max \left\{M_n^{K}(\pi_0): \pi_0 \in \Pi \right\}.
\]
The following proposition gives the asymptotic null distribution of the EM test statistic. Assumptions \ref{assn1}--\ref{assn3} are presented in Appendix \ref{app:assumptions} and correspond to Assumptions A1--A3 in \citet{chowhite10joe} except that we explicitly state assumptions on $C_t$. In addition, our \Cref{assn_3.6} strengthens their Assumption A3.6 to be consistent with Assumption A.5(iii) in \citet{chowhite07em}. These assumptions allow for dependent data and are applicable to autoregressive conditional duration models.
\begin{proposition} \label{EM_stat}
Suppose that Assumptions \ref{assn1}--\ref{assn3} hold and $\delta_1^*=\delta_2^*$. Suppose that $p(\pi)$ is nonpositive and continuous, $p(0.5)=0$, and $p(\pi) \to -\infty$ as $\pi \to 0$ or $\pi \to 1$. Suppose further that $0.5 \in \Pi$. Then, for any fixed finite $K$,
\[
\text{EM}_{n}^{K} \rightarrow_d (\max\{0,N(0,1)\})^2.
\]
\end{proposition}

\subsection{LRT for unobserved heterogeneity}\label{sec:LRT}

\citet{chowhite10joe} derive the asymptotic null distribution of the LRT statistic under fixed and random censoring. This section reviews their results. Let $A = [\alpha_L,\alpha_H]$, where $0 < \alpha_L \leq 1 \leq \alpha_H < \infty$, denote the set of admissible values for $\alpha_1=\delta_1/\delta^*$ and $\alpha_2=\delta_2/\delta^*$.

Define the MLE of the two-point discrete mixture model as
\[
(\widehat \pi, \widehat \delta_1, \widehat \delta_2, \widehat{\boldsymbol{\beta}}_{a}, \widehat \gamma_{a}) \coloneq \mathop{\arg\max}_{(\pi, \delta_1,\delta_2,\boldsymbol{\beta},\gamma) \in [0,1] \times \widehat{A} \times \widehat{A} \times B \times \Gamma} L_{n} (\pi, \delta_{1}, \delta_{2}, \boldsymbol{\beta}, \gamma) ,
\]
where $L_{n} (\pi, \delta_{1}, \delta_{2}, \boldsymbol{\beta}, \gamma)$ is defined in \eqref{loglik}, and $\widehat{A} \coloneq [\alpha_L \widehat\delta, \alpha_H \widehat\delta]$. The LRT statistic for testing $H_0$ is
\[
LR_n \coloneq 2 \left[ L_{n} (\widehat \pi, \widehat \delta_1, \widehat \delta_2, \widehat{\boldsymbol{\beta}}_{a}, \widehat \gamma_{a}) - L_{0n}(\widehat \delta, \widehat{\boldsymbol{\beta}}, \widehat \gamma) \right].
\]

\citet[][p.\ 465]{chowhite10joe} show that, if $\inf A > 1/2$, the asymptotic null distribution of $LR_n$ is given by
\begin{equation} \label{eq:cho_white_g}
\sup_{\alpha \in A}\left( \max[0,\mathcal{G}^c(\alpha)] \right)^2,
\end{equation}
where $\mathcal{G}^c(\alpha)$ is a Gaussian process with covariance kernel $E[\mathcal{G}^c(\alpha)\mathcal{G}^c(\alpha')]$. As noted by \citet[][p.\ 465]{chowhite10joe}, this covariance depends on the joint distribution of $(\boldsymbol{X}_t, C_t)$ and $(\alpha,\alpha')$. Due to the difficulty of consistently estimating $E[\mathcal{G}^c(\alpha)\mathcal{G}^c(\alpha')]$ across all grid points $(\alpha,\alpha')$, \citet{chowhite10joe} omit direct analysis of the covariance kernel and instead use the weighted bootstrap procedure of \citet{hansen96em} to obtain critical values. A parametric bootstrap would require generating censored samples under the null. This is feasible under fixed censoring and random censoring. Under covariate-dependent censoring, however, it would require the conditional distribution $G(\cdot|\boldsymbol{x})$, which the model leaves unspecified and whose estimation would require a nonparametric first step that smooths over $\boldsymbol{x}$. The weighted bootstrap avoids this step by reweighting the scores computed from the observed sample rather than generating new observations.

\section{Monte Carlo experiments}\label{sec:MC}

In this section, we use Monte Carlo simulations to compare the size and power properties of the EM test, the LRT, the IM test of \citet{Chesher1984} and \citet{Lancaster1984}, and the LM tests of \citet{Sharma1987}.

\subsection{Covariate-independent censoring}
We first consider the setting in which the censoring process is independent of the covariates. Specifically, we assume that, conditional on the covariate $\boldsymbol{X}_t$ and the heterogeneity parameter $\delta$, the uncensored duration $Y_t$ follows a Weibull distribution and the censoring variable $C_t$ follows an exponential distribution. Their respective density functions are given by
\begin{equation}\label{dgp_independent}
\begin{aligned}
	f(y | \boldsymbol{x}; \delta, \boldsymbol{\beta}, \gamma) &=
	\begin{cases} \delta \gamma \exp(\boldsymbol{x}^{\top}\boldsymbol{\beta}) y^{\gamma - 1} \exp \left(-\delta \exp(\boldsymbol{x}^{\top}\boldsymbol{\beta})y^{\gamma} \right) & y > 0, \\
	0 & y \leq 0,
	\end{cases} \\
	g(c|\boldsymbol{x}; \lambda) &=
	\begin{cases}
	\lambda \exp(-\lambda c) & c \geq 0, \\
	0 & c < 0.
	\end{cases}
\end{aligned}
\end{equation}
We assume that $C_t$ is independent of $(Y_t, \boldsymbol{X}_t)$, and that $\delta$ is i.i.d.\ across observations and independent of $(\boldsymbol{X}_t, C_t)$. The covariate $\boldsymbol{X}_t$ is one-dimensional and drawn from the standard normal distribution. The parameters $(\boldsymbol{\beta}, \gamma)$ are unknown. The parameter $\delta$ captures unobserved heterogeneity: under the null hypothesis of homogeneity, $\delta$ is constant, whereas under the alternative, $\delta$ follows a non-degenerate distribution, so that the distribution of $Y_t$ given $\boldsymbol{X}_t$ is a mixture of Weibull distributions. We specify the distribution of $\delta$ under the alternative below. For each sample size $n$, we generate $(Y_t, C_t, \boldsymbol{X}_t)$ from this process and conduct the tests on the observed data $\{(Y^c_t, D_t, \boldsymbol{X}_t)\}_{t=1}^n$.

We describe the implementation details of the EM test, the LRT, the IM test, and the LM tests. In the EM test, we set $\Pi = \{ 0.05, 0.1, 0.15, 0.2, 0.25, 0.3, 0.35, 0.4, 0.45, 0.5 \}$ and use the penalty function $p(\pi) = 10 \log (1 - |1 - 2 \pi|)$, following the functional form suggested by \citet{lcm09bm}. We consider $K = 1, 2, 3$, corresponding to zero, one, and two EM steps.

The LRT is implemented following the procedure described in Section \ref{sec:LRT}. For the specification of the set of admissible values $A$, we adopt three options from the simulation setup in \citet{chowhite10joe}: a narrower range $[7/9, 2]$, a moderate range $[2/3, 3]$, and a wider range $[5/9, 4]$. Each of these intervals contains 1 and satisfies $\inf A >1/2$. As the asymptotic null distribution in equation \eqref{eq:cho_white_g} is intractable, we obtain critical values using the weighted bootstrap procedure of \citet{hansen96em}. First, we discretize $A$ using equally spaced grid points $0.01$ apart. For each grid point $\alpha \in A$, we compute the score $\widehat S_{nt} (\alpha) \coloneq \{ \widehat D_{nt} (\alpha) \}^{-1/2} \widehat W_{nt} (\alpha)$, where
\begin{align*}
\widehat D_{nt} (\alpha) \coloneq \ &\frac{1}{n} \sum_{t = 1}^n (1 - \widehat R_{nt} (\alpha))^2 - \frac{1}{n} \sum_{t = 1}^n (1 - \widehat R_{nt} (\alpha)) \widehat{\boldsymbol{U}}_{nt}^{\top} \left( \frac{1}{n} \sum_{t = 1}^n \widehat{\boldsymbol{U}}_{nt} \widehat{\boldsymbol{U}}_{nt}^{\top} \right)^{-1} \\
	&\times \frac{1}{n} \sum_{t = 1}^n \widehat{\boldsymbol{U}}_{nt} (1 - \widehat R_{nt} (\alpha)), \\
	\widehat W_{nt} (\alpha) \coloneq \ &(1 - \widehat R_{nt} (\alpha)) - \widehat{\boldsymbol{U}}_{nt}^{\top} \left( \frac{1}{n} \sum_{t = 1}^n \widehat{\boldsymbol{U}}_{nt} \widehat{\boldsymbol{U}}_{nt}^{\top} \right)^{-1} \frac{1}{n} \sum_{t = 1}^n \widehat{\boldsymbol{U}}_{nt} (1 - \widehat R_{nt} (\alpha)), \\
	\widehat R_{nt} (\alpha) \coloneq \ &f^{c} (Y_t^c, D_t| \boldsymbol{X}_t; \alpha \widehat \delta, \widehat{\boldsymbol{\beta}}, \widehat \gamma) / f^{c} (Y_t^c, D_t| \boldsymbol{X}_t; \widehat \delta, \widehat{\boldsymbol{\beta}}, \widehat \gamma), \\
	\widehat{\boldsymbol{U}}_{nt} \coloneq \ &\nabla_{(\delta,\boldsymbol{\beta}, \gamma)^{\top}} \log f^{c} (Y_t^c, D_t| \boldsymbol{X}_t; \widehat \delta, \widehat{\boldsymbol{\beta}}, \widehat \gamma),
\end{align*}
and $f^c$ is defined in \eqref{eq:model1}. We then generate a sequence of i.i.d.\ standard normal random variables $\{ Z_{jt} : j=1,\ldots, J; t = 1, \ldots, n \}$ and simulate the asymptotic null distribution of $LR_n$ using the empirical distribution of
\[
\mathcal{LR}_{jn}(A) \coloneq \sup_{\alpha \in A} \left( \max \left[ 0, \frac{1}{\sqrt{n}} \sum_{t = 1}^n \widehat S_{nt} (\alpha) Z_{jt} \right] \right)^2, \quad j = 1, \dots, J.
\]
Following \citet{chowhite10joe}, we set the bootstrap sample size to $J=500$ in all simulations.

For the IM test, we follow \citet{chowhite10joe} and implement a version of the IM test statistic of \citet{Lancaster1984} that focuses on the diagonal element associated with $\delta$. The IM test statistic is defined as $IM_n \coloneq \boldsymbol{\iota}^{\top} \boldsymbol{W}(\boldsymbol{W}^{\top} \boldsymbol{W})^{-1}\boldsymbol{W}^{\top} \boldsymbol{\iota}$, where $\boldsymbol{\iota}$ is an $n \times 1$ vector of ones, and $\boldsymbol{W}$ is defined as
\begin{equation}
	\boldsymbol{W} \coloneq
	\begin{pmatrix}
		\widehat{\boldsymbol{S}}_1 ^{\top} & \widehat H_1 \\
		\vdots & \vdots \\
		\widehat{\boldsymbol{S}}_n ^{\top} & \widehat H_n
	\end{pmatrix}, \notag
\end{equation}
with $\widehat{\boldsymbol{S}}_t \coloneq \nabla_{(\delta, \boldsymbol{\beta}, \gamma)^{\top}} \log f^{c} (Y_t^c, D_t| \boldsymbol{X}_t; \widehat \delta, \widehat{\boldsymbol{\beta}}, \widehat \gamma)$ and
\[
\widehat H_{t} \coloneq \left[ \frac{\partial}{\partial \delta} \log f^{c} (Y_t^c, D_t| \boldsymbol{X}_t; \widehat \delta, \widehat{\boldsymbol{\beta}}, \widehat \gamma) \right]^2 + \frac{\partial^2}{\partial \delta^2} \log f^{c} (Y_t^c, D_t| \boldsymbol{X}_t; \widehat \delta, \widehat{\boldsymbol{\beta}}, \widehat \gamma).
\]
Under the null hypothesis, $IM_n$ converges in distribution to $\chi^2_1$.

For the LM tests, we follow \citet{chowhite10joe} and implement a version of the LM test statistic of \citet{Sharma1987} that accounts for parameter estimation error, as noted by \citet{Prieger00wp}. For $t = 1, \dots, n$ and $q=2,3$, define
\[
\widehat{\boldsymbol{L}}_{q, t} \coloneq \left[ L_2 \left(\{ \widehat \delta \exp( \boldsymbol{X}_t^{\top} \widehat{\boldsymbol{\beta}})\}^{1/\widehat \gamma}Y_t^c; D_t, \widehat \gamma \right), \dots, L_q \left(\{ \widehat \delta \exp( \boldsymbol{X}_t^{\top} \widehat{\boldsymbol{\beta}})\}^{1/\widehat \gamma}Y_t^c; D_t, \widehat \gamma \right) \right]^{\top},
\]
where the subscript $q$ denotes the order of the Laguerre polynomials, and the censoring-adjusted Laguerre polynomials are defined as
\begin{align*}
	L_2(x; D, \gamma) &\coloneq \frac{1}{2} (x^{2 \gamma} - 4 x^{\gamma} + 2) + (1 - D) (x^{\gamma} - 1), \\
	L_3(x; D, \gamma) &\coloneq \frac{1}{6} (- x^{3 \gamma} + 9 x^{2 \gamma} - 18 x^{\gamma} + 6) + \frac{1}{2} (1 - D) (-x^{2 \gamma} + 4 x^{\gamma} - 2).
\end{align*}
Define $\overline{\boldsymbol{L}}_{q, n} \coloneq n^{-1} \sum_{t = 1}^n \widehat{\boldsymbol{L}}_{q, t}$, $\widehat{\boldsymbol{W}}_{q, n} \coloneq \widehat{\boldsymbol{V}}_{q, n} - \widehat{\boldsymbol{H}}_{q, n} \widehat{\boldsymbol{B}}_n^{-1} \widehat{\boldsymbol{H}}_{q, n}^{\top}$, where $\widehat{\boldsymbol{V}}_{q, n} \coloneq n^{-1} \sum_{t = 1}^n (\widehat{\boldsymbol{L}}_{q,t} - \overline{\boldsymbol{L}}_{q, n}) (\widehat{\boldsymbol{L}}_{q, t} - \overline{\boldsymbol{L}}_{q, n})^{\top}$, $\widehat{\boldsymbol{H}}_{q, n} \coloneq n^{-1} \sum_{t = 1}^n (\widehat{\boldsymbol{L}}_{q, t} - \overline{\boldsymbol{L}}_{q, n}) \widehat{\boldsymbol{S}}_{t}^{\top}$, and $\widehat{\boldsymbol{B}}_n \coloneq n^{-1} \sum_{t = 1}^n \widehat{\boldsymbol{S}}_t\widehat{\boldsymbol{S}}_t^{\top}$. Define the LM test statistics as $LM_{q,n} \coloneq n \overline{\boldsymbol{L}}_{q, n}^{\top} \{ \widehat{\boldsymbol{W}}_{q, n} \}^{-1} \overline{\boldsymbol{L}}_{q, n}$. Under the null hypothesis, $LM_{q,n}$ converges in distribution to $\chi_{q - 1}^2$. For $q = 2$, because $\widehat{\boldsymbol{L}}_{2,t} = (\widehat\delta^2/2)\widehat H_t + \widehat\delta\, \widehat S_{\delta,t}$, where $\widehat S_{\delta,t}$ is the first element of $\widehat{\boldsymbol{S}}_t$, and $\sum_{t=1}^n \widehat{\boldsymbol{S}}_t = \boldsymbol{0}$ at the null MLE, we have $LM_{2,n} = IM_n/(1 - IM_n/n)$. Consequently, the IM and LM$_2$ tests have identical size-adjusted power.

Table \ref{table level independent} reports the empirical rejection rates of the tests under the null hypothesis. The parameters of the data-generating process \eqref{dgp_independent} are set to $(\delta, \boldsymbol{\beta}, \gamma, \lambda) = (1,1,1,1)$, so that about half of the observations are censored, and the nominal significance level is fixed at 5\%. The EM, IM, and LM tests use asymptotic critical values, while the LRT employs bootstrapped critical values.

The EM test exhibits mild size distortion in small samples: its rejection rates are 11\% to 12\% at $n = 50$ and about 8\% at $n = 100$, regardless of $K$. The empirical size quickly approaches the nominal level as the sample size increases. For $n \geq 500$, rejection rates are close to 5\%. In contrast, the LRT is conservative across all sample sizes, particularly when the admissible set $A$ is narrow (e.g., $[7/9, 2]$). Although rejection rates increase as the range of $A$ widens, they remain slightly below the nominal level even at $n = 5000$. The IM and LM tests exhibit size distortions, particularly in small and moderate samples. For instance, the IM test rejects the null in 16.18\% of replications at $n = 50$ and 12.40\% at $n = 100$, while LM$_2$ performs similarly. LM$_3$ severely over-rejects the null even in large samples.

Table \ref{table level robust} examines how the size of the EM test reported in Table \ref{table level independent} depends on the censoring fraction and the covariate distribution. We set the censoring fraction to 25\%, 50\%, or 75\% by changing the parameter $\lambda$ in the distribution of the censoring variable $C_t$ in \eqref{dgp_independent} and draw $\boldsymbol{X}_t$ from either the standard normal distribution or the $\chi^2_1$ distribution standardized to have mean zero and unit variance. For $n \geq 500$, the rejection rates lie between 4.3\% and 5.7\% in every design, confirming that the EM test controls size well at these sample sizes. At $n = 50$, the size distortion increases with the censoring fraction.

To evaluate the power properties of the tests, we consider the following six alternative distributions for the unobserved heterogeneity parameter $\delta$ in \eqref{dgp_independent}:
\begin{enumerate}
    \item Discrete mixture 1: $\delta \sim \text{DM}(0.7370, 1.9296;\ 0.5)$
    \item Discrete mixture 2: $\delta \sim \text{DM}(0.7370, 2.4222;\ 0.95)$
    \item Gamma mixture: $\delta \sim \text{Gamma}(5,\ 5)$
    \item Log-normal mixture: $\delta \sim \text{Log-normal}(-\log(1.2)/2,\ \log(1.2))$
    \item Uniform mixture 1: $\delta \sim \text{Uniform}[0.30053,\ 2.3661]$
    \item Uniform mixture 2: $\delta \sim \text{Uniform}[1,\ 5/3]$.
\end{enumerate}
Here, $\text{DM}(a, b; p)$ denotes a two-point discrete distribution with $\Pr(\delta = a) = p$ and $\Pr(\delta = b) = 1 - p$. Discrete mixture 2 is designed to assess whether the EM test loses power when the mixing proportion is close to zero or one; the other five specifications follow \citet{chowhite10joe}. For the other parameters in \eqref{dgp_independent}, we set $(\boldsymbol{\beta}, \gamma, \lambda) = (1,1,1)$.

Tables \ref{table power independent discrete 1}--\ref{table power independent uniform 2} show the power properties of the tests.
For the comparison of the EM test and the LRT, we focus on the size-unadjusted power for sample sizes of 500 or more, where the EM test controls the level. This comparison is the practically relevant one, because size adjustment requires knowledge of the data-generating process and is therefore infeasible in practice. In this comparison, the EM test generally performs better than the LRT, although the difference is small when the power of both tests is close to 100\%. The difference is also small under discrete mixture 2 and uniform mixture 2, where no test has appreciable power, and under discrete mixture 2 the LRT rejects slightly more often than the EM test at larger sample sizes, particularly with the wider sets $A$. The IM and LM$_2$ tests have lower power than the EM test in nearly all cases, while the high rejection rates of LM$_3$ reflect its large size distortion.

With size adjustment, the EM test and the LRT have higher power than the IM and LM tests in every design for $n \geq 500$. The size-adjusted comparison is the one more favorable to the LRT: because the EM test is already correctly sized at these sample sizes, the adjustment leaves its power essentially unchanged, whereas it raises the power of the conservative LRT substantially.
The power of the EM test for $K=2$ and $K=3$ is very similar to that for $K=1$; for brevity, Tables \ref{table power independent discrete 1}--\ref{table power independent uniform 2} report results only for $K=1$.

\subsection{Covariate-dependent censoring}\label{sec:MC_dependent}
We next examine the performance of the tests in the setting where the censoring process depends on covariates.
We specify the conditional density of $Y$ given $\boldsymbol{X}$ and $\delta$ and the censoring time $C$ as
\begin{align}
	f(y | \boldsymbol{x}; \delta, \boldsymbol{\beta}, \gamma) &= \delta \gamma \exp(\boldsymbol{x}^{\top}\boldsymbol{\beta})y^{\gamma - 1} \exp(- \delta \exp(\boldsymbol{x}^{\top}\boldsymbol{\beta}) y^{\gamma}) \quad \text{for} \ y > 0, \notag \\
	C &= 2.5 \exp(-\eta_1 X_1^2 + \eta_2 X_2), \label{dgp dependent}
\end{align}
where $\boldsymbol{X} \coloneq (X_1, X_2)^{\top}$, and $X_1$ and $X_2$ are mutually independent covariates following the chi-squared distribution with one degree of freedom and the standard normal distribution, respectively. Because $C$ is a deterministic function of the covariates, the design is covered by model \eqref{eq:model1} as the special case in which the conditional distribution of $C$ given $\boldsymbol{X}=\boldsymbol{x}$ is degenerate, so that $\nu_{\boldsymbol{x}}$ is counting measure on a single point and $g(\cdot|\boldsymbol{x})$ equals one there.
The specification of $\delta$, the observed data, and the implementation of the tests are the same as in the independent-censoring case, except that $\boldsymbol{X}_t$ is now two-dimensional.

Unlike the IM and LM tests, whose $\chi^2$ asymptotic null distributions follow from Assumptions \ref{assn1}--\ref{assn3} regardless of whether censoring depends on $\boldsymbol{X}_t$, the LRT's asymptotic null distribution has been formally established by \citet{chowhite10joe} only under fixed or random censoring.
Its validity under covariate-dependent censoring is, at present, an open question. We nonetheless compute the LRT's critical values using the same weighted bootstrap procedure as in the covariate-independent case above. The LRT results reported below therefore apply the procedure outside the setting covered by its asymptotic theory.

Table \ref{table level dependent} reports the empirical rejection rates of the tests under the null hypothesis.
We set $\delta = 1$, $\beta_1 = -0.3$, $\beta_2 = 1$, $\gamma = 1$, $\eta_1 = 0.5$, and $\eta_2 = 1$, so that about 42\% of the observations are censored.
The nominal significance level is fixed at 5\%.
The results are similar to those under independent censoring: the EM test has size close to 5\% for $n \geq 500$, the LRT is conservative, and the IM and LM tests over-reject.

To assess power, we consider the same six specifications for $\delta$ as in the independent-censoring setting. The value of $(\beta_1, \beta_2, \gamma, \eta_1, \eta_2)$ is the same as in the null setting.
Tables \ref{table power dependent discrete 1}--\ref{table power dependent uniform 2} show the power properties of the tests.
As in the independent-censoring case, the IM and LM$_2$ tests lag behind the EM test and the LRT in nearly all cases, and the higher rejection rates of LM$_3$ reflect its size distortion.
The comparison between the EM test and the LRT also follows the independent-censoring pattern: for $n \geq 500$, the EM test generally has higher size-unadjusted power than the LRT. As in the independent-censoring case, the power for $K=2$ and $K=3$ is very similar to that for $K=1$, so Tables \ref{table power dependent discrete 1}--\ref{table power dependent uniform 2} report results only for $K=1$.

\section{Real-world data analysis}\label{sec:realdata}
In this section, we compare the proposed EM test with the other tests by analyzing real-world data from the Stanford Heart Transplant Program.
The dataset is taken from \citet{kf2011book} and is also available as ``jasa'' in the R package \textbf{survival}. The dataset records the possibly censored survival times (in days) of 103 patients along with their censoring status, which takes the value one if the patient died (uncensored) and zero if the patient is censored.
The dataset also records patients' characteristics, of which we use three: the age (in years) at the time of acceptance, the year of acceptance into the program, and an indicator of prior heart surgery. These are the baseline covariates used by \citet{kf2011book} in their analysis of these data.
We do not use the patients' transplant status. A patient is eligible for a transplant from acceptance onward and joins the treatment group only when a donor heart becomes available, so transplant status is realized after acceptance. \citet{kf2011book} accordingly model it as a time-dependent covariate, which lies outside the model of Section \ref{sec:Weibull_model}.
If transplantation alters mortality, patients with the same covariates but different transplant histories have different survival distributions, which can by itself produce a rejection of the homogeneous Weibull model.

Censoring in this dataset is administrative. Patients entered the program over a period of several years, and those still alive when follow-up ended were censored.
A patient's potential censoring time is thus the time from acceptance to the common end of follow-up, which varies with the date of acceptance.
Because 26 of the 28 censored patients share the same end-of-follow-up date, the correlation between the potential censoring time and the year of acceptance is $-0.988$.
The program also changed over this period: \citet{kf2011book} find the year of acceptance to be a significant predictor of survival in these data and attribute it to a relaxation of the admission requirements.
The tests require that $Y_t$ and $C_t$ be conditionally independent given $\boldsymbol{X}_t$. Because the potential censoring time is determined by the acceptance date, this amounts to requiring that survival be independent of the acceptance date given the covariates. We return to this requirement after introducing the covariate specifications. When the covariates include the year of acceptance, the application falls in the covariate-dependent censoring regime of Section \ref{sec:MC_dependent}, in which the EM test remains valid under Assumptions \ref{assn1}--\ref{assn3} while the asymptotic validity of the LRT has not been established; the LRT results reported below should be read with this caveat in mind.

Using this dataset, we test the specified homogeneous Weibull regression against a scale-mixture alternative.
As introduced in Section \ref{sec:Weibull_model}, the tests are based on the Weibull density function of the survival time
\begin{equation}
	f(y | \boldsymbol{x}; \delta, \boldsymbol{\beta}, \gamma) = \delta \gamma \exp(\boldsymbol{x}^{\top}\boldsymbol{\beta})y^{\gamma - 1} \exp(-\delta \exp(\boldsymbol{x}^{\top}\boldsymbol{\beta})y^{\gamma}), \notag
\end{equation}
where, for the covariate vector $\boldsymbol{X}$, we consider the following five specifications
\begin{align}
	&\text{(I)} \ (X_{\mathrm{age}}), \quad
	\text{(II)} \ (X_{\mathrm{age}}, X_{\mathrm{age}}^2), \quad
	\text{(III)} \ (X_{\mathrm{age}}, X_{\mathrm{yr}}), \quad
	\text{(IV)} \ (X_{\mathrm{age}}, X_{\mathrm{age}}^2, X_{\mathrm{yr}}), \notag \\
	&\text{(V)} \ (X_{\mathrm{age}}, X_{\mathrm{age}}^2, X_{\mathrm{yr}}, X_{\mathrm{srg}}), \label{specification}
\end{align}
where $X_{\mathrm{age}}$ is the age at the time of acceptance, $X_{\mathrm{yr}}$ is the year of acceptance minus 1967, and $X_{\mathrm{srg}}$ indicates prior heart surgery. Specification (V) uses the full set of baseline covariates of \citet{kf2011book}. Only 16 of the 103 patients had prior surgery and only 9 of those 16 died, so the coefficient on $X_{\mathrm{srg}}$ is imprecisely estimated; \citet{kf2011book} report it as significant only at the 10\% level and caution that it rests on these 16 cases. For specifications (III)--(V), which include the year of acceptance, conditional independence of survival and censoring requires that the acceptance date within a year carry no further information about survival given the other covariates. We maintain this assumption. We also assume that the two patients censored before the common end of follow-up, for reasons the data do not record, were censored independently of their survival given the covariates. Specifications (I) and (II), by contrast, omit the year of acceptance. Because the year of acceptance predicts survival given age and also determines the censoring time, survival and censoring are unlikely to be conditionally independent given age alone: in a Weibull regression, the coefficient on the year of acceptance is significant at the 5\% level given age ($p = 0.04$) and at the 10\% level given age and its square ($p = 0.09$). We report specifications (I) and (II) for comparison, but they lie outside the maintained assumption, and their results should be interpreted with caution.

For the EM test, we use the same $\Pi$ and $p(\pi)$ as in Section \ref{sec:MC} and again consider $K = 1, 2, 3$. The other tests for comparison are the LRT, the IM test, and the LM tests, implemented using the same procedure as in Section \ref{sec:MC}.

Table \ref{real-world data analysis} summarizes the results of the analysis.
The upper panel reports $p$-values computed from the asymptotic null distribution of each test statistic.
Because the sample size of this dataset is $n = 103$ and Table \ref{table level dependent} shows that none of the tests has accurate size at $n = 100$, the lower panel reports $p$-values computed from the empirical distribution of each test statistic under the null setting with $n = 100$ in that simulation.
These simulation-based $p$-values are not a valid finite-sample correction for the present data. Although the asymptotic null distributions of the EM, IM, and LM test statistics do not depend on the data-generating process, their finite-sample distributions do, as the differences between Tables \ref{table level independent} and \ref{table level dependent} show. We therefore report the lower panel as a sensitivity analysis, which asks whether each test's conclusion survives the finite-sample null distribution observed in a related design.
We do not report simulation-based $p$-values for the LRT because, unlike those of the other tests, the asymptotic null distribution of the LRT statistic depends on the joint distribution of $(\boldsymbol{X}_t, C_t)$, and the critical values cannot be carried over from the simulation design in Section \ref{sec:MC_dependent} to this dataset.

In the upper panel, the EM test rejects the null hypothesis of homogeneity at the 1\% level in all five specifications.
The LRT fails to reject at the 5\% level for the narrowest admissible set $A$ in every specification.
Widening $A$ recovers rejection in specifications (I), (III), (IV), and (V), but in specification (II) all three choices of $A$ fail to reject.
In many cases, the LRT has a smaller $p$-value when $A$ is wide, which is consistent with Tables \ref{table power dependent discrete 1}--\ref{table power dependent uniform 2}, where the LRT has higher power with a wider $A$ than with a narrower one.
The EM test exhibits no comparable sensitivity to the choice of $K$.
The IM and LM$_2$ tests reject the null hypothesis in four of the five specifications, failing only in specification (II). LM$_3$ rejects in all five.

In the lower panel, the EM test continues to reject the null hypothesis at the 5\% level in all five specifications, and at the 1\% level in all specifications except (II), where its simulation-based $p$-value is $0.013$.
The IM and LM tests behave very differently.
Their simulation-based $p$-values are considerably larger than their asymptotic ones, and all three fail to reject the null hypothesis at the 5\% level except for specification (III).
This pattern is consistent with the substantial over-rejection of the IM and LM tests at $n = 100$ in Table \ref{table level dependent}, but the lower panel does not establish that their rejections in specifications (I), (IV), and (V), and that of LM$_3$ in (II), reflect size distortion for the present data. The IM and LM$_2$ tests have identical simulation-based $p$-values because $LM_{2,n} = IM_n/(1 - IM_n/n)$.

The same comparison is not available for the LRT. Table \ref{table level dependent} nonetheless indicates that the LRT is conservative at this sample size, consistent with its failure to reject the null in many cases.
Taken together, the two panels show that the EM test is the only test whose rejection of the null hypothesis of homogeneity survives both this sensitivity analysis and the choice of specification.
The EM test rejects at the 5\% level in every specification in both panels, whereas the conclusions of the IM and LM tests change between the two panels, and those of the LRT strongly depend on the choice of $A$ and of the specification.

\section{Conclusion}\label{sec:conclusion}

Unobserved heterogeneity poses a serious problem in duration analysis: ignoring it can bias the estimated structural parameters of the hazard function and invalidate inference. In censored duration models, the existing tests for unobserved heterogeneity offer an unsatisfactory remedy, as the IM and LM tests suffer from severe size distortion and the LRT requires a computationally demanding bootstrap for inference.

This paper develops an EM test for unobserved heterogeneity in censored Weibull duration models. The test builds on the EM approach of \citet{lcm09bm}, and its statistic converges in distribution to the square of $\max\{0, N(0,1)\}$ under the null hypothesis, so that critical values are obtained directly from the standard normal distribution without simulation or bootstrap. Unlike the LRT of \citet{chowhite10joe}, the EM test accommodates covariate-dependent censoring of arbitrary form without additional adjustment.

Monte Carlo simulations show that the EM test controls size well for sample sizes of 500 or more, under both covariate-independent and covariate-dependent censoring and for censoring fractions from 25\% to 75\%. Its size-adjusted power is comparable to that of the LRT and higher than that of the IM and LM tests. Size adjustment is not available in practice, however, and the LRT is markedly conservative in finite samples, so the EM test attains higher power than the LRT in most designs when the two are applied as a practitioner would apply them.

An application to survival data from the Stanford Heart Transplant Program illustrates the same pattern. The EM test rejects the null hypothesis of homogeneity in all five covariate specifications, whereas the LRT's conclusion depends on the choice of the admissible set in many cases. The rejections by the IM and LM tests, in contrast, generally do not survive when $p$-values are computed from the finite-sample null distributions in the simulations.

The test maintains the assumption that the duration and the censoring time are conditionally independent given the covariates. Under endogenous censoring, the duration distribution is not point identified without further structure. \citet{KhanTamer09joe}, \citet{Szydlowski19jae}, and \citet{Sakaguchi24jae} use instruments or report identified sets; extending the EM test in that direction is left for future work.

\section*{Funding}

This work was supported by JSPS KAKENHI Grant Number JP24K04814.

\section*{Declaration of generative AI and AI-assisted technologies in the manuscript preparation process}

During the preparation of this work the authors used Claude (Anthropic) and Codex (OpenAI) in order to proofread and copyedit the manuscript text, correct internal cross-references and notation, and check the simulation code for errors. After using these tools, the authors reviewed and edited the content as needed and take full responsibility for the content of the published article.

\bibliography{mixture_duration.bib}