EconBase
← Back to paper

Surrogate-powered Causal Inference with Censored Outcomes

The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.

80,414 characters

Surrogate-powered Causal Inference with Censored Outcomes


\maketitle
\begin{spacing}{1.15}
\begin{abstract}
Clinical trials with survival endpoints lose information when participants are censored before death is observed.
We develop target-preserving estimators that use posttreatment disease history, such as recurrence or progression, to recover information lost to censoring for marginal survival and restricted mean survival effects.
The difficulty is that the intermediate event is downstream of treatment: naive adjustment can change the causal estimand, and the useful information enters only through the observed coarsening.
We derive observed-data influence functions with and without recurrence history and obtain an exact gain identity.
The identity shows that efficiency improvement is driven by the censoring hazard, the split of the alive risk set into recurrence states, and the residual-survival separation between those states.
In the no-covariate illness--death model, the Aalen--Johansen estimator realizes the recurrence-augmented efficient score after standardization to the marginal target.
With covariates, correctly specified Cox--Breslow transition hazards provide a root-\(N\) plug-in benchmark, while a hazard-induced one-step estimator gives rate robustness and, under primitive transition and censoring learner rates, canonical inference with a second-order product remainder.
A semi-synthetic metastatic breast cancer study calibrated from digitized progression-free survival and overall survival curves illustrates the gain identity.
The framework applies broadly to censored time-to-event studies with informative intermediate histories.
\end{abstract}

\noindent{\it Keywords:} censoring; intermediate event history; semiparametric efficiency; marginal survival; restricted mean survival.
\newpage

\section{Introduction}\label{sec:introduction}

Cancer clinical trials are expensive because survival outcomes take time.
A day of an oncology trial has been estimated to cost an average of \$0.54 million.
Yet reviews report that 30--40\% of enrolled patients are censored before their survival outcome is observed
\citep{smith2024new,rosen2020censored,vitale2026censoring,moore2018estimated}.
A censored patient is not uninformative: we know that the patient survived up to the censoring time.
But the trial does not observe the terminal event that the study was designed to compare.
This loss of information matters because the sample size required to detect a treatment effect scales with the efficiency bound for the target estimand.
More information per enrolled patient means more power, shorter studies, and faster evidence.
The same statistical problem appears well beyond oncology, whenever causal effects are defined through censored event times, including public-health interventions, labor-market spells, startup lifetimes, and other duration outcomes
\citep{hernan00,honore2006bounds,honore2018new,deaner2024causal,hellmann02,alvarez2024decomposing}.

Causal inference with a survival endpoint is a compound missing-data problem.
Under the potential-outcomes framework, each unit has potential survival times $T^0,T^1$, but treatment assignment reveals at most one of them
\citep{neyman23,rubin05,rubin76,ding18}.
Survival data add a second coarsening: even the factual event time may be hidden by right censoring.
If $C$ denotes the censoring time, the observed data contain $\widetilde T=T\wedge C$ and $\Delta=I(T\le C)$ rather than necessarily observing $T$ itself
\citep{andersen12}.
Thus a survival trial has two statistical tasks.
The first is identification: the analyst needs assumptions on treatment assignment and censoring that connect the observed law to the marginal potential-outcome estimand.
The second is precision: after identification, the analyst wants to recover as much of the information lost to censoring as possible.
This paper is primarily about the second task, while also showing how richer disease history can make the censoring condition more plausible.

The guiding principle is simple: auxiliary data improve efficiency only when they recover information hidden by sampling or censoring.
If $T_i^a$ is fully observed, measuring auxiliary variables on each unit does not improve first-order efficiency for estimating a functional of the marginal distribution $P_{T^a}$.
This holds regardless of the strength of its association with $T^a_i$.
The outcome itself is already the relevant contribution to the marginal parameter.
Gains become possible when some contributions are missing or coarsened.
Then auxiliary variables improve precision by refining the conditional mean of the missing residual contribution.
This information-recovery view underlies augmented estimators for missing and censored data
\citep{robins92,robins94,robins95,murray1996nonparametric}.

Covariate adjustment is the familiar application of this idea.
To estimate an arm-specific parameter $\theta^a$, units assigned to the opposite arm have $T_i^a$ missing.
But their baseline covariates $X_i$ are observed.
Augmenting the arm-specific estimator with $X_i$ can recover part of the missing counterfactual information and improve precision
\citep{hahn98,robins95,li2023estimating}.
This is safe for total-effect estimands because $X_i$ is measured before treatment and is therefore not a mediator.
The same $X$-based information helps estimate both $\theta^0$ and $\theta^1$.
In contrast, the observed outcome under the opposite treatment, $T_i^{1-a}$, cannot be used as ordinary auxiliary information for $\theta^a$: it is never observed jointly with $T_i^a$, and using it would effectively change the problem into estimating a treatment contrast.
The limitation in survival studies is that baseline covariates are frozen at time zero.
Long-run survival often depends on post-randomization disease events, leaving substantial conditional variation in $T^a$ even after rich baseline adjustment.

This motivates using a posttreatment intermediate event time $R^a$.
In oncology, $T^a$ and $R^a$ may be time to death and time to progression.
Knowing whether progression has occurred before a censoring time can strongly change the conditional residual-survival probability for the event $\{T^a>\tau\}$.
Our goal is to use this later disease-state information to recover power lost to censoring while preserving the marginal survival estimand.
The target remains, for example,

\begingroup
\setstretch{1}
\setlength{\jot}{0pt}
\begin{align}\label{eq:parameters}
  \theta^a_\tau=P(T^a>\tau),
  \qquad
  \mu^a_\tau=E(T^a\wedge \tau)=\int_0^\tau P(T^a>t)\,dt .
\end{align}
\endgroup
The intermediate event is not promoted to the endpoint.
It is used only as observed-data information about the terminal event.

A posttreatment variable cannot be used like a baseline covariate without changing the causal question.
The counterfactual recurrence time $R^{1-A}$ is missing by definition, and the factual recurrence time is itself observed only until death or censoring.
Thus recurrence cannot be used to borrow information across treatment arms in the same way as pretreatment $X$.
Moreover, including recurrence as a time-dependent covariate in a treatment-effect Cox model generally changes the estimand: since recurrence is affected by treatment, conditioning on it targets a conditional or direct-effect-like quantity rather than the marginal total effect on survival.
This target drift is entangled with the noncollapsibility and limited causal interpretation of hazard-ratio estimands
\citep{hernan10,luo25}.
Our construction avoids this problem.
It uses recurrence armwise, inside the observed-data score, and then standardizes back to the marginal survival or restricted-mean estimand.

There is also an identification opportunity.
Causal survival analysis requires an assumption on the censoring mechanism.
If dropout depends on latent survival beyond the observed history used in the analysis, a marginal treatment-effect estimator can be inconsistent.
Recording posttreatment disease events expands the observed history: instead of requiring censoring to be independent of latent death given only baseline variables and being alive, one can condition on the current recurrence state.
This formulation makes explicit which disease information must be observed for the censoring condition to be plausible, while allowing for the possibility of unmeasured dropout. In oncology, where censoring bias can distort estimated treatment effects, this distinction is practically important
\citep{rosen2020censored,vitale2026censoring}.
We focus on efficiency, but the same representation clarifies how recurrence history can relax the censoring assumption.

Our contribution is threefold.
First, we show that recurrence history has no automatic efficiency value for marginal survival inference: without coarsening or model restrictions, it cannot improve first-order efficiency.
Second, we show that right censoring changes this conclusion.
We derive the observed-data influence functions (IFs) with and without recurrence history and obtain an exact efficiency-gain identity.
The gain is large when censoring hazard is high, recurrence splits the alive risk set into well-populated states, those states have different conditional residual-survival means, and sufficient subjects remain alive at the relevant times.
Third, we develop estimators that attain the resulting efficiency bounds.
In the no-covariate illness--death model, the Aalen--Johansen (AJ) estimator realizes the recurrence-augmented efficient score after standardization to the marginal survival or restricted mean survival target.
With covariates, correctly specified Cox--Breslow transition hazards give a root-$n$ plug-in benchmark, while our one-step residual-augmented estimator accommodates flexible conditional-hazard learners.
The one-step correction has two roles: it gives a rate-robust mean when censoring is correctly learned, and it yields canonical root-$N$ inference when the transition and censoring errors satisfy the usual product-rate condition.


\subsection{Related literature}\label{subsec:related-literature}

Semiparametric missing-data theory shows how inverse weighting, augmentation, and observed-data projection recover information lost to coarsening
\citep{robins92,robins94,robins95}.
Our setting is different from a routine augmented inverse-probability-weighting problem.
The auxiliary information is a stopped posttreatment illness--death history: recurrence lies downstream of treatment, can be precluded by death, and is observed only until censoring.
The main step is therefore to project the full-data survival score $I(T>\tau)-\theta$ onto the observed filtration generated by the censored transition processes.
The task is distinct from augmenting with a fully observed covariate regression.
This projection yields an inverse-censoring-weighted martingale integral; recurrence changes the efficient influence function by replacing the alive-only residual-survival mean with the state-specific means.
The resulting gain is therefore a structural consequence of censoring, state occupancy, and conditional-mean separation, not a generic augmentation term.

Our work is distinct from surrogate endpoint validation.
Classical surrogate endpoint methods ask whether an intermediate endpoint can validate, replace, or stand in for the clinical endpoint
\citep{prentice1989surrogate,fleming96,buyse2000validation}.
Surrogate-index and limited-outcome methods use short-term outcomes to learn long-term effects under surrogacy, comparability, or missing-primary-outcome assumptions \citep{athey25,kallus25}.
We do not make recurrence the endpoint and do not estimate a surrogate-index effect.
The target remains marginal survival or restricted mean survival.
Recurrence enters only armwise, inside the observed-data score, and is then integrated out.
The goal is target-preserving information recovery, not surrogate validation.

The distinction from covariate adjustment is equally important.
Baseline covariates can improve precision because they are observed before treatment and can be used in both arms without changing the causal estimand
\citep{hahn98,robins95}.
Semi-supervised learning, prediction-powered inference and its refinements have the same target-preservation logic: auxiliary predictions are useful only after a correction restores validity for the original target
\citep{zhang2022high,angelopoulos2023ppi++,zrnic24crossppi}.
Our correction takes the form of a martingale residual adjustment induced by the fitted illness--death transition law.
This construction differs from calibration of black-box predictions on unlabeled data and is essential because recurrence is a posttreatment variable.
Naively conditioning on it in a treatment-effect model can change a marginal total-effect question into a conditional or direct-effect-like question.
Furthermore, in survival settings, hazard ratios are noncollapsible and sensitive to conditioning choices
\citep{hernan10,martinussen2022causality,luo25}.

\section{Problem setup}\label{sec:problem-setup}

Let $A\in\{0,1\}$ be treatment and let $X$ be baseline covariates.
For arm $a$, let $(T^a,R^a,C^a)$ denote the potential death time, recurrence time, and censoring time, following the potential-outcomes notation of \citet{neyman23} and \citet{rubin05}.
The semi-competing-risk convention is $R^a<T^a$ if recurrence occurs before death, while $R^a=\tau_0$ denotes no recurrence before death or the study horizon.
The factual variables are $(T,R,C)=(T^A,R^A,C^A)$, and the observed coarsening is

\begingroup
\setstretch{1}
\setlength{\jot}{0pt}
\begin{align} \label{eq:C-coarsening-map}
  (\tilde T,\Delta)&=(T\wedge C,\,I(T\le C)),
  \qquad
  (\tilde R,\Gamma)=(R\wedge \tilde T,\,I(R\le \tilde T)).
\end{align}
\endgroup
For each unit, the observed data are represented by the six-tuple $ O=(\tilde T,\Delta,\tilde R,\Gamma,A,X) $.
The theory below is developed armwise; treatment effects are obtained by applying the same construction in each randomized or ignorable treatment arm and taking contrasts.

\begin{assumption} \label{ass:time-reg}
For each arm $a$, $(T^a,R^a,C^a)$ is supported on $[0,\tau_0]^3$ and is jointly absolutely continuous on $[0,\tau_0)^3$.
The death time $T^a$ is continuous, so $P(T^a=\tau_0)=0$, while point masses at $\tau_0$ are allowed for $R^a$ and $C^a$.
\end{assumption}

For an event time $U$, write $S_U(t)=P(U>t)$, $F_U=1-S_U$, and $\lambda_U=f_U/S_U$ when a density exists.
We write $S\equiv S_T$, $G\equiv S_C$, $\lambda\equiv\lambda_T$, and $\lambda_{01}\equiv\lambda_R$ in a fixed arm.
Throughout, $\theta$ and $\mu_\tau$ denote a generic single-arm survival probability and restricted mean survival time (RMST), \eqref{eq:parameters}, $0<\tau<\tau_0$.
Censoring coarsens the endpoint without eliminating all prognostic information carried by the subject.
At the censoring time, the observed history still contains both the partial death history and the intermediate disease state.
The latter is inferentially relevant only through its effect on the conditional residual-survival mean at that time.
Hence recurrence can improve efficiency only when censoring occurs with nonnegligible probability in regions where the current disease state predicts subsequent survival.
We formalize this censoring-supported information-recovery principle and characterize when recurrence recovers information about the survival endpoint lost to censoring.

\begin{assumption} \label{ass:A-ignore}
Consistency holds, $(T,R,C)=(T^A,R^A,C^A)$ almost surely.
Moreover, $A\mathrel{\text{{\rotatebox[origin=c]{90}{\resizebox{2.00ex}{1.55ex}{$\vDash$}}}}}(T^0,T^1,R^0,R^1,C^0,C^1)\mid X$, with $e\le P(A=1\mid X)\le1-e$ a.s. for some $e>0$.
\end{assumption}

\begin{assumption} \label{ass:C-ignore}
For each arm $a$, $C^a\mathrel{\text{{\rotatebox[origin=c]{90}{\resizebox{2.00ex}{1.55ex}{$\vDash$}}}}} (T^a,R^a)\mid X$.
\end{assumption}

Together with Assumptions~\ref{ass:time-reg}  -- \ref{ass:A-ignore},
Assumption \ref{ass:C-ignore} identifies the recurrence and death transition hazards from the coarsened observations,
as $S^a(\tau) =E\exp\left\{-\int_0^\tau \lambda^a(t\mid X)dt\right\}$, and $ \lambda^a(t\mid X)dt =P\{t\le \tilde T<t+dt\mid A=a,\,\tilde T\ge t,\,X\} $.
Identification reduces the causal problem to an armwise censored survival problem.




\section{Surrogate-powered efficiency} \label{sec:efficiency}


Under an unrestricted joint law, $R$ provides no first-order information for a marginal functional of fully observed $T$.

\begin{theorem}\label{thm:no-gain}
Suppose $(T,R)$ is fully observed, $P_{R\mid T}$ is unrestricted,
and $\psi(P)=\psi(P_T)$ is a scalar pathwise differentiable parameter with marginal IF $\varphi_T$.
Then
$ \varphi_{\mathrm{joint}}(T,R)=\varphi_T(T). $
\end{theorem}
With censoring, $R$ can improve efficiency by refining the projection of the unobserved residual-survival contribution onto the history available at censoring.
The semiparametric definitions and proof are deferred to Appendix~\ref{app:full-data-efficiency};
see also the general efficiency frameworks of
\citep{ibragimov1981statistical,newey1990semiparametric,vandervaart98}.
The theorem separates the full-data and coarsened-data problems: any efficiency gain derived below arises from coarsening, not from imposing a surrogate-validity restriction on the joint law.

\subsection{Efficient influence functions}


Fix a treatment arm and suppress $X$.
We derive observed-data scores using death history alone and using recurrence history as well.
Recurrence changes the information about subsequent death without itself being the endpoint.
We therefore track whether a patient is alive before recurrence, alive after recurrence, or dead through the illness--death process

\begingroup
\setstretch{1}
\setlength{\jot}{0pt}
\begin{align} \label{eq:S-state}
  \mathcal S(t) =0\,I(t<R\wedge T)+1\,I(R\le t<T)+2\,I(T\le t), \qquad \mathcal S(0)=0
\end{align}
\endgroup
\citep{aalen78,andersen12}.
The state space is ordered by the transitions $0\to1$, $0\to2$, and $1\to2$, where state $0$ denotes alive and recurrence-free, state $1$ denotes alive after recurrence, and state $2$ denotes death.
Thus the transition $0\to1$ updates the observed alive-state history, whereas the transitions $0\to2$ and $1\to2$ determine the death endpoint.
Censoring acts as a coarsening mechanism on this event history: it stops observation of the current alive state rather than defining an additional clinical state.

Let $\mathcal J=\{01,02,12\}$.
For $\ell=(j,k)\in\mathcal J$, write $j(\ell)=j$ and $k(\ell)=k$.
In the full data, the alive risk set at time $t$ is partitioned into those who have not yet recurred, $Y_0^*(t)=I(t\le R,\,t\le T)$, and those who have recurred, $Y_1^*(t)=I(R<t\le T)$.
Hence $Y_T^*(t)=Y_0^*(t)+Y_1^*(t)$.
The full-data transition counts distinguish recurrence before death,
death before recurrence, and death after recurrence, respectively:
$N_{01}^*(t) =I(R\le t,\,R<T)$,
$N_{02}^*(t) =I(T\le t,\,T<R)$,
$N_{12}^*(t) =I(R<T\le t)$.
The full-data recurrence risk indicator and event count are
$Y_R^*(t)=Y_0^*(t)$ and $N_R^*(t)=N_{01}^*(t)$.
Death can occur from either alive state, so
$N_T^*(t)=N_{02}^*(t)+N_{12}^*(t)$.
Finally, $N_C^*(t)=I(C\le t)$ counts censoring.

For $0\le t\le \tau$, define the augmented and marginal event-history filtrations

\begingroup
\setstretch{1}
\setlength{\jot}{0pt}
\begin{align} \label{eq:F-filtrations}
  \mathcal F^{
  \mathord{\mathchoice
    {\vcenter{\hbox{\scalebox{0.8}{$\displaystyle\oldblacklozenge$}}}}
    {\vcenter{\hbox{\scalebox{0.8}{$\textstyle\oldblacklozenge$}}}}
    {\vcenter{\hbox{\scalebox{0.8}{$\scriptstyle\oldblacklozenge$}}}}
    {\vcenter{\hbox{\scalebox{0.8}{$\scriptscriptstyle\oldblacklozenge$}}}}
  }
}(t)
  =
  \sigma\{N_T^*(s),N_R^*(s):s\le t\},
  \qquad
  \mathcal F^\bullet(t)
  =
  \sigma\{N_T^*(s):s\le t\}.
\end{align}
\endgroup
The filtration $\mathcal F^{
  \mathord{\mathchoice
    {\vcenter{\hbox{\scalebox{0.8}{$\displaystyle\oldblacklozenge$}}}}
    {\vcenter{\hbox{\scalebox{0.8}{$\textstyle\oldblacklozenge$}}}}
    {\vcenter{\hbox{\scalebox{0.8}{$\scriptstyle\oldblacklozenge$}}}}
    {\vcenter{\hbox{\scalebox{0.8}{$\scriptscriptstyle\oldblacklozenge$}}}}
  }
}(t)$ records both the recurrence and death histories, whereas $\mathcal F^\bullet(t)$ records only the death history.
Under Assumption~\ref{ass:C-ignore}, adding censoring history to the conditioning information leaves the residual-survival means below unchanged.
For $\theta=P(T>\tau)$, recall the full-data IF $\phi^*(T)=I(T>\tau)-\theta$.
The recurrence information relevant for the marginal endpoint is the projection contrast

\begingroup
\setstretch{1}
\setlength{\jot}{0pt}
\begin{align} \label{eq:IF-contrast}
  \Delta_\theta(t)
  =
  M_\theta^{
  \mathord{\mathchoice
    {\vcenter{\hbox{\scalebox{0.8}{$\displaystyle\oldblacklozenge$}}}}
    {\vcenter{\hbox{\scalebox{0.8}{$\textstyle\oldblacklozenge$}}}}
    {\vcenter{\hbox{\scalebox{0.8}{$\scriptstyle\oldblacklozenge$}}}}
    {\vcenter{\hbox{\scalebox{0.8}{$\scriptscriptstyle\oldblacklozenge$}}}}
  }
}(t)-M_\theta^\bullet(t),
  \quad   M_\theta^{
  \mathord{\mathchoice
    {\vcenter{\hbox{\scalebox{0.8}{$\displaystyle\oldblacklozenge$}}}}
    {\vcenter{\hbox{\scalebox{0.8}{$\textstyle\oldblacklozenge$}}}}
    {\vcenter{\hbox{\scalebox{0.8}{$\scriptstyle\oldblacklozenge$}}}}
    {\vcenter{\hbox{\scalebox{0.8}{$\scriptscriptstyle\oldblacklozenge$}}}}
  }
}(t)
  =
  E\{\phi^*(T)\mid \mathcal F^{
  \mathord{\mathchoice
    {\vcenter{\hbox{\scalebox{0.8}{$\displaystyle\oldblacklozenge$}}}}
    {\vcenter{\hbox{\scalebox{0.8}{$\textstyle\oldblacklozenge$}}}}
    {\vcenter{\hbox{\scalebox{0.8}{$\scriptstyle\oldblacklozenge$}}}}
    {\vcenter{\hbox{\scalebox{0.8}{$\scriptscriptstyle\oldblacklozenge$}}}}
  }
}(t)\}, \  M_\theta^\bullet(t)
  =
  E\{\phi^*(T)\mid \mathcal F^\bullet(t)\}.
\end{align}
\endgroup
The contrast vanishes when recurrence adds no information about residual survival beyond death history.
Under censoring, any efficiency gain depends on its value at censoring times.

\begin{assumption}  \label{ass:Markov-state}
The current state suffices for residual survival:
 for   $0\le t\le \tau$ and $j=0,1$,
$
  E \{I(T>\tau)\mid \mathcal F^{
  \mathord{\mathchoice
    {\vcenter{\hbox{\scalebox{0.8}{$\displaystyle\oldblacklozenge$}}}}
    {\vcenter{\hbox{\scalebox{0.8}{$\textstyle\oldblacklozenge$}}}}
    {\vcenter{\hbox{\scalebox{0.8}{$\scriptstyle\oldblacklozenge$}}}}
    {\vcenter{\hbox{\scalebox{0.8}{$\scriptscriptstyle\oldblacklozenge$}}}}
  }
}(t) \}
  =
  v_0(t)Y_0^*(t+)+v_1(t)Y_1^*(t+), $
with
$ v_j(t)=P\{T>\tau\mid \mathcal S(t)=j\}.
$
\end{assumption}


Under Assumption~\ref{ass:Markov-state}, recurrence affects residual
survival through the current state. Suppose the transitions
$0\to1$, $0\to2$, and $1\to2$ have state-dependent hazards
$\lambda_{01}$, $\lambda_{02}$, and $\lambda_{12}$, respectively.
For illness--death transitions $jk\in \mathcal J$, define the $\mathcal F^{
  \mathord{\mathchoice
    {\vcenter{\hbox{\scalebox{0.8}{$\displaystyle\oldblacklozenge$}}}}
    {\vcenter{\hbox{\scalebox{0.8}{$\textstyle\oldblacklozenge$}}}}
    {\vcenter{\hbox{\scalebox{0.8}{$\scriptstyle\oldblacklozenge$}}}}
    {\vcenter{\hbox{\scalebox{0.8}{$\scriptscriptstyle\oldblacklozenge$}}}}
  }
}$-martingales $ M^{
  \mathord{\mathchoice
    {\vcenter{\hbox{\scalebox{0.8}{$\displaystyle\oldblacklozenge$}}}}
    {\vcenter{\hbox{\scalebox{0.8}{$\textstyle\oldblacklozenge$}}}}
    {\vcenter{\hbox{\scalebox{0.8}{$\scriptstyle\oldblacklozenge$}}}}
    {\vcenter{\hbox{\scalebox{0.8}{$\scriptscriptstyle\oldblacklozenge$}}}}
  }
}_{jk}(t) =N^*_{jk}(t)-\int_0^tY_j^*(u)\lambda_{jk}(u)\,du. $
The death count $N_T^*=N^*_{02}+N^*_{12}$ has
$\mathcal F^\bullet$-martingale $ M_T^\bullet(t) =N_T^*(t)-\int_0^tY_T^*(u)\lambda_T(u)\,du, $
where $\lambda_T$ is its marginal death hazard.
The state-specific residual-survival probabilities satisfy

\begingroup
\setstretch{1}
\setlength{\jot}{0pt}
\begin{align*}
  \frac{d}{dt}v_1(t)=\lambda_{12}(t)v_1(t),
 \qquad \mbox{and} \qquad
  \frac{d}{dt}v_0(t)
  =
  \lambda_{01}(t)\{v_0(t)-v_1(t)\}
  +
  \lambda_{02}(t)v_0(t),
\end{align*}
\endgroup
with the terminal conditions $v_0(\tau)=v_1(\tau)=1$ and $v_2(t)=0$.
The transition hazards are otherwise unrestricted.
Under Assumption~\ref{ass:Markov-state}, and with $ v(t)=P\{T>\tau\mid T>t\} $,

\begingroup
\setstretch{1}
\setlength{\jot}{0pt}
\begin{equation}
  \label{eq:M-blacklozenge}
  M_\theta^{
  \mathord{\mathchoice
    {\vcenter{\hbox{\scalebox{0.8}{$\displaystyle\oldblacklozenge$}}}}
    {\vcenter{\hbox{\scalebox{0.8}{$\textstyle\oldblacklozenge$}}}}
    {\vcenter{\hbox{\scalebox{0.8}{$\scriptstyle\oldblacklozenge$}}}}
    {\vcenter{\hbox{\scalebox{0.8}{$\scriptscriptstyle\oldblacklozenge$}}}}
  }
}(t)
 =
  \sum_{jk\in\mathcal J}\int_0^t\{v_k(s)-v_j(s)\}\,
  dM_{jk}^{
  \mathord{\mathchoice
    {\vcenter{\hbox{\scalebox{0.8}{$\displaystyle\oldblacklozenge$}}}}
    {\vcenter{\hbox{\scalebox{0.8}{$\textstyle\oldblacklozenge$}}}}
    {\vcenter{\hbox{\scalebox{0.8}{$\scriptstyle\oldblacklozenge$}}}}
    {\vcenter{\hbox{\scalebox{0.8}{$\scriptscriptstyle\oldblacklozenge$}}}}
  }
}(s)
  =
  \sum_{j=0}^1 v_j(t)Y_j^*(t+)-\theta,
\end{equation}
\begin{equation}
  \label{eq:M-bullet}
  M_\theta^\bullet(t)
  =
  \int_0^t\{0-v(s)\}\,dM_T^\bullet(s)
 =
  v(t)Y_T^*(t+)-\theta.
\end{equation}
\endgroup

The marginal analysis assigns $v(t)$ to all survivors,
whereas the recurrence-augmented analysis assigns $v_0(t)$ to recurrence-free survivors and $v_1(t)$ to post-recurrence survivors.
Their difference is

\begingroup
\setstretch{1}
\setlength{\jot}{0pt}
\begin{align}\label{eq:IF-contrast2}
\Delta_\theta(t)
=\{v_0(t)-v(t)\}Y_0^*(t+)
 +\{v_1(t)-v(t)\}Y_1^*(t+).
\end{align}
\endgroup
A gain is possible where both alive states have positive probability,
their residual-survival means differ, and censoring can occur.
The martingale property of $\Delta_\theta$ is not immediate from \eqref{eq:IF-contrast}--\eqref{eq:IF-contrast2}; the integral representations \eqref{eq:M-blacklozenge} and \eqref{eq:M-bullet} establish it.

The observed processes and filtrations are stopped at $\tilde T=T\wedge C$:

\begingroup
\setstretch{1}
\setlength{\jot}{0pt}
\begin{align}
  \label{eq:observed-sa}
  \mathcal F^{
  \mathord{\mathchoice
    {\vcenter{\hbox{\scalebox{0.8}{$\displaystyle\oldlozenge$}}}}
    {\vcenter{\hbox{\scalebox{0.8}{$\textstyle\oldlozenge$}}}}
    {\vcenter{\hbox{\scalebox{0.8}{$\scriptstyle\oldlozenge$}}}}
    {\vcenter{\hbox{\scalebox{0.8}{$\scriptscriptstyle\oldlozenge$}}}}
  }
}(t)=\mathcal F^{
  \mathord{\mathchoice
    {\vcenter{\hbox{\scalebox{0.8}{$\displaystyle\oldblacklozenge$}}}}
    {\vcenter{\hbox{\scalebox{0.8}{$\textstyle\oldblacklozenge$}}}}
    {\vcenter{\hbox{\scalebox{0.8}{$\scriptstyle\oldblacklozenge$}}}}
    {\vcenter{\hbox{\scalebox{0.8}{$\scriptscriptstyle\oldblacklozenge$}}}}
  }
}(t\wedge\tilde T)
  \qquad \text{and} \qquad
  \mathcal F^\circ(t)=\mathcal F^\bullet(t\wedge\tilde T) .
\end{align}
\endgroup
The observed at-risk and events processes are:
$ Y(t)=Y_T^*(t)Y_C^*(t), $
$ Y_j(t)=Y_j^*(t)Y_C^*(t),$
$ N_T(t)=N_T^*(t\wedge\tilde T), $
$ N_{jk}(t)=N_{jk}^*(t\wedge\tilde T). $
If death is observed, the endpoint is known.
If censoring occurs first, the observed-data IF inserts a conditional mean.
The $\mathcal F^\circ$-score inserts the alive-only conditional mean;
the $\mathcal F^{
  \mathord{\mathchoice
    {\vcenter{\hbox{\scalebox{0.8}{$\displaystyle\oldlozenge$}}}}
    {\vcenter{\hbox{\scalebox{0.8}{$\textstyle\oldlozenge$}}}}
    {\vcenter{\hbox{\scalebox{0.8}{$\scriptstyle\oldlozenge$}}}}
    {\vcenter{\hbox{\scalebox{0.8}{$\scriptscriptstyle\oldlozenge$}}}}
  }
}$-score inserts the state-specific conditional mean.

\begin{assumption}
\label{ass:eif-gain-primitives}
Assumptions~\ref{ass:time-reg},
\ref{ass:C-ignore}, and~\ref{ass:Markov-state} hold, $ \inf_{0\le t\le \tau}G(t)>0, $ and   $M_{jk}^{
  \mathord{\mathchoice
    {\vcenter{\hbox{\scalebox{0.8}{$\displaystyle\oldblacklozenge$}}}}
    {\vcenter{\hbox{\scalebox{0.8}{$\textstyle\oldblacklozenge$}}}}
    {\vcenter{\hbox{\scalebox{0.8}{$\scriptstyle\oldblacklozenge$}}}}
    {\vcenter{\hbox{\scalebox{0.8}{$\scriptscriptstyle\oldblacklozenge$}}}}
  }
}$, $jk\in\mathcal J$,  $M_T^\bullet$, and the censoring martingale $M_C$ are square integrable on $[0,\tau]$.
\end{assumption}

Under Assumptions~\ref{ass:time-reg} and~\ref{ass:C-ignore}, censoring is orthogonal to the event martingales.
In particular,
$
  \langle M_C,M_{jk}^{
  \mathord{\mathchoice
    {\vcenter{\hbox{\scalebox{0.8}{$\displaystyle\oldblacklozenge$}}}}
    {\vcenter{\hbox{\scalebox{0.8}{$\textstyle\oldblacklozenge$}}}}
    {\vcenter{\hbox{\scalebox{0.8}{$\scriptstyle\oldblacklozenge$}}}}
    {\vcenter{\hbox{\scalebox{0.8}{$\scriptscriptstyle\oldblacklozenge$}}}}
  }
}\rangle(t)=0,
  \  jk\in\mathcal J,
$
and
$
  \langle M_C,M_T^\bullet\rangle(t)=0,
  \  0\le t\le\tau .
$
This orthogonality is used below to separate censoring variation from recurrence and death variation in the observed-data efficient scores.

\begin{theorem}
\label{thm:eif-main}
Under Assumption~\ref{ass:eif-gain-primitives}, the influence functions for $\theta=P\{T>\tau\}$ are

\begingroup
\setstretch{1}
\setlength{\jot}{0pt}
\begin{align}
  \label{eq:IF-star}
  \phi^*(T)
  &=I(T>\tau)-\theta
  =\int_0^\tau dM^\bullet_\theta(t)
  =\int_0^\tau dM^{
  \mathord{\mathchoice
    {\vcenter{\hbox{\scalebox{0.8}{$\displaystyle\oldblacklozenge$}}}}
    {\vcenter{\hbox{\scalebox{0.8}{$\textstyle\oldblacklozenge$}}}}
    {\vcenter{\hbox{\scalebox{0.8}{$\scriptstyle\oldblacklozenge$}}}}
    {\vcenter{\hbox{\scalebox{0.8}{$\scriptscriptstyle\oldblacklozenge$}}}}
  }
}_\theta(t),
  \\
  \label{eq:IF-Delta}
  \phi^\Delta(\tilde T,\Delta)
  &=\frac{\Delta}{G(\tilde T)}\{\phi^*(\tilde T)-\phi_0\}+\phi_0,
  \\
  \label{eq:IF-circ}
  \phi^\circ(\tilde T,\Delta)
  &=\int_0^\tau \frac{Y_C^*(t)}{G(t)}\,dM^\bullet_\theta(t)
  =\frac{\Delta_{C,\tau}}{G(\tilde T\wedge\tau)}\phi^*(\tilde T)
    +\int_0^{\tilde T\wedge\tau}M^\bullet_\theta(t-)\frac{dM_C^\bullet(t)}{G(t)},
  \\
  \label{eq:IF-lorenge}
  \phi^{
  \mathord{\mathchoice
    {\vcenter{\hbox{\scalebox{0.8}{$\displaystyle\oldlozenge$}}}}
    {\vcenter{\hbox{\scalebox{0.8}{$\textstyle\oldlozenge$}}}}
    {\vcenter{\hbox{\scalebox{0.8}{$\scriptstyle\oldlozenge$}}}}
    {\vcenter{\hbox{\scalebox{0.8}{$\scriptscriptstyle\oldlozenge$}}}}
  }
}(\tilde T,\Delta,\tilde R,\Gamma)
  &=\int_0^\tau \frac{Y_C^*(t)}{G(t)}\,dM^{
  \mathord{\mathchoice
    {\vcenter{\hbox{\scalebox{0.8}{$\displaystyle\oldblacklozenge$}}}}
    {\vcenter{\hbox{\scalebox{0.8}{$\textstyle\oldblacklozenge$}}}}
    {\vcenter{\hbox{\scalebox{0.8}{$\scriptstyle\oldblacklozenge$}}}}
    {\vcenter{\hbox{\scalebox{0.8}{$\scriptscriptstyle\oldblacklozenge$}}}}
  }
}_\theta(t)
  =\frac{\Delta_{C,\tau}}{G(\tilde T\wedge\tau)}\phi^*(\tilde T)
    +\int_0^{\tilde T\wedge\tau}M^{
  \mathord{\mathchoice
    {\vcenter{\hbox{\scalebox{0.8}{$\displaystyle\oldblacklozenge$}}}}
    {\vcenter{\hbox{\scalebox{0.8}{$\textstyle\oldblacklozenge$}}}}
    {\vcenter{\hbox{\scalebox{0.8}{$\scriptstyle\oldblacklozenge$}}}}
    {\vcenter{\hbox{\scalebox{0.8}{$\scriptscriptstyle\oldblacklozenge$}}}}
  }
}_\theta(t-)\frac{dM_C^{
  \mathord{\mathchoice
    {\vcenter{\hbox{\scalebox{0.8}{$\displaystyle\oldblacklozenge$}}}}
    {\vcenter{\hbox{\scalebox{0.8}{$\textstyle\oldblacklozenge$}}}}
    {\vcenter{\hbox{\scalebox{0.8}{$\scriptstyle\oldblacklozenge$}}}}
    {\vcenter{\hbox{\scalebox{0.8}{$\scriptscriptstyle\oldblacklozenge$}}}}
  }
}(t)}{G(t)},
\end{align}
\endgroup
where $\Delta_{C,\tau}=I\{C>T\wedge\tau\}$ and $\phi_0=E\{\phi^*(T)\omega(T)\}$.
\end{theorem}

The superscripts $*,\Delta,\circ,{
  \mathord{\mathchoice
    {\vcenter{\hbox{\scalebox{0.8}{$\displaystyle\oldlozenge$}}}}
    {\vcenter{\hbox{\scalebox{0.8}{$\textstyle\oldlozenge$}}}}
    {\vcenter{\hbox{\scalebox{0.8}{$\scriptstyle\oldlozenge$}}}}
    {\vcenter{\hbox{\scalebox{0.8}{$\scriptscriptstyle\oldlozenge$}}}}
  }
}$ denote, respectively, the full, missing-outcome, marginal-censored, and recurrence-augmented-censored data.
The theorem separates three sources of variation.
In the full-data experiment, recurrence does not alter the IF for the marginal survival probability:
$ \phi^*(T) $ is obtained under both the marginal death filtration and the recurrence-augmented filtration.
Second, right censoring replaces the unavailable full-data contribution after the censoring time by a predictable projection on the observed history.
The marginal censored-data IF $\phi^\circ$ is the usual right-censoring IF written in martingale form.
Third, the recurrence-augmented IF $\phi^{
  \mathord{\mathchoice
    {\vcenter{\hbox{\scalebox{0.8}{$\displaystyle\oldlozenge$}}}}
    {\vcenter{\hbox{\scalebox{0.8}{$\textstyle\oldlozenge$}}}}
    {\vcenter{\hbox{\scalebox{0.8}{$\scriptstyle\oldlozenge$}}}}
    {\vcenter{\hbox{\scalebox{0.8}{$\scriptscriptstyle\oldlozenge$}}}}
  }
}$ has the same inverse-censoring-weighted complete contribution as $\phi^\circ$.
The only change is in the augmentation term: the marginal residual-survival process $M_\theta^\bullet(t-)$ is replaced by the finer projection $M_\theta^{
  \mathord{\mathchoice
    {\vcenter{\hbox{\scalebox{0.8}{$\displaystyle\oldblacklozenge$}}}}
    {\vcenter{\hbox{\scalebox{0.8}{$\textstyle\oldblacklozenge$}}}}
    {\vcenter{\hbox{\scalebox{0.8}{$\scriptstyle\oldblacklozenge$}}}}
    {\vcenter{\hbox{\scalebox{0.8}{$\scriptscriptstyle\oldblacklozenge$}}}}
  }
}(t-)$, which conditions on the current illness--death state.
Hence recurrence can improve efficiency only by improving the prediction of the missing residual contribution at censoring times.
In particular, there is no gain without censoring, and there is no gain when $ M_\theta^{
  \mathord{\mathchoice
    {\vcenter{\hbox{\scalebox{0.8}{$\displaystyle\oldblacklozenge$}}}}
    {\vcenter{\hbox{\scalebox{0.8}{$\textstyle\oldblacklozenge$}}}}
    {\vcenter{\hbox{\scalebox{0.8}{$\scriptstyle\oldblacklozenge$}}}}
    {\vcenter{\hbox{\scalebox{0.8}{$\scriptscriptstyle\oldblacklozenge$}}}}
  }
}(t)=M_\theta^\bullet(t) $ on the time region where censoring occurs.
Under the current-state formulation, this condition is equivalent to the absence of a relevant state-specific residual-survival contrast between recurrence-free and post-recurrence survivors.
We state and interpret that comparison in \Cref{subsec:gain-snapshot}.

\subsection{No-covariate AJ estimation}
\label{subsec:estimation}

The recurrence-augmented score suggests a natural estimator: estimate the transition law of the observed illness--death process and standardize the resulting state-occupation probabilities to the marginal survival or RMST target.
In the no-covariate model, this construction coincides with the classical AJ estimator
\citep{aalen78,andersen12}.
The estimator learns the transition dynamics $0\to1$, $0\to2$, and $1\to2$, and then averages over the induced paths to recover the marginal death-time functional.
It can therefore be viewed as the empirical counterpart of the recurrence-augmented efficient score.

Let $A_{jk}(t)=\int_0^t\lambda_{jk}(u)\,du$ and define the cumulative-hazard matrix and transition matrix

\begingroup
\setstretch{1}
\setlength{\jot}{0pt}
\begin{align} \label{eq:matrix-product-integral}
  d\mathbf A(t)=
  \displaystyle
  \begin{psmallmatrix}
  \displaystyle
  -dA_{01}(t)-dA_{02}(t) &
  \displaystyle
  dA_{01}(t) &
  \displaystyle
  dA_{02}(t)\\
  \displaystyle
  0 &
  \displaystyle
  -dA_{12}(t) &
  \displaystyle
  dA_{12}(t)\\
  \displaystyle
  0 &
  \displaystyle
  0 &
  \displaystyle
  0
  \end{psmallmatrix},
  \quad
  \mathbf P(s,t) = \prod_{(s,t]} \big\{I+d\mathbf A(u) \big\},
\end{align}
\endgroup
with all increments evaluated at $t$.
Since $\mathcal S(0)=0$, the state-occupation probabilities are $P_j(t)=P_{0j}(0,t)$, and the survival and RMST estimands are

\begingroup
\setstretch{1}
\setlength{\jot}{0pt}
\begin{align*}
  \theta(t)=P_0(t)+P_1(t),
  \qquad
  \mu_\tau=\int_0^\tau\theta(t)\,dt.
\end{align*}
\endgroup
For $u\le t$, write $v_0(u,t)=P_{00}(u,t)+P_{01}(u,t)$ and $v_1(u,t)=P_{11}(u,t)$ for survival to $t$ conditional on the state at $u$.
Their integrals $V_j(u,t)=\int_u^t v_j(u,s)\,ds$ give the corresponding expected survival time over $[u,t]$.




For $n$ subjects in a fixed treatment arm,
let $N_{jk,i}$ and $Y_{j,i}$ denote subject $i$'s observed transition count and risk indicator,
and write $\bar N_{jk}=n^{-1}\sum_{i=1}^nN_{jk,i}$ and $\bar Y_j=n^{-1}\sum_{i=1}^nY_{j,i}$.
For $jk\in\mathcal J$, the Nelson--Aalen (NA) estimator and the AJ transition matrix $(\widehat P_{jk}(s,t))$ are

\begingroup
\setstretch{1}
\setlength{\jot}{0pt}
\begin{align*}
  \widehat A_{jk}^{\textsf{AJ}}(t) =\int_0^t\frac{I\{\bar Y_j(u)>0\}}{\bar Y_j(u)} \,d\bar N_{jk}(u),
  \qquad
  \widehat{\mathbf P}(s,t) =\prod_{(s,t]}\{I+d\widehat{\mathbf A}^{\textsf{AJ}}(u)\},
\end{align*}
\endgroup
with the ratio defined as zero when $\bar Y_j(u)=0$.
Replacing $A_{jk}$ by $\widehat A_{jk}^{\textsf{AJ}}$ in $\mathbf A$ defines $\widehat{\mathbf A}^{\textsf{AJ}}$.
Since $\mathcal S(0)=0$, the AJ survival and RMST estimators are

\begingroup
\setstretch{1}
\setlength{\jot}{0pt}
\begin{equation}\label{eq:AJ}
\widehat\theta(t)
=\widehat P_{00}(0,t)+\widehat P_{01}(0,t),
\qquad
\widehat\mu_\tau^{\textsf{AJ}}
=\int_0^\tau\widehat\theta(t)\,dt.
\end{equation}
\endgroup

\begin{assumption}
\label{ass:aj-main-integrated}
Assumption~\ref{ass:eif-gain-primitives} holds.
For $\ell=jk\in\mathcal J$, $M_{\ell,i}(t)=N_{\ell,i}(t)-\int_0^tY_{j,i}(u)\,dA_{\ell}(u)$ is square integrable.
Write $q_j(u)=E\{Y_j(u)\}=P_j(u-)G(u-)$, $j=0,1$.
\begin{enumerate}[label=\textnormal{(AJ\arabic*)},leftmargin=2em]
\item\label{ajmain:support}
  With $\mathcal T_j=\{u\in(0,\tau]:q_j(u)>0\}$, the infimum of $q_j$ on every compact subset of $\mathcal T_j$ is positive,
  and $\int_0^\tau I\{q_{j(\ell)}(u)=0\}\,dA_\ell(u)=0$.
\item\label{ajmain:inverse}
  For each $\ell$,
\[
   V_\ell^\textsf{AJ}(\tau)= \int_{(0,\tau]}\frac{I\{q_{j(\ell)}(u)>0\}}{q_{j(\ell)}(u)}\,dA_\ell(u)<\infty.
\]
\end{enumerate}
\end{assumption}

These conditions impose no uniform lower bound on the state-specific risk probabilities;
in particular, $q_1$ may vanish near time zero because recurrence has not yet occurred.
Without uniform positivity, the NA proof requires additional control.
We localize to $\{q_j\ge\varepsilon\}$, where standard linearization applies, and bound the remaining martingale contribution through
$\int_{\{0<q_j(u)<\varepsilon\}}q_j(u)^{-1}\,dA_\ell(u)$,
$j=j(\ell)$.
This integral vanishes as $\varepsilon\downarrow0$ by \ref{ajmain:inverse},
which also ensures that the contribution from empty empirical risk sets is negligible.
Condition~\ref{ajmain:support} ensures that $dA_\ell$ assigns no mass to $\{q_{j(\ell)}=0\}$.

Integrating the recurrence-augmented survival EIF in \eqref{eq:IF-lorenge} over the time horizon gives the EIF for RMST,
denoted by $\phi_\mu^{
  \mathord{\mathchoice
    {\vcenter{\hbox{\scalebox{0.8}{$\displaystyle\oldlozenge$}}}}
    {\vcenter{\hbox{\scalebox{0.8}{$\textstyle\oldlozenge$}}}}
    {\vcenter{\hbox{\scalebox{0.8}{$\scriptstyle\oldlozenge$}}}}
    {\vcenter{\hbox{\scalebox{0.8}{$\scriptscriptstyle\oldlozenge$}}}}
  }
}$:

\begingroup
\setstretch{1}
\setlength{\jot}{0pt}
\begin{equation} \label{eq:main-nocov-rmst-eif}
  \phi_\mu^{
  \mathord{\mathchoice
    {\vcenter{\hbox{\scalebox{0.8}{$\displaystyle\oldlozenge$}}}}
    {\vcenter{\hbox{\scalebox{0.8}{$\textstyle\oldlozenge$}}}}
    {\vcenter{\hbox{\scalebox{0.8}{$\scriptstyle\oldlozenge$}}}}
    {\vcenter{\hbox{\scalebox{0.8}{$\scriptscriptstyle\oldlozenge$}}}}
  }
}(O)=\sum_{\ell\in\mathcal J}\int_0^\tau H_\ell^{
  \mathord{\mathchoice
    {\vcenter{\hbox{\scalebox{0.8}{$\displaystyle\oldlozenge$}}}}
    {\vcenter{\hbox{\scalebox{0.8}{$\textstyle\oldlozenge$}}}}
    {\vcenter{\hbox{\scalebox{0.8}{$\scriptstyle\oldlozenge$}}}}
    {\vcenter{\hbox{\scalebox{0.8}{$\scriptscriptstyle\oldlozenge$}}}}
  }
}(u)dM_\ell(u),
\end{equation}
where
\begin{equation} \label{eq:main-nocov-rmst-H}
  H_{01}^{
  \mathord{\mathchoice
    {\vcenter{\hbox{\scalebox{0.8}{$\displaystyle\oldlozenge$}}}}
    {\vcenter{\hbox{\scalebox{0.8}{$\textstyle\oldlozenge$}}}}
    {\vcenter{\hbox{\scalebox{0.8}{$\scriptstyle\oldlozenge$}}}}
    {\vcenter{\hbox{\scalebox{0.8}{$\scriptscriptstyle\oldlozenge$}}}}
  }
}(u)=\frac{V_1(u,\tau)-V_0(u,\tau)}{G(u-)},
  \quad
  H_{02}^{
  \mathord{\mathchoice
    {\vcenter{\hbox{\scalebox{0.8}{$\displaystyle\oldlozenge$}}}}
    {\vcenter{\hbox{\scalebox{0.8}{$\textstyle\oldlozenge$}}}}
    {\vcenter{\hbox{\scalebox{0.8}{$\scriptstyle\oldlozenge$}}}}
    {\vcenter{\hbox{\scalebox{0.8}{$\scriptscriptstyle\oldlozenge$}}}}
  }
}(u)=-\frac{V_0(u,\tau)}{G(u-)},
  \quad
  H_{12}^{
  \mathord{\mathchoice
    {\vcenter{\hbox{\scalebox{0.8}{$\displaystyle\oldlozenge$}}}}
    {\vcenter{\hbox{\scalebox{0.8}{$\textstyle\oldlozenge$}}}}
    {\vcenter{\hbox{\scalebox{0.8}{$\scriptstyle\oldlozenge$}}}}
    {\vcenter{\hbox{\scalebox{0.8}{$\scriptscriptstyle\oldlozenge$}}}}
  }
}(u)=-\frac{V_1(u,\tau)}{G(u-)}.
\end{equation}
\endgroup
Each integrand is the change in expected remaining-survival time at the corresponding transition,
importance-weighted by $G(u-)^{-1}$ to account for censoring \citep{robins94}.
Recurrence changes this expectation from $V_0(u,\tau)$ to
$V_1(u,\tau)$, whereas death from state $j$ reduces
$V_j(u,\tau)$ to zero.

\begin{theorem} \label{thm:aj-integrated-main}
Under Assumption~\ref{ass:aj-main-integrated}, the AJ RMST estimator satisfies $\widehat\mu^\textsf{AJ}_\tau-\mu_\tau = (\mathbb P_n-P)\phi_\mu^{
  \mathord{\mathchoice
    {\vcenter{\hbox{\scalebox{0.8}{$\displaystyle\oldlozenge$}}}}
    {\vcenter{\hbox{\scalebox{0.8}{$\textstyle\oldlozenge$}}}}
    {\vcenter{\hbox{\scalebox{0.8}{$\scriptstyle\oldlozenge$}}}}
    {\vcenter{\hbox{\scalebox{0.8}{$\scriptscriptstyle\oldlozenge$}}}}
  }
} + o_p(n^{-1/2})$.
If $0<\sigma_\mu^2=\mathrm{Var}\{\phi_\mu^{
  \mathord{\mathchoice
    {\vcenter{\hbox{\scalebox{0.8}{$\displaystyle\oldlozenge$}}}}
    {\vcenter{\hbox{\scalebox{0.8}{$\textstyle\oldlozenge$}}}}
    {\vcenter{\hbox{\scalebox{0.8}{$\scriptstyle\oldlozenge$}}}}
    {\vcenter{\hbox{\scalebox{0.8}{$\scriptscriptstyle\oldlozenge$}}}}
  }
}(O)\}<\infty$, then $ \sqrt n(\widehat\mu^\textsf{AJ}_\tau-\mu_\tau)\rightsquigarrow N(0,\sigma_\mu^2), $ and the AJ estimator attains the semiparametric efficiency bound.
\end{theorem}

Recurrence history changes the observed-data efficient score.
It replaces the marginal residual-survival mean by state-specific residual-survival means for recurrence-free and post-recurrence survivors.
Thus, in this model, the AJ estimator is the empirical implementation of the optimal observed-data projection.
The comparison with Kaplan--Meier (KM) is therefore a comparison between two projections: one that conditions only on the death history, and one that also uses the disease-state history at the censoring times where the efficiency bound depends on it.

\begin{remark}
The Markovian trajectories part of \Cref{ass:Markov-state} is a semiparametric restriction but is not the source of the gain described here.
The IFs of \Cref{thm:eif-main} and the gain identity \eqref{eq:gain} below do not require Markovian transitions.
With non-Markovian data, the AJ remains consistent  and more efficient than the KM, but is not semiparametrically efficient
\citep{datta2001validity, munch2023targeted}.
\end{remark}

\subsection{When is the gain largest?} \label{subsec:gain-snapshot}

We now return to the variance comparison.
Under Assumption~\ref{ass:eif-gain-primitives},

\begingroup
\setstretch{1}
\setlength{\jot}{0pt}
\begin{align}
  \label{eq:gain}
  \mathrm{Var}(\phi^\circ) - \mathrm{Var}(\phi^{
  \mathord{\mathchoice
    {\vcenter{\hbox{\scalebox{0.8}{$\displaystyle\oldlozenge$}}}}
    {\vcenter{\hbox{\scalebox{0.8}{$\textstyle\oldlozenge$}}}}
    {\vcenter{\hbox{\scalebox{0.8}{$\scriptstyle\oldlozenge$}}}}
    {\vcenter{\hbox{\scalebox{0.8}{$\scriptscriptstyle\oldlozenge$}}}}
  }
})
  =\int_0^\tau \frac{\lambda_C(t)S_T(t)}{G(t)}
  \mathrm{Var}\big\{M^{
  \mathord{\mathchoice
    {\vcenter{\hbox{\scalebox{0.8}{$\displaystyle\oldblacklozenge$}}}}
    {\vcenter{\hbox{\scalebox{0.8}{$\textstyle\oldblacklozenge$}}}}
    {\vcenter{\hbox{\scalebox{0.8}{$\scriptstyle\oldblacklozenge$}}}}
    {\vcenter{\hbox{\scalebox{0.8}{$\scriptscriptstyle\oldblacklozenge$}}}}
  }
}_\theta(t-)-M^\bullet_\theta(t-)\mid T\ge t\big\}\,dt.
\end{align}
\endgroup
The efficiency-gain identity in \eqref{eq:gain} gives the design rule for when the augmented AJ estimator should improve most over the marginal KM estimator.
Under \Cref{ass:Markov-state}  we can write,

\begingroup
\setstretch{1}
\setlength{\jot}{0pt}
\begin{align} \label{eq:M-theta-variance}
  \mathrm{Var}\big\{M^{
  \mathord{\mathchoice
    {\vcenter{\hbox{\scalebox{0.8}{$\displaystyle\oldblacklozenge$}}}}
    {\vcenter{\hbox{\scalebox{0.8}{$\textstyle\oldblacklozenge$}}}}
    {\vcenter{\hbox{\scalebox{0.8}{$\scriptstyle\oldblacklozenge$}}}}
    {\vcenter{\hbox{\scalebox{0.8}{$\scriptscriptstyle\oldblacklozenge$}}}}
  }
}_\theta(t-)-M^\bullet_\theta(t-)\mid T\ge t\big\}
  =\frac{P_{00}(0,t)P_{01}(0,t)}{S_T(t)^2}\{v_0(t)-v_1(t)\}^2.
\end{align}
\endgroup
Thus recurrence helps precisely when censoring occurs at times when recurrence has split the alive population and the two alive states have different residual survival.
The gain is largest when the following four conditions hold concurrently:
\begin{enumerate}[label=\textnormal{(\roman*)},leftmargin=2em]
\item censoring is appreciable at times that matter for the target, so that $\lambda_C(t)/G(t)$ is large;
\item recurrence has already split the alive population into two non-negligible groups, so that both $P_{00}(0,t)/S_T(t)$ and $P_{01}(0,t)/S_T(t)$ are substantial;
\item the two alive states have sharply different residual survival, so that $|v_1(t)-v_0(t)|$, or $|V_1(t,\tau)-V_0(t,\tau)|$ for RMST, is large;
\item enough subjects remain alive, so that the factor $S_T(t)$ has not already collapsed.
\end{enumerate}
Thus, weak censoring, rare recurrence, very late recurrence, non-prognostic recurrence, or near-complete mortality all force the gain to be small.

\begin{figure}[H]
  \centering
  \includegraphics[width=0.92\linewidth]{gain_snapshot_combined.png}
  \caption{
  \footnotesize   Comparing the variance of two estimators: KM uses only $(\tilde T,\Delta)$, while AJ adds $(\tilde R,\Gamma)$.
    Left col: data designs varying censoring hazard (top), recurrence hazard (mid), post-recurrence hazard (bottom).
    Center col: AJ asymptotic variance reduction over KM across varying data designs.
    Right col: KM vs AJ fixed-$n$ variance across varying data designs.
  }
  \label{fig:gain-snapshot-combined}
\end{figure}

We illustrate the gain identity \eqref{eq:gain} with synthetic data experiments.
We generate $800$ samples according to a parametrized illness-death law:
\begin{align*}
  \lambda_C(t)    = A\,e^{-(t-8)^2/4},
  \qquad
  \lambda_{01}(t) = B\,e^{-(t-4)^2/3},
  \qquad
  \lambda_{02}(t) = 0.02\,e^{-(t-2)^2/18},
  \\
  \lambda_{12}(t)  = \lambda_{02}(t) + \kappa\, b(t),
  \qquad
  b(t) = \bigl(0.03 + 0.0005t\bigr) \big\{ 1 +  {2.4} \{1+e^{-(t-11)/1.7}\}^{-1} \big\}
\end{align*}
on a finite time horizon $\tau=20$, with $n=1000$ subjects, and compute the $\mathcal{F}^{
  \mathord{\mathchoice
    {\vcenter{\hbox{\scalebox{0.8}{$\displaystyle\oldlozenge$}}}}
    {\vcenter{\hbox{\scalebox{0.8}{$\textstyle\oldlozenge$}}}}
    {\vcenter{\hbox{\scalebox{0.8}{$\scriptstyle\oldlozenge$}}}}
    {\vcenter{\hbox{\scalebox{0.8}{$\scriptscriptstyle\oldlozenge$}}}}
  }
}$-measurable AJ and the $\mathcal{F}^\circ$-measurable KM estimators of the RMST.
We report empirical variances via boxplots in Fig. \ref{fig:gain-snapshot-combined}, third column.
We also compare the empirical AJ-vs-KM variance reduction, expressed as a percentage of the KM variance,
against our theoretical prediction \ref{eq:M-theta-variance}, computed numerically from the input $\lambda$'s.
See Fig. \ref{fig:gain-snapshot-combined}, second column.
Simulation details are provided in \Cref{app:no-X-simulations}.

To illustrate the impact of censoring, recurrence, and post-recurrence prognostic lift on variance reduction,
we vary the corresponding input hazards, see Fig. \ref{fig:gain-snapshot-combined}, first column.

We find:
(i) The gain increases with the strength of censoring (row one, column two), but so does the variance of both the KM and the AJ estimators (row one, column three): the stronger the censoring, the harder the estimation problem becomes.
(ii) As recurrence increases from zero, subjects begin transitioning into state $1$, and the gain rises sharply; as recurrence increases further, we observe a slow decay, since subjects in state $1$ are much more likely to transition onward to state $2$, decreasing the occupancies for both states $0$ and $1$.
(iii) While in theory the efficiency gain is always positive, in practice, if the surrogate is not predictive of the true outcome, the AJ estimator can be noisier than KM, as we observe in the empirical curve of the second column.
This, however, is not the case as the surrogate becomes even a little predictive, and we observe a linear proportional relationship between the prognostic lift and the efficiency gain.



\section{Covariate-adjusted estimation}
\label{sec:covariates}

The no-covariate result is clean because the transition law is a finite-dimensional-in-state, nonparametric-in-time object that AJ estimates directly.
Continuous covariates change the problem.
Now the analyst must learn how the whole transition system varies with $X$.
This is not ordinary baseline covariate adjustment for a scalar mean; the nuisance is a conditional illness--death law.
With finite well-populated strata, one may fit AJ within each stratum and standardize \citep{aalen78}.
With continuous or multi-dimensional $X$, however, the object is the conditional transition law $A_0(\cdot\mid X)$.

\subsection{Cox transitional hazard estimator}

The Cox--Breslow setting is used as a benchmark rather than as the general method \citep{cox72,breslow72,andersen12}.
Its role is to illustrate the best-case plug-in regime: when each transition model is correctly specified and finite-dimensional,
the conditional transition law is estimated at a sufficiently fast rate.
Thus ordinary standardization is already first-order valid.
For each transition $\ell=(j,k)\in\mathcal J$, suppose

\begingroup
\setstretch{1}
\setlength{\jot}{0pt}
\begin{equation}
\label{eq:cox-main-transition-model}
  E\{dN_\ell(u)\mid\mathcal F_{u-},X\}
  =Y_j(u)\exp(\beta_{\ell,0}^\top X)d\Lambda_{0,\ell,0}(u).
\end{equation}
Let $s_{\ell,0}(u)=E\{Y_j(u)\exp(\beta_{\ell,0}^\top X)\}$.
\endgroup

\begin{assumption}
\label{ass:cox-main-primitives}
The training data are i.i.d., Assumption~\ref{ass:eif-gain-primitives}   holds conditionally on $X$.
The model \eqref{eq:cox-main-transition-model} is correctly specified.
In addition: for each $\ell\in\mathcal J$,
\begin{enumerate}[label=\textnormal{(CX\arabic*)},leftmargin=2em]
  \item\label{cxmain:bounded}  For $K < \infty$,  $\|X\|\le K$ almost surely, and $\|\beta_{\ell,0}\|<R_\ell<\infty$.
  \item The baseline cumulative hazard is supported on the predictable source support and has finite inverse-risk mass:

  \begingroup
  \setstretch{1}
  \setlength{\jot}{0pt}
  \begin{equation} \label{eq:cox-main-source-support}
    \int I\{s_{\ell,0}(u)=0\}d\Lambda_{0,\ell,0}(u)=0,
    \qquad
    \int\frac{I\{s_{\ell,0}(u)>0\}}{s_{\ell,0}(u)}d\Lambda_{0,\ell,0}(u)<\infty.
  \end{equation}
  \endgroup
  \item\label{cxmain:information} The Cox information matrix is nonsingular, i.e., for $\bar X_\ell(\beta,u) = {s_\ell^{(1)}(\beta,u)} / {s_\ell^{(0)}(\beta,u)}.$

  \begingroup
  \setstretch{1}
  \setlength{\jot}{0pt}
  \begin{align*}
    \lambda_{\min}\left\{\int V_\ell(\beta_{\ell,0},u)s_{\ell,0}(u)d\Lambda_{0,\ell,0}(u)\right\}>0 ,
    \quad
    V_\ell(\beta,u) = \frac{s_\ell^{(2)}(\beta,u)}{s_\ell^{(0)}(\beta,u)} - \bar X_\ell(\beta,u)\bar X_\ell(\beta,u)^\top .
  \end{align*}
  \endgroup
\end{enumerate}
\end{assumption}

The source-support formulation is important in the illness--death setting.
The post-recurrence risk set is empty early in follow-up, so requiring a uniform lower bound for that risk set would rule out the natural process.
The condition instead asks only that the baseline hazard put mass where the corresponding source state has population risk and that the inverse-risk variation be finite.

Let $\widehat\beta_\ell$ be the partial-likelihood estimator and $\widehat\Lambda_{0,\ell}$ the predictable-support Breslow estimator.
The fitted conditional hazards are $d\widehat A_{\ell,n}(u\mid x)=\exp(\widehat\beta_\ell^\top x)d\widehat\Lambda_{0,\ell}(u)$, $\ell\in\mathcal J$.
For fixed $x$, the state probabilities are computed by the product-integral recursion associated with the fitted
transition hazards.
At a jump time $u$,

\begingroup
\setstretch{1}
\setlength{\jot}{0pt}
\begin{align*}
  \widehat P_0(u\mid x)
  &=
  \widehat P_0(u-\mid x)
  \big\{1-\Delta\widehat A_{01,n}(u\mid x)-\Delta\widehat A_{02,n}(u\mid x)\big\},
  \\[1em]
  \widehat P_1(u\mid x)
  &=
  \widehat P_1(u-\mid x)
  \big\{1-\Delta\widehat A_{12,n}(u\mid x)\big\}
  +
  \widehat P_0(u-\mid x)\Delta\widehat A_{01,n}(u\mid x).
\end{align*}
\endgroup
Let $I_1,I_2$ be a partition of $N=n+m$ i.i.d.\ observations, $n=|I_1|$, and $m=|I_2|$.
Let $\widehat A_n$ be fitted on $I_1$, and let $I_2$ be an independent evaluation sample.
The Cox plug-in RMST is

\begingroup
\setstretch{1}
\setlength{\jot}{0pt}
\begin{align} \label{eq:Cox-plug-in}
   \widehat\mu_{n,m}^{\textsf{Cox},\textsf{pi}}(\tau)
  =
  m^{-1}\sum_{i=1}^m
  \mu_{\widehat A_n}(\tau\mid X_i^{\rm ev}), \qquad  \mu_{\widehat A_n}(\tau\mid x)
  =
  \int_0^\tau
  \Big\{ \widehat P_0(t\mid x)+\widehat P_1(t\mid x) \Big\}\,dt .
\end{align}
\endgroup

\begin{proposition} \label{prop:cox-main-regime}
Under Assumption~\ref{ass:cox-main-primitives},
$ \text{for each } \ell\in\mathcal J $,
$ \|\widehat\beta_\ell-\beta_{\ell,0}\|=O_p(n^{-1/2}), $\;
$ \sup_{t\le\tau} \big|\widehat\Lambda_{0,\ell}(t)-\Lambda_{0,\ell,0}(t)\big| =O_p(n^{-1/2}), $\;
and\; $ \widehat\mu_{n,m}^{\textsf{Cox},\textsf{pi}}(\tau)-\mu_\tau =O_p(m^{-1/2}+n^{-1/2}). $
\end{proposition}


\subsection{Flexible conditional hazards: why one-step estimator matters}


The plug-in estimator directly averages the fitted conditional RMST, hence is generally sensitive to first-order errors in the transition hazards.
The one-step estimator adds a correction based on observed-minus-fitted transition increments.
Each increment is weighted by the fitted change in remaining RMST at that transition and adjusted for censoring.
We use the same estimated transition law to construct the fitted RMST, the transition residuals, and their weights.

Estimate the conditional transition law and censoring survival by $\widehat A_n$ and $\widehat G_n$ using a training sample of size $n$.
Let $P_Xg=\int g\,dP_X$, $Pf=\int f\,dP$ and let $\mathbb P_{{\rm ev},m}f=m^{-1}\sum_{i=1}^m f(O_i)$ denote, respectively,
averaging over the population law of $X$, of $O$, and an independent evaluation sample.
Let $\widehat P_{jk,n}(u,t\mid x)$ be the transition probabilities obtained from $\widehat A_n(u\mid x)$ by product integration \eqref{eq:matrix-product-integral}.
Define the fitted residual RMST by

\begingroup
\setstretch{1}
\setlength{\jot}{0pt}
\[
  \widehat V_{j,n}(u,\tau\mid x)
  =\int_u^\tau\sum_{r=0}^1
    \widehat P_{jr,n}(u,t\mid x)\,dt,
  \qquad j=0,1,
\]
\endgroup
and set $\widehat V_{2,n}=0$.
Since all subjects enter in state $0$, write
$\widehat\mu_n(x):=\mu_{\widehat A_n}(\tau\mid x)
=\widehat V_{0,n}(0,\tau\mid x)$.
For $\ell\in\mathcal J$, define
\[
  \widehat H_{\ell,n}(u,x)
  =
  \frac{\widehat V_{k(\ell),n}(u,\tau\mid x)
        -\widehat V_{j(\ell),n}(u,\tau\mid x)}
       {\widehat G_n(u-\mid x)},
\]
where $\widehat G_n(u-\mid x)>0$ on the estimation interval.
Each weight is the fitted change in remaining RMST at transition
$\ell$, adjusted for censoring.

The one-step score adds weighted transition residuals to the
fitted RMST:

\begingroup
\setstretch{1}
\setlength{\jot}{0pt}
\begin{equation}
\label{eq:main-psihat-def}
  \widehat\psi_n(O)
  =
  \widehat\mu_n(X)
  +\sum_{\ell\in\mathcal J}\int_0^\tau
    \widehat H_{\ell,n}(u,X)
    \big\{dN_\ell(u)
    -Y_{j(\ell)}(u)\,d\widehat A_{\ell,n}(u\mid X)\big\}.
\end{equation}
\endgroup
The plug-in and one-step estimators are therefore

\begingroup
\setstretch{1}
\setlength{\jot}{0pt}
\begin{align}
\label{eq:os-est-integrated-main}
  \widehat\mu_{n,m}^{\textsf{pi}}(\tau)
  &=\mathbb P_{{\rm ev},m}\{\widehat\mu_n(X)\},
  \qquad
  \widehat\mu_{n,m}^{\textsf{os}}(\tau)
  =\mathbb P_{{\rm ev},m}\{\widehat\psi_n(O)\}.
\end{align}
\endgroup

The theorem below first controls the population bias $P\widehat\psi_n-\mu_\tau$, conditional on the training sample,
under bounded transition fits and consistent censoring estimation.
It then gives the plug-in convergence rate and, under stronger nuisance-rate conditions,
the efficient expansion of the one-step estimator with a second-order transition--censoring product remainder.

The actual implementation uses cross-fitting.
Let $I_1,\ldots,I_K$ be a partition of $N$ i.i.d.\ observations, with $K$ fixed, $m_k=|I_k|$, and $N=\sum_{k=1}^K m_k$.
For each fold $k$, fit $\widehat A^{(-k)}$ and $\widehat G^{(-k)}$ on the observations outside $I_k$, and evaluate the fitted plug-in and one-step summands only on observations in $I_k$.
For $i\in I_k$, let $\widehat\psi_{-k}(O_i)$ denote \eqref{eq:main-psihat-def} with $(\widehat A_n,\widehat G_n)$ replaced by $(\widehat A^{(-k)},\widehat G^{(-k)})$.
The cross-fitted estimators are
\begin{align} \label{eq:crossfit-plugin-os-main}
  \widehat\mu_{N}^{\textsf{cf},\textsf{pi}}(\tau)
  &=
  \frac1N\sum_{k=1}^K\sum_{i\in I_k} \mu_{\widehat A^{(-k)}}(\tau\mid X_i),
  \qquad
  \widehat\mu_{N}^{\textsf{cf},\textsf{os}}(\tau)
  =
  \frac1N\sum_{k=1}^K\sum_{i\in I_k} \widehat\psi_{-k}(O_i).
\end{align}

Let $\mathscr A_\tau$ be the class of admissible finite-variation cumulative-hazard triplets on $[0,\tau]$, with nondecreasing off-diagonal transition coordinates and row increments at most one.
For $A,B\in\mathscr A_\tau$, let $H=B-A$, and let $d\mathbf H$ be the corresponding signed transition-matrix increment.
Define

\begingroup
\setstretch{1}
\setlength{\jot}{0pt}
\begin{align*}
  \kappa_{\rm PI}(A,B)&=\sup_{0\le s\le t\le\tau}\|P_B(s,t)-P_A(s,t)\|_\infty,
  \\
  \omega_{\rm PI}(A,B)&=\sup_{0\le s\le t\le\tau}\left\|P_B(s,t)-P_A(s,t)-\int_{(s,t]}P_A(s,u-)d\mathbf H(u)P_A(u,t)\right\|_\infty.
\end{align*}
\endgroup
With a sample partition $I_1,\ldots,I_K$, fixed $K$, and $m_k=|I_k|$,
assume there are constants $0<\underline\pi\le\overline\pi<1$ such that $\underline\pi\le m_k/N\le\overline\pi$ for all large $N$ and all folds $k$.
For each fold $k$, let $\mathcal D_{-k}$ be the observations outside $I_k$, and let $\widehat A^{(-k)}$ and $\widehat G^{(-k)}$ be $\mathcal D_{-k}$-measurable estimators.
For $M<\infty$, write
$ \mathscr A_\tau(M)=\{A\in\mathscr A_\tau: |A|_\tau\le M\}, $
with $ |A|_\tau=\sum_{\ell\in\mathcal J}A_\ell(\tau). $
For each fold define

\begingroup
\setstretch{1}
\setlength{\jot}{0pt}
\[
  \kappa_{-k}(x)=\kappa_{\rm PI} \big\{A_0(\cdot\mid x),\,\widehat A^{(-k)}(\cdot\mid x)\big\},
  \qquad
  \omega_{-k}(x)=\omega_{\rm PI} \big\{ A_0(\cdot\mid x),\,\widehat A^{(-k)}(\cdot\mid x) \big\}.
\]
\endgroup

For the sharper local comparison, also write
$
  D_{k,\ell}(u\mid x)
  =
  \widehat A_{\ell}^{(-k)}(u\mid x)-A_{\ell,0}(u\mid x),
$  and $
  \delta_{A,-k}(x)
  =
  \max_{\ell\in\mathcal J}\sup_{0\le u\le\tau}|D_{k,\ell}(u\mid x)|.
$
Define the signed censoring-survival-ratio error

\begingroup
\setstretch{1}
\setlength{\jot}{0pt}
\begin{align*}
  e_{G,-k}(u,x)
  =
  1-\frac{G(u-\mid x)}{\widehat G^{(-k)}(u-\mid x)},
  \qquad
  \eta_{-k}(x)=\sup_{0\le u\le\tau}|e_{G,-k}(u,x)|.
\end{align*}
\endgroup
For an admissible transition law $A$, put

\begingroup
\setstretch{1}
\setlength{\jot}{0pt}
\[
  W_{\ell,A}(u,x)
  =
  V^A_{k(\ell)}(u,\tau\mid x)-V^A_{j(\ell)}(u,\tau\mid x),
  \qquad
  L_{\ell,0}(u,x)
  =
  P^{A_0}_{j(\ell)}(u-\mid x)W_{\ell,A_0}(u,x).
\]
\endgroup

\begin{assumption}
\label{ass:main-transition-martingale-linearization}
For a deterministic sequence $a_N\downarrow0$, suppose that for every fold $k$, relative to the training-sample calendar-time filtration $\mathbb F_{-k}=\{\mathcal F_{-k}(u):0\le u\le\tau\}$,
\begin{align*}
  D_{k,\ell}(u\mid x)
  =
  U_{k,\ell}(u\mid x)+R_{k,\ell}(u\mid x),
  \qquad \ell\in\mathcal J,
\end{align*}
where $U_{k,\ell}(\cdot\mid x)$ is a square-integrable $\mathbb F_{-k}$-martingale and the stochastic Fubini interchanges below are valid.
The process $e_{G,-k}(u,x)$ is left-continuous and $\mathbb F_{-k}$-predictable.
For every bounded $\mathbb F_{-k}$-predictable multiplier $z$, write
$
\|z\|_{2,\infty}
=
\left[
P_X\left\{\sup_{0\le u\le\tau}|z(u,X)|^2\right\}
\right]^{1/2},
$
and define
\begin{align*}
  \mathcal U_{k,z}(t)
  &=
  \sum_{\ell\in\mathcal J}
  P_X\int_0^t
  L_{\ell,0}(u,X)z(u,X)\,dU_{k,\ell}(u\mid X),
\\
  \mathcal R_{k,z}
  &=
  \sum_{\ell\in\mathcal J}
  P_X\int_0^\tau
  L_{\ell,0}(u,X)z(u,X)\,dR_{k,\ell}(u\mid X).
\end{align*}
There are random variables $C_{k,N}=O_p(1)$, uniformly over the fixed number of folds, such that
$
  \langle\mathcal U_{k,z}\rangle(\tau)
  \le C_{k,N}a_N^2\|z\|_{2,\infty}^2,
$ and $
  |\mathcal R_{k,z}|
  \le C_{k,N}a_N\|z\|_{2,\infty}.
$
\end{assumption}

Assumption~\ref{ass:main-transition-martingale-linearization} is the conditional analogue of the weighted Nelson--Aalen linearization.
Taking $z=e_{G,-k}$, using $\|e_{G,-k}\|_{2,\infty}^2=P_X\{\eta_{-k}(X)^2\}$, and applying Lenglart's inequality gives

\begingroup
\setstretch{1}
\setlength{\jot}{0pt}
\begin{align} \label{eq:main-AG-transform-rate}
  \max_{1\le k\le K}
  \left|
  \sum_{\ell\in\mathcal J}P_X\int_0^\tau
  L_{\ell,0}(u,X)e_{G,-k}(u,X)\,dD_{k,\ell}(u\mid X)
  \right|
  =O_p(a_Nc_N)
\end{align}
\endgroup
whenever
$\max_k P_X\{\eta_{-k}(X)^2\}=O_p(c_N^2)$.

\begin{theorem}  \label{thm:plugin-onestep-integrated-summary}
Assume conditionally independent censoring.
Suppose that, for some constants $g_0>0$ and $M<\infty$, the event
\begin{align*}
  \mathcal E_N=
  \bigcap_{k=1}^K
  \left\{
  \begin{array}{l}
    A_0(\cdot\mid x) \text{ and } \widehat A^{(-k)}(\cdot\mid x) \in\mathscr A_\tau(M)\quad \text{for }P_X\text{-a.e. }x,
    \\[2mm]
    G(\cdot\mid x) \text{ and } \widehat G^{(-k)}(\cdot\mid x) \text{ are survival functions},
    \\[1mm]
    \displaystyle\inf_{0\le u\le\tau,\, x}G(u-\mid x)\ge g_0,
    \qquad
    \inf_{0\le u\le\tau,\, x}\widehat G^{(-k)}(u-\mid x)\ge g_0/2
  \end{array}
\right\}
\end{align*}
satisfies $P(\mathcal E_N)\to1$, where the infima over $x$ are understood $P_X$-essentially.
Suppose that, for a deterministic sequence $c_N\downarrow0$,

\begingroup
\setstretch{1}
\setlength{\jot}{0pt}
\begin{align} \label{eq:main-censoring-rate}
  \max_{1\le k\le K}P_X\{\eta_{-k}(X)^2\}=O_p(c_N^2).
\end{align}
\endgroup
Let $\widehat\mu_N^{\textsf{cf},\textsf{pi}}$ and $\widehat\mu_N^{\textsf{cf},\textsf{os}}$ be the
cross-fitted estimators in \eqref{eq:crossfit-plugin-os-main}, with hazard-induced fold-specific one-step weights.
Then the following statements hold.

\begin{enumerate}[label=\textnormal{(\alph*)},leftmargin=2em]
\item \textit{Global one-step robustness.}
The hazard-induced cross-fitted one-step estimator satisfies

\begingroup
\setstretch{1}
\setlength{\jot}{0pt}
\begin{equation} \label{eq:main-os-rate-with-Ghat}
  \widehat\mu_N^{\textsf{cf},\textsf{os}}(\tau)-\mu_\tau
  =
  O_p(N^{-1/2}+c_N).
\end{equation}
\endgroup
This conclusion uses no convergence of the transition learner.

\item \textit{Plug-in rate.}
Suppose, in addition, that for deterministic sequences $a_N,b_N\downarrow0$,

\begingroup
\setstretch{1}
\setlength{\jot}{0pt}
\begin{align} \label{eq:main-plugin-precise-rates}
  \max_{1\le k\le K}P_X\{\delta_{A,-k}(X)^2\}=O_p(a_N^2),
  \qquad
  \max_{1\le k\le K}P_X\{\omega_{-k}(X)\}=O_p(b_N).
\end{align}
\endgroup
Then
the cross-fitted plug-in estimator satisfies
\begingroup
\setstretch{1}
\setlength{\jot}{0pt}
\begin{align}
  \label{eq:main-plugin-rate}
  \widehat\mu_N^{\textsf{cf},\textsf{pi}}(\tau)-\mu_\tau
  =
  O_p(N^{-1/2}+a_N+b_N).
\end{align}
\endgroup

\item \textit{Local second-order one-step rate.}
Suppose \eqref{eq:main-plugin-precise-rates} and Assumption~\ref{ass:main-transition-martingale-linearization} hold.
Let

\begingroup
\setstretch{1}
\setlength{\jot}{0pt}
\begin{align*}
  \psi_0(O)=
  \mu_{A_0}(\tau\mid X)
  +
  \sum_{\ell\in\mathcal J}\int_0^\tau H_{\ell,0}(u,X)dM_{\ell,0}(u),
  \qquad
  \phi^{
  \mathord{\mathchoice
    {\vcenter{\hbox{\scalebox{0.8}{$\displaystyle\oldlozenge$}}}}
    {\vcenter{\hbox{\scalebox{0.8}{$\textstyle\oldlozenge$}}}}
    {\vcenter{\hbox{\scalebox{0.8}{$\scriptstyle\oldlozenge$}}}}
    {\vcenter{\hbox{\scalebox{0.8}{$\scriptscriptstyle\oldlozenge$}}}}
  }
}(O)=\psi_0(O)-\mu_\tau.
\end{align*}
\endgroup
Then

\begingroup
\setstretch{1}
\setlength{\jot}{0pt}
\begin{align} \label{eq:main-os-canonical-expansion}
  \widehat\mu_N^{\textsf{cf},\textsf{os}}(\tau)-\mu_\tau
  =
  (\mathbb P_N-P)\phi^{
  \mathord{\mathchoice
    {\vcenter{\hbox{\scalebox{0.8}{$\displaystyle\oldlozenge$}}}}
    {\vcenter{\hbox{\scalebox{0.8}{$\textstyle\oldlozenge$}}}}
    {\vcenter{\hbox{\scalebox{0.8}{$\scriptstyle\oldlozenge$}}}}
    {\vcenter{\hbox{\scalebox{0.8}{$\scriptscriptstyle\oldlozenge$}}}}
  }
}
  +O_p(a_Nc_N)
  +O_p\{N^{-1/2}(a_N+c_N)\},
\end{align}
\endgroup
and with it

\begingroup
\setstretch{1}
\setlength{\jot}{0pt}
\begin{equation} \label{eq:main-os-canonical-rate}
  \widehat\mu_N^{\textsf{cf},\textsf{os}}(\tau)-\mu_\tau
  =
  O_p(N^{-1/2}+a_Nc_N).
\end{equation}
\endgroup
\end{enumerate}
\end{theorem}

Part (a) is the global robustness statement: bounded admissibility and coherence give a valid one-step mean at rate $N^{-1/2}+c_N$, even when the transition learner is inconsistent.
Parts (b) and (c) are local statements.
The plug-in retains the first-order transition error $a_N$, whereas the coherent one-step correction removes that term and leaves the mixed transition--censoring remainder $a_Nc_N$.    If the nuisance product-rate condition
$
  a_Nc_N=o(N^{-1/2})
$
holds,
then the canonical expansion
$
\widehat\mu_N^{\textsf{cf},\textsf{os}}(\tau)-\mu_\tau
=
(\mathbb P_N-P)\phi^{
  \mathord{\mathchoice
    {\vcenter{\hbox{\scalebox{0.8}{$\displaystyle\oldlozenge$}}}}
    {\vcenter{\hbox{\scalebox{0.8}{$\textstyle\oldlozenge$}}}}
    {\vcenter{\hbox{\scalebox{0.8}{$\scriptstyle\oldlozenge$}}}}
    {\vcenter{\hbox{\scalebox{0.8}{$\scriptscriptstyle\oldlozenge$}}}}
  }
}+o_p(N^{-1/2})
$
holds and the one-step reaches the semiparametric efficiency bound.


In particular, if $a_N\asymp c_N\asymp r_N$ and
$b_N=O(r_N^2)$, then

\begingroup
\setstretch{1}
\setlength{\jot}{0pt}
\[
\widehat\mu_N^{\textsf{cf},\textsf{pi}}(\tau)-\mu_\tau
=O_p(N^{-1/2}+r_N),
\qquad
\widehat\mu_N^{\textsf{cf},\textsf{os}}(\tau)-\mu_\tau
=O_p(N^{-1/2}+r_N^2).
\]
\endgroup
For the standard integrated nonparametric rate $r_N=N^{-\beta/(2\beta+d)}$, the one-step remainder is $N^{-2\beta/(2\beta+d)}$, and canonical root-$N$ inference follows when $\beta>d/2$.
More generally, if the transition and censoring learners have rates $N^{-\alpha_A}$ and $N^{-\alpha_G}$, respectively, the product-rate condition is $\alpha_A+\alpha_G>1/2$.


\section{Numerical Work}

We contrast finite-$N$ performance of eight RMST estimators, obtained by crossing three binary design choices:
First, either death-only data or the death-and-recurrence data \eqref{eq:observed-sa} are used;
marked \textsf{KM} and \textsf{AJ}.
Second, either transition hazards are unadjusted for $X$, as in the no-$X$-AJ \eqref{eq:AJ},
or fit conditional on $X$ with Cox \eqref{eq:cox-main-transition-model} or a flexible learner;
marked ``\textsf{Cov}.''
Third, either the fitted transition law is plugged-in \eqref{eq:Cox-plug-in},
or corrected by one-step augmentation \eqref{eq:os-est-integrated-main};
marked ``\textsf{pi}'' or ``\textsf{os}.''
Crossing recurrence (\textsf{KM/AJ}), covariates (\textsf{$\emptyset$/Cov}), and debiasing (\textsf{pi/os}) yields

\begingroup
\setstretch{1}
\setlength{\jot}{0pt}
\begin{align} \label{eq:eight-estimators}
  \big\{
  \widehat\mu^{\textsf{KM}},\quad
  \widehat\mu^{\textsf{AJ}},\quad
  \widehat\mu^{\textsf{KM-Cov},\textsf{pi}},\quad
  \widehat\mu^{\textsf{AJ-Cov},\textsf{pi}},\quad
  \widehat\mu^{\textsf{KM},\textsf{os}},\quad
  \widehat\mu^{\textsf{AJ},\textsf{os}},\quad
  \widehat\mu^{\textsf{KM-Cov},\textsf{os}},\quad
  \widehat\mu^{\textsf{AJ-Cov},\textsf{os}}
  \big\}.
\end{align}
\endgroup
Each matched pair differs in exactly one feature and isolates the value of recurrence, covariates, or the one-step correction.
\Cref{app:simulations}
provides complete experimental details.


\subsection{Synthetic simulation studies} \label{sec:simulations}

\begin{figure}[t]
  \centering
  \includegraphics[width=\linewidth]{covariate_scenarios_boxplots.pdf}
  \caption{Fixed-$N$ estimator distributions; eight RMST estimators; three experiments.
  }
  \label{fig:sim-boxplots}
\bigskip
\end{figure}

Three simulation designs vary, in turn, misspecification in the post-recurrence hazard,
the flexibility of the nuisance learner, and the amount and timing of recurrence information.
Figure~\ref{fig:sim-boxplots}
reports the $N=800$ replicate error distributions for the strongest cell of each design:
each box is one scenario, sample-size, and estimator cell,
hatching distinguishes one-step from plug-in estimators, and color distinguishes the KM and AJ families.

We find:
(i) $\widehat\mu^{\textsf{AJ-Cov},\textsf{pi}}$ is inconsistent under misspecified $\widehat A_{12}$, but $\widehat\mu^{\textsf{AJ-Cov},\textsf{os}}$ restores consistency, top row.
(ii) $\widehat\mu^{\textsf{AJ-Cov},\textsf{os}}$ removes finite sample bias of $\widehat\mu^{\textsf{AJ-Cov},\textsf{pi}}$ in truly nonlinear design with nuisances learned by gradient-boosted Poisson-loss regression, middle row.
(iii) the efficiency gain from recurrence is lowered by covariate adjustment and when censoring arrives ahead of recurrence,
bottom row.

\Cref{app:sim-covariate-scenarios} details the three mechanisms and their scaling on $N$, \Cref{fig:sim-contrasts}.
Read together, the three designs separate the two questions that the population gain identity of \Cref{subsec:gain-snapshot} conflates:
Simulations 1 and 2 show that the one-step correction protects the target under a wrong or merely flexible nuisance model;
Simulation 3 shows that, even with a correctly specified model, the AJ-versus-KM gain tracks how strongly recurrence has split the risk set and separated the state-specific
residual-survival means, eroding its finite-sample benefit once censoring hides recurrence before it can be observed.



\subsection{Application to semi-synthetic clinical trials data}\label{sec:application}

We use metastatic breast cancer (MBC) clinical-trial curves to build semi-synthetic experiments to contrast the performance of eight estimators \eqref{eq:eight-estimators}.
The data come from a repository of 1,865 drug-therapy trials assembled by \citet{silberholz2019}.
Our setting requires a two-arm randomized trial with armwise curves for the intermediate (progression/PFS) and terminal (death/OS) events.
We conducted a quality-control audit to verify that the repository representation of each trial was sufficiently consistent with the corresponding source publication for our analysis.
This yields 40 comparisons we can simulate from.
Table~\ref{tab:rct_summary} in the Appendix summarizes key features of the patients and their outcomes.

From each trial, we calibrate a latent arm-specific event-time law to the published censoring-adjusted survival curves.
Because the public files do not contain the underlying patient-level censoring process, censoring is generated from a separately specified design.
The resulting application is semi-synthetic: published trial curves determine the event-time marginals, and an explicit generative model determines the joint illness--death path, patient covariates, and censoring mechanism.

\subsubsection{Semi-synthetic data-generation}

We run two sets of experiments.
In both we simulate the illness--death process \eqref{eq:S-state} with transitions $\mathcal J$.
The estimand throughout is the arm-wise RMST with $\tau=24$ months, and the treatment effect is $\mu^1-\mu^0$.
Full details are given in Appendix~\ref{app:datagen}.

We first fix the marginals at prespecified smooth Weibull PFS and OS curves.
The distributions for these curves are chosen to resemble the MBC trials in scale but not fitted specifically to any of them.
We calibrate a single illness--death process such that its induced marginal curves match these two Weibull distributions.
Holding both marginals smooth and fixed allows us to isolate the effect of the coupling mechanism alone.

We then apply the coupling with the marginals from an actual trial from the processed digitized PFS and OS curves.
For clarity of exposition we present one comparison from \cite{conte1996chemotherapy} in detail.
We select this example because it displays the surrogate paradox: at $\tau = 24$ months, treatment improves PFS by 1.322 months on average, while worsening OS by 0.749 months.
Later, in Section~\ref{sec:allTrials}, we further examine all 40 comparisons.
We take only the arm-specific digitized survival values as trial inputs.

The pair of curves alone cannot determine the process that generated it.
PFS counts both recurrence and death without recurrence without distinction.
We therefore construct a family of joint laws for recurrence and death, all of which reproduce the PFS and OS marginals, and present the results across that family.
Let $\pi_j$ denote the share of recurrence-free exits in interval $j$ allocated to recurrence; choosing $\pi_j$ fixes the remaining transition flows, so the compatible laws form a one-dimensional family in each interval.
We evaluate five couplings sweeping the feasible range, \emph{Lower}, \emph{Quarter}, \emph{Midpoint}, \emph{Three-quarter}, and \emph{Upper}.
In the trial-anchored experiment, we evaluate two additional couplings \emph{Smooth} and \emph{Constant projection}.
The couplings differ in how many patients occupy the post-recurrence state, when they enter it, and how quickly they then die.

Each simulated patient carries a baseline covariate $X\sim\mathcal N(0,1)$, drawn independently of the treatment.
It enters all three transition hazards through nonlinear, time-varying log-risk terms,
with per-interval intercepts calibrated so that averaging over $X$ still reproduces the same PFS and OS marginals.

Censoring is imposed after the latent paths are drawn.
We draw censoring times independently of recurrence, death, and the baseline covariate given the treatment arm.
Each arm follows a single two-point law, censoring at a fixed landmark with a probability calibrated so that the pooled-arm expected chance of censoring before death is $50\%$.

\subsubsection{Results and Main takeaways}

\begin{figure}[t]
  \centering
  \includegraphics[width=\linewidth]{CONTE_sampling_and_heatmap.pdf}
  \caption{
    CONTE0796-anchored experiment.
    (a): distributions of no-$X$ KM and AJ plug-in estimators under \emph{Upper}.
    (b): variance reduction for estimator-by-coupling combinations.
  }
  \label{fig:main-mc}
\end{figure}

Figure~\ref{fig:main-mc} shows fixed-$N$ comparisons for the trial-anchored couplings.
Panel~(a) enlarges a single AJ-vs-KM pair.
The two error distributions are centered but differ in spread, a $6.7\%$ variance reduction.
Panel~(b) reports the reduction in coupling-by-estimator combinations.
All twelve are positive.
The estimators preserve the same ordering in their gains: under \emph{Upper}, the no-$X$ estimators improve by $6.7\%$ and $6.2\%$, compared with $5.0\%$ for their covariate-adjusted counterparts.
The smaller adjusted gain is consistent with the baseline covariate already capturing part of the relevant prognostic information.
Collectively, these results show that observing and incorporating recurrence increases efficiency, but the gain depends on the joint law connecting recurrence and death.
This cannot simply be determined from effect sizes and PFS and OS curves.
Importantly, the estimator exploits recurrence without requiring the joint law to be known in advance: in the worst case the benefit is small, and in the best case it is substantial.

These gains translate directly into trial size.
A variance reduction of $\mathbb{V}$ means a Kaplan--Meier analysis requires $1/(1-\mathbb{V})$ times as many patients to reach the same precision.
Under the most informative construction, where $\mathbb{V}\approx5\%$, a trial that would otherwise enroll $600$ patients reaches the same precision with roughly $570$.
At \$50{,}000 per participant \citep{moore2018estimated}, that is a saving of approximately \$1.5 million.
The saving comes from information already collected in the course of the trial.

One might naturally ask whether individual-level and trial-level correlation used to validate surrogate endpoints
\citep{prentice1989surrogate, buyse2000validation} would tell us in advance the efficiency gain with recurrence.
However, they cannot, because the gain also depends on when recurrence occurs relative to when follow-up ends.
These measures were developed to assess whether a surrogate can substitute for the true endpoint.
Our question is different: overall survival remains the estimand and recurrence enters only as auxiliary information.


\subsubsection{Impact Across Trials}\label{sec:allTrials}

\begin{figure}[t]
  \centering
  \begin{minipage}[c]{0.49\textwidth}
    \centering
    \includegraphics[angle=270,width=\linewidth,keepaspectratio]{fig_03_trial_law_gain_heatmap.pdf}

  \end{minipage}
  \hfill
  \begin{minipage}[c]{0.49\textwidth}
    \centering
    \includegraphics[width=\linewidth,]{fig_01_primary_smooth_effect_map.pdf}
  \end{minipage}
  \caption{
    (a):
      variance reduction with recurrence (\%),
      for 40 trials (columns) under each seven transition couplings (rows).
    (b):
      trial's own 12-month PFS (x-axis) and OS (y-axis) treatment effects (months),
      and variance reduction with recurrence (\%, color),
      for 40 trials under \emph{Smooth} coupling.
  }
  \label{fig:all-40-trials}
  \bigskip
\end{figure}

To assess whether the CONTE0796 pattern generalizes, we repeat the same construction across all 40 eligible comparisons,
holding the estimators and the seven transition couplings fixed and varying only the trial.
Figure~\ref{fig:all-40-trials}
reports the resulting empirical variance gain for every trial-by-coupling combination, along with the target-vs-surrogate treatment effects comparison for each trial.






\section{Discussion}
Survival data are important targets for causal inference and pose challenges for identification and efficiency.
Surrogate outcomes carry no first-order efficiency value with complete data,
but recover information lost to censoring and relax the ignorability assumption by refining the predictable residual-survival projection.
We derive the information efficiency bound and the exact gap between the marginal KM and the augmented AJ scores.
In the no-covariate setting, this leads to the natural and efficient standardized AJ estimator.
In the setting with covariates, we show that a cross-fitted hazard-induced one-step estimator achieves rate-robustness and canonical root-$N$ inference.
Interesting follow-up questions are how to select surrogate events toward optimal efficiency and how to employ surrogate events in the design of experiments.




\clearpage