The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.
189,085 characters
Recursive-Head Geometry and Order-Free Efficient Inference in Finite-State Nested Markov Models
\begin{frontmatter}
\title{Recursive-Head Geometry and Order-Free Efficient Inference in
Finite-State Nested Markov Models}
\runtitle{Order-Free Inference in Nested Markov Models}
\begin{aug}
\author[A]{\fnms{Haoyu}~\snm{Wei}\ead[label=e1]{[email removed]}}
\address[A]{Department of Economics, University of California, San Diego\printead[presep={,\ }]{e1}}
\end{aug}
\begin{abstract}
Exploiting the equality restrictions that nested Markov models encode requires their tangent-space geometry. For strictly positive finite-state models on arbitrary acyclic directed mixed graphs (ADMGs), we differentiate the recursive-head chart and prove that the range of its score map is the full tangent space. Intrinsic-set coordinate blocks form an algebraic direct sum, blocks of distinct districts are orthogonal, and the resulting Gram projection needs neither mb-shieldedness nor a district order. Exact counterexamples show that observational centering and kernel normalization alone do not certify tangency.
For the node, complete-source edge, and compatible path-specific intervention targets considered here, boundary substitution yields a normalized configured active law and an order-free canonical-gradient formula, extended to finite mixtures by independent source redraws. A coherent one-step estimator with an exact remainder identity and a model-valid chart-flow targeted maximum likelihood estimator are efficient under stated local conditions. On a narrower source-isolated fixed-node subclass, a sequential estimator has an exact transition-factorized drift and up to $2^m$ nuisance-correctness regimes over $m$ active-district transitions; its all-correct influence function projects onto the canonical gradient, and an exact rational law exhibits a strict variance gap. These results separate order-free efficiency on arbitrary finite-state ADMGs from multiple robustness of a narrower construction.
\end{abstract}
\begin{keyword}[class=MSC]
\kwdgroup[type=primary]{\kwd{62H22}
\kwd{62G20}}
\kwdgroup[type=secondary]{\kwd{62D20}}
\end{keyword}
\begin{keyword}
\kwd{nested Markov model}
\kwd{acyclic directed mixed graph}
\kwd{efficient influence function}
\kwd{targeted maximum likelihood}
\end{keyword}
\end{frontmatter}
\tableofcontents
\section{Introduction}\label{sec:introduction}
Nested Markov models encode equality constraints induced by latent-variable structure that are invisible to ordinary conditional-independence models \citep{richardson2023nested}. Such constraints matter for inference even when the causal functional of interest is already identified: an estimator that ignores an equality restriction satisfied by the data-generating law can forgo precision, because the efficiency bound of the restricted model is never larger, and can be strictly smaller, than that of the model with the restriction dropped. Exploiting this gain requires the tangent space of the restricted model and the projection that yields the target's canonical gradient. This paper supplies that geometry for strictly positive finite-state nested Markov models on arbitrary ADMGs, without mb-shieldedness or an ordering of districts, and develops the efficient and robust estimators it supports.
Equality constraints of latent-variable margins beyond conditional independence were noted in \citep{robins1986new,verma1990equivalence}, and the c-component, or district, factorization for causal identification was developed in \citep{tian2002general}. The Markov properties of ADMGs were established in \citep{richardson2003markov}, and the discrete ordinary Markov model was parameterized in \citep{evans2014markovian}. Richardson, Evans, Robins and Shpitser \citep{richardson2023nested} defined the nested Markov model through fixing, proved that margins of DAG models lie in it, and characterized identifiable node interventions by fixing; the discrete latent-variable model is algebraically equivalent to the corresponding nested Markov model and has the same dimension \citep{evans2018margins}. Evans and Richardson \citep{evans2019smooth} proved that the recursive-head parameterization is a smooth identifiable chart of the finite-state nested model, which is therefore a curved exponential family; that chart is the object differentiated here. Identification of node interventions under latent confounding is complete \citep{shpitser2006identification,shpitser2008complete}; node, edge and path interventions were organized into a hierarchy in \citep{shpitser2016causal}, and graphical conditions for experimental identification of path-specific effects were given in \citep{avin2005identifiability}. The complete-source edge and path-specific classes of Section~\ref{sec:intervention-bridge} are identified under the conditions of \citep{shpitser2013counterfactual,shpitser2018identification}; that section applies those results under their stated hypotheses and does not enlarge them.
On the estimation side, semiparametric efficiency theory and doubly robust estimation \citep{bang2005doubly,bickel1993efficient,robins1994estimation,tsiatis2006semiparametric} were extended to mediation and path-specific functionals in \citep{fulcher2020robust,miles2020semiparametric,tchetgen2012semiparametric,zhou2022semiparametric}, with estimator-specific combinations of correctly specified nuisance components. The efficiency calculations for these mediation functionals use unrestricted observed-data models and do not incorporate additional nested Markov restrictions. Multiple robustness in factorized likelihood models was characterized in \citep{molina2017multiple}. Closest to the present work, Bhattacharya, Nabi and Shpitser \citep{bhattacharya2022semiparametric} obtain doubly robust augmented inverse probability weighted estimators under primal fixability, without mb-shieldedness; their tangent-space decomposition and corresponding efficiency results assume mb-shieldedness. Targeted maximum likelihood estimation \citep{van2006targeted} uses locally least favorable updates; for a scalar target, the universal least favorable submodel \citep{vanderlaan2016universal} follows the canonical gradient along the path. The present work supplies the tangent-space geometry of the nested model on an arbitrary finite-state ADMG---an order-free orthogonal decomposition by native districts with an explicit projection---and a targeting path that stays inside the model, constructed in recursive-head coordinates.
The setting is finite-state and interior throughout: every variable has a fixed finite state space, and the operative statistical model is the strictly positive part $\mathcal N_+(\mathcal G)$ of the nested Markov model $\mathcal N(\mathcal G)$, not the latent-margin model; the arbitrary-cardinality theorems do not rely on the binary worked examples. On this support the recursive-head parameterization of \cite{evans2019smooth} is a smooth identifiable chart and the model is a curved exponential family \citep[Corollary~5.6]{evans2019smooth}, so regular finite-dimensional inference is available in principle. What that fact does not supply is the geometry needed to use it on a graph: which functions of the data are model scores, how the scores attached to different parts of the graph relate to one another, and how an intervention target and its canonical gradient are expressed in the chart. The difficulty is the identification of the correct score subspaces, not dimension counting. A function centered under an observational conditional distribution need not be centered under a fixed intrinsic kernel, and a normalized tilt of a district kernel need not preserve the internal nested constraints of that district; Section~\ref{sec:counterexample} exhibits exact failures of both on a non-mb-shielded graph.
We therefore differentiate the forward recursive-head chart itself: the derivative of the chart map in a coordinate direction, divided pointwise by the mass function at that chart point, is a model-valid score by construction, and the coordinate blocks indexed by intrinsic sets give score blocks whose position relative to one another in $L_2(P)$ is the object of study. Section~\ref{sec:setup} defines this chart score $J_P$ and its blocks $\mathcal S_C(P)$.
\paragraph*{Contributions}
The parameterization is due to \cite{evans2019smooth}; the contributions of this paper are graph-specific consequences of its differential geometry, each stated here through its principal result.
\begin{enumerate}[label=(\roman*)]
\item \emph{Native-district separation without an order (Theorem~\ref{thm:district-orthogonality}).} At every strictly positive law the tangent space is the $L_2(P)$-orthogonal direct sum of the score spaces of the native districts, on every finite-state ADMG, without mb-shieldedness, a district order, or acyclicity of the district quotient. The metric projection onto the tangent space is therefore a sum of districtwise Gram projections (Proposition~\ref{prop:gram-projection}), which is what makes an explicit order-free efficient projection possible.
\item \emph{Configured active-chart selection (Theorem~\ref{thm:normalized-intervention-bridge}).} For the node, complete-source edge, and compatible path-specific targets of classes (I1)--(I3) in Section~\ref{sec:intervention-bridge}, substituting the boundary values of the intervention selects existing tail-indexed coordinates and reassembles a normalized nested Markov law of the active graph, throughout $\mathcal N_+(\mathcal G)$. The canonical gradient of the identified target is then a direct chart derivative solved district by district (Theorem~\ref{thm:intervention-chart-eif}), with no vertex order and no change-of-measure argument; finite mixtures defined by independent source redraws are covered by a product rule (Theorem~\ref{thm:independent-source-mixture-eif}).
\item \emph{Efficient one-step estimation and chart targeting (Theorems~\ref{thm:coherent-one-step} and~\ref{thm:chart-tmle}).} For an initial fit that is not an interior likelihood solution, a coherent one-step estimator built at a single fitted chart law has an exact remainder identity and is efficient under an $o_p(n^{-1/4})$ chart rate. Under the local conditions of Theorem~\ref{thm:chart-tmle}, a least-favorable flow in recursive-head coordinates moves an initial model law to an empirical efficient-influence-function root on an event of probability tending to one; the resulting targeted estimator preserves substitution and range on every sample through its prespecified fallback and is asymptotically equivalent to the same-sample one-step estimator.
\item \emph{A construction-specific efficiency--robustness comparison (Theorems~\ref{thm:sequential-multiple-robustness} and~\ref{thm:sequential-projection-comparison}).} On the narrower source-isolated fixed-node subclass of Section~\ref{sec:sequential-union}, a full-history sequential estimator has an exact transition-factorized drift and is consistent on up to $2^m$ nuisance-correctness regimes, $m$ being the number of active-district transitions, but its all-correct influence function is generally not canonical: its projection onto the tangent space is the order-free canonical gradient, and an exact law exhibits a strict variance gap (Proposition~\ref{prop:sequential-strict-variance-gap}). The comparison is between two constructions, not an impossibility theorem, and no union-model result is claimed for the full range of interventions in (I1)--(I3).
\end{enumerate}
\paragraph*{Organization}
Section~\ref{sec:setup} fixes notation and the finite-state chart; Section~\ref{sec:geometry} proves the tangent-space theorem, the district decomposition, and the Gram projection; Section~\ref{sec:counterexample} records the exact failure of the larger raw spaces. Section~\ref{sec:intervention-bridge} states the causal identification input and proves the normalized active chart, its canonical gradient, and the independent-draw mixture extension. Section~\ref{sec:targeted} develops the coherent one-step estimator and the chart-flow targeted estimator. Section~\ref{sec:sequential-union} gives the restricted sequential alternative, its exact union drift, and the variance comparison, which Section~\ref{sec:finite-sample} illustrates at one law; Section~\ref{sec:discussion} discusses what an extension beyond fixed finite support would require. \ifarxivversion Appendix~A and Appendix~\ref{app:prototype-ledger} are included below, with their sections, results, equations, figures and tables all numbered continuously with the article\else The Supplementary Material carries Appendix~A and Appendix~\ref{app:prototype-ledger}, whose sections, results and equations are numbered continuously with this article and whose figures and tables carry an S prefix\fi: proofs of the results of Sections~\ref{sec:setup}--\ref{sec:sequential-union} are collected in Sections~\ref{app:finite-state-parameterization-proof} and~\ref{app:proofs}, with the result-by-result scope summary in Section~\ref{app:results-scope} and the notation index in Section~\ref{app:notation}, and Appendix~\ref{app:prototype-ledger} keeps the exact computational verification separate from the general proofs.
\section{Finite-state nested Markov setup}\label{sec:setup}
Let $\mathcal G=(V,E)$ be an acyclic directed mixed graph on a finite vertex set $V$, and let $X_V=(X_v:v\in V)$ take values in $\mathfrak X_V=\mathop{\mathchoice{\vcenter{\hbox{\scalebox{1.7}{$\times$}}}}{\vcenter{\hbox{\scalebox{1.25}{$\times$}}}}{\vcenter{\hbox{\scalebox{0.9}{$\times$}}}}{\vcenter{\hbox{\scalebox{0.7}{$\times$}}}}}\displaylimits_{v\in V}\mathfrak X_v$, where every $\mathfrak X_v$ is \emph{finite} and \emph{has at least two elements}.
A conditional ADMG (CADMG) $\mathcal H=(R,W,E_{\mathcal H})$ \citep[Definition~2.1]{evans2019smooth} has disjoint sets $R$ of \emph{random} vertices and $W$ of \emph{fixed} vertices.
Its directed edges belong to $(R\cup W)\times R$, its bidirected edges have both endpoints in $R$, and its directed part is acyclic; in particular, no edge has an arrowhead at a fixed vertex.
Under the standing convention immediately following Definition 2.1 of \cite{evans2019smooth}, each displayed fixed vertex also has at least one child in $R$; Richardson-style fixing graphs may temporarily retain childless fixed vertices, which the reduction defined below removes.
A probability kernel for $X_R$ given $X_W$ is a nonnegative array $q_R(x_R\mid x_W)$ satisfying
\[
\sum_{x_R\in\mathfrak X_R}q_R(x_R\mid x_W)=1 \quad\text{ for every } \quad x_W\in\mathfrak X_W.
\]
An ADMG is the special case $W=\varnothing$.
For a law $P$, write $p$ for its probability mass function and $Pf=\mathrm E_Pf$. Probabilities of events are written with an upright $\mathrm P$: $\mathrm P(\cdot)$ when the law is generic or clear from the context, and $\mathrm P_{Q}(\cdot)$ under a named law $Q$, so that $\mathrm P_P(X_A=x_A)=\sum_{x_{V\setminus A}}p(x_V)$; for the chart laws $P_\theta$ introduced below we abbreviate $\mathrm P_{P_\theta}$ to $\mathrm P_\theta$. Italic $P$, $P_0$, $P_\theta$ and $P^{\mathsf I}$ denote laws, lower-case $p$ and $q$ denote mass functions and kernels, and $\mathbb P_n$ denotes an empirical measure. We write $\mathcal N(\mathcal G)$ for the finite-state nested Markov model of \cite[Definition~3.1]{evans2018margins} and define its strictly positive part by
\[
\mathcal N_+(\mathcal G) :=\{P\in\mathcal N(\mathcal G):p(x_V)>0 ~~ \text{ for every } ~~ x_V\in\mathfrak X_V\}.
\]
We use the CADMG recursive factorization of \cite[Definition 3.3]{evans2019smooth} for coordinate synthesis.
Proposition~\ref{prop:er-equivalence-rrs-compatibility} below establishes the bridge in the direction this paper uses: every law in $\mathcal N_+(\mathcal G)$ is nested Markov in the sense of \cite[Definitions~27--28, Proposition~29 and Theorem~38]{richardson2023nested}, so the fixing calculus of that source applies to such laws throughout --- the recursive formulation constructs the forward chart, and the fixing formulation identifies its reachable kernels. The converse direction is neither claimed nor used.
Write $\mathcal D(\mathcal G)$ for the districts and $\mathcal I(\mathcal G)$ for the intrinsic sets \citep[see Definition 3 and Definition 33 in][]{richardson2023nested}; the sans-serif symbol $\mathsf I$ is reserved for interventions.
For $A\subseteq V$, write $\mathcal G_A$ for the ordinary vertex-induced ADMG on $A$: its vertex set is $A$, and it retains exactly those directed and bidirected edges of $\mathcal G$ whose endpoints both lie in $A$. Applied to a CADMG the same construction additionally preserves each vertex's random or fixed status \citep{richardson2023nested}; for an ADMG that clause is vacuous, since every vertex is random. Square brackets are reserved throughout this paper for the reduced CADMG of \cite[Definition 2.8]{evans2019smooth}, so $\mathcal G[C]$ never abbreviates an unreduced fixing graph.
Following \cite[Definition 2.2]{evans2019smooth}, graphical relations are applied disjunctively to sets; in particular, $\operatorname{pa}_{\mathcal H}(A):=\bigcup_{a\in A}\operatorname{pa}_{\mathcal H}(a)$.
For $C\in\mathcal I(\mathcal G)$, its recursive head and tail are
\[
H(C):=C\setminus\operatorname{pa}_{\mathcal G}(C), \qquad T(C):=\operatorname{pa}_{\mathcal G}(C).
\]
Thus, $H(C)=\operatorname{sterile}_{\mathcal G}(C)$ is the \emph{sterile} subset of $C$---the set of sink nodes in the induced subgraph on $C$---and $T(C)\cap C=C\setminus H(C)$.
Recursive heads and intrinsic sets are in one-to-one correspondence \citep[Definition 4.1, Lemma 4.6, and the following remark]{evans2019smooth}. Lemma~\ref{lem:intrinsic-set-bridge} of the Supplementary Material shows that the intrinsic sets of Definition~33 of \cite{richardson2023nested} are exactly those of that source, so the correspondence applies to $\mathcal I(\mathcal G)$ as defined here.
Every intrinsic set is bidirected-connected and therefore lies in a unique native district of $\mathcal G$.
More generally, the districts of $\mathcal G$ are its bidirected-connected components, so every nonempty bidirected-connected $S\subseteq V$ lies in exactly one of them; write $\delta(S)$ for that district.
Besides every intrinsic set, this covers every district of a vertex-induced subgraph $\mathcal G_A$, because $\mathcal G_A$ retains exactly the bidirected edges of $\mathcal G$ with both endpoints in $A$, so a bidirected path witnessing connectedness in $\mathcal G_A$ is one in $\mathcal G$.
For $D\in\mathcal D(\mathcal G)$, write $\mathfrak d_D(\mathcal G)$ for the district CADMG of \cite[Definition 2.5]{evans2019smooth}, whose random vertices are $D$ and whose fixed vertices are $\operatorname{pa}_{\mathcal G}(D)\setminus D$.
It retains every edge of $\mathcal G$ with an arrowhead at a vertex in $D$.
The forward map defined in \eqref{eq:general-forward-map} applies verbatim to a CADMG, conditionally on each value of its fixed vertices; we write that map as $\Phi_{\mathfrak d_D(\mathcal G)}$.
More generally, for $C\subseteq V$, let $\mathcal G[C]$ denote the reduced CADMG of \cite[Definition 2.8]{evans2019smooth}: its random vertices are $C$, its fixed vertices are $\operatorname{pa}_{\mathcal G}(C)\setminus C$, its bidirected edges are those with both endpoints in $C$, and its directed edges run from $C\cup\operatorname{pa}_{\mathcal G}(C)$ into $C$.
Equivalently, it is the subgraph containing precisely the edges whose arrowheads are all in $C$.
When $C$ is intrinsic, this is the reduced CADMG used for its intrinsic kernel.
A random vertex $v\in R$ of a CADMG $\mathcal H=(R,W,E_{\mathcal H})$ is \emph{fixable} when no other vertex of its district is a directed descendant of it, $\operatorname{de}_{\mathcal H}(v)\cap\operatorname{dis}_{\mathcal H}(v)=\{v\}$; write $\mathbb F(\mathcal H)$ for the set of fixable vertices. Fixing a fixable $v$ deletes every edge with an arrowhead at $v$ and makes $v$ a fixed vertex. For a strictly positive kernel $q$ that is Markov with respect to $\mathcal H$, the kernel operation divides by the conditional of $X_v$ given its Markov blanket $\operatorname{mb}_{\mathcal H}(v):=\{\operatorname{dis}_{\mathcal H}(v)\cup\operatorname{pa}_{\mathcal H}(\operatorname{dis}_{\mathcal H}(v))\}\setminus\{v\}$ in the current graph:
\[
\phi_v(q;\mathcal H)(x_{R\setminus\{v\}}\mid x_{W\cup\{v\}}) :=\frac{q(x_R\mid x_W)}{q(x_v\mid x_{\operatorname{mb}_{\mathcal H}(v)})},
\]
where the divisor is $q(x_v\mid x_{\operatorname{mb}_{\mathcal H}(v)\cap R},x_W)$, formed by marginalizing and conditionalizing $q$ over random variables; the Markov property makes it independent of the fixed arguments indexed by $W\setminus\operatorname{mb}_{\mathcal H}(v)$, which the displayed shorthand suppresses \citep[Definitions~17 and~19]{richardson2023nested}.
A fixing sequence is \emph{valid} when each vertex is fixable in the graph current at its turn, iterated fixing $\phi_\omega$ is the composition along a valid sequence $\omega$, and a set is \emph{reachable} when some valid sequence produces it as the random set \citep[Definition~26]{richardson2023nested}; a set is \emph{intrinsic} when it is in addition a district in a reachable CADMG \citep[Definition~33]{richardson2023nested}. Reachability alone is not sufficient --- in $a\to b$ with a third isolated vertex, $\{a,b\}$ is reachable but not intrinsic --- whereas in a directed acyclic graph every district is a singleton, so the intrinsic sets are exactly the singletons $\{v\}$, with $T(\{v\})=\operatorname{pa}_{\mathcal G}(v)$, and the chart below reduces to the ordinary conditional-probability parameterization.
Every intrinsic $C$ is reachable as the entire random set of a terminal CADMG whose single district is $C$ (Section~\ref{app:er-rrs-compatibility} of the Supplementary Material). Intrinsicness is what licenses the intrinsic kernel $q_C$ on $\mathcal G[C]$ and the coordinate block indexed by $H(C)$ and $T(C)$: those two sets are defined by the displayed set operations for any $C$, but only for intrinsic $C$ do they index a block of the chart.
For a reachable random set $R$ with raw fixing pair $(\widetilde{\mathcal H}_R,\widetilde q_R)$, define the \emph{reduction}
\[
\operatorname{red}(\widetilde{\mathcal H}_R,\widetilde q_R) :=(\mathcal H_R,q_R)
\]
by deleting the childless --- hence isolated --- fixed vertices of $\widetilde{\mathcal H}_R$ and suppressing the corresponding kernel arguments; the operation is well defined whenever $\widetilde q_R$ is constant in those arguments.
This is the general operator $\operatorname{red}_S$, the subscript naming the surviving random set; $\operatorname{red}_C$, $\operatorname{red}_D$, $\operatorname{red}_\Delta$ and the stagewise $\operatorname{red}_{R_j}$ below are its instances.
The reduction is pair-valued; $\operatorname{red}_R(\widetilde q_R)$ denotes its kernel component, the raw graph being understood.
\begin{proposition}[Compatibility of the recursive and fixing formulations]\label{prop:er-equivalence-rrs-compatibility}
Let $\mathcal G$ be an ADMG on a finite $V$ with every $\mathfrak X_v$ finite, and let $p>0$ be a mass function on $\mathfrak X_V$ with law $P$; for $V=\varnothing$, every clause is read through the empty reduced-CADMG convention of Appendix~\ref{app:finite-state-parameterization-proof} of the Supplementary Material, under which all three descriptions in \textup{(i)} hold trivially.
\begin{enumerate}[label=(\roman*)]
\item The following are equivalent: \textup{(a)} $P\in\mathcal N(\mathcal G)$, the model of \cite[Definition~3.1]{evans2018margins}; \textup{(b)} $p$ is representable in the head--tail parameterization of that source (\S3.2, opening paragraph); \textup{(c)} $p$ recursively factorizes according to $\mathcal G$ in the sense of \cite[Definition~3.3]{evans2019smooth}. Here \textup{(b)}$\iff$\textup{(c)} is \cite[Theorem~5.4]{evans2019smooth}, whose general finite-state form is Proposition~\ref{prop:standing-finite-state-import}\textup{(i)}.
\item If the equivalent conditions in \textup{(i)} hold, then $P$ satisfies the nested Markov property of \cite[Definition~27]{richardson2023nested}. Consequently, for every reachable random set $R$, the raw fixing pair $(\phi_\omega(\mathcal G),\phi_\omega(p;\mathcal G))$ is the same for every valid sequence $\omega$ reaching $R$ \citep[Theorem~31]{richardson2023nested}; the raw kernel $\widetilde q_R:=\phi_\omega(p;\mathcal G)$ is strictly positive, and it is constant in the arguments indexed by the childless fixed vertices of $\phi_\omega(\mathcal G)$ \citep[Corollary~32]{richardson2023nested}; and its reduction $q_R:=\operatorname{red}_R(\widetilde q_R)$ is a strictly positive kernel that recursively factorizes according to the reduced CADMG $\mathcal H_R$, hence lies in $\mathcal P^c(\mathcal H_R)$ by Lemma~\ref{lem:rf-implies-markov}.
\item If the equivalent conditions in \textup{(i)} hold, then for every reduced reachable pair $(\mathcal H_0,q_{R_0})$ reached by a valid sequence and every $C\in\mathcal I(\mathcal H_0)$, the kernel $r_C$ recursively derived from $(\mathcal H_0,q_{R_0})$ as in Proposition~\ref{prop:standing-finite-state-import}\textup{(iii)} equals the intrinsic kernel $q_C$ of \eqref{eq:fixing-kernel-bridge} below, for every admissible derivation.
\end{enumerate}
\end{proposition}
Clause \textup{(i)} is by citation, as stated. Clauses \textup{(ii)}--\textup{(iii)}, the transfer of Markov membership across constant extension by childless fixed vertices, and the derivation-level form of \textup{(iii)} are stated in full and proved in Section~\ref{app:er-rrs-compatibility} of the Supplementary Material (Lemma~\ref{lem:rf-implies-markov} and Proposition~\ref{prop:er-rrs-compatibility-full}), by induction along the fixing sequence.
The intrinsic kernel of $C\in\mathcal I(\mathcal G)$ is the reduction of its fixing kernel. Let $P\in\mathcal N_+(\mathcal G)$ have mass function $p$ and put $\widetilde\mathcal G_C:=\phi_{V\setminus C}(\mathcal G)$ and $\widetilde q_C:=\phi_{V\setminus C}(p;\mathcal G)$, the unreduced fixing graph and kernel, which by Proposition~\ref{prop:er-equivalence-rrs-compatibility}(ii) do not depend on the valid sequence; following the convention fixed at the start of this section, the unreduced graph is written without brackets.
Deleting the childless fixed vertices $Z_C:=(V\setminus C)\setminus\operatorname{pa}_{\mathcal G}(C)$ of $\widetilde\mathcal G_C$ leaves exactly $\mathcal G[C]$ (Section~\ref{app:er-rrs-compatibility} of the Supplementary Material). By Proposition~\ref{prop:er-equivalence-rrs-compatibility}(ii), $\widetilde q_C$ is constant in $x_{Z_C}$, with the external parents $\operatorname{pa}_{\mathcal G}(C)\setminus C$ as its retained argument set, so the reduction operator applies and its kernel component defines the parent-indexed intrinsic kernel on the reduced CADMG $\mathcal G[C]$,
\begin{equation}\label{eq:fixing-kernel-bridge}
q_C(x_C\mid x_{\operatorname{pa}_{\mathcal G}(C)\setminus C}) :=\operatorname{red}_C(\widetilde q_C)(x_C\mid x_{\operatorname{pa}_{\mathcal G}(C)\setminus C}),
\end{equation}
equivalently $\widetilde q_C(x_C\mid x_{V\setminus C})=q_C(x_C\mid x_{\operatorname{pa}_{\mathcal G}(C)\setminus C})$ with the right side read as its constant extension; no claim is made that every retained argument is functionally active at every law.
For every district $D\in\mathcal D(\mathcal G)$ the two definitions agree, $\mathfrak d_D(\mathcal G)=\mathcal G[D]$ (Section~\ref{app:er-rrs-compatibility} of the Supplementary Material).
\subsection{General finite-state recursive-head coordinates}
For each $v\in V$, choose a \emph{corner state} $k_v\in\mathfrak X_v$ and set
\[
\widetilde\mathfrak X_v:=\mathfrak X_v\setminus\{k_v\}, \qquad \widetilde\mathfrak X_A:=\mathop{\mathchoice{\vcenter{\hbox{\scalebox{1.7}{$\times$}}}}{\vcenter{\hbox{\scalebox{1.25}{$\times$}}}}{\vcenter{\hbox{\scalebox{0.9}{$\times$}}}}{\vcenter{\hbox{\scalebox{0.7}{$\times$}}}}}\displaylimits_{v\in A}\widetilde\mathfrak X_v.
\]
In the binary examples $k_v=1$, so $\widetilde\mathfrak X_v=\{0\}$ and the single indexed head event is the all-zero event.
The corner state must not be confused with that indexed event.
For $C\in\mathcal I(\mathcal G)$, the coordinate block of $C$ is the family of real numbers
\begin{equation}\label{eq:general-coordinate-block}
\theta_C(a_H\mid x_T), \qquad a_H\in\widetilde\mathfrak X_{H(C)},\quad x_T\in\mathfrak X_{T(C)},
\end{equation}
one for each displayed pair of arguments. This display names the block and fixes its index ranges; away from a law the entries are free coordinates, and at a law they are defined by \eqref{eq:coordinate-extraction} below.
For $C\in\mathcal I(\mathcal G)$, define the coordinate index set and its associated array space by
\[
\mathcal A_C :=\widetilde\mathfrak X_{H(C)}\times\mathfrak X_{T(C)}, \qquad U_C:=\mathbb R^{\mathcal A_C}, \qquad d_C:=|\mathcal A_C|.
\]
Form the \emph{tagged} disjoint union
\[
\mathcal A_{\mathcal G} :=\bigsqcup_{C\in\mathcal I(\mathcal G)}\bigl(\{C\}\times\mathcal A_C\bigr), \qquad U_{\mathcal G}:=\mathbb R^{\mathcal A_{\mathcal G}}, \qquad d:=|\mathcal A_{\mathcal G}|=\sum_{C\in\mathcal I(\mathcal G)}d_C.
\]
At a law $P\in\mathcal N_+(\mathcal G)$, let $q_C$ be its intrinsic kernel defined by the fixing bridge in \eqref{eq:fixing-kernel-bridge}. Every numerator and conditional divisor in the one-step fixing operation is positive when $p$ is strictly positive; induction over a valid fixing sequence therefore gives $\widetilde q_C>0$, and hence $q_C>0$. The coordinate is then \emph{defined} as the kernel conditional probability
\begin{equation}\label{eq:coordinate-extraction}
\theta_C(a_H\mid x_T) = \frac{q_C(a_H,x_{C\setminus H}\mid x_{T\setminus C})} { \sum_{z_H\in\mathfrak X_H} q_C(z_H,x_{C\setminus H}\mid x_{T\setminus C})}, \qquad H=H(C),\ T=T(C).
\end{equation}
Here, $x_T=(x_{T\cap C},x_{T\setminus C})$ and $T\cap C=C\setminus H$, so $x_T$ uniquely supplies both the internal tail configuration $x_{C\setminus H}$ and the external fixed-parent configuration $x_{T\setminus C}$. Strict positivity makes the denominator positive.
Collecting these extracted arrays as the $C$-tagged blocks gives the map $\Theta_{\mathcal G}:\mathcal N_+(\mathcal G)\to U_{\mathcal G}$, characterized by $\{\Theta_{\mathcal G}(P)\}_C=\theta_C(P)$ for every $C\in\mathcal I(\mathcal G)$.
Let $\iota_C:U_C\to U_{\mathcal G}$ extend an array by zero outside the coordinates tagged by $C$.
For $\theta\in U_{\mathcal G}$, write $\theta_C\in U_C$ for its restriction to the coordinates tagged by $C$. Then $\theta=\sum_{C\in\mathcal I(\mathcal G)}\iota_C(\theta_C)$ uniquely. The images of distinct blocks have disjoint tagged supports, so, by construction,
\[
U_{\mathcal G}=\bigoplus_{C\in\mathcal I(\mathcal G)}\iota_C(U_C).
\]
After fixing an ordering of the finite set $\mathcal A_{\mathcal G}$, we identify $U_{\mathcal G}$ with $\mathbb R^d$; the same ordering identifies each $U_C$ with $\mathbb R^{d_C}$. The direct sum just displayed is a property of the tagged index set, and $d$ is the model dimension of \cite[Corollary~5.6]{evans2019smooth}, whose sum over recursive heads becomes the sum over intrinsic sets under the head--intrinsic-set correspondence. Explicitly,
\[
d_C =\left\{\prod_{v\in H(C)}(|\mathfrak X_v|-1)\right\} \left\{\prod_{v\in T(C)}|\mathfrak X_v|\right\},
\]
which reduces to $d_C=2^{|T(C)|}$ for binary variables.
\subsection{The forward M\"obius map}
For each recursive head $H$, let $C(H)$ denote its unique intrinsic set. For $B\subseteq V$, Definition 4.3 of \cite{evans2019smooth} defines $I_{\mathcal G}(B)$ as the stable random vertex set obtained by alternately restricting to the random ancestors of $B$ and to the districts meeting $B$; when this set is bidirected-connected, that definition calls it the \emph{intrinsic closure} of $B$.
In particular, Lemma 4.6 of that paper gives $I_{\mathcal G}(H)=C(H)$ for every recursive head. Define the strict head order
\[
H_1\prec H_2 \quad\Longleftrightarrow\quad I_{\mathcal G}(H_1)\subsetneq I_{\mathcal G}(H_2).
\]
For $A\subseteq V$, let $\mathfrak P_{\mathcal G}(A)$ be the recursive-head partition of \cite[Definition 4.11]{evans2019smooth}: select the $\prec$-maximal recursive heads contained in $A$, remove their vertices, and repeat on the remainder; the recursion is well defined because the selected maximal heads are disjoint and $\prec$ is partition-suitable \citep[Propositions B.1 and B.5]{evans2019smooth}.\footnote{\cite{evans2019smooth} use $\Phi_{\mathcal G}$ for the maximal-head selector inside this recursion. We reserve $\Phi_{\mathcal G}$ for the forward synthesis map and write $\mathfrak P_{\mathcal G}$ for the resulting partition.}
For a configuration of the random vertices $R$, set
\[
\operatorname{nc}(x_R) :=\{v\in R:x_v\in\widetilde\mathfrak X_v\},
\]
the set of coordinates taking a noncorner state. All the preceding definitions extend to a finite-state CADMG $\mathcal H=(R,W,E_{\mathcal H})$ by taking intrinsic sets and graphical relations relative to $\mathcal H$. Thus, with the graph clear from context, we use the same notation $H(C)$, $T(C)$, $\mathfrak P_{\mathcal H}$ and $\mathcal A_C$, and form the tagged array space
\[
\mathcal A_{\mathcal H} :=\bigsqcup_{C\in\mathcal I(\mathcal H)}(\{C\}\times\mathcal A_C), \qquad U_{\mathcal H}:=\mathbb R^{\mathcal A_{\mathcal H}}.
\]
Corner states are required only for the random vertices. If a reachable reduction $\mathcal H[A]$ has random set $A\subseteq R$, then $\theta|_A\in U_{\mathcal H[A]}$ denotes restriction to the tagged blocks whose intrinsic sets are contained in $A$. For $A\subseteq R$, define the partition product
\[
F_A^{\mathcal H}(\theta;y_A,x_R,x_W) :=\prod_{H\in\mathfrak P_{\mathcal H}(A)} \theta_{C(H)}(y_H\mid x_{T(C(H))}),
\]
with the empty product equal to one. The CADMG forward map is
\begin{align}
\Phi_{\mathcal H}(\theta)(x_R\mid x_W) := {}& \sum_{\operatorname{nc}(x_R)\subseteq A\subseteq R} (-1)^{|A\setminus\operatorname{nc}(x_R)|} \sum_{\substack{y_A\in\widetilde\mathfrak X_A\\ y_{\operatorname{nc}(x_R)} =x_{\operatorname{nc}(x_R)}}} F_A^{\mathcal H}(\theta;y_A,x_R,x_W). \label{eq:general-forward-map}
\end{align}
For an ADMG $\mathcal G$, this specializes to $\Phi_{\mathcal G}:U_{\mathcal G}\to\mathbb R^{\mathfrak X_V}$. Formula \eqref{eq:general-forward-map} is the general finite-state extension announced in Appendix C of \cite{evans2019smooth}; in the binary case its inner sum has a single term. It is a generalized M\"obius synthesis map from recursive-head coordinates to a possibly signed cell array. For a one-vertex ADMG it reduces to
\[
\Phi_{\mathcal G}(\theta)(x_v)
= \begin{cases}
\theta_{\{v\}}(x_v),&x_v\in\widetilde\mathfrak X_v,\\[2pt]
1-\displaystyle\sum_{a\in\widetilde\mathfrak X_v} \theta_{\{v\}}(a),&x_v=k_v,
\end{cases}
\]
so the alternating sum supplies the omitted corner cell.
A concrete four-vertex instance of the district forward map, for the cyclic-quotient graph of Section~\ref{sec:cyclic-control}, is recorded in Section~\ref{app:B-cyclic} of the Supplementary Material.
\begin{remark}[Cellwise polynomial and multiaffine structure]\label{rem:forward-map-multiaffine}
For fixed $(x_R,x_W)$, saying that $\Phi_{\mathcal H}$ is polynomial means that
\[
\eta\longmapsto \Phi_{\mathcal H}(\eta)(x_R\mid x_W)
\]
is a real polynomial in the finitely many tagged coordinates of $U_{\mathcal H}$. More strongly, every monomial in \eqref{eq:general-forward-map} is square-free. Indeed, each partition $\mathfrak P_{\mathcal H}(A)$ contains distinct recursive heads; the head--intrinsic-set bijection and the tagged coordinate construction then place the corresponding factors in distinct coordinate blocks. Thus no coordinate can occur twice in a partition product. Finite summation preserves this property, so every cell is affine in each coordinate separately.
For $i\in\mathcal A_{\mathcal H}$, let $e_i\in U_{\mathcal H}$ be the corresponding coordinate vector. For every $\eta\in U_{\mathcal H}$ and $t\in\mathbb R$, the polynomial map is Fr\'echet differentiable; write $\mathbb D\Phi_{\mathcal H,\eta}$ for its differential. Multiaffinity gives the array identity
\[
\Phi_{\mathcal H}(\eta+t e_i) =\Phi_{\mathcal H}(\eta) +t\{\Phi_{\mathcal H}(\eta+e_i)-\Phi_{\mathcal H}(\eta)\}, \quad \mathbb D\Phi_{\mathcal H,\eta}[e_i] =\Phi_{\mathcal H}(\eta+e_i)-\Phi_{\mathcal H}(\eta).
\]
These are algebraic identities on the whole ambient coordinate space; the synthesized array need not be nonnegative at either $\eta$ or $\eta+e_i$.
\end{remark}
\begin{lemma}[Sterile-vertex marginalization]\label{lem:sterile-marginalization}
Let $\mathcal H=(R,W,E_{\mathcal H})$ be a finite-state CADMG with $R\ne\varnothing$, and let $a\in R$ be sterile among the random vertices, so $\operatorname{ch}_{\mathcal H}(a)\cap R=\varnothing$. Put $R_a:=R\setminus\{a\}$ and $W_a:=\operatorname{pa}_{\mathcal H}(R_a)\setminus R_a\subseteq W$. Then, for every $\theta\in U_{\mathcal H}$, $x_{R_a}\in\mathfrak X_{R_a}$ and $x_W\in\mathfrak X_W$,
\begin{equation}\label{eq:sterile-marginalization}
\sum_{x_a\in\mathfrak X_a} \Phi_{\mathcal H}(\theta)(x_{R_a},x_a\mid x_W) =\Phi_{\mathcal H[R_a]}(\theta|_{R_a}) (x_{R_a}\mid x_{W_a}).
\end{equation}
In particular, the left side is independent of $x_{W\setminus W_a}$.
\end{lemma}
\begin{corollary}[Algebraic normalization]\label{lem:algebraic-normalization}
For every finite-state CADMG $\mathcal H=(R,W,E_{\mathcal H})$, every $\theta\in U_{\mathcal H}$ and every $x_W\in\mathfrak X_W$,
\begin{equation}\label{eq:algebraic-normalization}
\sum_{x_R\in\mathfrak X_R}\Phi_{\mathcal H}(\theta)(x_R\mid x_W)=1.
\end{equation}
\end{corollary}
\begin{proposition}[Finite-state recursive-head parameterization]\label{prop:standing-finite-state-import}
Let $\mathcal H=(R,W,E_{\mathcal H})$ be a reduced finite-state CADMG, with a corner state fixed for each random vertex, and let $\Phi_{\mathcal H}$ be given by \eqref{eq:general-forward-map}.
\begin{enumerate}[label=(\roman*)]
\item \emph{Representation.} A strictly positive probability kernel $q_R(\cdot\mid\cdot)$ recursively factorizes according to $\mathcal H$ if and only if $q_R=\Phi_{\mathcal H}(\eta)$ for some $\eta\in U_{\mathcal H}$.
\item \emph{Uniqueness.} If $q_R$ is strictly positive and $\Phi_{\mathcal H}(\eta)=\Phi_{\mathcal H}(\eta')=q_R$, then $\eta=\eta'$.
\item \emph{Identification.} For a strictly positive recursively factorizing $q_R$, let $r_C$ be the unique recursively derived kernel on an intrinsic set $C$, and put $H=H(C)$ and $T=T(C)$. The unique coordinate vector $\Theta_{\mathcal H}(q_R)\in U_{\mathcal H}$ is given by
\begin{equation}\label{eq:cadmg-coordinate-extraction}
\{\Theta_{\mathcal H}(q_R)\}_C(a_H\mid x_T) =\frac{r_C(a_H,x_{C\setminus H}\mid x_{T\setminus C})} {\displaystyle\sum_{z_H\in\mathfrak X_H} r_C(z_H,x_{C\setminus H}\mid x_{T\setminus C})}.
\end{equation}
If $q_R$ is a reachable kernel obtained by fixing a strictly positive nested Markov law on an ADMG, then, by Proposition~\ref{prop:er-equivalence-rrs-compatibility}(iii), the recursively derived $r_C$ equals the corresponding intrinsic fixing kernel $q_C$ in \eqref{eq:fixing-kernel-bridge}; in that setting \eqref{eq:cadmg-coordinate-extraction} is exactly \eqref{eq:coordinate-extraction}.
\item \emph{Ambient smooth recovery.} Fix one recursive derivation scheme for every intrinsic set and reference values for any arguments that the scheme discards on the model. On the open positive orthant
\[
\mathcal O_{\mathcal H} :=(0,\infty)^{\mathfrak X_R\times\mathfrak X_W},
\]
the same finite sequence of marginalizations, conditionalizations and coordinate ratios defines a rational, hence $C^\infty$, map $\widetilde\Theta_{\mathcal H}:\mathcal O_{\mathcal H}\to U_{\mathcal H}$. Its restriction to strictly positive recursively factorizing kernels is $\Theta_{\mathcal H}$. Off the model, its value may depend on the chosen derivation scheme.
\item \emph{Inverse identities.} The two inverse identities are
\begin{equation}\label{eq:standing-import-inverses}
\Phi_{\mathcal H}\{\Theta_{\mathcal H}(q_R)\}=q_R, \qquad \widetilde\Theta_{\mathcal H}\{\Phi_{\mathcal H}(\eta)\}=\eta.
\end{equation}
The first identity holds for every strictly positive recursively factorizing kernel; the second holds whenever $\Phi_{\mathcal H}(\eta)$ is a strictly positive probability kernel.
\end{enumerate}
\end{proposition}
Define its strict-positivity region by
\[
\Omega_{\mathcal G} :=\{\theta\in U_{\mathcal G}: \Phi_{\mathcal G}(\theta)(x_V)>0 ~~ \text{ for every }~~ x_V\in\mathfrak X_V\}.
\]
For $\theta\in\Omega_{\mathcal G}$, Corollary~\ref{lem:algebraic-normalization} and Proposition~\ref{prop:standing-finite-state-import} show that $p_\theta:=\Phi_{\mathcal G}(\theta)$ is a strictly positive recursively factorizing mass function --- hence, by Proposition~\ref{prop:er-equivalence-rrs-compatibility}(i), a nested Markov one; write $P_\theta$ for its law. Theorem \ref{thm:forward-chart} will package the resulting bijection and its smooth differential consequences: under the probability-mass-function identification, $\theta\mapsto P_\theta$ maps $\Omega_{\mathcal G}$ diffeomorphically onto $\mathcal N_+(\mathcal G)$, with inverse $P\mapsto\Theta_{\mathcal G}(P)$. For a CADMG $\mathcal H$ with random vertices $R$ and fixed vertices $W$, set
\begin{equation}\label{eq:cadmg-positive-coordinate-domain}
\Omega_{\mathcal H} :=\{\theta\in U_{\mathcal H}: \Phi_{\mathcal H}(\theta)(x_R\mid x_W)>0 \text{ for every }(x_R,x_W)\in\mathfrak X_R\times\mathfrak X_W\}.
\end{equation}
\subsection{Scores and tangent spaces}
Fix $P\in\mathcal N_+(\mathcal G)$ with mass function $p$, and equip functions on $\mathfrak X_V$ with
\[
\langle f,g\rangle_P:=\mathrm E_P\{f(X_V)g(X_V)\}, \qquad \|f\|_{P,2}:=\langle f,f\rangle_P^{1/2}.
\]
Write $L_2(P)$ for the resulting finite-dimensional Hilbert space and $L_2^0(P)$ for its subspace of \emph{mean-zero} functions. Following the standard pathwise convention \citep[Section~18.1]{kosorok2008introduction}, the tangent space $\mathcal T_P\mathcal N(\mathcal G)$ is the $L_2(P)$-closure of the linear span of scores of regular two-sided paths $\{P_t:|t|<\epsilon\}\subseteq\mathcal N(\mathcal G)$ satisfying $P_0=P$, a regular path $t\mapsto p_t$ being cellwise differentiable at zero with score $s(x_V)=\dot p_0(x_V)/p(x_V)$. On a finite support with $p>0$, cellwise differentiability at zero is equivalent to differentiability in quadratic mean at zero with the same score, since each cell's square root is then differentiable and the quadratic-mean remainder is a finite sum of cellwise remainders; cellwise continuity places every such path inside $\mathcal N_+(\mathcal G)$ after its parameter interval is shortened if necessary, so $\mathcal N(\mathcal G)$ and $\mathcal N_+(\mathcal G)$ have the same tangent space at $P$, and that space is finite-dimensional, so the closure adds no further directions.
Put $\theta_0=\Theta_{\mathcal G}(P)$. Remark~\ref{rem:forward-map-multiaffine} shows cellwise that $\Phi_{\mathcal G}$ is polynomial, hence Fr\'echet differentiable, between finite-dimensional spaces. Throughout, $\mathbb D$ denotes the differential of a map between finite-dimensional spaces and $\mathbb Df[u]$ its value in the direction $u$, which is the ordinary directional derivative; the letter $D$ is reserved for districts, and $D_P$ in Remark~\ref{rem:chart-not-projection} is multiplication by $p$, not a differential. For $u\in U_{\mathcal G}$,
\[
\mathbb D\Phi_{\mathcal G,\theta_0}[u] :=\left.\frac{\mathrm d}{\mathrm dt}\Phi_{\mathcal G}(\theta_0+tu)\right|_{t=0}.
\]
Define
\[
J_Pu(x_V) :=\frac{\mathbb D\Phi_{\mathcal G,\theta_0}[u](x_V)}{p(x_V)}.
\]
Pointwise division by $p$ turns the derivative of the cell probabilities into a log-density derivative. Thus $J_Pu$ is the score of the coordinate path $t\mapsto\Phi_{\mathcal G}(\theta_0+tu)$ whenever that line remains in $\Omega_{\mathcal G}$; Theorem~\ref{thm:forward-chart} supplies a two-sided interval for every $u$. Using the fixed identification $U_{\mathcal G}\cong\mathbb R^d$, regard the canonical map $\iota_C:U_C\to U_{\mathcal G}$ above as coordinate insertion into $\mathbb R^d$. For a linear map $A$, $\operatorname{ran}(A)$ denotes its range. Set
\[
J_{C,P}:=J_P\iota_C, \qquad \mathcal S_C(P):=\operatorname{ran}(J_{C,P}).
\]
Thus $J_{C,P}$ inserts a perturbation in the $C$-tagged coordinate block and maps it to its model-valid score, while $\mathcal S_C(P)$ collects all such scores. These are the operative intrinsic-set score spaces. Membership in them, and in $\mathcal T_P\mathcal N(\mathcal G)$, is decided by the chart and is a different property from observational centering, the vanishing of a conditional mean given an observational conditioning set; the two are kept apart throughout, and Section~\ref{sec:counterexample} exhibits an observationally centered direction that is not a model score.
\section{Chart-induced tangent geometry}\label{sec:geometry}
For a finite-state CADMG $\mathcal H=(R,W,E_{\mathcal H})$, let
\begin{align*}
\mathcal K_{\mathcal H} &:={} \left\{q\in\mathbb R^{\mathfrak X_R\times\mathfrak X_W}: \sum_{x_R}q(x_R\mid x_W)=1\text{ for every }x_W\right\},\\
\mathcal Q_+(\mathcal H) &:={} \left\{q\in\mathcal K_{\mathcal H}:q>0 \text{ and }q\text{ recursively factorizes according to }\mathcal H \right\}.
\end{align*}
Thus $\mathcal K_{\mathcal H}$ is the affine space of normalized real kernel arrays, whereas $\mathcal Q_+(\mathcal H)$ is the set of strictly positive recursively factorizing probability kernels.
\begin{theorem}[Forward finite-state chart]\label{thm:forward-chart}
Let $\mathcal H=(R,W,E_{\mathcal H})$ be a reduced CADMG, and suppose that every $\mathfrak X_v$, $v\in R\cup W$, is finite, with every random-variable state space having at least two elements. Fix the derivation scheme and reference values used to construct $\widetilde\Theta_{\mathcal H}$ in Proposition~\ref{prop:standing-finite-state-import}(iv). Then $\Omega_{\mathcal H}$ is open and
\[
\Phi_{\mathcal H}:\Omega_{\mathcal H}\longrightarrow \mathcal Q_+(\mathcal H)
\]
is a bijection whose inverse is $\Theta_{\mathcal H} =\widetilde\Theta_{\mathcal H}|_{\mathcal Q_+(\mathcal H)}$. Moreover, $\Phi_{\mathcal H}$ is a smooth embedding into $\mathcal K_{\mathcal H}$ with image $\mathcal Q_+(\mathcal H)$. Consequently, with the induced embedded-manifold structure on its image, $\Phi_{\mathcal H}$ and $\Theta_{\mathcal H}$ are mutually inverse $C^\infty$ maps.
For $m\geq1$, $\eta_0\in\Omega_{\mathcal H}$ and $u_1,\ldots,u_m\in U_{\mathcal H}$, there is an $\epsilon>0$ such that
\begin{equation}\label{eq:finite-direction-rectangle}
q_{t_1,\ldots,t_m} =\Phi_{\mathcal H}\!\left(\eta_0+\sum_{j=1}^m t_ju_j\right), \qquad |t_j|<\epsilon,
\end{equation}
is a smooth rectangle of strictly positive recursively factorizing kernels through $\Phi_{\mathcal H}(\eta_0)$. The same $\epsilon$ works simultaneously for every fixed-variable configuration. If $u_1,\ldots,u_m$ are linearly independent, this rectangle is an $m$-dimensional submodel.
In the ADMG case $W=\varnothing$, write $\mathcal H=\mathcal G$ and let $P_\eta$ denote the law with mass function $\Phi_{\mathcal G}(\eta)$. Then $\eta\mapsto P_\eta$ is a diffeomorphism from $\Omega_{\mathcal G}$ onto $\mathcal N_+(\mathcal G)$ under the probability-mass-function identification. We likewise identify the kernel notation $\Theta_{\mathcal G}(p)$ with the law notation $\Theta_{\mathcal G}(P)$ used in Section~\ref{sec:setup}. For every $P\in\mathcal N_+(\mathcal G)$, with mass function $p$ and $\theta_0=\Theta_{\mathcal G}(P)$, the score differential
\[
J_P:U_{\mathcal G}\cong\mathbb R^d\longrightarrow L_2^0(P)
\]
is a linear isomorphism from $U_{\mathcal G}$ onto $\mathcal T_P\mathcal N(\mathcal G)$.
\end{theorem}
Corollary~5.6 of \cite{evans2019smooth} exhibits the strictly positive recursively factorizing kernels as a curved exponential family. Theorem~\ref{thm:forward-chart} is the chart-level form of that smoothness for the general finite-state forward map \eqref{eq:general-forward-map}; what the rest of the paper uses is its last clause, the identification of the tangent space with $\operatorname{ran}(J_P)$, and the next two corollaries are the standard finite-dimensional consequences of a smooth chart, recorded in the form in which they are applied.
\begin{corollary}[Fisher information as the chart pullback]\label{cor:fisher-chart-pullback}
Fix the ordered-coordinate identification $U_{\mathcal G}\cong\mathbb R^d$ and equip it with the Euclidean inner product $\langle u,v\rangle_{\mathrm E}:=u^\top v$. Let $e_1,\ldots,e_d$ be its coordinate basis. For $P=P_{\theta_0}\in\mathcal N_+(\mathcal G)$, let $p=p_{\theta_0}$ be its mass function and define the Fisher information matrix of the chart by
\begin{equation}\label{eq:fisher-chart-matrix}
\{I_{\mathrm F}(\theta_0)\}_{ij} :=\sum_{x_V\in\mathfrak X_V} \frac{\mathbb D\Phi_{\mathcal G,\theta_0}[e_i](x_V)\, \mathbb D\Phi_{\mathcal G,\theta_0}[e_j](x_V)}{p(x_V)} =\langle J_Pe_i,J_Pe_j\rangle_P.
\end{equation}
Then, for all $u,v\in U_{\mathcal G}$,
\[
\langle J_Pu,J_Pv\rangle_P =u^\top I_{\mathrm F}(\theta_0)v, \qquad I_{\mathrm F}(\theta_0)=J_P^*J_P.
\]
Here $J_P^*:L_2^0(P)\to U_{\mathcal G}$ is the adjoint relative to $\langle\cdot,\cdot\rangle_P$ and $\langle\cdot,\cdot\rangle_{\mathrm E}$, and the final equality uses the fixed coordinate identification. In particular, $I_{\mathrm F}(\theta_0)$ is positive definite.
\end{corollary}
\begin{corollary}[Intrinsic-block directness]\label{cor:block-directness}
Under Theorem \ref{thm:forward-chart},
\[
\mathcal T_P\mathcal N(\mathcal G) =\bigoplus_{C\in\mathcal I(\mathcal G)}\mathcal S_C(P), \qquad \dim\mathcal S_C(P)=d_C.
\]
The sum is \emph{algebraically direct}, but its summands need \emph{not} be mutually orthogonal.
\end{corollary}
\subsection{District separation}
For $D\in\mathcal D(\mathcal G)$, define
\begin{equation}\label{eq:district-space}
\begin{aligned}
U_D^{\mathrm{nat}} &:=\bigoplus_{C:\,\delta(C)=D}\iota_C(U_C) \subseteq U_{\mathcal G},\\
d_D^{\mathrm{nat}} &:=\dim U_D^{\mathrm{nat}} =\sum_{C:\,\delta(C)=D}d_C,\\
\mathcal S_D^{\mathrm{nat}}(P) &:=J_P(U_D^{\mathrm{nat}}) =\bigoplus_{C:\,\delta(C)=D}\mathcal S_C(P).
\end{aligned}
\end{equation}
The superscript ``nat'' abbreviates \emph{native}: the grouping is by the districts of the observational graph $\mathcal G$ itself, as opposed to the districts of an intervention-specific active graph $\mathcal H_{\mathsf I}$ introduced in Section~\ref{sec:intervention-bridge}, which may split a native district. The same superscript is used for the coordinate space, score space, dimension, basis and Gram matrix of a native district.
\begin{lemma}[Algebraic district factorization]\label{lem:algebraic-district-factorization}
Let $\mathcal H=(R,W,E_{\mathcal H})$ be a finite-state CADMG. For $D\in\mathcal D(\mathcal H)$ and $\eta\in U_{\mathcal H}$, write
\[
\eta_{[D]}:=\eta|_D \in U_{\mathfrak d_D(\mathcal H)}.
\]
Then every $\eta\in U_{\mathcal H}$, without any positivity assumption, satisfies
\begin{equation}\label{eq:top-level-factorization}
\Phi_{\mathcal H}(\eta)(x_R\mid x_W) =\prod_{D\in\mathcal D(\mathcal H)} r_{D,\eta_{[D]}} \bigl(x_D\mid x_{\operatorname{pa}_{\mathcal H}(D)\setminus D}\bigr), \qquad r_{D,\eta_{[D]}} :=\Phi_{\mathfrak d_D(\mathcal H)}(\eta_{[D]}).
\end{equation}
The tagged coordinate set of the $D$ factor is exactly $\{C\in\mathcal I(\mathcal H):C\subseteq D\}$. Thus no recursive-head coordinate occurs in two distinct factors, although the factors may be signed.
\end{lemma}
We call \eqref{eq:top-level-factorization} the \emph{top-level} district factorization; ``top-level'' is shorthand of this paper for the outermost product over native districts in a recursive factorization and is not terminology of \cite{evans2019smooth}.
\begin{lemma}[Reachable-kernel chart realization]\label{lem:reachable-chart-realization}
For an intrinsic set $C\in\mathcal I(\mathcal G)$, put
\[
\theta_{\downarrow C} :=(\theta_S:S\in\mathcal I(\mathcal G),\ S\subseteq C).
\]
Then, for every $\theta\in\Omega_{\mathcal G}$,
\begin{equation}\label{eq:reachable-chart-realization}
\operatorname{red}_C\{\phi_{V\setminus C}(p_\theta;\mathcal G)\} =\Phi_{\mathcal G[C]}(\theta_{\downarrow C}), \qquad \theta_{\downarrow C}\in\Omega_{\mathcal G[C]}.
\end{equation}
The equality is an identity of kernels, pointwise in every fixed-parent configuration.
\end{lemma}
\begin{proposition}[Districtwise feasibility and score locality]\label{prop:district-locality}
For $\theta\in U_{\mathcal G}$, write $\theta_{[D]}=(\theta_C:C\in\mathcal I(\mathcal G),\ \delta(C)=D)$. Under $U_{\mathcal G}=\bigoplus_DU_D^{\mathrm{nat}}\cong\mathbb R^d$, the exact feasible domain is
\begin{equation}\label{eq:district-domain-product}
\Omega_{\mathcal G} =\prod_{D\in\mathcal D(\mathcal G)}\Omega_{\mathfrak d_D(\mathcal G)}.
\end{equation}
For $\theta_0\in\Omega_{\mathcal G}$, let $P=P_{\theta_0}$. Every factor in \eqref{eq:top-level-factorization} is a strictly positive district kernel lying in $\mathcal Q_+(\mathfrak d_D(\mathcal G))$, the image of its positive recursive-head chart, and, in its canonical parent-indexed form,
\begin{equation}\label{eq:district-factor-as-fixing}
r_{D,\theta_{0,[D]}} =\Phi_{\mathfrak d_D(\mathcal G)}(\theta_{0,[D]}) =\operatorname{red}_D\{\phi_{V\setminus D}(p_{\theta_0};\mathcal G)\}>0.
\end{equation}
Holding the other district blocks of $\theta_0$ fixed and changing only its $U_D^{\mathrm{nat}}$ component changes only this factor. The resulting chart-path scores at $P$ are exactly $\mathcal S_D^{\mathrm{nat}}(P)$, and every such score has a two-sided realizing path.
\end{proposition}
\begin{remark}[Local compatibility is not arbitrary kernel tilting]
Equation~\eqref{eq:district-domain-product} permits independent selection, from the district domains, of district factors lying in $\mathcal Q_+(\mathfrak d_D(\mathcal G))=\Phi_{\mathfrak d_D(\mathcal G)}(\Omega_{\mathfrak d_D(\mathcal G)})$. It does not permit independently chosen arbitrary normalized district kernels: each $r_{D,\theta_{[D]}}$ must lie in that chart image and retain all internal nested constraints. In general, $\Omega_{\mathfrak d_D(\mathcal G)}$ is variation dependent and need not factor over the individual intrinsic-set blocks within $D$. Openness nevertheless gives a Euclidean ball at every interior point, and hence a rectangle for any finite collection of intrinsic- or district-block directions.
\end{remark}
\subsection{Cross-district orthogonality}
\begin{theorem}[Orthogonal district decomposition]\label{thm:district-orthogonality}
For every $P\in\mathcal N_+(\mathcal G)$,
\[
\mathcal T_P\mathcal N(\mathcal G) =\bigoplus_{D\in\mathcal D(\mathcal G)}^{\perp}\mathcal S_D^{\mathrm{nat}}(P).
\]
No orthogonality is asserted between distinct intrinsic-set blocks inside the same district.
\end{theorem}
For the scope statements and diagnostics of this subsection and of Section~\ref{sec:counterexample}, the \emph{district quotient} of $\mathcal G$ is the directed graph whose vertex set is $\mathcal D(\mathcal G)$, with an arrow $D\to D'$ whenever $D\ne D'$ and some $v\in D$ is a parent of some $v'\in D'$ in $\mathcal G$. Although the directed part of an ADMG is acyclic, its district quotient may contain a directed cycle; it is a diagnostic device only and enters neither the chart nor the argument above.
\begin{remark}[Cyclic-quotient diagnostic]\label{rem:cyclic-quotient-diagnostic}
For an intrinsic set $C$, define the observationally centered head-family space
\[
\mathcal L_C^{\mathrm{obs}}(P) :=\left\{g(X_{H(C)\cup T(C)}): \mathrm E_P\{g\mid X_{T(C)}\}=0\right\}.
\]
At the law $P_{\mathrm{cyc}}$ of the cyclic-quotient control (Section~\ref{sec:cyclic-control}; specified in Section~\ref{app:B-cyclic} of the Supplementary Material), the sum of these spaces over $\mathcal I(\mathcal G_{\mathrm{cyc}})$ has dimension $11$ and contains the ten-dimensional tangent space $\mathcal T_{P_{\mathrm{cyc}}}\mathcal N(\mathcal G_{\mathrm{cyc}})$ strictly, with an exact rank certificate in Section~\ref{app:B-cyclic}. The two directions of Section~\ref{sec:cyclic-control} lie in such spaces for different districts and outside the tangent space, so cross-products between these raw spaces say nothing about Theorem~\ref{thm:district-orthogonality}. This $11$-versus-$10$ comparison is distinct from the $30$-versus-$23$ comparison in Section~\ref{sec:counterexample}, where the larger space is the score space of arbitrary normalized $q_\Delta(\cdot\mid W)$ paths, centered under $q_\Delta$ separately for each $W=w$.
\end{remark}
\subsection{Explicit metric projection}
Fix $P\in\mathcal N_+(\mathcal G)$. For a closed linear subspace $\mathcal S\subseteq L_2(P)$, write $\Pi_{\mathcal S}f$ for the unique $L_2(P)$-orthogonal projection of $f$ onto $\mathcal S$, equivalently the unique minimizer of $g\mapsto\|f-g\|_{P,2}$ over $g\in\mathcal S$.
For each $D\in\mathcal D(\mathcal G)$, choose a basis $\{e_{Dj}^{\mathrm{nat}}:1\leq j\leq d_D^{\mathrm{nat}}\}$ of $U_D^{\mathrm{nat}}$ and define the column vector of score functions
\[
b_D^{\mathrm{nat}}(X_V) :=\bigl(J_Pe_{Dj}^{\mathrm{nat}}(X_V): 1\leq j\leq d_D^{\mathrm{nat}}\bigr)^\top, \qquad G_D^{\mathrm{nat}}(P) :=\mathrm E_P\{b_D^{\mathrm{nat}}(b_D^{\mathrm{nat}})^\top\}.
\]
\begin{proposition}[District Gram projection]\label{prop:gram-projection}
With these choices, each $G_D^{\mathrm{nat}}(P)$ is positive definite and, for every $f\in L_2(P)$,
\begin{equation}\label{eq:district-gram-projection}
\Pi_{\mathcal S_D^{\mathrm{nat}}(P)}f =(b_D^{\mathrm{nat}})^\top \{G_D^{\mathrm{nat}}(P)\}^{-1}\mathrm E_P(b_D^{\mathrm{nat}}f).
\end{equation}
For the full tangent space,
\begin{equation}\label{eq:global-gram-projection}
\Pi_{\mathcal T_P\mathcal N(\mathcal G)}f =\sum_{D\in\mathcal D(\mathcal G)}(b_D^{\mathrm{nat}})^\top \{G_D^{\mathrm{nat}}(P)\}^{-1}\mathrm E_P(b_D^{\mathrm{nat}}f).
\end{equation}
\end{proposition}
The full within-district Gram matrix is, in general, required: $\Pi_{\mathcal S_D^{\mathrm{nat}}(P)}f$ is not the sum of separate projections onto the generally nonorthogonal intrinsic-block spaces $\mathcal S_C(P)$ with $\delta(C)=D$, because within-district cross-block Gram entries need not vanish; an explicit direction at the law $P_{\mathrm{cyc}}$ of Section~\ref{app:B-cyclic} of the Supplementary Material has different joint and blockwise projections, so the two operators differ there.
\begin{remark}[Chart synthesis is not projection]\label{rem:chart-not-projection}
Recovering chart coordinates from a perturbed law and synthesizing a model score from them is a retraction onto the tangent space and need not equal the orthogonal projection of Proposition~\ref{prop:gram-projection}. Let $p$ be the mass function of $P\in\mathcal N_+(\mathcal G)$ and put $\theta_0=\Theta_{\mathcal G}(P)$. Because $p>0$, multiplication by $p$ is a linear isomorphism $D_P:L_2^0(P)\to T_p\mathcal K_{\mathcal G}$, $D_Pf:=pf$, onto the tangent space $T_p\mathcal K_{\mathcal G}=\{a\in\mathbb R^{\mathfrak X_V}:\sum_{x_V}a(x_V)=0\}$ of the affine normalization space at $p$. For a derivation scheme $\sigma$ --- one complete recursive derivation per intrinsic set, with reference states for the arguments the scheme discards --- write $\widetilde\Theta^{\sigma}_{\mathcal G}$ for the ambient recovery map of Proposition~\ref{prop:standing-finite-state-import}(iv) built from $\sigma$; the selected default scheme is $\sigma_0$. Define
\[
\mathcal R^{\sigma}_P :=J_P\circ \mathbb D\widetilde\Theta^{\sigma}_{\mathcal G,p}\circ D_P: L_2^0(P)\longrightarrow\mathcal T_P\mathcal N(\mathcal G);
\]
and write $\mathcal R_P:=\mathcal R^{\sigma_0}_P$ for its default-scheme instance. If $f=J_Pu\in\mathcal T_P\mathcal N(\mathcal G)$, then $D_Pf=\mathbb D\Phi_{\mathcal G,\theta_0}[u]$, and the left-inverse identity \eqref{eq:left-inverse-differential} of the Supplementary Material gives $\mathbb D\widetilde\Theta^{\sigma}_{\mathcal G,p}\mathbb D\Phi_{\mathcal G,\theta_0}[u]=u$ for every scheme, because $\widetilde\Theta^{\sigma}_{\mathcal G}$ restricts to $\Theta_{\mathcal G}$ on the model (Proposition~\ref{prop:standing-finite-state-import}(iv)--(v)); hence $\mathcal R^{\sigma}_Pf=f$, so $\mathcal R^{\sigma}_P$ is idempotent with range $\mathcal T_P\mathcal N(\mathcal G)$: a generally oblique, scheme-dependent retraction. Tangent spaces and metric projections are scheme-invariant, and only the orthogonal projection \eqref{eq:global-gram-projection} is canonical for efficiency. That the two operators differ is shown by exact rational computation at the cyclic-quotient law $P_{\mathrm{cyc}}$: for the direction $f_1$ of \eqref{eq:cyclic-raw-directions} and the derivation scheme $\sigma_*$ recorded in Section~\ref{app:B-cyclic}, $\mathcal R^{\sigma_*}_{P_{\mathrm{cyc}}}f_1\ne\Pi_{\mathcal T_{P_{\mathrm{cyc}}}\mathcal N(\mathcal G_{\mathrm{cyc}})}f_1$, with the squared difference recorded there.
\end{remark}
\section{Exact non-mb-shielded counterexample}\label{sec:counterexample}
Consider the binary ADMG whose directed and bidirected edges are
\begin{equation}\label{eq:verma-counterexample-graph}
\begin{aligned}
\text{directed: }&W\to Z,\quad B\to Z,\quad Z\to Y,\\
\text{bidirected: }&A\leftrightarrow B,\quad A\leftrightarrow Z,\quad B\leftrightarrow Y.
\end{aligned}
\end{equation}
Its native districts are $\{W\}$ and $\Delta=\{A,B,Z,Y\}$, with quotient $\{W\}\to\Delta$; see Figure~\ref{fig:verma-counterexample-graph}. With the Markov blanket $\operatorname{mb}_{\mathcal G}(v)$ of Section~\ref{sec:setup}, an ADMG is \emph{mb-shielded} when any two vertices are adjacent whenever either belongs to the other's Markov blanket \citep[Theorem~2]{bhattacharya2022semiparametric}. The displayed graph is not mb-shielded: $W$ and $A$ are nonadjacent although $W\in\operatorname{mb}_{\mathcal G}(A)=\{B,Z,Y,W\}$.
On strictly positive laws, its nested Markov model imposes no restriction beyond ordinary conditional independence. The reachable $\{B,Y\}$ kernel $q_{BY,p}(b,y\mid z,w)=p(b\mid w)\,p(y\mid b,z,w)$, obtained by fixing $A$, $Z$ and $W$, is invariant in $w$ at every model law --- equivalently $F_{BY}(p)=0$ in \eqref{eq:BY-constraint-map} of the Supplementary Material --- and that invariance holds exactly when $B\mathrel{\perp\!\!\!\perp} W$ and $Y\mathrel{\perp\!\!\!\perp} W\mid(B,Z)$; Appendix~\ref{app:prototype-ledger} shows that $\mathcal N_+(\mathcal G)$ is the strictly positive ordinary Markov model $\{p>0:\ W\mathrel{\perp\!\!\!\perp}(A,B),\ Y\mathrel{\perp\!\!\!\perp} W\mid(B,Z)\}$ of $\mathcal G$, in which the joint independence $W\mathrel{\perp\!\!\!\perp}(A,B)$ is part of the model and is not implied by the kernel invariance. No DAG on $V$ represents this model (Section~\ref{app:B-primary}). Adding the single arrow $W\to B$ produces a genuine Verma restriction; see Remark~\ref{rem:verma-variant}.
The example matters because $\mathcal G$ is not mb-shielded: the mb-shielded decomposition theorem \citep[Theorem~2]{bhattacharya2022semiparametric} does not apply to it, whereas Theorem~\ref{thm:district-orthogonality} does, and the observationally centered direction $h$ defined next is not a model score.
\begin{figure}[t]
\centering
\begin{tikzpicture}[
-Latex, semithick,
state/.style={circle, draw, minimum width=0.62cm, inner sep=1pt},
bidirected/.style={Latex-Latex, dashed}]
\node[state] (W) at (0,0) {$W$};
\node[state] (Z) at (2,0) {$Z$};
\node[state] (Y) at (4,0) {$Y$};
\node[state] (A) at (1,1.35) {$A$};
\node[state] (B) at (3,1.35) {$B$};
\draw (W) -- (Z);
\draw (B) -- (Z);
\draw (Z) -- (Y);
\draw[bidirected] (A) -- (B);
\draw[bidirected] (A) -- (Z);
\draw[bidirected] (B) -- (Y);
\end{tikzpicture}
\caption{The binary ADMG of \eqref{eq:verma-counterexample-graph}. Solid arrows are directed edges; dashed double arrows are bidirected edges.}\label{fig:verma-counterexample-graph}
\end{figure}
Let $P$ be the strictly positive binary law specified by \eqref{eq:verma-rational-sem} in Appendix~\ref{app:prototype-ledger} of the Supplementary Material, and define
\[
h(X_V):=X_Y-\mathrm E_P(X_Y\mid X_Z).
\]
By construction, $\mathrm E_P(h\mid X_Z)=0$.
\begin{example}[Exact failure of the raw direction]\label{prop:raw-failure}
At the law $P$, $h\notin\mathcal T_P\mathcal N(\mathcal G)$ and $\|h-\Pi_{\mathcal T_P\mathcal N(\mathcal G)}h\|_{P,2}^2=7.84\times10^{-6}>0$, and the residual has the closed form $h-\Pi_{\mathcal T_P\mathcal N(\mathcal G)}h=0.0028\,(2W-1)(2B-1)$: at this law $\mathrm P_P(W{=}1)=\mathrm P_P(B{=}1)=0.5$, the restriction $B\mathrel{\perp\!\!\!\perp} W$ makes $(2W-1)(2B-1)$ orthogonal to every model score, and $\langle h,(2W-1)(2B-1)\rangle_P=0.0028\ne0$ certifies non-tangency. The exact projection is verified in Appendix~\ref{app:proofs-counterexample} of the Supplementary Material; Section~\ref{app:B-primary} records the relative residual at $P$ and verifies non-tangency of the corresponding observational residual $Y-\mathrm E_{P'}(Y\mid Z)$ at a second law $P'$.
\end{example}
\subsection{Intrinsic-kernel versus observational centering}
The intrinsic kernel of the singleton set $\{Y\}$ is $q_{\{Y\}}(y\mid z)$, obtained by fixing the other four vertices, for instance in the valid order $A,Z,W,B$. At the terminal stage, where $Y$ is the only random vertex, this kernel is also the fixing divisor; the divisor for fixing $Y$ in the original graph is instead $p(y\mid a,b,z,w)$. At $P$, $q_{\{Y\}}(0\mid 0)=0.55$ and $\mathrm P_P(Y=0\mid Z=0)=0.718$. Both values follow from \eqref{eq:verma-rational-sem} of the Supplementary Material, as recorded in Section~\ref{app:B-primary}. Thus, although $Y$ is a directed sink and can be placed last in a topological order, its intrinsic kernel differs from the observational conditional distribution. In particular, the observationally centered direction $h=Y-\mathrm E_P(Y\mid Z)$ satisfies $\sum_{y=0}^1 q_{\{Y\}}(y\mid 0)\{y-\mathrm E_P(Y\mid Z=0)\}=0.45-0.282=0.168\ne0$. This calculation shows why $P$-conditional centering cannot be transferred directly to intrinsic-kernel centering.
\subsection{The 30-versus-23 gap}
For any $P\in\mathcal N_+(\mathcal G)$, let $\theta_P:=\Theta_{\mathcal G}(P)$ and $q_{\Delta,P}:=\operatorname{red}_\Delta\{\phi_{V\setminus\Delta}(p;\mathcal G)\}$. Define the normalized district-kernel score space
\begin{equation}\label{eq:verma-normalized-district-score-space}
\mathcal K_\Delta(P) :=\left\{g\in L_2(P): \sum_{x_\Delta}q_{\Delta,P}(x_\Delta\mid w) g(x_\Delta,w)=0\quad\text{for every }w\right\}.
\end{equation}
We first establish the structural containment
\begin{equation}\label{eq:verma-district-score-containment}
\mathcal S_\Delta^{\mathrm{nat}}(P)\subseteq\mathcal K_\Delta(P).
\end{equation}
Indeed, take $g=J_Pu$ with $u\in U_\Delta^{\mathrm{nat}}$ and use the two-sided chart path $p_t=\Phi_\mathcal G(\theta_P+tu)$ for sufficiently small $|t|$. By Proposition~\ref{prop:district-locality}, this path holds the $\{W\}$ factor fixed and changes only the normalized district kernel:
\[
p_t(w,x_\Delta)=p(w)q_{\Delta,t}(x_\Delta\mid w), \qquad g(x_\Delta,w) =\left.\frac{\mathrm d}{\mathrm dt}\log q_{\Delta,t}(x_\Delta\mid w)\right|_{t=0}.
\]
Differentiating $\sum_{x_\Delta}q_{\Delta,t}(x_\Delta\mid w)=1$ separately for each $w$ gives \eqref{eq:verma-district-score-containment}. This argument uses a model-valid chart path; it does not assert that an arbitrary normalized kernel tilt lies in $\mathcal Q_+(\mathfrak d_\Delta(\mathcal G))=\Phi_{\mathfrak d_\Delta(\mathcal G)}(\Omega_{\mathfrak d_\Delta(\mathcal G)})$, the image of the positive recursive-head chart of the $\Delta$ district.
Conversely, every $g\in\mathcal K_\Delta(P)$ is realized by the local kernel path $q_{\Delta,t}:=q_{\Delta,P}(1+tg)$: the centering condition in \eqref{eq:verma-normalized-district-score-space} keeps $q_{\Delta,t}$ normalized for every $w$, finiteness and strict positivity keep it positive for all sufficiently small two-sided $t$, and $(\mathrm d/\mathrm dt)\log q_{\Delta,t}|_{t=0}=g$. Thus $\mathcal K_\Delta(P)$ is exactly the tangent space of unrestricted normalized district kernels at $q_{\Delta,P}$; it generally exceeds the tangent space of the chart image $\mathcal Q_+(\mathfrak d_\Delta(\mathcal G))$.
The dimensions make the containment strict at every positive model law. For each of the two values of $w$, the positive kernel $q_{\Delta,P}(\cdot\mid w)$ imposes one nonzero linear constraint on the 16-dimensional space of functions of $x_\Delta$. Therefore $\dim\mathcal K_\Delta(P)=2(2^4-1)=30$.
On the other hand, $\mathcal G$ has nine intrinsic sets, one of which is $\{W\}$; the eight in the $\Delta$ block, with their recursive heads, tails and binary block dimensions $d_C=2^{|T(C)|}$, are tabulated in Section~\ref{app:B-primary} of the Supplementary Material, and their dimensions sum to $23$. Thus, using the intrinsic-set-indexed definition $d_\Delta^{\mathrm{nat}}=\sum_{C:\delta(C)=\Delta}d_C$ from \eqref{eq:district-space} and injectivity of $J_P$, $\dim\mathcal S_\Delta^{\mathrm{nat}}(P)=23$.
Combining \eqref{eq:verma-district-score-containment} with the dimension count just displayed proves, for every $P\in\mathcal N_+(\mathcal G)$,
\begin{equation}\label{eq:verma-district-codimension-seven}
\mathcal S_\Delta^{\mathrm{nat}}(P)\subsetneq\mathcal K_\Delta(P), \qquad \dim\!\left(\mathcal K_\Delta(P)/ \mathcal S_\Delta^{\mathrm{nat}}(P)\right)=7.
\end{equation}
The restriction $W\mathrel{\perp\!\!\!\perp}(A,B)$ contributes $(2-1)(4-1)=3$ scalar constraints, and $Y\mathrel{\perp\!\!\!\perp} W\mid(B,Z)$ contributes $4(2-1)(2-1)=4$. Their differentials are independent on the simplex tangent at every positive model law: the four conditional constraints have independent differentials along perturbations of $p(y\mid w,a,b,z)$, which leave the $(W,A,B)$ margin fixed, and the three marginal constraints have independent differentials along perturbations of that margin. Thus $31-7=24$, the chart dimension recorded in Section~\ref{app:B-primary} of the Supplementary Material.
Hence kernel normalization is necessary for a valid district path but is not sufficient: the quotient in \eqref{eq:verma-district-codimension-seven} contains seven linearly independent normalized-kernel directions that are excluded by the linearized Markov restrictions of $\mathcal G$ at $P$, here the seven ordinary-independence constraints $W\mathrel{\perp\!\!\!\perp}(A,B)$ and $Y\mathrel{\perp\!\!\!\perp} W\mid(B,Z)$. This does not by itself rule out every alternative factorized nuisance representation; see the discussion closing Section~\ref{sec:targeted}. This is a structural statement, not a feature of either rational law used below.
\subsection{What is structural and what is evidence}
The containment and codimension-seven conclusion above hold at every \(P\in\mathcal N_+(\mathcal G)\). The exact calculations at the primary law $P$ and at the second active-latent law $P'$ of \eqref{eq:verma-active-latent-sem} of the Supplementary Material, recorded in Section~\ref{app:B-primary}, are finite-law corroborations of Theorems~\ref{thm:forward-chart} and~\ref{thm:district-orthogonality}, not proofs for arbitrary graphs or arbitrary positive laws; properties holding at one law but not the other are not used in the theory.
\subsection{Cyclic district quotients}\label{sec:cyclic-control}
Let $\mathcal G_{\mathrm{cyc}}$ be the binary ADMG on $\{A,A',B,B'\}$ with bidirected edges $A\leftrightarrow A'$ and $B\leftrightarrow B'$ and directed edges $A\to B$ and $B'\to A'$: its directed vertex graph is acyclic, but its two native districts form a directed two-cycle in the district quotient (Section~\ref{app:B-cyclic} of the Supplementary Material: the edge-list display \eqref{eq:cyclic-counterexample-graph}, Figure~\ref{fig:cyclic-counterexample-graph}, and the detailed development). At the strictly positive rational law $P_{\mathrm{cyc}}$ of \eqref{eq:cyclic-rational-coordinates}, all $25$ entries of the cross-district Gram block vanish exactly \eqref{eq:cyclic-cross-district-gram-zero}, as Theorem~\ref{thm:district-orthogonality} requires, whereas the two observationally centered directions of \eqref{eq:cyclic-raw-directions} have a nonzero inner product and strictly positive exact squared projection residuals \eqref{eq:cyclic-raw-direction-failure}; Remark~\ref{rem:cyclic-quotient-diagnostic} records the head-family comparison at the same law. The law is thus a positive control for Theorem~\ref{thm:district-orthogonality} and a negative control for observational centering as a tangency criterion.
\subsection{A structurally misleading pairwise diagnostic}\label{sec:three-district-raw-control}
Let $\mathcal G_3$ be the six-node control of \eqref{eq:three-district-control-graph}, whose three bidirected pairs form a directed three-cycle in the district quotient although the directed vertex graph is acyclic (Section~\ref{app:B-three} of the Supplementary Material: Figure~\ref{fig:three-district-control-graph} and the detailed development). For the raw residuals of \eqref{eq:three-district-raw-residuals}, the pairwise orthogonality \eqref{eq:three-district-pairwise-orthogonality} is \emph{structural}: it holds at every $P\in\mathcal N_+(\mathcal G_3)$, through the $m$-separation global Markov property. Non-tangency is \emph{law-specific}: at the rational law $P_3$ of \eqref{eq:three-district-rational-coordinates}, the three exact squared projection residuals \eqref{eq:three-district-projection-residuals} are strictly positive and the third moment \eqref{eq:three-district-third-moment} is nonzero. The projection residuals, not the third moment, certify the non-tangency claims: the graph-implied conditional independences together with observational centering force the displayed pairwise inner products to vanish even though the directions are not model scores.
\section{A chart-level intervention bridge}\label{sec:intervention-bridge}
This section separates the causal identification input, which is used only at the true causal law, from the statistical geometry of Sections~\ref{sec:setup}--\ref{sec:geometry}: the same fixing functional is defined on all of $\mathcal N_+(\mathcal G)$ and shown to be a smooth normalized chart functional there. Remark~\ref{rem:statistical-not-causal-extension} states the distinction precisely.
\subsection{The causal identification input}
Let $Y\subseteq V$ be the outcome set. For a fixed node intervention, let $A_{\mathsf I}$ be the intervened vertices. For a fixed edge intervention $\alpha$, let $A_{\mathsf I}$ be the sources of its intervened arrows; for a path-specific intervention of class (I3) below, let it be the treatment-source set. Assume throughout that $Y\cap A_{\mathsf I}=\varnothing$. We call $\alpha$ \emph{complete-source} if, once an arrow out of $a\in A_{\mathsf I}$ is included, every directed arrow out of $a$ is included in $\alpha$. This is the convention under which Theorem~1 of \cite{shpitser2018identification} is stated. Put
\begin{equation}\label{eq:active-ancestral-set}
R_{\mathsf I} :=\operatorname{an}_{\mathcal G_{V\setminus A_{\mathsf I}}}(Y), \qquad \mathcal H_{\mathsf I}:=\mathcal G_{R_{\mathsf I}}.
\end{equation}
For a system-wide response, $Y=V\setminus A_{\mathsf I}$ and hence $R_{\mathsf I}=V\setminus A_{\mathsf I}$.
The intervention classes considered are the following; the common configuration conditions above apply to all of them.
\begin{enumerate}[label=(I\arabic*),leftmargin=4em]
\item A fixed-value node intervention.
\item A complete-source fixed-value edge intervention: every source $a\in A_{\mathsf I}$ assigns one common value $a_{a,D}\in\mathfrak X_a$ to all arrows from $a$ into any one district $D\in\mathcal D(\mathcal H_{\mathsf I})$, and values may differ across active districts. Its causal model is stated in \textup{(H2)} and its identification hypotheses in \textup{(H3)}--\textup{(H4)}.
\item A path-specific intervention, by one of two routes.
\emph{Direct route.} A standard path-specific intervention along specified proper causal paths from a treatment set to $Y$, with two fixed treatment-value vectors, treated under the causal model of \textup{(H2)} and the route-specific hypotheses of \textup{(H3)}--\textup{(H4)}; these are the cases covered directly by \cite[Theorems~3--4]{shpitser2013counterfactual}.
\emph{Reduced route.} A path intervention that an independently justified edge-consistent reduction turns into an intervention satisfying every condition in \textup{(I2)}, which is then treated under every \textup{(I2)} hypothesis; the requirements on its path set and assignment are stated in \textup{(H4)}, and the cited DAG results and their limits are stated immediately after the hypotheses.
\end{enumerate}
In the boundary notation below, $a_{w,D}$ means the node value $a_w$ in (I1), the common district-compatible edge value in (I2), and the active or reference value assigned to district $D$ by the path formula of (I3). In every case we require
\begin{equation}\label{eq:active-intrinsicness}
\mathcal D(\mathcal H_{\mathsf I})\subseteq\mathcal I(\mathcal G).
\end{equation}
For $D\in\mathcal D(\mathcal H_{\mathsf I})$, define the boundary configuration $c_D^{\mathsf I}(x_{R_{\mathsf I}})$ on $\operatorname{pa}_{\mathcal G}(D)\setminus D$ by
\[
c_{D,w}^{\mathsf I}(x_{R_{\mathsf I}})
:= \begin{cases}
x_w, & w\in R_{\mathsf I},\\
a_{w,D}, & w\in A_{\mathsf I}.
\end{cases}
\]
Ancestrality in \eqref{eq:active-ancestral-set} places every non-source parent of an active district in $R_{\mathsf I}$, and the class-specific assignment --- the node value in (I1), the complete-source common value in (I2), and the path-formula value in (I3) --- fixes every remaining parent, so these are the only two possibilities; Appendix~\ref{app:proofs-bridge} of the Supplementary Material records why this closure is a consequence here rather than a condition attributed to Theorem~1 of \cite{shpitser2018identification}.
\paragraph*{Hypotheses of the causal identification assertion}
The causal content of Proposition~\ref{prop:causal-identification-interface} --- that the configured product identifies a causal response --- is asserted under the following hypotheses.
\begin{enumerate}[label=(H\arabic*),leftmargin=4em]
\item $\mathcal G$ is the latent projection of a hidden-variable causal DAG and the observed law $P$ is strictly positive; together with the causal model in \textup{(H2)}, this places $P$ in $\mathcal N_+(\mathcal G)$ and makes every fixing divisor below positive.
\item $\mathsf I$ belongs to one of the classes \textup{(I1)}--\textup{(I3)}, $Y\cap A_{\mathsf I}=\varnothing$, and $P$ is generated by the causal model attached to the class: the hidden-variable causal DAG model under which \cite[Theorem~48]{richardson2023nested} is stated for \textup{(I1)}; Pearl's functional model with mutually independent response-function families for \textup{(I2)} \citep[Theorem~1]{shpitser2018identification}; the nonparametric structural equation model with independent errors for the direct branch of \textup{(I3)} \citep[Supplement~A.2]{shpitser2013counterfactual}; the branch of \textup{(I3)} reduced to \textup{(I2)} inherits every \textup{(I2)} hypothesis.
\item Active intrinsicness \eqref{eq:active-intrinsicness} holds. In \textup{(I2)}, and hence in the reduced branch of \textup{(I3)}, every district of $\mathcal G_{V\setminus A_{\mathsf I}}$ is intrinsic in $\mathcal G$ as well, a condition that coincides with \eqref{eq:active-intrinsicness} in the system-wide case $Y=V\setminus A_{\mathsf I}$ (Appendix~\ref{app:proofs-bridge} of the Supplementary Material). In the direct branch of \textup{(I3)}, no recanting district exists and the total-effect district factors required by the cited path-specific result are identified, which \eqref{eq:active-intrinsicness} supplies through \cite[Theorem~48]{richardson2023nested}.
\item The class-specific compatibility holds: complete-source assignment with one common value $a_{a,D}\in\mathfrak X_a$ per source--district pair in \textup{(I2)}; the same within-district coherence and a $P$-independent value in $\mathfrak X_w$ for every excluded-source tail coordinate in \textup{(I3)}; in every class every boundary value satisfies $a_{w,D}\in\mathfrak X_w$. In the direct branch of \textup{(I3)}, the path set is proper, live and consistent for $Y$. In the reduced branch of \textup{(I3)}, the path set is proper, $Y$-live and $Y$-consistent with an edge-consistent assignment in the formal sense of \cite{shpitser2016causal}, the equality with the induced intervention of class \textup{(I2)} is independently justified, and every \textup{(I2)} hypothesis --- its causal model, the complete-source convention, the common assignment, and the intrinsicness of every district of $\mathcal G_{V\setminus A_{\mathsf I}}$ --- is inherited.
\end{enumerate}
The cited results are stated on a DAG: Section~5.3 of \cite{shpitser2016causal} defines the edge-consistent path intervention, its Lemma~5.7 gives the equality with the induced edge intervention, and its Corollary~5.2 the identification result, all under the multiple-world model on a DAG; no general observed-ADMG path-to-edge reduction is inferred from those underlying-DAG statements.
The causal interpretation of Proposition~\ref{prop:causal-identification-interface} uses the class-specific identification hypotheses \textup{(H1)}--\textup{(H4)} at the causal law. The statistical extension that follows, from Theorem~\ref{thm:normalized-intervention-bridge} onward, ranges over every law in $\mathcal N_+(\mathcal G)$ for the fixed graph and intervention configuration --- strict positivity remains part of its domain --- and requires neither a latent realization nor a causal reading of nearby laws (Remark~\ref{rem:statistical-not-causal-extension}). Its operative graphical and configuration conditions are active intrinsicness \eqref{eq:active-intrinsicness}, the first clause of \textup{(H3)}, and the fixed, compatible, single-valued boundary selection $c_D^{\mathsf I}$ with $a_{w,D}\in\mathfrak X_w$, the boundary clause of \textup{(H4)}; together they define the selection map $\operatorname{Sel}_{\mathsf I}$ below.
\begin{proposition}[Configured product and identified response]\label{prop:causal-identification-interface}
Under \textup{(H1)}--\textup{(H4)}, with $p$ the mass function of $P$ and
\[
q_{D,P}:=\operatorname{red}_D\{\phi_{V\setminus D}(p;\mathcal G)\},
\]
the array
\begin{equation}\label{eq:imported-intervention-product}
p^{\mathsf I}(x_{R_{\mathsf I}}) =\prod_{D\in\mathcal D(\mathcal H_{\mathsf I})} q_{D,P}\!\left(x_D\mid c_D^{\mathsf I}(x_{R_{\mathsf I}})\right)
\end{equation}
is a normalized strictly positive law on $R_{\mathsf I}$, by Theorem~\ref{thm:normalized-intervention-bridge} below, and
\textup{(a)} its $Y$ margin, obtained by summing over $R_{\mathsf I}\setminus Y$, is the causal response identified by the cited result of the class;
\textup{(b)} in \textup{(I1)}--\textup{(I2)} the whole array is the identified joint response law on $R_{\mathsf I}$, whereas in \textup{(I3)} only the $Y$ margin is asserted to be causal and no joint counterfactual reading of \eqref{eq:imported-intervention-product} is claimed.
The cited identification theorems, including their target-set applications, are documented in Appendix~\ref{app:proofs-bridge} of the Supplementary Material.
\end{proposition}
\begin{example}[Why a mixed source boundary is not identified]\label{ex:split-boundary-nonidentification}
Consider $A\to B$, $A\to B'$, and $B\leftrightarrow B'$ with $A\sim\operatorname{Bernoulli}(p)$, $0<p<1$. Section~\ref{app:B-intervention} of the Supplementary Material specifies two latent-variable models for this graph that induce the same strictly positive observed law. Setting only the arrow $A\to B$ to zero while $A\to B'$ retains the natural value of $A$ asks for the joint law of the split responses $B(0)$ and $B'(1)$ --- $B$ with the arrow $A\to B$ set to $0$ and $B'$ with the arrow $A\to B'$ at the natural value --- given the natural value $A=1$, the counterfactual that a kernel indexed by the mixed boundary $(a_{A\to B},a_{A\to B'})=(0,1)$ would have to represent. Write $K_i(a):=\operatorname{Law}_i\{B(0),B'(a)\mid A=a\}$ for the natural-retention response in model $i$. The two models agree on every observed-kernel slice $\mathrm P(B=b,B'=b'\mid A=a)$, but $K_1(1)$ and $K_2(1)$ are at total-variation distance $0.64$, while $K_1(0)=K_2(0)$; averaging over the natural draw of $A$ leaves the two split-response laws at distance $0.64p>0$. Setting both arrows to zero gives the same law in the two models. The split intervention deliberately violates the complete-source condition of (I2): it shows why enlarging the theorem to such mixed boundaries would be invalid. Thus a kernel boundary that asks one source coordinate to be simultaneously natural and intervened within one district need not be a functional of the observed law.
\end{example}
\subsection{Active-coordinate selection and normalized reassembly}
Fix an intervention of one of the classes (I1)--(I3) and abbreviate $R=R_{\mathsf I}$, $\mathcal H=\mathcal H_{\mathsf I}$, and $A=A_{\mathsf I}$. Every intrinsic set $S$ of $\mathcal H$ lies in a unique $D\in\mathcal D(\mathcal H)$. By \eqref{eq:active-intrinsicness}, $D$ is reachable in $\mathcal G$. Lemma~4.9 of \cite{evans2019smooth}, applied within $\mathcal H$, first identifies $S$ as intrinsic in $\mathfrak d_D(\mathcal H)$. The CADMGs $\mathcal G[D]$ and $\mathfrak d_D(\mathcal H)$ have exactly the same directed and bidirected edges among their random vertices $D$. Fixability, intrinsic subsets, and recursive heads depend only on this random graph. Hence $S$ is intrinsic in $\mathcal G[D]$; applying the same lemma in $\mathcal G$ then identifies $S$ as a global intrinsic set with the same recursive head. Moreover,
\begin{equation}\label{eq:active-tail-restriction}
T_{\mathcal H}(S) =T_{\mathcal G}(S)\cap R =T_{\mathcal G}(S)\setminus A,
\end{equation}
because any $w\in T_{\mathcal G}(S)\setminus A$ is either internal to $S$ or has an arrow into $S$. The internal case gives $w\in S\subseteq R$ immediately; in the external case, appending that arrow to an $S$-to-$Y$ directed path in $\mathcal G_{V\setminus A}$ gives $w\in R$. The reverse inclusion follows from $\mathcal H=\mathcal G_R$. Define the linear \emph{tail-slice map} $\operatorname{Sel}_{\mathsf I}:U_{\mathcal G}\to U_{\mathcal H}$ by
\begin{equation}\label{eq:tail-slice-map}
\bigl(\operatorname{Sel}_{\mathsf I}\theta\bigr)_S (a_{H(S)}\mid x_{T_{\mathcal H}(S)}) :=\theta_S\!\left( a_{H(S)}\mid x_{T_{\mathcal H}(S)}, a_{T_{\mathcal G}(S)\cap A,D} \right), \qquad S\subseteq D.
\end{equation}
District compatibility makes the selected source-tail entry single-valued. Here $U_{\mathcal H}=\mathbb R^{\mathcal A_{\mathcal H}}$ is the tagged active-chart coordinate space defined in Section~\ref{sec:setup}. Relative to the displayed coordinate bases, $\operatorname{Sel}_{\mathsf I}$ is represented by a $0/1$ row-selection matrix: every output row contains exactly one $1$, and every input column contains at most one $1$. Thus it selects already existing finite-state coordinate entries; it neither extrapolates a kernel nor changes a coordinate value.
\begin{lemma}[Configured factor reassembly]\label{lem:configured-factor-reassembly}
For every $\theta\in\Omega_{\mathcal G}$ and $D\in\mathcal D(\mathcal H)$,
\begin{align}
&\operatorname{red}_D\{\phi_{V\setminus D}(p_\theta;\mathcal G)\} \left(x_D\mid c_D^{\mathsf I}(x_R)\right) \label{eq:configured-factor-reassembly}\\
&\qquad={} \Phi_{\mathfrak d_D(\mathcal H)} \left((\operatorname{Sel}_{\mathsf I}\theta)_{\downarrow D}\right) \left(x_D\mid x_{\operatorname{pa}_{\mathcal H}(D)\setminus D}\right). \nonumber
\end{align}
Both sides are strictly positive.
\end{lemma}
\begin{corollary}[Active-factor locality]\label{cor:active-factor-locality}
For $D\in\mathcal D(\mathcal H)$, let $\delta(D)$ be its native district in $\mathcal G$, as in Section~\ref{sec:setup}. If $u\in U_K^{\mathrm{nat}}$ for a native district $K\ne\delta(D)$, then, for every $\theta\in\Omega_{\mathcal G}$,
\[
\mathbb D_\theta\!\left[ \operatorname{red}_D\{\phi_{V\setminus D}(p_\theta;\mathcal G)\} \bigl(x_D\mid c_D^{\mathsf I}(x_R)\bigr) \right][u]=0 \quad\text{for every }x_R.
\]
Equivalently, the active-$D$ rows of $\operatorname{Sel}_{\mathsf I}u$ vanish. Hence active channels may be grouped by their native district before the Gram projection.
\end{corollary}
\begin{theorem}[Active nested-law closure and normalized intervention bridge]\label{thm:normalized-intervention-bridge}
For $p_\theta=\Phi_{\mathcal G}(\theta)$ with $\theta\in\Omega_{\mathcal G}$, define
\begin{equation}\label{eq:statistical-intervention-functional}
\Gamma_{\mathsf I}(p_\theta)(x_R) :=\prod_{D\in\mathcal D(\mathcal H)} \operatorname{red}_D\{\phi_{V\setminus D}(p_\theta;\mathcal G)\} \left(x_D\mid c_D^{\mathsf I}(x_R)\right).
\end{equation}
Then the global identity
\begin{equation}\label{eq:active-chart-bridge}
\Gamma_{\mathsf I}(p_\theta) =\Phi_{\mathcal H}(\operatorname{Sel}_{\mathsf I}\theta)
\end{equation}
holds on all of $\Omega_{\mathcal G}$, and
\begin{equation}\label{eq:active-chart-domain}
\operatorname{Sel}_{\mathsf I}\theta\in\Omega_{\mathcal H}, \qquad \Gamma_{\mathsf I}(p_\theta)\in\mathcal N_+(\mathcal H).
\end{equation}
Thus $\Gamma_{\mathsf I}$ maps $\mathcal N_+(\mathcal G)$ into $\mathcal N_+(\mathcal H)$. Consequently, $\theta\mapsto\Gamma_{\mathsf I}(p_\theta)$ is polynomial and $p\mapsto\Gamma_{\mathsf I}(p)$ is $C^\infty$ on $\mathcal N_+(\mathcal G)$.
\end{theorem}
For $Q\in\mathcal N_+(\mathcal H)$, let $\eta=\Theta_{\mathcal H}(Q)$ and let $q=\Phi_{\mathcal H}(\eta)$ be its mass function. Define the active-chart score operator
\begin{equation}\label{eq:active-chart-score-operator}
J_{Q,\mathcal H}:U_{\mathcal H}\longrightarrow L_2^0(Q), \qquad J_{Q,\mathcal H}v :=\frac{\mathbb D\Phi_{\mathcal H,\eta}[v]}{q}.
\end{equation}
\begin{corollary}[Active-side nested geometry]\label{cor:active-side-nested-geometry}
For every class (I1)--(I3), the configured active law $\Gamma_{\mathsf I}(P)$ of Theorem~\ref{thm:normalized-intervention-bridge} is a strictly positive nested Markov law of the active graph $\mathcal H$. In particular, at $P^{\mathsf I}=\Gamma_{\mathsf I}(P)$,
\[
\mathcal T_{P^{\mathsf I}}\mathcal N(\mathcal H) =\operatorname{ran}(J_{P^{\mathsf I},\mathcal H}),
\]
where the active score operator is defined in \eqref{eq:active-chart-score-operator}. Hence the complete finite-state chart geometry of Theorem~\ref{thm:forward-chart} is available on the active side.
\end{corollary}
\begin{remark}[Statistical, not neighborhood-causal, extension]\label{rem:statistical-not-causal-extension}
At the true causal law, $\Gamma_{\mathsf I}(P)$ is the configured active law of Proposition~\ref{prop:causal-identification-interface}: in classes (I1)--(I2) it is additionally identified as the joint causal response on $R_{\mathsf I}$, while in (I3) only its $Y$ margin carries the causal interpretation. Away from that law, equation~\eqref{eq:active-chart-bridge} defines the statistical parameter on $\mathcal N_+(\mathcal G)$. We do not assert that every nearby nested law admits a compatible latent realization, or that the causal identification theorem itself transports to that law. Only the graph, factor indices, boundary values, and the linear map $\operatorname{Sel}_{\mathsf I}$ are fixed throughout the chart. This is the distinction that prevents the latent-margin model from re-entering the operative statistical scope.
\end{remark}
\subsection{Order-free differentiation and the canonical gradient}
Let $\theta_0=\Theta_{\mathcal G}(P)$, $\eta_0=\operatorname{Sel}_{\mathsf I}\theta_0$, and $P^{\mathsf I}=\Phi_{\mathcal H}(\eta_0)$. For $g:\mathfrak X_Y\to\mathbb R$, define
\[
\Psi_{\mathsf I,g}(P) :=\mathrm E_{\Gamma_{\mathsf I}(P)}\{g(X_Y)\}.
\]
For a global chart direction $u\in U_{\mathcal G}$, put
\begin{equation}\label{eq:active-chart-score}
h_{\mathsf I}^u :=J_{P^{\mathsf I},\mathcal H}(\operatorname{Sel}_{\mathsf I}u).
\end{equation}
Equivalently, differentiating the product in \eqref{eq:statistical-intervention-functional} gives
\begin{equation}\label{eq:active-factor-score-sum}
h_{\mathsf I}^u(x_R) =\sum_{D\in\mathcal D(\mathcal H)} \frac{ \mathbb D\Phi_{\mathcal G[D],\theta_{0,\downarrow D}} [u_{\downarrow D}] \bigl(x_D\mid c_D^{\mathsf I}(x_R)\bigr)} {\Phi_{\mathcal G[D]}(\theta_{0,\downarrow D}) \bigl(x_D\mid c_D^{\mathsf I}(x_R)\bigr)}.
\end{equation}
Equation~\eqref{eq:active-chart-score} is the structural form used in the proof, whereas \eqref{eq:active-factor-score-sum} is the computationally convenient form: it evaluates the same score through separately configured active factors.
\begin{theorem}[Direct chart derivative and finite-state EIF]\label{thm:intervention-chart-eif}
For every regular model path $P_t$ through $P$, let $u=(\mathrm d/\mathrm dt)\Theta_{\mathcal G}(P_t)|_{t=0}$. Then
\begin{equation}\label{eq:intervention-chart-derivative}
\left.\frac{\mathrm d}{\mathrm dt}\Psi_{\mathsf I,g}(P_t)\right|_{t=0} =\mathrm E_{P^{\mathsf I}} \left[ \{g(X_Y)-\Psi_{\mathsf I,g}(P)\}h_{\mathsf I}^u(X_R) \right].
\end{equation}
For each native district $K\in\mathcal D(\mathcal G)$, use the basis $\{e_{Kj}^{\mathrm{nat}}:1\leq j\leq d_K^{\mathrm{nat}}\}$, score vector $b_K^{\mathrm{nat}}$, and Gram matrix $G_K^{\mathrm{nat}}(P)$ from Proposition~\ref{prop:gram-projection}, and define
\begin{equation}\label{eq:intervention-coordinate-derivative}
(a_K)_j :=\mathrm E_{P^{\mathsf I}} \left[ \{g(X_Y)-\Psi_{\mathsf I,g}(P)\} J_{P^{\mathsf I},\mathcal H} (\operatorname{Sel}_{\mathsf I}e_{Kj}^{\mathrm{nat}}) \right].
\end{equation}
The efficient influence function in $\mathcal N_+(\mathcal G)$ is
\begin{equation}\label{eq:direct-intervention-eif}
\phi_{P,\mathsf I,g}^{\mathrm{eff}} =\sum_{K\in\mathcal D(\mathcal G)} (b_K^{\mathrm{nat}})^\top \{G_K^{\mathrm{nat}}(P)\}^{-1}a_K.
\end{equation}
No active-district, source-aware, or observational vertex order is required. All active contributions descending from the same native district are already aggregated in $a_K$ before the Gram inverse is applied.
\end{theorem}
\begin{corollary}[Excluded source-only native blocks]\label{cor:excluded-source-native-block}
If a native district $K\in\mathcal D(\mathcal G)$ satisfies $K\cap R=\varnothing$, then $\operatorname{Sel}_{\mathsf I}U_K^{\mathrm{nat}}=\{0\}$, $a_K=0$, and the $K$-summand in \eqref{eq:direct-intervention-eif} vanishes. In particular, when an excluded source is a singleton native district containing no active vertex, the target derivative vanishes along its native score block, and that block contributes no term to the efficient influence function.
\end{corollary}
\subsection{Independent-draw stochastic and source-standardized targets}
The last extension averages the configured active laws over one boundary assignment drawn afresh from a possibly population-dependent law on a fixed finite assignment set, independently of the natural source variables and of the structural disturbances, although the coordinates of the drawn assignment vector may be dependent (Definition~\ref{def:admissible-source-law}). When the weights depend on the population, differentiating the mixture mean of \eqref{eq:independent-source-mixture-and-mean} adds the weight derivative $\alpha_K^\mu$ of Theorem~\ref{thm:independent-source-mixture-eif} to the averaged fixed-component derivative. Every component is a configured active law with the common excluded set, active set and active graph and is differentiated through its own active chart; the mixture law itself receives no chart, since it need not belong to the nested model of the active graph, and identification of the mixture mean is distinguished below from identification of the whole mixture law.
Let $\mathcal V_{\mathsf I}$ be a fixed nonempty finite set of boundary assignments. For every $a\in\mathcal V_{\mathsf I}$, assume that $\mathsf I(a)$ is a fixed intervention of one of the classes (I1)--(I3) satisfying Proposition~\ref{prop:causal-identification-interface}, with the same excluded set $A$, active set $R$, and active graph $\mathcal H$. Write its selection map as $\operatorname{Sel}_{\mathsf I,a}$ and put
\[
P_{\theta}^{\mathsf I,a} :=\Phi_{\mathcal H}(\operatorname{Sel}_{\mathsf I,a}\theta), \qquad \psi_a(\theta) :=\mathrm E_{P_{\theta}^{\mathsf I,a}}\{g(X_Y)\}.
\]
For a node intervention one may take $\mathcal V_{\mathsf I}=\mathfrak X_A$. For an edge or path intervention, the index set instead consists only of boundary assignments satisfying the within-district compatibility condition of \textup{(H4)}.
\begin{definition}[Admissible independent-draw source law]\label{def:admissible-source-law}
A family $\mu=(\mu_\theta:\theta\in\Omega_{\mathcal G})$ is admissible on $\mathcal V_{\mathsf I}$ if
\[
\mu_\theta(a)\geq0, \qquad \sum_{a\in\mathcal V_{\mathsf I}}\mu_\theta(a)=1,
\]
for every $\theta\in\Omega_{\mathcal G}$, and every map $\theta\mapsto\mu_\theta(a)$ is $C^\infty$. The intervention draws one boundary assignment afresh from $\mu_\theta$, independently of the natural source variables and of the structural disturbances, and then applies $\mathsf I(a)$. Coordinates of a jointly drawn source vector may be dependent under $\mu_\theta$. The fresh independent draw is part of admissibility (Remark~\ref{rem:independent-not-natural-source}).
\end{definition}
The associated configured mixture law and mean are
\begin{equation}\label{eq:independent-source-mixture-and-mean}
\begin{aligned}
\Gamma_{\mu}(p_\theta) &:=\sum_{a\in\mathcal V_{\mathsf I}} \mu_\theta(a)P_{\theta}^{\mathsf I,a},\\
\Psi_{\mu,g}(p_\theta) &:=\sum_{a\in\mathcal V_{\mathsf I}} \mu_\theta(a)\psi_a(\theta).
\end{aligned}
\end{equation}
At the true causal law, this composes the configured component laws of Proposition~\ref{prop:causal-identification-interface}; no additional causal identification result is invoked. The mean $\Psi_{\mu,g}$ is causally identified componentwise in every class (I1)--(I3), because each $\psi_a$ uses only the causally identified component functional. The cited identification theorems justify a joint causal interpretation of $\Gamma_\mu$ when every positive-weight component is jointly identified, as in (I1)--(I2); (I3) margin identification alone does not justify such a reading.
\begin{proposition}[Positive smooth mixture]\label{prop:positive-smooth-source-mixture}
For every admissible $\mu$, the mass function in the first line of \eqref{eq:independent-source-mixture-and-mean} is normalized and strictly positive, and both $\Gamma_\mu$ and $\Psi_{\mu,g}$ are $C^\infty$ on $\mathcal N_+(\mathcal G)$. If the weights $\mu_\theta(a)$ are polynomial in $\theta$, as for a fixed exogenous law or the marginal source-standardization rule below, then both maps are polynomial in $\theta$.
\end{proposition}
Unlike a fixed component, the mixture is not generally a chart law of $\mathcal H$, and membership in $\mathcal N_+(\mathcal H)$ is not asserted. Section~\ref{app:B-mixture} of the Supplementary Material exhibits the loss with a strictly positive two-component counterexample whose equal mixture is not Markov for its active graph. The active chart is therefore used componentwise below, not assigned to the mixture itself.
Let $\theta_0=\Theta_{\mathcal G}(P)$, $P^{\mathsf I,a}=P_{\theta_0}^{\mathsf I,a}$, $\psi_a=\psi_a(\theta_0)$, and $\mu_0=\mu_{\theta_0}$. For each native district $K$ define the component and weight coordinate vectors by
\begin{equation}\label{eq:mixture-coordinate-contributions}
\begin{aligned}
(\alpha_K^{\mathrm{cmp}})_j &:={} \sum_{a\in\mathcal V_{\mathsf I}}\mu_0(a) \mathrm E_{P^{\mathsf I,a}} \left[ \{g(X_Y)-\psi_a\} J_{P^{\mathsf I,a},\mathcal H} (\operatorname{Sel}_{\mathsf I,a}e_{Kj}^{\mathrm{nat}}) \right],\\
(\alpha_K^{\mu})_j &:={} \sum_{a\in\mathcal V_{\mathsf I}} \mathbb D\mu_{\theta_0}[e_{Kj}^{\mathrm{nat}}](a)\,\psi_a.
\end{aligned}
\end{equation}
Since the weights sum to one, $\sum_a\mathbb D\mu_{\theta_0}[u](a)=0$; hence $\psi_a$ in the weight line may equivalently be replaced by $\psi_a-\Psi_{\mu,g}(P)$.
\begin{theorem}[EIF for an admissible finite independent-draw mixture]\label{thm:independent-source-mixture-eif}
For every regular path $P_t$ whose chart velocity is $u=\sum_Ku_K$,
\begin{equation}\label{eq:independent-source-mixture-derivative}
\left.\frac{\mathrm d}{\mathrm dt}\Psi_{\mu,g}(P_t)\right|_{t=0} =\sum_{K\in\mathcal D(\mathcal G)} (\alpha_K^{\mathrm{cmp}}+\alpha_K^\mu)^\top u_K.
\end{equation}
Its efficient influence function in $\mathcal N_+(\mathcal G)$ is
\begin{equation}\label{eq:independent-source-mixture-eif}
\phi_{P,\mu,g}^{\mathrm{eff}} =\sum_{K\in\mathcal D(\mathcal G)} (b_K^{\mathrm{nat}})^\top\{G_K^{\mathrm{nat}}(P)\}^{-1} (\alpha_K^{\mathrm{cmp}}+\alpha_K^\mu).
\end{equation}
No active-graph membership or vertex ordering for the mixture is required.
\end{theorem}
\begin{corollary}[Exogenous and degenerate source laws]\label{cor:exogenous-source-mixture}
If $\mu_\theta\equiv\mu$ is fixed, then $\alpha_K^\mu=0$ and
\[
\phi_{P,\mu,g}^{\mathrm{eff}} =\sum_{a\in\mathcal V_{\mathsf I}} \mu(a)\phi_{P,\mathsf I(a),g}^{\mathrm{eff}}.
\]
If $\mu$ is a point mass at $a$, this identity reduces exactly to Theorem~\ref{thm:intervention-chart-eif} for $\mathsf I(a)$.
\end{corollary}
\begin{corollary}[Source-standardized independent redraw]\label{cor:source-standardized-eif}
Suppose $\mathcal V_{\mathsf I}=\mathfrak X_A$ and
\begin{equation}\label{eq:source-standardization-law}
\mu_\theta(a)=\mathrm P_\theta(X_A=a).
\end{equation}
Define the observed-data function $\psi_{X_A}:=\sum_a\psi_a\mathbf 1\{X_A=a\}$. Then
\begin{equation}\label{eq:source-standardized-weight-coordinate}
\alpha_K^\mu =\mathrm E_P\!\left[ \{\psi_{X_A}-\Psi_{\mu,g}(P)\}b_K^{\mathrm{nat}} \right],
\end{equation}
and, writing $\Pi_{\mathcal T_P\mathcal N(\mathcal G)}$ for the $L_2(P)$ projection onto the model tangent space,
\begin{equation}\label{eq:source-standardized-eif}
\phi_{P,\mu,g}^{\mathrm{eff}} =\sum_a \mathrm P_P(X_A=a)\, \phi_{P,\mathsf I(a),g}^{\mathrm{eff}} +\Pi_{\mathcal T_P\mathcal N(\mathcal G)} \{\psi_{X_A}-\Psi_{\mu,g}(P)\}.
\end{equation}
\end{corollary}
\begin{remark}[Independent redraw is not natural-value retention]\label{rem:independent-not-natural-source}
Equation~\eqref{eq:source-standardization-law} draws a fresh joint source vector independently from its observed marginal. It does not set the response to the unit's naturally realized source vector, which retains the natural source--disturbance coupling and is not represented in general by the first line of \eqref{eq:independent-source-mixture-and-mean}. The latter natural-value construction is outside the present theorem.
\end{remark}
\begin{corollary}[Complementary blocks for a root singleton source]\label{cor:source-standardized-complementary-blocks}
Let $\mu$ be either a fixed exogenous law or the marginal source-standardization rule \eqref{eq:source-standardization-law}, and suppose $A=\{w\}$, $\operatorname{pa}_{\mathcal G}(w)=\varnothing$, and $\operatorname{dis}_{\mathcal G}(w)=\{w\}$. For the native source district $K_w=\{w\}$,
\[
\alpha_{K_w}^{\mathrm{cmp}}=0, \qquad \alpha_K^\mu=0\quad(K\ne K_w).
\]
Thus a fixed or exogenous intervention has no $K_w$ contribution, whereas under \eqref{eq:source-standardization-law} every $K_w$ contribution comes from the weight term, which is nonzero whenever $w\mapsto\psi_w$ is nonconstant. For a general admissible $P$-dependent rule, $\alpha_K^\mu$ need not vanish for $K\ne K_w$; see Remark~\ref{rem:p-dependent-source-law}.
\end{corollary}
\begin{remark}[The source-law chain-rule term]\label{rem:p-dependent-source-law}
For a general admissible $P$-dependent law, the correct weight coordinate is the weight line of \eqref{eq:mixture-coordinate-contributions}; the observed-data formula \eqref{eq:source-standardized-weight-coordinate} is specific to marginal source standardization. Dropping $\alpha_K^\mu$ is a chain-rule error and in general destroys the Riesz identity on the native directions on which the source-law derivative is nonzero. For instance, take binary $W,Y,Z$ in the DAG $W\to Y$ with $Z$ isolated. The admissible rule $\mu_\theta(1)=\mathrm P_\theta(Z=1)$, with $\mu_\theta(0)=1-\mu_\theta(1)$, gives $\alpha^\mu_{\{Z\}}=\mathbb D\mu_\theta[e_{\{Z\}}](1)(\psi_1-\psi_0)\ne0$ whenever $\psi_1\ne\psi_0$, where $e_{\{Z\}}$ is the singleton native coordinate direction.
\end{remark}
\section{Consequences for efficient and targeted learning}\label{sec:targeted}
Let $\Psi:\mathcal N_+(\mathcal G)\to\mathbb R$ be pathwise differentiable, and suppose an ambient gradient $\widetilde\phi_P\in L_2^0(P)$ satisfies
\[
\left.\frac{\mathrm d}{\mathrm dt}\Psi(P_t)\right|_{t=0} =\mathrm E_P(\widetilde\phi_P s)
\]
for every regular model score $s$. The efficient influence function in the nested model is
\begin{equation}\label{eq:eif-projection}
\phi_P^{\mathrm{eff}} =\Pi_{\mathcal T_P\mathcal N(\mathcal G)}\widetilde\phi_P.
\end{equation}
\begin{corollary}[Explicit finite-state EIF]\label{cor:explicit-eif}
Under Theorem~\ref{thm:district-orthogonality},
\[
\phi_P^{\mathrm{eff}} =\sum_{D\in\mathcal D(\mathcal G)} (b_D^{\mathrm{nat}})^\top\{G_D^{\mathrm{nat}}(P)\}^{-1} \mathrm E_P\{b_D^{\mathrm{nat}}\widetilde\phi_P\}.
\]
If several intervention-induced active channels descend from one native district, their coordinate derivative vectors are summed within that district and the same native Gram inverse is applied to every channel. By linearity, summation may precede or follow this operation; treating the channels as separate orthogonal blocks with their own metrics is not valid.
\end{corollary}
For the fixed-node, complete-source edge, and path-specific functionals of the classes (I1)--(I3), Theorem~\ref{thm:normalized-intervention-bridge} supplies the identification-to-chart bridge on the entire positive coordinate domain. Theorem~\ref{thm:intervention-chart-eif} then differentiates the active chart and solves the finite-dimensional Riesz problem directly. Thus no ordered change-of-measure argument or separately chosen saturated-model gradient is needed for the abstract EIF. A saturated off-model extension would be noncanonical unless specified; the canonical gradient in \eqref{eq:direct-intervention-eif} is intrinsic to the operative nested model.
This order-free result does not by itself validate an anchored or sequential observed-data representation based on a single vertex order; Section~\ref{sec:sequential-union} states what such a construction additionally requires. Admissible finite independent-draw mixtures of the configured component laws of those classes are covered by Theorem~\ref{thm:independent-source-mixture-eif}, which adds the source-law weight derivative to the averaged fixed-component derivative, and by its source-standardized specialization, Corollary~\ref{cor:source-standardized-eif}, within the scope fixed in Remark~\ref{rem:independent-not-natural-source}.
Applied to the direction $h$ of Section~\ref{sec:counterexample}, which is observationally centered but not a model score, the projection removes a nonzero component $h-\Pi_{\mathcal T_P\mathcal N(\mathcal G)}h$, whose squared relative norm $\|h-\Pi_{\mathcal T_P\mathcal N(\mathcal G)}h\|_{P,2}^2/\|h\|_{P,2}^2$ is strictly positive at the law of that section (exact value in Section~\ref{app:B-primary} of the Supplementary Material). That component is a projection diagnostic, not an efficiency correction: $h$ is a diagnostic direction, not an influence function identified for a specified target, and describing it as one would require first exhibiting a functional whose derivative it represents.
\subsection{Coherent chart plug-in and local stability}
On fixed finite support the chart information matrix is positive definite at every interior law (Corollary~\ref{cor:fisher-chart-pullback}), so a regular interior nested-model maximum likelihood estimator is root-$n$ and its plug-in target is efficient by the delta method: efficiency itself needs no targeting in this idealization. The constructions of this section use coherent chart laws as initial estimators, including regularized, cross-fitted, or approximately solved fits that are not interior full-chart likelihood solutions. Proposition~\ref{prop:modular-eif-remainder} additionally allows approximations of the Gram and derivative coefficients at a single such law. Fits assembled from mutually incompatible modules, an estimated score basis, or a separately estimated target are not analysed. The coherent one-step estimator and the chart-flow TMLE differ in their exact guarantees; that comparison is made in Section~\ref{subsec:one-step-tmle-contrast}, and the null-update facts for the two estimators at an interior likelihood fit are recorded in the remarks following Corollaries~\ref{cor:same-sample-one-step} and~\ref{cor:tmle-substitution}.
For the remainder of this section, $\Psi$ is either a fixed-response mean of the classes (I1)--(I3) from Theorem~\ref{thm:intervention-chart-eif} or an admissible independent-draw mixture mean from Theorem~\ref{thm:independent-source-mixture-eif}. The same arguments apply to any functional whose chart representation and canonical-gradient map have the smoothness asserted below.
Throughout this section, the graph, vertex state spaces, chart dimension, finite intervention-assignment set, and (when used) number of folds are fixed. We equip $U_{\mathcal G}$ with the Euclidean norm induced by its fixed coordinate ordering and write it as $\|\cdot\|$; because $d$ is fixed, any other norm on $U_{\mathcal G}$ gives the same rate conditions. None of the arguments below covers a growing-support or growing-dimension regime.
Write $F(\theta)=\Psi(P_\theta)$. For the fixed coordinate basis $\{e_{Kj}^{\mathrm{nat}}:1\leq j\leq d_K^{\mathrm{nat}}\}$ of a native district $K$, let $b_{K,\theta}^{\mathrm{nat}}$ and $G_K^{\mathrm{nat}}(\theta)$ denote the score vector and Gram matrix at $P_\theta$. Define the coordinate-gradient vector
\[
\begin{aligned}
\gamma_K(\theta) &:=\bigl(\mathbb DF(\theta)[e_{Kj}^{\mathrm{nat}}]\bigr)_{j=1}^{d_K^{\mathrm{nat}}},\\
\gamma_K(\theta) &=\begin{cases} a_K(\theta),&\text{for a fixed intervention},\\ \alpha_K^{\mathrm{cmp}}(\theta)+\alpha_K^\mu(\theta), &\text{for an admissible independent-draw mixture}, \end{cases}
\end{aligned}
\]
where the second line uses \eqref{eq:mixture-coordinate-contributions} evaluated at $P_\theta$. We call the two summands the \emph{component contribution} and the \emph{weight contribution}, respectively. Put
\begin{equation}\label{eq:coherent-eif-map}
\beta_K(\theta):=\{G_K^{\mathrm{nat}}(\theta)\}^{-1}\gamma_K(\theta), \qquad \phi_\theta:=\sum_{K\in\mathcal D(\mathcal G)} (b_{K,\theta}^{\mathrm{nat}})^\top\beta_K(\theta).
\end{equation}
Thus $\phi_\theta$ is the exact canonical gradient at the single fitted law $P_\theta$; the target, basis, Gram matrix, component means, source law, and source-law derivative are not fitted at mutually incompatible laws.
\begin{proposition}[Local stability of the coordinate EIF]\label{prop:estimated-eif-stability}
Fix $\theta_0\in\Omega_{\mathcal G}$. There is a closed convex neighborhood $\mathcal K\Subset\Omega_{\mathcal G}$ of $\theta_0$ and constants $c,\kappa,C>0$ such that, for every $\theta,\eta\in\mathcal K$,
\begin{equation}\label{eq:local-eif-stability-bounds}
\begin{aligned}
\min_{x_V}p_\theta(x_V)&\geq c, &\lambda_{\min}\{G_K^{\mathrm{nat}}(\theta)\}&\geq\kappa \quad(K\in\mathcal D(\mathcal G)),\\
\max_{x_V}|\phi_\theta(x_V)-\phi_\eta(x_V)| &\leq C\|\theta-\eta\|.
\end{aligned}
\end{equation}
Every cell probability of every configured active law is also uniformly bounded away from zero on $\mathcal K$. The maps $b_{K,\theta}^{\mathrm{nat}}$, $G_K^{\mathrm{nat}}(\theta)$, $\{G_K^{\mathrm{nat}}(\theta)\}^{-1}$, $\gamma_K(\theta)$, $\beta_K(\theta)$, and $\phi_\theta$ are smooth on $\mathcal K$. These conclusions hold for every admissible $C^\infty$ source-law rule, whether or not that rule is polynomial. More generally, the same conclusions and bounds hold, with set-dependent constants, on every compact convex subset of $\Omega_{\mathcal G}$ on which the target is defined.
\end{proposition}
The coherent construction uses the model-based finite sum
\[
\widehat G_K^{\mathrm{nat}} :=G_K^{\mathrm{nat}}(\widehat\theta) =\sum_{x_V}p_{\widehat\theta}(x_V) b_{K,\widehat\theta}^{\mathrm{nat}}(x_V) \{b_{K,\widehat\theta}^{\mathrm{nat}}(x_V)\}^\top,
\]
not an evaluation-sample empirical Gram matrix. In the fixed-state model this quantity is exactly computable once $\widehat\theta$ is fitted.
\begin{lemma}[Mixture error bookkeeping]\label{lem:mixture-error-bookkeeping}
For an admissible mixture, define
\[
d_{K,a}(\theta) :=\bigl(\mathbb D\psi_{a,\theta}[e_{Kj}^{\mathrm{nat}}]\bigr) _{j=1}^{d_K^{\mathrm{nat}}}, \qquad m_{K,a}(\theta) :=\bigl(\mathbb D\mu_\theta[e_{Kj}^{\mathrm{nat}}](a)\bigr) _{j=1}^{d_K^{\mathrm{nat}}}.
\]
Then
\[
\alpha_K^{\mathrm{cmp}}(\theta)=\sum_a\mu_\theta(a)d_{K,a}(\theta), \qquad \alpha_K^\mu(\theta)=\sum_a m_{K,a}(\theta)\psi_a(\theta).
\]
For $\eta,\theta\in\mathcal K$, let $\Delta$ denote evaluation at $\eta$ minus evaluation at $\theta$. The exact identities
\begin{equation}\label{eq:mixture-error-decompositions}
\begin{aligned}
\Delta\alpha_K^{\mathrm{cmp}} &=\sum_a\{\mu_\theta(a)\Delta d_{K,a} +d_{K,a}(\theta)\Delta\mu(a) +\Delta\mu(a)\Delta d_{K,a}\},\\
\Delta\alpha_K^\mu &=\sum_a\{m_{K,a}(\theta)\Delta\psi_a +\psi_a(\theta)\Delta m_{K,a} +\Delta m_{K,a}\Delta\psi_a\},\\
F(\eta)-F(\theta) &=\sum_a\{\mu_\theta(a)\Delta\psi_a +\psi_a(\theta)\Delta\mu(a) +\Delta\mu(a)\Delta\psi_a\}
\end{aligned}
\end{equation}
hold. In particular, the bilinear term in the value line of \eqref{eq:mixture-error-decompositions} is distinct from the $\Delta m_{K,a}\Delta\psi_a$ term in the estimated weight contribution. Moreover, for every direction $u$,
\begin{equation}\label{eq:smooth-mixture-second-derivative}
\mathbb D^2F_\theta[u,u] =\sum_a\left[ \mu_\theta(a)\mathbb D^2\psi_{a,\theta}[u,u] +2\mathbb D\mu_\theta[u](a)\mathbb D\psi_{a,\theta}[u] +\mathbb D^2\mu_\theta[u,u](a)\psi_a(\theta) \right].
\end{equation}
\end{lemma}
The cross-products in \eqref{eq:mixture-error-decompositions} are second order for a coherent chart plug-in. They do not establish double robustness: the two pure-curvature terms in \eqref{eq:smooth-mixture-second-derivative}, together with score, Gram, and coordinate errors, remain in general.
\subsection{Exact one-step remainder and limit theory}
For any $\eta,\theta_0\in\Omega_{\mathcal G}$ at which the target is defined, set
\begin{equation}\label{eq:coherent-second-order-remainder}
R_2(\eta,\theta_0) :=F(\eta)-F(\theta_0)+\mathrm E_{P_{\theta_0}}(\phi_\eta).
\end{equation}
The next lemma works on an arbitrary compact convex subset of the positive chart domain. The compact set supplies a uniform local bound; the definition itself is global on that domain.
\begin{lemma}[Exact remainder identity and quadratic bound]\label{lem:exact-one-step-remainder}
Let $\mathcal K_\star\Subset\Omega_{\mathcal G}$ be compact and convex, and suppose the target is defined on an open neighborhood of $\mathcal K_\star$. For $\theta_0,\eta\in\mathcal K_\star$, let $\delta=\eta-\theta_0$ and $\theta_s=\theta_0+s\delta$. Then
\begin{equation}\label{eq:exact-one-step-remainder}
R_2(\eta,\theta_0) =\int_0^1 \mathrm E_{P_{\theta_s}}\!\left[ \{\phi_{\theta_s}-\phi_\eta\} J_{P_{\theta_s}}\delta \right]\mathrm ds.
\end{equation}
Consequently,
\begin{equation}\label{eq:quadratic-one-step-bound}
|R_2(\eta,\theta_0)|\leq C_{\mathcal K_\star} \|\eta-\theta_0\|^2 \qquad(\theta_0,\eta\in\mathcal K_\star).
\end{equation}
\end{lemma}
\begin{proposition}[Exact source-block cancellation of the coherent one-step drift]\label{prop:source-block-drift-cancellation}
Suppose the target uses marginal source standardization as in Corollary~\ref{cor:source-standardized-eif}, with $A=\{w\}$, $\operatorname{pa}_{\mathcal G}(w)=\varnothing$, and $\operatorname{dis}_{\mathcal G}(w)=\{w\}$. Let $K_w=\{w\}$ be the native source district. If $\eta,\theta_0\in\Omega_{\mathcal G}$ satisfy
\[
\eta-\theta_0\in U_{K_w}^{\mathrm{nat}},
\]
then
\begin{equation}\label{eq:exact-source-block-cancellation}
R_2(\eta,\theta_0)=0, \qquad\text{equivalently}\qquad F(\eta)+\mathrm E_{P_{\theta_0}}(\phi_\eta)=F(\theta_0).
\end{equation}
Thus the coherent population one-step correction removes an arbitrary interior error in the source marginal exactly, provided every complementary chart block is correct.
\end{proposition}
\begin{remark}[One-sided population robustness, not double robustness]
Proposition~\ref{prop:source-block-drift-cancellation} is a global identity on the positive source-block slice, not merely a local quadratic bound. There is no reverse robustness statement: correctness of the source marginal alone does not force $R_2=0$, and simultaneous errors in complementary outcome-side coordinates can leave a nonzero drift. The result is therefore one-sided source-block robustness of the coherent population one-step correction, not double or multiple robustness.
Nor does the proposition establish robustness of the chart-flow TMLE under a fixed, arbitrarily misspecified source marginal. The targeting field in \eqref{eq:targeting-coordinate-field} generally has nonzero complementary coordinates, so its integral curve need not remain on the source-only slice; its updated-law remainder is $R_2(\theta^\star,\theta_0)$, not $R_2(\eta,\theta_0)$. The local chart-TMLE theorem remains valid under its stated $o_p(n^{-1/4})$ full-chart rate, but no union-model TMLE claim is made.
\end{remark}
Let $P=P_{\theta_0}$, and let $X_1,\ldots,X_n$ be i.i.d. from $P$. Partition the sample, independently of the observations, into a fixed number $L$ of folds $I_v$, with $n_v/n\to\rho_v\in(0,1)$; throughout, condition on the realized fold assignment. Let $\widehat\theta_{-v}$ use only observations outside $I_v$, and compute every object in \eqref{eq:coherent-eif-map} at the single fitted law $\widehat P_{-v}=P_{\widehat\theta_{-v}}$. Write $\mathbb P_{n,v}$ for the empirical law on $I_v$. The compact set $\mathcal K$ in Proposition \ref{prop:estimated-eif-stability} is convex. Thus every segment joining $\theta_0$ to an admitted $\widehat\theta_{-v}\in\mathcal K$ remains in $\mathcal K$, as required by Lemma~\ref{lem:exact-one-step-remainder}; a proposed step that leaves this compact positive-chart neighborhood is not covered by the rate theorem.
\begin{theorem}[Coherent cross-fitted one-step estimator]\label{thm:coherent-one-step}
Assume that, with probability tending to one, $\widehat\theta_{-v}\in\mathcal K$ for every $v$, and
\begin{equation}\label{eq:one-step-initial-rate}
\max_{1\leq v\leq L} \|\widehat\theta_{-v}-\theta_0\|=o_p(n^{-1/4}).
\end{equation}
Define
\begin{equation}\label{eq:cross-fitted-one-step}
\widehat\Psi_{\mathrm{cf}} :=\sum_{v=1}^L\frac{n_v}{n} \left\{F(\widehat\theta_{-v}) +\mathbb P_{n,v}\phi_{\widehat\theta_{-v}}\right\}.
\end{equation}
Then the exact expansion
\begin{align}
\widehat\Psi_{\mathrm{cf}}-F(\theta_0) =(\mathbb P_n-P)\phi_{\theta_0} +\sum_{v=1}^L\frac{n_v}{n} \Bigl[R_2(\widehat\theta_{-v},\theta_0) +(\mathbb P_{n,v}-P)\{\phi_{\widehat\theta_{-v}}-\phi_{\theta_0}\}\Bigr] \label{eq:cross-fitted-exact-expansion}
\end{align}
holds, and
\begin{equation}\label{eq:one-step-asymptotic-linearity}
\sqrt n\{\widehat\Psi_{\mathrm{cf}}-F(\theta_0)\} =\frac1{\sqrt n}\sum_{i=1}^n\phi_{\theta_0}(X_i)+o_p(1).
\end{equation}
If $\sigma^2=\mathrm E_P(\phi_{\theta_0}^2)>0$, the limit distribution is $N(0,\sigma^2)$. In particular, with
\[
\overline\phi_{\mathrm{cf}} :=\sum_{v=1}^L\frac{n_v}{n} \mathbb P_{n,v}\phi_{\widehat\theta_{-v}}, \qquad \widehat\sigma^2_{\mathrm{cf}} :=\sum_{v=1}^L\frac{n_v}{n} \mathbb P_{n,v} \{\phi_{\widehat\theta_{-v}}-\overline\phi_{\mathrm{cf}}\}^2,
\]
$\widehat\sigma^2_{\mathrm{cf}}\to_p\sigma^2$.
\end{theorem}
\begin{corollary}[Exact foldwise unbiasedness on the source-only slice]\label{cor:source-slice-foldwise-unbiased}
Assume the conditions of Proposition \ref{prop:source-block-drift-cancellation}. For the fixed, data-independent fold partition above, let $\mathcal F_{-v}:=\sigma(X_i:i\notin I_v)$, and suppose $\widehat\theta_{-v}$ is $\mathcal F_{-v}$-measurable. Suppose also that there is a fixed compact set $\mathcal K_s\Subset\Omega_{\mathcal G}$ such that
\[
\widehat\theta_{-v}\in\mathcal K_s, \qquad \widehat\theta_{-v}-\theta_0\in U_{K_w}^{\mathrm{nat}} \quad\text{almost surely}.
\]
Then, without a rate condition on the fitted source block,
\begin{equation}\label{eq:source-slice-foldwise-unbiasedness}
\mathrm E\!\left[ F(\widehat\theta_{-v}) +\mathbb P_{n,v}\phi_{\widehat\theta_{-v}} \,\middle|\, \mathcal F_{-v} \right] =F(\theta_0) \qquad(v=1,\ldots,L).
\end{equation}
Consequently, the cross-fitted estimator in \eqref{eq:cross-fitted-one-step}, formed with the exact coherent gradients $\phi_{\widehat\theta_{-v}}$, obeys the exact representation
\begin{equation}\label{eq:source-slice-cross-fitted-representation}
\widehat\Psi_{\mathrm{cf}}-F(\theta_0) =\sum_{v=1}^L\frac{n_v}{n} (\mathbb P_{n,v}-P)\phi_{\widehat\theta_{-v}}.
\end{equation}
It is therefore exactly unbiased for every $n$ and, under the standing fixed $L$ and fold-proportion conditions,
\begin{equation}\label{eq:source-slice-mean-square-rate}
\mathrm E\bigl[\{\widehat\Psi_{\mathrm{cf}}-F(\theta_0)\}^2\bigr]=O(n^{-1}).
\end{equation}
No convergence of the fitted source blocks is required, although every complementary chart block is assumed exactly equal to its truth.
\end{corollary}
\begin{remark}
Corollary~\ref{cor:source-slice-foldwise-unbiased} is an exact finite-sample statement conditional on every complementary chart block being fixed at its true value. It is an oracle comparison that isolates error in the source marginal. When the complementary blocks are estimated, their error is governed by Theorem~\ref{thm:coherent-one-step} and Proposition~\ref{prop:modular-eif-remainder} under the rate conditions stated there, rather than by the exact identity. Without convergence of the fitted source blocks the corollary does not by itself provide asymptotic normality, the efficient limit law, or a stable asymptotic variance in Theorem~\ref{thm:coherent-one-step}. The conclusion does not extend automatically to same-sample fitting or to a modular approximation of the canonical gradient.
\end{remark}
\begin{corollary}[Cross-fitting is optional on fixed finite support]\label{cor:same-sample-one-step}
Let $\widehat\theta$ be a same-sample estimator with $\widehat\theta\in\Omega_{\mathcal G}$ almost surely, so that the display below is defined on every sample, and suppose $\|\widehat\theta-\theta_0\|=o_p(n^{-1/4})$. Because $\mathcal K$ is a neighborhood of $\theta_0$, $\widehat\theta\in\mathcal K$ with probability tending to one, and the bounds of Proposition~\ref{prop:estimated-eif-stability} are used only there. Then
\[
\widehat\Psi_{\mathrm{os}} :=F(\widehat\theta)+\mathbb P_n\phi_{\widehat\theta}
\]
obeys \eqref{eq:one-step-asymptotic-linearity}.
\end{corollary}
\begin{remark}[The finite-dimensional likelihood benchmark]
Under the usual regular likelihood expansion, any consistent interior local solution $\widehat\theta$ of the likelihood score equations in a neighborhood of $\theta_0$ satisfies \eqref{eq:one-step-initial-rate}, and its same-sample score equations give $\mathbb P_n b_{K,\widehat\theta}^{\mathrm{nat}}=0$ for every $K$, so the coherent one-step correction $\mathbb P_n\phi_{\widehat\theta}$ is exactly zero and $\widehat\Psi_{\mathrm{os}}$ is the likelihood plug-in. Corollary~\ref{cor:same-sample-one-step} thus contains the efficient plug-in as the case of a null correction; its content lies in the non-likelihood initial fits described at the start of this section.
\end{remark}
\begin{proposition}[Centered modular approximations]\label{prop:modular-eif-remainder}
For a fold-specific fitted law $Q=P_\eta$ with $\eta\in\mathcal K$, retain the exact chart basis $b_{K,\eta}^{\mathrm{nat}}$ but replace $G_K^{\mathrm{nat}}(\eta)$ and $\gamma_K(\eta)$ by training-sample quantities $\widetilde G_K^{\mathrm{nat}}$ and $\widetilde\gamma_K$, with each $\widetilde G_K^{\mathrm{nat}}$ symmetric. Let $E_\kappa$ be the eigenvalue event on which $\lambda_{\min}(\widetilde G_K^{\mathrm{nat}})\geq\kappa/2$ for every $K$, and assume $\mathrm P(E_\kappa)\to1$. Set
\[
\widetilde\beta_K :=(\widetilde G_K^{\mathrm{nat}})^{-1}\widetilde\gamma_K \text{ on }E_\kappa, \quad \widetilde\beta_K :=\beta_K(\eta) \text{ on }E_\kappa^c, \qquad \phi_\eta^{\mathrm{mod}} :=\sum_K(b_{K,\eta}^{\mathrm{nat}})^\top\widetilde\beta_K,
\]
so that $\phi_\eta^{\mathrm{mod}}$ is defined on every sample and coincides with the exact coherent gradient $\phi_\eta$ off $E_\kappa$. Then $Q\phi_\eta^{\mathrm{mod}}=0$, and on $E_\kappa$ the exact coefficient identity
\begin{equation}\label{eq:modular-coefficient-identity}
\widetilde\beta_K-\beta_K(\eta) =(\widetilde G_K^{\mathrm{nat}})^{-1}\left[ \widetilde\gamma_K-\gamma_K(\eta) +\{G_K^{\mathrm{nat}}(\eta)-\widetilde G_K^{\mathrm{nat}}\} \beta_K(\eta) \right]
\end{equation}
holds. If
\[
\epsilon_n :=\max_K\{\|\widetilde G_K^{\mathrm{nat}}-G_K^{\mathrm{nat}}(\eta)\|_{\mathrm{op}} +\|\widetilde\gamma_K-\gamma_K(\eta)\|\},
\]
then $\|\phi_\eta^{\mathrm{mod}}-\phi_\eta\|_{L_2(P)}\leq C\epsilon_n$ on every sample: the bound is proved on $E_\kappa$, and off $E_\kappa$ its left side is zero. Write $(P-Q)(f):=Pf-Qf$ for any two probability or empirical measures. Replacing the coherent fold-specific EIF in \eqref{eq:cross-fitted-one-step} by $\phi_\eta^{\mathrm{mod}}$ adds exactly
\begin{equation}\label{eq:modular-eif-extra-remainder}
(\mathbb P_{n,v}-P)(\phi_\eta^{\mathrm{mod}}-\phi_\eta) +(P-Q)(\phi_\eta^{\mathrm{mod}}-\phi_\eta)
\end{equation}
in that fold. In conjunction with \eqref{eq:one-step-initial-rate}, the additional modular term is $o_p(n^{-1/2})$ if, uniformly over the fixed number of folds,
\begin{equation}\label{eq:modular-eif-rate}
\epsilon_n=o_p(1), \qquad \|\eta-\theta_0\|\epsilon_n=o_p(n^{-1/2}).
\end{equation}
If the estimated direction is not centered under $Q$, the additional term $Q\phi_\eta^{\mathrm{mod}}$ must be added to \eqref{eq:modular-eif-extra-remainder}; it is generally first order.
\end{proposition}
An empirical Gram matrix may be used as $\widetilde G_K^{\mathrm{nat}}$ provided its error is carried through \eqref{eq:modular-eif-rate}. The conditional argument in the proof computes it on the training sample; on fixed finite support the same conclusion holds for validation-dependent or same-sample coefficients, because $|(\mathbb P_{n,v}-P)(\phi_\eta^{\mathrm{mod}}-\phi_\eta)|\leq\|\mathbb P_{n,v}-P\|_1\,\|\phi_\eta^{\mathrm{mod}}-\phi_\eta\|_\infty=O_p(n^{-1/2})\,\epsilon_n$ needs no independence between the two factors. Likewise, a data-adaptively learned source-law rule defines a random target and is not covered by the theorem above.
\subsection{Model-valid targeting through the forward chart}
A targeted update must remain inside $\mathcal N_+(\mathcal G)$. Directly tilting a normalized district kernel does not guarantee this: in the example of Section~\ref{sec:counterexample} the unrestricted normalized-kernel score space has dimension $30$, whereas the district score space $\mathcal S_\Delta^{\mathrm{nat}}(P)$, the image of the district chart's differential, has dimension $23$. We instead lift the canonical gradient to the recursive-head coordinate chart.
Relative to the fixed native basis, define the aggregate synthesis map
\[
\iota_K^{\mathrm{nat}}:\mathbb R^{d_K^{\mathrm{nat}}}\longrightarrow U_{\mathcal G}, \qquad \iota_K^{\mathrm{nat}}c :=\sum_{j=1}^{d_K^{\mathrm{nat}}}c_j e_{Kj}^{\mathrm{nat}}.
\]
Thus $\iota_K^{\mathrm{nat}}$ is distinct from the intrinsic-block insertion $\iota_C:U_C\to U_{\mathcal G}$. Define the \emph{targeting coordinate field}
\begin{equation}\label{eq:targeting-coordinate-field}
\zeta(\theta) :=\sum_{K\in\mathcal D(\mathcal G)}\iota_K^{\mathrm{nat}}\beta_K(\theta) =\sum_{K\in\mathcal D(\mathcal G)} \iota_K^{\mathrm{nat}} [\{G_K^{\mathrm{nat}}(\theta)\}^{-1}\gamma_K(\theta)].
\end{equation}
For a mixture target, $\gamma_K=\alpha_K^{\mathrm{cmp}}+\alpha_K^\mu$; thus both the component and weight contributions enter the targeting field.
\begin{proposition}[Coordinate lift of the canonical gradient]\label{prop:targeting-coordinate-lift}
For every $\theta\in\Omega_{\mathcal G}$ at which the target is defined,
\[
\phi_\theta=J_{P_\theta}\zeta(\theta).
\]
The map $\zeta$ is $C^\infty$ on an open neighborhood of every compact subset of $\Omega_{\mathcal G}$ on which the target is defined; in particular the open-set hypothesis of Theorem~\ref{thm:least-favorable-chart-flow} holds automatically for the fixed-response and admissible mixture targets.
\end{proposition}
\begin{theorem}[Universal least-favorable chart flow]\label{thm:least-favorable-chart-flow}
Let $\mathcal K_0$ and $\mathcal K_1$ be compact sets satisfying
\[
\mathcal K_0\Subset\operatorname{int}(\mathcal K_1), \qquad \mathcal K_1\Subset\Omega_{\mathcal G},
\]
with $\mathcal K_1$ convex. Assume further that there is an open set $O$ with $\mathcal K_1\Subset O\subset\Omega_{\mathcal G}$ on which the target is defined and $\zeta\in C^1(O;U_{\mathcal G})$; the fixed-response and admissible independent-draw mixture targets satisfy this automatically. There is $\epsilon_0>0$, independent of $\eta\in\mathcal K_0$, such that the initial-value problem
\begin{equation}\label{eq:least-favorable-chart-ode}
\dot\vartheta_\eta(\epsilon) =\zeta\{\vartheta_\eta(\epsilon)\}, \qquad \vartheta_\eta(0)=\eta,
\end{equation}
has a unique solution for $|\epsilon|<\epsilon_0$, and that solution remains in $\mathcal K_1$. Throughout this interval,
\begin{equation}\label{eq:least-favorable-score-identity}
p_{\vartheta_\eta(\epsilon)}\in\mathcal N_+(\mathcal G), \qquad \frac{\mathrm d}{\mathrm d\epsilon} \log p_{\vartheta_\eta(\epsilon)}(x_V) =\phi_{\vartheta_\eta(\epsilon)}(x_V) \quad(x_V\in\mathfrak X_V).
\end{equation}
\end{theorem}
The path in Theorem~\ref{thm:least-favorable-chart-flow} is the finite-dimensional chart realization of a universal least-favorable one-dimensional submodel: its score equals the canonical gradient at every point, not only at the initial law \citep[Definition~1 and Appendix]{vanderlaan2016universal}.
\begin{corollary}[Exact empirical EIF equation]\label{cor:exact-empirical-eif-equation}
For an initial coordinate $\eta\in\mathcal K_0$, set
\[
\ell_n^\eta(\epsilon) :=\mathbb P_n\log p_{\vartheta_\eta(\epsilon)}.
\]
If $\widehat\epsilon$ is an interior stationary point---in particular, an interior local maximizer---of $\ell_n^\eta$, and $\theta^\star=\vartheta_\eta(\widehat\epsilon)$, then
\begin{equation}\label{eq:exact-empirical-eif-equation}
\mathbb P_n\phi_{\theta^\star}=0
\end{equation}
exactly.
\end{corollary}
The corollary is conditional on an interior local solution. The next theorem shows that the relevant root exists near zero with probability tending to one; no global concavity of the likelihood along the flow is asserted.
\begin{theorem}[Chart TMLE: local root, limit law, and one-step equivalence]\label{thm:chart-tmle}
Let $P=P_{\theta_0}$ and suppose $\theta_0\in\operatorname{int}(\mathcal K_0)$ and
\[
\sigma^2:=P\phi_{\theta_0}^2>0.
\]
Let $\widehat\theta$ be a same-sample initial estimator with $\widehat\theta\in\Omega_{\mathcal G}$ almost surely and, with probability tending to one,
\[
\widehat\theta\in\mathcal K_0, \qquad \|\widehat\theta-\theta_0\|=o_p(n^{-1/4}).
\]
Fix a sufficiently small deterministic $\delta\in(0,\epsilon_0)$. Let $E_n$ be the event that $\widehat\theta\in\mathcal K_0$ and the score $\epsilon\mapsto\mathbb P_n\phi_{\vartheta_{\widehat\theta}(\epsilon)}$ has strictly negative derivative on $[-\delta,\delta]$, is positive at $-\delta$, and is negative at $\delta$. This event has probability tending to one. On $E_n$, let $\widehat\epsilon$ be its unique root in $(-\delta,\delta)$; set $\widehat\epsilon=0$ on $E_n^c$. The root on $E_n$ is an interior strict local maximizer of $\ell_n^{\widehat\theta}$, and
\begin{equation}\label{eq:tmle-root-rate}
|\widehat\epsilon| =O_p\{n^{-1/2}+\|\widehat\theta-\theta_0\|\}.
\end{equation}
For a fixed $\theta_{\mathrm{ref}}\in\mathcal K_0$, define
\[
\theta^\star :=
\begin{cases}
\vartheta_{\widehat\theta}(\widehat\epsilon),&\text{on }E_n,\\
\widehat\theta,&\text{on }E_n^c\text{ if }\widehat\theta\in\mathcal K_0,\\
\theta_{\mathrm{ref}},&\text{otherwise},
\end{cases}
\qquad \widehat\Psi_{\mathrm{tmle}}:=F(\theta^\star).
\]
The empirical EIF equation \eqref{eq:exact-empirical-eif-equation} holds on $E_n$. The estimator is defined on every sample, with $p_{\theta^\star}\in\mathcal N_+(\mathcal G)$, and satisfies
\begin{align}
\|\theta^\star-\theta_0\|&=o_p(n^{-1/4}), \label{eq:tmle-updated-rate}\\
\sqrt n\{\widehat\Psi_{\mathrm{tmle}}-F(\theta_0)\} &=\frac1{\sqrt n}\sum_{i=1}^n\phi_{\theta_0}(X_i)+o_p(1). \label{eq:tmle-asymptotic-linearity}
\end{align}
Consequently,
\begin{equation}\label{eq:tmle-one-step-equivalence}
\widehat\Psi_{\mathrm{tmle}}-\widehat\Psi_{\mathrm{os}} =o_p(n^{-1/2}),
\end{equation}
where $\widehat\Psi_{\mathrm{os}}$ is the same-sample one-step estimator in Corollary~\ref{cor:same-sample-one-step}, formed from the same initial $\widehat\theta$ and, by the almost-sure chart requirement there, defined on every sample. Moreover, $\widehat\sigma^2_{\mathrm{tmle}}:=\mathbb P_n\phi_{\theta^\star}^2$ converges in probability to $\sigma^2$.
\end{theorem}
\begin{corollary}[Substitution and range preservation]\label{cor:tmle-substitution}
With the fallback convention in Theorem~\ref{thm:chart-tmle}, $p_{\theta^\star}\in\mathcal N_+(\mathcal G)$ on every sample, and $\widehat\Psi_{\mathrm{tmle}}$ is the target functional evaluated at that law. If the response transform satisfies $a\leq g\leq b$, then
\[
a\leq\widehat\Psi_{\mathrm{tmle}}\leq b.
\]
In particular, an indicator-response TMLE lies in $[0,1]$.
\end{corollary}
\begin{remark}[Null update at an interior likelihood fit]
If $\widehat\theta$ solves the full interior likelihood score equations, then $\mathbb P_n b_{K,\widehat\theta}^{\mathrm{nat}}=0$ for every $K$ and hence $\mathbb P_n\phi_{\widehat\theta}=0$. Thus zero is exactly the targeting score root. Under the local strictness in Theorem~\ref{thm:chart-tmle}, the near-zero root selected there is $\widehat\epsilon=0$, and the TMLE equals the MLE plug-in. Targeting is therefore not needed for efficiency in the regular fixed-dimensional MLE benchmark. Its roles here are to update coherent non-MLE, regularized, or approximately solved initial fits, to preserve the substitution property, and to provide a model-valid template for later regimes in which the MLE benchmark disappears.
\end{remark}
\begin{remark}[Straight-line and numerical implementations]\label{rem:straight-line-numerical-targeting}
The fixed-direction path $\eta+\epsilon\zeta(\eta)$ is model-valid for sufficiently small $\epsilon$, and its score at zero equals $\phi_\eta$. Away from zero its score is $J_{P_{\eta+\epsilon\zeta(\eta)}}\zeta(\eta)$, not generally the canonical gradient at the updated law. Iterative targeting must therefore recompute $\zeta$ after each update; convergence to an exact EIF root requires its own numerical argument.
Likewise, a discretized solution of \eqref{eq:least-favorable-chart-ode} solves the exact empirical equation only up to numerical error. To inherit \eqref{eq:tmle-asymptotic-linearity}, a terminal score residual by itself is not sufficient: it does not control numerical error transverse to the exact flow. It is sufficient, for example, to make the terminal coordinate error relative to $\theta^\star$ $o_p(n^{-1/2})$. Equivalently, one may require a numerical endpoint $\widetilde\theta$ to be $o_p(n^{-1/2})$-close to a point $\vartheta_{\widehat\theta}(\widetilde\epsilon)$ on the exact local flow with $|\widetilde\epsilon|\leq\delta$ with probability tending to one, and also require $\mathbb P_n\phi_{\widetilde\theta}=o_p(n^{-1/2})$. Writing $H_n(\epsilon):=\mathbb P_n\phi_{\vartheta_{\widehat\theta}(\epsilon)}$, the sup-norm line of \eqref{eq:local-eif-stability-bounds} gives $|H_n(\widetilde\epsilon)|\leq|\mathbb P_n\phi_{\widetilde\theta}|+C\|\widetilde\theta-\vartheta_{\widehat\theta}(\widetilde\epsilon)\|=o_p(n^{-1/2})$, and the slope bound in \eqref{eq:uniform-targeting-slope} of the Supplementary Material then gives $|\widetilde\epsilon-\widehat\epsilon|\leq4|H_n(\widetilde\epsilon)|/\sigma^2=o_p(n^{-1/2})$, after which bounded flow velocity completes the comparison with $\theta^\star$. The restriction $|\widetilde\epsilon|\leq\delta$ is needed: the flow exists on $(-\epsilon_0,\epsilon_0)$, whereas the negative-slope bound is established only on the smaller interval $[-\delta,\delta]$, so a root outside it is not excluded. Replacing $G_K^{\mathrm{nat}}$ or $\gamma_K$ by modular approximations similarly targets the modular score rather than $\phi_\theta$; the errors in Proposition \ref{prop:modular-eif-remainder} must then be carried into the targeting equation.
\end{remark}
\begin{remark}[Robustness boundary]
Neither the metric projection, the empirical EIF equation, nor the mixture product expansions imply double or multiple robustness: such a claim requires a separately proved drift factorization into products of errors of separately estimable nuisance functionals, which the exact remainder \eqref{eq:exact-one-step-remainder} is not, and the pure-curvature terms in \eqref{eq:smooth-mixture-second-derivative} remain. Section \ref{sec:sequential-union} proves such a factorization for a different, source-isolated fixed-node sequential construction; it does not strengthen the order-free one-step or TMLE claims of this section.
\end{remark}
\begin{example}[The universal flow need not preserve the source-only slice]\label{ex:tmle-source-slice-failure}
Consider the binary DAG $W\to X$ and $W\to Z$, with no edge between $X$ and $Z$. Source-standardize the intervention on $W$ and take $g(X,Z)=XZ$. Write
\[
\begin{gathered}
\pi=\mathrm P_\theta(W=1),\qquad r_w=\mathrm P_\theta(X=1\mid W=w),\\
s_w=\mathrm P_\theta(Z=1\mid W=w),\qquad m_w=r_ws_w.
\end{gathered}
\]
Then
\[
F(\theta)=(1-\pi)m_0+\pi m_1, \qquad \phi_\theta =m_W-F(\theta)+s_W(X-r_W)+r_W(Z-s_W).
\]
In these probability coordinates, the targeting field is
\[
\dot\pi=\pi(1-\pi)(m_1-m_0),\qquad \dot r_w=r_w(1-r_w)s_w, \qquad \dot s_w=s_w(1-s_w)r_w.
\]
Let $\theta_0$ be interior, write $r_{0w}:=r_w(\theta_0)$ and $s_{0w}:=s_w(\theta_0)$, and let $\eta$ differ from it only through $\pi_\eta\ne\pi_0$, with $m_1\ne m_0$, where $\pi_\eta=\mathrm P_\eta(W=1)$, $\pi_0=\mathrm P_{\theta_0}(W=1)$, and $P_0:=P_{\theta_0}$. Although Proposition \ref{prop:source-block-drift-cancellation} gives $R_2(\eta,\theta_0)=0$, one has
\[
P_0\phi_\eta=(\pi_0-\pi_\eta)(m_1-m_0)\ne0.
\]
Because $m_1\ne m_0$ and $\theta_0$ is interior, $P_0\phi_{\theta_0}^2>0$. Continuity and the implicit-function theorem therefore give, for $\eta$ sufficiently close to $\theta_0$, a nonzero local population EIF root along the universal flow. Both $r_w$ and $s_w$ move off their true values in the same direction, for either sign of the root time, because $\dot r_w=r_w(1-r_w)s_w>0$ and $\dot s_w=s_w(1-s_w)r_w>0$ at interior coordinates. At the resulting population root $\theta^\dagger$, writing $r_w^\dagger:=r_w(\theta^\dagger)$ and $s_w^\dagger:=s_w(\theta^\dagger)$, the root equation gives
\[
R_2(\theta^\dagger,\theta_0) =F(\theta^\dagger)-F(\theta_0) =-\sum_w\mathrm P_{\theta_0}(W=w) (r_w^\dagger-r_{0w})(s_w^\dagger-s_{0w})<0.
\]
Every factor pair has the same sign, so each summand is positive.
Thus exact cancellation at the initial source-only error does not imply exact robustness after universal-flow targeting. Under Theorem \ref{thm:chart-tmle}'s local full-chart rate the displayed drift is still second order; arbitrary fixed source misspecification is outside that theorem.
\end{example}
\subsection{One-step correction and chart targeting: different exact guarantees}\label{subsec:one-step-tmle-contrast}
The coherent one-step estimator and chart-flow TMLE use the same order-free canonical gradient. Under their common local regime, the chart-flow TMLE and the same-sample coherent one-step estimator formed from the same initial law are asymptotically equivalent; both one-step versions have the same efficient first-order expansion. Thus there is no first-order efficiency trade-off between them. Their exact guarantees nevertheless differ, as summarized in Table~\ref{tab:one-step-tmle-contrast}.
\begin{table}[!t]
\centering
\caption{Proved comparison of the coherent one-step estimator and the exact chart-flow TMLE.}\label{tab:one-step-tmle-contrast}
\small
\begin{tabular}{@{}
>{\raggedright\arraybackslash}p{0.29\textwidth}
>{\raggedright\arraybackslash}p{0.31\textwidth}
>{\raggedright\arraybackslash}p{0.28\textwidth}@{}}
\toprule
Property & Coherent one-step & Chart-flow TMLE \\
\midrule
Graph-order requirement
& No active-block chain or sequential vertex order for the classes (I1)--(I3)
& No active-block chain or sequential vertex order for the classes (I1)--(I3) \\
First-order efficiency
& Yes, under the local full-chart conditions of
Theorem~\ref{thm:coherent-one-step}
& Yes, under the local root conditions of
Theorem~\ref{thm:chart-tmle} \\
Exact marginal-source-block property
& Population cancellation; for the cross-fitted coherent version, exact
foldwise conditional and unconditional unbiasedness under the source-only
conditions of Proposition
\ref{prop:source-block-drift-cancellation} and Corollary
\ref{cor:source-slice-foldwise-unbiased}
& No general identity; the flow can leave the source-only slice by Example
\ref{ex:tmle-source-slice-failure} \\
Exact empirical EIF equation
& The additive correction constructs no updated law and does not enforce
this equation; at an interior full-chart likelihood fit, the initial law
already solves it and the correction is zero
& Yes at an interior stationary point of the exact flow, by Corollary
\ref{cor:exact-empirical-eif-equation} \\
Substitution and response-range preservation
& Not guaranteed by the additive correction; it holds in null-update cases
such as an interior full-chart likelihood fit
& Yes, including the fallback, by Corollary
\ref{cor:tmle-substitution} \\
\bottomrule
\end{tabular}
\end{table}
The efficiency and source-block entries concern different regimes: efficiency requires full-chart convergence to $\theta_0$ (Theorems~\ref{thm:coherent-one-step} and~\ref{thm:chart-tmle}), whereas the exact cancellation allows an arbitrary interior source-block error only because every complementary block is held at its true value (Proposition~\ref{prop:source-block-drift-cancellation} and the remark following it). Neither full-chart efficiency under fixed source misspecification nor an exact source-robust TMLE follows: the exact chart flow enforces model membership and the inherited response range, but its nonlinear update can leave the source-only slice (Example~\ref{ex:tmle-source-slice-failure}). This is a limited contrast between one-sided drift cancellation and substitution, not an efficiency-versus-double- or multiple-robustness theorem.
\begin{remark}[Comparison with existing doubly robust ADMG estimators]
Under primal fixability, Theorem~7 of \cite{bhattacharya2022semiparametric} establishes variational independence of the portions of the observed-data law used by the primal and dual IPW representations, and \cite[Theorem~9]{bhattacharya2022semiparametric} establishes double robustness of augmented primal IPW under the associated union model; neither result requires mb-shieldedness. Mb-shieldedness has a different role in their analysis: it makes the equality restrictions DAG-like and supplies the vertexwise tangent decomposition used for the constrained-model efficient projection \citep[Theorem~2, Lemma~3, and Theorem~12]{bhattacharya2022semiparametric}, and it is not the source of their double robustness. Extending the efficiency geometry beyond mb-shielded ADMGs therefore neither forecloses a doubly robust construction nor confers one on the order-free estimators studied here.
The two contributions are complementary but logically distinct. Those authors establish union-model double robustness for their augmented primal IPW estimator in the primal-fixable setting; we establish an order-free canonical gradient and model-valid efficient inference in the strictly positive finite-state nested model for the intervention classes (I1)--(I3) on arbitrary ADMGs. Neither result implies the other, and we prove no double- or multiple-robustness guarantee for that full class: the transitionwise union result of Section~\ref{sec:sequential-union} holds only under Assumption~\ref{ass:mr-sequential-transport}, and Proposition~\ref{prop:source-block-drift-cancellation} is the exact one-sided source-block result available here, with the scope stated in the remark following it, not an analogue of the primal/dual union model. Such robustness is not needed for first-order efficiency in the likelihood benchmark described at the start of this section, but it remains meaningful for regularized or approximately solved fits, for which it requires a drift factorization and rate conditions appropriate to those fits. The 30-versus-23 dimension gap in Section~\ref{sec:counterexample} rules out arbitrary normalized district-kernel tilting, not every alternative nuisance decomposition or factorized drift; recursive-head-compatible factorized drifts over the full non-mb-shielded class are not characterized here.
\end{remark}
\section{A restricted sequential union-model alternative}\label{sec:sequential-union}
The order-free estimator in Section~\ref{sec:targeted} is built from the canonical gradient and applies to every intervention covered there. This section constructs a second estimator for a narrower fixed-node subclass. Its purpose is different: it preserves a transitionwise factorization of the population drift and is consequently consistent on a union of nuisance correctness regimes. At the all-correct intersection its influence function is generally not canonical. Thus the comparison below is between two specific constructions---an efficient recursive-head estimator and a restricted union-model sequential estimator; what that comparison does and does not establish is stated after Proposition~\ref{prop:sequential-strict-variance-gap}.
The restrictions stated below---source isolation, an acyclic literal block order with source precedence, and source-tail coverage---are the conditions this ordinary-propensity construction uses; they are not claimed to be necessary for every possible robust estimator. The next subsection shows by example that three natural shortcuts fail, and the examples motivate the conditions that follow; the full-history transport identity is then proved and every estimation claim is derived from it.
\subsection{Three shortcuts that fail}\label{sec:shortcuts}
The first shortcut deletes earlier source values from the histories of later propensities. The product of the resulting inverse propensities need not have mean one, so it need not define a change of probability. At the exact binary law for $A_1\prec A_2$, the inverse-probability product anchored at $(1,1)$, using the two marginal denominators, has expectation $1.6$, whereas keeping $A_1=1$ in the second denominator replaces $0.5$ by $0.8$ and makes the expectation one.
The second shortcut allows an intervention source inside a transition block by restricting the order only among active vertices. The graph is $L\leftrightarrow Y$ with $L\to A\to Y$ under $\operatorname{do}(A=1)$. The active graph has the single district $\{L,Y\}$, so the literal block-order condition is vacuous after restricting to the active vertices. Nevertheless every observational topological order has $L\prec A\prec Y$: the source lies inside the active block. A candidate weight may then depend on $L$, part of the current transition, and the true district residual need not be centered against that weight.
The third shortcut substitutes conditioning for fixing at a source that is bidirected-confounded with its district, even when the source precedes the block. In the graph $A\leftrightarrow Y$, fixing $A$ under a node intervention gives the kernel $q_Y(y)=p(y)$, whereas conditioning on $A$ gives $p(y\mid a)$. At a strictly positive binary law the two probabilities of $Y=1$ are $0.5$ and $0.8$, respectively, so the ordinary weight $\mathbf 1(A=1)/\mathrm P(A=1)$ transports to the wrong law. Placing the source before the block does not exclude this example: what removes $A$ from a bidirected district is the fixing divisor $p(a\mid y)$, not a propensity measurable with respect to the past. Section~\ref{app:B-shortcuts} of the Supplementary Material states the three exact laws and carries the examples in full: the anchored expectation $1.6$ rather than one, the drift $0.05625$ rather than zero, and the fixed-against-conditioned probabilities $0.5$ and $0.8$.
The examples motivate the three conditions imposed below: propensities conditioned on the full strict past, precedence of every source relevant to a transition over that transition, and, for the present theorem, intervention sources that are singleton observational districts. The last condition is sufficient rather than necessary; removing it requires a fixing-based or primal--dual weighting theory and is outside the present paper.
\subsection{The sequential scope conditions}
Fix a \emph{system-wide fixed-node intervention} $\mathsf I$ from (I1), put $A=A_{\mathsf I}$, $V^*=V\setminus A$, and let $\mathcal H=\mathcal G_{V^*}$. Write $\beta\in\mathfrak X_A$ for the intervention's assigned value vector, with coordinates $\beta_a$, $a\in A$.
\paragraph*{Response set and margins in this section}
Throughout this section the response set of Section~\ref{sec:intervention-bridge} is $Y=V^*$, so $R_{\mathsf I}=V^*$, $\mathcal H_{\mathsf I}=\mathcal H$, and $P^{\mathsf I}_\theta=\Gamma_{\mathsf I}(P_\theta)$ is the configured law on all of $V^*$. A target $g$ that depends only on the coordinates in a subset $Y_0\subseteq V^*$ is read as a function on $\mathfrak X_{V^*}$ through coordinate projection, so that $\Psi_{\mathsf I,g}(P)=\mathrm E_{\Gamma_{\mathsf I}(P)}\{g(X_{Y_0})\}$. This agrees with the ancestry-restricted convention of Section~\ref{sec:intervention-bridge}: with $R_0:=\operatorname{an}_{\mathcal H}(Y_0)$, the $R_0$ margin of $\Gamma_{\mathsf I}(P)$ equals, at every $P\in\mathcal N_+(\mathcal G)$, the configured law obtained by applying Theorem~\ref{thm:normalized-intervention-bridge} to the same node intervention with response set $Y_0$. Consequently the target functional and its canonical gradient in Theorem~\ref{thm:intervention-chart-eif} are the same under the two conventions; the proof is in Section~\ref{app:proofs-sequential} of the Supplementary Material.
The first restriction is
\begin{equation}\label{eq:source-isolation}
\operatorname{dis}_{\mathcal G}(a)=\{a\} \qquad(a\in A).
\end{equation}
For every active district $D\in\mathcal D(\mathcal H)$, choose $A_D\subseteq A$ so that the following \emph{source-tail coverage} condition holds:
\begin{equation}\label{eq:source-tail-coverage}
\operatorname{pa}_{\mathcal G}(D)\setminus V^*\subseteq A_D.
\end{equation}
The minimal choice is $A_D=\operatorname{pa}_{\mathcal G}(D)\setminus V^*$, the set of excluded parents of $D$: exactly the sources occupying an external tail slot of the native $D$-factor. Larger source sets are permitted. The node intervention assigns the value sub-vector $\beta_D:=(\beta_a)_{a\in A_D}$ to the sources in $A_D$.
To make the order restriction checkable, form a directed block graph whose vertices are
\begin{equation}\label{eq:mr-block-partition}
\mathcal B_{\mathsf I} :=\mathcal D(\mathcal H)\cup\bigl\{\{a\}:a\in A\bigr\}.
\end{equation}
Under source isolation no bidirected edge meets $A$, so the active districts are the districts of $\mathcal G$ other than the singletons $\{a\}$, and $\mathcal B_{\mathsf I}$ is the district partition of $\mathcal G$ itself. Insert every observational directed arrow between distinct blocks and, as a formal precedence constraint, insert $\{a\}\to D$ whenever $a\in A_D$. With the minimal source sets every precedence arrow duplicates an observational arrow, so the block graph is the district quotient of $\mathcal G$ (Section~\ref{sec:geometry}) and its acyclicity is acyclicity of that quotient; each additional source placed in some $A_D$ adds one precedence constraint, which is the reason for the general formulation. The precedence arrows constrain an order; they are not causal edges.
\begin{proposition}[MR-compatible block-order screen]\label{prop:mr-block-order-screen}
The directed block graph in \eqref{eq:mr-block-partition} is acyclic if and only if the observational directed graph has a topological order $\prec$ in which every block in $\mathcal B_{\mathsf I}$ is a literal interval and every $a\in A_D$ precedes every vertex of $D$.
\end{proposition}
\begin{assumption}[MR-compatible sequential scope]\label{ass:mr-sequential-transport}
The intervention obeys source isolation \eqref{eq:source-isolation} and source-tail coverage \eqref{eq:source-tail-coverage}. The directed block graph in \eqref{eq:mr-block-partition} is acyclic. Fix an observational topological order $\prec$ supplied by Proposition~\ref{prop:mr-block-order-screen}; thus every block is a literal interval and every $a\in A_D$ precedes every vertex of $D$.
\end{assumption}
For the remainder of this section work under Assumption~\ref{ass:mr-sequential-transport} and write the active districts in their induced order as $D_1,\ldots,D_m$. Put
\[
B_0:=\varnothing, \qquad B_k:=\bigcup_{j=1}^kD_j.
\]
For $v\in V$, let $\operatorname{Past}(v)=\{u:u\prec v\}$. If $A_k:=A_{D_k}=\{a_{k1}\prec\cdots\prec a_{kr_k}\}$, put $\beta_{kj}:=\beta_{a_{kj}}$ and $\beta_k:=(\beta_{k1},\ldots,\beta_{kr_k})$, and define the full-history propensities and ladder
\begin{equation}\label{eq:full-history-propensity-ladder}
\begin{aligned}
e_{kj,P}(X_{\operatorname{Past}(a_{kj})}) &:=\mathrm P_P\!\left( X_{a_{kj}}=\beta_{kj} \mid X_{\operatorname{Past}(a_{kj})} \right),\\
W_{k,P}^{\mathrm{pre}} &:=\prod_{j=1}^{r_k} \frac{\mathbf 1\{X_{a_{kj}}=\beta_{kj}\}} {e_{kj,P}(X_{\operatorname{Past}(a_{kj})})}.
\end{aligned}
\end{equation}
Thus earlier selected sources remain in every later strict past.
The \emph{full history} of the $k$th transition is $\operatorname{Past}(D_k):=\bigcap_{v\in D_k}\operatorname{Past}(v)$, the set of vertices preceding the entire block $D_k$; let $\mathcal F_k^-:=\sigma(X_{\operatorname{Past}(D_k)})$. By Proposition~\ref{prop:mr-block-order-screen}, $A_k\subseteq\operatorname{Past}(D_k)$ and $B_{k-1}\subseteq\operatorname{Past}(D_k)$, so $W_{k,P}^{\mathrm{pre}}$ and $X_{B_{k-1}}$ are $\mathcal F_k^-$-measurable. A history configuration $x_{\operatorname{Past}(D_k)}$ is \emph{compatible with the intervention} when $x_{A_k}=\beta_k$. The next lemma proves, rather than assumes, the graphical-to-statistical transition identity on which the sequential theory depends.
\begin{lemma}[Source-isolated active-block slice]\label{lem:source-isolated-active-block-slice}
Under Assumption~\ref{ass:mr-sequential-transport}, let $K_{k,\theta}^{\mathsf I}$ be the $k$th conditional kernel of $P_\theta^{\mathsf I}=\Gamma_{\mathsf I}(P_\theta)$ in the active-district order. Then, for every $\theta\in\Omega_{\mathcal G}$, every $x_{D_k}$, and every history configuration $x_{\operatorname{Past}(D_k)}$ compatible with the intervention,
\begin{equation}\label{eq:mr-sequential-kernel-interface}
\mathrm P_\theta\!\left( X_{D_k}=x_{D_k} \mid X_{\operatorname{Past}(D_k)}=x_{\operatorname{Past}(D_k)} \right) =K_{k,\theta}^{\mathsf I} (x_{D_k}\mid x_{B_{k-1}}).
\end{equation}
The identity is pointwise in all displayed arguments and global on the positive chart, not merely local at $P_0$.
\end{lemma}
This lemma, rather than acyclicity alone, is the transport statement on which the sequential theory rests; the three examples of the previous subsection show that full strict-past propensities, source precedence, and source isolation cannot simply be omitted from this ordinary-propensity construction.
Strict positivity on the finite support makes every propensity in the first line of \eqref{eq:full-history-propensity-ladder} positive. Uniform bounds needed for estimation follow after restricting fitted laws to a compact subset of the positive chart, as in Proposition~\ref{prop:estimated-eif-stability}.
By backward conditioning in the ladder line of \eqref{eq:full-history-propensity-ladder},
\begin{equation}\label{eq:corrected-ladder-normalization}
P W_{k,P}^{\mathrm{pre}}=1.
\end{equation}
The normalization \eqref{eq:corrected-ladder-normalization} is not yet a transport statement. Reweighting by $W_{k,P}^{\mathrm{pre}}$ enforces the assigned values $\beta_k$ of the sources relevant to the $k$th transition, but it can leave the preceding active variables $X_{B_{k-1}}$ under a marginal different from $P^{\mathsf I}_{B_{k-1}}$. Because $W_{k,P}^{\mathrm{pre}}$ is $\mathcal F_k^-$-measurable, reweighting by it does not change the conditional law of $X_{D_k}$ given the full history at any history configuration compatible with the intervention, and Lemma~\ref{lem:source-isolated-active-block-slice} identifies that conditional with the target kernel $K_{k,\theta}^{\mathsf I}(\cdot\mid x_{B_{k-1}})$. What remains to be corrected is the marginal of $X_{B_{k-1}}$; this is the role of the prefix ratio defined next, which replaces that marginal by $P^{\mathsf I}_{B_{k-1}}$ and equals one for the first transition.
Let $\widetilde P_{k,P}$ denote the resulting probability law, $\widetilde P_{k,P}f:=P(W_{k,P}^{\mathrm{pre}}f)$. On the fixed finite support, strict positivity implies mutual absolute continuity of its $B_{k-1}$ marginal and $P_{B_{k-1}}^{\mathsf I}$. Define
\begin{equation}\label{eq:sequential-prefix-bridge}
\Pi_{k-1,P} :=\frac{\mathrm dP_{B_{k-1}}^{\mathsf I}} {\mathrm d\widetilde P_{k,P,B_{k-1}}}(X_{B_{k-1}}), \qquad H_{k,P}:=W_{k,P}^{\mathrm{pre}}\Pi_{k-1,P},
\end{equation}
with $\Pi_{0,P}\equiv1$.
\begin{lemma}[Transport and robust centering]\label{lem:sequential-transport-centering}
Under Assumption~\ref{ass:mr-sequential-transport}, for every bounded $f(X_{B_k})$,
\begin{equation}\label{eq:sequential-prefix-transport}
P\{H_{k,P}f(X_{B_k})\}=P^{\mathsf I}f(X_{B_k}).
\end{equation}
Define $Q_{m,P}:=g(X_Y)$ and, recursively,
\begin{equation}\label{eq:sequential-outcome-recursion}
Q_{k-1,P} :=\mathrm E_{P^{\mathsf I}}(Q_{k,P}\mid X_{B_{k-1}}), \qquad R_{k,P}:=Q_{k,P}-Q_{k-1,P}.
\end{equation}
If
\begin{equation}\label{eq:admissible-candidate-weight}
\bar H_k=\mathbf 1\{X_{A_k}=\beta_k\}\bar G_k, \qquad \bar G_k\ \text{is }\mathcal F_k^-\text{-measurable},
\end{equation}
then
\begin{equation}\label{eq:robust-transition-centering}
P(\bar H_kR_{k,P})=0.
\end{equation}
In particular, the true $H_{k,P}$ is an admissible candidate weight.
\end{lemma}
\subsection{Exact drift and transitionwise union model}
Let $P_0\in\mathcal N_+(\mathcal G)$ denote the true law, with probability mass function $p_0$; as above, $P_0f=\mathrm E_{P_0}f$.
For arbitrary $B_k$-measurable candidates $\bar Q_k$, impose only $\bar Q_m=g(X_Y)$ and let $\bar R_k=\bar Q_k-\bar Q_{k-1}$; $\bar Q_0$ is a constant. Candidate weights always obey \eqref{eq:admissible-candidate-weight}.
\begin{theorem}[Exact factorized sequential drift]\label{thm:exact-sequential-drift}
At $P_0$, write $H_{0k}=H_{k,P_0}$, $R_{0k}=R_{k,P_0}$, $Q_{k,0}:=Q_{k,P_0}$, and $\psi_0=P_0^{\mathsf I}g$. Then
\begin{equation}\label{eq:exact-sequential-drift}
P_0\!\left[ \bar Q_0-\psi_0+\sum_{k=1}^m\bar H_k\bar R_k \right] =\sum_{k=1}^m P_0\{(\bar H_k-H_{0k})(\bar R_k-R_{0k})\}.
\end{equation}
\end{theorem}
\paragraph*{Sampling and candidate-nuisance conditions}
Let $X_1,\ldots,X_n$ be i.i.d. from $P_0$ on the fixed finite support. Let $I_1,\ldots,I_L$ be fixed, data-independent validation folds, with $L$ fixed and $n_\ell/n\to\rho_\ell\in(0,1)$, and let every fitted nuisance with superscript $(-\ell)$ be measurable with respect to the observations outside $I_\ell$. Each $\widehat Q_k^{(-\ell)}$ is $B_k$-measurable, with $\widehat Q_0^{(-\ell)}$ constant and $\widehat Q_m^{(-\ell)}=g(X_Y)$; each fitted weight satisfies \eqref{eq:admissible-candidate-weight} foldwise; and
\[
\max_{1\leq\ell\leq L}\Bigl\{\max_{1\leq k\leq m}\|\widehat H_k^{(-\ell)}\|_\infty +\max_{0\leq k\leq m}\|\widehat Q_k^{(-\ell)}\|_\infty\Bigr\} =O_p(1).
\]
Write $\widehat R_k^{(-\ell)}:=\widehat Q_k^{(-\ell)}-\widehat Q_{k-1}^{(-\ell)}$ and define
\begin{equation}\label{eq:sequential-cross-fitted-estimator}
\widehat\psi_{\mathrm{seq}} :=\sum_{\ell=1}^L\frac{n_\ell}{n}\, \mathbb P_{n,\ell}\!\left[ \widehat Q_0^{(-\ell)} +\sum_{k=1}^m \widehat H_k^{(-\ell)} \{\widehat Q_k^{(-\ell)}-\widehat Q_{k-1}^{(-\ell)}\} \right].
\end{equation}
For $\omega=(\omega_1,\ldots,\omega_m)\in\{H,R\}^m$, define the population correctness regime
\begin{equation}\label{eq:transition-union-regime}
\mathcal U_\omega
:=\bigcap_{k=1}^m \begin{cases}
\{\bar H_k=H_{0k}\},&\omega_k=H,\\
\{\bar R_k=R_{0k}\},&\omega_k=R.
\end{cases}
\end{equation}
The exponent in \eqref{eq:transition-union-regime} counts \emph{choice vectors}: for each of the $m$ active-district transitions, either the weight or the outcome residual is nominated as correct, where weight correctness concerns the complete transported weight $H_{0k}=W_{k,P_0}^{\mathrm{pre}}\Pi_{k-1,P_0}$, prefix ratio included. Because $\bar R_k$ and $\bar R_{k+1}$ both involve $\bar Q_k$, the regimes need not be distinct, so $2^m$ is an upper bound on their number; they are nuisance-correctness regimes for candidate functions, not necessarily distinct or variation-independent law-indexed submodels of $\mathcal N_+(\mathcal G)$. The count is available because the drift \eqref{eq:exact-sequential-drift} is a sum of per-transition products with no term coupling distinct transitions, which Theorem~\ref{thm:exact-sequential-drift} obtains from Assumption~\ref{ass:mr-sequential-transport} together with Lemmas~\ref{lem:source-isolated-active-block-slice} and~\ref{lem:sequential-transport-centering} and the admissible candidate structure \eqref{eq:admissible-candidate-weight}; it does not follow from the product identification formula \eqref{eq:imported-intervention-product} alone. Throughout, \emph{union model} is shorthand for the union $\bigcup_{\omega\in\{H,R\}^m}\mathcal U_\omega$ of these regimes; it is not a law-indexed statistical submodel of $\mathcal N_+(\mathcal G)$.
Correctness of a residual is correctness of a difference: $\bar R_k=R_{0k}$ holds exactly when the errors $\bar Q_k-Q_{k,0}$ and $\bar Q_{k-1}-Q_{k-1,0}$ coincide. Adjacent residuals share a candidate function; each residual-correctness clause constrains $\bar Q_k-\bar Q_{k-1}$, and in general neither adjacent regression need be individually correct. For $m=1$, $R_{01}=g(X_Y)-\psi_0$, and correctness of the residual is the scalar condition $\bar Q_0=\psi_0$. Three things are kept distinct: the algebraic regimes $\mathcal U_\omega$, which concern population candidates; the estimation conditions of Theorem~\ref{thm:sequential-multiple-robustness}, which concern the fitted nuisances; and a fitting procedure producing candidates that lie in a regime, which this paper does not specify. The fixtures of Section~\ref{app:B-sequential} of the Supplementary Material verify the algebra of the regimes at an exact law by offsetting the true nuisances; they do not demonstrate an implemented misspecification-robust estimator.
\begin{theorem}[Transitionwise multiple robustness and intersection theory]\label{thm:sequential-multiple-robustness}
Under the sampling and candidate-nuisance conditions above, the following hold.
\begin{enumerate}[label=(\roman*)]
\item \emph{Consistency on the union.} If, uniformly over folds,
\begin{equation}\label{eq:sequential-consistency-product}
\sum_{k=1}^m \|\widehat H_k-H_{0k}\|_{P_0,2} \|\widehat R_k-R_{0k}\|_{P_0,2}=o_p(1),
\end{equation}
then \eqref{eq:sequential-cross-fitted-estimator} is consistent. In particular, this holds when, for every transition, one nuisance is consistent and the other remains bounded; exact membership in any $\mathcal U_\omega$ makes the population drift zero.
\item \emph{Inference at the all-correct intersection.} Define
\begin{equation}\label{eq:sequential-influence-function}
\phi_{\mathrm{seq}} :=Q_{0,0}-\psi_0+\sum_{k=1}^mH_{0k}R_{0k}
\end{equation}
and the foldwise estimated summand $\widehat M^{(-\ell)} :=\widehat Q_0^{(-\ell)} +\sum_{k=1}^m\widehat H_k^{(-\ell)}\widehat R_k^{(-\ell)}$. If, in addition to \eqref{eq:sequential-consistency-product}, $\widehat M^{(-\ell)}$ converges in $L_2(P_0)$ to $\psi_0+\phi_{\mathrm{seq}}$ uniformly over folds and the left-hand side of \eqref{eq:sequential-consistency-product} is $o_p(n^{-1/2})$, then
\begin{equation}\label{eq:sequential-asymptotic-linearity}
\widehat\psi_{\mathrm{seq}}-\psi_0 = (\mathbb P_n-P_0)\phi_{\mathrm{seq}}+o_p(n^{-1/2}),
\end{equation}
and, writing $\ell(i)$ for the index of the fold containing observation $i$, the variance estimator
\[
\widehat\sigma^2_{\mathrm{seq}} :=\frac1n\sum_{i=1}^n \bigl\{\widehat M^{(-\ell(i))}(X_i) -\widehat\psi_{\mathrm{seq}}\bigr\}^2
\]
consistently estimates $P_0\phi_{\mathrm{seq}}^2$.
\end{enumerate}
\end{theorem}
For $m=1$ the terminal function is the known response, $\bar Q_1=g(X_Y)$, so the union statement in (i) has two regimes: correctness of the weight ($\bar H_1=H_{01}$), which is consistency of inverse-probability weighting, and correctness of the residual $\bar R_1=g(X_Y)-\bar Q_0$, which is exactly the equality $\bar Q_0=\psi_0$ of the scalar initial value with the target. No covariate-dependent fitted outcome regression enters at $m=1$; the only fitted outcome quantity is that scalar, and its population correctness is a statement about the candidate, distinct from the estimation and rate conditions of the theorem. The exact two-transition example in \eqref{eq:two-transition-rational-sem}--\eqref{eq:two-transition-both-wrong-fixture} of Section~\ref{app:B-sequential} of the Supplementary Material has both sides of the drift identity equal, and nonzero, under a state-dependent joint misspecification, with the drift carried entirely by the second transition; under the same law the telescoping-constant construction verifies all four correctness regimes, and the three controls in \eqref{eq:two-transition-control-fixtures} separately activate the first drift summand, the second, and both. That section states every offset and weight scale explicitly, so that each calculation is reproducible from the paper.
Part (ii) is intentionally an intersection result. Under fixed misspecification of the unneeded side of a union component, consistency may remain true, but a root-$n$ limit requires separate rates and its influence function generally depends on the nuisance probability limits. Union-model consistency must not be reported as automatically efficient or as having the all-correct influence function \eqref{eq:sequential-influence-function}.
\subsection{Projection and the efficiency comparison}
\begin{theorem}[Canonical projection of the sequential influence function]\label{thm:sequential-projection-comparison}
Assume that the transition identity \eqref{eq:mr-sequential-kernel-interface} holds throughout a positive chart neighborhood of $P_0$, as it does under Assumption~\ref{ass:mr-sequential-transport} by Lemma~\ref{lem:source-isolated-active-block-slice}. Then $\phi_{\mathrm{seq}}$ in \eqref{eq:sequential-influence-function} is an influence function for the fixed-response mean in the nested Markov model, and the order-free canonical gradient from Theorem~\ref{thm:intervention-chart-eif} satisfies
\begin{equation}\label{eq:sequential-canonical-projection}
\phi_{\mathrm{eff}} =\Pi_{\mathcal T_{P_0}\mathcal N(\mathcal G)}\phi_{\mathrm{seq}} =\sum_{K\in\mathcal D(\mathcal G)} (b_K^{\mathrm{nat}})^\top\{G_K^{\mathrm{nat}}(P_0)\}^{-1} P_0(b_K^{\mathrm{nat}}\phi_{\mathrm{seq}}).
\end{equation}
Consequently,
\begin{equation}\label{eq:sequential-pythagoras}
P_0\phi_{\mathrm{seq}}^2 =P_0\phi_{\mathrm{eff}}^2 +\|\phi_{\mathrm{seq}}-\phi_{\mathrm{eff}}\|_{P_0,2}^2.
\end{equation}
Equality with the efficiency bound holds if and only if $\phi_{\mathrm{seq}}\in\mathcal T_{P_0}\mathcal N(\mathcal G)$.
\end{theorem}
\begin{proposition}[A strict finite-state variance gap]\label{prop:sequential-strict-variance-gap}
For the binary graph of Section~\ref{sec:counterexample},
\[
W\to Z\leftarrow B, \qquad Z\to Y, \qquad A\leftrightarrow B, \quad A\leftrightarrow Z, \quad B\leftrightarrow Y,
\]
the fixed intervention $\operatorname{do}(W=0)$ satisfies the restricted sequential conditions. At the strictly positive rational law specified by \eqref{eq:verma-rational-sem} in Appendix~\ref{app:prototype-ledger} of the Supplementary Material, with $g(Y)=Y$,
the target and the sequential and efficient influence-function variances are $\psi_0=0.54725$, $P_0\phi_{\mathrm{seq}}^2\approx0.49553$ and $P_0\phi_{\mathrm{eff}}^2\approx0.28803$. The squared projection residual is $\|\phi_{\mathrm{seq}}-\phi_{\mathrm{eff}}\|_{P_0,2}^2\approx0.20750>0$; the exact rational constants are established in the proof of this proposition, in Section~\ref{app:proofs-sequential} of the Supplementary Material.
Here
\[
\phi_{\mathrm{seq}} =\frac{\mathbf 1(W=0)}{\mathrm P_{P_0}(W=0)}(Y-\psi_0)
\]
satisfies all $24$ recursive-head coordinate Riesz equations but lies outside the nested-model tangent space. The canonical gradient is its strict metric projection.
\end{proposition}
In this one-transition example $\phi_{\mathrm{seq}}$ is a pure inverse-probability-weighted residual: since $B_0=\varnothing$, it contains no covariate-dependent outcome regression, although the estimator remains augmented through $\widehat Q_0$ and keeps the two regimes of Theorem~\ref{thm:sequential-multiple-robustness}(i), which for $m=1$ are weight correctness and the scalar condition $\bar Q_0=\psi_0$ noted after that theorem. Its canonical projection uses the recursive-head restrictions and removes the orthogonal residual recorded in Proposition~\ref{prop:sequential-strict-variance-gap}; by \eqref{eq:sequential-pythagoras}, equal-level asymptotic Wald intervals are wider for the sequential estimator by the factor $(P_0\phi_{\mathrm{seq}}^2/P_0\phi_{\mathrm{eff}}^2)^{1/2}=1.3116\ldots$, about $31\%$ (Section~\ref{sec:finite-sample}). Two sources of gain are distinct. Here the nested and ordinary Markov models coincide (Section~\ref{sec:counterexample}), so the gain comes from ordinary conditional independences in a graph that is not mb-shielded; in the one-edge variant of Remark~\ref{rem:verma-variant}, the additional Verma restriction alone reduces the efficiency bound at the law $P^{+}$ by approximately $22.45\%$, a reduction of variance rather than of interval width (width factor approximately $1.1356$).
The proposition does not say that every sequential influence function is inefficient, that no efficient estimator can be robust on a special submodel, or that efficiency and robustness can never coexist: the comparison is between two specific constructions, and no impossibility theorem is asserted. When the observed-data model is nonparametric the influence function is unique, so an estimator that is regular with respect to the full nonparametric observed-data model and asymptotically linear at the all-correct law necessarily has the canonical gradient, and no projection gap can arise; regularity within one nuisance regime does not suffice for this conclusion. On the graph of Section~\ref{sec:counterexample} the operative model $\mathcal N_+(\mathcal G)$ is a strict submodel of the observed-data law space, influence functions are correspondingly nonunique, and Theorem~\ref{thm:sequential-projection-comparison} gives only the inequality $P_0\phi_{\mathrm{seq}}^2\geq P_0\phi_{\mathrm{eff}}^2$ together with its equality condition; strictness is established at the law of Proposition~\ref{prop:sequential-strict-variance-gap} and is not asserted for every graph in the scope of this paper. The two-transition calculation in Section~\ref{app:B-sequential} of the Supplementary Material separately checks all four union regimes and a second strict projection gap at a law generated through a latent SEM rather than from recursive-head coordinates.
\begin{table}[tbp]
\centering
\caption{Proved comparison between the order-free canonical construction and the restricted sequential alternative.}\label{tab:efficient-sequential-comparison}
\small
\setlength{\tabcolsep}{4pt}
\begin{tabular}{>{\raggedright\arraybackslash}p{0.20\linewidth}
>{\raggedright\arraybackslash}p{0.36\linewidth}
>{\raggedright\arraybackslash}p{0.36\linewidth}}
\toprule
Property & Canonical one-step/TMLE & Sequential union estimator \\
\midrule
Graph and intervention scope
& Every ADMG; fixed interventions of the classes (I1)--(I3) and admissible independent-draw mixtures
& Source-isolated system-wide fixed-node subclass with the acyclic literal
block order, source precedence, and source-tail coverage in Assumption
\ref{ass:mr-sequential-transport} \\
Population drift
& Exact chart-integral remainder for one-step and for TMLE on its root
event; source-only cancellation for the coherent one-step
(Proposition \ref{prop:source-block-drift-cancellation});
the flow can leave the slice (Example \ref{ex:tmle-source-slice-failure})
& Transitionwise product in \eqref{eq:exact-sequential-drift} \\
Robustness guarantee
& Local efficiency; no general union model
& Consistency on up to $2^m$ transitionwise correctness regimes \\
All-correct influence function
& Canonical gradient $\phi_{\mathrm{eff}}$
& Ambient representer $\phi_{\mathrm{seq}}$ \\
Variance at the intersection
& Nested-model efficiency bound
& Efficiency bound plus the squared projection residual in
\eqref{eq:sequential-pythagoras} \\
Substitution
& Exact-flow TMLE only
& Not automatic for the additive sequential estimator \\
\bottomrule
\end{tabular}
\end{table}
\begin{remark}[Reading the count $2^m$]\label{rem:robustness-count-comparison}
Counts reported elsewhere index different objects and do not order robustness. For a generalized mediation functional with $K_{\mathrm{med}}$ causally ordered mediators, \cite{zhou2022semiparametric} obtains estimators that are $K_{\mathrm{med}}+2$ robust in the sense that consistency follows if any one of $K_{\mathrm{med}}+2$ \emph{specified sets} of nuisance functions is correctly specified and consistently estimated, and cautions (Section~1 there, writing $K$ for its mediator count) that this convention ``does not imply that a `$K+2$-robust' estimator is necessarily more robust than, for example, a `$K+1$-robust' estimator'', since the estimators may correspond to different estimands that require modeling different parts of the likelihood. That result, like Theorem~\ref{thm:sequential-multiple-robustness}, concerns a single component functional rather than a contrast. We make no comparison of robustness strength between the two constructions and no impossibility claim for either; the count that the present representation reports for a mediation graph is recorded in Section~\ref{app:B-sequential} of the Supplementary Material.
\end{remark}
\begin{remark}[A one-edge Verma variant]\label{rem:verma-variant}
The nested model of the graph in Proposition~\ref{prop:sequential-strict-variance-gap} coincides with its ordinary Markov model (Section~\ref{sec:counterexample} and Appendix~\ref{app:prototype-ledger} of the Supplementary Material), so the gap in Proposition~\ref{prop:sequential-strict-variance-gap} is produced by ordinary conditional independences. Adding the arrow $W\to B$ gives a graph $\mathcal G^{+}$ with the same districts and the same root singleton source, in which the reachable $\{B,Y\}$ kernel is no longer constrained to be invariant in $w$, while its $Y$-margin, obtained by summing over $B$, is constrained to be invariant: $\sum_b p(b\mid w)p(y\mid b,z,w)$ does not depend on $w$. This is a Verma restriction; the ordinary Markov model of $\mathcal G^{+}$ is $\{W\mathrel{\perp\!\!\!\perp} A\}$, and the chart dimensions are $28<30<31$. At the strictly positive law $P^{+}$ of \eqref{eq:verma-variant-rational-sem} in Appendix~\ref{app:prototype-ledger}, $\operatorname{do}(W=0)$ satisfies Assumption~\ref{ass:mr-sequential-transport} with $m=1$, $\psi_0=0.542$ and $P^{+}\phi_{\mathrm{seq}}^2=0.496472$. At this law, imposing the ordinary Markov restriction leaves the efficiency bound unchanged, whereas the additional Verma restriction reduces it by approximately $22.45\%$, from $0.496472$ to approximately $0.38500$ (exact values in \eqref{eq:verma-variant-values}); the corresponding interval widths differ by the factor approximately $1.1356$.
\end{remark}
\section{Finite-sample illustration of the variance comparison}\label{sec:finite-sample}
We illustrate the strict comparison in Proposition~\ref{prop:sequential-strict-variance-gap} at its rational law. The target is $\psi_0=\mathrm E_{P_0}\{Y\mid\operatorname{do}(W=0)\}=0.54725$; the samples are i.i.d.\ draws of size $n=1000$, $2500$, and $5000$, with $500$ Monte Carlo replicates at each size and fixed seed $20260722$, and every estimator is evaluated on every replicate. Three estimators are compared: the same-sample one-step estimator $F(\widetilde\theta)+\mathbb P_n\phi_{\widetilde\theta}$ of Corollary~\ref{cor:same-sample-one-step}, the practical iterated chart-targeted estimator of Remark~\ref{rem:straight-line-numerical-targeting}, and the sequential estimator in its same-sample pooled-ratio form, whose influence function at $P_0$ is the all-correct sequential influence function $\phi_{\mathrm{seq}}$ of Proposition~\ref{prop:sequential-strict-variance-gap}; Section~\ref{app:B-finite-sample} of the Supplementary Material records the pooled-ratio identity and its distinction from the cross-fitted estimator \eqref{eq:sequential-cross-fitted-estimator}, the standard errors, the initialization, targeting and fallback settings, and the complete results.
The influence-function variance constants are exact rationals, established in the proof of Proposition~\ref{prop:sequential-strict-variance-gap} in Section~\ref{app:proofs-sequential} of the Supplementary Material. To five decimals, and so that $v/n$ is the leading asymptotic variance of the corresponding estimator, $v_{\mathrm{eff}}\approx0.28803$ and $v_{\mathrm{seq}}\approx0.49553$, with $(v_{\mathrm{seq}}/v_{\mathrm{eff}})^{1/2}=1.311647\ldots$.
Thus the exact limit theory predicts equal-level asymptotic Wald intervals about $31\%$ wider for the sequential estimator.
The observed sequential interval widths (Table~\ref{tab:finite-sample-variance-comparison} of the Supplementary Material) exceed those of the one-step and targeted estimators by approximately $31.2\%$ to $31.4\%$ across the three sample sizes, computed from the unrounded replicate summaries recorded in the supplement, in agreement with the exact ratio above; the coverage values are descriptive rather than acceptance criteria.
This experiment illustrates the efficiency gap under correct specification; it does not test nuisance misspecification or establish the sequential union model, whose two-transition bookkeeping and four correctness regimes are checked exactly in Section~\ref{app:B-sequential} of the Supplementary Material. Code, replicate-level results, and all diagnostics are included in the computational supplement.
\needspace{6\baselineskip}
\section{Discussion}\label{sec:discussion}
\subsection*{What the theory establishes}
Three notions concerning a direction $f\in L_2(P)$ are kept apart throughout: observational centering, membership in the tangent space $\mathcal T_P\mathcal N(\mathcal G)$, and orthogonal projection onto it. The first does not imply the second: the observationally centered direction of Example~\ref{prop:raw-failure}, in a graph that is not mb-shielded, is not a model score, and its projection residual is a diagnostic, not an efficiency correction. Three conclusions follow. The forward recursive-head chart settles which directions are model scores (Theorem~\ref{thm:forward-chart}), and because native-district score spaces are orthogonal at every positive law (Theorem~\ref{thm:district-orthogonality}) the efficient projection is computed district by district, with no district order and no acyclicity condition on the district quotient. For the interventions of the classes (I1)--(I3) the canonical gradient is a district-by-district Riesz solution (Theorems~\ref{thm:normalized-intervention-bridge} and~\ref{thm:intervention-chart-eif}), extended to independent-draw mixtures by a product rule; the coherent one-step and chart-flow targeted estimators are first-order efficient under the local conditions of Theorems~\ref{thm:coherent-one-step} and~\ref{thm:chart-tmle}, and neither carries a union-model guarantee for the full class (I1)--(I3) (Table~\ref{tab:one-step-tmle-contrast}). The double robustness of \cite{bhattacharya2022semiparametric} comes from the primal--dual structure under primal fixability, not from mb-shieldedness. The restricted sequential construction is transitionwise multiply robust only under Assumption~\ref{ass:mr-sequential-transport} (Theorem~\ref{thm:sequential-multiple-robustness})---intervention sources that are singleton observational districts, a literal block order in which each relevant source precedes its district, and source-tail coverage, whose separate necessity is not asserted---and its all-correct influence function has variance at least the efficiency bound, strictly larger at the law of Proposition~\ref{prop:sequential-strict-variance-gap} (Theorem~\ref{thm:sequential-projection-comparison}). No general trade-off between efficiency and robustness is asserted (Remark~\ref{rem:robustness-count-comparison}).
\subsection*{Limits and extensions}
The finite-state restriction is what makes the chart a finite-dimensional smooth manifold, the paths explicit, and the projection a matrix computation; it is also what makes the full-chart likelihood benchmark available. Boundary laws, vanishing fixing divisors, and nonregular targets lie outside the interior theory. Three questions are left open. First, extending this route beyond fixed finite support would need a smooth parameterization with an explicit tangent space, a computable projection onto it, and a replacement for strict positivity; whether other routes need the same ingredients is not asserted, and the benchmark of Section~\ref{sec:targeted}, under which a regular interior maximum likelihood plug-in is efficient without targeting, would no longer be available. Second, the efficient projection separates across native districts, but within a district the full Gram matrix across its intrinsic blocks is in general required (Section~\ref{sec:geometry}), so scalable computation of the projection at larger finite state spaces is open; the chart dimension alone is not an established complexity bound. Third, the order-free estimators lack a constructive nuisance decomposition supporting robustness: one would need a factorized drift compatible with recursive-head coordinates, which Section~\ref{sec:targeted} does not characterize and which Section~\ref{sec:sequential-union} obtains only under Assumption~\ref{ass:mr-sequential-transport}.
The intervention classes are a matter of identification scope rather than of support. Natural-value source retention, partially mixed edge boundaries (Example~\ref{ex:split-boundary-nonidentification}), and general observed-ADMG path-to-edge reductions are outside the theorems already on fixed finite support, and are not unfinished cases of them.