Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
56,505 characters · 23 sections · 20 citation commands
Interventional Score Geometry for Causal Inference: A Dictionary for Structural Identification
Structural causal analysis distinguishes between the joint distribution of observables and the causal mechanism that generates it. Distinct structural models can generate the same reduced-form distribution while disagreeing about causal direction; this classical non-identification result motivates instrumental variables, exclusion restrictions, and randomized designs haavelmo1943,pearl2009,imbens2015.
Information geometry equips statistical models with a Riemannian structure derived from the Fisher information metric amari2016, and recent work studies causal and predictive relationships through the geometry of manifolds built from observational data surasinghe2020. These are, by construction, functions of $p(x)$ alone, and therefore face the same identification limit as $p(x)$ itself.
This paper's organizing idea is that causal asymmetry has to come from outside $p(x)$, the joint density of the observed variables: from an explicit intervention. A distribution, however carefully you stare at it, tells you what tends to go with what; it does not tell you what would happen if you reached in and moved something. That has to be supplied, and once it is, the interesting question is how to write it down without smuggling in more than was actually given. Section (ref) builds the resulting interventional score correctly, on the coordinates left non-degenerate by a hard intervention, and works through a fully explicit example. Section (ref) shows why an admissible set of intervention directions cannot, by itself, convert an observational score into a causal one, and introduces the interventional response field that must replace it. Section (ref) builds a causal metric as ordinary Fisher information on a family of interventions sharing a common target, so that the underlying densities live on a common space. Section (ref) restates randomization, instrumental variables, and conditional-independence designs in this vocabulary, stating precisely what follows from each assumption and what does not.
It helps to have a picture in mind before the definitions start. The intuition I are borrowing, loosely, is the one general relativity uses for gravity: gravity is not an extra force bolted onto an otherwise fixed background, but something read off from how the geometry of spacetime governs motion. What is locally felt as gravitational pull reflects a larger structure, not a property sitting inside one object. I want to treat causal influence the same way: not as a label attached to a single observational distribution, but as something read off from how a whole family of distributions moves as an intervention parameter changes. The relevant object is therefore never $(M,g,\psi)$ sitting still; it is the indexed family $\xi\mapsto p_{-k}^{(k,\xi)}$, and how that family deforms as $\xi$ varies.
The analogy is worth exactly as much as it is limited, and it is worth being blunt about the limit. I are not claiming that causal influence literally is curvature, and Proposition (ref) says something closer to the opposite of that claim: causality is provably not hiding anywhere inside the static observational geometry, however carefully that geometry is built. What the analogy is meant to convey is only a shift in where to look: from a single shape to a family of shapes and their deformation. The rest of the paper is the attempt to make that shift precise rather than leave it as a metaphor.
The admissible intervention set at $x$ (the directions in which the system can, in principle, be perturbed by a given design or policy; Section (ref) makes this precise), and the interventional response field built on it, must be supplied from outside the model: by design, institutional knowledge, or a structural assumption. Nothing in this framework manufactures causal content from $p(x)$ and a choice of manipulable coordinates. I do not claim new asymptotic results for meta-learners or double machine learning; Section (ref) restates known results and is explicit that "score" is used in two unrelated technical senses across the two literatures.
Pearl's Ladder of Causation separates causal reasoning into three levels: association, intervention, and counterfactual reasoning, and the objects of this paper sort cleanly onto the first two. The observational geometry $(M,g,\psi)$ belongs to the associational rung: it is determined entirely by $p(x)$ and answers questions about what tends to occur together, not about what would happen under manipulation. The interventional score $\psi_{-k}^{(k,\xi)}$, the sensitivities $S_{k\to j}$, and the response fields $S(v)$ belong to the interventional rung, since all three are indexed by an explicit $\operatorname{do}(X_k=\xi)$ and cannot generally be recovered from $p(x)$ alone. Proposition (ref) can be read as a geometric statement of exactly the gap between these two rungs: no transformation of the observational distribution, however sophisticated, substitutes for the intervention operator.
The framework does not reach the third rung, and I do not try to stretch it there. Counterfactual reasoning requires coupling factual and hypothetical outcomes at the level of a single unit, using structure such as shared exogenous disturbances in a structural causal model, and that goes beyond a family of interventional marginals indexed only by $\xi$. A geometry of that third rung is left for future work.
Information geometry and causal manifolds. amari2016 develops the Fisher information geometry of parametric families. surasinghe2020 construct manifolds from time-delayed observations and relate geometric distances to transfer entropy and Granger causality; this is an observational and predictive program in spirit. dominguez2023 show that structural causal models with $k$ parents induce data on $k$-dimensional submanifolds, characterizing the support of an SCM. We differ by studying how that support and its score deform under an explicit intervention operator.
Score-based generative models. song2021 estimate $\nabla_x\log p(x)$ directly from data via score matching and diffusion. This supplies estimates of the observational score $\psi$, not of the interventional response field of Section (ref); Section (ref) is explicit about that gap.
Instrumental variables, meta-learners, and double machine learning. angrist1996 give the canonical LATE identification result under monotonicity; heckman2005 develop the continuous-instrument marginal treatment effect framework, invoked in Remark (ref). kunzel2019 introduce the X-learner; chernozhukov2018 develop double machine learning via Neyman-orthogonal moments; nie2021 develop the R-learner. Section (ref) is explicit that the "score" in Neyman-orthogonal estimation is not the density score used elsewhere in this paper.
Let $X=(X_1,\dots,X_d)$ have density $p$ on $\mathcal X\subseteq\mathbb R^d$, smooth enough for the derivatives below to exist. let $M\subseteq\mathbb R^d$ denote the smooth support of $p$, equipped with a background Riemannian metric $g$ chosen independently of $p$; typically this is just the ambient Euclidean metric in the given coordinates. We are deliberately modest about what $M$ is: it is the support of $p$ carrying a metric we supply, not a Riemannian structure that $p$ itself induces. Define the score field \[ \psi(x)=\nabla_x\log p(x), \qquad x\in M,\ p(x)>0. \] Intuitively, $\psi(x)$ points in the direction a particle sitting at $x$ would need to move to climb the log-density landscape fastest: where $p$ is large and flat, $\psi$ is small, and where $p$ rises or falls steeply, $\psi$ is large. The triple $(M,g,\psi)$, the observational geometry, is a reparameterization of $p$: support, plus the local direction of increasing log-density, and nothing more than that.
Most of what follows turns on getting one object right, so it is worth being unhurried about it here rather than fixing it later by patching.
For coordinate $k$ and level $\xi$, the intervention $\operatorname{do}(X_k=\xi)$ replaces the structural equation for $X_k$ with the constant $\xi$: every unit is assigned exactly the value $\xi$ on that coordinate, whatever it would otherwise have been. It helps to picture $p(x)$ as a cloud of probability mass spread over $\mathbb R^d$. Pinning one coordinate to a single value does not merely reshape that cloud; it flattens it entirely onto the $(d-1)$-dimensional slice $\{x \in \mathbb R^d : x_k = \xi\}$, much as pressing a lump of dough flat against a table leaves it spread across the tabletop rather than filling any volume above it. A distribution living on a slice like that has no density with respect to ordinary $d$-dimensional volume, since there is no thickness left in the $x_k$ direction to divide by. An expression such as $\nabla_x \log p(x \mid \operatorname{do}(X_k=\xi))$, taken naively over all $d$ coordinates, is therefore not well defined. I work instead with the density of the coordinates that remain random, which is exactly where the $d-1$ dimensions that did not get flattened away still live.
When $d=2$, $X_{-k}$ is a single coordinate and $p^{(k,\xi)}_{-k}$ is already an ordinary one-dimensional density; Section (ref) is exactly this case.
Figure (ref) draws the picture for that same two-coordinate case. The left panel shows the ordinary two-dimensional density $p(x,y)$, with the vertical line marking the slice $\{X=\xi\}$ that a hard intervention on $X$ leaves behind. The right panel shows what actually survives the intervention: an ordinary one-dimensional density in $y$ alone, living along that slice. The interventional score of Definition (ref) is the score of that one-dimensional density, not of anything defined on the original two-dimensional picture.
The diagnostic used to detect this must be built from the same marginal object, not from the joint score of all non-intervened coordinates. A nonzero derivative of $\partial_{x_j}\log p^{(k,\xi)}_{-k}(x_{-k})$, the joint score in the $x_j$ direction, can arise purely from a change in the dependence between $X_j$ and some other non-intervened coordinate $X_\ell$, with the marginal law of $X_j$ itself left unchanged. For instance, an intervention could tighten or loosen how closely $X_j$ tracks $X_\ell$ without shifting $X_j$'s own typical behavior at all, and the joint score would register exactly that tightening, wrongly making it look as though $X_k$ had reached all the way to $X_j$. A change of that kind would not witness $X_k\rightsquigarrow X_j$ in the sense of Definition (ref). I therefore define the diagnostic directly on the marginal.
Because $S_{k\to j}$ is now built from the same marginal $p_j^{(k,\xi)}$ that Definition (ref) refers to, the implication above is immediate and does not depend on the behavior of coordinates other than $X_j$. All later uses of "interventional score sensitivity" refer to this marginal $S_{k\to j}$.
Combined with the sufficiency direction already noted in Definition (ref), Proposition (ref) says that on a fixed, connected, strictly positive support, $S_{k\to j}\equiv 0$ throughout $I\times\text{supp}(X_j)$ if and only if $X_k$ has no causal influence on $X_j$ over $I$: the score diagnostic is exactly as good as the distributional definition once the support itself is not moving. The gap between the two definitions, in other words, is entirely a support phenomenon. In practice this means the diagnostic can be trusted whenever an intervention is not plausibly pushing $X_j$ into values it could never have taken before; the gap is only worth worrying about when a policy might introduce genuinely new values of $X_j$ that fall outside anything observed under the status quo.
The definitions above are abstract enough that it helps to see them worked out with actual numbers before going further. The pair of models below does exactly what Proposition (ref) says is possible: two structural stories that agree on absolutely everything observational, yet disagree on the one thing this paper cares about, namely how $Y$ responds when $X$ is manipulated.
Let $X\sim N(0,1)$ and, under Model A, \[ Y=\beta X+\varepsilon_Y, \qquad \varepsilon_Y\sim N(0,\sigma^2),\ \varepsilon_Y\perp X, \] with $\beta=0.8$, $\sigma^2=1$. The joint law is bivariate normal with covariance \[ \Sigma=
, \qquad \operatorname{Var}(Y)=\beta^2+\sigma^2=1.64. \] Because $(X,Y)$ is jointly Gaussian, the same $\Sigma$ is generated by the reverse linear model. Under Model B, \[ Y=\tilde\varepsilon_Y,\quad \tilde\varepsilon_Y\sim N(0,1.64), \qquad X=\gamma Y+\tilde\varepsilon_X,\quad \tilde\varepsilon_X\sim N(0,\tau^2),\ \tilde\varepsilon_X\perp Y, \] with $\gamma=\operatorname{Cov}(X,Y)/\operatorname{Var}(Y)=0.8/1.64\approx 0.488$ and $\tau^2=\operatorname{Var}(X)-\gamma^2\operatorname{Var}(Y)=1-0.8^2/1.64\approx 0.610$. A direct computation confirms Model B also induces $\Sigma$: a bivariate Gaussian admits a valid linear-Gaussian SCM representation in either causal direction. Models A and B are observationally equivalent, hence by Proposition (ref) induce the same $(M,g,\psi)$.
Here $d=2$, so intervening on $X$ leaves exactly one non-intervened coordinate, $Y$; the joint interventional density $p^{(X,\xi)}_{-X}(y)$ of Definition (ref) and the marginal interventional density $p_Y^{(X,\xi)}(y)$ of Definition (ref) coincide in this case, since there is nothing else to marginalize over.
Under Model A, $Y=\beta X+\varepsilon_Y$ is untouched by the intervention on $X$'s own equation, so \[ p_Y^{(X,\xi)}(y) = N(\beta\xi,\ \sigma^2), \qquad \partial_y\log p_Y^{(X,\xi)}(y) = -\frac{y-\beta\xi}{\sigma^2}, \qquad S_{X\to Y}(\xi,y) = \frac{\beta}{\sigma^2} = 0.8 \neq 0. \] Under Model B, $Y=\tilde\varepsilon_Y$ does not involve $X$, so intervening on $X$ leaves $Y$'s marginal untouched: \[ p_Y^{(X,\xi)}(y) = N(0,\ 1.64) \ \text{ for every }\xi, \qquad S_{X\to Y}(\xi,y) = 0. \]
Figure (ref) makes both halves of this concrete. Panel (a) plots the shared observational score field $\psi(x,y)$, the single vector field both models agree on since they generate the identical joint density. Panel (b) plots $E[Y\mid \operatorname{do}(X=\xi)]$ against the intervention level $\xi$ under each model: a line of slope $0.8$ under Model A, and a flat line at zero under Model B, exactly the two numbers computed above.
Models A and B share the identical observational score but disagree exactly on $S_{X\to Y}$: $0.8$ under A, $0$ under B. Model A is score-detectably consistent with $X\rightsquigarrow Y$; Model B is not. This same pair reappears in Section (ref), where it is used to show that a common admissible set for both models still yields a common projected score, even though the true responses differ.
Not all coordinates are equally manipulable. I record the feasible directions at $x\in M$ as a set \[ A(x) \subseteq T_xM, \] representing directions in which the system can, in principle, be perturbed by design or policy. I deliberately call this an admissible intervention set rather than a cone: treating $A(x)$ as closed and convex silently assumes that mixtures and positive rescalings of feasible interventions are themselves feasible. That is a reasonable idealization for continuously divisible interventions (e.g., a tax rate or a dosage that can be set to any level in a range), but it need not hold for discrete or binary interventions (e.g., "treated" versus "untreated," or a fixed menu of policy options), where $A(x)$ may be a finite set with no natural convex or scalar structure. I use the convex-cone case only where it is the natural description of continuously divisible interventions; otherwise $A(x)$ should be read as an unstructured admissible set. In other words, $A(x)$ is best thought of as a menu of moves the analyst is willing to entertain at $x$, not as a piece of geometric structure that comes for free once $M$ and $g$ are fixed.
The admissible set tells us which directions are feasible to perturb; it does not tell us how the distribution responds. Knowing that a lever can be pulled says nothing about what happens when it is pulled, and that is the whole content of this distinction. The response is a separate primitive.
For intuition, $v$ might stand for "increase the dosage." The family $p^{\,\operatorname{do},\varepsilon}_v$ is then the sequence of outcome distributions produced by dosages of increasing intensity $\varepsilon$, and $S(v)$ measures how fast that sequence moves the outcome's score right at the status quo, $\varepsilon=0$. Causal information lives in the map $v\mapsto S(v)(\cdot)$, which requires the family $\{p^{\,\operatorname{do},\varepsilon}_v\}$ as an input: structural or experimental knowledge of how the system actually responds to manipulation, not merely which directions are available to move in. It is not recoverable from $\psi$ and $A(x)$ alone (Proposition (ref)).
The response field of Definition (ref) is deliberately abstract, so it is worth seeing it worked out in a case simple enough to hold in one's head. Many interventions of practical interest, such as a subsidy that adds a fixed amount to everyone's income or a dosage shift applied uniformly across patients, act by sliding a distribution along the number line without reshaping it. Suppose an intervention indexed by \(\xi\) changes the marginal distribution of a response variable \(Y\) through exactly such a location shift:
where \(m(\xi)\) is differentiable and gives the size of the shift at intervention level $\xi$. Let
The score of the interventional distribution is then
Differentiating with respect to the intervention level gives
Equation (ref) separates the interventional response into two components with a clean reading. The term \(m'(\xi)\) measures the rate at which the intervention translates the distribution (how fast the subsidy grows, say), while \(\ell_0''\) describes the local shape of the log-density, i.e.\ how sharply peaked the untouched distribution $p_0$ is at the point being shifted through. A location intervention moves the probability landscape without changing its shape, so the score's response to it factors cleanly into "how fast are we moving it" times "how steep is the landscape we are moving it through." Nothing about dependence structure or higher moments enters at all, which is exactly why this special case is simple enough to compute by hand.
For a Gaussian location family,
we have
Therefore,
In the bivariate Gaussian example of Section (ref), \(m(\xi)=\beta \xi\), so
This is precisely the interventional score sensitivity computed by hand in Section (ref), which is reassuring: it means that earlier calculation was not a coincidence specific to that example but an instance of this general location-family formula.
A metric earns its keep only if it answers a simple question in a single number: how much does a small twist of the intervention knob actually move the distribution? Fisher information is the natural way to measure that, but it needs a well-defined family of distributions to measure it within. A causal metric built directly from cross-target derivatives runs into a further issue beyond the ones already corrected: an intervention on $X_k$ leaves a density on $X_{-k}$, while an intervention on a different coordinate $X_{k'}$ leaves a density on $X_{-k'}$, and these are not the same coordinate space. Fisher information requires a single dominated family on a common measurable space, so comparing sensitivities across different intervention targets is not simply a matter of writing a joint expectation. I therefore build the metric one target at a time.
A large $G^{(k)}(\eta)$ means that even a small change in the intervention level $\eta$ produces a density that is easy to tell apart, statistically, from the one just before it: the intervention is, in a precise sense, informative there. A small $G^{(k)}(\eta)$ means the opposite: nearby intervention levels are hard to distinguish from data, so an experiment run at that $\eta$ will need a much larger sample before anything shows up.
\paragraph{Fisher geometry of a location-intervention family.}
The location family of Equation (ref) is transparent enough to compute the fixed-target causal metric by hand, which is worth doing once before trusting Definition (ref) in messier settings. Differentiating the log-density with respect to the intervention parameter gives
The corresponding Fisher information metric is
After the change of variable
we obtain
where
is the Fisher information for the location parameter.
If the intervention is a direct translation, \(m(\xi)=\xi\), then
which is constant in \(\xi\). For a Gaussian distribution with variance \(\sigma^2\),
so the line element on the intervention-parameter space is
More generally, for any regular one-dimensional intervention family, define the arc-length coordinate
In the coordinate \(s\), the line element becomes
Thus every regular one-dimensional intervention family is locally Euclidean after arc-length reparameterization. This flatness is worth not over-reading: any smooth, strictly positive metric on a one-dimensional space becomes the standard Euclidean line element after an arc-length change of coordinates. That is a generic fact of one-dimensional Riemannian geometry, true of every regular curve whatsoever, and it is not a special discovery about interventions or about causal structure. What the calculation actually contributes is $G(\xi)$ itself, computed before that reparameterization flattens it away: the decomposition into the intervention's own "speed" $m'(\xi)$ and the ambient sharpness of the log-density, $I_{\mathrm{loc}}(p_0)$, is the substantive content, and it is a statement about the geometry of the intervention-parameter space, not about the curvature of the observational manifold or of the complete causal system.
This is the ordinary Fisher information metric of the family $\{p^{(k)}_\eta\}$, well defined under the stated regularity conditions, and it is defined separately for each target $k$, so every member of the family it describes lives on the same space $\mathbb R^{d-1}$. Comparing sensitivity across two different targets $k\neq k'$ requires either restricting attention to the coordinates common to both $X_{-k}$ and $X_{-k'}$ (i.e., excluding both $k$ and $k'$), or moving to soft or stochastic interventions that preserve the full $d$-dimensional support and hence a common space throughout; I flag the latter as the more natural fix but do not develop it here.
Let $(T,Y,Z)$ be treatment, outcome, and covariates, with $T$ and $Z$ continuously distributed with positive smooth density. Define $\kappa_{T,Z}(t,z)=\partial_t\partial_z\log p(t,z)$.
The quantity $\kappa_{T,Z}$ is not a separate object worth tracking on its own; it is simply the derivative, along the treatment coordinate, of the covariate component of the joint score field
Under random assignment, $T\perp Z$ gives $p(t,z)=p_T(t)p_Z(z)$, so $\log p(t,z)=\log p_T(t)+\log p_Z(z)$, and the score above separates block-wise into a piece depending only on $t$ and a piece depending only on $z$:
Moving along the treatment coordinate therefore leaves the covariate component of the score untouched, and moving along the covariate coordinates leaves the treatment component untouched, which is exactly $\kappa_{T,Z}\equiv 0$ again, now visible directly in the block structure of $\psi_{T,Z}$ rather than just asserted. Geometrically, the treated and control arms begin from the same covariate score field, the same pretreatment probability landscape; whatever happens to the outcome afterward cannot be traced back to a systematic difference in how $Z$ was distributed at the moment of assignment.
This block-separation is easy to over-read, so it is worth being blunt about what it does not give us. Random assignment delivers $T\perp Z$; it does not, and should not, deliver $T\perp Y$: the entire point of running the trial is to let $T$ affect $Y$, so $\partial_t\partial_y\log p(t,y,z)$ is generally, and intentionally, nonzero. Nor should this be confused with vanishing Riemannian curvature of the full manifold. What randomization buys geometrically is exactly the block-separation above, applied to the assignment mechanism only, and it says nothing about the treatment-outcome relationship the trial exists to measure; in particular, it does not flatten that relationship.
Randomization and intervention therefore play distinct geometric roles, and it is worth closing this subsection by naming the difference plainly. Randomization flattens the coupling in the assignment mechanism, $\partial_t\nabla_z\log p(t,z)=0$: it certifies that the two arms started from a level field. An actual intervention on $X_k$, by contrast, generates a trajectory of response distributions $\xi\mapsto p_j^{(k,\xi)}(x_j)$, whose local deformation is exactly $S_{k\to j}$ from Definition (ref). A randomized trial tells you the starting line was fair; the response field is what tells you how far the race actually moved.
Under $Y(t)\perp T\mid Z$ for $t\in\{0,1\}$, standard potential-outcomes arguments rosenbaum1983,imbens2015 show that $E[Y\mid T=t,Z=z]$ identifies $E[Y(t)\mid Z=z]$. I do not have a geometric restatement of this that is both non-circular and adds content beyond the standard statement, and I do not offer one.
Let $Z$ be a candidate instrument for treatment $T$ on outcome $Y$, with $v_Z$ a direction of variation in $Z$. Since $T$ and $Y$ are random variables rather than deterministic functions of $Z$, the relevant well-defined objects are the conditional mean functions $m_Y(z)=E[Y\mid Z=z]$, $m_T(z)=E[T\mid Z=z]$, which are ordinary functions on the $Z$-manifold with genuine gradients.
In a DAG-based structural causal model, Pearl's $\operatorname{do}$-operator replaces the structural equation of the intervened node with a constant and deletes its incoming edges. This is exactly the map from $p$ to $p^{(k,\xi)}_{-k}$ of Definition (ref): the intervened coordinate is fixed, not merely reweighted, and drops out of the density.
It is tempting to think that a sufficiently good generative model, trained on enough observational data, would eventually "learn" causal structure as a byproduct of learning the density well. Section (ref) rules this out directly, and it is worth restating why in this more concrete setting before moving on. Score-matching and diffusion-based generative models estimate $\nabla_x\log p(x)$ from data by fitting $s_\theta(x)$ and sampling via $dX_t = s_\theta(X_t)\,dt+\sqrt2\,dW_t$ song2021, supplying $\hat\psi(x)$, an estimate of the observational score. Given Proposition (ref), projecting $\hat\psi$ onto an admissible set does not yield a causal estimate, no matter how accurate $\hat\psi$ becomes: the family $\{p^{\,\operatorname{do},\varepsilon}_v\}$ that causal content actually lives in is simply not a fact about $p(x)$, so no amount of fitting $p(x)$ more precisely brings it into view. If interventional or quasi-experimental data are available (samples from $p^{\,\operatorname{do},\varepsilon}_v$ for $\varepsilon$ near $0$), one could in principle estimate the interventional response field $S(v)$ of Definition (ref) by score-matching each member of the interventional family and differencing. This is a data requirement beyond observational score estimation, not a byproduct of it.
The word "score" is overloaded across the literatures this paper touches, and it is worth being explicit about that rather than building a table that only papers over it. In this paper, "score" means $\nabla_x\log p(x)$, the gradient of a log-density. In semiparametric efficiency theory and in double machine learning, "score" (or "estimating equation") means a function $\psi(W;\theta,\eta)$ satisfying a moment condition $E[\psi(W;\theta_0,\eta_0)]=0$, chosen for Neyman orthogonality $\partial_\eta E[\psi(W;\theta_0,\eta)]|_{\eta=\eta_0}=0$ chernozhukov2018. In classical statistics, "score" can also mean the likelihood score $\partial_\theta \log p_\theta(x)$ of a parametric family (related to our usage, but indexed by a parameter rather than by the data coordinates themselves). These are three distinct objects, related only by an accident of terminology, and conflating them is an easy way to make a false claim sound like a derivation.
With that said: S-, T-, and X-learners kunzel2019, double machine learning chernozhukov2018, and the R-learner nie2021 are all, under their respective identifying assumptions, estimators of the same conditional average treatment effect $E[Y(1)-Y(0)\mid X=x]$, a contrast rather than a derivative in the binary-treatment case that motivates most of these methods. (For a continuously parameterized treatment, an analogous derivative object can be defined under additional smoothness assumptions, but the R-learner's usual target is the discrete contrast.) Their consistency results belong to the cited papers; I do not derive anything new about them here, and I would rather leave them uncatalogued than dress them in geometric language that adds nothing beyond what a plain restatement already says.
Observational geometry cannot identify causal direction, because it is a function of $p(x)$ alone (Proposition (ref)). Building the interventional analogue requires working on the coordinates left non-degenerate by a hard intervention (Definition (ref)), defining causal influence distributionally with a matching marginal-score diagnostic (Definitions (ref) and (ref)), and recognizing that an admissible intervention set cannot by itself convert an observational score into a causal one (Proposition (ref)); the causal content instead lives in an interventional response field that must be supplied structurally (Definition (ref)). A causal metric built from Fisher information on a fixed-target intervention family is well posed without the cross-target space mismatch of a naive construction (Definition (ref)). Applied to randomized trials, instrumental variables, and conditional-independence designs, this vocabulary reproduces exactly what each design implies and nothing more: randomization decouples treatment from covariates but says nothing about the treatment-outcome relationship it is designed to reveal, and the instrumental-variables identity is a ratio of gradients of conditional mean functions rather than of the outcome and treatment variables themselves. The framework is a common notation for relating interventions, admissible directions, and score fields across structural econometrics, experimental design, and score-based generative modeling; it does not add identification power beyond what the underlying assumptions already supply. One loose end is now tied off: the score diagnostic is not merely sufficient but also necessary for causal influence once the support of the affected variable holds still (Proposition (ref)), so the diagnostic and the definition can only part ways when a support itself moves, shrinks, or splits apart under intervention. A metric that compares sensitivity across different intervention targets, rather than one target at a time, remains open, and would be the natural next thing to build. In the language of Pearl's Ladder of Causation, the paper geometrically separates association from intervention while leaving the counterfactual rung, which requires cross-world, unit-level structural information, as a natural extension.