EconBase
← Back to paper

Position: Prioritize Identifying Structure, Not Complex Models, for Scientific Discovery

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

53,858 characters · 11 sections · 38 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Position: Prioritize Identifying Structure, Not Complex Models, for Scientific Discovery

\twocolumn[ \icmltitle{Position: Prioritize Identifying Structure, Not Complex Models, for Scientific Discovery}

icmlauthorlist\icmlauthor{Tyler H. McCormick}{uw}

\icmlaffiliation{uw}{Departments of Statistics and Sociology, University of Washington, Seattle, WA, USA} \icmlcorrespondingauthor{T. H. McCormick}{[email removed]}

\icmlkeywords{Identifiability, Machine Learning, Miasma, Scientific Discovery}

\vskip 0.3in ]

\printAffiliationsAndNotice

abstractModern Machine Learning (ML) and Artificial Intelligence (AI) models, especially large language models (LLMs), are increasingly used to generate scientific hypotheses and mechanistic explanations from observational data. This position paper argues that in the high-dimensional proxy regimes where modern ML excels, mechanistic learning is generically underdetermined: many incompatible mechanisms induce essentially the same observational relationships on the support of the data, so predictive success and coherent explanations are insufficient evidence of mechanism discovery. This underdetermination becomes uniquely hazardous with large language models (LLMs), which tend to collapse large equivalence classes of explanations into a single fluent narrative. This paper proposes concrete standards for “mechanistic ML,” and argues these norms are necessary if LLM-centered workflows are to support science rather than merely simulate it.

Introduction

Position: Absent explicit identifying structure, high-dimensional “scientific discovery” from observed features is generically underdetermined; predictive success and fluent explanations are insufficient evidence of mechanism discovery. Therefore, research should prioritize articulating and evaluating identifying assumptions and discriminating regimes, rather than building ever more complex models. There is a substantial and growing body of work that frames complex AI and ML models as tools for scientific discovery, including symbolic regression Schmidt2009,Udrescu2020, sparse identification of dynamics and PDEs Brunton2016,Rudy2017, physics-informed inverse problems Raissi2019, and methods that extract interpretable hypotheses from high-dimensional text via sparse autoencoders and LLMs pmlr-v267-movva25a. In the social sciences, the Fragile Families Challenge tested the value of prediction systems for social outcomes Salganik2020, and “agnostic” approaches argue for more inductive, sequential discovery workflows Grimmer2021. There is great promise in incorporating AI/ML tools in scientific discovery.

This position paper, however, contends that using modern AI/ML tools agnostically to uncover mechanisms is inherently fraught. Put simply, there are many representations of the relationship between observed features and outcomes that are consistent with any observed dataset. A representation can yield strong in-domain prediction and even convincing simulations, yet it provides no guarantee that the learned representation corresponds to the underlying, intervention-stable driver of the system. The remedy is to make identifying structure explicit. By identifying structure, the paper means joint specification of (a) restrictions on the latent mechanism (what the system is allowed to do under interventions) and (b) restrictions on the observation process (how latent drivers generate measured proxies), together with (c) a data-collection regime that supplies discriminating variation (new environments, interventions, or measurement channels). Identifying structure and complex models are not opposites. In many successful scientific-ML systems, complex models are exactly what make it possible to exploit identifying structure once it exists. This point echoes Cartwright's nomological machine: law-like regularities arise when components with stable capacities are arranged in a sufficiently stable environment so that repeated operation yields repeatable behavior cartwright1999nomological.

The argument starts from a key observation: scientific workflows rarely observe the drivers of scientific processes directly. Instead, they observe proxies that are associated with one or more latent drivers. In the apocryphal story, Newton didn't see gravity; he felt a bump on his head from the falling apple. mendel1866 observed phenotypes, traits such as seed shape and seed color, recorded for parent plants and their offspring. These phenotypes are proxies that, when observed across a set of controlled crosses, gave clues about the underlying scientific mechanism, the transmission of discrete hereditary “factors” across generations. Crucially, Mendel's success did not come from black-box prediction on high-dimensional proxies, but from identifying structure: favorable biological architecture, carefully chosen discrete traits, controlled matings, replication, and a narrow hypothesis class (segregation and independent assortment) against which the proxy patterns could be tested and falsified mendel1866,fisher1936,curtis2023mendel. For the purpose of this paper, a mechanistic query is a concrete question about an underlying data-generating process, such as what would happen under an intervention, whether a relationship is stable across settings, which latent driver explains an observed pattern, or what sign a causal effect has. Such a query \(q(M)\) is identified when every mechanism compatible with the observed proxy evidence and stated assumptions gives the same answer, and partially identified when the compatible mechanisms give a bounded or otherwise structured set of possible answers.

The fundamental challenge is that mechanistic inquiry needs to understand both how good the proxy is (i.e., how strong the relationship is between the proxy and driving mechanism) and what it is a proxy for (i.e., the underlying scientific mechanism). The position here is that, in high-dimensional observational regimes, both cannot be identified without additional structure. The shift from bureaucratic statistical regimes toward brokerage regimes of passively collected digital traces makes this proxy problem increasingly central in social and administrative data fourcade2026census. Unstructured discovery, at best, identifies an equivalence class of mechanisms based on plausible relationships between the proxies and the mechanisms. Formally, the proxy observational law \(P_X\) is the population distribution of the observed proxies under the observational regime at hand. For a fixed proxy observational law \(P_X\), let \(\mathcal M(P_X)\) denote the set of mechanisms and measurement channels that induce \(P_X\). Without identifying structure, \(\mathcal M(P_X)\) is typically not a singleton. ogburn2021warning gives examples in the context of causal graphs.

Improving model performance or even the existence of a plausible, consistent explanatory channel is also not sufficient. Coherence is not, by itself, scientific evidence. What's more, fixating on the wrong mechanism can have dramatic, longstanding consequences. Take, for example, the miasma, or “bad air” theory of disease. The proxies are environmental correlates of illness, such as odor, proximity to swamps/sewers, stagnant water, signs of decaying material, or meteorological conditions. Many of these proxies are strongly correlated with outbreaks, providing seemingly ironclad evidence for the theory. The strength of these associations, along with the tendency for path dependence in science kuhn1997structure,soler2025would, meant that the miasma theory remained pervasive even in the presence of mounting evidence for germ-based disease transmission. Mendel, on the other hand, succeeded because favorable biology, selected measurements, and discriminating design effectively made \(\mathcal M(P_X)\) small for the mechanistic queries he asked. Throughout, the term design is used to refer to the structure of the evidence-generating regime, including which variables are measured, which labels or environments are recorded, which naturally occurring contrasts are available, and which interventions or controlled protocols, if any, produced the observations.

A richer feature space might seem capable of identifying underlying mechanisms from proxies, but high-dimensional proxies often concentrate data on a thin effective support shaped by institutions, selection, and measurement. As the proxy dimension increases, the number of possible patterns that, spuriously or otherwise, give a representation that's consistent with the data also increases. The inevitable result is that many incompatible mechanisms can agree closely on what they support yet diverge sharply under interventions or modest domain shifts. Formally, Appendix (ref) shows that if $P_X$ is supported on a lower-dimensional set, then off-support behavior is not identified for any rich function class without additional constraints. LLMs add a distinctive hazard. Recent work formalizes an “inevitability” result for open-world querying: for any fixed computable learner, there exist computable target functions on which it must systematically err on infinitely many inputs Xu2025. This paper argues that an analogous inevitability holds for explanations: for any observational evidence based on proxies, there exist multiple compatible but conflicting mechanistic stories. What is new about LLM workflows is not underidentification itself, as identification limits are classical, but that LLMs industrialize the collapse of underidentification into a single fluent narrative at scale.

This contention does not depend on a narrow definition of mechanism. The term means stable structural features that explain how a system responds to interventions: derivative constraints (which variables matter and with what sign), conservation laws, invariances, and causal pathways supporting policy counterfactuals Machamer2000,Woodward2003,Heckman2000. Economists distinguish reduced-form associations, which may be accurate yet mechanistically agnostic Angrist2010,AtheyImbens2019, from structural models encoding policy-relevant counterfactuals Heckman2000.

The call to action is simple: for ML systems to contribute to scientific discovery, “mechanistic learning” must be treated as an identification problem, with research prioritizing the discovery and declaration of the identifying structure that makes mechanistic questions answerable from proxy data. Accordingly, any mechanistic claim must be paired with at least one of: (i) a clear statement of this identifying structure (and what remains unidentified), (ii) mechanism-discriminating evaluations that can shrink the proxy-compatible set (interventions, cross-environment invariance tests, derivative/shape constraints), or (iii) multiplicity reporting that characterizes the surviving equivalence class, including explicit falsifiers and sensitivity to the stated assumptions.

The paper proceeds as follows. Section (ref) sets up mechanisms, proxies, and models; Section (ref) shows that absent identifying structure, proxies identify mechanism equivalence classes rather than unique mechanisms; Section (ref) explains how high dimensionality and thin support aggravate those classes; Section (ref) introduces LLM narrative collapse; Section (ref) relates the ambiguity to aleatoric, epistemic, and structural identification uncertainty; Section (ref) gives the empirical example; Section (ref) states the call to action; and Section (ref) provides alternative views.

Problem Setup and Terminology

Let \(S\) denote latent scientific drivers (“mechanism-level” variables). The observed data consist of proxies \(X\) generated from \(S\) by a measurement process. Consider deterministic measurement \(X=f(S)\) and stochastic measurement channels \(X\mid S\sim K(\cdot\mid S)\), where \(K\) is a Markov kernel from the latent space to the proxy space, covering additive noise, discretization, aggregation, censoring, and other coarsenings fuller1987,carroll2006. Given a latent law \(P_S\) and measurement channel \(K\), the induced proxy observational law is \(P_X(A)=\int K(A\mid s)\,P_S(ds)\) for measurable proxy events \(A\). The proxy gap is the distinction between this observable law and the latent mechanism \(M=(P_S,K)\), or a mechanistic query \(q(M)\), when distinct mechanisms induce the same \(P_X\). A modern workflow trains a predictor \(h\), often through a learned representation \(Z=\phi(X)\), to optimize an in-domain objective, such as predicting an outcome \(Y\) from proxies \(X\), which is distinct from recovering the latent \(S\).

To move from proxies to mechanisms, many methods pool “similar” observations, either explicitly (clustering, trees, mixtures, nearest neighbors) or implicitly (representation learning that induces neighborhood smoothing). The workhorse assumption behind pooling is a symmetry choice often framed as exchangeability: declaring certain differences irrelevant so that averaging is meaningful. Exchangeability yields powerful representation theorems definetti1937,aldous1981,kallenberg2005,diaconis1980. But representation is not identification: the theorems specify what must be true if a symmetry holds, but do not guarantee the symmetry is correct, nor that a unique mechanism is pinned down by proxies. Kleinberg’s clustering impossibility theorem makes this concrete kleinberg2002. This paper makes three claims. Proofs of the three forthcoming Propositions are in Appendix (ref). For clarity, the objects relate as: \[ S \xrightarrow[\text{measurement}]{K \text{ or } f} X \xrightarrow[\text{encoding}]{\phi} Z \xrightarrow[\text{predictor}]{h} \widehat Y, \] Many predictors or encodings can map the same proxies to similarly good predictions, but that happens to the right of \(X\). The proxy gap lives to the left of \(X\): many latent mechanisms and measurement channels can generate the same proxies while implying different mechanistic answers. Identification concerns what can be inferred about \(S\) or \(q(M)\) from \(X\), not merely how well \(h\) predicts \(Y\) in-domain.

Claim I: The Proxy Gap Creates Mechanism Equivalence Classes

Proxies constrain mechanisms only through a many-to-one measurement process (deterministic or stochastic). Without additional identifying structure, the evidence generally identifies an equivalence class of mechanisms rather than a unique mechanism. Causal-quartet examples in mcgowan2024causal show this in the context of causal inference. For 19th-century epidemic disease, the day-to-day evidence available to observers---symptoms, odors, local air conditions, seasonality, crowding, sanitation, neighborhood mortality rates, and water/waste correlates---was often compatible with both miasmatic and germ-theoretic accounts. Miasmatic theories were empirically plausible because foul environments were genuinely predictive of disease, and sanitation often improved health; germ-theoretic accounts became compelling only when new evidence decoupled these proxies from specific microbial exposure and transmission mechanisms. With only the original proxy channel and no new measurement or identification restriction, observing more cases need not resolve the mechanistic ambiguity.

The argument begins with an intentionally “easy” setting. Assume the latent state decomposes into \(B\) blocks \(S=(S_1,\dots,S_B)\). A proxy block \(X_b\) is the collection of observed measurements intended to capture (possibly noisily) the corresponding latent block \(S_b\). Assume a modular measurement structure in which each proxy block depends only on its corresponding latent block (e.g., multiple sensors targeting the same driver), either deterministically \(X_b=f_b(S_b)\) or via a channel \(X_b\mid S_b\sim K_b(\cdot\mid S_b)\). The propositions below show that even if this modular structure holds and measurement is invertible within blocks, proxies alone still do not uniquely identify the latent mechanism without additional identifying assumptions or discriminating regimes.

proposition[Proxy identifiability] Let $X=(X_1,\dots,X_B)$ be observed proxy blocks. Suppose an admissible deterministic representation consists of latent blocks $S=(S_1,\dots,S_B)$ and bimeasurable bijections $f_b:\mathcal S_b\to\mathcal X_b^{\mathrm{supp}}$ such that $X_b=f_b(S_b)$ a.s. for each $b=1,\dots,B$. If a second admissible representation of the same proxy blocks is given by $X_b=\tilde f_{\pi(b)}(\tilde S_{\pi(b)})$ a.s. for each $b$, where $\pi$ is a permutation of $\{1,\dots,B\}$ and each $\tilde f_{\pi(b)}$ is also a bimeasurable bijection onto $\mathcal X_b^{\mathrm{supp}}$, then there exist bimeasurable bijections satisfying, for each $b$, \[ \begin{gathered} g_b:=\tilde f_{\pi(b)}^{-1}\circ f_b,\qquad \tilde S_{\pi(b)}=g_b(S_b)\ \text{a.s.},\\[-0.25ex] \tilde f_{\pi(b)}=f_b\circ g_b^{-1}. \end{gathered} \]

In this deterministic invertible setting, the equivalence class induced by $P_X$ is exactly within-block reparameterizations plus, when labels are not intrinsic, block permutations. Thus only quantities invariant under those transformations (e.g., certain conditional independences or invariance relations) are identified from $P_X$ without further structure. This population-level ambiguity is not a finite-sample issue and does not disappear as parameter uncertainty shrinks; Proposition (ref) shows it also survives arbitrary conditionally independent measurement noise.

samepage\begin{proposition}[Non-identification persists under noise] Let $S=(S_1,\dots,S_B)$ and let $K(x\mid s)=\prod_{b=1}^B K_b(x_b\mid s_b)$ be a conditionally independent proxy channel. For any blockwise bimeasurable maps $g_b$, define \[ \tilde S_b=g_b(S_b),\qquad \tilde K_b(\cdot\mid \tilde s_b):=K_b(\cdot\mid g_b^{-1}(\tilde s_b)). \] Then $(S,K)$ and $(\tilde S,\tilde K)$ induce the same distribution of $X$. Thus, blockwise reparameterization non-identification survives arbitrary conditionally independent measurement noise. \end{proposition}

In short: richer proxies can improve estimation of a chosen model, but they do not, by themselves, guarantee identification of a mechanism.

Claim II: High Dimensionality Aggravates the Equivalence-Class Problem

The proxy gap already creates mechanism equivalence classes. High dimensionality and modern ML practice aggravate this problem by making disagreement easier to hide on thin support and harder to diagnose in realistic workflows. High-dimensional proxy vectors rarely fill their ambient space; they concentrate on a thin effective support shaped by selection, measurement, and correlation structure. At the same time, the curse of dimensionality implies that learning general functions (not to mention mechanism-relevant counterfactual behavior) requires sample sizes that explode with dimension unless strong structure is assumed Bellman1957DynamicProgramming,Stone1982OptimalRates.

Two geometric consequences are especially relevant. First, in high dimensions, distances can concentrate: “nearest” and “farthest” neighbors can become nearly indistinguishable under broad conditions beyer1999,aggarwal2001. When neighborhoods lose meaning, any method that depends on local pooling (nearest neighbors, kernels, clustering, or representation-induced neighborhood smoothing) becomes sensitive to modeling choices that are not themselves identified by the data. Second, even coarse partitioning of a $d$-dimensional proxy space yields exponentially many regions as $d$ grows. Unless sample sizes grow commensurately, most regions contain too little data to discriminate mechanisms that agree on the observed support but differ elsewhere. Operationally, this is how a large mechanism equivalence class shows up in finite samples: many incompatible mechanisms are observationally indistinguishable on the thin support where data live, yet can diverge sharply off-support.

In sum, Sections (ref)--(ref) argue that in proxy-rich, high-dimensional regimes, mechanistic claims should be treated as claims about a set of observationally compatible mechanisms unless accompanied by identifying structure or mechanism-discriminating evaluation.

Claim III: An Additional LLM-Specific Hazard---Narrative Collapse as False Resolution

Sections (ref)--(ref) argue that, without identifying structure, proxy evidence often supports a set of observationally compatible mechanisms rather than a unique mechanism. LLM-centered workflows create a distinctive hazard when they represent that set as a resolved explanatory account. This paper calls this narrative collapse: a many-to-one mapping from a mechanism compatibility set to an explanatory output---a single story, a ranked list of factors, or a polished synthesis---that is naturally read as more identified than the evidence warrants. Humans also do this---scientific communities routinely converge on “the” story under ambiguity, but LLMs industrialize the process by making single-story explanations fast and easy to produce at scale. Narrative collapse is not the claim that every LLM response names exactly one cause. Many systems, especially when prompted, enumerate several factors or caveats. A factor list avoids collapse only if it preserves the relevant scientific multiplicity: which mechanisms are mutually incompatible, which assumptions make each admissible, which answers to \(q(M)\) are stable across \(\mathcal M(P_X)\), and what evidence would distinguish them.

This use of “collapse” is intentionally different from two nearby concerns. Model collapse refers to degenerative dynamics that can arise when generative models are trained recursively on model-generated content, leading to distributional narrowing and loss of tails shumailov2024modelcollapse. Algorithmic monoculture refers to population-level harms when many decision-makers converge on the same algorithm, reducing system robustness even when the algorithm is locally better for each agent kleinberg2021monoculture. Both phenomena can amplify narrative collapse, but neither is required for it: narrative collapse can occur in a single interaction with a single model on a fixed dataset, whenever the evidence does not identify a unique mechanism and the interface reports a resolved explanation without representing what remains compatible.

How often this occurs in present systems is an empirical question, and the rate will vary by model, prompt, interface, task, and user population. The structural claim is conditional: whenever a workflow asks for, or rewards, a resolved mechanistic answer while \(\mathcal M(P_X)\) contains mechanisms with different values of \(q(M)\), no such answer can be uniformly warranted. Better interfaces can reduce this risk by surfacing the compatible mechanism set, tagging assumptions, and proposing discriminating tests. That is precisely the design recommendation rather than an exception to the argument. Recent critiques of uncertainty quantification for LLM agents emphasize that the standard aleatoric/epistemic dichotomy becomes strained in open, interactive settings kirchhof2025uq. The point is complementary: in proxy-driven science, the dominant ambiguity is often structural identification uncertainty, and a resolved narrative is a lossy representation of that uncertainty.

To make the issue precise without over-formalizing, fix an observed proxy distribution \(P_X\) and let \(D=(X^{(1)},\dots,X^{(n)}) \sim P_X^n\) denote the observed dataset. Let \(\mathcal M(P_X)\) denote the set of mechanisms (latent data-generating processes for \(S\) together with measurement processes \(K\)) that are observationally compatible with \(P_X\) under the setup of Sections (ref)--(ref). Let \(q:\mathcal M \to \mathcal A\) be a mechanistic query mapping a mechanism to an answer space \(\mathcal A\) (e.g., “Which variables are causally upstream?” “What intervention increases \(Y\)?” “Which invariance should hold out of domain?”). Everything below holds verbatim if \(D\) includes outcomes \(Y\) or other observables: simply replace \(P_X\) by the joint observational law of whatever variables the workflow conditions on. Proposition (ref) records the basic obstruction: if two proxy-compatible mechanisms imply different answers to \(q\), then no rule that sees only the proxies can always give the right single answer.

proposition[Narrative collapse as minimax ambiguity] If there exist \(M_1,M_2\in \mathcal M(P_X)\) such that \(q(M_1)\neq q(M_2)\), then no single-valued explanation rule \(\widehat a=\widehat a(D)\) can be uniformly correct over \(\mathcal M(P_X)\) for the query \(q\). In particular, under 0--1 loss \(\ell(\widehat a,q(M))=\mathbf 1\{\widehat a\neq q(M)\}\), the worst-case risk satisfies \[ \inf_{\widehat a}\ \sup_{M\in\mathcal M(P_X)} \ \mathbb E\big[\ell(\widehat a(D),q(M))\big] \ \ge \ \tfrac{1}{2}, \] where the expectation is over \(D\sim P_X^n\) (equivalently, over any procedure that only has access to the proxies).

Proposition (ref) is deliberately illustrative: when the evidence admits multiple incompatible answers to a mechanistic query, any single reported answer must be wrong for at least one observationally compatible mechanism. The scientifically responsible response is therefore not “be more eloquent”; it is to change what is reported. In interactive use, the risk is amplified when follow-up questions move from description to “why” and “what-if” claims. Those questions can require extrapolating beyond what the proxies identify. Conversational norms can also reward consistency: once a workflow commits to a mechanism in an early turn, later turns may rationalize and elaborate that commitment rather than reopening the mechanism set. This coherence pressure is useful for ordinary assistance, but under underidentification it can turn a set-valued scientific state of knowledge into an increasingly entrenched point narrative. Narrative collapse is therefore the mechanism-level analogue of overconfident extrapolation: it converts a set-valued scientific state of knowledge into a resolved explanatory object. Operationalizing the concept requires benchmarks that specify (i) proxy data, (ii) a mechanistic query \(q\), (iii) the set of observationally compatible answers under stated assumptions, and (iv) at least one discriminating test that would shrink that set. The unit of evaluation should not be whether a response names one cause or many factors, but whether it preserves the identified set: compatible-answer coverage, separation of mutually incompatible mechanisms, assumption tagging, discriminating-test quality, and identified-set calibration---whether the system makes point claims only when the benchmark enters an identified regime.

Where This Sits Relative to Aleatoric and Epistemic Uncertainty

Uncertainty quantification in ML is often organized around a dichotomy: aleatoric uncertainty, attributed to irreducible randomness in outcomes, and epistemic uncertainty, attributed to an agent's lack of knowledge about parameters, models, or hypotheses kendall2017uncertainties,huellermeier2021aleatoricepistemic. This split is useful for prediction problems, but it does not cleanly capture the central phenomenon in Sections 3--5. Even with unlimited data about proxies \(X\), mechanism-level claims about \(S\) can remain underdetermined unless one adds explicit identifying structure. This paper isolates a particularly important form of epistemic uncertainty: uncertainty induced by non-identification from proxies.

To make this precise (with a more formal presentation in Appendix (ref)), fix an observational distribution \(P_X\). Consider the set of mechanism stories that could have produced it: $ \mathcal{M}(P_X) \;=\; \bigl\{ (P_S,\text{measurement map/channel}) \,:\, (S \to X)\ \text{induces}\ P_X \bigr\}. $ Sections 3--4 argue that \(\mathcal{M}(P_X)\) is typically large in modern regimes: proxies are many-to-one summaries of latent drivers, and high dimensionality makes disagreement outside the observed support easier to hide and harder to diagnose. This is the structural identification uncertainty emphasized here. In the context of causal inference, gelman2024causalquartets show that even when a scalar average causal effect is held fixed, very different heterogeneous effect patterns can remain compatible with that same summary, for example. In contrast, aleatoric uncertainty concerns variability of outcomes conditional on a mechanism, while common uses of epistemic uncertainty quantify uncertainty relative to a chosen model or hypothesis class. The proxy gap matters because it questions whether that class and its measurement geometry identify the mechanistic query at all.

This perspective clarifies how these claims relate to (and differ from) the Rashomon effect and underspecification. Rashomon-style multiplicity concerns many predictors achieving similar risk fisher2019allmodels,xin2022rashomon,venkateswaran2024robustly; underspecification emphasizes that many models can fit training objectives yet behave differently under deployment damour2022. Structural identification uncertainty is more basic: from \(X\) alone, the mapping back to mechanism \(S\) is not pinned down without additional assumptions on measurement and structure. Rashomon and underspecification are important amplifiers in finite samples and deployment, but the proxy gap can generate mechanism equivalence classes even before optimization and architecture enter the picture.

These distinctions matter for LLM-centered scientific workflows. Many uncertainty quantification (UQ) tools quantify dispersion around a well-defined predictive target. But when the user asks a mechanism question—“what explains this pattern?”—the dominant uncertainty may be which members of \(\mathcal{M}(P_X)\) remain plausible. In that regime, a single calibrated interval or scalar “epistemic uncertainty” can be actively misleading: it suggests that what is missing is merely more data or better parameter estimates, when the binding constraint is lack of identification structure. This issue is sharper in interactive settings, where queries expand into counterfactuals and the system is forced to pick a story off-support; recent critiques argue that the aleatoric/epistemic split becomes strained for LLM agents in open interaction kirchhof2025uq. The point here is complementary: even perfect predictive calibration can leave mechanistic claims non-identified.

An Example with Pea Plants

Mendel crossed a pure-breeding wrinkled-green line with a pure-breeding round-yellow line. The first generation (F$_1$) was uniformly round-yellow, and those F$_1$ plants were crossed again to produce the F$_2$ generation analyzed here. The observed labels are visible traits---round-yellow (RY), round-green (RG), wrinkled-yellow (WY), and wrinkled-green (WG)---not the hereditary state itself. The hidden mechanism is the rule by which parental hereditary factors produce those visible categories. The example is intentionally favorable: Mendel worked with unusually discrete, selected, experimentally tractable traits; many genetic traits involve many loci, interactions, pleiotropy, and environmental dependence curtis2023mendel,bapty2023mendel,mackay2024pleiotropy. This section asks what happens when that structure is removed from the evidence regime.

The simulation is designed to make three points concrete. First, under a phenotype-only observational channel, distinct mechanisms can be observationally equivalent on the support of the data (Claim I). Second, adding experimental structure (here: labeling the cross type and including additional crosses) shrinks this mechanism equivalence class (Claim II). Third, modern “black-box discoverers” can achieve similar predictive performance while still yielding unstable mechanistic counterfactuals when the design metadata is removed (Claim III). The simulation compares two regimes. Regime A gives only pooled F$_2$ phenotype counts: “here are the offspring categories.” Regime B also records which controlled cross produced each observation (F2_self, testcross, monoA, or monoB) and includes additional labeled crosses. This cross label is the identifying structure: phenotype counts alone record what the offspring looked like, but the mating protocol tells us whether those counts came from an F$_2$ self-cross, a testcross, or a single-locus cross, which are different tests of the inheritance rule.

The example compares two stories that look identical if only the pooled F$_2$ counts are inspected. The Mendelian story says those counts come from a real inheritance rule: two traits, dominance, and independent assortment, which produce the familiar \(9{:}3{:}3{:}1\) pattern. The deliberately simple alternative ignores that inheritance rule and instead treats shape and color as two separate weighted coin flips, each producing the dominant trait with probability \(3/4\). This alternative is constructed to match the pooled F$_2$ counts exactly. But once the data record which cross produced each offspring, the two stories no longer agree. In Regime A (phenotype-only), the simulation generates \(n_{\text{F2}}=6000\) offspring from an F$_2$-self distribution (parents RY \(\times\) RY), and the analysis evaluates fit only at the level of the pooled phenotype distribution. In Regime B (design), the simulation generates additional labeled crosses, each with \(n_{\text{per cross}}=1200\) offspring: F2_self, testcross, \texttt{monoA}, and \texttt{monoB}. The key identifying structure is the cross-type label (e.g., testcross vs mono cross), which encodes experimental lineage/design information that is not recoverable from phenotype proxies alone but radically changes what the same proxy observations imply. The analysis fits a two-hidden-layer Multilayer Perceptron (MLP) classifier with architecture \((32,32)\) to predict offspring phenotype from parent phenotypes. In the figure labels, “no design” is Regime A: the MLP trains without cross labels and sees only F$_2$ self-cross data. “With design” is Regime B: the MLP trains with the cross label and sees the queried testcross support. A seed is one independently trained MLP run with a different random-state value controlling random initialization and training randomness; the train/test split is fixed across seeds. The “all” bars use all 60 such runs in the relevant regime; the “near-opt” bars restrict to seeds whose held-out test log loss is within \(\varepsilon=0.01\) of the best seed for that same regime.

figure[figure omitted — 790 chars of source]

The prediction task asks: for the same observed parent phenotype pair, RY\(\times\)WG, what offspring probabilities should the model assign to RG, RY, WG, and WY? Without cross labels, this question is underdetermined because the same parent phenotypes can arise from different experimental lineages; with labels, it becomes a specific testcross prediction (\texttt{cross=testcross}). The mean \(\ell_1\) spread summarizes how much the fitted models disagree: \(\frac{1}{N_{\text{seeds}}}\sum_{r=1}^{N_{\text{seeds}}} \| \hat p_r - \bar p \|_1\), where \(\bar p := \frac{1}{N_{\text{seeds}}}\sum_{r=1}^{N_{\text{seeds}}}\hat p_r\), computed either across all seeds or only near-optimal seeds. Zero would mean every fitted network gives the same four probabilities. Figure (ref) shows the point: in Regime A, equally good predictors disagree more about the counterfactual cross; in Regime B, the added cross labels and testcross support make the answer substantially more stable. The near-optimal spread falls from \(0.095\) without design to \(0.029\) with design, while the all-seed spread falls from \(0.189\) to \(0.095\). Additional concrete probabilities and plots are in Appendix (ref); a continuous low-dimensional simulation is in Appendix (ref).

Call to Action: Mechanistic ML

The core prescription is to treat identification, not expressiveness, as the binding constraint in proxy-rich scientific discovery. In this paper, that means identifying the answer to a stated mechanistic query from a compatibility set of mechanisms, not necessarily estimating a treatment effect from a randomized or quasi-random treatment assignment. First-principles restrictions are archetypal identifying structure: conservation laws, symmetries, monotonicity, geometry, and intervention semantics all shrink the compatible mechanism set and make mechanistic claims auditable. Identification is a common concept in economics and causal inference, among other areas. For example, unsupervised disentanglement is provably non-identifiable without inductive biases and/or supervision pmlr-v97-locatello19a, while identifiability can be recovered under additional observed variables, grouping structure, or restricted ICA-style assumptions pmlr-v108-khemakhem20a,pmlr-v235-morioka24a. More broadly, the logic of no-free-lunch and Kolmogorov complexity arguments makes the same meta-point: without inductive bias aligned with the world, learning cannot select the right explanation class goldblum2024nfl. In the presence of proxies, mechanistic ML should adopt the same norms in pursuit of scientific discovery. The central question is not “how flexible is the learner?” but “under what assumptions and what data design is the mechanistic query $q(M)$ identified?” Appendix (ref) gives six examples.

Appendix (ref) defines identifying structure with mathematical formalism, while Appendix (ref) gives an example of how this conceptualization can be used to identify minimal identification assumptions for a specific class of models. This is also why several successful scientific-ML systems should be viewed as positive cases rather than counterexamples. Equation-discovery and inverse-problem methods such as symbolic regression, SINDy, and PINNs succeed precisely when they combine expressive function classes with explicit structural restrictions on the mechanism, the observation process, or the experimental regime. Their success comes not from black-box flexibility alone, but from the identifying structure that makes mechanistic queries auditable. In practice, identifying structure yields three auditable manifestations. First, a workflow can declare assumptions that restrict both the mechanism and the measurement channel, so that readers can see which equivalence classes have been ruled in and ruled out. Second, it can collect and evaluate on discriminating data: interventions, invariances across environments, and derivative constraints that are implied by the mechanism but not by proxy correlations alone. This is where design enters: the call is not just for more data, but for new data that breaks observational equivalence by changing environments, manipulating inputs, or measuring closer to the latent drivers. Without a notion of identifying structure, there is no guidance about how to collect data. Third, when assumptions and available data still leave ambiguity, a workflow can report multiplicity: characterize the remaining identified set of mechanistic answers (or mechanisms) rather than collapsing it into a single narrative.

Mendel's case worked because favorable biology and discriminating design lined up. On the mechanism side, his hypothesis class was sharply constrained: particulate inheritance with segregation and independent assortment. On the observation side, he used selected discrete phenotypes whose relationship to latent hereditary factors was stable enough to generate sharp, falsifiable implications. These restrictions were paired with controlled crosses, replication, and targeted contrasts, such as F$_2$ ratios and testcrosses, that would look different under alternative mechanisms mendel1866,curtis2023mendel. The lesson is not that design alone made genetics simple, but that mechanistic discovery requires an evidence regime in which relevant alternatives imply different observable patterns. Miasmatic theories of epidemic disease, by contrast, were supported by rich but indirect proxies: odor, dampness, crowding, poor drainage, sewage, low elevation, seasonality, and neighborhood mortality. These proxies were genuinely predictive, and sanitary reforms motivated by miasmatic reasoning often improved health, which made the theory empirically plausible. But the same proxy patterns were also compatible with germ-theoretic accounts, because waste, water, housing, and occupational environments also structured microbial exposure and host susceptibility. The eventual shift toward specific pathogen explanations required evidence that decoupled these mechanisms: water-source and exposure contrasts in cholera, laboratory isolation and culture, experimental inoculation, and later microbiological assays snow1855,pasteur1861,koch1884. The lesson for mechanistic ML is therefore not “collect big data,” but “collect the right data to shrink $\mathcal M(P_X)$.”

This paper therefore proposes the following auditable standards as community expectations for any workflow, especially LLM-centered workflows, that present outputs as mechanistic. First, Mechanism cards: authors should combine every mechanistic explanation with a short structured appendix that (i) states the mechanism-side assumptions (what is invariant under interventions, which functional forms/constraints are being imposed, what is excluded), (ii) states the observation-side assumptions (how proxies relate to latent drivers; what is assumed stable vs.\ artifactual in measurement), and (iii) states what the assumptions actually identify: which aspects of the mechanism are pinned down, and which remain an equivalence class (Appendix (ref)). A mechanism card is incomplete without at least one explicit falsifier: an additional environment, intervention, new measurement channel, or invariance/derivative test that would rule out the proposed mechanism. Second, Multiplicity statements: when identification is unavailable, the system must represent the identified set rather than output a single story. Concretely, it should report a set of distinct compatible mechanisms or a set of answers to the query $q$ (e.g., a sign-robust derivative range, a range of counterfactual effects), and it should label which qualitative conclusions are stable across that set. Third, \textbf{Mechanism-discriminating evaluation}: mechanistic claims must be evaluated on targets that directly probe identifying structure like interventions, invariances across environments, proxy-invariance checks tied to measurement, and derivative constraints.

If a paper or product claims a black box has “discovered a mechanism,” reviewers and readers should ask: what is assumed about the mechanism, what is assumed about the measurement channel, what data or contrasts actually discriminate among alternatives, and what else remains compatible? If the answer is unclear, the authors should phrase the results as bounded evidence: a coherent hypothesis, simulation, or partially identified set consistent with the observed proxies on their realized support, together with a clear statement of what the evidence rules out and what remains compatible. In these settings, human experts remain central because the key tasks are precisely to state defensible assumptions, design discriminating experiments, and decide which claims are prediction-only versus mechanistic.

Alternative Views

\paragraph{Alternative View 1: “This is just epistemic uncertainty; ensembles and calibration solve it.”} Ensembles and calibration are valuable, but they primarily address uncertainty about predictions under a fixed observational target distribution. They do not, by themselves, identify mechanisms when proxies admit multiple observationally compatible causal stories (Sections (ref)--(ref)). The relevant ambiguity here is structural identification uncertainty: even with unlimited proxy data from the same measurement channel, mechanistic queries \(q(M)\) can differ across \(M\in\mathcal M(P_X)\). Diversity helps precisely when it surfaces disagreement, but that requires an interface that does not collapse into one story.

\paragraph{Alternative View 2: “Scaling and better data will select the true mechanism.”} This objection is partly right: better data help when they add mechanistically informative variation---new environments, interventions, or measurement channels closer to the latent drivers---because these can shrink \(\mathcal M(P_X)\). However, scale alone can also improve in-domain fit while leaving \(\mathcal M(P_X)\) large. In this framework, this objection is best understood as a claim about the observational regime \(r\): if scaling implicitly changes \(r\) (e.g., by adding environments or new measurement channels), then it is adding identifying structure. The request is that such structure be made explicit and evaluated via discriminating tests, rather than assumed. Corollary (ref) makes this precise: if two observationally compatible mechanisms disagree on \(q\), no increase in sample size within the same observational regime can point-identify \(q\). \paragraph{Alternative View 3: “Identification is too strong a standard; useful science often proceeds with partial evidence.”} This objection is right, and it is part of the position here rather than an exception to it. Many scientific and policy tasks do not require point-identifying a general mechanism. Partial identification, triangulation, process tracing, negative controls, and falsification tests can rule out candidate explanations, bound plausible effects, shift beliefs, guide data collection, and justify provisional action under uncertainty. The standard advocated here is narrower: when a workflow claims to have discovered a generalizable, intervention-stable mechanism, it should either supply the identifying structure that makes that claim answerable or report the remaining multiplicity. Evidence short of identification is valuable precisely when it is labeled as such.

\paragraph{Alternative View 4: “Narrative collapse is a UI issue, not a scientific issue.”} The point is well taken: narrative collapse is partly an interface issue. When the interface encourages single-story explanations in underidentified settings, it invites overclaiming. Related phenomena like model collapse and algorithmic monoculture can amplify the problem in ecosystems, but narrative collapse can occur even in a single interaction on a fixed dataset. The appropriate response is not only better UX; it is reporting and evaluation norms that force mechanistic claims to be conditional on discriminating tests and to represent multiplicity when identification is unavailable.

\paragraph{Alternative View 5: “Occam/simplicity will pick the right mechanism among observationally equivalent ones.”} Simplicity biases (explicit or implicit) can be effective inductive biases in practice, and they often guide scientific progress. The point here is that simplicity is itself an identifying assumption: it selects one member of a proxy-compatible equivalence class. When a workflow reports a single mechanism on the basis of simplicity, it should declare (i) the simplicity criterion being used (description length, sparsity, functional form, etc.), (ii) which alternatives remain compatible absent that criterion, and (iii) falsifiers or regime shifts that would discriminate between the selected mechanism and plausible competitors.

\paragraph{Alternative View 6: “Mechanisms are only defined up to query-specific equivalence; reporting one representative is fine.”} Mechanism claims are indeed query-specific: if two mechanistic stories imply the same answer to the intended interventional query \(q\), then they are equivalent for that query. The concern is precisely that LLM-centered workflows tend to output a rich narrative that contains many additional mechanistic commitments not warranted by the evidence (or by the intended query), and to present those commitments as uniquely supported. A safer norm is to separate (i) query-stable conclusions from (ii) narrative embellishments that vary across \(M\in\mathcal M(P_X)\).

\paragraph{When ML/LLM-centered workflows can genuinely support discovery.} This paper argues against treating fluent narratives or predictive success as evidence of mechanism discovery absent identifying structure, not against complex models. AlphaFold2, for example, pairs model complexity with a narrow target query, exact sequence inputs, evolutionary signal from alignments/templates, experimentally determined training structures, geometry-aware inductive biases, and blind evaluation against withheld structures jumper2021alphafold,nobel2024chemistry. LLM-centered workflows can be constructive when they generate candidate mechanisms with explicit identifying assumptions, propose discriminating experiments, measurements, or invariance tests that shrink \(\mathcal M(P_X)\), and summarize evidence as identified sets rather than single narratives when identification is unavailable. Relatedly, recent work on open-ended autonomous discovery proposes steering hypothesis search by Bayesian surprise agarwal2025autodiscovery; the concern here is orthogonal: even principled exploration must either represent multiplicity or introduce discriminating structure when the same observational channel admits multiple incompatible mechanisms.

Conclusion

Without explicit identifying structure, it is not possible to tell whether a fluent model has learned a mechanism or merely a representation. The position here would be weakened by a demonstrated class of proxy-rich observational regimes in which (i) mechanistic claims can be reliably distinguished, tightly bounded, or otherwise made robust using only declared observational structure, and (ii) LLM-centered workflows reliably surface those limits, bounds, and falsifiers rather than collapsing to a single narrative. High-performing predictive models can still be scientifically valuable for emulation, screening, design, measurement, and hypothesis generation even when they are not mechanistically faithful. The requirement is claim-evidence alignment: predictive or descriptive success should not be automatically elevated into mechanism discovery. Absent such evidence, the position is twofold: build models that predict well, and require mechanistic claims to be stated in a form that supports falsification, sensitivity analysis, and explicit accounting of multiplicity. Otherwise, the field may well end up “discovering” the next miasma.

Acknowledgements

Support for this research came from the Eunice Kennedy Shriver National Institute of Child Health and Human Development under Award Number R01HD107015 and Award Number R21HD119931, as well as from Grant N000142512270 from the Office of Naval Research. Thanks to Izabel Aguiar, Arun Chandrasekhar, Jishnu Das, Avi Feller, Markus Goldstein, Rachel Heath, Jason Kerwin, Jeff Lockhart, Bodhisattwa Majumder, Harsh Parikh, Karl Rohe, Stephen Salerno, Elizabeth Stewart, and the anonymous referees for constructive comments on earlier versions of the draft. All errors and shortcomings remain mine despite their guidance.