EconBase
← All papers

Position: Prioritize Identifying Structure, Not Complex Models, for Scientific Discovery

Tyler H. McCormick

arXiv 30 May 2026 · Statistics — Machine Learning

arXiv:2606.02632 · PDF · DOI · OpenAlex · Extracted main text

Abstract

Modern Machine Learning (ML) and Artificial Intelligence (AI) models, especially large language models (LLMs), are increasingly used to generate scientific hypotheses and mechanistic explanations from observational data. This position paper argues that in the high-dimensional proxy regimes where modern ML excels, mechanistic learning is generically underdetermined: many incompatible mechanisms induce essentially the same observational relationships on the support of the data, so predictive success and coherent explanations are insufficient evidence of mechanism discovery. This underdetermination becomes uniquely hazardous with large language models (LLMs), which tend to collapse large equivalence classes of explanations into a single fluent narrative. This paper proposes concrete standards for “mechanistic ML,” and argues these norms are necessary if LLM-centered workflows are to support science rather than merely simulate it.

Citation extraction

60
references
84
in-text mentions
60
distinct cited
2
self-citations
7,618
main-text words

appendix boundary found by appendix_command · 50% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Curtis, David (2023) Mendel Did Not Study Common, Naturally Occurring Phenotypes0.84333100%
2Mendel, Gregor (1866) Versuche über Pflanzen-Hybriden0.73732100%
3Kirchhof, Michael and Kasneci, Gjergji and Kasneci, Enkelejda (2025) Position: Uncertainty Quantification Needs Reassessment for Large Language Model Agents0.64422100%
4Morioka, Hiroshi and Hyvarinen, Aapo (2024) Causal Representation Learning Made Identifiable by Grouping of Observational Variables0.5115220%
5Brunton, Steven L. and Proctor, Joshua L. and Kutz, J. Nathan (2016) Discovering governing equations from data by sparse identification of nonlinear dynamical systems0.5113233%
6Raissi, M. and Perdikaris, P. and Karniadakis, G.E (2019) Physics-informed neural networks: A deep learning framework for solving forward and inverse problems involving nonlinear partial…0.5113233%
7Schmidt, Michael and Lipson, Hod (2009) Distilling Free-Form Natural Laws from Experimental Data0.5113233%
8Khemakhem, Ilyes and Kingma, Diederik and Monti, Ricardo and Hyvarin… (2020) Variational Autoencoders and Nonlinear ICA: A Unifying Framework0.5113233%
9Locatello, Francesco and Bauer, Stefan and Lucic, Mario and Raetsch,… (2019) Challenging Common Assumptions in the Unsupervised Learning of Disentangled Representations0.5113233%
10Gelman, Andrew and Hullman, Jessica and Kennedy, Lauren (2024) Causal Quartets: Different Ways to Attain the Same Average Treatment Effect0.5112250%

Showing the top 10 of 60 scored citations.