EconBase
← All papers

Causality Elicitation from Large Language Models

Takashi Kameyama, Masahiro Kato, Yasuko Hio, Yasushi Takano, Naoto Minakawa

arXiv 4 Mar 2026 · Machine Learning

arXiv:2603.04276 · PDF · DOI · OpenAlex · Extracted main text

Abstract

Large language models (LLMs) are trained on enormous amounts of data and encode knowledge in their parameters. We propose a pipeline to elicit causal relationships from LLMs. Specifically, (i) we sample many documents from LLMs on a given topic, (ii) we extract an event list from from each document, (iii) we group events that appear across documents into canonical events, (iv) we construct a binary indicator vector for each document over canonical events, and (v) we estimate candidate causal graphs using causal discovery methods. Our approach does not guarantee real-world causality. Rather, it provides a framework for presenting the set of causal hypotheses that LLMs can plausibly assume, as an inspectable set of variables and candidate graphs.

Citation extraction

19
references
20
in-text mentions
19
distinct cited
0
self-citations
4,785
main-text words

appendix boundary found by none_found · 100% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Judea Pearl (2000) Causality: Models, Reasoning, and Inference0.51121100%
2David Maxwell Chickering (2002) Optimal structure identification with greedy search0.40511100%
3P. Christen (2012) Data Matching: Concepts and Techniques for Record Linkage, Entity Resolution, and Duplicate Detection0.40511100%
4Agata Cybulska and Piek Vossen (2014) Using a sledgehammer to crack a nut? lexical diversity and event coreference resolution0.40511100%
5Ahmed K. Elmagarmid, Panagiotis G. Ipeirotis, and Vassilios S. Veryk… (2007) Duplicate record detection: A survey0.40511100%
6Ivan P. Fellegi and Alan B. Sunter (1969) A theory for record linkage0.40511100%
7Matthew Gentzkow, Bryan Kelly, and Matt Taddy (2019) Text as data0.40511100%
8Justin Grimmer and Brandon M. Stewart (2013) Text as data: The promise and pitfalls of automatic content analysis methods for political texts0.40511100%
9Oktie Hassanzadeh, Debarun Bhattacharjya, Mark Feblowitz, Kavitha Sr… (2020) Causal knowledge extraction through large-scale text mining0.40511100%
10Guido W. Imbens and Donald B. Rubin (2015) Causal Inference for Statistics, Social, and Biomedical Sciences: An Introduction0.40511100%

Showing the top 10 of 19 scored citations.