EconBase
← All papers

How Many Samples Are Needed to Determine Causal Direction? Sharp Minimax Bounds for Bivariate LiNGAM

Jikai Jin

arXiv 16 Aug 2026 · Mathematics — Statistics Theory

arXiv:2608.15840 · PDF · Extracted main text

Abstract

We study how many observations are needed to determine the causal direction between two linearly related variables. Classical LiNGAM theory shows that independent non-Gaussian disturbances identify the direction, but does not quantify the difficulty when the causal effect is weak or the disturbances are nearly Gaussian. Let $β$ bound the absolute structural coefficient from below, let $ν$ measure each standardized disturbance's distance from Gaussianity, and let the disturbance scales lie in $[\underlineσ,\overlineσ]$. We prove the sharp local minimax law \[ N_2^\star(β,ν,δ) \asymp \frac{\log(1/δ)} {d_β^2+β^2ν^2}, \qquad d_β= \left[β^2- \left(1-\frac{\underlineσ^2}{\overlineσ^2}\right)\right]_+. \] Previous theory established population identifiability or assumed a fixed separation between the two directions. By contrast, we establish the sharp sample complexity as a joint function of edge strength, distance from Gaussianity, and scale uncertainty, and characterize when identification comes from non-Gaussian dependence or from covariance alone. The proof was independently generated with GPT-5.6 Sol in Codex's Ultra mode during a two-hour session. The human author supplied the prompt and was responsible only forchecking the proof and revising and polishing the manuscript.

Citation extraction

15
references
29
in-text mentions
15
distinct cited
0
self-citations
17,334
main-text words

appendix boundary found by none_found · 100% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1A. M. Kagan, Yu. V. Linnik, and C. R. Rao. Characterization Problems… (1973)0.84333100%
2G. M. Feldman and P. Graczyk. The Skitovich–Darmois theorem for loca… (2010)0.64422100%
3M. I. Gabovich. Stability of the characterization of Gaussian distri… (1974)0.64422100%
4M. I. Gabovich. Stability of a characterization theorem for the norm… (1981)0.64422100%
5K. Genin and C. Mayo-Wilson. Success concepts for causal discovery:… (2024) https://doi.org/10.1007/s41237-022-00188-60.64422100%
6A. Hyvärinen and S. M. Smith. Pairwise likelihood ratios for estimat… (2013)0.64422100%
7F. Laplante, C. Ambroise, and P. Humbert. Contrast-free ICA and caus… (2026)0.64422100%
8P.-L. Loh and P. Bühlmann. High-dimensional learning of linear causa… (2014)0.64422100%
9S. Oh, S. Han, and G. Park. Optimal estimation of linear non-Gaussia… (2025)0.64422100%
10J. Peters and P. Bühlmann. Identifiability of Gaussian structural eq… (2014)0.64422100%

Showing the top 10 of 15 scored citations.