Hong Kiat Tan, Isaac-Neil Zanoria, James Chen, Haoyang Lyu, Mihai Cucuringu
arXiv 4 Oct 2026 · Statistics — Methodology
arXiv:2610.05618 · PDF · Extracted main text
Finding which variables cause which others in multivariate time series, and at what lags, is central to science and policy, yet existing methods force a choice between flexible confounder adjustment, data-driven lag selection, and inference that controls the false discovery rate (FDR). ORACLE-VARX does all three in one pipeline. First, double/debiased machine learning (DML) removes nonlinear confounder effects from the outcomes and the lagged series. Second, adaptive causal lag estimation (ACLE) picks the lag order at each time step by sequential significance tests, tracking regime changes. Third, entry-wise $z$-tests with Benjamini--Hochberg correction select directed edges at a target FDR. We prove that in each rolling window, the debiased coefficients are asymptotically normal around a window-averaged target, so their $z$-tests are asymptotically valid. On a synthetic benchmark with time-varying structure and nonlinear confounding, ORACLE-VARX (LightGBM) tracks the true lag order best (RMSE $0.96$ vs $1.1$--$1.5$), has edge FDR $0.047$, close to PCMCI ($0.045$) and below VAR ($0.129$) and VAR-LiNGAM ($0.187$), and forecasts better than all three. On nine U.S. sector ETFs with macroeconomic confounders, it yields interpretable causal graphs whose lag order rises in high-volatility regimes.
appendix boundary found by appendix_command · 21% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Victor Chernozhukov, Denis Chetverikov, Mert Demirer, Esther Duflo,… (2018) Double/debiased machine learning for treatment and structural parameters | 0.928 | 5 | 4 | 80% |
| 2 | Helmut Lütkepohl (2005) New Introduction to Multiple Time Series Analysis | 0.843 | 3 | 3 | 100% |
| 3 | Aapo Hyvärinen, Kun Zhang, Shohei Shimizu, and Patrik O. Hoyer (2010) Estimation of a structural vector autoregression model using non-Gaussianity | 0.737 | 4 | 4 | 50% |
| 4 | Jakob Runge, Peer Nowack, Marlene Kretschmer, Seth Flaxman, and Dino… (2019) Detecting and quantifying causal associations in large nonlinear time series datasets | 0.737 | 4 | 4 | 50% |
| 5 | Clive W. J. Granger (1969) Investigating causal relations by econometric models and cross-spectral methods | 0.737 | 4 | 3 | 50% |
| 6 | Yoav Benjamini and Yosef Hochberg (1995) Controlling the false discovery rate: a practical and powerful approach to multiple testing | 0.737 | 3 | 2 | 100% |
| 7 | Daniele Ballinari and Alexander Wehrli (2024) Semiparametric inference for impulse response functions using double/debiased machine learning | 0.644 | 2 | 2 | 100% |
| 8 | Roxana Pamfil, Nisara Sriwattanaworachai, Shaan Desai, Philip Pilger… (2020) DYNOTEARS: Structure learning from time-series data | 0.644 | 2 | 2 | 100% |
| 9 | Agathe Sadeghi, Achintya Gopal, and Mohammad Fesanghary (2025) Causal discovery from nonstationary time series | 0.644 | 2 | 2 | 100% |
| 10 | Christopher A. Sims (1980) Macroeconomics and reality | 0.644 | 2 | 2 | 100% |
Showing the top 10 of 43 scored citations.