EconBase
← All papers

Rethinking Distributional IVs: KAN-Powered D-IV-LATE & Model Choice

Charles Shaw

arXiv 15 Jun 2025 · Econometrics

arXiv:2506.12765 · PDF · DOI · OpenAlex · Extracted main text

Abstract

The double/debiased machine learning (DML) framework has become a cornerstone of modern causal inference, allowing researchers to utilise flexible machine learning models for the estimation of nuisance functions without introducing first-order bias into the final parameter estimate. However, the choice of machine learning model for the nuisance functions is often treated as a minor implementation detail. In this paper, we argue that this choice can have a profound impact on the substantive conclusions of the analysis. We demonstrate this by presenting and comparing two distinct Distributional Instrumental Variable Local Average Treatment Effect (D-IV-LATE) estimators. The first estimator leverages standard machine learning models like Random Forests for nuisance function estimation, while the second is a novel estimator employing Kolmogorov-Arnold Networks (KANs). We establish the asymptotic properties of these estimators and evaluate their performance through Monte Carlo simulations. An empirical application analysing the distributional effects of 401(k) participation on net financial assets reveals that the choice of machine learning model for nuisance functions can significantly alter substantive conclusions, with the KAN-based estimator suggesting more complex treatment effect heterogeneity. These findings underscore a critical "caveat emptor". The selection of nuisance function estimators is not a mere implementation detail. Instead, it is a pivotal choice that can profoundly impact research outcomes in causal inference.

Citation extraction

26
references
38
in-text mentions
24
distinct cited
0
self-citations
8,953
main-text words

appendix boundary found by none_found · 100% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Kratsios, A., & Furuya, T (2025) Kolmogorov-Arnold Networks: Approximation and Learning Guarantees for Functions and their Derivatives0.92844100%
2Liu, Z., Wang, Y., Vaidya, S., Ruehle, F., Halverson, J., Soljacić,… (2024) Kan: Kolmogorov-arnold networks0.92843100%
3Chernozhukov, V., Chetverikov, D., Demirer, M., Duflo, E., Hansen, C… (2018) Double/debiased machine learning for treatment and structural parameters0.84333100%
4Abadie, A (2002) Bootstrap tests for distributional treatment effects in instrumental variable models0.64422100%
5Angrist, J. D., Imbens, G. W., & Rubin, D. B (1996) Identification of causal effects using instrumental variables0.64422100%
6Imbens, G. W., & Angrist, J. D (1994) Identification and estimation of local average treatment effects0.64422100%
7Imbens, G. W., & Rubin, D. B (1997) Estimating outcome distributions for compliers in instrumental variables models0.64422100%
8Chernozhukov, V., & Hansen, C (2004) The impact of 401 (k) participation on the wealth distribution: an instrumental quantile regression analysis0.58531100%
9Athey, S., Tibshirani, J., & Wager, S (2019) Generalized random forests0.40511100%
10Belloni, A., Chernozhukov, V., & Hansen, C (2014) Inference on treatment effects after selection among high-dimensional controls0.40511100%

Showing the top 10 of 24 scored citations.