EconBase
← All papers

On the causality-preservation capabilities of generative modelling

Yves-Cédric Bauwelinckx, Jan Dhaene, Tim Verdonck, Milan van den Heuvel

arXiv 3 Jan 2023 · Machine Learning · publishedJournal of Computational and Applied Mathematics (2024) · 3 citations (OpenAlex)

arXiv:2301.01109 · PDF · DOI · OpenAlex · Extracted main text

Abstract

Modeling lies at the core of both the financial and the insurance industry for a wide variety of tasks. The rise and development of machine learning and deep learning models have created many opportunities to improve our modeling toolbox. Breakthroughs in these fields often come with the requirement of large amounts of data. Such large datasets are often not publicly available in finance and insurance, mainly due to privacy and ethics concerns. This lack of data is currently one of the main hurdles in developing better models. One possible option to alleviating this issue is generative modeling. Generative models are capable of simulating fake but realistic-looking data, also referred to as synthetic data, that can be shared more freely. Generative Adversarial Networks (GANs) is such a model that increases our capacity to fit very high-dimensional distributions of data. While research on GANs is an active topic in fields like computer vision, they have found limited adoption within the human sciences, like economics and insurance. Reason for this is that in these fields, most questions are inherently about identification of causal effects, while to this day neural networks, which are at the center of the GAN framework, focus mostly on high-dimensional correlations. In this paper we study the causal preservation capabilities of GANs and whether the produced synthetic data can reliably be used to answer causal questions. This is done by performing causal analyses on the synthetic data, produced by a GAN, with increasingly more lenient assumptions. We consider the cross-sectional case, the time series case and the case with a complete structural model. It is shown that in the simple cross-sectional scenario where correlation equals causation the GAN preserves causality, but that challenges arise for more advanced analyses.

Citation extraction

60
references
74
in-text mentions
60
distinct cited
0
self-citations
8,816
main-text words

appendix boundary found by appendix_command · 100% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Goodfellow, I. J., J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farle… (2014) Generative adversarial networks1.00053100%
2Xu, L., M. Skoularidou, A. Cuesta-Infante, and K. Veeramachaneni (2019) Modeling tabular data using conditional gan0.64422100%
3Feng, Q., C. Guo, F. Benitez-Quiroz, and A. Martinez (2022) When do gans replicate? on the choice of dataset size0.64422100%
4Scholkopf, B., F. Locatello, S. Bauer, N. R. Ke, N. Kalchbrenner, A.… (2021) Toward causal representation learning0.64422100%
5Yoon, J., D. Jarrett, and M. van der Schaar (2019) Time-series generative adversarial networks0.64422100%
6Wen, B., L. Colon, K. Subbalakshmi, and R. Chandramouli (2021, 04) (2021) Causal-tgan: Generating tabular data using causal generative adversarial networks0.64422100%
7Kocaoglu, M., C. Snyder, A. G. Dimakis, and S. Vishwanath (2017) Causalgan: Learning causal implicit generative models with adversarial training0.64422100%
8Yoon, J., J. Jordon, and M. van der Schaar (2019) PATE-GAN: Generating synthetic data with differential privacy guarantees0.64422100%
9van Breugel, B., T. Kyono, J. Berrevoets, and M. van der Schaar (2021) Decaf: Generating fair synthetic data using causally-aware generative networks0.51121100%
10Radford, A., L. Metz, and S. Chintala (2015) Unsupervised representation learning with deep convolutional generative adversarial networks0.51121100%

Showing the top 10 of 60 scored citations.