EconBase
← All papers

Synthetic Potential Outcomes and Causal Mixture Identifiability

Bijan Mazaheri, Chandler Squires, Caroline Uhler

arXiv 29 May 2024 · Machine Learning

arXiv:2405.19225 · PDF · DOI · OpenAlex · Extracted main text

Abstract

Heterogeneous data from multiple populations, sub-groups, or sources is often represented as a “mixture model” with a single latent class influencing all of the observed covariates. Heterogeneity can be resolved at multiple levels by grouping populations according to different notions of similarity. This paper proposes grouping with respect to the causal response of an intervention or perturbation on the system. This definition is distinct from previous notions, such as similar covariate values (e.g. clustering) or similar correlations between covariates (e.g. Gaussian mixture models). To solve the problem, we “synthetically sample” from a counterfactual distribution using higher-order multi-linear moments of the observable data. To understand how these “causal mixtures” fit in with more classical notions, we develop a hierarchy of mixture identifiability.

Citation extraction

45
references
57
in-text mentions
45
distinct cited
5
self-citations
8,344
main-text words

appendix boundary found by appendix_command · 67% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Allman, E. S., Matias, C., and Rhodes, J. A (2009) Identifiability of parameters in latent structure models with many observed variables0.7373367%
2Pearl, J (2009) Causality0.64441100%
3Gordon, S., Mazaheri, B., Schulman, L. J., and Rabani, Y (2020) The sparse hausdorff moment problem, with application to topic models self0.5112250%
4Kim, Y., Koehler, F., Moitra, A., Mossel, E., and Ramnarayan, G (2019) How many subpopulations is too many? exponential lower bounds for inferring population histories0.5112250%
5Imbens, G. W. and Rubin, D. B (2015) Causal inference in statistics, social, and biomedical sciences0.51121100%
6Miao, W., Shi, X., Li, Y., and Tchetgen Tchetgen, E. J (2024) A confounding bridge approach for double negative control inference on causal effects0.51121100%
7Suk, Y., Kim, J.-S., and Kang, H (2021) Hybridizing machine learning methods and finite mixture models for estimating heterogeneous treatment effects in latent classes0.51121100%
8Tchetgen, E. J. T., Ying, A., Cui, Y., Shi, X., and Miao, W (2020) An introduction to proximal causal learning0.51121100%
9Kossaifi, J., Panagakis, Y., Anandkumar, A., and Pantic, M (2019) Tensorly: Tensor learning in python0.51121100%
10Abadie, A., Diamond, A., and Hainmueller, J (2010) Synthetic control methods for comparative case studies: Estimating the effect of california’s tobacco control program0.40511100%

Showing the top 10 of 45 scored citations.