Yizhi Liu, Balaji Padmanabhan, Siva Viswanathan
arXiv 2 Mar 2026 · Artificial Intelligence
arXiv:2603.02359 · PDF · DOI · OpenAlex · Extracted main text
Digital advertising increasingly relies on visual content, yet marketers lack rigorous methods for understanding how specific visual attributes causally affect consumer engagement. This paper addresses a fundamental methodological challenge: estimating causal effects when the treatment, such as a model's skin tone, is an attribute embedded within the image itself. Standard approaches like Double Machine Learning (DML) fail in this setting because vision encoders entangle treatment information with confounding variables, producing severely biased estimates. We develop DICE-DML (Deepfake-Informed Control Encoder for Double Machine Learning), a framework that leverages generative AI to disentangle treatment from confounders. The approach combines three mechanisms: (1) deepfake-generated image pairs that isolate treatment variation; (2) DICE-Diff adversarial learning on paired difference vectors, where background signals cancel to reveal pure treatment fingerprints; and (3) orthogonal projection that geometrically removes treatment-axis components. In simulations with known ground truth, DICE-DML reduces root mean squared error by 73-97% compared to standard DML, with the strongest improvement (97.5%) at the null effect point, demonstrating robust Type I error control. Applying DICE-DML to 232,089 Instagram influencer posts, we estimate the causal effect of skin tone on engagement. Standard DML produces diagnostically invalid results (negative outcome R^2), while DICE-DML achieves valid confounding control (R^2 = 0.63) and estimates a marginally significant negative effect of darker skin tone (-522 likes; p = 0.062), substantially smaller than the biased standard estimate. Our framework provides a principled approach for causal inference with visual data when treatments and confounders coexist within images.
appendix boundary found by appendix_command · 91% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Victor Chernozhukov, Denis Chetverikov, Mert Demirer, Esther Duflo,… (2018) Double/debiased machine learning for treatment and structural parameters | 1.000 | 5 | 4 | 100% |
| 2 | Osei Appiah (2001) Ethnic identification on adolescents' evaluations of advertisements | 0.644 | 2 | 2 | 100% |
| 3 | Jochen Hartmann, Mark Heitmann, Christian Siebert, and Christina Sch… (2023) More than a feeling: Accuracy and application of sentiment analysis | 0.644 | 2 | 2 | 100% |
| 4 | Sven Klaassen, Julius Teichert-Kluge, Philipp Bach, Victor Chernozhu… (2024) DoubleMLDeep: Estimation of causal effects with multimodal data | 0.644 | 2 | 2 | 100% |
| 5 | Xiao Liu, Bin Zhang, Anjana Susarla, and Rema Padman (2020) Go to YouTube and call me in the morning: Use of social media for chronic conditions | 0.644 | 2 | 2 | 100% |
| 6 | Joann Peck and Barbara Loken (2004) When will larger-sized female models in advertisements be viewed positively? | 0.644 | 2 | 2 | 100% |
| 7 | Victor Veitch, Dhanya Sridhar, and David Blei (2020) Adapting text embeddings for causal inference. In Proceedings of the Conference on Uncertainty in Artificial Intelligence. 919–928 | 0.644 | 2 | 2 | 100% |
| 8 | Shunyuan Zhang, Dokyun Lee, Param Vir Singh, and Kannan Srinivasan (2022) What makes a good image? Airbnb demand analytics leveraging interpretable image features | 0.644 | 2 | 2 | 100% |
| 9 | Adrien Bardes, Jean Ponce, and Yann LeCun (2022) VICReg: Variance-invariance-covariance regularization for self-supervised learning. In Proceedings of the International Conferen… | 0.511 | 2 | 2 | 50% |
| 10 | Alain Chardon, Isabelle Cretois, and Colette Hourseau (1991) Skin colour typology and suntanning pathways | 0.511 | 2 | 1 | 100% |
Showing the top 10 of 37 scored citations.