EconBase
← Back to paper

Visual Polarization Measurement Using Counterfactual Image Generation

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

144,951 characters · 25 sections · 91 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Visual Polarization Measurement Using Counterfactual Image Generation

abstractPolitical polarization is a significant issue in American politics, influencing public discourse, policy, and consumer behavior. While studies on polarization in news media have extensively focused on verbal content, non-verbal elements, particularly visual content, have received less attention due to the complexity and high dimensionality of image data. Traditional descriptive approaches often rely on feature extraction from images, leading to biased polarization estimates due to information loss. In this paper, we introduce the Polarization Measurement using Counterfactual Image Generation (PMCIG) method, which combines economic theory with generative models and multi-modal deep learning to fully utilize the richness of image data and provide a theoretically grounded measure of polarization in visual content. Applying this framework to a decade-long dataset featuring 30 prominent politicians across 20 major news outlets, we identify significant polarization in visual content, with notable variations across outlets and politicians. At the news outlet level, we observe significant heterogeneity in visual slant. Outlets such as Daily Mail, Fox News, and Newsmax tend to favor Republican politicians in their visual content, while The Washington Post, USA Today, and \textit{The New York Times} exhibit a slant in favor of Democratic politicians. At the politician level, our results reveal substantial variation in polarized coverage, with Donald Trump and Barack Obama among the most polarizing figures, while Joe Manchin and Susan Collins are among the least. Finally, we conduct a series of validation tests demonstrating the consistency of our proposed measures with external measures of media slant that rely on non-image-based sources. {\bf Keywords:} Polarization, News Media, Politics, Generative Models, Computer Vision, Counterfactual Reasoning

\thispagestyle{empty}

Introduction

Political polarization has emerged as a central issue in American politics over the past few decades. Three out of ten Americans now consider polarization one of the most significant challenges facing the country skelley_fuong_2022. When asked to describe the current political climate, the term “divisive” was the most frequent response, and similar terms like “polarized” and “partisan” were among the most common responses by Americans pew2023. As individuals become more deeply rooted in their political identities, the potential for cross-party dialogue, compromise, and effective governance diminishes. Indeed, the impact of polarization extends beyond politics; it affects public policy, social cohesion, and the overall functioning of democratic institutions. Thus, it is important to have precise and systematic measures to study polarization.

Several studies have focused on the verbal content in news articles as a rich source of information to propose measures for media bias and polarization and examine the demand-driven motives for news outlets to create content that matches the partisan preferences of their readers gentzkow2006media, gentzkow2010drives. The research in this domain suggests that the choice of words and framing of issues can reveal accurate information about the speaker or the writer gentzkow2019measuring. The existence of ideological slanting in the choice of words and framing poses important questions about the non-verbal aspect of news content. The non-verbal content conveys meaning through channels other than language, such as facial expression and body language. In news articles, the non-verbal content is often communicated through visual content such as images. Editors often pay great attention to the choice of visual content, as visuals are more memorable, processed more rapidly, and elicit stronger emotional responses than text sullivan1988happy, TownsendKahn2014, blanchard2023extraction.\footnote{A large body of work in visual marketing and eye-tracking data demonstrates the value of visual information wedel_pieters_2007, chandon2009does, wedel2023modeling.} Visuals not only shape engagement but also play a key role in spreading misinformation matatov2022stop and influencing voter perceptions of competence and trustworthiness in politics hoegg2011impact. Moreover, younger generations of consumers are showing a growing preference for news content that relies less on verbal and more on visual elements. Nevertheless, only a few studies have focused on the visual content to study polarization peng2018same, boxell2021slanted, ash2021visual, caprini2023visual, and no prior work has utilized the richness and dimensionality of visual information.

In this paper, we bridge this gap, and develop a framework to measure polarization and slant in visual content. Figure (ref) shows a small sample of images used by {\it CNN} and {\it Fox News} to portray the 2024 presidential nominees, Donald Trump and Kamala Harris. Visually, we can see a more positive portrayal of Donald Trump (Kamala Harris) by {\it Fox News} ({\it CNN}). Our goal in this paper is to propose a method to systematically quantify this form of visual slanting. In particular, we seek to answer the following questions:

enumerate• How can we quantify political polarization in visual content, particularly in the images used in news articles? • What is the extent of visual slant and polarization in mainstream media outlets? • How does the polarization in visual content vary across politicians and news outlets?
figure[figure omitted — 231 chars of source]

There are three key challenges that we need to address to answer these questions satisfactorily. First, we need a formal definition of a metric or parameter of interest that captures polarization in visual content. To address this challenge, we turn to the very structure of the problem: the editor's choice of visual content. An ideologically slanted choice of image intuitively means that a news outlet prefers a more positive (negative) portrayal of a politician from the same (opposite) ideological side, compared to a neutral news outlet. We construct a flexible utility framework that models the editor's choice of images from a set of available options and focus on the smile as a focal feature we use to measure visual polarization given the consistent finding that a smile leads to a more positive portrayal of a politician sulflow2019power. Using a utility framework allows us to compare the utility derived from two images that are on all aspects except the focal feature (e.g., smile) on which the visual polarization is measured. This comparison enables us to build a polarization measure centered around the focal feature of interest. For instance, consider two images of Donald Trump that differ only in whether he is smiling. For each outlet, we can define the difference in the utility from using the image with and without a smile. Intuitively, the greater the variability in this utility difference across outlets, the higher the level of political polarization in visual content. To capture this notion of variability for a given pair of news outlets, we define the difference in this utility difference as the visual polarization measure between any pair of outlets. When one news outlet in the pair is a neutral outlet, the visual polarization measure reflects how slanted the image choice is by the other outlet, allowing us to characterize a visual slant measure. Together, we develop a utility framework that allows us characterize both visual slant and visual polarization.

Our second challenge stems from the richness and high dimensionality of the visual content. The common approach in the literature is to use an off-the-shelf machine learning model that extracts a certain feature (e.g., smile in an image) from the image peng2018same, boxell2021slanted. However, this feature extraction approach relies solely on the selected feature and disregards other information in the image, which can introduce both omitted variable bias and extraction bias in the analysis wei2022unstructured. We address this challenge by employing a generative approach that creates counterfactual versions of the same image with and without the feature of interest. This method enables us to retain all the information in the image and minimizes the degree to which other factors can confound our polarization measure.

The third challenge lies in identifying our polarization parameter from the observed data, which is defined based on the utility functions of news outlets. This challenge arises because these utility functions are not directly identified from the data as we only observe the image selected by each outlet, but not the full choice set available to them. To address this issue, we leverage the variation in image choices across outlets for similar events (e.g., a specific press hearing), under the assumption that the choice set is nearly identical for all outlets covering the same event. For instance, for a similar event, both {\it Fox News} and {\it CNN} likely have access to almost identical sets of images sourced from common providers such as Getty Images or Associated Press, which allows us to assess whether one outlet derives greater utility from a positive portrayal of a politician compared to the other. We then theoretically link the identification of the polarization parameter to a news outlet prediction problem and develop a multi-modal deep learning model for this task. The model is designed to capture both the clustering structure in similar events and the subtle facial features that may reflect outlets' ideological preferences (if any). Finally, the estimates derived from this prediction task allow us to measure polarization at any desired level of granularity.

Together, we build a unified framework that combines the economic structure of the problem with generative models to generate comparable counterfactual images and measure polarization in visual content. The key advantage of our framework in comparison to the traditional regression-based approaches lies in its ability to overcome the potential confounding and mis-estimation of polarization due to information loss from feature extraction. Additionally, our generative strategy extends beyond the study of polarization and can be applied to other contexts involving visual data. Another key benefit of our framework is its ability to provide individual-level measures, enabling us to quantify the heterogeneity in visual polarization across politicians and news outlets.

We apply our framework to a comprehensive dataset comprising over 60,000 images of 30 prominent politicians from both the Republican and Democratic parties across 20 major news outlets over a 10-year span from 2011 to 2021. First, we employ a multi-modal deep learning model to predict the outlet of an image based on its visual information, the textual content of the article, and contextual data about the politician, year, and other relevant factors. We then employ Generative Adversarial Networks (GANs) to create sets of counterfactual images for all politicians with and without a smile. Using the model estimated in the initial step, we measure how the predicted probability of the image belonging to a particular outlet changes between these counterfactual images. Finally, we connect these differences to our visual slant measure and quantify the extent of polarization at both the aggregate and individual levels.

Our results suggest that both Democratic- and Republican-leaning outlets ideologically slant their visual content. Using Reuters as the base neutral news outlet, we find that compared to a neutral image of a Republican politician, a smiling image increases utility for a Republican-leaning news outlet and decreases utility for a Democratic-leaning outlet. Conversely, for Democratic politicians, a smiling image generates lower utility for Republican-leaning outlets and higher utility for Democratic-leaning outlets compared to the neutral image. These results suggest that news outlets exhibit a positive visual slant when covering politicians who share their ideological leanings and a negative visual slant when covering those with opposing views. We perform formal statistical tests and show that the distributions of visual slant are significantly different, highlighting the divide in their portrayal of politicians from either side.

We then examine the heterogeneity of visual slant across news outlets, documenting substantial variation in how politicians are portrayed. For Democratic politicians, USA Today, The New York Times, and The Washington Post exhibit the highest positive visual slant in their favor, while Daily Mail and Fox News display the strongest negative slant. For Republican politicians, the most positive visual slant scores appear in Daily Mail and \textit{Newsmax}, whereas \textit{The Washington Post} and \textit{CNN} show the most negative slant. To quantify overall outlet-specific visual slant, we introduce a new measure, \textit{Conservative Visual Slant (CVS)}, which captures the degree to which an outlet's visual content favors Republican politicians while disadvantaging Democratic ones. {\it CVS} measure calculates the difference between an outlet's visual slant for Republican and Democratic politicians. Based on this measure, we identify \textit{Daily Mail}, \textit{Fox News}, and \textit{Newsmax} as outlets with the highest \textit{CVS} values, while \textit{The Washington Post}, \textit{USA Today}, and \textit{The New York Times} as the outlets the most negative \textit{CVS} values.

We also document significant variation in the degree of polarization across individual politicians. To capture this, we introduce a metric called {\it Overall Visual Polarization (OVP)}, which quantifies the standard deviation of a politician’s visual slant measures across media outlets. A higher {\it OVP} indicates greater dispersion in how different outlets visually present a politician, suggesting a more polarizing figure. On the Republican side, Donald Trump emerges as the most polarizing figure. That is, the gap in the extent of ideological slanting is remarkably large, with the Republican-leaning outlets receiving higher utility from using a positive and smiley portrayal of him compared to Democratic-leaning outlets. On the Democratic side, we find Barack Obama, Bernie Sanders, and Kamala Harris to be among the most polarizing figures who are portrayed very differently in Democratic-leaning and Republican-leaning outlets. Interestingly, Joe Manchin (D) and Susan Collins (R) rank among the least polarizing politicians, reflecting their reputations as moderates within their respective parties. We also observe low polarization levels for Liz Cheney, whose fallout with Donald Trump resulted in increased favorability among liberal outlets and decreased favorability among conservative outlets, ultimately contributing to reduced levels of visual polarization.

Lastly, we conduct a series of tests to validate the measures derived from our algorithm. First, we compare our {\it Conservative Visual Slant (CVS)} measure with existing measures of media slant from the prior literature flaxman2016filter, faris2017partisanship using a series of correlation tests. Second, we compare the performance of our {\it CVS} measure against the visual slant metric proposed by boxell2021slanted and demonstrate that our measure is more effective at capturing variations in external media slant indicators. Finally, we validate our {\it Overall Visual Polarization (OVP)} measure by showing a strong correlation between {\it OVP} scores and the ideological alignment of a politician’s constituency. Together, these validation tests underscore our algorithm’s ability to capture media polarization using images used by news outlets.

In summary, our paper makes several contributions to the literature. Methodologically, we propose a framework that combines economic theory with generative models to provide robust measures of media bias and polarization in visual content. A key innovation of our framework is the use of Generative Adversarial Networks (GANs) in a way consistent with experimentation to generate counterfactual images based on a feature of interest, addressing the bias due to information loss present in social science studies that utilize image data. As such, our framework is general and applicable to all settings where researchers seek to quantify polarization on a given feature across a given set of images and outlets. Substantively, our work demonstrates the existence of ideological slanting in the visual content used by media outlets, with substantial heterogeneity observed across news outlets and politicians.

Related Literature

Our paper relates to the study of political polarization in news media, a phenomenon well documented using text data. groseclose2005measure quantifies media bias by examining the frequency with which different media outlets cite various think tanks, uncovering a persistent liberal bias. Similarly, gentzkow2010drives investigate the factors driving media slant, highlighting the significant roles of consumer preferences and political affiliations by analyzing the alignment of newspaper language with political parties. jensen2012political study the polarization of political discourse by analyzing records of Congressional speech and the Google Ngrams corpus, discovering a notable increase in discourse polarization since the late 1990s. Furthermore, gentzkow2019measuring focus on the linguistic divide in Congressional speeches, showing how Democrats and Republicans increasingly use distinct vocabularies. While these studies predominantly focus on textual data, our work shifts the focus to polarization in visual content by examining how images in news articles contribute to this phenomenon.

More recently, research in this area has expanded from textual to visual content due to advances in computer vision techniques. peng2018same use a two-stage approach with Azure Microsoft to extract facial expressions from 13,000 images of the 2016 election in the first stage and then analyze them through a regression model in the second stage, revealing bias toward politicians aligned with a news outlet's stance. boxell2021slanted expand this by analyzing 70,000 images across more politicians and outlets. caprini2023visual further extend this by generating and analyzing textual descriptions of images alongside news articles using Azure, applying gentzkow2019measuring to show how the alignment of visual and textual bias amplifies polarization. Our study advances previous research in three significant ways. First, we enhance the traditional two-stage approach by addressing its inherent biases and issues with omitted variables, introducing a more robust methodology called Polarization Measurement Using Counterfactual Image Generation (PMCIG) that uses the rich information in the image content to measure the polarization in visual content. Second, our method delivers results that are both precise and available at a lower level of granularity, enabling us to quantify polarization at both the news outlet and politician levels and to track their evolution over the past decade. Third, we leverage a long term dataset spanning from 2011 to 2021, which allows us to capture long-term trends in political coverage that earlier studies may have missed.

From a methodological perspective, our work aligns with the growing trend in social science research that leverages computer vision techniques to analyze unstructured image data. One stream of work uses image data to generate descriptive insights on consumers and firms using novel measurement techniques dew2022letting, liu2020visual. Another stream of work employs a two-stage approach and tries to connect image features to economic/marketing outcomes of interest by first, extracting features from images to create structured data and then applying statistical analysis to these features davenport2017analytics. However, as discussed earlier, this approach has limitations, including potential information loss, endogeneity issues, and oversimplification due to parametric assumptions. Recently, researchers have started paying attention to this problem and proposing potential solutions. wei2022unstructured identify biases in econometric models using machine-learned variables from unstructured data and propose solutions to improve accuracy. singh2023causal introduce the RieszIV estimator, which incorporates high-dimensional unstructured data directly into causal analysis to manage endogeneity. xu2024unstructured propose a debiased embedding framework that integrates representation learning with causal inference, addressing biases inherent in traditional embedding-then-inference frameworks. Additionally, luo2024using employ GANs for Controllable Stimuli Generation (CSG), enabling precise manipulation of image attributes to isolate causal effects. Similarly, li2024product advances AI-driven product design by integrating consumer preferences from internal data and external user-generated content into a generative framework, addressing the limitations of traditional GAN-based approaches. Building on these advancements, our PMCIG method combines GAN-based image manipulation with non-parametric deep learning models to more accurately quantify polarization in an image feature, holding all other features constant.

Setting and Data

We collect publicly available data on images of politicians from media websites spanning a 10-year period for the study. Below, we describe our data collection, cleaning, and labeling strategy.

Data Collection Strategy

We collect data on 30 politicians and 20 news outlets over a span of ten years, from 2011 to 2021. The set of politicians consists of those who ran for important public offices and/or held important national roles during the ten-year span of 2011--2021 and were commonly searched on Google (based on their popularity on Google searches using Google Trends data). The news outlets used in our study consist of the list of popular outlets that have been used in earlier studies on media polarization; see flaxman2016filter for details. We refer readers to Web Appendix $\S$(ref) for a full list of politicians and outlets.

The data collection is done using the SerpAPI application serpapi. SerpAPI facilitates efficient large-scale image scraping by leveraging Google's image search to extract relevant images and metadata. The process involves generating search queries based on specific criteria, such as targeted news outlets and date ranges, to extract images, article links, titles, and dates.\footnote{For each politician-outlet combination, we use a two-year rolling window that advances by one year at a time, resulting in 10 queries spanning a total of 10 years.} This approach was chosen because it provides a structured, efficient, and reliable interface for large-scale data collection, automating the process while ensuring compliance with Google's data access policies.

figure[figure omitted — 254 chars of source]

Figure (ref) presents an example of a query we use for data collection, where we use “site:" to restrict searches to specific news websites, “before:" and “after:" to filter results by publication date, and exact phrase searches to capture relevant content precisely. For each query, we aim to collect 80 news articles for each politician from each news outlet for the specified period.\footnote{Some queries retrieve fewer articles/images for certain politicians-outlet combinations because the outlet may not have published 80 images for that specific politician.} This information for each image in the query is compiled into a structured data frame with the following fields:

list{$\bullet$} { {0pt} {1pt} {1pt} {1pt} {1.5em} {1em} {0.5em} } • Image: A URL pointing to the image. • Alt: A short description or alternate text corresponding to the image. • Href: A hyperlink where the related news article can be found. • Title: The title of the news article or caption associated with the image. • Query Parameter: The search query string utilized to retrieve the image and associated news, emphasizing the political figure and the source website along with a defined temporal span.

Overall, our data collection strategy gives us a comprehensive data set of 287,275 images for 30 politicians from 20 news outlets.

Data Cleaning and Identifying Politicians

A crucial step is cleaning the data to ensure that each image contains the face of the politician referred to in the query. We face the following challenges in the data-cleaning step:

list{$\bullet$} { {0pt} {1pt} {1pt} {1pt} {1.5em} {1em} {0.5em} } • Some images do not feature the intended politician but instead capture relevant scenes or contexts of the news without the individual's presence. • Certain images, although retrieved under a specific individual's search query (e.g., Joe Biden), may inadvertently include another politician (e.g., Donald Trump), introducing cross-representation. • Several images include multiple politicians, complicating the analysis of each politician's records.\footnote{We focus on single-face images to ensure that each news outlet’s selection reflects its portrayal of the intended politician, avoiding confounding effects from multiple individuals in the same image.}

To address these challenges, we design a two-phase computer vision framework. We provide a brief overview of this framework here and refer readers to Web Appendix $\S$(ref) for the technical details. In the first phase, we use a series of computer vision models to keep only images with one face presented in them. In the second phase, we build and deploy a face-verification tool to ensure that the face in an image belongs to the intended politician. First, we manually select 20 high-quality, single-face images for each of the 30 politicians. These selected images serve as true labels identifying the correct politician, ensuring a reliable foundation for training the face verification model. Next, we use the trained verification model for each instance in our one-face sample obtained from the first phase to verify that the predicted label for the face is the same as the intended politician. After applying the above data cleaning procedure, we are left with a set of 63,188 images, where each image shows a single face that belongs to the intended politician. Web Appendix $\S$(ref) provides further details on our two-step model and includes a comprehensive table (Figure (ref)) listing the number of images for each politician-outlet combination after the cleaning.

Problem Definition

figure[figure omitted — 273 chars of source]

Recall that our objective is to see if and how images of politicians in media outlets can be used to measure political polarization. As such, our goal here is to develop measures of {\it visual slant} and {\it visual polarization} that can be estimated from data on news articles (with images of politicians).

Consider a dataset of news articles, $\mathcal{D}$. Each news article $i$ in this dataset is characterized by a tuple $(X_i, P_i, Y_i, Z_i)$, where $X_i$ represents the features associated with the key aspects of the article, such as its title text, topic, and publication date. $P_i$ refers to the politician who is the focal subject of the article, $Y_i$ is the news outlet that produces the article, and $Z_i$ is the image. $Z_i$ can be interpreted as detailed pixel-level information capturing all the relevant aspects of the image, such as its background, brightness, other objects present, and the facial expression of the politician $P_i$.

Next, we define the image choice problem for a given article $i$ from the perspective of a news editor. Prior research has shown that images can have a significant impact on the extent to which readers engage with and click on news content MatiasEtAl2021. As such, editors must carefully select images for each article, which involves choosing an image that appeals to the target audience and aligns with the editorial stance while ensuring that the visual elements complement the textual content jakesch2022belief. Figure (ref) shows an example news article, highlighting the decision stage where an editor picks an accompanying image for the given article.

Formally, let the editor or news outlet receive utility $U_i$ from producing article $i$ with characteristics $(X_i,P_i,Y_i,Z_i)$, which is defined as follows:

equation[equation omitted — 68 chars of source]

where $u(Z_i, X_i, P_i, Y_i)$ is the deterministic component and represents the editor's expected utility from producing the article, and $\xi_i$ denotes the idiosyncratic error term, which follows an i.i.d. Type 1 Extreme Value distribution.\footnote{We model the news outlet's decision-making using a utility function rather than a conventional profit function, as ideological preferences may drive them to deviate from profit-maximizing choices. For other theoretical models that micro-found agents' decision-making to examine polarization in equilibrium, see gentzkow2010drives, iyer_yoganarasimhan_2021, amaldoss2021media and bondi2023privacy.} This utility function captures all the main considerations of the editor when making image choices, such as how well the image supports the outlet's ideological stance (alignment with editorial policy, i.e., alignment of features of $Z_i$ with $Y_i$), the image's ability to attract and retain readers (reader-engagement), and the aesthetic quality and/or emotional impact of the image (visual appeal).

Intuitively, ideological slanting in images means that the outlet receives a higher (lower) utility from a more positive visual portrayal of a politician from the same (opposite) ideological side compared to a neutral outlet. The key challenge in developing a measure of ideological slant in visual content stems from the ambiguity around the definition of a “positive” portrayal: an image is an unstructured and high-dimensional object, and there are presumably numerous ways for the outlet to choose a more or less positive image. As such, we need to define ideological slanting for a certain image feature. In our analysis, we focus on the presence of smile as the feature of interest because there is consensus that smile is a feature that contributes to a more positive portrayal sulflow2019power. However, our framework is not restricted to this particular feature and can easily extend to other features or sets of features.

Let $T_i$ denote whether the subject in image $Z_i$ is smiling or not. To define the ideological slant measure for feature $T$, we start with a utility difference measure that compares the utility from two articles that are identical in all respects, except for the smile feature. Similar to the causal inference literature, we therefore consider two versions of an image $Z$: $Z(T=1)$ and $Z(T=0)$, where $Z(T=1)$ is the image with a smile and $Z(T=0)$ is the same image without a smile (with a neutral expression). For ease of exposition, let $Z^{(-T)}$ denote all the information in the image $Z$, except the smile feature $T$. This implies that $Z(T=1) \equiv (T=1, Z^{(-T)})$ and $Z(T=0) \equiv (T=0, Z^{(-T)})$, i.e., the two versions of the image are identical in all features other than smile \footnote{The notation assumes that $Z^{(-T)}$ is independent of $T$, meaning all non-smile features remain identical across conditions. This holds if the images are constructed such that only the smile varies, ensuring $Z^{(-T)} \mid T=1 \sim Z^{(-T)} \mid T=0$.}. Using this set of two images, we can define a measure of utility difference for a given news outlet $y$ and politician $p$ with respect to visual feature $T$ as follows:

equation[equation omitted — 163 chars of source]

This utility difference helps isolate the utility increase or decrease for a news outlet that solely comes from the presence of a smile in the image of any given politician. However, this is still not a measure of ideological slant in visual content because we cannot fully attribute this utility difference to ideological preferences. For instance, there can be vertical preferences for the smile feature where all news outlets prefer a smiling image of a politician. As such, the difference in the utility difference $\Delta^T u (p,y)$ across two news outlets helps cancel out the vertical preference for the feature and the remaining difference can be linked to ideological preferences. For example, we expect a news outlet with a higher conservative audience share (e.g., {\it Fox News}) to have a higher $\Delta^T u (p,y)$ for a conservative politician (e.g., Donald Trump) than a news outlet with lower conservative audience share (e.g., {\it CNN}) even if they both have a vertical preference for smiling images. Since this difference is defined for any two news outlets, it measures visual polarization of the two outlets with respect to the feature of interest. However, it is important to note that this is not a measure of visual slant, because both outlets may have slanted preferences. In what follows, we present formal definitions for both visual polarization and visual slant. We first present the formal definition for visual polarization as follows:

defnFor a given politician $p$ and any pair of news outlets $y_1$ and $y_2$, {\bf visual polarization} with respect to feature $T$ is defined as follows: \begin{equation} \rho^T (p,y_1,y_2) = \Delta^T u (p,y_1) - \Delta^T u (p,y_2). \end{equation}

This definition of {\it visual polarization}, $\rho^T (p,y_1,y_2)$, is flexible and allows for detailed analysis at the level of individual politicians and news outlets. For example, by examining the above measure for a specific politician like Donald Trump across specific outlets such as {\it CNN} and {\it Fox News}, we can measure the extent to which these outlets are polarized or differentiated in their visual portrayal of Trump. Naturally, for the measure of visual polarization to capture visual slant, we need the news outlet $y_2$ to be neutral. Let $y_n$ denote the neutral news outlet. We can formally define visual slant as follows:

defnFor a given politician $p$ and any news outlet $y_1$, {\bf visual slant} with respect to feature $T$ is denoted by $\rho_s^T$ and defined as follows: \begin{equation} \rho^T_s (p,y_1) = \Delta^T u (p,y_1) - \Delta^T u (p,y_n), \end{equation} where $y_n$ is the neutral news outlet.

To further clarify the difference between visual polarization and visual slant, we return to the example with the portrayal of Donald Trump in {\it CNN} and {\it Fox News}, but add a neutral news outlet like Reuters. Suppose that the visual polarization between {\it Fox News} and {\it CNN} for Donald Trump is equal to one, i.e., $\rho^T (\text{Donald Trump},\text{Fox News},\text{CNN})=1$. This measure is a composite of both Fox News' preference for a positive portrayal of Trump and {\it CNN}'s preference for a negative portrayal of him. Our visual slant measure helps decompose the visual polarization measure. For example, we may find that there is a positive visual slant for Donald Trump at {\it Fox News} such that $\rho^T_s (\text{Donald Trump},\text{Fox News})=0.4$, but a negative \textit{visual slant} for him at {\it CNN} such that $\rho^T_s (\text{Donald Trump},\text{CNN})=-0.6$. It is easy to verify the following relationship between the two definitions:

equation[equation omitted — 77 chars of source]

Lastly, a notable feature of our measures is that they are defined at a high degree of granularity, which allows us to capture the varying degrees of ideological slant in visual content specific to each outlet or each politician. This specification allows us to consider different types of aggregation:

list{$\bullet$} { {0pt} {1pt} {1pt} {1pt} {1.5em} {1em} {0.5em} } • First, by aggregating the {\it visual slant} measure over Democratic (Republican) politicians for a given outlet, we can obtain insights into how a given outlet portrays liberal (conservative) politicians. This analysis can tell us how conservative outlets like {\it Fox News} and liberal outlets like {\it CNN} portray the two groups of politicians differently. • Second, by aggregating a given politician's {\it visual slant} measure across all the outlets and then comparing these measures across politicians, we can derive insights into how different politicians are portrayed in the media. For instance, this can help us understand questions such as whether the portrayal of Donald Trump in the media is more (or less) polarizing than that of Joe Biden.

Standard Reduced-Form Approach and its Limitations

In this section, we discuss the standard reduced-form approach used to measure polarization in visual content. In $\S$(ref), we describe the two-step reduced-form approach and connect the estimated parameter under this approach to the measure of polarization in visual content as defined in Equation (ref)). Next, in $\S$(ref), we provide a theoretical and conceptual discussion of the drawbacks of this approach.

Two-step Model

Recall that our goal is to measure {\it visual polarization}, \( \rho^T (p,y_1,y_2) \), with respect to a focal feature $T$, such as whether the politician in the image is smiling or not. However, this is inherently challenging when dealing with unstructured image data, because unlike the standard causal inference literature, where treatment is observed imbens_rubin_2015, here we do not directly observe $T_i$, i.e., the presence or absence of the treatment (smile in this case) in a given article $i$. As a result, a common approach in social science settings is to use a {\it Two-step Approach}. This approach addresses the high dimensionality and complexity of unstructured data by first extracting meaningful features and then using these features in a structured econometric model. See davenport2017analytics for a general discussion of this approach in the broader social sciences literature, and boxell2021slanted and peng2018same for applications of this approach to the context of image polarization in media.\footnote{Beyond the political polarization context, this two-step approach is now commonly used in marketing research involving images in other settings as well. For instance, unstructured data such as video and audio streams (\( Z \)) are analyzed to extract features like facial expressions (\( T \)), which are then used to study their impact on user engagement (\( Y \)) lu2021larger. In the online labor market, profile pictures (\( Z \)) are examined to identify features such as perceived race or attire (\( T \)) to assess job offer likelihood (\( Y \)), highlighting visual biases in employment opportunities troncoso2022look. Additionally, research on social media posts about e-cigarettes (\( Z \)) extracts demographic features (\( T \)) to study tax policy compliance (\( Y \)), revealing demographic responses to legislative changes anand2024frontiers.} Methodologically, the two-step approach can be outlined as follows:

enumerate• Feature extraction using a machine learning model: The first step involves using an off-the-shelf machine learning model, denoted as \( f_1 \), to extract relevant feature(s) \( T \) from the unstructured data \( Z \). This model \( f_{1}: Z \rightarrow T \) is used to derive the feature(s) \( T \): \begin{equation} \hat{T} = f_{1}(Z) \end{equation} • Econometric analysis: After extracting features \( \hat{T} \), the second step involves analyzing these features within a structured econometric model. There are two potential ways to relate \( \hat{T} \) and \( Y \): \begin{subequations} \begin{align} Y &= f_{2}^{T-independent}(\hat{T}, \boldsymbol{X}) + \epsilon, \;\;\; or, \\ \hat{T} &= f_{2}^{T-dependent}(Y, \boldsymbol{X}) + \epsilon, \end{align} \end{subequations} where $f_{2}^{\text{T-independent}}$ is the second-stage econometric model where extracted feature $\hat{T}$ is used as an independent variable, and $f_{2}^{\text{T-dependent}}$ is the second-stage model where $\hat{T}$ is used as dependent variable. Social scientists typically prefer the second form because \( \hat{T} \) is an estimated variable that could include measurement error, and therefore using it as the independent variable is not appropriate bollen2009causal.

We now describe how this approach can be used to quantify polarization in visual content with respect to feature \(T\), defined as \(\rho^T (p,y_1,y_2)\) in Equation (ref). Consider a simple example where news outlets from Democratic-leaning outlets like the {\it The New York Times}, {\it CNN}, and {\it BBC} (\(y_1\)) and Republican-leaning outlets like {\it Fox News}, {\it Newsmax}, and {\it Daily Mail} (\(y_2\)) choose images for a politician, say Hillary Clinton (\(p\)). In this context, the first step involves estimating the binary indicator \(\hat{T}\), which denotes whether the focal politician (Hillary Clinton) is smiling in the image (\(\hat{T} = 1\)) or not (\(\hat{T} = 0\)). This can be obtained from standard off-the-shelf emotion recognition models such as Face++ faceplusplus_emotion, Google Vision google_ml_kit\footnote{microsoft_azure_vision’s Face API, a popular model previously used in studies such as boxell2021slanted, caprini2023visual, has been unavailable for use since June 2022 due to updated responsible AI policies microsoft2023responsible.}, or through in-house model training where we first label a sample of the images using human subjects and then train a model to predict the label given the image.

For the second step, we can use a logistic regression where we regress \(\hat{T}_i\) on characteristics \(X_i\) and \(Y_i\). For notational simplicity, consider the case where \(Y\) is a binary variable (e.g., \(Y=1\) for Democratic-leaning outlets, and \(Y=0\) for Republican-leaning outlets). We can write the log odds ratio for the resulting logistic regression as follows:

equation[equation omitted — 197 chars of source]

If the model in the equation above is well-specified, the log odds ratio characterizes the utility difference \(\Delta^T u(p, y)\) between choosing an image with a smile versus one without a smile, as defined in Equation (ref). In that event, we can connect the parameter \(\beta\) to the polarization measure defined in Equation (ref) as follows:

equation[equation omitted — 230 chars of source]

Thus, the coefficient \( \beta \) in the logistic regression model serves as an empirical estimate of the polarization measurement \( \rho^T(p, Y=1, Y=0) \). If \( \beta > 0 \), it suggests that Democratic-leaning outlets (\(Y = 1\)) derive a higher utility from using a smiling image of Hillary Clinton compared to Republican-leaning outlets, implying a greater \(\Delta^T u\) for the Democratic-leaning outlets. Conversely, if \( \beta < 0 \), Republican-leaning outlets (\(Y = 0\)) have a higher utility difference compared to Democratic-leaning outlets, indicating a greater \(\Delta^T u\). Therefore, \( \rho^T(p, Y=1, Y=0) \) can be approximated by \( \beta \), and provides a quantitative measure of the ideological polarization across news outlets.

Drawbacks of the Two-Step Approach

While the two-step model described in $\S$(ref) is easy to interpret and apply, it relies on the assumption that the regression model is well-specified. However, this assumption can fail due to two natural reasons and lead to incorrect inference: (1) extraction bias, and (2) omitted variable bias. In this section, we discuss these biases and explain how they can manifest in our setting.

{\it Extraction bias} occurs due to imperfections in the feature extraction process, or the first step of the two-step process wei2022unstructured. That is, the true label $T$ may differ from the extracted label \(\hat{T} = f_1(Z)\) as follows:

equation[equation omitted — 79 chars of source]

The error in Equation (ref) can be viewed as a measurement error. It is well-known that systematic dependence of this error on other relevant factors in the model can introduce biases or inconsistencies in the estimates; see Chapter 4 of wooldridge2010econometrics. As such, we decompose this error into two parts, such that $\epsilon_1 = \epsilon_r + \epsilon_e$, where $\epsilon_r$ is the random i.i.d. noise in the measurement process, and $\epsilon_e$ is the extraction error that can be correlated with the other relevant image features contained in \(Z^{(-T)}\). For instance, a machine learning model designed to detect the facial expressions (e.g., whether a person is smiling) may also inadvertently capture brightness levels in the image, as brighter images are often associated with happier expressions. We can rewrite the relationship between $T$ and $\hat{T}$ as:

equation[equation omitted — 76 chars of source]

Then, the predicted feature $\hat{T}$ is a combination of the true feature of interest $T$ (e.g., smile) and a correlated nuisance feature $\epsilon_e$ (e.g., brightness). Thus, when we perform the second stage estimation, the estimated effect ($\beta$) captures the effect of the outlet on the biased measurement $\hat{T}$ rather than the true feature $T$.

Next, omitted variable bias arises from relevant factors omitted from the second step. This issue can arise even if we have the true labels $T$. To illustrate this point, consider the second-step relationship between the true label and covariates:

equation[equation omitted — 60 chars of source]

where $\epsilon_2$ consists of unobserved variables that are not captured by the observed covariates $Y$ and $X$ through the semi-parametric function $f_2$. If this error term is independent of the covariates, we can identify function $f_2$ correctly. However, this is often a very strong assumption given the amount of information in \(Z^{(-T)}\) that is omitted from the model. For example, consider a collection of images by Donald Trump in articles by {\it CNN} and {\it Fox News}, where the smile labels are accurate, i.e., $\hat{T} = T$. Now, consider an image feature, such as the presence of the US flag. It is likely that the US flag is present more often in images from {\it Fox News}, given their higher share of nationalist viewers. At the same time, it is more likely for a politician to smile in formal settings with flags present in the background. In this scenario, estimating the parameters of Equation (ref) will lead to function $f_2$ picking up the association between the presence of the flag and the smile in the image because the presence of the flag is omitted from the model. Given the high-dimensional and unstructured nature of the images, numerous other image features can result in omitted variable bias in our estimates. To characterize the endogenous part of the error term $\epsilon_2$, we rewrite Equation (ref) as follows:

equation[equation omitted — 80 chars of source]

where $\epsilon_0$ is the part that is independent of covariates, and $\epsilon_{ov}$ is the part that is correlated with $Y$ or $X$, which can lead to omitted variable bias. It is worth noting that image features that lead to extraction bias can also lead to omitted variable bias. For example, omitting image features like brightness can lead to omitted variable bias if brightness is correlated with both the actual presence of a smile in the image (not the extracted one) and the news outlet.

Together, if we want to estimate the second step of the two-step model in $\S$(ref) by using $\hat{T}$ as the outcome, the errors in Equation (ref) also appear in the second step equation as follows:

equation[equation omitted — 112 chars of source]

where $\epsilon_{ov}$ leads to omitted variable bias, $\epsilon_e$ leads to extraction bias due to the use of a machine learning model in the first step, and $\epsilon_r$ contributes to higher uncertainty in model estimates. In summary, the information loss resulting from extracting a single feature from an image can introduce both extraction bias and omitted variable bias.

We present a more formal characterization of these biases and their derivations in Web Appendix $\S$(ref). We also provide empirical evidence of how the two-step approach can be biased in estimating visual polarization in our setting in Web Appendix $\S$(ref).

Our Approach: Polarization Measurement Using Counterfactual Image Generation

As discussed in $\S$(ref), the two-step approach is subject to information loss since it ignores the high-dimensional information in the image, which in turn can bias the estimated measure of polarization. Therefore, we develop a novel algorithm -- Polarization Measurement Using Counterfactual Image Generation (PMCIG) -- that recovers the polarization measure defined in $\S$(ref) without incurring information loss.

Recall that the {\it visual polarization} measure is defined as (see Definition (ref)): \[\rho^T(p, y_1, y_2) = \Delta^T u(p,y_1) - \Delta^T u(p,y_2),\] where $\Delta^T u (p,y) = \mathbb{E}_{X}\left[u \left( Z(T=1), X,P,Y \right) - u \left(Z(T=0),X,P,Y \right) \mid P = p, Y = y\right]$. Before outlining our approach to measuring visual polarization, we first present a thought experiment of what an ideal experiment to measure $\rho^T(p, y_1, y_2)$ would look like. In order to measure $\Delta^T u (p,y)$ for a given politician-outlet pair $(p,y)$ for a feature/treatment $T$, we need to observe the difference in the editors' utilities (or choice probabilities) for two images that are exactly the same on all dimensions except $T$. Thus, an ideal experiment to measure $\rho^T(p, y_1, y_2)$ would involve presenting the editors of outlets $y_1$ and $y_2$ two images of politician $p$, that are exactly the same on all dimensions except the feature of interest, i.e., one where the politician is smiling (or the feature of interest $T$ is turned on, i.e., $Z(T=1) \equiv (T=1, Z^{(-T)})$) and one where s/he is not smiling (i.e., $Z(T=0) \equiv (T=0, Z^{(-T)})$). Then, the observed difference in relative choice probabilities of the two images ($Z(T=1)$ and $Z(T=0$)) across the two outlets can be linked to polarization measure $\rho^T(p, y_1, y_2)$.

However, our data comes from an observational setting rather than this ideal experiment. Therefore, we only observe realized combinations of $Z$ and $y$. For example, for an image $Z$ with $T=1$, we may see that it was chosen by outlet $y_1$, as in the case where Biden's smiling image was chosen by {\it CNN} in Figure (ref). However, we do not see what would be the probability of this image being chosen by $y_1$ if everything else about it was held the same, with only the feature/treatment of interest turned off (i.e., a non-smiling version of the same image of Biden with $T=0$). Similarly, we also do not see what would have been the likelihood of outlet $y_2$ choosing this image for the two cases: when $T=1$ and when $T=0$. Thus, we only observe one realized combination of $y$-$T$ for an image $Z$, as shown in Table (ref). However, to quantify/measure polarization, we need to be able to reliably estimate/model the three other counterfactual outcomes for each image $Z$ in the data. This limitation is similar in spirit to the well-established unobservability challenge in the potential outcomes framework angrist2009mostly.

table[table omitted — 434 chars of source]

As we can see from Table (ref), there are two key challenges that we need to overcome to measure {\it visual polarization}.

list{$\bullet$} { {0pt} {1pt} {1pt} {1pt} {1.5em} {1em} {0.5em} } • Challenge 1: Counterfactual Image Generation \\ Our polarization measure is defined for two versions of an image that only differ in one feature ($T$). Using two versions of an image where the only difference is the presence of a smile ($T=1$ vs. $T=0$) addresses the issue of information loss in reduced-form approaches by using the rich information contained in images as opposed to mapping the image to a single feature. However, for any given image, we only have one version, where the image either has a smile or does not. Thus, for any focal image $Z$, the first challenge consists of obtaining two versions of the image that are exactly the same in all aspects except the smile. • Challenge 2: Identification of the Polarization Parameter from Observed Data \\ Second, even after we have two versions of each image ($Z(T=1)$ and $Z(T=0)$), we need to identify the polarization parameter $\rho^T(p, y_1, y_2)$ for the pair of news outlets ($y_1$, $y_2$), which is defined based on the news outlets' utility. However, identifying the utility function is not feasible because we do not observe the news outlets' choice set. Note that, in standard discrete choice models, identification of the utility function comes from observing which option an agent chooses from a set of alternatives train_2009. However, in our setting, we only observe the chosen image. Therefore, we need a systematic approach to map the observed data to the polarization parameter without directly observing/identifying the utility functions.

The rest of this section is organized as follows. First, in $\S$(ref), we present the overview of our solution to these two challenges. Next, in $\S$(ref), we present our unified algorithm for quantifying visual polarization.

Overview of the Solution

In this section, we present an overview of our solution to the two challenges described earlier. We first present the generative component in $\S$(ref) (to address Challenge 1). We then discuss our identification strategy in $\S$(ref) (to address Challenge 2).

Counterfactual Image Generation

To address the first challenge, we leverage the recent developments in generative image models that allow us to manipulate images such that we can modify one specific aspect of an image (e.g., feature $T$) while keeping everything else constant. Generative models have been employed in several recent studies on image analysis: athey2022smiles use generative models to adjust features such as smiles in profile images to examine their effects on user preferences and economic transactions in a micro-lending platform. ludwig2024machine utilize these models to manipulate facial features and explore how these alterations affect judicial decisions. Finally, in luo2024using, these models create facial images to assess the impact of perceived gender traits on discrimination within online marketplaces.

We define two operators, \(\pi^1\) and \(\pi^0\), which transform an image \(Z\) into either an image where the treatment/feature of interest is turned on \(Z(T=1)\) or one where it is turned off \(Z(T=0)\). The purpose of these operators is to create a set of control and treatment images that only differ in the feature of interest (e.g., the facial expression or smile).\footnote{We can easily extend this operation to a continuous case, where we manipulate the extent to which the feature is activated, e.g., the extent to which the politician smiles.}

list{$\bullet$} { {0pt} {1pt} {1pt} {1pt} {1.5em} {1em} {0.5em} } • \(\pi^1\): This operator ensures that the feature or emotion \(T\) is turned on. For example, when the treatment of interest is smile, then \(\pi^1\) transforms an image of a politician into one where the politician is smiling, regardless of the initial state of \(T\) in the original image. Formally, applying \(\pi^1\) to an image \(Z\) results in: \begin{equation} \pi^1(Z) = (T = 1, Z^{(-T)}) = Z^1. \end{equation} • \(\pi^0\): This operator ensures that the feature or emotion \(T\) is turned off. For instance, \(\pi^0\) transforms an image of a politician into one where the politician has a neutral expression, regardless of the initial state of \(T\). Formally, applying \(\pi^0\) to an image \(Z\) results in: \begin{equation} \pi^0(Z) = (T = 0, Z^{(-T)}) = Z^0. \end{equation}

For a set of original images, together, these two operators give us pairs of treated and control images.

figure[figure omitted — 223 chars of source]

{\bf Implementation Details:} We now discuss how we generate counterfactual images in our study. For each politician $p$, we first select three neutral images \( \tilde{Z}^{0}_{p1}, \tilde{Z}^{0}_{p2}, \tilde{Z}^{0}_{p3} \) (where $T=0$) from our dataset to ensure a balanced and unbiased representation of the politician.\footnote{Using three instead of one image ensures that the findings are not driven by the peculiarities of a specific image. In principle, it is possible to use a larger/fewer number of images for this task. We choose three because it balances noise reduction with the cost of generating counterfactual images. Focusing on three images also allows us to manually assess the quality of the generated images and avoid cases where the smile looks unnatural or unappealing.} As such, the function $\pi^0(\cdot)$ to generate the neutral image without the smile feature is an identity function. To generate the counterfactual version of each image with a smile, we employ \href{https://www.ailabtools.com/image-editor/}{AILabTools} as our \(\pi^1\) operator; this tool utilizes conditional Generative Adversarial Networks (cGANs) mirza2014conditional and produces high quality and realistic outputs.\footnote{While many such tools are available, we use AILabTools because it allows us to precisely modify the focal images (by adding a smile) while maintaining the authenticity of the other features in the original image.} Thus, for each neutral image, we generate corresponding smiling versions \( \tilde{Z}^{1}_{p1}, \tilde{Z}^{1}_{p2}, \tilde{Z}^{1}_{p3} \), where the only difference is the added attribute (smile). Figure (ref) illustrates how we apply $\pi^1$ to three neutral images of Donald Trump, which in turn gives us three corresponding neutral images of Donald Trump with a smile. We repeat this process for all the 30 politicians in our study to construct the set $\{ (\tilde{Z}^{0}_{p1}, \tilde{Z}^{0}_{p2}, \tilde{Z}^{0}_{p3} ,\tilde{Z}^{1}_{p1}, \tilde{Z}^{1}_{p2}, \tilde{Z}^{1}_{p3}) \}_p$.

Identification: Linking Polarization to the News Outlet Prediction Problem

We begin with a high-level overview of our identification strategy before formalizing our approach. The core challenge in this task lies in the fact that while we observe the image selected by a news outlet, we do not observe the full set of available choices. Our identification strategy leverages variation from similar events covered by multiple outlets (e.g., press conferences). We argue that in these cases, outlets have access to the same choice set, making the identification task analogous to a news outlet prediction problem: the selected image is the one that provides the highest utility to the outlet compared to all other similar images chosen by other outlets.

Recall that our measure is defined for a politician $p$ and any pair of news outlets $y_1$ and $y_2$. Consider a pair of images $\{ Z^{0}_{p}, Z^{1}_{p}\}$ for politician $p$, where $Z^{0}_{p}$ is the version of the image $Z$ with the feature $T$ turned off and $Z^{1}_{p}$ is the version with the feature $T$ turned on. Then, we can write the sample analogue estimator for the polarization parameter for this pair of images as follows:

equation[equation omitted — 237 chars of source]

where $\mathcal{X}_p$ is the set of all articles featuring politician $p$, so the sample analog estimator takes the average over all these articles featuring politician $p$. Clearly, if the utility function is known, we can directly estimate $\hat{\rho}(p,y_1,y_2)$ by integrating the above equation over all the articles. However, as highlighted in Challenge 2, the underlying utility function is not identified with our data because we do not observe the editors' choice sets of images.

We propose a solution to this identification challenge that directly links the {\it visual polarization} parameter to an outlet-prediction model that can be identified from the data given without observing/identifying utility functions. We now introduce some additional notation to help with this task. In a setting with two outlets $y_1$ and $y_2$, we define $\tilde{Y}$ as a pseudo-variable that denotes the news outlet with he highest utility for producing an article with the content $x$, politician $p$, and image $z$. Formally, we can define $\tilde{Y}$ as follows:

equation[equation omitted — 114 chars of source]

where $\xi_1$ and $\xi_2$ are independent and identically distributed terms that come from the Type 1 Extreme Value distribution. We can link the elements of the right-hand side (RHS) of Equation (ref) to pieces identifiable from the data at hand by re-writing the sample analog estimator for {\it visual polarization} as:

equation[equation omitted — 1,018 chars of source]

In deriving Equation (ref), we switch two elements (second line), apply $\log(\exp(\cdot))$ transformation (third line), and use a log-odds interpretation that connects the polarization parameter to a news outlet choice problem. As such, Equation (ref) allows us to define the polarization parameter as the difference between two log-odds ratios related to the pseudo-variable $\tilde{Y}$.

We now link the identification of the polarization parameter to an outlet prediction problem (given article and image features). This link depends on the link between the pseudo-variable $\tilde{Y}$ and the actual $Y$ variable that determines the news outlet for an article. The following proposition characterizes the assumption needed for our identification:

propositionSuppose that we have data $\mathcal{D} = \{(X_i, P_i, Y_i, Z_i) \}_i$, where if an article $i$ is produced by news outlet $Y_i$, then $Y_i$ has higher utility from this article than other outlets. Then, the identification of the predicted probabilities for news outlet prediction task results in the identification of the polarization parameter.

The proof of this proposition is simple: under the assumption that the outlet producing an article $i$ has the highest utility from it (compared to other outlets), we have $\tilde{Y} = Y$. Therefore, the identification of log-odds in Equation (ref) becomes equivalent to log-odds for variable $Y$. Let $\hat{g}(y \mid z,x,p)$ denote the probability of an article with image $z$, characteristics $x$, and politician $p$ is produced by news outlet $y$. If the condition in Proposition (ref) is satisfied, we can write {\it visual polarization} as follows:

equation[equation omitted — 349 chars of source]

As such, the critical question is where in our data the condition in Proposition (ref) is satisfied, that is, the production of an article by an outlet implies having the highest utility from producing it compared to other outlets. Such areas in the data satisfy $\tilde{Y} = Y$ and characterize our identifying variation. For that purpose, we focus on important events that are covered by all news outlets, such as important events like press conferences by the politician. For instance, consider different images of Donald Trump during the “Operation Warp Speed” speech in our dataset, as shown in Figure (ref). Despite being captured from the same event, these images present Trump in varying ways, ranging from serious to expressive. We argue that in such events, our assumption that the producing outlet has the highest utility is more reasonable because all outlets have access to the same set of images. Therefore, the news outlet prediction from the set \( \mathcal{Y} \) can potentially mimic the true choice model over the set of images \(\mathcal{Z}\). Later in $\S$(ref), we present results highlighting how our prediction model utilizes this identifying variation in the data.

figure[figure omitted — 259 chars of source]

A notable feature of our data is the presence of such similar patterns that form a clustering structure, where we observe extensive choice variation by outlets within a cluster of very similar events. In what follows, we discuss how we design the structure of our news outlet prediction model such that it will be able to fully exploit the variation in similar events.

{\bf Model Architecture for the Multi-Modal News Outlet Prediction Problem:} We now discuss the estimation of our news outlet prediction model, \( g(y_i \mid Z_i, \boldsymbol{X}_i, P_i; \theta) \). Given the multi-modal nature of inputs (e.g., text, image, categorical variables), we can use any flexible semi-parametric model that can capture complex patterns in the data. However, there are important challenges that we need to address. First, we need to ensure that the model utilizes the variation in similar events to satisfy the condition in Proposition (ref). Second, we need to ensure that our predictive model correctly identifies the link between the use of smile in an image and the outlet. For example, a purely loss-minimizing objective may end up estimating a model that underestimates the strength of the link between smile and the outlet by misattributing its link to features correlated with a smile. Therefore, instead of simply using a naive off-the-shelf prediction model, we impose some structure on the architecture of machine learning models we consider, as illustrated in Figure (ref).

figure[figure omitted — 218 chars of source]

The multi-modal deep learning model has the following architecture:

list{$\bullet$} { {0pt} {1pt} {1pt} {1pt} {1.5em} {1em} {0.5em} } • First, to leverage the variation in similar events, we use the contextual information in news articles. This includes textual data (\( X^{\text{text}} \)), publication dates (\( X^{\text{date}} \)), politicians' names (\( P \)), political affiliations (\( P^\text{aff} \)), and image data (\( Z \)). Textual data is processed using Latent Dirichlet Allocation (LDA), which encodes each news title as a 40-dimensional topic vector. Second, we capture metadata such as dates, names, and affiliations as categorical variables by embedding them into dense vectors. Then, a Attention Mechanism processes the structured data (\( X^{\text{text}} \), \( X^{\text{date}} \), \( P \), \( P^\text{aff} \)), and learns to prioritize the most relevant text and metadata features for the classification task vaswani2017attention. • Second, ResNet-101 processes the entire image \(Z\) by leveraging its capabilities from general image recognition tasks and extracts hierarchical and context-rich features from the broader visual content and scene information he2016deep. Intuitively, given similarities in the event/news covered, these two parts of the model architecture effectively capture a given outlet's stylistic preference for certain styles of text and image features. • Finally, to accurately capture smiles and link them to news outlets, we use MTCNN performs face detection and crops the face alone from the rest of the image zhang2016joint. Then, the detected face is passed through the VGG-Face network, which is pre-trained on facial expression data. This makes it highly effective at capturing facial attributes such as smiles parkhi2015deep. Finally, we apply Chunk attention to VGG-Face embeddings, and combine them with categorical data (politicians' names \( P \) and affiliations \( P^\text{aff} \)) to capture correlations between facial features and structured metadata related to the image. This step of separately capturing facial features and emotions ensures that the model accurately captures the relationship between smiles and news outlets, mitigating the risk of misattribution to correlated features.

The final classification layer combines all modalities. The model is trained using the AdamW optimizer loshchilov2017decoupled, which ensures efficient optimization and robust regularization. To train the model, we maximize the following entropy function over the training data:

equation[equation omitted — 237 chars of source]

The resulting model $\hat{g}$ is then used to estimate the polarization parameter using Equation (ref). Please see Web Appendix $\S$(ref) for additional details on implementation and model training.

Algorithm

We now present our algorithm, Polarization Measurement Using Counterfactual Image Generation (PMCIG) in Algorithm (ref). We split our data $\mathcal{D}$ into training and test sets, denoted by $\mathcal{D}_{\text{train}}$ and $\mathcal{D}_{\text{test}}$, respectively. We use the training data to train the deep learning model, and use the estimated model to calculate the polarization measure on the test data. Since our main goal is polarization measurement, this form of cross-fitting helps avoid overfitting bias. The algorithm consists of three steps as described below:

algorithm[algorithm omitted — 3,212 chars of source]
list{$\bullet$} { {0pt} {1pt} {1pt} {1pt} {1.5em} {1em} {0.5em} } • In the first step, we generate the counterfactual image versions for each politician. To do so, we take three neutral images from each politician $p$ from the test data and apply the function $\pi^1$ using the procedure outlined in $\S$(ref). An important consideration in choosing the neutral images is that the images are within the joint distribution of the training data. If the image is very different from those in the training data, our $\hat{g}$ estimates can be inaccurate. • In the second step of our algorithm, we use the model architecture presented in $\S$(ref) to estimate the model $\hat{g}$ using training data $\mathcal{D}_{\emph{train}}$, which is a random subset of 85% of all articles in data $\mathcal{D}$. • Finally, in the third step, we measure polarization on the held-out test data set $\mathcal{D}_{\emph{test}}$, which is not used in the process of model building. This step incorporates the idea of cross-fitting presented in the literature on the intersection of machine learning and econometrics, and ensures that our estimates do not exhibit overfitting bias gentzkow2019measuring, chernozhukov2018double. In the final part of our algorithm, we aggregate over all articles and counterfactual image versions to measure polarization for each politician $p$ in each outlet $y^k$ relative to a baseline outlet $y^0$.

In our empirical application, we use Reuters as the baseline news outlet given its reputation to be a non-partisan news source. However, we can easily change that baseline to any other outlet to obtain polarization measures. To the extent that Reuters is the neutral point zero on the spectrum, we can interpret our polarization estimates as measures of visual slant.

Empirical Evaluation and Findings

We now present the results from our empirical analysis. First, in $\S$(ref), we provide evidence demonstrating that the counterfactual images generated in the Step 1 of our algorithm are consistent with the original images, preserving all features except the added smile. Next, $\S$(ref), we present results on the performance of our news outlet prediction model from Step 2 of our algorithm. Finally, in $\S$(ref), we share the Step 3 results on estimates of political polarization in visual content, and present findings at both the news outlet and politician levels.

Validating Counterfactual Image Consistency

We briefly discuss the validation of the counterfactual images using our cGANs toolkit. It is important to ensure that adding a smile does not systematically change other contextual features of the image. Ensuring this consistency is essential for isolating the effect of the smile without introducing biases from other image characteristics, such as brightness or colorfulness (see $\S$(ref)). To that end, we measure a few other image characteristics (such as brightness and colorfulness) before and after applying the smile operator \( \pi^{1} \) and use the Kolmogorov-Smirnov (K-S) test to compare the distributions of these characteristics for the original and smiley images. Our tests suggest that there are no statistically significant differences between the original and smiley images for other contextual features. This confirms that the operator \( \pi^{1} \) does not introduce significant changes in other characteristics, ensuring that the counterfactual images remain consistent with the original ones. Detailed histograms and a complete analysis of the test results are shown in the Web Appendix $\S$(ref).

Performance of the News Outlet Prediction Model

In this section, we present the performance of our multi-modal multi-class classification problem for news outlet prediction. As mentioned earlier, we use an 85%-15% split between training and test data. We consider four measures to evaluate the model's predictive performance -- (1) Accuracy, (2) Precision, (3) Recall, and (4) Weighted Cross-Entropy (WCE). Detailed explanations of these metrics are provided in Web Appendix $\S$(ref).

We use the above performance measures to evaluate different versions of our deep learning model on the test data and present the results in Table (ref). Each row shows a more complex model that progressively adds more explanatory variables/features. The first row considers a model that only uses categorical inputs such as politician identity, party affiliation, and date. We then add textual information in the second row and add image features in the final row. Two key points emerge from Table (ref). First, our final outlet prediction model achieves a remarkably good performance of approximately 44%. It is worth emphasizing that a random benchmark only achieves a 5% accuracy (given that there are 20 news labels). Second, we notice that the image information greatly contributes to the predictive accuracy of the model: while the non-image models achieve a maximum accuracy of 16.72%, the multi-modal model achieves an accuracy of 43.81%. This aligns with prior research that highlights the importance of visual data in predictive tasks dzyabura2023leveraging.

table[table omitted — 588 chars of source]

We present additional details on the model's performance measures across outlets in Web Appendix (ref). In the rest of this section, we provide some additional interpretation and intuition for the results from the black-box predictive model. In $\S$(ref), we discuss how our model captures the information in a smile, and in $\S$(ref), we examine how outlet predictions shift when we add a smile to an image.

How Does the Model Account for Smiles?

Understanding how a multi-modal model learns and utilizes visual features is essential for interpreting predictions, especially in complex tasks like news outlet classification. However, this can be challenging since visual features are high-dimensional. In particular, it is important to examine whether the model is really utilizing information from facial features and emotions or whether it is simply utilizing the broader contextual information in the image for prediction (such as background color, text, etc.).

We employ a recently developed, popular tool -- Grad-CAM -- to visualize the contribution of specific visual features to the model's predictions by helping us understand the attention patterns of different components in the model selvaraju_etal_2017. Grad-CAM highlights the regions of an image that drive the decision-making process by examining the gradient of the model's prediction function \( g(\cdot) \) with respect to its inputs. This technique is particularly helpful in our architecture, where two different CNN branches—Face-VGG and ResNet101—process facial features and global image context, respectively. We can therefore use Grad-CAM to analyze attention patterns in both CNN branches separately.

figure[figure omitted — 451 chars of source]

As an illustrative example, consider the last image of Donald Trump in Figure (ref), from Washington Post. Figure (ref) shows visualizations for two versions of this image: the original image (top row) and a smile-added version (bottom row) generated using the transformation function \(\pi^1\). The Grad-CAM visualizations illustrate two distinct patterns of model attention. First, the modified ResNet101 heatmaps predominantly focus on broader contextual features, such as the lectern, the logo, and background elements, including the event name (“Operation Warp Speed"). The gradients of \(g(\cdot)\) with respect to the global image embeddings emphasize that ResNet101 captures cues related to the overall spatial and contextual layout of the image, such as positioning and visual saliency. Second, the modified Face-VGG heatmaps display heightened sensitivity to facial features, with particular emphasis on the mouth region. For the original image, the gradient of \(g(\cdot)\) with respect to the facial embeddings shows moderate activation around the mouth, corresponding to a neutral expression. When a smile is introduced, the activation in the mouth region intensifies and extends to other facial areas associated with smile dynamics. This pattern suggests that Face-VGG is able to successfully detect smile-related features and subtle changes in facial expressions, which provides strong evidence in support of including Face-VGG in our network architecture.

How Does Adding a Smile Change Predictions?

In the previous section, we saw that our outlet prediction model can capture the information in the smile/facial features of the focal politician. We now examine how adding a smile to a focal image changes news outlet predictions. As discussed in $\S$(ref), our identifying assumption requires that the dataset contains similar images (e.g., from the same/similar events, as in the case of Operation Warp Speed) from a diverse set of news outlets, where the main difference between the images is the facial expression of the focal politician. As such, our goal in this section is to demonstrate how our multi-modal ML model $g$ not only learns the structure of the comparable images for similar events but also captures the editorial preferences of different news outlets for smiles versus neutral expressions.

To do so, we turn to unsupervised learning techniques. First, we extract embeddings for all images using ResNet101. Each image is represented as a high-dimensional vector $\mathbf{e}_i \in \mathbb{R}^d$, where $d$ denotes the dimensionality of the embedding space. These embeddings capture general visual features of the images, such as texture, color, and background features, without being tied to specific news outlets. To make the embeddings interpretable and suitable for visualization, we reduce their dimensionality to two dimensions using Principal Component Analysis (PCA) and t-Distributed Stochastic Neighbor Embedding (t-SNE), followed by clustering to identify patterns. We present the details of PCA and t-SNE in our context in Web Appendix $\S$(ref).

figure[figure omitted — 656 chars of source]

Figure (ref) shows the t-SNE projection for all images of Donald Trump, grouped into 20 clusters based on their visual similarity. Each cluster captures images with shared contextual or visual features and provides a conceptualization of “similar events". For instance, consider Cluster 18, which is highlighted in Figure (ref). Here, each point represents an image, and the color indicates the conservative score of the news outlet that selected the image, ranging from liberal (blue) to conservative (red). The diversity in news outlets within the cluster suggests that multiple news outlets show images that are similar in structure/background (likely covering similar events). This variation is critical since it allows the outlet prediction model to capture how variation in facial features influences the likelihood of an image being chosen by a given outlet, conditional on the other image features (e.g., background, text, texture) being similar.

Figure (ref) illustrates examples from Cluster 18, showcasing images used by different news outlets, while Figure (ref) displays the corresponding images for the highlighted IDs. Within this cluster, consider the image corresponding to ID 196491 in our dataset by the Washington Post, which is the same image of Donald Trump used in the previous section. In this image, Trump appears unhappy, leaning slightly towards an angry expression. As shown in Figure (ref), this image is initially positioned on the top right-hand side of the plot. Next, we modify the original image (ID 196491) by applying the transformation function \(\pi^1\), which adds a smile to Donald Trump while preserving the contextual elements of the image. After computing the embedding for this modified image, its coordinates in the t-SNE space are updated, as illustrated by the gold dot in the scatter plot. Interestingly, the smile-added version of the original image shifts to the left. Now, the updated position results in a shorter distance to images such as ID 36529 from the Daily Mail and ID 72022 from Newsmax, compared to its distance to the original image ID 196491 by the Washington Post and ID 12736 by ABC News. This demonstrates how the addition of a smile alters the embedding position to align with images that share similar facial expressions, even though the overall context remains unchanged.

figure[figure omitted — 794 chars of source]

Although the t-SNE coordinates provide interpretable insights into how our model works, they do not capture the full complexity of the image features that our model $\hat{g}$ captures. Therefore, in Table (ref), we present our model's outlet predictions for the fourth image from Figure 4 (the one that was originally shown in Washington Post) for two cases -- the original image and smile-added image version. We see substantial differences in the outlet predictions between the original image and the counterfactual version with the smile. Notably, we find that adding a smile substantially increases the predicted probability for Daily Mail (from 2.89% to 20.97%) and Newsmax (from 7.10% to 16.52%), slightly decreases the predicted probability for ABS News (from 2.59% to 1.70%), and substantially decreases the predicted probability for Washington Post (from 30.98% to 3.32%).

table[table omitted — 693 chars of source]

In summary, we see that adding a smile systematically shifts the outlet prediction probabilities, consistent with the finding in the prior work in the advertising domain that shows facial expressions and emotions can shape audience engagement by capturing attention and influencing perception teixeira2012emotion. Specifically, we see an increase (decrease) in the predicted probabilities for right-leaning (left-leaning) news outlets when we add a smile to an image of Donald Trump. Nevertheless, this is just a single instance in the data that we use to illustrate the intuition behind how our algorithm works. In the next section, we present more systematic measures to quantify visual polarization.

Visual Slant and Polarization

We now present our main results on visual slant. All results are direct applications of Algorithm (ref) in $\S$(ref). Before presenting the results, we review a few important considerations in applying our algorithm. First, recall that our visual slant parameter requires a neutral outlet denoted by $y_n$ in Definition (ref). We use {\it Reuters} as the baseline outlet $y_n$ to measure visual slant, as {\it Reuters} is a largely neutral and fact-based outlet.\footnote{It is worth emphasizing that one could easily change the choice of base. Although the visual slant measure will change with a different choice, the overall visual polarization between two outlets remains unchanged.} Second, we use a train-test split, and all the polarization and slant measures are shown for the test data $\mathcal{D}_{\emph{test}}$ using the estimates obtained from the training data $\mathcal{D}_{\emph{train}}$. Third, for each politician, we start with three neutral images, generate counterfactual smiling versions of these three images, and use these six images in our polarization and slant measurement for each politician.

This section is organized as follows. In $\S$(ref), we present the overall distribution of visual slant across all outlets and politicians. We then explore the extent of heterogeneity across news outlets in $\S$(ref) to see which outlets exhibit higher levels of visual slant and validate our measure using external measures of media slant. Finally, in $\S$(ref), we document the heterogeneity in visual slant at the politician level.

Distribution of Visual Slant

An important feature of our algorithm is its ability to produce an individual-level measure of visual slant \(\hat{\rho}^T_{i}(p_i, y^{k}_{i}, y^{Reuters}_{i}) \). As such, for each article $i$ featuring politician $p$, we can estimate 19 visual slant measures corresponding to all the outlets other than Reuters. Based on the scores from faris2017partisanship, allsides2024, and flaxman2016filter, we categorize the three most Republican-leaning news outlets as: Fox News, Newsmax, and Daily Mail, and the three most Democratic-leaning news outlets as: Washington Post, CNN, and The New York Times. Figure (ref) shows the distributions of visual slant for Democratic and Republican politicians as featured in Democratic- and Republican-leaning news outlets. If there were no visual slant or polarization, all distributions would be tightly centered around zero. However, we observe that for both Republican- and Democratic-leaning outlets, the individual-level visual slant measures deviate from zero and exhibit significant dispersion. This indicates the presence of a visual slant in these outlets. Additionally, we notice a clear divide in the distributions of Democratic- and Republican-leaning outlets for both Republican and Democratic politicians, pointing to evidence of visual polarization.

figure[figure omitted — 665 chars of source]

Next, we examine if there is a systematic difference between Republican- and Democratic-learning outlets by examining their averages in each panel. We denote the set of Republican and Democratic politicians by $\mathcal{P}_R$ and $\mathcal{P}_D$, respectively. Similarly, we denote the sets of Republican- and Democratic-leaning outlets defined above are denoted by $\mathcal{Y}_R$ and $\mathcal{Y}_D$, respectively. For Republican politicians, in Figure (ref), the distributions of \( \hat{\rho}^T_{i} \) show distinct differences between left- and Republican-leaning news outlets. The visual slant measurement for left-leaning news outlets is -0.45, implying that, on average, smiling images of Republican politicians decrease the utility of a news outlet classified as Democratic-leaning compared to Reuters. In contrast, the visual slant measurement for Republican-leaning news outlets is 0.43, suggesting that, on average, smiling images of Republican politicians significantly increase their utility compared to Reuters. Therefore, for Republican politicians, we can conclude that:

equation[equation omitted — 165 chars of source]

For Democratic politicians, in Figure (ref), the visual slant distribution for Democratic and Republican news outlets also shows clear differences. The average visual slant for Democratic news outlets is 0.43, indicating that, on average, images of smiling Democratic politicians increase the utility of a Democratic-leaning outlet compared to Reuters (neutral baseline). Conversely, the average visual slant for Republican-leaning new outlets is -0.38, suggesting that, on average, smiling images of democratic politicians decrease their utility compared to Reuters. Therefore, for Democratic politicians, we can conclude that:

equation[equation omitted — 165 chars of source]

In Web Appendix $\S$(ref), we employ two statistical tests to analyze these differences: the Kolmogorov-Smirnov (K-S) test and one-sample t-tests. Both confirm that there are significant differences between the distributions and means of Republican- and Democratic-leaning outlets for both Democratic and Republican politicians.

Visual Slant Across News Outlets

We now document the extent to which visual slant varies across outlets. Figure (ref) shows our visual slant measure for each news outlet ($y^k \in \mathcal{Y}$) for Republican politicians as \( \hat{\rho}^T(p \in \mathcal{P}_R, y^{k}, y^{Reuters}) \)), and for Democratic politicians as \( \hat{\rho}^T(p \in \mathcal{P}_D, y^{k}, y^{Reuters}) \), ranked in bar charts. In Figure (ref), we see that Democratic news outlets have a positive visual slant measurement for Democratic politicians and a negative visual slant for Republican politicians. Conversely, in Figure (ref), we see that Republican news outlets exhibit a positive visual slant measurement for Republican politicians and a negative visual slant measurement for Democratic politicians.

figure[figure omitted — 671 chars of source]

Further, outlets such as Fox News and Daily Mail display a more positive visual slant measurement for Republican politicians, indicating a more favorable portrayal of these politicians. Simultaneously, these outlets show a more negative visual slant measure of Democratic politicians, displaying a less favorable portrayal of these politicians. On the other hand, outlets like CNN, Washington Post, and The New York Times exhibit a more positive visual slant measurement for Democratic politicians and a more negative visual slant for Republican politicians. Interestingly, some outlets, such as the Wall Street Journal and \textit{CBS News}, demonstrate a positive visual slant for both Democratic and Republican politicians. This suggests that their editorial approach may balance portrayals of politicians from both parties, reflecting a more centrist ideological positioning rather than strong partisan alignment.

To establish a single outlet-specific visual slant measurement, we compute the difference between the visual slant measures for Republican and Democratic politicians within each outlet. A larger difference indicates a stronger conservative visual slant. Accordingly, we define our unified measure as the Conservative Visual Slant (CVS) as follows:

equation[equation omitted — 207 chars of source]

Intuitively, this metric captures the degree to which an outlet portrays Republican politicians more favorably compared to Democratic politicians. For example, if a Republican-leaning outlet scores high on this measure, it implies that it emphasizes showing Republican politicians with a smile while depicting Democratic politicians without a smile. Figure (ref) illustrates the conservative visual slant scores for all outlets in our dataset, sorted in increasing order. The results reveal that outlets such as Daily Mail and Fox News demonstrate high conservative visual slant, strongly favoring Republican politicians. On the other hand, outlets like Washington Post, USA Today, and NY Times exhibit a pronounced liberal slant, favoring Democratic politicians. Meanwhile, outlets such as Wall Street Journal, \textit{BBC News}, and \textit{CBS News} exhibit relatively neutral or low levels of slant. Overall, there is significant heterogeneity in visual slant across outlets, providing insights into which outlets are the most and least polarized visually.

figure[figure omitted — 237 chars of source]

Lastly, we perform two validation tests to show how our proposed {\it CVS} measure relates to the existing measures of media slant, and how it performs relative to benchmarks that measure visual slant. We present a brief summary of our validation exercise below and refer readers to Web Appendix $\S$(ref) for a comprehensive analysis of outlet-level visual slant, including individual histograms, statistical summaries, and additional tests to quantify bias and polarization across news outlets.

list{$\bullet$} { {0pt} {1pt} {1pt} {1pt} {1.5em} {1em} {0.5em} } • Consistency with external measures of media slant: We begin by assessing the correlation between our {\it CVS} measure and existing media slant measures derived from independent sources not used in our algorithm. Specifically, we consider three existing measures from faris2017partisanship, flaxman2016filter, and allsides2024. For instance, faris2017partisanship quantifies media slant based on the proportion of a media outlet's stories shared on Twitter by users who predominantly retweet conservative-leaning sources. Since our algorithm does not incorporate the data used to construct these external slant measures, a positive correlation between our {\it CVS} measure and these benchmarks would provide validation for our approach. For the 20 outlets in our dataset, we find significant and positive correlations of 0.79, 0.55, and 0.81 with the measures from faris2017partisanship, flaxman2016filter, and allsides2024, respectively. We visualize these correlations and conduct formal statistical tests in Web Appendix $\S$(ref) to further support this validation. • Comparison to existing measures of visual slant: Second, we compare our {\it CVS} measure with existing outlet-specific visual slant measures from prior literature. Specifically, we examine the visual slant measure proposed by boxell2021slanted, which employs a reduced-form approach similar to that discussed in $\S$(ref). Our objective is to determine which measure better aligns with existing polarization benchmarks discussed earlier. In the main text, we focus on the conservative share score by faris2017partisanship, a widely recognized benchmark in media slant research, including in boxell2021slanted. To quantify the alignment between these visual slant measures and the conservative share score from faris2017partisanship, we conduct both Pearson and Spearman correlation analyses, assessing how well each measure captures the overall slant of news outlets. Given that our dataset and boxell2021slanted share 11 common outlets, we focus on this subset for direct comparison. \begin{figure}[htp!] \begin{subfigure}[t]{0.50\textwidth} \caption{ {\it Conservative Visual Slant (CVS)} from PMCIG} \end{subfigure} \begin{subfigure}[t]{0.50\textwidth} \caption{ Visual Slant Measure from boxell2021slanted.} \end{subfigure} \caption{ Comparison of {\it CVS} and visual slant measure by boxell2021slanted against Conservative Share Score from faris2017partisanship.} \end{figure} Figure (ref) presents a comparative analysis of visual slant measures across news outlets. In both plots, the x-axis represents each outlet’s conservative share score from faris2017partisanship, while the y-axis represents either our {\it CVS} metric (Figure (ref)) or the visual slant measure from boxell2021slanted (Figure (ref)). Each point corresponds to a news outlet, with a red trendline illustrating the relationship between visual slant and conservative share scores. As shown in these figures, our {\it CVS} measure exhibits a statistically significant and strong Pearson and Spearman correlation with the conservative share score from faris2017partisanship, whereas the visual slant measure from boxell2021slanted shows only a weak, statistically insignificant correlation. This finding further validates our method, demonstrating that our {\it CVS} measure more accurately captures media slant compared to prior visual slant metrics. Additionally, in Web Appendix $\S$(ref), we replicate the analysis presented in this part using two other widely recognized conservative share scores and demonstrate the robustness of our findings.

In summary, these results demonstrate the validity of our approach in comparison to existing benchmark measures of media bias.

Visual Polarization Across Politicians

figure[figure omitted — 247 chars of source]

We now examine how different politicians are portrayed across news outlets and identify the politicians with the most polarizing depictions in media. Figure (ref) presents the visual slant measures for two prominent figures: Barack Obama and Donald Trump. Consistent with our earlier findings, we observe a clear divide in visual slant scores -- Obama receives more favorable portrayals in Democratic-leaning outlets, while Trump is depicted more positively in Republican-leaning outlets. We extend this analysis to other politicians, with detailed results provided in Web Appendix $\S$(ref).

Next, we develop a measure for the overall extent of visual polarization for each politician. Intuitively, we expect to observe a greater extent of variability in visual slant measures across outlets for a more polarizing politician. As such, we quantify the {\it Overall Visual Polarization (OVP)} for each politician, which is measured by the standard deviation of the polarization across all news outlets as:

equation[equation omitted — 294 chars of source]
figure[figure omitted — 327 chars of source]

Figure (ref) ranks the politicians from most polarized to least polarized based on this criterion. We see that Donald Trump and Barack Obama are the top two politicians with the most visually polarizing portrayal across all outlets. Given that both were the president/presidential candidate for significant chunks of our observation period and at the forefront of multiple polarizing discussions and events, this is understandable. Further, Rand Paul and Ted Cruz, both prominent Republican politicians, also show high levels of polarization, likely attributed to their roles in policy debates and media prominence during key political events. On the Democratic side, figures such as Bernie Sanders, Hillary Clinton, and Kamala Harris display significant polarization, reflecting their leadership roles and ideological positions within the party. Further, on the right end of Figure (ref), we observe politicians with lower {\it OVP} scores. Notably, Susan Collins, Liz Cheney, and Joe Manchin rank among the least polarized, aligning with their reputations for bipartisanship and moderation cnn_2018. These findings have important implications for political strategy, particularly in how politicians shape their election campaigns when deciding whether to mobilize their base or appeal to a broader, centrist electorate.

Finally, we seek to validate our politician-specific {\it OVP} measure by comparing it with external indicators of a politician's level of polarization. However, there exists no widely accepted, politician-specific polarization metric. To overcome this issue, we propose that a politician’s ideological alignment with their primary constituency serves as a meaningful proxy, since more polarizing politicians are likely to perform better in ideologically aligned constituencies and struggle in misaligned ones. To quantify this alignment, we use the politician’s party's success in the 2016 election within their state, by measuring their party's percentage point advantage in that election. We then analyze the correlation between this measure and our {\it OVP} metric. Figure (ref) presents the results from this exercise. We see a strong and significant correlation between the two measures, which further supports the validity of our {\it OVP} measure.

figure[figure omitted — 281 chars of source]

This finding supports the idea that politicians from states with strong partisan leanings tend to adopt more pronounced partisan positions without seeking to appeal to a broad electorate. For example, Bernie Sanders (Vermont) and Ted Cruz (Texas) come from states with clear ideological identities and exhibit high {\it OVP} values, likely reflecting their strong partisan stances and the resulting polarized media portrayals. In contrast, Joe Manchin (West Virginia) and Susan Collins (Maine), who represent states where their party is in the minority, have lower {\it OVP} values, which reflects their moderate positions.

Conclusion

In this paper, we present a framework for measuring slant and polarization in the visual content accompanying news articles. We propose the Polarization Measurement Using Counterfactual Image Generation (PMCIG) algorithm, which quantifies news outlets' preference for slanted imagery -- such as smiling images -- to convey positive (vs. negative) representation of politicians. Our framework combines the economic structure of the problem with generative models to generate comparable counterfactual images and measure polarization in visual content. Notably, our algorithm overcomes the key limitations in traditional descriptive methods due to information loss in the feature extraction phase by using the rich information contained in images.

In our empirical analysis, we apply the PMCIG framework to a decade-long dataset that covers 20 major news outlets and 30 prominent politicians. We identify clear patterns of ideological slanting and political polarization in the visual representation of political figures. We validate our measure of visual slant by demonstrating a high correlation between our measure and the existing measures used for media slant and partisanship. Our framework measures visual slant and polarization with detailed granularity, highlighting differences both at the outlet level and for individual politicians. Among outlets, we find that {\it Daily Mail} and {\it Fox News} display the strongest Republican-leaning visual slant, while {\it Washington Post} and {\it The New York Times} exhibit the strongest Democratic-leaning slant. In contrast, {\it CBS News} and {\it Wall Street Journal} are among the outlets with the lowest overall visual slant. At the individual level, Donald Trump and Barack Obama stand out as the most polarizing figures, whereas Joe Manchin, Liz Cheney, and Susan Collins are among the least polarizing in their visual portrayal across news outlets.

In summary, the PMCIG framework offers a systematic approach to analyze how ideological preferences shape visual content in news media and contribute to polarization. Nevertheless, our paper has limitations that serve as excellent avenues for future research. For instance, our analysis is based on data from the United States, and extending the framework to other regions, such as Europe, could uncover cross-cultural differences in visual polarization. Additionally, while the framework measures visual slant and polarization, it does not examine the downstream impact on individuals' beliefs or behavior, which could be fruitful avenues for future research. Future studies could also apply the PMCIG framework to other forms of media, such as social media or advertising, and investigate the extent of ideological slanting/bias in these settings.

Competing Interests Declaration

Author(s) have no competing interests to declare.

thebibliography{77} \expandafter\ifx\csname urlstyle\endcsname\relax \else \fi \bibitem[AllSides(2024)]{allsides2024} AllSides. \newblock Media bias rating methods, 2024. \newblock URL https://www.allsides.com/media-bias/media-bias-rating-methods. \bibitem[Amaldoss et al.(2021)Amaldoss, Du, and Shin]{amaldoss2021media} Wilfred Amaldoss, Jinzhao Du, and Woochoel Shin. \newblock Media platforms’ content provision strategies and sources of profits. \newblock Marketing Science, 40\penalty0 (3):\penalty0 527--547, 2021. \bibitem[Anand and Kadiyali(2024)]{anand2024frontiers} Piyush Anand and Vrinda Kadiyali. \newblock Frontiers: Smoke and mirrors: Impact of e-cigarette taxes on underage social media posting. \newblock Marketing Science, 2024. \bibitem[Angrist and Pischke(2009)]{angrist2009mostly} Joshua D. Angrist and Jörn-Steffen Pischke. \newblock Mostly Harmless Econometrics: An Empiricist's Companion. \newblock Princeton University Press, 2009. \bibitem[Ash et al.(2021)Ash, Durante, Grebenschikova, and Schwarz]{ash2021visual} Elliott Ash, Ruben Durante, Maria Grebenschikova, and Carlo Schwarz. \newblock Visual representation and stereotypes in news media. \newblock 2021. \bibitem[Athey et al.(2022)Athey, Karlan, Palikot, and Yuan]{athey2022smiles} Susan Athey, Dean Karlan, Emil Palikot, and Yuan Yuan. \newblock Smiles in profiles: Improving fairness and efficiency using estimates of user preferences in online marketplaces. \newblock Technical report, National Bureau of Economic Research, 2022. \bibitem[Azure(2023)]{microsoft_azure_vision} Microsoft Azure. \newblock Face api - cognitive services, 2023. \newblock URL https://azure.microsoft.com/en-us/services/cognitive-services/face/. \newblock Accessed on September, 2023. \bibitem[Belkina et al.(2019)Belkina, Ciccolella, Anno, Halpert, Spidlen, and Snyder-Cappione]{belkina2019automated} Anna C Belkina, Christopher O Ciccolella, Rina Anno, Richard Halpert, Josef Spidlen, and Jennifer E Snyder-Cappione. \newblock Automated optimized parameters for t-distributed stochastic neighbor embedding improve visualization and analysis of large datasets. \newblock Nature communications, 10\penalty0 (1):\penalty0 5415, 2019. \bibitem[Bird et al.(2009)Bird, Klein, and Loper]{bird2009natural} Steven Bird, Ewan Klein, and Edward Loper. \newblock \emph{Natural Language Processing with Python: Analyzing Text with the Natural Language Toolkit}. \newblock " O'Reilly Media, Inc.", 2009. \bibitem[Blanchard et al.(2023)Blanchard, Noseworthy, Pancer, and Poole]{blanchard2023extraction} Simon J Blanchard, Theodore J Noseworthy, Ethan Pancer, and Maxwell Poole. \newblock Extraction of visual information to predict crowdfunding success. \newblock \emph{Production and Operations Management}, 32\penalty0 (12):\penalty0 4172--4189, 2023. \bibitem[Blei et al.(2003)Blei, Ng, and Jordan]{blei2003latent} David M Blei, Andrew Y Ng, and Michael I Jordan. \newblock Latent dirichlet allocation. \newblock \emph{Journal of Machine Learning Research}, 3\penalty0 (Jan):\penalty0 993--1022, 2003. \bibitem[Bollen and Davis(2009)]{bollen2009causal} Kenneth A Bollen and Walter R Davis. \newblock Causal indicator models: Identification, estimation, and testing. \newblock \emph{Structural Equation Modeling: A Multidisciplinary Journal}, 16\penalty0 (3):\penalty0 498--522, 2009. \bibitem[Bondi et al.(2023)Bondi, Rafieian, and Yao]{bondi2023privacy} Tommaso Bondi, Omid Rafieian, and Yunfei Jesse Yao. \newblock Privacy and polarization: An inference-based framework. \newblock \emph{Available at SSRN 4641822}, 2023. \bibitem[Boxell(2021)]{boxell2021slanted} Levi Boxell. \newblock Slanted images: Measuring nonverbal media bias during the 2016 election. \newblock \emph{Available at SSRN 3837521}, 2021. \bibitem[Caprini(2023)]{caprini2023visual} Giulia Caprini. \newblock Visual bias. \newblock 2023. \bibitem[Chandon et al.(2009)Chandon, Hutchinson, Bradlow, and Young]{chandon2009does} Pierre Chandon, J Wesley Hutchinson, Eric T Bradlow, and Scott H Young. \newblock Does in-store marketing work? effects of the number and position of shelf facings on brand attention and evaluation at the point of purchase. \newblock \emph{Journal of marketing}, 73\penalty0 (6):\penalty0 1--17, 2009. \bibitem[Chernozhukov et al.(2018)Chernozhukov, Chetverikov, Demirer, Duflo, Hansen, Newey, and Robins]{chernozhukov2018double} Victor Chernozhukov, Denis Chetverikov, Mert Demirer, Esther Duflo, Christian Hansen, Whitney Newey, and James Robins. \newblock Double/debiased machine learning for treatment and structural parameters, 2018. \bibitem[Davenport(2017)]{davenport2017analytics} Thomas H Davenport. \newblock How analytics has changed in the last 10 years (and how it’s stayed the same). \newblock \emph{Harvard Business Review}, 28\penalty0 (08):\penalty0 2017, 2017. \bibitem[Deng et al.(2019)Deng, Guo, Xue, and Zafeiriou]{deng2019arcface} Jiankang Deng, Jia Guo, Niannan Xue, and Stefanos Zafeiriou. \newblock Arcface: Additive angular margin loss for deep face recognition. \newblock In \emph{Proceedings of the IEEE/CVF conference on computer vision and pattern recognition}, pages 4690--4699, 2019. \bibitem[Devlin et al.(2018)Devlin, Chang, Lee, and Toutanova]{devlin2018bert} Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. \newblock Bert: Pre-training of deep bidirectional transformers for language understanding. \newblock \emph{arXiv preprint arXiv:1810.04805}, 2018. \bibitem[Dew et al.(2022)Dew, Ansari, and Toubia]{dew2022letting} Ryan Dew, Asim Ansari, and Olivier Toubia. \newblock Letting logos speak: Leveraging multiview representation learning for data-driven branding and logo design. \newblock \emph{Marketing Science}, 41\penalty0 (2):\penalty0 401--425, 2022. \bibitem[Doherty et al.(2023)Doherty, Kiley, Asheer, and Price]{pew2023} Carroll Doherty, Jocelyn Kiley, Nida Asheer, and Talie Price. \newblock Americans’ feelings about politics, polarization and the tone of political discourse, 2023. \newblock URL \texttt{https://www.pewresearch.org/wp-content/uploads/sites/20/2023/09/PP_2023.09.19_views-of-politics_REPORT.pdf}. \bibitem[Dvorkin et al.(2021)Dvorkin, S{\'a}nchez, Sapriza, and Yurdagul]{dvorkin2021sovereign} Maximiliano Dvorkin, Juan M S{\'a}nchez, Horacio Sapriza, and Emircan Yurdagul. \newblock Sovereign debt restructurings. \newblock \emph{American Economic Journal: Macroeconomics}, 13\penalty0 (2):\penalty0 26--77, 2021. \bibitem[Dzyabura et al.(2023)Dzyabura, El Kihal, Hauser, and Ibragimov]{dzyabura2023leveraging} Daria Dzyabura, Siham El Kihal, John R Hauser, and Marat Ibragimov. \newblock Leveraging the power of images in managing product return rates. \newblock \emph{Marketing Science}, 42\penalty0 (6):\penalty0 1125--1142, 2023. \bibitem[Face++(2023)]{faceplusplus_emotion} Face++. \newblock Emotion recognition, 2023. \newblock URL \texttt{https://www.faceplusplus.com/emotion-recognition/}. \newblock Accessed on September, 2023. \bibitem[Faris et al.(2017)Faris, Roberts, Etling, Bourassa, Zuckerman, and Benkler]{faris2017partisanship} Robert Faris, Hal Roberts, Bruce Etling, Nikki Bourassa, Ethan Zuckerman, and Yochai Benkler. \newblock Partisanship, propaganda, and disinformation: Online media and the 2016 us presidential election. \newblock \emph{Berkman Klein Center Research Publication}, 6, 2017. \bibitem[Flaxman et al.(2016)Flaxman, Goel, and Rao]{flaxman2016filter} Seth Flaxman, Sharad Goel, and Justin M Rao. \newblock Filter bubbles, echo chambers, and online news consumption. \newblock \emph{Public opinion quarterly}, 80\penalty0 (S1):\penalty0 298--320, 2016. \bibitem[Fu et al.(2019)Fu, Liu, Tian, Li, Bao, Fang, and Lu]{fu2019dual} Jun Fu, Jing Liu, Haijie Tian, Yong Li, Yongjun Bao, Zhiwei Fang, and Hanqing Lu. \newblock Dual attention network for scene segmentation. \newblock In \emph{Proceedings of the IEEE/CVF conference on computer vision and pattern recognition}, pages 3146--3154, 2019. \bibitem[Gentzkow and Shapiro(2006)]{gentzkow2006media} Matthew Gentzkow and Jesse M Shapiro. \newblock Media bias and reputation. \newblock \emph{Journal of political Economy}, 114\penalty0 (2):\penalty0 280--316, 2006. \bibitem[Gentzkow and Shapiro(2010)]{gentzkow2010drives} Matthew Gentzkow and Jesse M Shapiro. \newblock What drives media slant? evidence from us daily newspapers. \newblock \emph{Econometrica}, 78\penalty0 (1):\penalty0 35--71, 2010. \bibitem[Gentzkow et al.(2019)Gentzkow, Shapiro, and Taddy]{gentzkow2019measuring} Matthew Gentzkow, Jesse M Shapiro, and Matt Taddy. \newblock Measuring group differences in high-dimensional choices: method and application to congressional speech. \newblock \emph{Econometrica}, 87\penalty0 (4):\penalty0 1307--1340, 2019. \bibitem[Groseclose and Milyo(2005)]{groseclose2005measure} Tim Groseclose and Jeffrey Milyo. \newblock A measure of media bias. \newblock \emph{The quarterly journal of economics}, 120\penalty0 (4):\penalty0 1191--1237, 2005. \bibitem[Hansen(2022)]{hansen2022econometrics} Bruce Hansen. \newblock \emph{Econometrics}. \newblock Princeton University Press, 2022. \bibitem[He et al.(2016)He, Zhang, Ren, and Sun]{he2016deep} Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. \newblock Deep residual learning for image recognition. \newblock \emph{Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition}, pages 770--778, 2016. \bibitem[He et al.(2017)He, Gkioxari, Doll{\'a}r, and Girshick]{he2017mask} Kaiming He, Georgia Gkioxari, Piotr Doll{\'a}r, and Ross Girshick. \newblock Mask r-cnn. \newblock In \emph{Proceedings of the IEEE international conference on computer vision}, pages 2961--2969, 2017. \bibitem[Hoegg and Lewis(2011)]{hoegg2011impact} Joandrea Hoegg and Michael V Lewis. \newblock The impact of candidate appearance and advertising strategies on election results. \newblock \emph{Journal of Marketing Research}, 48\penalty0 (5):\penalty0 895--909, 2011. \bibitem[Imbens and Rubin(2015)]{imbens_rubin_2015} Guido W Imbens and Donald B Rubin. \newblock \emph{Causal inference in statistics, social, and biomedical sciences}. \newblock Cambridge university press, 2015. \bibitem[Iyer and Yoganarasimhan(2021)]{iyer_yoganarasimhan_2021} Ganesh Iyer and Hema Yoganarasimhan. \newblock Strategic polarization in group interactions. \newblock \emph{Journal of Marketing Research}, 58\penalty0 (4):\penalty0 782--800, 2021. \bibitem[Jakesch et al.(2022)Jakesch, Naaman, and Michael]{jakesch2022belief} Maurice Jakesch, Mor Naaman, and MACY Michael. \newblock Belief in partisan news depends on favorable content more than on a trusted source. \newblock 2022. \bibitem[Jensen et al.(2012)Jensen, Naidu, Kaplan, Wilse-Samson, Gergen, Zuckerman, and Spirling]{jensen2012political} Jacob Jensen, Suresh Naidu, Ethan Kaplan, Laurence Wilse-Samson, David Gergen, Michael Zuckerman, and Arthur Spirling. \newblock Political polarization and the dynamics of political language: Evidence from 130 years of partisan speech [with comments and discussion]. \newblock \emph{Brookings Papers on Economic Activity}, pages 1--81, 2012. \bibitem[Killough(2018)]{cnn_2018} Ashley Killough. \newblock Moderate senators feel boost as shutdown ends. \newblock \emph{CNN}, 2018. \newblock URL \texttt{https://www.cnn.com/2018/01/22/politics/moderate-senators-shutdown-end-bipartisan-group-senate/index.html}. \bibitem[Li et al.(2024)Li, Ni, and Yang]{li2024product} Hui Li, Jian Ni, and Fangzhu Yang. \newblock Product design using generative adversarial network: Incorporating consumer preference and external data. \newblock \emph{arXiv preprint arXiv:2405.15929}, 2024. \bibitem[Liang et al.(2024)Liang, Zadeh, and Morency]{liang2024foundations} Paul Pu Liang, Amir Zadeh, and Louis-Philippe Morency. \newblock Foundations & trends in multimodal machine learning: Principles, challenges, and open questions. \newblock \emph{ACM Computing Surveys}, 56\penalty0 (10):\penalty0 1--42, 2024. \bibitem[Lin et al.(2017)Lin, Doll{\'a}r, Girshick, He, Hariharan, and Belongie]{lin2017feature} Tsung-Yi Lin, Piotr Doll{\'a}r, Ross Girshick, Kaiming He, Bharath Hariharan, and Serge Belongie. \newblock Feature pyramid networks for object detection. \newblock In \emph{Proceedings of the IEEE conference on computer vision and pattern recognition}, pages 2117--2125, 2017. \bibitem[Liu et al.(2020)Liu, Dzyabura, and Mizik]{liu2020visual} Liu Liu, Daria Dzyabura, and Natalie Mizik. \newblock Visual listening in: Extracting brand image portrayed on social media. \newblock \emph{Marketing Science}, 39\penalty0 (4):\penalty0 669--686, 2020. \bibitem[Loshchilov and Hutter(2017)]{loshchilov2017decoupled} Ilya Loshchilov and Frank Hutter. \newblock Decoupled weight decay regularization. \newblock \emph{arXiv preprint arXiv:1711.05101}, 2017. \bibitem[Lu et al.(2021)Lu, Yao, Chen, and Grewal]{lu2021larger} Shijie Lu, Dai Yao, Xingyu Chen, and Rajdeep Grewal. \newblock Do larger audiences generate greater revenues under pay what you want? evidence from a live streaming platform. \newblock \emph{Marketing Science}, 40\penalty0 (5):\penalty0 964--984, 2021. \bibitem[Luca et al.(2022)Luca, Pronkina, and Rossi]{luca2022scapegoating} Michael Luca, Elizaveta Pronkina, and Michelangelo Rossi. \newblock Scapegoating and discrimination in times of crisis: Evidence from airbnb. \newblock Technical report, National Bureau of Economic Research, 2022. \bibitem[Ludwig and Mullainathan(2024)]{ludwig2024machine} Jens Ludwig and Sendhil Mullainathan. \newblock Machine learning as a tool for hypothesis generation. \newblock \emph{The Quarterly Journal of Economics}, 139\penalty0 (2):\penalty0 751--827, 2024. \bibitem[Luo and Toubia(2024)]{luo2024using} Lan E Luo and Olivier Toubia. \newblock Using ai for controllable stimuli generation: An application to gender discrimination with faces. \newblock \emph{Available at SSRN 4865798}, 2024. \bibitem[Matatov et al.(2022)Matatov, Naaman, and Amir]{matatov2022stop} Hana Matatov, Mor Naaman, and Ofra Amir. \newblock Stop the [image] steal: The role and dynamics of visual content in the 2020 us election misinformation campaign. \newblock \emph{Proceedings of the ACM on Human-Computer Interaction}, 6\penalty0 (CSCW2):\penalty0 1--24, 2022. \bibitem[Matias et al.(2021)Matias, Munger, Quere, and Ebersole]{MatiasEtAl2021} J. Nathan Matias, Kevin Munger, Marianne Aubin Le Quere, and Charles Ebersole. \newblock The upworthy research archive, a time series of 32,487 experiments in u.s. media. \newblock \emph{Scientific Data}, 8\penalty0 (195), 2021. \newblock doi: \begingroup \urlstyle{rm}\Url{10.1038/s41597-021-00934-7}. \bibitem[Microsoft(2022)]{microsoft2023responsible} Microsoft. \newblock Responsible ai investments and safeguards for facial recognition, 2022. \newblock URL \texttt{https://azure.microsoft.com/en-us/blog/responsible-ai-investments-and-safeguards-for-facial-recognition}. \bibitem[Mirza and Osindero(2014)]{mirza2014conditional} Mehdi Mirza and Simon Osindero. \newblock Conditional generative adversarial nets. \newblock \emph{arXiv preprint arXiv:1411.1784}, 2014. \bibitem[Parkhi et al.(2015)Parkhi, Vedaldi, and Zisserman]{parkhi2015deep} Omkar Parkhi, Andrea Vedaldi, and Andrew Zisserman. \newblock Deep face recognition. \newblock In \emph{BMVC 2015-Proceedings of the British Machine Vision Conference 2015}. British Machine Vision Association, 2015. \bibitem[Peng(2018)]{peng2018same} Yilang Peng. \newblock Same candidates, different faces: Uncovering media bias in visual portrayals of presidential candidates with computer vision. \newblock \emph{Journal of Communication}, 68\penalty0 (5):\penalty0 920--941, 2018. \bibitem[{\v R}eh{\r u}{\v r}ek and Sojka(2010)]{rehurek_lrec} Radim {\v R}eh{\r u}{\v r}ek and Petr Sojka. \newblock Software framework for topic modelling with large corpora. \newblock In \emph{Proceedings of the LREC 2010 Workshop on New Challenges for NLP Frameworks}, pages 45--50, Valletta, Malta, May 2010. ELRA. \newblock \texttt{http://is.muni.cz/publication/884893/en}. \bibitem[Selvaraju et al.(2017)Selvaraju, Cogswell, Das, Vedantam, Parikh, and Batra]{selvaraju_etal_2017} Ramprasaath R Selvaraju, Michael Cogswell, Abhishek Das, Ramakrishna Vedantam, Devi Parikh, and Dhruv Batra. \newblock Grad-cam: Visual explanations from deep networks via gradient-based localization. \newblock In \emph{Proceedings of the IEEE international conference on computer vision}, pages 618--626, 2017. \bibitem[SerpAPI(2023)]{serpapi} SerpAPI. \newblock Serpapi - real-time search engine results api, 2023. \newblock URL \texttt{https://serpapi.com/}. \newblock Accessed on January, 2023. \bibitem[Singh and Zheng(2023)]{singh2023causal} Amandeep Singh and Bolong Zheng. \newblock Causal regressions for unstructured data. \newblock In \emph{Causal Representation Learning Workshop at NeurIPS 2023}, 2023. \bibitem[Skelley and Fuong(2022)]{skelley_fuong_2022} Geoffrey Skelley and Holly Fuong. \newblock 3 in 10 americans named political polarization as a top issue facing the country, 2022. \newblock URL \texttt{https://fivethirtyeight.com/features/3-in-10-americans-named-political-polarization-as-a-top-issue-facing-the-country/}. \bibitem[S{\"u}lflow and Maurer(2019)]{sulflow2019power} Michael S{\"u}lflow and Marcus Maurer. \newblock The power of smiling. how politicians’ displays of happiness affect viewers’ gaze behavior and political judgments. \newblock \emph{Visual political communication}, pages 207--224, 2019. \bibitem[Sullivan and Masters(1988)]{sullivan1988happy} Denis G Sullivan and Roger D Masters. \newblock " happy warriors": Leaders' facial displays, viewers' emotions, and political support. \newblock \emph{American Journal of Political Science}, pages 345--368, 1988. \bibitem[Szegedy et al.(2016)Szegedy, Vanhoucke, Ioffe, Shlens, and Wojna]{szegedy2016rethinking} Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jon Shlens, and Zbigniew Wojna. \newblock Rethinking the inception architecture for computer vision. \newblock In \emph{Proceedings of the IEEE conference on computer vision and pattern recognition}, pages 2818--2826, 2016. \bibitem[Taigman et al.(2014)Taigman, Yang, Ranzato, and Wolf]{taigman2014deepface} Yaniv Taigman, Ming Yang, Marc'Aurelio Ranzato, and Lior Wolf. \newblock Deepface: Closing the gap to human-level performance in face verification. \newblock In \emph{Proceedings of the IEEE conference on computer vision and pattern recognition}, pages 1701--1708, 2014. \bibitem[Teixeira et al.(2012)Teixeira, Wedel, and Pieters]{teixeira2012emotion} Thales Teixeira, Michel Wedel, and Rik Pieters. \newblock Emotion-induced engagement in internet video advertisements. \newblock \emph{Journal of marketing research}, 49\penalty0 (2):\penalty0 144--159, 2012. \bibitem[Townsend and Kahn(2014)]{TownsendKahn2014} Claudia Townsend and Barbara E. Kahn. \newblock The visual preference heuristic: Effects of visual vs. verbal depiction on assessment of others, choice of hedonic products, and voting. \newblock \emph{Journal of Consumer Research}, 41\penalty0 (2):\penalty0 392--411, 2014. \bibitem[Train(2009)]{train_2009} Kenneth E Train. \newblock \emph{Discrete choice methods with simulation}. \newblock Cambridge university press, 2009. \bibitem[Troncoso and Luo(2022)]{troncoso2022look} Isamar Troncoso and Lan Luo. \newblock Look the part? the role of profile pictures in online labor markets. \newblock \emph{Marketing Science}, 2022. \bibitem[Vaswani(2017)]{vaswani2017attention} A Vaswani. \newblock Attention is all you need. \newblock \emph{Advances in Neural Information Processing Systems}, 2017. \bibitem[Vision(2023)]{google_ml_kit} Google Vision. \newblock Ml kit: Face detection api, 2023. \newblock URL \texttt{https://developers.google.com/ml-kit/vision/face-detection}. \newblock Accessed on September, 2023. \bibitem[Wedel and Pieters(2007)]{wedel_pieters_2007} Michel Wedel and Rik Pieters. \newblock \emph{Visual marketing: From attention to action}. \newblock Psychology Press, 2007. \bibitem[Wedel et al.(2023)Wedel, Pieters, and van der Lans]{wedel2023modeling} Michel Wedel, Rik Pieters, and Ralf van der Lans. \newblock Modeling eye movements during decision making: A review. \newblock \emph{psychometrika}, 88\penalty0 (2):\penalty0 697--729, 2023. \bibitem[Wei and Malik(2022)]{wei2022unstructured} Yanhao Wei and Nikhil Malik. \newblock Unstructured data, econometric models, and estimation bias. \newblock SSRN, 2022. \bibitem[Wooldridge(2010)]{wooldridge2010econometrics} Jeffrey M. Wooldridge. \newblock \emph{Econometric Analysis of Cross Section and Panel Data}. \newblock MIT Press, Cambridge, MA, 2nd edition, 2010. \bibitem[Xu et al.(2024)Xu, Zhang, Jiang, and Qi]{xu2024unstructured} Sikun Xu, Dennis J. Zhang, Zhenling Jiang, and Zhengling Qi. \newblock Causal inference when controlling for unstructured data. \newblock \emph{Working Paper}, 2024. \newblock Olin Business School, Washington University in St. Louis. Accessed: 2025-01-21. \bibitem[Zhang et al.(2016)Zhang, Zhang, Li, and Qiao]{zhang2016joint} Kaipeng Zhang, Zhanpeng Zhang, Zhifeng Li, and Yu Qiao. \newblock Joint face detection and alignment using multitask cascaded convolutional networks. \newblock \emph{IEEE signal processing letters}, 23\penalty0 (10):\penalty0 1499--1503, 2016.

\setcounter{table}{0} \setcounter{page}{0} \setcounter{figure}{0}