EconBase
← Back to paper

Generalization Issues in Conjoint Experiment: Attention and Salience

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

63,062 characters · 10 sections · 50 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

A Formal Theory of Survey Experiment Generalizability: Attention and Salience

abstractSurvey experiments are widely used to identify causal effects in political science and the social sciences. Yet researchers are typically interested in more than the internal validity of an experimentally induced contrast. They also want to know whether the estimated effect corresponds to the effect in the real world. We develop a formal theory of survey experiment generalizability grounded in behavioral microfoundations. The theory highlights two mechanisms. First, the survey environment shapes attention: it determines which considerations enter the respondent's active consideration set. Second, it shapes salience: conditional on consideration, it influences the relative weight assigned to those considerations. This framework yields two main results. Consideration-set compression generates amplification: survey-experimental effects can be larger in magnitude than their real-world counterparts, even for the same individuals, treatment content, and outcome. Context-dependent salience generates sign instability: the direction of the survey effect need not coincide with the direction of the corresponding real-world effect. The theory clarifies what survey experiments identify, when those effects are likely to generalize, and how survey designs can be modified to improve decision-environment transportability. \noindentKeywords: Survey Experiment, External Validity, Generalizability, Attention, Salience.

\thispagestyle{empty} \doublespacing

\doparttoc \faketableofcontents

\setcounter{page}{1}

Introduction

Survey experiments are a core empirical tool in political science and the broader social sciences. Researchers use them to study candidate choice, policy support, accountability, punishment, among other outcomes eggers2018corruption,carnes2016voters,sanbonmatsu2002gender. Because treatment is randomized, survey experiments are widely regarded as providing credible causal evidence. Between 2021 and 2025, a total of 204 survey experiments were published in leading political science journals, as shown in Table (ref). Following the classification proposed by EGAP, we categorize these studies into five types: conjoint, endorsement, list, priming, and vignette/factorial.\footnote{https://egap.org/resource/10-things-to-know-about-survey-experiments/} Figure (ref) presents the distribution of these categories over time. Among them, vignette/factorial and conjoint designs are the two dominant approaches.

However, in most substantive applications, internal validity is not the end of the story. Researchers are often interested in a stronger claim: how the same information, attribute, or intervention would operate in real-world environments beyond the survey setting. For example, a candidate trait that shifts support in a conjoint experiment is often interpreted as evidence of how voters would weigh that trait in actual elections. Similarly, a corruption message that affects respondents’ evaluations in a vignette is often taken to reflect how such information would influence real-world political judgments. Survey experiments are often viewed as a powerful tool for learning about real-world effects gaines2007logic. Yet, under what conditions is such generalizability warranted? More fundamentally, what mechanisms determine whether findings from survey experiments successfully translate to real-world behavior, and when might they fail to do so?

figure[figure omitted — 187 chars of source]
table[table omitted — 732 chars of source]

In this study, we develop a formal framework that integrates a behavioral model with the potential outcomes framework to characterize two mechanisms that shape the generalizability of survey experiments. In the existing literature, generalizability and external validity are typically understood as the ability to make inferences about a broader population findley2021external. By contrast, we take a more fundamental step by asking: even when we fix the same individual, the same treatment, and the same outcome measure, does the effect estimated in a survey experiment inform the corresponding real-world effect? If the answer is negative, then concerns about generalizing survey experimental results to other populations become secondary.

More specifically, following egami2023elements, we focus on two dimensions of generalizability: effect magnitude and effect sign. Effect magnitude generalization requires that the size of the estimated effect in the survey experiment corresponds to its magnitude in the real world. Effect sign generalization, by contrast, requires only that the direction (positive or negative) of the effect is preserved across settings. Importantly, an experiment need not satisfy both criteria simultaneously; the relevant notion of generalizability depends on the researcher’s objective. For example, policy evaluations often require accurate estimates of effect magnitudes to assess welfare implications, and thus demand magnitude consistency. In contrast, when experiments are used to test predictions derived from formal theory, preserving the direction of the effect is often sufficient. In some cases, experimental designs may even deliberately amplify treatment variation to increase statistical power, making sign consistency the more appropriate benchmark.

Building on our model, we characterize two mechanisms that govern the generalizability of both effect magnitude and effect sign. These mechanisms rest on two behavioral microfoundations.

The first is limited attention. Conventional models of decision-making typically assume that individuals possess unlimited attention—that is, they account for all relevant attributes when making choices. While such assumptions yield tractable and often sharp theoretical predictions, they have been increasingly challenged by empirical anomalies. Classic examples include the paradoxes documented by allais1953comportement and ellsberg1961risk, which illustrate systematic violations of expected utility theory.\footnote{The former shows that individuals frequently violate the independence axiom of expected utility theory, while the latter demonstrates ambiguity aversion inconsistent with subjective expected utility theory.} In practice, individuals rarely process all potentially relevant considerations. Instead, they rely on a selective and context-dependent consideration set. This notion has long been recognized in the social sciences. For example, consumer choices may favor option $y$ in the presence of $x$, yet this preference can reverse when $x$ is not readily available masatlioglu2016revealed. We incorporate this idea into the causal inference framework for survey experiments and show that differences in consideration sets across environments play a central role in shaping generalizability. In particular, the more restricted consideration sets typically induced by survey environments can mechanically amplify estimated effects relative to their real-world counterparts.

The second mechanism arises from salience. Even within a given consideration set, attributes are not weighted equally; some dimensions attract disproportionate attention. This phenomenon is well documented. For instance, chetty2009salience show that consumers underreact to taxes when tax-inclusive prices are not prominently displayed, but reduce demand by approximately eight percent when such prices are made salient. To formalize this mechanism, we draw on the psychologically grounded model of choice developed by bordalo2012salience,bordalo2013salience, in which decision weights are distorted by the relative salience of attributes. We show that salience, when combined with limited attention, can generate failures of effect sign generalization. The key intuition is that salience distorts the relative weighting of attributes, effectively inducing correlations among them in the decision process. When consideration sets differ between survey and real-world environments, these salience-driven distortions operate on different subsets of attributes, potentially reversing the direction of the estimated effect.

Although the core components of our theory—consideration sets and salience—are not directly observable, we assess the empirical relevance of our framework by deriving a set of testable implications from the theoretical results. A central prediction is that the estimated effect attenuates, and may even reverse, as the number of attributes increases in some experiments. Intuitively, expanding the attribute space alters both the consideration set and the relative salience of each attribute, thereby changing the mapping from experimental effects to their real-world counterparts. We examine this prediction using data from multiple existing studies and supplement the analysis with additional pre-registered conjoint experiments.\footnote{The study is pre-registered and approved by the authors’ Institutional Review Board; further details are provided in the SI (ref).} As the well-known adage suggests, “all models are wrong, but some are useful.” We do not claim that our framework exhausts all mechanisms underlying the empirical patterns we document. Rather, our goal is to isolate two theoretically grounded mechanisms—limited attention and salience—and to show that they generate distinctive, testable predictions.

Attention and salience have already emerged as important considerations in the study of conjoint experiments. For example, jenke2021using and bansak2025odd employ eye-tracking techniques to examine how respondents allocate attention across attributes. We build on this line of work by providing a formal framework that characterizes the conditions under which such patterns arise and how they affect causal interpretation. We emphasize that the definitions of attention and salience are quite versatile in the literature. Here, we define limited attention as a situation in which only a subset of attributes is under consideration, and we define salience as the extent to which the value of an attribute affects its importance. These definitions follow the framework of bordalo2012salience. This idea is closely related to the concept of extremity liu1998catastrophic, abelson2014attitude, krosnick1992case.

More broadly, our results suggest that, among survey-based designs, well-constructed conjoint experiments may offer advantages for generalizability. By presenting multiple attributes simultaneously, conjoint designs encourage respondents to consider a richer set of factors, thereby mitigating distortions induced by limited attention and salience. When the set of attributes included in the experiment sufficiently captures the core and relevant dimensions of real-world decision-making, concerns about generalizability are correspondingly reduced. This perspective helps rationalize prior empirical findings documenting the robustness of conjoint experiments jenke2021using,hainmueller2015validating,bansak2018number, bansak2021beyond, and provides a microfoundation for understanding when such robustness should be expected.

Our research intersects with and contributes to several important strands of literature. At its core, our study engages with a central question in the literature on external validity (barabas2010survey; list2005laboratory; gaines2007logic), particularly the external validity of survey experiments. slough2023external,slough2022sign develop a formal framework to clarify external validity by identifying conditions under which different studies are target-equivalent and target-congruent. Their framework emphasizes the importance of harmonizing studies along key dimensions, including treatment, control, and outcome measurement. Similarly, egami2023elements propose a unified framework encompassing multiple dimensions of external validity—population, context, treatment, and outcome—building on the typology of cook2002experimental. They further develop estimators for both effect magnitude and sign generalization. We complement this literature by providing a behavioral foundation for understanding when and why effect magnitude and sign generalization succeed or fail.

Much of this literature focuses on generalizing results to new populations (mullinix2015generalizability; huang2022sensitivity). By contrast, our paper examines a more fundamental form of external validity: the consistency between experimental effects and real-world effects, even when holding fixed the same individuals, treatments, and outcomes. A variety of factors may undermine such consistency, including well-known phenomena such as the Hawthorne effect adair1984hawthorne. We contribute to this literature by providing a new explanation tailored to settings involving multidimensional decision-making.

More specifically, concerns about the generalizability of survey experiments have been widely documented, with mixed empirical evidence barabas2010survey,findley2017external,hainmueller2014survey. Existing explanations for threats to external validity include features of experimental design de2020commensurability and information (non)equivalence across settings dafoe2018information. Departing from purely statistical approaches, we develop a behavioral framework that highlights the roles of limited attention and salience as key mechanisms underlying these discrepancies.

Our work also contributes to the emerging literature on the Theoretical Implications of Empirical Methods (TIEM) FU_SLOUGH_2026, slough2023external,de2020commensurability,slough2023phantom, which evaluates empirical methodologies through the lens of formal theory. In contrast to the Empirical Implications of Theoretical Models (EITM) approach, TIEM emphasizes how methodological choices themselves embed implicit theoretical assumptions. Within this perspective, our application of decision theory and behavioral economics to survey experiments reveals systematic discrepancies between experimental estimates and real-world effects, extending beyond the aggregate-level inconsistencies documented by abramson2019we.

Finally, we contribute to the literature on salience in political economy. While salience theory has been widely applied to topics such as electoral behavior, party competition, agenda setting, and judicial decision-making moniz2020issue, dragu2016agenda,riker1986art,ascencio2015endogenous,guthrie2000inside,viscusi2001jurors, bartels1986issue,niemi1985new,repass1929issue, its implications for research design have received less attention. By incorporating the framework developed by bordalo2012salience,bordalo2013salience,bordalo2016competition into the study of survey experiments, we highlight salience as a central consideration in experimental methodology, thereby extending its scope and applicability.

Model of Survey Experiments

Causal effects are typically formalized within the potential outcomes framework. To illustrate the mechanisms underlying generalizability, we develop a model that integrates behavioral microfoundations of attention and salience with the potential outcomes framework. We explicitly incorporate environmental features into the potential outcomes and emphasize that the target estimand in a survey experiment depends on the consideration set and the salience of attributes.

Attention and Salience

Consider a survey experiment conducted in a survey environment. Let \(E^s\) denote the survey environment, and \(E^r\) denote the target real-world environment. Let the informational content available to respondent \(i\) be summarized by an attribute vector $X_i = (X_{i1},\dots,X_{iK}) \in \mathcal{X}$. The components of \(X_i\) can be interpreted broadly. In a conjoint experiment, they correspond to profile attributes. In a vignette experiment, they represent features of a scenario. In an informational treatment, they capture the arguments, facts, or cues embedded in the message. In candidate or policy evaluation tasks, they describe the characteristics of the object being evaluated. Throughout the paper, we refer to $X_{ik}$ as attribute, and to its specific realization $x_{ik}$ as the value of the attribute.

Individuals have limited attention. A large body of research shows that, when making decisions—such as purchasing a computer—individuals do not evaluate the full set of available options or attributes hausman2008mindless. Instead, they focus on a subset of relevant information. In the literature, the subset of attributes or alternatives that a respondent attends to when making a decision is referred to as the consideration set hausman2008mindless,wright1977phased. It is well established that, due to cognitive constraints, individuals cannot allocate sufficient attention to all potentially relevant attributes stigler1961economics,jones2005politics,chetty2009salience. As a result, the consideration set is typically a strict subset of the full attribute space that may, in principle, influence a decision. Formally, let \[ C_i(E) \subseteq \{1,\dots,K\} \] denote respondent \(i\)'s active consideration set under environment \(E \in \{E^r,E^s\}\). If \(k \notin C_i(E)\), then attribute \(k\) does not enter the active evaluation in environment \(E\).

Following the salience model bordalo2012salience,bordalo2013salience, given an attribute vector $X_i$, respondent $i$ evaluates it in environment $E$ as: $$ V_i(X_i,E)=\sum_{k=1}^K \alpha_{ik}(X_i,E) u_{ik}(X_{ik})=\sum_{k \in C_i(E)} \alpha_{ik}(X_i,E) u_{ik}(X_{ik}) $$ where $\alpha_{ik}(X_i,E)>0$ denotes the salience weight assigned to attribute $k$, satisfying $\sum_{k=1}^K \alpha_{ik}(X_i,E)=1$. The function $u_{ik}(X_{ik})$ represents the utility contribution of attribute $k$. We assume that the baseline utility component $u_{ik}$ is invariant across environments. This restriction allows us to isolate the role of attention and salience—captured by $C_i(E)$ and $\alpha_{ik}(X_i,E)$—in shaping differences between the survey and real-world evaluations. It is worth noting that the environment $E$ may influence many aspects of the decision-making process. In this study, however, we focus exclusively on the roles of attention and salience.

The key idea is that salience $\alpha_{ik}(X_i,E)$ need not be fixed. As emphasized by taylor1982stalking, "salience refers to the phenomenon that when one’s attention is differentially directed to one portion on the environment rather than to others, the information contained in that portion will receive disproportionate weighing in subsequent judgment." This idea is formalized by a salience function $\sigma(x_k,\overline{X}_k)$, where $x_k$ denotes the realized value of attribute $X_k$, and $\overline{X}_k$ is a reference point. The reference may be individual-specific or determined by the experimental context. For example, in a vignette experiment where the attribute is the level of corruption, each respondent $i$ may have a baseline perception of the typical level of corruption among politicians. In a conjoint experiment, where respondents evaluate two hypothetical politicians, a natural reference point is the average level of corruption across the two profiles.

The salience function captures the “distance” between an attribute value and its reference point. Intuitively, attributes that deviate more from the reference point receive greater salience. A more formal treatment is provided in bordalo2012salience. The salience function is typically assumed to satisfy several axiomatic properties. First, for any two intervals $[x,y]$ and $[x',y']$, with $[x,y]$ fully contained within $[x',y']$, it holds that $\sigma(x,y) < \sigma(x',y')$. This reflects the psychological principle that larger contrasts are more perceptually salient. Second, the function satisfies homogeneity of degree zero: $\sigma(b x, b y) = \sigma(x,y)$ for any positive scalar $b$. A commonly used functional form that satisfies these properties is $\sigma(x,y) = \frac{|x-y|}{y}$.

Now we introduce how salience operates. Let $r_{ik}(X_i,E)\in\{1,\dots,|C_i(E)|\}$ denote the salience rank of attribute \(k\) among the active considerations in environment \(E\), where smaller values indicate greater salience. The effective salience weight $\alpha_{ik}$ takes the form

equation[equation omitted — 204 chars of source]

where \(\delta \in (0,1]\) indexes the strength of salience, and $\bar{\alpha}_{ik}$ denotes the baseline salience. When attribute $j$ is more salient than attribute $k$, as determined by the salience function, i.e., $\sigma(x_{ij},\overline{X}_{ij})>\sigma(x_{ik},\overline{X}_{ik})$, it follows that $r_{ij}<r_{ik}$. Consequently, under the effective salience weights, attribute $X_{k}$ is more heavily discounted, since $\delta^{r_{ik}-1}<\delta^{r_{ij}-1}$.

To illustrate the mechanism, consider a simple example. For ease of exposition, we suppress the individual subscript $i$. Suppose a candidate-choice conjoint experiment with two attributes: age $X_1$ and gender $X_2$. Let $\alpha_1$ and $\alpha_2$ denote the baseline salience weights, reflecting the importance of age and gender in candidate evaluation. When respondents observe two hypothetical candidates, salience is determined by how each attribute value deviates from its reference point. Without loss of generality, suppose that in the conjoint setting the reference for each attribute is given by the average across the two candidates $\overline{X}_j=\frac{x_1+x_2}{2}$. If, for a given comparison, gender is more salient than age, i.e., $\sigma(x_1,\overline{X}_1) < \sigma(x_2,\overline{X}_2)$, then $r_1 = 2$ and $r_2 = 1$. The resulting evaluation $V(X_1,X_2,E^s)$ with effective salience is

equation[equation omitted — 601 chars of source]

Accordingly, when an attribute, such as $x_1$, is relatively more salient, its weight in the utility function ($\alpha_1$) is effectively increased to $\frac{\alpha_1}{\alpha_1 + \delta \alpha_2}$, while the weight on the less salient attribute, $\alpha_2$, is reduced to $\frac{\delta \alpha_2}{\alpha_1 + \delta \alpha_2}$.The denominator ensures that the salience weights sum to one. The parameter $\delta$ governs the strength of the salience effect. When $\delta = 1$, there is no salience distortion, and attributes are weighted according to their baseline importance. As $\delta$ decreases below one, relatively more salient attributes receive disproportionately greater weight, while less salient attributes are increasingly discounted.

Causal Effects

We use $Z_i$ to denote treatment. In survey experiments, the treatment typically determines the values of the attributes that respondents observe. Accordingly, the attribute vector $X_i(Z_i)$ can be viewed as a function of $Z_i$. The individual causal effect on utility, comparing $Z_i=z$ and $Z_i=z'$ in environment $E$, is defined as $$ ICE_i(E)=V_i(X_i(z);E)-V_i(X_i(z');E). $$ where $=V_i(X_i(z);E)$ denotes the potential utility under treatment $z$.\footnote{For conjoint experiments, please see SI (ref). For factorial designs, the framework can be readily extended to accommodate multiple treatments.}

The outcome variable need not directly measure this latent evaluation. Instead, it is generally a function of the observed response, given by $ Y_i = g_i(V_i(X_i;E))$. This formulation nests several commonly used outcomes: forced-choice decisions, where \(g_i(v)=\mathbf{1}\{v\ge 0\}\); ratings or thermometer scales, where \(g_i\) is the identity or a discretized monotone transformation; and approval or support indicators, where \(g_i\) follows a threshold rule. Because \(g_i\) is typically weakly increasing, the latent comparison in $V_i$ remains the central theoretical object. Our results extend to any monotone transformation of $V_i$.

As discussed in the introduction, we examine whether causal effects identified in the survey environment generalize to the real-world environment while holding fixed the individuals, treatments, and outcome mappings. This constitutes a necessary precondition for any subsequent discussion of generalizability to a broader population. Following egami2023elements, we focus on two key dimensions of generalizability: effect magnitude and effect sign.

definition[Effect Magnitude Equivalence] $ICE_i(z,z';E^s)=ICE_i(z,z';E^r)$
definition[Effect Sign Congruence] $sign[ICE_i(z,z';E^s)]=sign[ICE_i(z,z';E^r)]$

Generalizability of average effects follows naturally once we aggregate individual-level causal effects across the population. We focus on ICE because it provides a clear way to illustrate the mechanism at the individual level and, as discussed earlier, because our notion of weaker generalizability involves holding individuals, treatments, and outcome mappings fixed.

Limited Attention and Survey Experiment Generalizability

This section studies how attention mechanisms affect generalizability, with a particular focus on the equivalence of effect magnitudes. To isolate the role of attention, we first suppress salience distortions and examine how generalizability is affected when the survey environment alters only the active consideration set. Accordingly, as a benchmark, we assume $\delta=1$ throughout this section. The evaluation function then simplifies to $V_i(X_i,E)=\sum_{h \in C_i(E)}\alpha_{ik} u_{ik}(X_{ik})$, where the (effective) salience $\alpha_{ik}$ are invariant to attribute values.

Individuals have limited attention. In most decision-making contexts, it is unrealistic for individuals to consider all potentially relevant features. Instead, decisions are based only on attributes that enter the active consideration set. A defining feature of the survey environment is that respondents’ attention is strongly shaped by the information explicitly provided in the experimental design. For example, in a candidate-choice experiment where only age and gender are presented, respondents are likely to focus primarily, if not exclusively, on these attributes. Ensuring that respondents attend to the provided information, rather than skim past it, is therefore a central concern in survey experiments, motivating the widespread use of attention checks to detect inattentive respondents.

By contrast, in the real-world environment, individuals are not constrained by the limited set of attributes specified by the researcher. When making analogous decisions outside the survey context, respondents may attend to a broader and potentially different set of features. To formalize this distinction, we analyze the attention mechanism under the following limited-attention assumption.

assumption[Limited Attention] $C_i(E^s) \subset C_i(E^r).$

The consideration set in survey experiments is often a subset of the real-world consideration set for several reasons. In many applications, this provides a natural benchmark. First, survey experiments are designed to address specific research questions. As a result, researchers typically include a selected set of substantively relevant attributes, along with additional attributes serving as controls. These attributes are therefore best understood as a subset of the real-world consideration set $C_i(E^r)$. Second, researchers rarely possess complete knowledge of all attributes that respondents consider in real-world decision-making. Even if such knowledge were available, it would generally be infeasible to present the full set of relevant attributes within the constraints of a single survey instrument.

Ideally, respondents attend primarily to the attributes presented in the survey experiment, as these are the sources of randomized variation. However, in practice, respondents may also infer unlisted characteristics or rely on prior beliefs dafoe2018information. A more general formulation therefore allows for $C_i(E^s) \neq C_i(E^r)$. In some special cases, it is even possible that $C_i(E^s) \supset C_i(E^r)$. Under such scenarios, our results continue to hold qualitatively, although the implications shift—from amplified effect magnitudes to potentially attenuated or unequal effects. Empirically, however, available evidence is more consistent with the limited-attention benchmark, although direct evidence remains limited.

Without loss of generality, we assume that the real-world consideration set is $C_I^{r} = \{X_1, X_2, ..., X_m\}$, with associated baseline salience weights denoted by $\alpha^r=(\alpha_{i1},...,\alpha_{im})$. The experimental consideration set is given by $C_i^{s} = \{X_1, X_2, ..., X_k\}$, with corresponding baseline salience $\alpha^s=(\alpha'_{i1},...,\alpha'_{ik})$. We assume the cardinality of the real-world consideration set, $m=|C^r_i|$, exceeds that of the experimental consideration set, $k=|C_i^{s}|$. In other words, the experimental consideration set consists of only the first $k$ attributes from the real-world consideration set.

Now, we provide intuition for how limited attention affects the magnitude of causal effects. For ease of exposition, suppose the treatment in the survey experiment changes only the first attribute $X_{i1}(Z_i)$, from $x_{i1}$ to $x'_{i1}$. Then, the utility changes from $V_i(x_{i1},x_{i-1},E^s)=\alpha_{i1}u_{i1}(x_{i1})+ \sum_{j=2}^m \alpha_j u_{ij}(x_{ij})$ to $V_i(x'_{i1},,x_{i-1},E^s)=\alpha_{i1}u_{i1}(x'_{i1})+ \sum_{j=2}^m \alpha_j u_{ij}(x_{ij})$. Therefore, the ICE in the survey experiment is $ICE_I(E^S)=V_i(x_{i1},x_{i-1},E^s)-V_i(x'_{i1},x_{i-1},E^s)=\alpha_{i1}[u_{i1}(x_{i1})-u_{i1}(x'_{i1})]$. Similarly, in the real-world environment, the ICE is $\alpha'_{i1}[u_{i1}(x_{i1})-u_{i1}(x'_{i1})]$. Unless the baseline salience weights coincide, the two effects will generally differ. Under limited attention, the survey consideration set contains fewer attributes than the real-world consideration set ($C_i(E^s) \subset C_i(E^r)$). As a result, each attribute in the smaller consideration set $C_i(E^s)$ tends to receive relatively greater weight than in $C_i(E^r)$. Consequently, it is unlikely that $\alpha_{i1}=\alpha'_{i1}$. Instead, it is more plausible that $\alpha_{i1}>\alpha'_{i1}$, implying that the causal effect is amplified in the survey environment, $ICE(E^s)>ICE(E^r)$.

Formally establishing this amplification result requires additional assumptions. In particular, when the real-world consideration set includes additional attributes, these attributes may be correlated with existing ones—especially the focal attribute $X_{i1}$-which can alter baseline salience weights and potentially attenuate or even reverse the effect. Although such scenarios are possible, our focus is on isolating the limited-attention mechanism. Moreover, well-designed survey experiments rarely omit attributes that are highly correlated with the focal attribute, as doing so would undermine the internal validity of the design. To formalize this idea, we impose that the relative importance of attributes is preserved even when additional attributes are introduced.

assumptionFor two consideration set $C^1_i$ and $C^2_i$, and any $j,k \in C^1_i \cap C^2_i$, $$ \frac{\alpha_{ij}}{\alpha_{ik}}=\frac{\alpha'_{ij}}{\alpha'_{ik}} $$

Specifically, if a respondent perceives attribute $X_1$ to be more important than attribute $X_2$ in one setting, then in a comparable setting with an expanded consideration set, the respondent continues to rank $X_1$ as more important than $X_2$. This assumption is more likely to hold in well-designed experiments that explicitly incorporate attributes highly correlated with the focal dimensions of interest, thereby ensuring that any additional attributes in the real-world consideration set are not strongly correlated with those included in the experimental design.

The following results hold for general treatment.Let $D^i=\{j\in \{1,...,k\}|u_{ij}(X_{ij}(z)) \neq u_{ij}(X_{ij}(z'))\}$ denote the set of attribute indices for which the treatment affects individual $i$.

proposition[Effect Magnitude Non-Generalizability] Given treatment assignment $Z=z$ and $Z=z'$, suppose the limited attention assumption (ref) hold, and $D_i \neq \emptyset$. (1) The causal effects in the survey experiment and the real world differ almost surely: $$ICE(E^s) \neq ICE(E^s)$$ if $\alpha_{ij} \neq \alpha'_{ij}$ for some $j \in C^r_s$ and $u_{ij}(X_{ij}(z)) \neq u_{ij}(X_{ij}(z'))$. (2) Moreover, if the assumption (ref) also holds, then the causal effect in the survey experiment are amplified relative to those in the real-world environment by a factor \(\delta = \frac{1}{\alpha'_1 + \alpha'_2 + \ldots + \alpha'_k} > 1\).
proofAll proofs are in the SI.

The first result follows directly. When baseline salience differs between the survey and real-world environments due to limited attention, the magnitude of causal effects estimated in the experimental setting will, in general, not be externally valid.

Under the additional stable salience assumption, we obtain a sharper result. If the inclusion of additional attributes in the real-world environment does not distort the relative importance of the attributes already present in the experiment, then the experimental causal effect is systematically amplified relative to the real-world effect. This condition is more likely to hold when the experimental design incorporates the most substantively important attributes—particularly those that are highly correlated with other relevant dimensions—so that omitted attributes do not substantially alter relative salience rankings.

The amplification factor \(\delta\) depends on the total salience of the attributes included in the survey experiment, as measured by their real-world salience weights $\alpha$. When respondents place substantial weight on attributes that are omitted from the experimental design, the share of total salience allocated to the included attributes is correspondingly reduced in the real world. As a result, the experimental estimate which implicitly redistributes attention over a smaller set of attributes can substantially overstate the true causal effect.

Empirical Evidence

The formal model and Proposition (ref) illustrate the mechanism through which limited attention distorts the magnitude of experimental causal effects. We now turn to empirical evidence. Our goal is not to “prove” the model—no formal model can be proven in that sense. The value of the model here lies in its ability to clarify and organize an empirically relevant phenomenon.

One implication of Proposition (ref) is that amplification bias should decline as more attributes are included in the consideration set. Of course, neither the true consideration set nor the real-world causal effect is directly observable. What researchers can manipulate, however, is the set of attributes presented in the survey experiment. When respondents attend to the experimental task, their attention is likely to be concentrated primarily on the attributes explicitly provided by the researcher. This is particularly true for conjoint experiments compared to other types of survey experiments. We therefore focus primarily on conjoint experiments. We argue that the attributes included in the experiment largely determine, or at least dominate, the active consideration set.

To examine this implication, we first study the 67 candidate-choice conjoint and vignette experiments compiled by schwarz2022have. In these experiments, a gender attribute is randomly assigned, and the outcome measures are comparable or can be consistently recoded across studies. This common structure allows us to isolate and compare the average causal effect of the gender attribute across studies. In addition, we collected information on the number of non-gender attributes presented in each experiment. Notably, in some studies, the number of attributes provided is as small as two.

To investigate the hypothesized relationship between the number of attributes and the AMCE of gender across these experiments, we conduct a meta-analysis. Figure (ref) presents the meta-regression results based on a random-effects model, both for the full sample of countries and for the United States subsample, given that the majority of experiments were conducted in the United States. The vertical axis reports the absolute value of the estimated effect, recognizing that the causal effect of gender may be negative in some studies.\footnote{Consistent with our theoretical framework, a true positive effect is expected to diminish as the number of attributes increases, whereas a true negative effect should move toward zero in magnitude (i.e., its absolute value decreases).} Across both specifications, we find a statistically significant negative relationship between the number of attributes and the experimental causal effect of gender, consistent with our theoretical predictions. The results are robust to excluding experiments with a small number of negative estimates as well as potential outliers.

figure[figure omitted — 545 chars of source]

To address cross-study heterogeneity and more directly test our hypothesis regarding the influence of the number of attributes on experimental effects, we designed and conducted an original candidate-choice experiment with a controlled and consistent attribute set. The experiment was fielded to 1,200 respondents in the United States via Lucid in September 2023. Lucid employs a quota sampling strategy to align participant demographics with those of the U.S. Census. Prior research by coppock2019validating shows that Lucid samples yield behavioral patterns comparable to those observed in nationally representative benchmark experiments.

Additional details on the experimental design are provided in SI (ref). Respondents were randomly assigned to one of five groups. Within each group, participants completed six paired candidate-choice tasks. The groups differed in the number of attributes presented, while holding the specific attribute set fixed within each group.

\fbox{ \parbox{\textwidth}{ - Group 1 was presented with only two attributes: gender and age.

- Group 2 included the previous attributes plus education and tax policy, totaling four attributes.

- Group 3 added race and income to the attributes in Group 2, resulting in six attributes.

- Group 4 included military service and religious beliefs, bringing the total to eight attributes.

- Group 5 encompassed ten attributes by adding children and marital status to those in Group 4. } }

Consistent with our meta-analytic focus, we examine the effect of gender across these groups. The results, presented in Figure (ref), reveal a statistically significant negative relationship between the number of attributes and the magnitude of the experimental gender effect, despite the limited number of observations (five groups). This pattern provides supporting evidence for our theory, which predicts that effect sizes diminish as the number of attributes increases.

An exception arises in Group 4 (with eight attributes), where the estimated gender effect is larger than expected. While this deviation may reflect sampling variability, an alternative explanation is that salience effects may be at play—a possibility we explore in subsequent sections.

figure[figure omitted — 508 chars of source]

Salience and Survey Experiment Generalizability

In this section, we incorporate the role of salience. Building on the framework developed by bordalo2012salience, we examine how salience, in conjunction with limited attention, further shapes the generalizability of survey experiments. In particular, we focus on the possibility of effect sign reversal across environments.

Not all attributes are equally important in decision-making. Evidence from survey experiments, including studies using eye-tracking techniques, shows that respondents process information selectively, focusing on attributes they perceive as important while ignoring less relevant ones as task complexity increases. For example, jenke2021using demonstrate that respondents in conjoint experiments allocate attention unevenly across attributes. In salience theory, attention is differentially allocated to attributes that stand out relative to a reference point, as captured by the salience function $\sigma(x_{ij},\overline{X}_{ij})$. To operationalize this idea, we rank attributes by salience, where a smaller rank $r$ indicates greater salience. The resulting mechanism is summarized by equation (ref), which we reproduce here for convenience: $$ \alpha_{ik}(X_i,E) = \frac{\bar{\alpha}_{ik}\,\delta^{\,r_{ik}(X_i,E)-1}} {\sum_{j\in C_i(E)} \bar{\alpha}_{ij}\,\delta^{\,r_{ij}(X_i,E)-1}} $$ Therefore, if an attribute lies close to its reference point, its salience is substantially discounted; conversely, attributes that deviate markedly from the reference receive relatively greater salience. The parameter $\delta$ governs the strength of this effect. To capture the possibility that perceived attribute salience varies across decision contexts, we introduce the following assumption:

assumption[Salience Effect] $\delta<1$.

The salience model emphasizes that the weight assigned to each attribute depends on the extent to which its realized value deviates from a reference point or prevailing expectation. Throughout this section, we maintain the assumption (ref). Absent this condition, changes in the attribute set could alter relative salience rankings in ways that confound the mechanism of interest, making it difficult to disentangle salience-driven distortions from broader shifts in underlying preferences.

Effect Sign Reversal

In the previous section, we showed that experimental estimates can be distorted by limited attention, leading to systematically amplified causal effects. Such amplification may, in some cases, be advantageous for researchers seeking to detect the direction of an effect, as it can reduce the sample size required to achieve a given level of statistical power.

A more conservative—and ultimately more fundamental—requirement for experimental evidence to meaningfully validate theoretical predictions is effect sign congruence: the experimental causal effect should have the same direction as the corresponding real-world effect slough2022sign. However, once salience effects are introduced, this requirement may fail. In particular, salience-induced reweighting of attributes can generate effect sign reversal, thereby undermining the reliability of experimental findings as indicators of underlying causal relationships.

As is evident from equation (ref), salience depends on the realized attribute levels. Accordingly, the salience rank for attribute $k$ can be expressed as a function of the treatment, $r_{k}(Z_i)$. The following proposition provides a sufficient condition under which effect sign reversal may occur.

proposition[Effect Sign Reversal] Suppose that assumptions (ref) and (ref) hold. The direction of the experimental effects may differ from their real-world counterparts if there exists an attribute $k$ such that $r_k(z_i) \neq r_k(z'_i)$.

The key condition in this proposition -- that there exists an attribute $k$ such that the salience ranking changes across treatment states -- captures a scenario in which a change in the treatment alters the relative salience of at least one attribute. As shown in the proof, such a reversal can arise regardless of the utility magnitudes associated with excluded attributes or the baseline salience of the treatment attribute itself.

To build intuition for the phenomenon of effect sign reversal, consider again a survey experiment in which the treatment affects only the first attribute. As illustrated in the left panel of Table (ref), under the control condition, the attribute values are $(x_1,x_2)$ with corresponding salience weights $(\alpha_1,\alpha_2)$. The treatment changes only $x_1$ to $\tilde{x}_1$, leading to updated salience weights $(\alpha'_1,\alpha'_2)$. In the real-world environment, suppose there is an additional attribute $X_3$ in the consideration set, as shown in the right panel of the table (ref). Under the same experiment, only the first attribute is affected. The individual causal effect in the survey environment is given by $$ICE(E^s)=\sum_{k=1}^2\alpha_{ik}u_{ik}(x_k)-\alpha_{ik}u_{ik}(\tilde{x}_k)$$ where $\tilde{x}_2=x_2$ since the treatment only affects the first attribute. In contrast, in the real-world environment, the individual causal effect is $$ ICE(E^r)=\sum_{k=1}^2\beta_{ik}u_{ik}(x_k)-\beta_{ik}u_{ik}(\tilde{x}_k)+[\beta_3u_{ik}(x_3)-\tilde{\beta}_3 u_{ik}(\tilde{x}_3)]. $$ We observe that salience affects the individual causal effect through two distinct channels. First, changes in salience weights imply that, even when the treatment does not alter the values of other attributes, those attributes still contribute to the treatment effect via reweighting. In particular, because the real-world environment includes an additional attribute, the salience weights assigned to the first two attributes $(\beta)$ differ from those in the survey environment $(\alpha)$. This reallocation of attention can lead to differences in effect magnitude and, potentially, in effect sign across environments. Second, in the real-world environment, $ICE(E^r)$ includes an additional term arising from the third attribute, $[\beta_3u_{ik}(x_3)-\tilde{\beta}_3 u_{ik}(\tilde{x}_3)]$. This additional component can be sufficiently large to overturn the direction of the treatment effect, thereby inducing effect sign reversal.

table[table omitted — 922 chars of source]

It is immediate that if there exist parameter configurations under which a positive effect can be reversed to a negative one, then it is even more likely that configurations exist under which the effect is attenuated to zero. We formalize this observation in the following corollary:

corollarySuppose assumptions (ref) and (ref) hold. If there exists $j$ and $k$, such that $r_j(z_i) \neq r_k(z'_i)$, then the experimental treatment effect may become null in the real-world environment, and vice versa.

It is important to emphasize that our results are existence results; they do not imply that effect sign reversal or attenuation must occur in practice. For example, if a particular attribute—or its associated salience—dominates the evaluation process, then reversal is unlikely. The intuition is straightforward. Suppose attribute $X_1$ is substantially more important than all other attributes. In that case, it yield a large utility contribution $u(X_1)$. Consequently, in the overall evaluation $V$, the term $\alpha_1 u(X_1)$ dominates. Even if salience rankings shift, this dominant component is unlikely to be offset by changes in other attributes. As a result, the sign of $V$, and hence of the individual causal effect, will largely be determined by $X_1$, making sign reversal unlikely in such settings.

Another scenario that precludes effect sign reversal or attenuation arises when the rank-change condition is not satisfied. As emphasized in Proposition (ref), variation in salience rankings across treatment states is essential for generating sign reversal. The following proposition formalizes that, in the absence of such rank changes, effect sign reversal cannot occur.

propositionAssuming conditions (ref) and (ref) hold. If the relative salience rankings satisfy $r_k(z) = r_k(z')$ for all $k$, then the effect sign reversal cannot occur.

Empirical Evidence

The preceding propositions illustrate how salience can generate effect sign reversal or effect attenuation. We cannot directly test the theory, because neither the real-world consideration set nor the exact salience weights are observed or experimentally controlled. What we can do, instead, is derive testable implications. As before, our goal is not to “prove” the model.

Although we cannot manipulate the real-world consideration set, we do expect to substantially influence the consideration set in the survey environment. Building on Proposition (ref) and Corollary (ref), we therefore derive the following testable implications:

hypothesisThe sign of the effect in a conjoint experiment may reverse or attenuate to null as the number of attributes increases.

The proposition also highlights that changes in salience rankings are necessary for this phenomenon to arise. This raises the question: under what conditions can we expect salience rankings to remain unchanged, as assumed in the preceding proposition? The mechanism underlying salience implies that rankings are determined by the salience function $\sigma(x_k,\overline{X}_k)$. If the treatment-induced change in the attribute $X_1$ is not sufficiently large to meaningfully alter the reference point $\overline{X}_k$, then $\sigma(x_k,\overline{X}_k)$ will be close to $\sigma(\tilde{x}_k,\overline{X}_k)$. In such cases, the change in salience is too small to affect the relative ranking of attributes. Motivated by this intuition, we propose the following implication:

hypothesisAttribute effect sign reversal or attenuation is less likely when changes in attribute levels are marginal.

We draw on data from a conjoint experiment on hotel rooms conducted by bansak2021beyond. The study identifies four core attributes of hotel rooms: “view from the room (ocean or mountain view), floor (top, club lounge, or gym and spa floor), bedroom furniture (1 king bed and 1 small couch or 1 queen bed and 1 large couch), and the type of in-room wireless internet (free standard or paid high-bandwidth wireless).” In addition to these core attributes, the authors include 18 supplementary attributes that are largely unrelated to the core set.

Respondents were asked to choose their preferred hotel room from 15 paired comparisons, where each profile included the four core attributes along with a randomly selected subset of additional attributes. As a result, respondents were randomly assigned to one of 11 experimental conditions, with profiles containing 4, 5, 6, 7, 8, 9, 10, 12, 14, 18, or up to 22 attributes.

Figure (ref) presents a heatmap of the results.\footnote{Figure (ref) in the SI reports statistical significance. Due to limited statistical power, relatively few estimates are statistically significant. As the sample size increases, confidence intervals would be expected to narrow, potentially revealing additional cases of statistically significant effect reversals. At the same time, null effects are also consistent with our theoretical predictions and align with implication (ref).} The horizontal axis indicates the number of attributes, while the vertical axis corresponds to attribute levels. Darker colors denote negative AMCEs, whereas lighter colors indicate positive AMCEs. The results provide support for our hypothesis on effect sign reversal (implication (ref)). In particular, the estimated effects of some attributes—such as menu, bar, closet, and pillow—exhibit substantial instability across conditions. By contrast, other attributes, including View, Towels, and Internet, display considerable stability. This pattern is consistent with our theoretical framework: these attributes likely carry high baseline salience and utility, allowing them to dominate the evaluation process and remain robust to changes in the attribute set. This evidence also reinforces the classification of View and Internet as core attributes, as emphasized by bansak2021beyond.

To test implication (ref), we adapted our candidate-choice experimental design. To better control for salience, rather than randomly assigning attribute levels across profiles, we constrained level differences to be minimal. For example, for the age attribute with five levels (40, 52, 60, 68, 75), only adjacent values were allowed to appear in a comparison. Thus, if one profile was assigned age 52, the other profile could only take values 40 or 60.

As in the previous experiment, this study was fielded to 1,200 U.S. respondents via Lucid in November 2023. Each respondent completed five paired choice tasks. Table (ref) reports the AMCEs for gender under both the original design and the modified, reduced-salience design. When attribute levels were assigned randomly—without controlling for salience—the corresponding AMCEs (Column 2) exhibit a sign reversal as the number of attributes increases from six to eight. By contrast, Column 3 presents results from the reduced-salience design. A key observation is the absence of a clear effect sign reversal under this specification. We emphasize that this finding does not imply that researchers should universally restrict conjoint designs to adjacent attribute levels. Rather, this design choice is intended as a targeted test of the theoretical mechanism.

table[table omitted — 401 chars of source]
figure[figure omitted — 472 chars of source]

We emphasize again that our results are not claims about inevitability or frequency. We do not assert that such inconsistencies must occur, nor do we quantify how often they arise in practice. Rather, our goal is to illustrate a plausible mechanism. Other mechanisms—and potentially offsetting countervailing forces—may also be at work.

Indeed, based on our theory, relative to other types of survey experiments, well-designed conjoint experiments may be less susceptible to generalizability concerns. Because they incorporate multiple attributes within a single design, they substantially reduce the likelihood of omitting attributes that are important in real-world decision-making. The robustness of conjoint experiments has been widely discussed in the literature jenke2021using,hainmueller2015validating.

Discussion and Concluding Remarks

An increasing number of studies in the social sciences employ experiments to identify causal effects (druckman2006growth). While such designs provide strong internal validity, external validity remains a longstanding concern. This paper develops a formal framework to explain why survey experiments may fail to satisfy this stronger notion of generalizability, even when treatment is randomized and even when the same individuals, treatment content, and outcome measures are held fixed. This limitation raises important concerns about the extent to which experimental findings can be extrapolated to real-world settings. For example, boas2019norms show that although voters appear to sanction corruption information in survey experiments, they do not take action when presented with similar information about their own mayor in a field setting.

We provide both theoretical and empirical evidence demonstrating how two mechanisms—limited attention and salience effects—can undermine the generalizability of survey experiments. In particular, we show that experimental effects may be systematically amplified in magnitude and, in some cases, may even diverge in direction from their real-world counterparts.

A caveat to our findings is that we do not claim that non-generalizability is inevitable in survey experiments. Many studies—for example, jenke2021using—as well as empirical evidence presented in this paper, suggest that certain attributes exhibit stability across contexts in conjoint experiments. Our theoretical framework also highlights that conjoint experiments possess particular advantages for achieving high levels of generalizability.

It is important to emphasize that salience effects are not inherently detrimental to experimental design. Whether individuals make decisions in real-world contexts or within experimental settings, salience is an integral component of multi-dimensional decision-making. Accordingly, the goal should not be to eliminate salience effects altogether. Rather, the objective is to ensure that the salience patterns induced in the experiment closely mirror those that arise in real-world environments. It is the artificially induced salience distortions—those that diverge from real-world conditions—that are most problematic.

Notably, in the absence of attention distortions—i.e., if the experimental environment perfectly replicates the real-world decision context—experimental ICEs would coincide with their real-world counterparts, precisely because salience patterns would align across environments. The challenge, of course, is that achieving such equivalence is rarely feasible in practice.

A more practical approach is to design attribute levels and profile combinations that closely approximate real-world scenarios. This calls for a phased research design that emphasizes realism, relevance, and precision. A natural starting point is a comprehensive review of the existing literature and available data on the decision-making context of interest. This preliminary step should be complemented by exploratory interviews or pilot surveys with individuals who engage in similar decisions in real-world settings. Such qualitative and descriptive evidence is essential for identifying the attributes and attribute levels that are most salient and substantively relevant to the decision-making process.

spacing{0.0}

\setcounter{page}{1}