Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
48,759 characters · 31 sections · 49 citation commands
Identifying Causal Effects in Information Provision Experiments
\addtocontents{toc}{\setcounter{tocdepth}{-1}}
{.25em} {.25em}
Information provision experiments have become a standard tool for studying the causal effects of beliefs wiswallDeterminantsCollege15, bottanChoosingYour22, jensenPerceivedReturns10. But standard panel and two-stage least squares (TSLS) estimators systematically misrepresent average effects because they overweight individuals who update their beliefs the most. This matters because individuals whose beliefs most strongly affect their choices tend to update their beliefs the least, perhaps because they already sought out information before the experiment began. I propose a local linear slopes (LLS) estimator that weights all individuals equally. In five of six recent studies I reanalyze, LLS yields substantially larger estimates; in two cases the estimates more than double.
This paper is about experiments that study the causal effects of beliefs: how beliefs affect behavior, policy preferences, and even other beliefs. In these experiments, researchers vary the information (\q{signal}) shown to participants, then estimate the effect of beliefs on behavior using panel or TSLS regressions. It is well known that such estimators target weighted averages of individual causal effects.\footnote{The weighted average interpretation of TSLS follows from imbensIdentificationEstimation94. Similar results apply to difference-in-differences and other settings sunEstimatingDynamic20, goodman-baconDifferenceindifferencesVariation21, callawayDifferenceinDifferencesMultiple21.} In information provision experiments, these weights are proportional to the first-stage effect of information on beliefs.
Strong dependence between belief updating and belief effects makes panel and TSLS estimators substantially misrepresent average effects. When belief updating is negatively correlated with belief effects, standard panel or TSLS estimators can severely understate the average effect. The central empirical finding of this paper is that belief effects and belief updating are systematically negatively correlated: individuals whose beliefs most strongly affect their choices tend to update their beliefs least when provided new information.
I therefore propose a local least squares (LLS) estimator that consistently estimates an unweighted average effect, even when there is strong dependence between belief updating and belief effects. Researchers may prefer targeting an unweighted average as it is a representative summary of heterogeneous effects. \footnote{In an early application, guentherPoliticalRepresentation25 use results from the working paper version of this paper infoiv_v4. They use LLS because unequal weights cause \q{2SLS [to] substantially misrepresent average effects} and so \q{we adopt [LLS] which identifies the unweighted average effect (p. 18-19)}} This estimator can be applied to panel, active control, and passive control experiments.\footnote{ The LLS estimator applies immediately in the panel experiment. In experiments with active control groups, the LLS estimator identifies an unweighted average under a learning rate updating assumption. In experiments with passive control groups, the LLS estimator identifies the unweighted average when the variance of the prior is elicited in addition to the mean and the learning rate comes from Bayesian updating. An alternative approach with a passive control imposes the strong assumption that covariates are sufficiently rich to predict the belief update and that there is no residual variation in beliefs that cannot be predicted (i.e. \q{selection on observables}.)}
I apply the LLS estimator to six recent information provision studies published in leading economics journals.\footnote{ These applications span diverse contexts: college major choice wiswallDeterminantsCollege15, housing investment armonaHomePrice19, gender policy preferences setteleHowBeliefs22, household rothHowExpectations20 and firm kumarEffectMacroeconomic23 responses to macroeconomic uncertainty, and protest participation cantoniProtestsStrategic19. These six studies include examples of within-person panel experiments, and between person experiments with both active and passive control groups. } In five of these six applications, the LLS estimates are meaningfully larger than the panel or TSLS estimates. In two cases the estimates more than double. To study mechanisms, I show how LLS can also be used to estimate effects of beliefs on outcomes conditional on the learning rate. Empirically, belief effects are generally larger for the groups with smaller learning rates. A simple model of endogenous information acquisition can rationalize this pattern. People whose beliefs strongly affect their decisions are incentived to form precise priors; when researchers provide new information, they update only modestly. When beliefs matter less, people start with noisier priors and update more.\footnote{mackowiakRationalInattentionForthcoming consider a similar model with rational inattention before and during the experiment. Since they argue that the rational inattention dynamics before the experiment dominate, their results are consistent with the model proposed in Appendix (ref) that does not include rational inattention during the experiment. }
The identification arguments in this paper use results in correlated random coefficients models from mastenIdentificationInstrumental16 and grahamIdentificationEstimation12, generalized here to a nonparametric potential outcomes framework. vz_aeri study TSLS in information provision experiments and provide conditions under which TSLS targets some non-negatively weighted average. This paper proposes an alternative to TSLS that targets the equally-weighted average.
The remainder of this paper is organized as follows. Section (ref) develops the conceptual framework. Section (ref) shows that standard panel and TSLS estimators target weighted averages of individual slopes; panel regressions have negative weights. Section (ref) proposes the LLS estimator, which identifies an unweighted average. Section (ref) shows that under linearity, the unweighted average can be used to extrapolate. Section (ref) shows that attenuation is empirically widespread. Section (ref) concludes.
This paper is about experiments that study how beliefs affect behavior. I analyze three leading experimental designs: panel experiments that compare the same individual before and after information provision, active control experiments that compare individuals receiving different signals, and passive control experiments that compare treated individuals to an untreated control group.\footnote{In between-subject experiments (with active or passive controls), I will focus on experimental designs where the information treatment is quantitative, for example \q{12 percent of the US population are immigrants} hopkinsMutedConsequences19,grigorieffDoesInformation20 and not treatments that are qualitative, for example \q{[t]he chances of a poor kid staying poor as an adult are extremely large} alesinaIntergenerationalMobility18. The results for within-person (panel) experiments extend to qualitative or other kinds of signals.}
The identification argument follows a simple causal chain: treatment assignment $Z$ determines the signal $S$ shown to participants, which affects their beliefs $X$, which in turn affects outcomes $Y$. This $Z \rightarrow S \rightarrow X \rightarrow Y$ structure allows us to study how exogenous variation in information provision translates into belief changes and ultimately behavioral responses. I formalize this causal chain in three parts: the outcome equation that links beliefs to behavior, the experimental designs that generate exogenous variation in beliefs, and the identifying assumptions that permit causal inference.
The outcome equation allows for arbitrary heterogeneity in how beliefs affect outcomes:
where $Y_i$ is the outcome or behavior of interest, $X_i$ is the belief, and $G_i(\cdot)$ is the individual-specific response function. The function $G_i(\cdot)$ generates potential outcomes: $Y_i(x) = G_i(x)$, where $Y_i(x)$ is $i$'s potential outcome when beliefs are exogenously set to $x$. This formulation places no restriction on treatment effect heterogeneity; agents can differ both in their average responsiveness to beliefs and in the shape of their response functions.
We assume that beliefs $X_i$ are endogenous in the sense that $\E[Y_i \mid X_i = x] \neq \E[G_i(x)]$ for at least some $x$. This says that the difference in outcomes at two values of $X$ is not a causal effect.\footnote{If $G_i(x) = c x + U_i$ this is a familiar expression of endogeneity bias $\E[U_i \mid X_i] \neq \E[U_i]$.} This occurs when unobserved determinants of outcomes also affect beliefs.
This paper considers three broad classes of information provision experiments. The first design uses within-person panel variation.
The second and third designs use between-person variation, but differ in the construction of the control group.
Within and between person designs use different kinds of identifying variation and rely on qualitatively different kinds of identifying assumptions. The within person design uses the panel structure on the outcome and does not rely on any assumption on {how} people update beliefs in response to new information. In contrast, between person designs use assumptions on belief updating to match treatment units to the appropriate control units.\footnote{ In principle, panel and active control designs could be combined by eliciting pre-treatment outcomes in an active control experiment. Exploring the identification implications of such hybrid designs is beyond the scope of this paper but is an interesting direction for future research. }
The identifying assumption is that outcomes follow a \q{panel} form. Let time $t$ have two periods, denoting pre ($t=0$) and post ($t=1$) information provision. Then, let
The response function $G_i(\cdot)$ is time-invariant but arbitrarily heterogeneous across individuals; the time effects $\gamma_t$ are additively separable. This is a nonparametric generalization of the standard panel model used in the literature (e.g. wiswallDeterminantsCollege15,armonaHomePrice19). The special case $G_i(x) = \tau_i x + U_i$ generates the classic linear panel model $Y_{it} = \tau_i X_i + U_i + \gamma_t$ with heterogeneous treatment effects.
The identifying assumption that different changes in outcomes are due only to different changes in beliefs.\footnote{The time trend $\gamma_t$ is commonplace in empirical practice armonaHomePrice19,wiswallDeterminantsCollege15. This allows for all respondents to, for example, respond with a higher number when the outcome is re-elicited, perhaps because of salience or other behavioral factors. The time trend $\gamma_t$ can be interacted with observables $W_i$ to allow for these time trends to vary across observables, like the prior belief. Without a time trend, the model implies that outcomes should not change when beliefs do not change: $\E \bs{\Delta Y_i | \Delta X_i = 0} = 0$. This restriction is testable in the data.} There are no assumptions on how beliefs are updated; researchers who do not wish to place structure on belief updating may find the panel design particularly appealing.
In active and passive control experiments, the relationship between beliefs and outcomes is completely flexible. The identifying assumption is that belief updating follows a simple {learning rate} structure. This includes the workhorse linear updating or \q{signal averaging} models like Bayesian updating. Randomization to a particular signal generates variation in posterior beliefs through this learning rate updating.
For exposition, the main text uses the familiar linear form throughout; potential beliefs are a linear function of the prior $X_i^0$ and an experimental signal $s$:
The heterogeneous coefficient on the signal $\alpha_i$ is often called the learning rate. In this model, posterior beliefs are a weighted average of the prior and the signal, with weight $\alpha_i$ on the signal. This updating rule is widely used in applied work and fits observed belief changes well in information provision experiments.\footnote{See for example cavalloInflationExpectations17, cullenHowMuch22, giaccobassoWhereMy22, cullenIncreasingDemand23,fusterExpectationsEndogenous22.}
This linear updating rule is often microfounded in a normal-normal Bayesian updating, but it also arises in several other behavioral models.This class of linear updating models includes rational inattention fusterExpectationsEndogenous22, base-rate neglect, over-reaction, under-reaction gretherBayesRule80 and anchoring on the prior or signal gabaixBehavioralInattention19. \footnote{ Linearity in belief updating can be relaxed as long as differences in updating are still driven only by the learning rate. Nonlinear learning rate models take the form $X_i(s) = \alpha_i f(s, X_i^0) + X^0_i$, where $f(\cdot, X_i^0)$ is any function monotonic in the signal with $f(X_i^0, X_i^0) = 0$. For example, $f$ could be a nonlinear \q{dampener} that discounts signals further away from the prior. Or, it could be asymmetric around zero so that people respond more to signals of a particular sign. The remainder of the paper uses the linear updating rule with $f(s, X_i^0) = s-X_i^0$ due to its overwhelming popularity in practice and because it can be microfounded in many popular models of belief updating.} See Appendix (ref) for further discussion.
Denote treatment arms by $Z_i$. In the active and passive control designs, assume that the researcher randomizes over two arms $Z_i \in \{A,B\}$. In the active design, arm $A$ will be the treatment arm that receives the \q{high} signal and arm $B$ will be the treatment arm that receives the \q{low} signal. In the passive design, arm $A$ will be the treatment arm that receives a signal and arm $B$ will be the control arm that does not receive a signal. Finally, $S_i(z)$ is the signal that is shown to individual $i$ in treatment arm $z$.\footnote{In the panel design, the researcher may randomly assign $Z_i$ in the same way, or may chose to show the information to all participants. If the panel design includes a treatment arm that receives no information, denote that arm with $B$. Since the panel design uses within-person contrasts, identification does not come from randomization across people. Thus it is sufficient to work with the realized signal $S_i$.}
Treatment is assigned randomly in the sense that $Z_i$ is independent of the potential outcomes: the outcome function $G_i(\cdot)$, the prior $X^0_i$, the potential signals $S_i(\cdot)$, and the learning rate $\alpha_i$. \footnote{While the treatment $Z_i$ will be randomly assigned, it is important to note that the realized signal $S_i(Z_i)$ can generally vary across individuals endogenously. In bottanBettingHouse22, $S_i(A)$ and $S_i(B)$ are high and low estimates of the home value and thus the realized signal is only randomly assigned conditional on the potential signals.} In passive designs, treatment arm $B$ does not receive any signal. For the sake of completeness, define $S_i(B) \equiv X^0_i$ in passive designs. It will be convenient to work with the following shorthand where potential beliefs are directly a function of the treatment assignment $z$. In a slight abuse of notation, we redefine
We will use this equation for potential beliefs along with the potential outcome equations (ref) and (ref) to study common empirical specifications.
The following three sections compare standard estimators to a local least squares (LLS) alternative. Standard estimators weight individuals by their belief updating; LLS weights all individuals equally. When belief updates are negatively correlated with causal effects, standard estimators understate the average effect. The current section begins by introducing the individual slopes, which are the causal building block of all the estimators considered in this paper, and then shows that standard panel and TSLS estimators recover weighted averages of these slopes.
Define the individual slope as the ratio of outcome change to belief change induced by the experiment:
This is the average rate of change in individual $i$'s outcome as beliefs move from $X_i(B)$ to $X_i(A)$, which are the individual-specific beliefs in treatment arms $A$ and $B$. Equivalently, this is the individual-specific average partial effect $G'_i(x)$ over the individual-specific interval of beliefs induced by the experiment. These individual slopes $\beta_i$ thus depend both on the individual response function $G_i(\cdot)$ and the variation in beliefs induced by the experiment $\bc{X_i(B), X_i(A)}$. In the panel design, define $X_i(A) \equiv X_{i1}$ and $X_i(B) \equiv X_{i0}$. The standard estimators used in the literature and the new LLS estimator aggregate these individual slopes differently. The differences between the parameters targeted by LLS and TSLS or panel estimators come entirely from differences in aggregation. The remainder of this section characterizes standard panel and TSLS estimators.
Standard estimators in information provision experiments yield weighted averages of individual effects $\beta_i$, with weights proportional to belief updating. In panels, individuals with below-average belief updates receive negative weights.
The precise form of these weights varies, but in all three cases, standard specifications weight individual effects $\beta_i$ in proportion to the first-stage belief updating. In all specifications, these weights integrate to one. Appendix (ref) contains derivations for all expressions in this section and Appendix (ref) provides a more general discussion of TSLS in information experiments.
We now examine three representative specifications and derive the implicit weights each places on different individuals.
armonaHomePrice19 use a regression in first-differences. Since there are only two time periods, this is equivalent to a panel regression with individual and time fixed effects. Let $\Delta X_i$ denote the difference between the post- and pre-treatment observations, $X_{i1} - X_{i0}$. The regression specification is simply
The regression of $\Delta Y_i$ on $\Delta X_i$ and a constant can assign negative weights to observations with $\Delta X_i$ between zero and the mean $\E[\Delta X_i]$.
\paragraph{Heterogeneity Bias Causes Negative Weights in Panel Regressions}
This negative weights result restates chamberlainMultivariateRegression82's classic (chamberlainMultivariateRegression82) \q{heterogeneity bias} as negative weights in a weighted average of individual effects. A closely related expression appears in Theorem 3.4c of callawayDifferenceinDifferencesContinuous25, who show that units with below-mean treatment intensity receive negative weights in difference-in-differences with continuous treatment. The panel regression here is analogous: it compares outcomes for big changers to small changers. Small changers act as the control group and their outcomes are subtracted from outcomes for big changers. Increasing the treatment effects of small changers thus decreases the slope estimate. This is what is means for them to have negative weights. Heterogeneity bias arises because these cross-update comparisons are contaminated by differences in treatment effects.
setteleHowBeliefs22 uses an IV specification where assignment to the \q{high} signal $T_i \equiv \1\bc{Z_i = A}$ is a binary instrument for beliefs. The estimand takes the canonical Wald form:
These weights are non-negative under learning rate updating with $\alpha_i \geq 0$ and in a general class of updating models when a monotonicity assumption holds such that $(X_i(A)-X_i(B))$ has the same sign for everyone.
cullenIncreasingDemand23 use an IV specification where the instrument is an indictor for assignment to the information treatment interacted with the initial gap in beliefs.\footnote{vz_aeri point out that similar specifications that also include the treatment indictor as an excluded instrument have negative weights.}
Since these specifications control for the exposure $S_i(A) - X_i^0$, the residual variation in the instrument is simply a re-centered version of the instrument.\footnote{To see this, notice that random assignment implies that $\E \bs{T^{ex}_i \mid S_i(A) - X_i^0} = \E \bs{T_i} \bp{S_i(A) - X_i^0} = \L \bs{T^{ex}_i \mid S_i(A) - X_i^0}$. By FWL $\wt{T}^{ex}_i \equiv T^{ex}_i - \L \bs{T^{ex}_i \mid S_i(A) - X_i^0} = (T_i - \E[T_i])(S_i(A) - X_i^0) $. }
The TSLS coefficient is then given by
These weights are non-negative under learning rate updating with $\alpha_i \geq 0$ and in a general class of updating models when monotonicity holds: $\text{sign}(X_i(A)-X_i(B)) = \text{sign}(S_i(A)-X_i^0)$.
The key takeaway from these expressions is that these standard specifications weight individual effects by the strength of belief updating. In the active and passive controls, weights are non-negative and thus are \q{weakly causal}.
This section presents a local least squares (LLS) estimator that recovers an equally weighted average of individual belief effects. With a linear outcome equation, these individual slopes have a structural interpretation as partial derivatives of the outcome with respect to beliefs and so the equally weighted average is the (structural) average partial effect (APE).
LLS is a control function estimator. It works by constructing a vector of controls that isolates the experimental variation in beliefs. In this setting, learning rate updating means that people who have the same prior, the same potential signals, and the same learning rate have the same potential beliefs; the only variation in their actual beliefs comes from the random assignment to the actual signal. The LLS approach aggregates many \q{local} regressions that use only this exogenous (i.e. experimental) variation in beliefs.\footnote{mastenIdentificationInstrumental16, grahamIdentificationEstimation12 show how to construct these \q{local} regressions in panel and IV settings more generally. I generalize their results from the linear random coefficients model to a more general nonparametric potential outcome model.}
The LLS estimator recovers equally weighted averages of individual slopes $\E \bs{\beta_i}$ by constructing local regressions that isolate purely experimental variation in beliefs. The ideal regression conditions on the potential beliefs $X_i(A)$ and $X_i(B)$, which isolates only the remaining variation in beliefs that comes from being assigned randomly to treatment $A$ or $B$. This ideal regression is:
This regression recovers a conditional average $\E \bs[\big]{\beta_i \mid X_i(A) = x_A, X_i(B) = x_B }$. Iterating expectations thus recovers the average individual slope $\E[\beta_i]$. This is an easily interpretable causal parameter: it answers the question, “On average, how much do outcomes change per unit change in beliefs, over the range of beliefs induced by the experiment?”.
The LLS estimation strategy also produces intermediate estimates $\E[\beta_i \mid \alpha_i]$ that reveal how causal effects vary with belief updating. Many behavioral models make strong predictions about the relationship between belief updating and belief effects mackowiakRationalInattentionForthcoming, yangDecisionRelevanceSubjective24, enkeBehavioralAttenuation24, fusterExpectationsEndogenous22. Section (ref) presents estimates of these conditional average slopes to document strong negative correlation between belief updates and causal effects across a range of settings.
The identification strategy in practice is then to condition on a set of controls that is as good as conditioning on the potential beliefs directly. The following section shows how to construct feasible local regressions.
The following sections show to construct feasible local regressions in the three experimental designs. Appendix (ref) provides proofs for the results in this section.
The panel approach works with any information treatment (including qualitative treatments or bundles of signals) because identification relies only on the panel structure, not on the content of the signal.
For any belief change $x \neq 0$:
The right hand side is a feasible local regression using only observations with $\Delta X_i = x$ or $\Delta X_i = 0$. Iterating over $x$ and averaging yields $\E[\beta_i]$. This requires that some individuals have (close to) zero change in beliefs.\footnote{This is an easily verifiable condition. It is satisfied if $P\bs{\Delta X_i = 0} > 0$, or more generally if $\Delta X_i$ has positive mass in any neighborhood around zero. See grahamIdentificationEstimation12 for detailed discussion of technical considerations with continuous $\Delta X_i$.}
Active designs rely on the Bayesian updating assumption (ref) and identify learning rates directly from observed belief updates: $\alpha_i = \bp{X_i - X_i^0}/\bp{S_i - X_i^0}$. Under Bayesian updating, people with the same learning rate, prior, and potential signals have the same potential beliefs; the only remaining variation comes from random assignment.
The control vector is $C_i \equiv \bs{\alpha_i \; X_i^0 \; S_i(A)\; S_i(B)}$. Conditional on $C_i = c$:
Iterating over $c$ and averaging yields $\E[\beta_i]$. The regression is feasible when $(S_i - X_i^0) \neq 0$ and $\var\bs{X_i \mid C_i = c} > 0$, which excludes cases with no learning ($\alpha_i = 0$) or identical signals ($S_i(A) = S_i(B)$)
Passive designs also rely on the Bayesian updating assumption (ref), but require additional assumptions because learning rates for the control group are unobserved. Consider two possible approaches to infer learning rates in the control group:
\paragraph{Case 1: Observed Prior Variance} In normal-normal Bayesian updating, $\alpha_i = {{\sigma^2_X}_i}/ \bp{{\sigma^2_X}_i + \sigma^2_S}$. If signal precision $\sigma^2_S$ is common across individuals, then conditioning on the rank of prior variance ${\sigma^2_X}_i$ is equivalent to conditioning on $\alpha_i$. The control vector becomes $C_i \equiv \bs{\text{rank}\bp{ {\sigma^2_X}_i} \; X_i^0 \; S_i(A)}$.
\paragraph{Case 2: Rich Observables} When researchers can predict beliefs from observables ballaelliott22,cantoniProtestsStrategic19, they can use predicted updates instead of observed updates. The implied predicted learning rate $\wt{\alpha}_i$ replaces the observed rate. The control vector becomes $C_i \equiv \bs{\wt{\alpha}_i \; X_i^0 \; S_i(A)}$.
In either case, under the linear outcome equation (ref) and Bayesian updating (ref):
Appendix (ref) formally states the assumptions in both of these cases.
The three experimental designs require progressively stronger assumptions to implement LLS. Panel designs impose no new behavioral assumptions. Active designs require Bayesian updating. Passive designs require Bayesian updating and also require either elicited prior variances or rich observables to infer unobserved learning rates.
The assumptions in the active case are weaker than in the passive case because in the active case researchers observe all participants update beliefs in response to new information. The experiment {reveals} heterogeneity in belief updating. In contrast, in a passive design, researchers need to use observables to {infer} heterogeneity in belief updating for a control group that the researcher never sees update their beliefs.\footnote{Recall that the learning rate is identified from the observed update $\alpha_i = \bp{X_i - X_i^0} / \bp{S_i(Z_i) - X_i^0}$, which is undefined for the passive control group that receives no information. Randomization is enough to ensure that the learning rates have the same distribution in both groups, but the individual learning rates are not directly identified in the passive control group. } This suggests that researchers interested in implementing an LLS estimator may find active designs more attractive since they reveal more information about belief updating.\footnote{There are many design considerations beyond the scope of this paper. haalandDesigningInformation23 discuss implementation considerations of active and passive control designs. listExperimentalistLooks25 discusses within- and between-subject experimental designs more generally.}
Conditioning on high-dimensional control vectors is often impractical in experimental samples. When belief updating is linear in the signal and prior, it is sufficient to control for $C_i$ semi-parametrically. The local regressions in between-person designs need only condition on the learning rate and can simply control linearly for the prior and signals in each local regression. In passive designs, or designs with person-specific high and low signals (i.e. rothRiskExposure22), it is also necessary to reweight by the inverse of the exposure. This weighted local regression recovers $\E \bs{\beta_i \mid \alpha_i}$. Appendix (ref) shows that this modified local regression is sufficient and Appendix (ref) provides general implementation guidance.
The estimators in Sections (ref) and (ref) target parameters that can be written as $\E[\beta_i \times \omega_i]$ for some weights $\omega_i$. The interpretation of these parameters depends on the interpretation of the individual slopes $\beta_i$, but the difference between estimators comes only from the weights $\omega_i$. LLS assigns equal weights. Under linearity, these equal weights deliver the APE, which has a structural interpretation that permits extrapolation. Appendix (ref) discusses the nonlinear case in greater detail.
With treatment effect heterogeneity, researchers must decide how to summarize heterogeneous effects. LLS recovers a simple average $\E[\beta_i]$. This equally weighted average $\E[\beta_i]$ answers the question: \q{On average, how much did outcomes change per unit change in beliefs?} Like the non-parametric ATE $\E[G'_i(X_i)]$, this parameter is local to the variation the experiment actually induced heckmanChapter7107a. A TSLS-weighted average may be policy-relevant when the intervention under consideration is information provision, since it captures effects among those whose beliefs would actually change.\footnote{If the policy question is whether to implement an information campaign, the reduced form (the effect of treatment assignment on outcomes) answers this directly.} However, the attenuation documented in Section (ref) suggests that relying on TSLS outside this narrow case is risky: researchers may conclude that belief effects are generally unimportant on the basis of an unrepresentative average.
In the linear outcome equation $G_i(x) = \tau_i x + U_i$, the parameter $\tau_i$ is structural: it fully characterizes $i$'s response to any hypothetical belief shift, not just those induced by the experiment. The average $\E[\tau_i]$ inherits this structural property: on average, a one-unit increase in beliefs causes an $\E[\tau_i]$-unit increase in outcomes, regardless of initial belief levels. This permits extrapolation; predictions for hypothetical interventions that shift beliefs by any amount can be formed by scaling the average effect appropriately.\footnote{Since $\tau_i$ is the individual partial effect, the average $\E[\tau_i]$ is also called the average partial effect (APE).}
This section demonstrates that attenuation due to dependence between belief updating and belief effect is empirically relevant. I compare standard panel and TSLS specifications to LLS estimates in six recent studies from leading economics journals .\footnote{ I searched the Web of Science database for papers in the top five economics journals, ReStat, AER: Insights, and all AEJs containing “beliefs,” “information,” or “perception” together with “experiment” or “treatment.” This yielded 116 potentially eligible experiments. I replicated the two most highly cited studies of each experimental design. To standardize the presentation of the results, I flip the sign of the outcome variable when necessary to ensure that mean effects are always positive. I also omit additional demographic controls and probability weights from all estimates for simplicity.} See Appendix (ref) for estimation details.
Table (ref) contrasts LLS estimate with estimates recovered by the standard specification in each study. In five of the six studies, standard estimators are substantially attenuated. Figure (ref) plots an estimate of conditional slopes for each study: $\E \bs{\beta_i \mid \abs{\Delta X_i}}$ in panel experiments (Panel A) and $\E \bs{\beta_i \mid \text{rank}(\alpha_i)}$ in active and passive control experiments (Panels B and C). These curves directly that people with the strongest causal effects tend to have smaller belief updates.
wiswallDeterminantsCollege15 study how beliefs about field-specific earnings affect college students' major choices. The panel estimate of 0.32 (s.e. 0.086) is substantially smaller than the LLS estimate of 0.721 (s.e. 0.33), with the LLS estimate being 125% larger. armonaHomePrice19 study how beliefs about home prices affect investment decisions. The panel estimate of 1.15 (s.e. 0.234) is smaller than the LLS estimate of 1.8 (s.e. 0.381), with the LLS estimate being over 50% larger.
setteleHowBeliefs22 studies how beliefs about the gender wage gap affect support for gender equality policies. The TSLS estimate of 0.096 (s.e. 0.033) is substantially smaller than the LLS estimate of 0.16 (s.e. 0.042), with the LLS estimate being 66% larger. rothRiskExposure22 study how recession expectations affect subjective personal unemployment risk. Their TSLS estimate of 0.755 (s.e. 0.433) is somewhat smaller than the LLS estimate of 0.882 (s.e. 0.379), with the LLS estimate being 17% larger.
kumarEffectMacroeconomic23 study how beliefs about GDP growth affect employment decisions. The TSLS estimate 0.466 (s.e. 0.19) is smaller than the LLS estimate 1.787 (s.e. 0.409), with the LLS estimate being 284% larger. cantoniProtestsStrategic19 study how beliefs about others' protest participation affect one's own willingness to participate. The TSLS estimate (0.68, s.e. 0.253) and the LLS estimate (0.18, s.e. 0.133) are both quite noisy, making it difficult to draw strong conclusions about the direction or magnitude of any difference. The difference between the TSLS and LLS estimates is suggestive evidence that people with larger belief effects had larger belief updates. However, the conditional effects in Panel C.ii of Figure (ref) reveals only modest variation across learning rate ranks, with quite wide confidence intervals.\footnote{Concerns of attenuation are only one reason among many to consider using the LLS estimator. The estimator consistently recovers the unweighted average regardless of the sign of dependence between belief updating and treatment effects. The pattern of attenuation observed in five of six applications is an empirical finding, not a mechanical feature of the estimator.}
The conditional effects in Figure (ref) reveal that individuals who update their beliefs the least have the strongest causal effects across many applications. This provides direct empirical support for models of endogenous information acquisition where people with decision-relevant beliefs invest in forming precise priors. Indeed, the mechanism where people are well informed about things that matter for their decisions and so respond less to new information is quite general and reflects basic features of rational inattention (Appendix (ref), see also mackowiakRationalInattentionForthcoming, fusterExpectationsEndogenous22, cavalloInflationExpectations17).
Standard empirical specifications in information provision experiments systematically understate the causal effects of beliefs on behavior. This paper demonstrates that in five of six high-profile studies in leading economics journals, ranging from college major choice to macroeconomic expectations, LLS estimates average effects of beliefs that are larger than estimates from standard specifications.
{ \singlespacing \printbibliography }