Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
79,381 characters · 15 sections · 64 citation commands
Identifying the Effect of Persuasion
\setcounter{footnote}{0}
\raggedbottom
How effectively one can persuade one's audience has been of interest to ancient Greek philosophers in the Lyceum of Athens,\footnote{See sep-aristotle-rhetoric for three technical means of persuasion in Aristotle's Rhetoric.} early--modern English preachers in St Paul's Cathedral,\footnote{See kirby2008public for historic details of the public persuasion at Paul's Cross, the open-air pulpit in St Paul's Cathedral in the 16th century.} and contemporary American news producers at Fox News in New York City.\footnote{dellavigna2007fox and martin2017bias measure the persuasive effects of slanted news using data on Fox News.} Recently, economists have been endeavoring to build theoretical models of persuasion kamenica2011,CDK,gentzkow2017,Prat,bergemann2017information and to quantify empirically the extent to which persuasive efforts affect the behavior of consumers, voters, donors, and investors dellavigna2010persuasion.
Since dellavigna2007fox proposed a measure of persuasion, it has been used and modified by many authors EPZ,gentzkow2011effect,DEMPZ,bassi2016persuasion,martin2017bias,CY2019 to quantify the persuasive effects of informational treatment. However, its precise interpretation has been obscure because of a lack of formal identification analysis. In fact, we show that the commonly used measure of persuasion does not estimate the causal rate of persuasion on any subpopulation in general. Therefore, it is misleading to call DK's measure the persuasion rate, although it is a common practice in the literature; instead, we will reserve the term of the persuasion rate for the population parameter that is properly defined by a conditional probability in the potential outcome framework.
The flaw of DK's measure arises from failing to distinguish a local average treatment effect imbens1994late from the the average treatment effect (ATE), where the difference between the two can be substantial for a heterogeneous population; specifically, DK's measure rescales a LATE with a factor that is relevant only for the ATE. For instance, in DK's example, even if Fox News has a high persuasive effect among those who will watch the channel if and only if it is available through a local cable package, it may not be persuasive at all when other people such as Democrats or Democrat-leaning independents are all included.
Focusing on the case of binary outcomes, we analyze the problem of measuring persuasive effects of informational treatment through the lens of the potential outcome framework. Specifically, we define the persuasion rate by a proper conditional probability of the agent taking an action of interest with a persuasive message given that the agent does not take the action without the conveyed message. We then formally study identification under a few empirically relevant scenarios of data availability. Our analysis will articulate what DK's measure of persuasion estimates, why it is misleading, and how we can fix the problem.
While the persuasion rate is concerned with the entire population (and hence related with the ATE), we also consider a local persuasion rate that focuses on the group of compliers. Our identification analysis shows that DK's measure estimates neither a local nor an average persuasion rate.
Our identification analyses are based on a few empirically relevant scenarios of data availability. We do this because the problem of data availability is particularly important in the context of measuring persuasive effects. For instance, individual-level partisan vote outcome data rarely exist due to confidentiality issues. Thus, the outcome variable is frequently measured only at an aggregate level. Also, it is not always the case to observe an actual exposure to a persuasive message in the same data set along with the outcome and instrument. Indeed, DK's analysis uses a micro dataset for the treatment and instrument and a separate aggregate dataset for the outcome and instrument. In order to address these challenges, we consider three different scenarios of data availability explicitly: given the instrument,
We obtain the sharp bounds on the persuasion rate for the entire population as well as other subpopulations of potential interest. Therefore, our work builds on the econometrics literature on partial identification Manski:03,Manski:07,Tamer:10 as well as the literature on program evaluation Heckman/Vytlacil:07,imbens2009recent.
The main findings of this paper can be summarized as follows. If there is no heterogeneity in the population, then DK's measure of persuasion estimates the rate of persuasion, provided that a simple monotonicity assumption is imposed; however, this case is an exception rather than the rule. Indeed, the rate of persuasion is only partially identified as an interval in general, where the sharp lower bound generally corresponds to DK's measure multiplied by the relative size of the complier group regardless of the specific data scenario. Therefore, DK's measure is strictly larger than the lower bound whenever there is partial compliance. It is also remarkable that the data scenarios matter only for the sharp upper bound. Therefore, the value of observing the treatment and outcome jointly only lies in obtaining a potentially more informative upper bound on the persuasion rate. We also investigate identification of the local persuasion rate (i.e.,\ the persuasion rate for the group of compliers) under the same three data scenarios. It is point-identified under the most favorable data scenario, but only partially identified under the other scenarios; even in the case of point identification, DK's measure generally differs from the local persuasion rate. If a continuous instrument is available, then we can target a marginal persuasion rate that is akin to the marginal treatment effect Heckman/Vytlacil:05. Therefore, having a continuous instrument opens up the possibility of point identification of the persuasion rate for a policy-relevant population if the instrument is sufficiently rich.
In order to illustrate our findings, we discuss two empirical examples in the main text, while we provide a few more in the online appendices. First, we revisit CY2019, where the interest is in the effect of Chinese students having access to uncensored media on behaviors, beliefs and attitudes; we use the same variables and setup as in the original paper for this exercise. Second, we analyze the voting behavior and newspaper readership using the data from gerber2009does. Overall, we show that DK's measure of persuasion tends to overstate the persuasive effects while masking underlying heterogeneity.
The remainder of the paper is organized as follows. In Section (ref), we recall some backgrounds and specifics of DK's measure of persuasion. Here, we properly define the persuasion rate at the population level by using the potential outcome framework. In Section (ref), we discuss identification of the persuasion rate. In Section (ref), we study the local and marginal versions of the persuasion rate. In Section (ref), we discuss our recommendations on what to do in practice, including practical inferential issues,\footnote{Stata commands for estimation and inference based on this paper's identification results are publicly available at \url{https://github.com/persuasio}. Alternatively, they can be installed from within Stata by typing “ssc install persuasio”. } and clarify the difference between a population version of DK's measure of persuasion and the persuasion rate. In Section (ref), we provide two empirical illustrations. We conclude in Section (ref). The online appendices include additional results and examples, including an extension to non-binary outcomes, a detailed discussion about methods for inference, and all of the proofs.
It is helpful to recall DK as a prototypical example, where they study the effect of an exposure to Fox News on the probability of voting for a Republican presidential candidate. Here, the informational treatment of interest is the viewership of the Fox News channel, and the persuasion rate of the media can be thought of as the proportion of the Fox News viewers who voted for a Republican candidate among those who would not have done so if they had not watched Fox News. It should be noted that the agents' decisions about whether to watch Fox News or not may be correlated with their political orientation. In order to address the endogeneity issue, DK use Fox News availability via local cable in the year of 2000 as an instrumental variable.
In efforts to measure the persuasion rate as explained above, DK and dellavigna2010persuasion propose the (infeasible) estimand $f$ defined as follows: for a binary outcome and by using DK's notation
where $\mathbb{T}$ and $\mathbb{C}$ represent an instrument assignment status such as having Fox News available via local cable or not.\footnote{Therefore, $\mathbb{T}$ and $\mathbb{C}$, which appear to denote treatment and control groups, should be understood as the status of the intent to treat, not the actual treatment status.} Here, for $j\in \{ \mathbb{T}, \mathbb{C} \}$, $y_j$ is the share of group $j$ taking the action of interest (e.g.,\ voting for a Republican candidate), and $e_j$ is the share of group $j$ exposed to a persuasive message. Further, $y_0$ is the share of those who would take the action of interest without listening to the persuasive message. So, $f$ is a rescaled version of the usual Wald statistic that estimates the LATE, where rescaling is apparently to obtain a “rate” that focuses on those who are to be persuaded. As DK noted, $y_0$ involves a counterfactual that is often unobserved. For this reason, $f$ is generally not a feasible estimand, and DK propose using $y_\mathbb{C}$ in place of $y_0$ as an approximation, which yields a feasible estimand, say $\tilde f$. We will refer to $f$ or its feasible approximation $\tilde f$ by DK's measure of the persuasion rate.
Without a formal justification, it is common in the literature to interpret $f$ as a conditional probability. For instance, DK explain, on page 1218 of their paper, “The key parameter is $f$, the fraction of the audience that is convinced by Fox News to vote Republican.” Similar interpretations are prevalent in the literature, as the following quotations demonstrate:
Also, in their survey, dellavigna2010persuasion use $\tilde f$, i.e., a feasible approximation of $f$, as a key summary statistic to compare persuasive effects across different studies.
However, as we mentioned in the introduction, neither $f$ nor $\tilde f$ estimates the persuasion rate, i.e., the intended conditional probability, on any subpopulation in general. Therefore, it is misleading to label $f$ or $\tilde f$ as a persuasion rate, and it is generally invalid to make comparisons across different studies. Indeed, $f$ or $\tilde f$ may not even be a proper conditional probability in a heterogeneous population: e.g., the approximation $\tilde f$ can even be larger than one. We will articulate under what assumptions $\tilde f$ turns out to be a proper conditional probability, and how we should interpret it in its relationship with the persuasion rate.
In order to facilitate our discussion, we start with formally defining the persuasion rate at the population level by using the potential outcome framework. Let $T_i$ denote the binary indicator that equals $1$ if individual $i$ is exposed to persuasive information such as Fox News. Let $Y_i(t)$ be a binary indicator, which shows agent $i$'s potential action when $T_i$ is set to $t \in \{0,1\}$. For example, $Y_i(1)$ equals $1$ if individual $i$ votes for a Republican candidate after watching Fox News. The econometrician never observes both $Y_i(0)$ and $Y_i(1)$ but can only observe either of the two: that is, $Y_i = T_iY_i(1) +(1-T_i) Y_i(0)$.\footnote{So, both $Y_i$ and $T_i$ are binary. In online (ref), we extend our results to the case where the potential outcomes are multinomial.} Then, the fraction of the people who take the action of interest with an exposure to a persuasive message, among those who would not without it, is given by
provided that the conditional probability is well-defined: $\theta_\textrm{pr}$ is the persuasion rate at the population level. Using conditional probability is to rule out the case of “preaching to the converted”;\ if $Y_i(0) = 1$, then those individuals are already “persuaded” to take the action of interest even without the persuasive treatment, and therefore we do not count them in defining the persuasion rate.\footnote{The idea of using conditional probability to define a parameter of interest can also be found in heckman1997making, though their context is quite different from ours.}
The common estimand $\tilde f$ (or even $f$) does not estimate $\theta_\textrm{pr}$ in a heterogeneous population; in fact, even in a homogeneous population, $\tilde f$ or $f$ cannot be (asymptotically) equated with $\theta_\textrm{pr}$ without an extra monotonicity assumption. However, the rescaled quantity $(e_\mathbb{T} - e_\mathbb{C} ) \tilde f $, which is always no greater than $\tilde f$, does provide valid information about $\theta_\textrm{pr}$ in that it corresponds to the sharp lower bound of the identified interval of $\theta_\textrm{pr}$ in general. The best way to clarify all the issues is to conduct a rigorous analysis on the identification of $\theta_\textrm{pr}$, which is our next topic. For quick takeaways, see (ref).
Identification of $\theta_\textrm{pr}$ is challenging for various reasons, including the fact that $\theta_\textrm{pr}$ depends on the joint distribution of the potential outcomes and that $T_i$ can be endogenous, and it is often difficult to observe $T_i$ and $Y_i$ jointly. To allow for potential endogeneity, we use a binary instrument, $Z_i$, throughout the paper unless otherwise specified. Exogenous covariates, $X_i$, could be observed, but we suppress $X_i$ from our identification analysis. In other words, we implicitly assume throughout the paper that all assumptions and results are conditional on the value of $X_i$. Therefore, the observed variables are $Y_i, T_i, Z_i$, all of which are binary throughout the main text. See (ref) for how to deal with $X_i$ in practice. Also, see (ref) for an extension to the case of multinomial outcomes.
The data issue on $T_i$ is addressed by considering three scenarios of data availability. Specifically, for the purpose of the identification analysis, we assume that for $(y,t,z)\in \{0,1\}^3$, (i) $\mathbb{P}(Y_i = y, T_i = t\mid Z_i=z)$ is known, (ii) $\mathbb{P}(Y_i = y\mid Z_i=z)$ and $\mathbb{P}(T_i = t\mid Z_i=z)$ are known, or (iii) $\mathbb{P}(Y_i = y\mid Z_i=z)$ is all that is known, depending on the specific scenario of interest. For example, DK use town-level election data to estimate $\mathbb{P}(Y_i = y \mid Z_i=z)$ and micro-level media audience data to infer $\mathbb{P}(T_i = t\mid Z_i=z)$, which corresponds to Case (ii).\footnote{Throughout the discussion, we assume that $T_i$ is correctly measured if it is observed. See latewithmismeasurement, nguimkeu2016estimation, and Ura2018miclassification for the issues of mismeasured treatment. Their subject matter is distinct from ours.}
It requires an additional assumption to address the challenge that $Y_i(1)$ and $Y_i(0)$ are never simultaneously observed while $\theta_\textrm{pr}$ depends on their joint distribution. Before we present our identification results, we discuss our key assumptions in the following subsection.
Our first key assumption is that the persuasive message is directional, which will be important to decouple $\theta_\textrm{pr}$ by the marginals of the potential outcomes.
(ref) is a binary version of the monotonic treatment response (MTR) assumption used in manski1997monotone and manski2000mtr.\footnote{In online Appendix A, we present a simple economic model that motivates (ref).} (ref) allows $Y_i(0) = Y_i(1)$ with probability one, and therefore it does not rule out the possibility that `watching Fox News' has no impact on the agent's behavior at all. The inequality in (ref) means that the messages Fox News delivers are biased or directional in favor of Republican candidates; i.e.,\ if a voter is going to vote for a Republican candidate without watching Fox News, then watching Fox News will not change that. In other words, (ref) rules out the possibility that the level of distrust that a voter has on Fox News is so high that she takes actions based on the opposite of the messages Fox News delivers.
Since (ref) is a key assumption in the paper, we first clarify how much we can hope for with and without (ref).
(ref) is a consequence of the Fr\'{e}chet--Hoeffding inequality on the probability of a joint event. Since the two potential outcomes are never observed simultaneously, all we can ever hope to identify is their marginals and (ref) expresses the (sharp) bounds on $\theta_\textrm{pr}$ in terms of the marginal probabilities of the potential outcomes.
Suppose that there are some voters who have such a high level of distrust on Fox News that rather not watching Fox News would help them to take favorable actions to a Republican candidate but that those voters are only minority and on average we still have a strict stochastic dominance relationship between $Y_i(1)$ and $Y_i(0)$ (i.e., $\mathbb{P}\{ Y_i(1) = 1 \} > \mathbb{P}\{ Y_i(0) = 1 \}$). Then, the lower bound on $\theta_\textrm{pr}$ will be ensured to be nontrivial.
(ref) is stronger than the stochastic dominance, but it delivers a stronger result. In fact, (ref) is necessary and sufficient to express $\theta_\textrm{pr}$ in terms of the marginal probabilities of the counterfactual outcomes. In this case, the conditional probability $\theta_\textrm{pr}$ is the average treatment effect (ATE) divided by $\mathbb{P}\{ Y_i(0) = 0\}$. Throughout the rest of the paper, we present most of our results by using (ref) because it is not only convenient but also being biased or directional seems to be the nature of persuasive effort. However, we emphasize that $\theta_\textrm{avg}$, the rescaled version of the ATE, is always a valid lower bound on $\theta_\textrm{pr}$ as (ref) shows.
The next assumption is concerned about the treatment assignment $T_i$ and the instrument $Z_i$.
(ref) is standard for causal inference using instrumental variables. The intent-to-treat (ITT), $Z_i$, is randomly assigned; however, $T_i$ can be endogenous via the dependence between $V_i$ and $Y_i(t)$. The function $e(\cdot)$ is the propensity score or, more descriptively in our context, it can be referred to as the exposure rate.
If $T_i(z)$ denotes counterfactual treatment when a binary instrument takes value $z$, then agent $i$ is called a complier when $T_i(1) = 1$ and $T_i(0)=0$; never-takers ($T_i(z) = 0$ for all $z$), always-takers ($T_i(z)=1$ for all $z$), and defiers ($T_i(1) = 0, T_i(0) = 1$) are similarly understood. Under (ref), $i$ is a complier if and only if $e(0)<V_i\leq e(1)$. Similarly, $i$ is an always-taker when $V_i\leq e(0)$, and she is a never-taker when $V_i> e(1)$. Therefore, under (ref), there are only three groups in the population: i.e.,\ always-takers, never-takers and compliers. Indeed, as vytlacil2002independence has shown, the threshold structure in (ref) is equivalent to assuming the absence of defiers, which is a popular assumption in econometrics to identify the local average treatment effect.
In the following two subsections, we present identification results for $\theta_\textrm{pr}$, which is the same as $\theta_\textrm{avg}$ under (ref). (ref) covers the simplest case, where everybody complies so that there is no difference between the actual treatment and the intent to treat (ITT), i.e., $T_i = Z_i$ for each $i$: this case is referred to as the sharp persuasion design, where there is no distinction among different data scenarios. In this case, not surprisingly, $\theta_\textrm{avg}$ is point identified from the distribution of $Y_i$ given $Z_i$. However, when $T_i$ and $Z_i$ are different, which we call the fuzzy persuasion design, we have only partial identification of $\theta_\textrm{avg}$, where each of the three data scenarios become relevant. Before we move on, we define
provided that $\mathbb{P}(Y_i = 1\mid Z_i=0)<1$: $\theta_L$ is an identified parameter from the distribution of $Y_i$ given $Z_i$. It turns out that first, $\theta_L$ is equal to $\theta_\textrm{avg}$ under the sharp persuasion design, and second, it is the sharp lower bound on $\theta_\textrm{avg}$ in the fuzzy persuasion design regardless of which of the three data scenarios applies.
The condition of $e(1) - e(0) = 1$ means that everybody is a complier, and hence there is essentially no difference between $T_i$ and $Z_i$; thus, the sharp design is equivalent to a situation where $T_i$ is observed and randomized. However, this is rather an exceptional situation in social sciences. The key identification question should be how far we can go when the design is not sharp (i.e.,\ not everybody is a complier). We answer this question in the following subsection.
Without (ref), the general sharp identified bounds on $\theta_\textrm{pr}$ are given by
Therefore, even in the sharp persuasion design, $\theta_\textrm{pr}$ is only partially identified and its sharp bounds can be trivial without (ref). Even so, it is worth noting that $\theta_L$ remains a valid lower bound.
In the fuzzy design, the three scenarios of data availability we mentioned earlier become pertinent.
Even the full joint distribution of $(Y_i,T_i,Z_i)$ does not point-identify the ATE. Recall that those with $Z_i = 0$ and $T_i =0$ comprise the compliers and the never-takers, while those with $Z_i = 0$ and $T_i = 1$ are the always-takers. Similarly, those with $Z_i = 1$ and $T_i = 0$ are the never-takers, while those with $Z_i = 1$ and $T_i=1$ consist of the compliers and the always-takers. Therefore, these four cases correspond to different subpopulations, and the only subpopulation that we can study for both $T_i=0$ and $T_i = 1$ in common is that of compliers, which explains why the Wald statistic estimates the LATE, not ATE. For the same reason, $\theta_\textrm{avg}$ cannot be point identified; however, we can derive its sharp bounds.
To prove (ref), we first derive the sharp identified bounds for $\mathbb{P}\{Y_i(1) = 1 \}$ and $\mathbb{P}\{Y_i(0) = 1 \}$ separately; we denote them by the intervals $[m_a, M_a]$ and $[m_b,M_b]$, respectively. These bounds are special cases of manski2000mtr under the MTR assumption coupled with the exogeneity of the instrument. Then, letting $a := \mathbb{P}\{ Y_i(1) = 1 \}$ and $b := \mathbb{P}\{ Y_i(0) = 1 \}$, we obtain the upper bound of the identified interval of $\theta_\textrm{avg}$ by solving
while the lower bound can be found by doing minimization instead of maximization. We then appeal to continuity and the intermediate value theorem for the sharpness result.
It is proved in online (ref) that $m_a = \mathbb{P}(Y_i=1\mid Z_i = 1)$ and $M_b = \mathbb{P}(Y_i=1\mid Z_i = 0)$. Therefore, an examination of (ref) reveals that the lower bound is attained when $a = m_a$ and $b = M_b$. To develop intuition behind (ref), we discuss what $a=m_a$ and $b = M_b$ means, for which the behavior of the never-takers and that of the always-takers matter. Since $m_a = \mathbb{P}\{Y_i(1) = 1, T_i = 1 \mid Z_i=1\} + \mathbb{P}\{Y_i(0) = 1, T_i=0 \mid Z_i=1\}$, we know that $a = m_a$ holds when $\mathbb{P}\{ Y_i(0)=1, T_i = 0\mid Z_i = 1\} = \mathbb{P}\{ Y_i(1)=1, T_i = 0\mid Z_i = 1\}$. Here, the event $Z_i = 1, T_i=0$ means that $i$ is a never-taker, because defiers are assumed to be non-existent. Therefore, $a=m_a$ means that the treatment has no effect on the group of never-takers, unless there are no never-takers at all. Similarly, $b = M_b$ holds when $\mathbb{P}\{Y_i(1)=1,T_i=1 \mid Z_i=0 \} = \mathbb{P}\{ Y_i(0)=1,T_i=1 \mid Z_i = 0 \}$. Since $Z_i = 0, T_i = 1$ means that $i$ is an always-taker, we know that $b=M_b$ holds when the treatment does not affect the behavior of the always-takers, unless there are no always-takers at all. Therefore, $\theta_\textrm{avg}$, which is the same the persuasion rate $\theta_\textrm{pr}$ under (ref), is smallest when there are null treatment effects for both the never-takers and the always-takers.
Intuition for the upper bound can also be obtained by considering the non-complier groups. The upper bound corresponds to the case where $a = M_a$ and $b = m_b$, where it is shown in (ref) that $M_a = \mathbb{P}(Y_i = 1, T_i=1 \mid Z_i=1)+1-e(1)$ and $m_b = \mathbb{P}(Y_i=1,T_i=0 \mid Z_i=0)$. Here, note that $a=M_a$ is equivalent to $\mathbb{P}\{Y_i(1)=0,T_i=0 \mid Z_i=1 \} =0$, and $b=m_b$ is to $\mathbb{P}\{ Y_i(0) = 1, T_i = 1 \mid Z_i=0\} = 0$. Therefore, we can see that the persuasion rate $\theta_\textrm{avg}$ equals the upper bound when every never-taker has $Y_i(1) = 1$ and none of the always-taker has $Y_i(0)=1$: e.g.,\ all those who never watch Fox News (whether it is available or not) would actually have voted for a Republican candidate if they had watched it and all those who always watch Fox News would not have voted for a Republican without watching the channel.
The bounds in (ref) shrink to a singleton as $\bigl( e(0), e(1) \bigr)$ approaches $(0,1)$, which is consistent with the result in (ref). Also, it is worth noting that the lower bound $\theta_L$ only depends on the distribution of $(Y_i, Z_i)$: observing $T_i$ along with $(Y_i, Z_i)$ helps only for the upper bound. If $e(1)$ is too small, then the upper bound will not be very informative: $\theta_U$ converges to $1$ as $e(1)$ approaches $0$; that is, if nobody reads a newspaper when they receive free subscriptions, then we do not learn much about how “persuading” the newspaper is. However, even if $e(1)$ approaches $1$, the upper bound does not necessarily shrink to the lower bound; e.g.,\ we do not necessarily pin down the persuasion rate of reading the newspaper even if everybody who has free subscriptions actually reads it.
We now establish partial identification of $\theta_\textrm{pr}$ without (ref), i.e., $\theta_\textrm{pr} \neq \theta_\textrm{avg}$, for the sake of completeness. Let $NT = \{V_i > e(1)\}$ and $AT = \{ V_i\leq e(0)\}$ be the event of $i$ being a never-taker and an always-taker, respectively.
(ref) shows that $\theta_L$ continues to be the sharp lower bound, provided that $\theta_L \geq 0$, i.e., $\mathbb{P}(Y_i = 1\mid Z_i = 1) \geq \mathbb{P}(Y_i = 1\mid Z_i = 0)$, and a stochastic dominance condition holds for the never-takers and always-takers. However, the upper bound will be larger than that in (ref) in general.
As in the case of DK, the researcher may not directly observe $T_i$ along with $(Y_i,Z_i)$ but may have auxiliary data from which the exposure rates $e(1)$ and $e(0)$ can be estimated.\footnote{The case in which the outcome and the treatment are separately observed belongs to an identification problem called the ecological inference problem. For instance, CrossManski and Manski2017longshort discuss bounding a “long regression” by using information from a “short regression”. Their substantive concerns are distinct from ours.} In this case, the sharp identified bounds on $\theta$ become generally wider than those of (ref).
Therefore, the upper bound in this case is nontrivial if and only if $e(1) > \mathbb{P}(Y_i = 1 \mid Z_i = 1)$.\footnote{The trivial case can occur in applications. See (ref) for such cases.} Note that it is the relative size of the take-up rate $e(1)$ (e.g.,\ the probability of reading a newspaper when a free subscription to it is offered) that determines how much we can hope to learn about the persuasion rate. For example, if the probability of watching Fox News is too small relative to the probability of voting for a Republican candidate when Fox News was introduced in the local cable, then it becomes difficult to pin down how successfully Fox News persuaded their audience to vote for a Republican candidate. Also, it is worth noting that $e(0) = 0$ is not uncommon as (ref) and (ref) show. In this case, the maximum in the expression of the upper bound is unnecessary. Intuition for the lower bound is the same as the case of (ref), because $\theta_L$ requires only the distribution of $Y_i$ given $Z_i$.
The final scenario is the least informative one, where $T_i$ is not observed at all. This is an almost trivial case, but we state it in a separate theorem for the sake of completeness.
The lower bound from (ref) depends only on the distribution of $(Y_i,Z_i)$, and therefore $\theta_L$ continues to be the lower bound in this case as well. Further, since no information for $e(1)$ and $e(0)$ is available, it suffices to note that the upper bound in (ref) equals one whenever $e(0)>\mathbb{P}(Y_i=1\mid Z_i=0)$ and $e(1)<\mathbb{P}(Y_i=1\mid Z_i=1)$.
In this section, we consider rates of persuasion on other subpopulations that have been considered in econometrics: i.e.,\
provided that the conditional probabilities are well-defined: $\theta_\textrm{local}$ is the persuasion rate for the compliers imbens1994late, whereas $\theta_\textrm{marginal}(v)$ is for the subpopulation such that $V_i = v$ Heckman/Vytlacil:05.
First, we obtain identification results for $\theta_\textrm{local}$ under the three sampling scenarios in the fuzzy persuasion design. The first step for this purpose is to note that the same reasoning as (ref) yields
where the numerator is the LATE, which has received great attention in the econometrics literature Deaton2010, Heckman2010, Imbens2010. The denominator that rescales the LATE is also conditioned on the same subpopulation of the compliers.
Recall that the identification of the LATE requires the joint distribution of $(T_i, Z_i)$ and that of $(Y_i, Z_i)$ separately but not the full joint distribution of $(Y_i, T_i, Z_i)$. Unlike the LATE, the point identification in (ref)(ref) demands the knowledge of the joint distribution of $(Y_i, T_i, Z_i)$.\footnote{The denominator of (ref) requires that we know the marginal distribution of $Y_i(0)$ for the compliers. imbens1997estimating show that the marginal distributions of $Y_i(1)$ and $Y_i(0)$ for the compliers are identified if the joint distribution of $(Y_i, T_i, Z_i)$ is known; however, they did not consider the local persuasion rate.} (ref)(ref) shows that this requirement is not only sufficient but also necessary to achieve the point identification of $\theta_\textrm{local}$.
Just like the LATE, it may be contentious whether or not $\theta_\textrm{local}$ should be the parameter of interest, because the compliers are concerned with an unidentified subgroup of the population. However, we take a practical view that the identification results on $\theta_\textrm{local}$ can complement the results obtained in (ref).
The local persuasion rate $\theta_\textrm{local}$ represents the average persuasive effect for a population that is different from the entire population. Given this caveat, it is interesting to note that, in (ref)(ref), the upper bound on $\theta_\textrm{local}$ is always trivial in contrast to $\theta_\textrm{avg}$, but the lower bound of $\theta_\textrm{local}$ can never be worse than that of $\theta_\textrm{avg}$. Therefore, in principle, the length of the identified interval of $\theta_\textrm{avg}$ can be smaller than that of $\theta_\textrm{local}$. If $T_i$ is not observed at all, then there is no advantage in focusing on the compliers. (ref)(ref) confirms the intuition that the bounds for $\theta_\textrm{local}$ are identical to those for $\theta_\textrm{avg}$ if the distribution of $(Y_i, Z_i)$ is the only piece of information available. This corresponds to an uninteresting case for $\theta_\textrm{local}$ though, as we have no information on compliers.
Data requirements for the identification of $\theta_\textrm{marginal}(v)$ are generally quite demanding: e.g.,\ a continuous instrument is needed. However, identifying $\theta_\textrm{marginal}(v)$ for various values of $v$ can open up the possibility of point identification of $\theta_\textrm{avg}$. Therefore, it is worth understanding what is sufficient for the identification of $\theta_\textrm{marginal}$.
If $Y_i$ and $T_i$ are jointly observed along with a continuous instrument $Z_i$, then $\theta_\textrm{marginal}(v)$ can be point identified as in Heckman/Vytlacil:05 and carneiro2011. Examples of continuous instruments can be found in the literature on the media effects on voting. For instance, EPZ and DEMPZ use the signal strength of NTV and Serbian radio as instruments, respectively; in both of the papers, $(Y_i, T_i, Z_i)$ are jointly observed. The following assumption describes the situation in which we can obtain point identification of $\theta_\textrm{marginal}(v)$. We use the standard results in the literature Heckman/Vytlacil:05 for the subsequent theorem.
Similarly to the case of $\theta_\textrm{avg}$ or $\theta_\textrm{local}$, (ref) enable us to rewrite $\theta_\textrm{marginal}(v)$ as $\mathbb{E}\{ Y_i(1) - Y_i(0) \mid V_i = v \} / \mathbb{P}\{Y_i(0) = 0\mid V_i = v\}$; i.e.,\ $\theta_\textrm{marginal}(v)$ is a rescaled version of the marginal treatment effect of Heckman/Vytlacil:05. (ref) is a direct consequence of that.
(ref) does not consider the other two scenarios of data availability. This is mainly because continuous instruments are relatively infrequent in the context of persuasion, and we are not aware of any applications where continuous instruments are available while the outcome and treatment are not jointly observed.
If the support of the exposure rate $e(Z_i)$ is equal to the unit interval $[0,1]$, then (ref) shows the identification of $\theta_\textrm{marginal}(v)$ for all $v$ in the unit interval. Then, we can use $\theta_\textrm{marginal}(v)$ to construct different policy-oriented quantities as in Heckman/Vytlacil:05 and carneiro2011. For instance, the persuasion rate of the entire population can be obtained by $\int_0^1 \theta_\textrm{marginal}(v) dF\{v\mid Y_i(0) = 0\}$, which is equal to
by Bayes' theorem.
In this section we articulate the relationships between $\theta_\textrm{avg}, \theta_\textrm{local}$, and DK's measures $f$ and $\tilde f$ defined in (ref), and we summarize the main takeaways of our identification results. We assume that (ref) hold throughout this section, so we have $\theta_\textrm{pr} = \theta_\textrm{avg}$. Also, in order to be consistent with the identification analysis, we work with the population versions of $f$ and $\tilde f$: i.e.,\
First, neither $\theta_{DK}$ nor $\tilde \theta_{DK}$ is generally equal to the persuasion rate $\theta_\textrm{avg}$, or that for the compliers, $\theta_\textrm{local}$, which is clear from the fact that \[ \mathbb{P}\{ Y_i(0)=0 \} \theta_{DK} = \mathbb{P}\{ Y_i=0\mid Z_i = 0 \} \tilde \theta_{DK} = \mathbb{P}\{Y_i(0) = 0\mid e(0)<V_i\leq e(1) \} \theta_\textrm{local} \] is equal to the LATE, while $\mathbb{P}\{ Y_i(0) = 0\} \theta_\textrm{avg}$ is equal to the ATE.\footnote{Recall that the population version of the Wald statistic is equal to the LATE under (ref).} For example, $\theta_{DK}$ rescales the LATE with an unconditional probability and hence it does not render a well-defined conditional probability in general. Similarly, $\tilde\theta_{DK}$ is not necessarily a `rate' in spite of the rescaling factor. So, making comparisons across different studies based on $\theta_{DK}$ or $\tilde\theta_{DK}$ can be misleading, although it is a common practice dellavigna2010persuasion.
There are some special cases of exception though:
That is, in cases ((ref)) and ((ref)), there is no endogeneity issue, whereas in case ((ref)), no one is affected by the persuasive message (i.e.,\ $Y_i(1) - Y_i(0) = 0$ for all $i$), or everybody is persuaded (i.e.\ $Y_i(1) - Y_i(0) = 1$ for all $i$). Note that the LATE is the same as the ATE when any of the three cases applies. As a result, we have that in case ((ref)), $\theta_{DK} = \tilde\theta_{DK} = \theta_\textrm{avg} = \theta_\textrm{local}$ holds; in case ((ref)), $\theta_{DK} = \theta_\textrm{avg} = \theta_\textrm{local} \leq \tilde\theta_{DK}$; and, in case ((ref)), $\theta_{DK} = \theta_\textrm{avg} = \theta_\textrm{local} = 1\leq \tilde\theta_{DK}$ if everybody is persuaded, or $\theta_{DK}=\tilde\theta_{DK} =\theta_\textrm{avg} = \theta_\textrm{local} = 0$, unless any of them are ill-defined, if no one is affected by the informational treatment. Therefore, the feasible version $\tilde f$ of DK's measure of persuasion does estimate the persuasion rate $\theta_\textrm{avg}$ in case ((ref)), although, as DK correctly pointed out, $\tilde\theta_{DK}$ does approximate $\theta_{DK}$ in case ((ref)) as well if either $e(0)$ or $\theta$ is close to zero.\footnote{In case ((ref)), we have $ \mathbb{P}(Y_i = 1 \mid Z_i = 0) = \mathbb{P}( Y_i = 1, T_i = 1 \mid Z_i = 0 ) + \mathbb{P}( Y_i = 1, T_i = 0\mid Z_i = 0 ) = \mathbb{P}\{ Y_i(0) = 1\} + \bigl[ \mathbb{P}\{ Y_i(1) = 1\} - \mathbb{P}\{ Y_i(0) = 1\} \bigr] e(0)$.}
Our identification results in the previous sections show that $\theta_L$ is the sharp lower bound of the identified interval of $\theta_\textrm{avg}$ in the fuzzy design regardless of whether the full joint distribution of the outcome, treatment, and instrument is available or not. The parameter $\theta_L$ has been reported in the literature without understanding that it is the sharp lower bound on $\theta_\textrm{avg}$. For instance, dellavigna2010persuasion extensively estimate $\tilde\theta_{DK}$ by using many examples but they report $\theta_L$ as a lower bound of $\tilde\theta_{DK}$ when $T_i$'s are unobserved and hence $e(1)$ and $e(0)$ are unknown. Our results show that $\theta_L$ is always a meaningful parameter, but $\tilde\theta_{DK}$ may not. Therefore, even when information about $e(0)$ and $e(1)$ is available, $\theta_L$ is a better parameter to estimate than $\tilde\theta_{DK}$.
Indeed, if the full joint distribution of $(Y_i,T_i,Z_i)$ is available, then we recommend reporting $[\theta_L,\ \theta_U]$ along with $\theta^*$; these can be consistently estimated by their sample analogs. If $(Y_i,Z_i)$ is observed with some auxiliary information for $e(0)$ and $e(1)$ available, then $[\theta_L,\ \theta_{U_e}]$ and $[\theta_L^*,\ 1]$ should be reported. If $T_i$ is not observed at all, then the interval $[\theta_L,\ 1]$ is the best we can hope for to study either $\theta_\textrm{avg}$ or $\theta_\textrm{local}$.
Note that $\theta_L$ should be estimated all the time; it only requires data on $(Y_i,Z_i)$. Because the actual $T_i$ can be difficult to observe, researchers have used an extra micro-level survey to obtain auxiliary data on $T_i$, which seems quite costly. However, the value of an attempt to observe $T_i$ can be limited, depending on which parameter the researcher wants to learn about. For instance, if the researcher cares about the persuasion rate of the entire population, then observing $T_i$ does not add any information for the lower bound, while it can potentially improve the upper bound. If the group of compliers is of interest, then whether we observe $T_i$ or not, and how we observe it, can be relevant issues; we have $\theta_L^*\geq \theta_L$ in the second data scenario and $\theta^*$ is point identified if $(Y_i, T_i, Z_i)$ is jointly observed. If $Z_i$ is continuously distributed, the value of observing $(Y_i, T_i, Z_i)$ jointly increases dramatically as well. In summary, our identification analysis shows that the value of observing $T_i$ depends crucially on which population the researcher is interested in.
In order to illustrate the difference between DK's measure and our bounds, we have calculated them in (ref). We focus on the results reported in dellavigna2010persuasion when the outcome variable is voter turnout. We have chosen this type of study as the turnout is among the most studied outcome variables in the literature and it is naturally a binary measure. (ref) provides estimates of $\mathbb{P}(Y_i=1 \mid Z_i=z)$ and $e(z)$ for $z=0,1$, thereby enabling us to obtain the bounds based on (ref) and (ref) (ii). It can be seen that using DK's persuasion rates alone may lead to misleading conclusions because the bounds on $\theta_\textrm{avg}$ as well as those on $\theta_\textrm{local}$ are in fact wide. Moreover, the results in (ref) suggest that identification power under (ref) in this example is limited, especially for the upper bounds on $\theta_\textrm{avg}$ and $\theta_\textrm{local}$. We further illustrate these points with empirical examples in (ref).
Finally, since the parameters are partially identified, inference should also account for that. The method proposed by Stoye:07 is useful for that purpose, at least in the most favorable data scenario, in which case the sample analog principle and the delta method show that we can construct the estimators $\hat \theta_L$ and $\hat \theta_U$ that are asymptotically jointly normal. Therefore, by Stoye:07, a $(1-\alpha)$ confidence interval for $\theta_\textrm{avg}$ can be obtained by $[\hat\theta_L-c_\alpha\hat\sigma_L,\, \hat \theta_U + c_\alpha \hat \sigma_U]$, where $\hat\sigma_L$ and $\hat \sigma_U$ are the estimated standard errors of $\hat \theta_L$ and $\hat \theta_U$, respectively, and $c_\alpha$ is chosen by solving \[ \Phi\Bigl( c_\alpha + \frac{\hat \Delta}{\max(\hat\sigma_L,\hat\sigma_U)} \Bigr) - \Phi(-c_\alpha) = 1-\alpha, \] where $\Phi$ is the distribution function of the standard normal and $\hat \Delta$ is the estimated length of the identified interval.
The second data scenario is slightly more complicated, because $\theta_{U_e}$ and $\theta_L^*$ contain the min or max function that is not smooth; so, the delta method does not apply. In online Appendices (ref) and (ref), we propose a two-step method for inference to overcome this problem, which we have applied to the empirical example we discuss in (ref). In the third data scheme, confidence intervals for $\theta_\textrm{avg}$ and $\theta_\textrm{local}$ always coincide, and they can be obtained by using a one-side critical value on $\hat \theta_L$. Specifically, they are given by $[\hat\theta_L - z_{1-\alpha}\hat \sigma_L,\, 1]$, where $z_{1-\alpha}$ is the $(1-\alpha)$ quantile of the standard normal distribution. Online Appendices (ref) and (ref) provide a more detailed discussion on inference. Furthermore, see online (ref) for semiparametrically efficient estimation of the two key parameters, i.e.,\ $\theta_L$ and $\theta^*$, when exogenous covariates $X_i$ are present and integrated out.
In sum, this paper clarifies identification issues when we insert exposure as a choice variable and employ a proper causal framework that is used in policy evaluation to model two causal links (i.e.,\ $Z_i \rightarrow T_i$ and $T_i \rightarrow Y_i$).\footnote{We are grateful to an anonymous referee who provided us with insightful comments.}
In this subsection, we revisit CY2019, who conducted a field experiment in China to measure the effects of providing students with internet access to the uncensored media on various outcome variables. Excluding the existing users, the subjects in their experiments consist of four groups: (i) the control group; (ii) the control group who were encouraged to visit foreign news websites blocked by the Great Firewall; (iii) students who received free access to uncensored internet; and (iv) students who received both the access and encouragement treatments. They followed the subjects over 18 months to collect outcomes on media-related behaviors, beliefs, and attitudes among other things. It turns out that there were no differences between groups (i) and (ii) and the effects were the largest for group (iv), i.e., the access plus encouragement group (the Group-AE students from now on). To benchmark their findings, CY computed DK's measure of persuasion for the Group-AE students (see online Appendix Table A.13 of their paper for the details) and commented that “their estimated persuasion rates are of a similar magnitude to those found in authoritarian regimes that typically have highly regulated media markets” (see pp. 2323--24 in CY).
In this section, we use the data from CY to illustrate how the common practice of reporting DK's measure can lead to misleading conclusions by contrasting DK's measure of persuasion with our proposed approaches. As in CY, we focus on the Group-AE students. That is, $Z_i = 1$ if the $i$th subject is randomly assigned to the Group-AE group, and $Z_i = 0$ if the $i$th subject is randomly assigned to the control or control-encouragement group, while dropping the access only group and the existing users. In CY's two-stage least squares analysis (Table 3 in their paper), the treatment variable is $T_i = 1$ if the $i$th subject is an active user of the censorship circumvention tool and and $T_i = 0$ otherwise. We use the same treatment variable in our analysis. To replicate the results in CY in a representative but succinct way, we focus on the 11 outcome variables listed in Panel A of online Appendix Table A.13 in their paper. They represent media-related behaviors, beliefs, and attitudes and are transformed to binary variables by CY.
Recall the population version of DK's measure $f$; see (ref). In their online Appendix Table A.13, CY measure $\mathbb{P}(Y_i = 1\mid Z_i = 1) - \mathbb{P}(Y_i = 1\mid Z_i = 0)$ by the intent-to-treat effects of the Group-AE assignment, and approximate $\mathbb{P}\{ Y_i(0) = 1 \}$ using variables collected at the time of the baseline survey or by the estimates of $\mathbb{P}( Y_i = 1 \mid Z_i = 0 )$ as in $\tilde \theta_{DK}$, if the former is unavailable. Table 2 in CY and the data provided by CY indicate that the change in the exposure rate is 45.5% if the treatment status is measured by being active users, and therefore we use $e(1) - e(0) = 0.455$ for our subsequent calculations.
(ref) summarizes the empirical results. In Column (2) of Table (ref), we recompute CY's persuasion rates for the 11 outcome variables listed in Panel A of online Appendix Table A.13 in their paper. As we explained above, these estimates are based on $e(1) - e(0) = 0.455$. Out of the 11 outcome variables, the median persuasion rate is 101%, and therefore these “persuasion rates” cannot be understood as conditional probabilities. Columns (3) and (4) report our estimates of the average and local persuasion rates along with the 95% confidence intervals (in curly braces) that were obtained via 10,000 bootstrap replications. Our estimates show that (i) the average persuasion rates are only partially identified and the widths of the identified intervals are substantial for most of the outcomes, (ii) the point-identified local persuasion rates are contained by the identified intervals for the average persuasion rates and are typically closer to the upper end points of the intervals. The persuasive effects are of a relatively large magnitude in that the smallest lower end point of the confidence interval for the average persuasion is 16%. However, the original estimates overstate the magnitude by a substantial factor and mask under-identification of average persuasion rates. In short, we find that the subjects in the experiments responded to exposure to uncensored internet highly heterogeneously, indicating that it is important to go beyond the benchmark measures of DK type.
We now illustrate our proposed methods by using data from gerber2009does, who report findings from a field experiment to measure the effect of political news. We have chosen this example because it contains a credible binary instrument from the field experiment and we can also illustrate all of the three sampling scenarios as well as the case of nonbinary outcomes; for the theory on the multinomial outcome case, see online (ref). In GKB, there are three statuses in the intention to treat: a control group, an offer of free subscription to The Washington Post, and one to The Washington Times. To illustrate the usefulness of our paper, we focus on The Washington Post and drop all observations from The Washington Times subscription. That is, $Z_i = 1$ if the $i$th individual received free subscription to The Washington Post, and $Z_i = 0$ if not.
Focusing on the ITT analysis, GKB have reported ITT estimates for various outcomes $Y_i$. dellavigna2010persuasion compute persuasion rates for GKB, for which they simply set $T_i = 1$ if the $i$th individual opted into the free subscription and $T_i = 0$ if they opted out of it.\footnote{We provide the empirical results of bound analysis using the opting-into-the-free-subscription treatment variable in online (ref).} In this section, for the purpose of illustrating our identification results, we consider a different treatment variable: $T_i = 1$ if the $i$th individual read a newspaper at least several times per week and $T_i = 0$ otherwise, which is a variable that GKB kept track of in a follow-up survey. Therefore, the relevant treatment we consider differs from that of dellavigna2010persuasion, but it is whether individuals have actually read the newspaper or not. The outcome variables we consider are as follows. For the binary case, $Y_i = 1$ if the $i$th individual reported voting for the Democratic candidate in the 2005 gubernatorial election, and $Y_i = 0$ if the subject did not vote for the Democratic candidate or did not vote at all. For the multinomial case, not voting at all is treated as an outside option. We use only a subsample of the GKB data with those who responded to the follow-up survey to use information on $(Y_i, T_i, Z_i)$ jointly. After dropping observations for The Washington Times subscription and removing missing data, we summarize the GKB data in (ref). Although the joint distribution of $(Y_i, T_i, Z_i)$ is observed in this example, we also consider using the two marginals of $(Y_i,Z_i)$ and $(T_i, Z_i)$ separately, to make a comparison. The estimates are summarized in (ref). Because the size of the sample extract we use is relatively modest ($n=701$) for an interval-identified object, we report the 80% confidence intervals obtained by the inference methods described in (ref) as well as in online (ref).
First, we discuss the case where the full joint distribution of $(Y_i, T_i, Z_i)$ is used. In this data scenario, the average effect of persuasion by reading the newspaper is bounded between $7\%$ and $63\%$. In contrast, the persuasion rate for the group of compliers is point-estimated by $81\%$. It is interesting to note that the estimate of $\theta_\textrm{local}$ is so large that it is greater than the upper bound of $\theta_\textrm{avg}$. This suggests that individuals are highly heterogeneous in this example, indicating that $\tilde \theta_{DK}$ might not be a well-defined conditional probability here. Indeed, the estimate of $\tilde \theta_{DK}$ in (ref) is $\hat{\tilde \theta}_{DK} = 1.1027$, which is greater than one.
When the marginals of $(Y_i, Z_i)$ and $(T_i, Z_i)$ are used separately, the upper bound on $\theta_\textrm{avg}$ increases from $63\%$ to $78\%$. Further, $\theta_\textrm{local}$ is no longer point estimated but we only know that it is bounded between $78\%$ and $100\%$. This difference illustrates the loss of identification power if we do not observe the joint distribution of $(Y_i, T_i, Z_i)$.
Finally, we estimate the lower bound on the average persuasion rate by additionally conditioning on those who would vote even without reading the newspaper (see online (ref) for details). The resulting lower bound on the average persuasion increases from 0.0707 (0.0289) to 0.0975 (0.0554), where the numbers in the parentheses are the left-end points of the 80% confidence intervals. Therefore, the (point identified) ITT effect is 5%, while the lower bound of the average persuasion rate is about 7%, or 10% if we further focus on those who would vote without reading the newspaper.
We have set up a simple econometric model of persuasion, introduced several parameters of interest, and analyzed their identification. The empirical examples in (ref) as well as the examples in online (ref) demonstrate that the persuasive effects are highly heterogeneous in the settings of media and fundraising.
We have focused on the case of binary outcomes and binary treatments. In online (ref), we extend our analysis to nonbinary outcomes. If the outcome is nonbinary, then we can condition on those who would not choose the outside option without the treatment. For instance, suppose that we have three options of voting for a Republican, voting for a Democrat, or not voting at all. Then, the persuasive effect of a message supporting a Republican can be measured in a couple of different ways: focusing on those who would not have voted for a Republican without the message is one way and conditioning on those who would have voted for a Democrat (i.e.,\ voted but not voted for a Republican) is the other. In the latter case, we show that the resulting lower bound is always no smaller than that of the binary outcome case.
In general, treatments are multivalued: unordered treatments (e.g.,\ watching Fox News, CNN or MSNBC) and ordered treatments (e.g.,\ numbers of hours watching Fox News) arise naturally in applications. It would be fruitful to build on recent developments in multivalued treatments HUV2006,HV2007-handbook-2,HUV2008,HeckmanPinto, LeeSalanie to investigate identification of persuasive effects. It would also be interesting to estimate deep parameters in an economic model of persuasion by using a more structural approach in the set-up of multivalued treatments. These are topics for future research.