EconBase
← Back to paper

Identifying the Effect of Persuasion

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

79,381 characters · 15 sections · 64 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Identifying the Effect of Persuasion

titlepage\begin{center} \LargeIdentifying the Effect of Persuasion\footnote{We would like to thank the editor, four anonymous referees, Eric Auerbach, Stefano DellaVigna, Leonard Goff, Marc Henry, Keisuke Hirano, Joel Horowitz, Charles Manski, Joris Pinkse, Imran Rasul, Roman Rivera, Myunghyun Song, David Yang, and seminar participants at Northwestern University, Rutgers University, Seoul National University and Vanderbilt University for helpful comments. This work was supported in part by the European Research Council (ERC-2014-CoG-646917-ROMIA) and by the UK Economic and Social Research Council (ESRC) through research grant (ES/P008909/1) to the Centre for Microdata Methods and Practice. This paper includes applications based on previously published articles, and we would like to thank the authors of these papers for making their data sets and replication files available via journal archives or personal web pages. } \end{center} \begin{center} \begin{tabular}{ccc} \largeSung Jae Jun\footnote{Department of Economics, Pennsylvania State University, 619 Kern Graduate Building, University Park, PA 16802. Email:\, [email removed]} & & \largeSokbae Lee\footnote{Department of Economics, Columbia University, 1022 International Affairs Building, 420 West 118th Street, New York, NY 10027; Centre for Microdata Methods and Practice, Institute for Fiscal Studies, 7 Ridgmount Street, London WC1E 7AE, UK. Email:\, [email removed]} \\ \smallPenn State Univ. & & \smallColumbia Univ. and IFS \end{tabular} \end{center} \begin{center} November 30, 2022 \end{center} Abstract. This paper examines a commonly used measure of persuasion whose precise interpretation has been obscure in the literature. By using the potential outcome framework, we define the causal persuasion rate by a proper conditional probability of taking the action of interest with a persuasive message conditional on not taking the action without the message. We then formally study identification under empirically relevant data scenarios and show that the commonly adopted measure generally does not estimate, but often overstates, the causal rate of persuasion. We discuss several new parameters of interest and provide practical methods for causal inference. \textbf{Key Words: } Communication, Media, Persuasion, Partial Identification, Treatment Effects \\ \textbf{JEL Classification Codes: } C21, D72, L82

\setcounter{footnote}{0}

\raggedbottom

Introduction

How effectively one can persuade one's audience has been of interest to ancient Greek philosophers in the Lyceum of Athens,\footnote{See sep-aristotle-rhetoric for three technical means of persuasion in Aristotle's Rhetoric.} early--modern English preachers in St Paul's Cathedral,\footnote{See kirby2008public for historic details of the public persuasion at Paul's Cross, the open-air pulpit in St Paul's Cathedral in the 16th century.} and contemporary American news producers at Fox News in New York City.\footnote{dellavigna2007fox and martin2017bias measure the persuasive effects of slanted news using data on Fox News.} Recently, economists have been endeavoring to build theoretical models of persuasion kamenica2011,CDK,gentzkow2017,Prat,bergemann2017information and to quantify empirically the extent to which persuasive efforts affect the behavior of consumers, voters, donors, and investors dellavigna2010persuasion.

Since dellavigna2007fox proposed a measure of persuasion, it has been used and modified by many authors EPZ,gentzkow2011effect,DEMPZ,bassi2016persuasion,martin2017bias,CY2019 to quantify the persuasive effects of informational treatment. However, its precise interpretation has been obscure because of a lack of formal identification analysis. In fact, we show that the commonly used measure of persuasion does not estimate the causal rate of persuasion on any subpopulation in general. Therefore, it is misleading to call DK's measure the persuasion rate, although it is a common practice in the literature; instead, we will reserve the term of the persuasion rate for the population parameter that is properly defined by a conditional probability in the potential outcome framework.

The flaw of DK's measure arises from failing to distinguish a local average treatment effect imbens1994late from the the average treatment effect (ATE), where the difference between the two can be substantial for a heterogeneous population; specifically, DK's measure rescales a LATE with a factor that is relevant only for the ATE. For instance, in DK's example, even if Fox News has a high persuasive effect among those who will watch the channel if and only if it is available through a local cable package, it may not be persuasive at all when other people such as Democrats or Democrat-leaning independents are all included.

Focusing on the case of binary outcomes, we analyze the problem of measuring persuasive effects of informational treatment through the lens of the potential outcome framework. Specifically, we define the persuasion rate by a proper conditional probability of the agent taking an action of interest with a persuasive message given that the agent does not take the action without the conveyed message. We then formally study identification under a few empirically relevant scenarios of data availability. Our analysis will articulate what DK's measure of persuasion estimates, why it is misleading, and how we can fix the problem.

While the persuasion rate is concerned with the entire population (and hence related with the ATE), we also consider a local persuasion rate that focuses on the group of compliers. Our identification analysis shows that DK's measure estimates neither a local nor an average persuasion rate.

Our identification analyses are based on a few empirically relevant scenarios of data availability. We do this because the problem of data availability is particularly important in the context of measuring persuasive effects. For instance, individual-level partisan vote outcome data rarely exist due to confidentiality issues. Thus, the outcome variable is frequently measured only at an aggregate level. Also, it is not always the case to observe an actual exposure to a persuasive message in the same data set along with the outcome and instrument. Indeed, DK's analysis uses a micro dataset for the treatment and instrument and a separate aggregate dataset for the outcome and instrument. In order to address these challenges, we consider three different scenarios of data availability explicitly: given the instrument,

inparaenum[(i)] • the outcome and treatment are jointly observed, • they are observed separately, or • the treatment is not observed at all.

We obtain the sharp bounds on the persuasion rate for the entire population as well as other subpopulations of potential interest. Therefore, our work builds on the econometrics literature on partial identification Manski:03,Manski:07,Tamer:10 as well as the literature on program evaluation Heckman/Vytlacil:07,imbens2009recent.

The main findings of this paper can be summarized as follows. If there is no heterogeneity in the population, then DK's measure of persuasion estimates the rate of persuasion, provided that a simple monotonicity assumption is imposed; however, this case is an exception rather than the rule. Indeed, the rate of persuasion is only partially identified as an interval in general, where the sharp lower bound generally corresponds to DK's measure multiplied by the relative size of the complier group regardless of the specific data scenario. Therefore, DK's measure is strictly larger than the lower bound whenever there is partial compliance. It is also remarkable that the data scenarios matter only for the sharp upper bound. Therefore, the value of observing the treatment and outcome jointly only lies in obtaining a potentially more informative upper bound on the persuasion rate. We also investigate identification of the local persuasion rate (i.e.,\ the persuasion rate for the group of compliers) under the same three data scenarios. It is point-identified under the most favorable data scenario, but only partially identified under the other scenarios; even in the case of point identification, DK's measure generally differs from the local persuasion rate. If a continuous instrument is available, then we can target a marginal persuasion rate that is akin to the marginal treatment effect Heckman/Vytlacil:05. Therefore, having a continuous instrument opens up the possibility of point identification of the persuasion rate for a policy-relevant population if the instrument is sufficiently rich.

In order to illustrate our findings, we discuss two empirical examples in the main text, while we provide a few more in the online appendices. First, we revisit CY2019, where the interest is in the effect of Chinese students having access to uncensored media on behaviors, beliefs and attitudes; we use the same variables and setup as in the original paper for this exercise. Second, we analyze the voting behavior and newspaper readership using the data from gerber2009does. Overall, we show that DK's measure of persuasion tends to overstate the persuasive effects while masking underlying heterogeneity.

The remainder of the paper is organized as follows. In Section (ref), we recall some backgrounds and specifics of DK's measure of persuasion. Here, we properly define the persuasion rate at the population level by using the potential outcome framework. In Section (ref), we discuss identification of the persuasion rate. In Section (ref), we study the local and marginal versions of the persuasion rate. In Section (ref), we discuss our recommendations on what to do in practice, including practical inferential issues,\footnote{Stata commands for estimation and inference based on this paper's identification results are publicly available at \url{https://github.com/persuasio}. Alternatively, they can be installed from within Stata by typing “ssc install persuasio”. } and clarify the difference between a population version of DK's measure of persuasion and the persuasion rate. In Section (ref), we provide two empirical illustrations. We conclude in Section (ref). The online appendices include additional results and examples, including an extension to non-binary outcomes, a detailed discussion about methods for inference, and all of the proofs.

Background

It is helpful to recall DK as a prototypical example, where they study the effect of an exposure to Fox News on the probability of voting for a Republican presidential candidate. Here, the informational treatment of interest is the viewership of the Fox News channel, and the persuasion rate of the media can be thought of as the proportion of the Fox News viewers who voted for a Republican candidate among those who would not have done so if they had not watched Fox News. It should be noted that the agents' decisions about whether to watch Fox News or not may be correlated with their political orientation. In order to address the endogeneity issue, DK use Fox News availability via local cable in the year of 2000 as an instrumental variable.

In efforts to measure the persuasion rate as explained above, DK and dellavigna2010persuasion propose the (infeasible) estimand $f$ defined as follows: for a binary outcome and by using DK's notation

equation[equation omitted — 125 chars of source]

where $\mathbb{T}$ and $\mathbb{C}$ represent an instrument assignment status such as having Fox News available via local cable or not.\footnote{Therefore, $\mathbb{T}$ and $\mathbb{C}$, which appear to denote treatment and control groups, should be understood as the status of the intent to treat, not the actual treatment status.} Here, for $j\in \{ \mathbb{T}, \mathbb{C} \}$, $y_j$ is the share of group $j$ taking the action of interest (e.g.,\ voting for a Republican candidate), and $e_j$ is the share of group $j$ exposed to a persuasive message. Further, $y_0$ is the share of those who would take the action of interest without listening to the persuasive message. So, $f$ is a rescaled version of the usual Wald statistic that estimates the LATE, where rescaling is apparently to obtain a “rate” that focuses on those who are to be persuaded. As DK noted, $y_0$ involves a counterfactual that is often unobserved. For this reason, $f$ is generally not a feasible estimand, and DK propose using $y_\mathbb{C}$ in place of $y_0$ as an approximation, which yields a feasible estimand, say $\tilde f$. We will refer to $f$ or its feasible approximation $\tilde f$ by DK's measure of the persuasion rate.

Without a formal justification, it is common in the literature to interpret $f$ as a conditional probability. For instance, DK explain, on page 1218 of their paper, “The key parameter is $f$, the fraction of the audience that is convinced by Fox News to vote Republican.” Similar interpretations are prevalent in the literature, as the following quotations demonstrate:

quotedellavigna2010persuasion: Whenever possible, we report the results in terms of the persuasion rate (DellaVigna & Kaplan 2007), which estimates the percentage of receivers that change the behavior among those that receive a message and are not already persuaded. \\ gentzkow2011effect: We can also translate our estimates into a “persuasion rate” (DellaVigna and Kaplan 2007) -- the number of eligible voters who changed their voting behavior as a result of the introduction of the newspaper, as a fraction of all those who could have changed their behavior. \\ DEMPZ: The persuasion rate is the fraction of the audience of a media outlet who are convinced to change their behavior (in this case, their vote) as a result of being exposed to this media outlet. \\ martin2017bias: Finally, we computed estimates of DellaVigna and Kaplan's (2007) concept of persuasion rates: the success rate of the channels at converting votes from one party to the other. The numerator in the persuasion rate here is the number of, for example, the Fox News Channel (FNC) viewers who are initially Democrats but by the end of an election cycle change to supporting the Republican party. The denominator is the number of FNC viewers who are initially Democrats.

Also, in their survey, dellavigna2010persuasion use $\tilde f$, i.e., a feasible approximation of $f$, as a key summary statistic to compare persuasive effects across different studies.

However, as we mentioned in the introduction, neither $f$ nor $\tilde f$ estimates the persuasion rate, i.e., the intended conditional probability, on any subpopulation in general. Therefore, it is misleading to label $f$ or $\tilde f$ as a persuasion rate, and it is generally invalid to make comparisons across different studies. Indeed, $f$ or $\tilde f$ may not even be a proper conditional probability in a heterogeneous population: e.g., the approximation $\tilde f$ can even be larger than one. We will articulate under what assumptions $\tilde f$ turns out to be a proper conditional probability, and how we should interpret it in its relationship with the persuasion rate.

In order to facilitate our discussion, we start with formally defining the persuasion rate at the population level by using the potential outcome framework. Let $T_i$ denote the binary indicator that equals $1$ if individual $i$ is exposed to persuasive information such as Fox News. Let $Y_i(t)$ be a binary indicator, which shows agent $i$'s potential action when $T_i$ is set to $t \in \{0,1\}$. For example, $Y_i(1)$ equals $1$ if individual $i$ votes for a Republican candidate after watching Fox News. The econometrician never observes both $Y_i(0)$ and $Y_i(1)$ but can only observe either of the two: that is, $Y_i = T_iY_i(1) +(1-T_i) Y_i(0)$.\footnote{So, both $Y_i$ and $T_i$ are binary. In online (ref), we extend our results to the case where the potential outcomes are multinomial.} Then, the fraction of the people who take the action of interest with an exposure to a persuasive message, among those who would not without it, is given by

equation[equation omitted — 102 chars of source]

provided that the conditional probability is well-defined: $\theta_\textrm{pr}$ is the persuasion rate at the population level. Using conditional probability is to rule out the case of “preaching to the converted”;\ if $Y_i(0) = 1$, then those individuals are already “persuaded” to take the action of interest even without the persuasive treatment, and therefore we do not count them in defining the persuasion rate.\footnote{The idea of using conditional probability to define a parameter of interest can also be found in heckman1997making, though their context is quite different from ours.}

The common estimand $\tilde f$ (or even $f$) does not estimate $\theta_\textrm{pr}$ in a heterogeneous population; in fact, even in a homogeneous population, $\tilde f$ or $f$ cannot be (asymptotically) equated with $\theta_\textrm{pr}$ without an extra monotonicity assumption. However, the rescaled quantity $(e_\mathbb{T} - e_\mathbb{C} ) \tilde f $, which is always no greater than $\tilde f$, does provide valid information about $\theta_\textrm{pr}$ in that it corresponds to the sharp lower bound of the identified interval of $\theta_\textrm{pr}$ in general. The best way to clarify all the issues is to conduct a rigorous analysis on the identification of $\theta_\textrm{pr}$, which is our next topic. For quick takeaways, see (ref).

Identification of the Persuasion Rate

Identification of $\theta_\textrm{pr}$ is challenging for various reasons, including the fact that $\theta_\textrm{pr}$ depends on the joint distribution of the potential outcomes and that $T_i$ can be endogenous, and it is often difficult to observe $T_i$ and $Y_i$ jointly. To allow for potential endogeneity, we use a binary instrument, $Z_i$, throughout the paper unless otherwise specified. Exogenous covariates, $X_i$, could be observed, but we suppress $X_i$ from our identification analysis. In other words, we implicitly assume throughout the paper that all assumptions and results are conditional on the value of $X_i$. Therefore, the observed variables are $Y_i, T_i, Z_i$, all of which are binary throughout the main text. See (ref) for how to deal with $X_i$ in practice. Also, see (ref) for an extension to the case of multinomial outcomes.

The data issue on $T_i$ is addressed by considering three scenarios of data availability. Specifically, for the purpose of the identification analysis, we assume that for $(y,t,z)\in \{0,1\}^3$, (i) $\mathbb{P}(Y_i = y, T_i = t\mid Z_i=z)$ is known, (ii) $\mathbb{P}(Y_i = y\mid Z_i=z)$ and $\mathbb{P}(T_i = t\mid Z_i=z)$ are known, or (iii) $\mathbb{P}(Y_i = y\mid Z_i=z)$ is all that is known, depending on the specific scenario of interest. For example, DK use town-level election data to estimate $\mathbb{P}(Y_i = y \mid Z_i=z)$ and micro-level media audience data to infer $\mathbb{P}(T_i = t\mid Z_i=z)$, which corresponds to Case (ii).\footnote{Throughout the discussion, we assume that $T_i$ is correctly measured if it is observed. See latewithmismeasurement, nguimkeu2016estimation, and Ura2018miclassification for the issues of mismeasured treatment. Their subject matter is distinct from ours.}

It requires an additional assumption to address the challenge that $Y_i(1)$ and $Y_i(0)$ are never simultaneously observed while $\theta_\textrm{pr}$ depends on their joint distribution. Before we present our identification results, we discuss our key assumptions in the following subsection.

The Key Assumptions

Our first key assumption is that the persuasive message is directional, which will be important to decouple $\theta_\textrm{pr}$ by the marginals of the potential outcomes.

assumption[Monotonic Treatment Response] The potential outcomes $Y_i(1)$ and $Y_i(0)$ are binary, and they satisfy $Y_i(0) \leq Y_i(1)$ with probability one.

(ref) is a binary version of the monotonic treatment response (MTR) assumption used in manski1997monotone and manski2000mtr.\footnote{In online Appendix A, we present a simple economic model that motivates (ref).} (ref) allows $Y_i(0) = Y_i(1)$ with probability one, and therefore it does not rule out the possibility that `watching Fox News' has no impact on the agent's behavior at all. The inequality in (ref) means that the messages Fox News delivers are biased or directional in favor of Republican candidates; i.e.,\ if a voter is going to vote for a Republican candidate without watching Fox News, then watching Fox News will not change that. In other words, (ref) rules out the possibility that the level of distrust that a voter has on Fox News is so high that she takes actions based on the opposite of the messages Fox News delivers.

Since (ref) is a key assumption in the paper, we first clarify how much we can hope for with and without (ref).

lemmaWe generally have \begin{equation} \max\Bigr[ 0,\ \dfrac{ \mathbb{P}\{ Y_i(1) = 1 \} - \mathbb{P}\{ Y_i(0) = 1 \} }{ 1 - \mathbb{P}\{ Y_i(0) = 1\} } \Bigr] \leq \theta_pr \leq \min\Bigl[ \dfrac{\mathbb{P}\{ Y_i(1) = 1 \}}{1-\mathbb{P}\{Y_i(0)=1\}} ,\ 1 \Bigr], \end{equation} where the bounds are sharp in that $\theta_\textrm{pr}$ can be anything between the bounds without changing the marginals $\mathbb{P}\{ Y_i(1) = 1\}$ and $\mathbb{P}\{ Y_i(0) = 1\}$. Further, (ref) holds if and only if $\theta_\textrm{pr} = \theta_\textrm{avg}$, where \begin{equation} \theta_avg := \dfrac{ \mathbb{P}\{ Y_i(1) = 1 \} - \mathbb{P}\{ Y_i(0) = 1 \} }{ 1 - \mathbb{P}\{ Y_i(0) = 1\} }. \end{equation}

(ref) is a consequence of the Fr\'{e}chet--Hoeffding inequality on the probability of a joint event. Since the two potential outcomes are never observed simultaneously, all we can ever hope to identify is their marginals and (ref) expresses the (sharp) bounds on $\theta_\textrm{pr}$ in terms of the marginal probabilities of the potential outcomes.

Suppose that there are some voters who have such a high level of distrust on Fox News that rather not watching Fox News would help them to take favorable actions to a Republican candidate but that those voters are only minority and on average we still have a strict stochastic dominance relationship between $Y_i(1)$ and $Y_i(0)$ (i.e., $\mathbb{P}\{ Y_i(1) = 1 \} > \mathbb{P}\{ Y_i(0) = 1 \}$). Then, the lower bound on $\theta_\textrm{pr}$ will be ensured to be nontrivial.

(ref) is stronger than the stochastic dominance, but it delivers a stronger result. In fact, (ref) is necessary and sufficient to express $\theta_\textrm{pr}$ in terms of the marginal probabilities of the counterfactual outcomes. In this case, the conditional probability $\theta_\textrm{pr}$ is the average treatment effect (ATE) divided by $\mathbb{P}\{ Y_i(0) = 0\}$. Throughout the rest of the paper, we present most of our results by using (ref) because it is not only convenient but also being biased or directional seems to be the nature of persuasive effort. However, we emphasize that $\theta_\textrm{avg}$, the rescaled version of the ATE, is always a valid lower bound on $\theta_\textrm{pr}$ as (ref) shows.

The next assumption is concerned about the treatment assignment $T_i$ and the instrument $Z_i$.

assumption[No Defiers and an Exogenous Instrument] The binary treatment $T_i$ has a threshold structure, i.e., \begin{equation} T_i = \mathbb{1}\{ V_i \leq e(Z_i) \}, \end{equation} where $V_i$ is an unobserved random variable that is uniformly distributed on $[0,1]$. The binary instrument $Z_i$ is independent of $\bigl(Y_i(t), V_i \bigr)$ for $t= 0,1$. Finally, we have $0\leq e(0) < e(1) \leq 1$.

(ref) is standard for causal inference using instrumental variables. The intent-to-treat (ITT), $Z_i$, is randomly assigned; however, $T_i$ can be endogenous via the dependence between $V_i$ and $Y_i(t)$. The function $e(\cdot)$ is the propensity score or, more descriptively in our context, it can be referred to as the exposure rate.

If $T_i(z)$ denotes counterfactual treatment when a binary instrument takes value $z$, then agent $i$ is called a complier when $T_i(1) = 1$ and $T_i(0)=0$; never-takers ($T_i(z) = 0$ for all $z$), always-takers ($T_i(z)=1$ for all $z$), and defiers ($T_i(1) = 0, T_i(0) = 1$) are similarly understood. Under (ref), $i$ is a complier if and only if $e(0)<V_i\leq e(1)$. Similarly, $i$ is an always-taker when $V_i\leq e(0)$, and she is a never-taker when $V_i> e(1)$. Therefore, under (ref), there are only three groups in the population: i.e.,\ always-takers, never-takers and compliers. Indeed, as vytlacil2002independence has shown, the threshold structure in (ref) is equivalent to assuming the absence of defiers, which is a popular assumption in econometrics to identify the local average treatment effect.

In the following two subsections, we present identification results for $\theta_\textrm{pr}$, which is the same as $\theta_\textrm{avg}$ under (ref). (ref) covers the simplest case, where everybody complies so that there is no difference between the actual treatment and the intent to treat (ITT), i.e., $T_i = Z_i$ for each $i$: this case is referred to as the sharp persuasion design, where there is no distinction among different data scenarios. In this case, not surprisingly, $\theta_\textrm{avg}$ is point identified from the distribution of $Y_i$ given $Z_i$. However, when $T_i$ and $Z_i$ are different, which we call the fuzzy persuasion design, we have only partial identification of $\theta_\textrm{avg}$, where each of the three data scenarios become relevant. Before we move on, we define

equation[equation omitted — 155 chars of source]

provided that $\mathbb{P}(Y_i = 1\mid Z_i=0)<1$: $\theta_L$ is an identified parameter from the distribution of $Y_i$ given $Z_i$. It turns out that first, $\theta_L$ is equal to $\theta_\textrm{avg}$ under the sharp persuasion design, and second, it is the sharp lower bound on $\theta_\textrm{avg}$ in the fuzzy persuasion design regardless of which of the three data scenarios applies.

The Sharp Persuasion Design

theoremSuppose that (ref) hold. If $e(1) - e(0) = 1$ (i.e.,\ $T_i = Z_i$ with probability one), then for $z\in\{0,1\}$, we have $\mathbb{P}\{ Y_i(z) = 1\} = \mathbb{P}( Y_i = 1 \mid Z_i = z)$, and hence $\theta_\textrm{pr} = \theta_\textrm{avg} = \theta_L$.

The condition of $e(1) - e(0) = 1$ means that everybody is a complier, and hence there is essentially no difference between $T_i$ and $Z_i$; thus, the sharp design is equivalent to a situation where $T_i$ is observed and randomized. However, this is rather an exceptional situation in social sciences. The key identification question should be how far we can go when the design is not sharp (i.e.,\ not everybody is a complier). We answer this question in the following subsection.

Without (ref), the general sharp identified bounds on $\theta_\textrm{pr}$ are given by

equation[equation omitted — 164 chars of source]

Therefore, even in the sharp persuasion design, $\theta_\textrm{pr}$ is only partially identified and its sharp bounds can be trivial without (ref). Even so, it is worth noting that $\theta_L$ remains a valid lower bound.

The Fuzzy Persuasion Design

In the fuzzy design, the three scenarios of data availability we mentioned earlier become pertinent.

Identification with the Joint Distribution of $(Y_i, T_i, Z_i)$

Even the full joint distribution of $(Y_i,T_i,Z_i)$ does not point-identify the ATE. Recall that those with $Z_i = 0$ and $T_i =0$ comprise the compliers and the never-takers, while those with $Z_i = 0$ and $T_i = 1$ are the always-takers. Similarly, those with $Z_i = 1$ and $T_i = 0$ are the never-takers, while those with $Z_i = 1$ and $T_i=1$ consist of the compliers and the always-takers. Therefore, these four cases correspond to different subpopulations, and the only subpopulation that we can study for both $T_i=0$ and $T_i = 1$ in common is that of compliers, which explains why the Wald statistic estimates the LATE, not ATE. For the same reason, $\theta_\textrm{avg}$ cannot be point identified; however, we can derive its sharp bounds.

assumption[Full Observability] The joint distribution of $(Y_i, T_i, Z_i)$ is known, where $\mathbb{P}(Y_i = 0 \mid Z_i = 0 ) > 0$.
theoremSuppose that (ref) are satisfied. Then, the sharp identified interval of $\theta_\textrm{pr} = \theta_\textrm{avg}$ is given by $[\theta_L, \theta_U]$, where $\theta_L$ is given in (ref) and \begin{equation*} \theta_U := \dfrac{\mathbb{P}(Y_i = 1, T_i = 1\mid Z_i = 1)-\mathbb{P}(Y_i = 1, T_i = 0\mid Z_i = 0) + 1-e(1)}{1-\mathbb{P}(Y_i = 1, T_i = 0 \mid Z_i = 0)}. \end{equation*}

To prove (ref), we first derive the sharp identified bounds for $\mathbb{P}\{Y_i(1) = 1 \}$ and $\mathbb{P}\{Y_i(0) = 1 \}$ separately; we denote them by the intervals $[m_a, M_a]$ and $[m_b,M_b]$, respectively. These bounds are special cases of manski2000mtr under the MTR assumption coupled with the exogeneity of the instrument. Then, letting $a := \mathbb{P}\{ Y_i(1) = 1 \}$ and $b := \mathbb{P}\{ Y_i(0) = 1 \}$, we obtain the upper bound of the identified interval of $\theta_\textrm{avg}$ by solving

equation[equation omitted — 150 chars of source]

while the lower bound can be found by doing minimization instead of maximization. We then appeal to continuity and the intermediate value theorem for the sharpness result.

It is proved in online (ref) that $m_a = \mathbb{P}(Y_i=1\mid Z_i = 1)$ and $M_b = \mathbb{P}(Y_i=1\mid Z_i = 0)$. Therefore, an examination of (ref) reveals that the lower bound is attained when $a = m_a$ and $b = M_b$. To develop intuition behind (ref), we discuss what $a=m_a$ and $b = M_b$ means, for which the behavior of the never-takers and that of the always-takers matter. Since $m_a = \mathbb{P}\{Y_i(1) = 1, T_i = 1 \mid Z_i=1\} + \mathbb{P}\{Y_i(0) = 1, T_i=0 \mid Z_i=1\}$, we know that $a = m_a$ holds when $\mathbb{P}\{ Y_i(0)=1, T_i = 0\mid Z_i = 1\} = \mathbb{P}\{ Y_i(1)=1, T_i = 0\mid Z_i = 1\}$. Here, the event $Z_i = 1, T_i=0$ means that $i$ is a never-taker, because defiers are assumed to be non-existent. Therefore, $a=m_a$ means that the treatment has no effect on the group of never-takers, unless there are no never-takers at all. Similarly, $b = M_b$ holds when $\mathbb{P}\{Y_i(1)=1,T_i=1 \mid Z_i=0 \} = \mathbb{P}\{ Y_i(0)=1,T_i=1 \mid Z_i = 0 \}$. Since $Z_i = 0, T_i = 1$ means that $i$ is an always-taker, we know that $b=M_b$ holds when the treatment does not affect the behavior of the always-takers, unless there are no always-takers at all. Therefore, $\theta_\textrm{avg}$, which is the same the persuasion rate $\theta_\textrm{pr}$ under (ref), is smallest when there are null treatment effects for both the never-takers and the always-takers.

Intuition for the upper bound can also be obtained by considering the non-complier groups. The upper bound corresponds to the case where $a = M_a$ and $b = m_b$, where it is shown in (ref) that $M_a = \mathbb{P}(Y_i = 1, T_i=1 \mid Z_i=1)+1-e(1)$ and $m_b = \mathbb{P}(Y_i=1,T_i=0 \mid Z_i=0)$. Here, note that $a=M_a$ is equivalent to $\mathbb{P}\{Y_i(1)=0,T_i=0 \mid Z_i=1 \} =0$, and $b=m_b$ is to $\mathbb{P}\{ Y_i(0) = 1, T_i = 1 \mid Z_i=0\} = 0$. Therefore, we can see that the persuasion rate $\theta_\textrm{avg}$ equals the upper bound when every never-taker has $Y_i(1) = 1$ and none of the always-taker has $Y_i(0)=1$: e.g.,\ all those who never watch Fox News (whether it is available or not) would actually have voted for a Republican candidate if they had watched it and all those who always watch Fox News would not have voted for a Republican without watching the channel.

The bounds in (ref) shrink to a singleton as $\bigl( e(0), e(1) \bigr)$ approaches $(0,1)$, which is consistent with the result in (ref). Also, it is worth noting that the lower bound $\theta_L$ only depends on the distribution of $(Y_i, Z_i)$: observing $T_i$ along with $(Y_i, Z_i)$ helps only for the upper bound. If $e(1)$ is too small, then the upper bound will not be very informative: $\theta_U$ converges to $1$ as $e(1)$ approaches $0$; that is, if nobody reads a newspaper when they receive free subscriptions, then we do not learn much about how “persuading” the newspaper is. However, even if $e(1)$ approaches $1$, the upper bound does not necessarily shrink to the lower bound; e.g.,\ we do not necessarily pin down the persuasion rate of reading the newspaper even if everybody who has free subscriptions actually reads it.

We now establish partial identification of $\theta_\textrm{pr}$ without (ref), i.e., $\theta_\textrm{pr} \neq \theta_\textrm{avg}$, for the sake of completeness. Let $NT = \{V_i > e(1)\}$ and $AT = \{ V_i\leq e(0)\}$ be the event of $i$ being a never-taker and an always-taker, respectively.

theoremSuppose that (ref) are satisfied. \begin{enumerate}[label=(\roman*)] • If $\mathbb{P}(Y_i = 0\mid Z_i = z)>0$ for $z=0,1$, then the sharp identified interval of $\theta_\textrm{pr}$ is given by \begin{multline*} \max\Bigl\{0,\ \frac{\mathbb{P}(Y_i=1,T_i=1\mid Z_i=1)-\mathbb{P}(Y_i=1,T_i=0\mid Z_i=0)-e(0)}{1-\mathbb{P}(Y_i=1,T_i=0\mid Z_i=0)-e(0)} \Bigr\} \\ \leq \theta_pr \leq \min\Bigl\{ \frac{\mathbb{P}(Y=1,T=1\mid Z=1) +1-e(1)}{1-\mathbb{P}(Y=1,T=0\mid Z=0)},\ 1 \Bigr\}. \end{multline*} • If $\mathbb{P}(Y_i = 0\mid Z_i = 0)>0$ and $\mathbb{P}\{ Y_i(1)=1,\ \mathcal{E} \} \geq \mathbb{P}\{ Y_i(0)= 1,\ \mathcal{E}\}$ for $\mathcal{E} \in \{ NT, AT\}$, then the sharp identified interval of $\theta_\textrm{pr}$ is given by \[ \max\{ 0,\ \theta_L\} \leq \theta_\textrm{pr} \leq \min\Bigl\{ \frac{\mathbb{P}(Y=1,T=1\mid Z=1) +1-e(1)}{1-\mathbb{P}(Y=1,T=0\mid Z=0)},\ 1 \Bigr\}. \] \end{enumerate}

(ref) shows that $\theta_L$ continues to be the sharp lower bound, provided that $\theta_L \geq 0$, i.e., $\mathbb{P}(Y_i = 1\mid Z_i = 1) \geq \mathbb{P}(Y_i = 1\mid Z_i = 0)$, and a stochastic dominance condition holds for the never-takers and always-takers. However, the upper bound will be larger than that in (ref) in general.

Identification with the Knowledge of the Exposure Rates

As in the case of DK, the researcher may not directly observe $T_i$ along with $(Y_i,Z_i)$ but may have auxiliary data from which the exposure rates $e(1)$ and $e(0)$ can be estimated.\footnote{The case in which the outcome and the treatment are separately observed belongs to an identification problem called the ecological inference problem. For instance, CrossManski and Manski2017longshort discuss bounding a “long regression” by using information from a “short regression”. Their substantive concerns are distinct from ours.} In this case, the sharp identified bounds on $\theta$ become generally wider than those of (ref).

assumption[Observability of Two Marginals] Only the distribution of $(Y_i, Z_i)$ and the exposure rates $\{ e(0), e(1) \}$ are known, where $\mathbb{P}(Y_i = 0 \mid Z_i = 0) > 0$.
theoremSuppose that (ref) are satisfied. Then, the sharp identified interval of $\theta_\textrm{pr} = \theta_\textrm{avg}$ is given by $[\theta_L,\ \theta_{U_e}]$, where $\theta_L$ is given in (ref) and \begin{equation} \theta_{U_e} := \dfrac{ \min\{ 1, \mathbb{P}(Y_i = 1\mid Z_i = 1)+1-e(1) \} - \max\{0,\ \mathbb{P}(Y_i = 1\mid Z_i = 0)-e(0) \}}{1 - \max\{0,\ \mathbb{P}(Y_i = 1\mid Z_i = 0)-e(0) \}}. \end{equation}

Therefore, the upper bound in this case is nontrivial if and only if $e(1) > \mathbb{P}(Y_i = 1 \mid Z_i = 1)$.\footnote{The trivial case can occur in applications. See (ref) for such cases.} Note that it is the relative size of the take-up rate $e(1)$ (e.g.,\ the probability of reading a newspaper when a free subscription to it is offered) that determines how much we can hope to learn about the persuasion rate. For example, if the probability of watching Fox News is too small relative to the probability of voting for a Republican candidate when Fox News was introduced in the local cable, then it becomes difficult to pin down how successfully Fox News persuaded their audience to vote for a Republican candidate. Also, it is worth noting that $e(0) = 0$ is not uncommon as (ref) and (ref) show. In this case, the maximum in the expression of the upper bound is unnecessary. Intuition for the lower bound is the same as the case of (ref), because $\theta_L$ requires only the distribution of $Y_i$ given $Z_i$.

Identification with No Information Associated with $T_i$

The final scenario is the least informative one, where $T_i$ is not observed at all. This is an almost trivial case, but we state it in a separate theorem for the sake of completeness.

assumption[Limited Observability] No information associated with $T_i$ is available (i.e.,\ the distribution of $(Y_i, Z_i)$ is all that is known), where $\mathbb{P}(Y_i = 0\mid Z_i = 0) > 0$.
theoremSuppose that (ref) are satisfied. Then, the sharp bound of $\theta_\textrm{pr} = \theta_\textrm{avg}$ is given by $[\theta_L,1]$, where $\theta_L$ is given in (ref).

The lower bound from (ref) depends only on the distribution of $(Y_i,Z_i)$, and therefore $\theta_L$ continues to be the lower bound in this case as well. Further, since no information for $e(1)$ and $e(0)$ is available, it suffices to note that the upper bound in (ref) equals one whenever $e(0)>\mathbb{P}(Y_i=1\mid Z_i=0)$ and $e(1)<\mathbb{P}(Y_i=1\mid Z_i=1)$.

The Local and Marginal Persuasion Rates

In this section, we consider rates of persuasion on other subpopulations that have been considered in econometrics: i.e.,\

equation[equation omitted — 294 chars of source]

provided that the conditional probabilities are well-defined: $\theta_\textrm{local}$ is the persuasion rate for the compliers imbens1994late, whereas $\theta_\textrm{marginal}(v)$ is for the subpopulation such that $V_i = v$ Heckman/Vytlacil:05.

First, we obtain identification results for $\theta_\textrm{local}$ under the three sampling scenarios in the fuzzy persuasion design. The first step for this purpose is to note that the same reasoning as (ref) yields

equation[equation omitted — 236 chars of source]

where the numerator is the LATE, which has received great attention in the econometrics literature Deaton2010, Heckman2010, Imbens2010. The denominator that rescales the LATE is also conditioned on the same subpopulation of the compliers.

theoremSuppose that (ref) are satisfied. \begin{enumerate}[label=(\roman*)] • Under (ref), $\theta_\textrm{local}$ is point identified by $\theta_\textrm{local} = \theta^*$, where \[ \theta^* := \dfrac{\mathbb{P}(Y_i = 1 \mid Z_i = 1) - \mathbb{P}(Y_i = 1 \mid Z_i = 0)}{\mathbb{P}(Y_i = 0, T_i = 0 \mid Z_i = 0) - \mathbb{P}( Y_i = 0, T_i = 0 \mid Z_i = 1)}. \] • Under (ref), the sharp identified interval of $\theta_\textrm{local}$ is given by $[\theta_L^*, 1]$, where \[ \theta_L^* := \max\Bigl\{ \theta_L,\ \dfrac{\mathbb{P}(Y_i = 1 \mid Z_i = 1) - \mathbb{P}(Y_i = 1\mid Z_i = 0)}{e(1)-e(0)} \Bigr\}. \] • Under (ref), the sharp identified interval of $\theta_\textrm{local}$ coincides with that of $\theta_\textrm{pr}=\theta_\textrm{avg}$, i.e.,\ $\bigl[ \theta_L,\ 1 \bigr]$. \end{enumerate}

Recall that the identification of the LATE requires the joint distribution of $(T_i, Z_i)$ and that of $(Y_i, Z_i)$ separately but not the full joint distribution of $(Y_i, T_i, Z_i)$. Unlike the LATE, the point identification in (ref)(ref) demands the knowledge of the joint distribution of $(Y_i, T_i, Z_i)$.\footnote{The denominator of (ref) requires that we know the marginal distribution of $Y_i(0)$ for the compliers. imbens1997estimating show that the marginal distributions of $Y_i(1)$ and $Y_i(0)$ for the compliers are identified if the joint distribution of $(Y_i, T_i, Z_i)$ is known; however, they did not consider the local persuasion rate.} (ref)(ref) shows that this requirement is not only sufficient but also necessary to achieve the point identification of $\theta_\textrm{local}$.

Just like the LATE, it may be contentious whether or not $\theta_\textrm{local}$ should be the parameter of interest, because the compliers are concerned with an unidentified subgroup of the population. However, we take a practical view that the identification results on $\theta_\textrm{local}$ can complement the results obtained in (ref).

The local persuasion rate $\theta_\textrm{local}$ represents the average persuasive effect for a population that is different from the entire population. Given this caveat, it is interesting to note that, in (ref)(ref), the upper bound on $\theta_\textrm{local}$ is always trivial in contrast to $\theta_\textrm{avg}$, but the lower bound of $\theta_\textrm{local}$ can never be worse than that of $\theta_\textrm{avg}$. Therefore, in principle, the length of the identified interval of $\theta_\textrm{avg}$ can be smaller than that of $\theta_\textrm{local}$. If $T_i$ is not observed at all, then there is no advantage in focusing on the compliers. (ref)(ref) confirms the intuition that the bounds for $\theta_\textrm{local}$ are identical to those for $\theta_\textrm{avg}$ if the distribution of $(Y_i, Z_i)$ is the only piece of information available. This corresponds to an uninteresting case for $\theta_\textrm{local}$ though, as we have no information on compliers.

Data requirements for the identification of $\theta_\textrm{marginal}(v)$ are generally quite demanding: e.g.,\ a continuous instrument is needed. However, identifying $\theta_\textrm{marginal}(v)$ for various values of $v$ can open up the possibility of point identification of $\theta_\textrm{avg}$. Therefore, it is worth understanding what is sufficient for the identification of $\theta_\textrm{marginal}$.

If $Y_i$ and $T_i$ are jointly observed along with a continuous instrument $Z_i$, then $\theta_\textrm{marginal}(v)$ can be point identified as in Heckman/Vytlacil:05 and carneiro2011. Examples of continuous instruments can be found in the literature on the media effects on voting. For instance, EPZ and DEMPZ use the signal strength of NTV and Serbian radio as instruments, respectively; in both of the papers, $(Y_i, T_i, Z_i)$ are jointly observed. The following assumption describes the situation in which we can obtain point identification of $\theta_\textrm{marginal}(v)$. We use the standard results in the literature Heckman/Vytlacil:05 for the subsequent theorem.

assumption[Marginal Treatment Effects] \begin{enumerate}[label=(\roman*)] • The joint distribution of $(Y_i, T_i, Z_i)$ is known. • $T_i$ has the threshold structure in (ref), where $V_i$ is uniformly distributed on $[0,1]$, and $Z_i$ is independent of $\bigl(Y_i(t), V_i \bigr)$ for $t= 0,1$. • The distribution of $e(Z_i)$ is absolutely continuous with respect to Lebesgue measure. \end{enumerate}
theoremSuppose that (ref) are satisfied. Then, for $v$ such that $v$ is in the interior of the support of $e(Z_i)$, $\theta_\textrm{marginal}(v)$ is point identified by \begin{align} \theta_marginal(v) = \dfrac{\partial \mathbb{P} \{ Y_i = 1 \mid e(Z_i) = e \}/\partial e \big|_{e=v}}{1 + \partial \mathbb{P}\{ Y_i = 1, T_i = 0 \mid e(Z_i) = e \}/\partial e \big|_{e=v}}, \end{align} provided that $\mathbb{P} \{ Y_i = 1\mid e(Z_i) = e \}$ and $\mathbb{P} \{ Y_i = 1, T_i = 0 \mid e(Z_i) = e \}$ are continuously differentiable with respect to $e$.

Similarly to the case of $\theta_\textrm{avg}$ or $\theta_\textrm{local}$, (ref) enable us to rewrite $\theta_\textrm{marginal}(v)$ as $\mathbb{E}\{ Y_i(1) - Y_i(0) \mid V_i = v \} / \mathbb{P}\{Y_i(0) = 0\mid V_i = v\}$; i.e.,\ $\theta_\textrm{marginal}(v)$ is a rescaled version of the marginal treatment effect of Heckman/Vytlacil:05. (ref) is a direct consequence of that.

(ref) does not consider the other two scenarios of data availability. This is mainly because continuous instruments are relatively infrequent in the context of persuasion, and we are not aware of any applications where continuous instruments are available while the outcome and treatment are not jointly observed.

If the support of the exposure rate $e(Z_i)$ is equal to the unit interval $[0,1]$, then (ref) shows the identification of $\theta_\textrm{marginal}(v)$ for all $v$ in the unit interval. Then, we can use $\theta_\textrm{marginal}(v)$ to construct different policy-oriented quantities as in Heckman/Vytlacil:05 and carneiro2011. For instance, the persuasion rate of the entire population can be obtained by $\int_0^1 \theta_\textrm{marginal}(v) dF\{v\mid Y_i(0) = 0\}$, which is equal to

equation*[equation* omitted — 219 chars of source]

by Bayes' theorem.

Discussion

In this section we articulate the relationships between $\theta_\textrm{avg}, \theta_\textrm{local}$, and DK's measures $f$ and $\tilde f$ defined in (ref), and we summarize the main takeaways of our identification results. We assume that (ref) hold throughout this section, so we have $\theta_\textrm{pr} = \theta_\textrm{avg}$. Also, in order to be consistent with the identification analysis, we work with the population versions of $f$ and $\tilde f$: i.e.,\

align[align omitted — 417 chars of source]

First, neither $\theta_{DK}$ nor $\tilde \theta_{DK}$ is generally equal to the persuasion rate $\theta_\textrm{avg}$, or that for the compliers, $\theta_\textrm{local}$, which is clear from the fact that \[ \mathbb{P}\{ Y_i(0)=0 \} \theta_{DK} = \mathbb{P}\{ Y_i=0\mid Z_i = 0 \} \tilde \theta_{DK} = \mathbb{P}\{Y_i(0) = 0\mid e(0)<V_i\leq e(1) \} \theta_\textrm{local} \] is equal to the LATE, while $\mathbb{P}\{ Y_i(0) = 0\} \theta_\textrm{avg}$ is equal to the ATE.\footnote{Recall that the population version of the Wald statistic is equal to the LATE under (ref).} For example, $\theta_{DK}$ rescales the LATE with an unconditional probability and hence it does not render a well-defined conditional probability in general. Similarly, $\tilde\theta_{DK}$ is not necessarily a `rate' in spite of the rescaling factor. So, making comparisons across different studies based on $\theta_{DK}$ or $\tilde\theta_{DK}$ can be misleading, although it is a common practice dellavigna2010persuasion.

There are some special cases of exception though:

inparaenum[(i)] • everybody is a complier as in the sharp persuasion design, • $T_i$ is independent of the potential outcomes $Y_i(t)$ for $t=0,1$ conditional on $Z_i$, or • there is no heterogeneity in the treatment effect in that $Y_i(1) - Y_i(0)$ is a constant.

That is, in cases ((ref)) and ((ref)), there is no endogeneity issue, whereas in case ((ref)), no one is affected by the persuasive message (i.e.,\ $Y_i(1) - Y_i(0) = 0$ for all $i$), or everybody is persuaded (i.e.\ $Y_i(1) - Y_i(0) = 1$ for all $i$). Note that the LATE is the same as the ATE when any of the three cases applies. As a result, we have that in case ((ref)), $\theta_{DK} = \tilde\theta_{DK} = \theta_\textrm{avg} = \theta_\textrm{local}$ holds; in case ((ref)), $\theta_{DK} = \theta_\textrm{avg} = \theta_\textrm{local} \leq \tilde\theta_{DK}$; and, in case ((ref)), $\theta_{DK} = \theta_\textrm{avg} = \theta_\textrm{local} = 1\leq \tilde\theta_{DK}$ if everybody is persuaded, or $\theta_{DK}=\tilde\theta_{DK} =\theta_\textrm{avg} = \theta_\textrm{local} = 0$, unless any of them are ill-defined, if no one is affected by the informational treatment. Therefore, the feasible version $\tilde f$ of DK's measure of persuasion does estimate the persuasion rate $\theta_\textrm{avg}$ in case ((ref)), although, as DK correctly pointed out, $\tilde\theta_{DK}$ does approximate $\theta_{DK}$ in case ((ref)) as well if either $e(0)$ or $\theta$ is close to zero.\footnote{In case ((ref)), we have $ \mathbb{P}(Y_i = 1 \mid Z_i = 0) = \mathbb{P}( Y_i = 1, T_i = 1 \mid Z_i = 0 ) + \mathbb{P}( Y_i = 1, T_i = 0\mid Z_i = 0 ) = \mathbb{P}\{ Y_i(0) = 1\} + \bigl[ \mathbb{P}\{ Y_i(1) = 1\} - \mathbb{P}\{ Y_i(0) = 1\} \bigr] e(0)$.}

Our identification results in the previous sections show that $\theta_L$ is the sharp lower bound of the identified interval of $\theta_\textrm{avg}$ in the fuzzy design regardless of whether the full joint distribution of the outcome, treatment, and instrument is available or not. The parameter $\theta_L$ has been reported in the literature without understanding that it is the sharp lower bound on $\theta_\textrm{avg}$. For instance, dellavigna2010persuasion extensively estimate $\tilde\theta_{DK}$ by using many examples but they report $\theta_L$ as a lower bound of $\tilde\theta_{DK}$ when $T_i$'s are unobserved and hence $e(1)$ and $e(0)$ are unknown. Our results show that $\theta_L$ is always a meaningful parameter, but $\tilde\theta_{DK}$ may not. Therefore, even when information about $e(0)$ and $e(1)$ is available, $\theta_L$ is a better parameter to estimate than $\tilde\theta_{DK}$.

Indeed, if the full joint distribution of $(Y_i,T_i,Z_i)$ is available, then we recommend reporting $[\theta_L,\ \theta_U]$ along with $\theta^*$; these can be consistently estimated by their sample analogs. If $(Y_i,Z_i)$ is observed with some auxiliary information for $e(0)$ and $e(1)$ available, then $[\theta_L,\ \theta_{U_e}]$ and $[\theta_L^*,\ 1]$ should be reported. If $T_i$ is not observed at all, then the interval $[\theta_L,\ 1]$ is the best we can hope for to study either $\theta_\textrm{avg}$ or $\theta_\textrm{local}$.

Note that $\theta_L$ should be estimated all the time; it only requires data on $(Y_i,Z_i)$. Because the actual $T_i$ can be difficult to observe, researchers have used an extra micro-level survey to obtain auxiliary data on $T_i$, which seems quite costly. However, the value of an attempt to observe $T_i$ can be limited, depending on which parameter the researcher wants to learn about. For instance, if the researcher cares about the persuasion rate of the entire population, then observing $T_i$ does not add any information for the lower bound, while it can potentially improve the upper bound. If the group of compliers is of interest, then whether we observe $T_i$ or not, and how we observe it, can be relevant issues; we have $\theta_L^*\geq \theta_L$ in the second data scenario and $\theta^*$ is point identified if $(Y_i, T_i, Z_i)$ is jointly observed. If $Z_i$ is continuously distributed, the value of observing $(Y_i, T_i, Z_i)$ jointly increases dramatically as well. In summary, our identification analysis shows that the value of observing $T_i$ depends crucially on which population the researcher is interested in.

In order to illustrate the difference between DK's measure and our bounds, we have calculated them in (ref). We focus on the results reported in dellavigna2010persuasion when the outcome variable is voter turnout. We have chosen this type of study as the turnout is among the most studied outcome variables in the literature and it is naturally a binary measure. (ref) provides estimates of $\mathbb{P}(Y_i=1 \mid Z_i=z)$ and $e(z)$ for $z=0,1$, thereby enabling us to obtain the bounds based on (ref) and (ref) (ii). It can be seen that using DK's persuasion rates alone may lead to misleading conclusions because the bounds on $\theta_\textrm{avg}$ as well as those on $\theta_\textrm{local}$ are in fact wide. Moreover, the results in (ref) suggest that identification power under (ref) in this example is limited, especially for the upper bounds on $\theta_\textrm{avg}$ and $\theta_\textrm{local}$. We further illustrate these points with empirical examples in (ref).

Finally, since the parameters are partially identified, inference should also account for that. The method proposed by Stoye:07 is useful for that purpose, at least in the most favorable data scenario, in which case the sample analog principle and the delta method show that we can construct the estimators $\hat \theta_L$ and $\hat \theta_U$ that are asymptotically jointly normal. Therefore, by Stoye:07, a $(1-\alpha)$ confidence interval for $\theta_\textrm{avg}$ can be obtained by $[\hat\theta_L-c_\alpha\hat\sigma_L,\, \hat \theta_U + c_\alpha \hat \sigma_U]$, where $\hat\sigma_L$ and $\hat \sigma_U$ are the estimated standard errors of $\hat \theta_L$ and $\hat \theta_U$, respectively, and $c_\alpha$ is chosen by solving \[ \Phi\Bigl( c_\alpha + \frac{\hat \Delta}{\max(\hat\sigma_L,\hat\sigma_U)} \Bigr) - \Phi(-c_\alpha) = 1-\alpha, \] where $\Phi$ is the distribution function of the standard normal and $\hat \Delta$ is the estimated length of the identified interval.

The second data scenario is slightly more complicated, because $\theta_{U_e}$ and $\theta_L^*$ contain the min or max function that is not smooth; so, the delta method does not apply. In online Appendices (ref) and (ref), we propose a two-step method for inference to overcome this problem, which we have applied to the empirical example we discuss in (ref). In the third data scheme, confidence intervals for $\theta_\textrm{avg}$ and $\theta_\textrm{local}$ always coincide, and they can be obtained by using a one-side critical value on $\hat \theta_L$. Specifically, they are given by $[\hat\theta_L - z_{1-\alpha}\hat \sigma_L,\, 1]$, where $z_{1-\alpha}$ is the $(1-\alpha)$ quantile of the standard normal distribution. Online Appendices (ref) and (ref) provide a more detailed discussion on inference. Furthermore, see online (ref) for semiparametrically efficient estimation of the two key parameters, i.e.,\ $\theta_L$ and $\theta^*$, when exogenous covariates $X_i$ are present and integrated out.

In sum, this paper clarifies identification issues when we insert exposure as a choice variable and employ a proper causal framework that is used in policy evaluation to model two causal links (i.e.,\ $Z_i \rightarrow T_i$ and $T_i \rightarrow Y_i$).\footnote{We are grateful to an anonymous referee who provided us with insightful comments.}

Empirical Examples

Effects of Uncensored Media

In this subsection, we revisit CY2019, who conducted a field experiment in China to measure the effects of providing students with internet access to the uncensored media on various outcome variables. Excluding the existing users, the subjects in their experiments consist of four groups: (i) the control group; (ii) the control group who were encouraged to visit foreign news websites blocked by the Great Firewall; (iii) students who received free access to uncensored internet; and (iv) students who received both the access and encouragement treatments. They followed the subjects over 18 months to collect outcomes on media-related behaviors, beliefs, and attitudes among other things. It turns out that there were no differences between groups (i) and (ii) and the effects were the largest for group (iv), i.e., the access plus encouragement group (the Group-AE students from now on). To benchmark their findings, CY computed DK's measure of persuasion for the Group-AE students (see online Appendix Table A.13 of their paper for the details) and commented that “their estimated persuasion rates are of a similar magnitude to those found in authoritarian regimes that typically have highly regulated media markets” (see pp. 2323--24 in CY).

In this section, we use the data from CY to illustrate how the common practice of reporting DK's measure can lead to misleading conclusions by contrasting DK's measure of persuasion with our proposed approaches. As in CY, we focus on the Group-AE students. That is, $Z_i = 1$ if the $i$th subject is randomly assigned to the Group-AE group, and $Z_i = 0$ if the $i$th subject is randomly assigned to the control or control-encouragement group, while dropping the access only group and the existing users. In CY's two-stage least squares analysis (Table 3 in their paper), the treatment variable is $T_i = 1$ if the $i$th subject is an active user of the censorship circumvention tool and and $T_i = 0$ otherwise. We use the same treatment variable in our analysis. To replicate the results in CY in a representative but succinct way, we focus on the 11 outcome variables listed in Panel A of online Appendix Table A.13 in their paper. They represent media-related behaviors, beliefs, and attitudes and are transformed to binary variables by CY.

Recall the population version of DK's measure $f$; see (ref). In their online Appendix Table A.13, CY measure $\mathbb{P}(Y_i = 1\mid Z_i = 1) - \mathbb{P}(Y_i = 1\mid Z_i = 0)$ by the intent-to-treat effects of the Group-AE assignment, and approximate $\mathbb{P}\{ Y_i(0) = 1 \}$ using variables collected at the time of the baseline survey or by the estimates of $\mathbb{P}( Y_i = 1 \mid Z_i = 0 )$ as in $\tilde \theta_{DK}$, if the former is unavailable. Table 2 in CY and the data provided by CY indicate that the change in the exposure rate is 45.5% if the treatment status is measured by being active users, and therefore we use $e(1) - e(0) = 0.455$ for our subsequent calculations.

(ref) summarizes the empirical results. In Column (2) of Table (ref), we recompute CY's persuasion rates for the 11 outcome variables listed in Panel A of online Appendix Table A.13 in their paper. As we explained above, these estimates are based on $e(1) - e(0) = 0.455$. Out of the 11 outcome variables, the median persuasion rate is 101%, and therefore these “persuasion rates” cannot be understood as conditional probabilities. Columns (3) and (4) report our estimates of the average and local persuasion rates along with the 95% confidence intervals (in curly braces) that were obtained via 10,000 bootstrap replications. Our estimates show that (i) the average persuasion rates are only partially identified and the widths of the identified intervals are substantial for most of the outcomes, (ii) the point-identified local persuasion rates are contained by the identified intervals for the average persuasion rates and are typically closer to the upper end points of the intervals. The persuasive effects are of a relatively large magnitude in that the smallest lower end point of the confidence interval for the average persuasion is 16%. However, the original estimates overstate the magnitude by a substantial factor and mask under-identification of average persuasion rates. In short, we find that the subjects in the experiments responded to exposure to uncensored internet highly heterogeneously, indicating that it is important to go beyond the benchmark measures of DK type.

Effects of Political News

We now illustrate our proposed methods by using data from gerber2009does, who report findings from a field experiment to measure the effect of political news. We have chosen this example because it contains a credible binary instrument from the field experiment and we can also illustrate all of the three sampling scenarios as well as the case of nonbinary outcomes; for the theory on the multinomial outcome case, see online (ref). In GKB, there are three statuses in the intention to treat: a control group, an offer of free subscription to The Washington Post, and one to The Washington Times. To illustrate the usefulness of our paper, we focus on The Washington Post and drop all observations from The Washington Times subscription. That is, $Z_i = 1$ if the $i$th individual received free subscription to The Washington Post, and $Z_i = 0$ if not.

Focusing on the ITT analysis, GKB have reported ITT estimates for various outcomes $Y_i$. dellavigna2010persuasion compute persuasion rates for GKB, for which they simply set $T_i = 1$ if the $i$th individual opted into the free subscription and $T_i = 0$ if they opted out of it.\footnote{We provide the empirical results of bound analysis using the opting-into-the-free-subscription treatment variable in online (ref).} In this section, for the purpose of illustrating our identification results, we consider a different treatment variable: $T_i = 1$ if the $i$th individual read a newspaper at least several times per week and $T_i = 0$ otherwise, which is a variable that GKB kept track of in a follow-up survey. Therefore, the relevant treatment we consider differs from that of dellavigna2010persuasion, but it is whether individuals have actually read the newspaper or not. The outcome variables we consider are as follows. For the binary case, $Y_i = 1$ if the $i$th individual reported voting for the Democratic candidate in the 2005 gubernatorial election, and $Y_i = 0$ if the subject did not vote for the Democratic candidate or did not vote at all. For the multinomial case, not voting at all is treated as an outside option. We use only a subsample of the GKB data with those who responded to the follow-up survey to use information on $(Y_i, T_i, Z_i)$ jointly. After dropping observations for The Washington Times subscription and removing missing data, we summarize the GKB data in (ref). Although the joint distribution of $(Y_i, T_i, Z_i)$ is observed in this example, we also consider using the two marginals of $(Y_i,Z_i)$ and $(T_i, Z_i)$ separately, to make a comparison. The estimates are summarized in (ref). Because the size of the sample extract we use is relatively modest ($n=701$) for an interval-identified object, we report the 80% confidence intervals obtained by the inference methods described in (ref) as well as in online (ref).

First, we discuss the case where the full joint distribution of $(Y_i, T_i, Z_i)$ is used. In this data scenario, the average effect of persuasion by reading the newspaper is bounded between $7\%$ and $63\%$. In contrast, the persuasion rate for the group of compliers is point-estimated by $81\%$. It is interesting to note that the estimate of $\theta_\textrm{local}$ is so large that it is greater than the upper bound of $\theta_\textrm{avg}$. This suggests that individuals are highly heterogeneous in this example, indicating that $\tilde \theta_{DK}$ might not be a well-defined conditional probability here. Indeed, the estimate of $\tilde \theta_{DK}$ in (ref) is $\hat{\tilde \theta}_{DK} = 1.1027$, which is greater than one.

When the marginals of $(Y_i, Z_i)$ and $(T_i, Z_i)$ are used separately, the upper bound on $\theta_\textrm{avg}$ increases from $63\%$ to $78\%$. Further, $\theta_\textrm{local}$ is no longer point estimated but we only know that it is bounded between $78\%$ and $100\%$. This difference illustrates the loss of identification power if we do not observe the joint distribution of $(Y_i, T_i, Z_i)$.

Finally, we estimate the lower bound on the average persuasion rate by additionally conditioning on those who would vote even without reading the newspaper (see online (ref) for details). The resulting lower bound on the average persuasion increases from 0.0707 (0.0289) to 0.0975 (0.0554), where the numbers in the parentheses are the left-end points of the 80% confidence intervals. Therefore, the (point identified) ITT effect is 5%, while the lower bound of the average persuasion rate is about 7%, or 10% if we further focus on those who would vote without reading the newspaper.

Conclusions

We have set up a simple econometric model of persuasion, introduced several parameters of interest, and analyzed their identification. The empirical examples in (ref) as well as the examples in online (ref) demonstrate that the persuasive effects are highly heterogeneous in the settings of media and fundraising.

We have focused on the case of binary outcomes and binary treatments. In online (ref), we extend our analysis to nonbinary outcomes. If the outcome is nonbinary, then we can condition on those who would not choose the outside option without the treatment. For instance, suppose that we have three options of voting for a Republican, voting for a Democrat, or not voting at all. Then, the persuasive effect of a message supporting a Republican can be measured in a couple of different ways: focusing on those who would not have voted for a Republican without the message is one way and conditioning on those who would have voted for a Democrat (i.e.,\ voted but not voted for a Republican) is the other. In the latter case, we show that the resulting lower bound is always no smaller than that of the binary outcome case.

In general, treatments are multivalued: unordered treatments (e.g.,\ watching Fox News, CNN or MSNBC) and ordered treatments (e.g.,\ numbers of hours watching Fox News) arise naturally in applications. It would be fruitful to build on recent developments in multivalued treatments HUV2006,HV2007-handbook-2,HUV2008,HeckmanPinto, LeeSalanie to investigate identification of persuasive effects. It would also be interesting to estimate deep parameters in an economic model of persuasion by using a more structural approach in the set-up of multivalued treatments. These are topics for future research.

table[table omitted — 2,239 chars of source]
table[table omitted — 3,124 chars of source]
table[table omitted — 992 chars of source]
table[table omitted — 1,787 chars of source]