EconBase
← Back to paper

Bounding the Effect of Persuasion with Monotonicity Assumptions: Reassessing the Impact of TV Debates

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

44,020 characters · 11 sections · 52 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.
center[center omitted — 275 chars of source]
center[center omitted — 347 chars of source]
center[center omitted — 61 chars of source]

Abstract.

Televised debates between presidential candidates are often regarded as the exemplar of persuasive communication. Yet, recent evidence from TVdebate:23 indicates that they may not sway voters as strongly as popular belief suggests. We revisit their findings through the lens of the persuasion rate and introduce a robust framework that does not require exogenous treatment, parallel trends, or credible instruments. Instead, we leverage plausible monotonicity assumptions to partially identify the persuasion rate and related parameters. Our results reaffirm that the sharp upper bounds on the persuasive effects of TV debates remain modest.

\noindentKey Words: Persuasion rates; Monotonicity assumptions; Partial identification; Media; TV debates \\ \noindentJEL Classification Codes: C21, D72

\raggedbottom

Introduction

Economists have extensively documented empirical evidence demonstrating how persuasive efforts can affect the decisions and actions of consumers, voters, donors, and investors dellavigna2010persuasion. Indeed, the art of persuasion has captivated observers for centuries, from ancient rhetoricians to modern communication experts. In today's media-driven world, few events draw more attention than a live TV debate between U.S. presidential candidates, often billed as a make-or-break moment that can sway millions of undecided voters. However, recent evidence from TVdebate:23 challenges this common belief. Their large-scale study examines 56 TV debates across 31 elections in seven countries, and they find no discernible impact of those debates on vote choices, beliefs, or policy preferences.

In this paper, we revisit LP’s findings to reassess the effectiveness of TV debates from a different, robust perspective. Our framework focuses on the persuasion rate and related parameters, yet does not rely on linearity, additive separability, exogenous treatments, parallel trends, or credible instruments. Instead, it draws on a pair of monotonicity assumptions in the spirit of manski2000monotone. Because the sharp lower bounds are always trivially zero, we concentrate on the sharp upper bounds on the persuasion rate and other measures. Our results reaffirm LP’s conclusion in that these upper bounds remain small, implying that the overall effects of TV debates appear limited.

The persuasion rate, first introduced by dellavigna2007fox to evaluate the efficacy of persuasive efforts, has been extensively used to make a comparison across diverse settings. Building upon this concept, jun2023identifying formulate the persuasion rate as a causal parameter that quantifies how a persuasive message affects recipients' behavior, and they conduct a formal identification analysis based on exogenous treatment assignments or credible instruments. Extending this line of research, jun2024aprt propose the forward and backward average persuasion rates on the treated as additional causal parameters of interest, and they develop a difference-in-differences (DID) framework to identify, estimate, and conduct inference for them. Additional research on persuasion rates includes an analysis of persuasion types among compliers yu2023binary, identification of the average persuasion rate under sample selection possebom2022probability, an extension of the persuasion effect to continuous decisions kaji2023assessing, and covariate-assisted bounds on the persuasion rate without monotonicity assumptions ji2023model.

In this paper, we explore implications of two monotonicity assumptions: monotone treatment response manski1997 and monotone treatment selection manski2000monotone. The MTR assumption states that the effect of the informational treatment of interest is directional, and it has also been employed by jun2023identifying,jun2024aprt in the context of persuasion rates. In our application to TV debates, MTR implies that exposure to the debate never increases uncertainty regarding vote intentions: i.e., debates provide information that helps voters crystallize their preferences. For identifying the persuasion rate, MTR links the joint probability of two potential outcomes with the average treatment effect. The MTS assumption is new in studying the persuasion rate. In our context, it requires that the probability of vote choice consistency (voting exactly as intended at the time of the pre-election survey) not decrease as the survey date approaches election day, which seems natural.

Since the foundational work of manski1997 and manski2000monotone, employing MTR, MTS, and other types of monotonicity assumptions has been fruitful in the literature on partial identification. To sample a few, see Kreider2007JASA, vytlacil2007dummy, BSV-2012, Kreider-et-al:2012, okumura2014concave, Kim-et-al:2018, MSV-2019, and Jun2024JBES among others. Our paper is another example that shows the usefulness of the MTR and MTS assumptions.

It is worth noting that using MTR and MTS does not alter the estimands one would compute under MTR and exogenous treatments. What changes is how we interpret these estimands: the population quantities that point-identify the persuasion rate and its reverse version under MTR with exogenous treatments become the sharp upper bounds on these same measures when the exogenous treatment assumption is replaced with MTS. We illustrate these findings by reanalyzing LP’s data.

This robust interpretation of our estimands can be extended further. Both the persuasion rate and its reverse version are defined by conditioning on a counterfactual event. However, pearl1999probabilities,pearl2009causality,Pearl2015causes has advocated conditioning on an observable event rather than a counterfactual one. We show that this conceptual difference is inconsequential in our approach: our estimands serve as sharp upper bounds not only on the persuasion rate and its reverse version, but also on the probabilities of causation proposed by Pearl, when we solely rely on MTR and MTS for identification.

Our empirical findings suggest that watching TV debates has, at most, a modest impact on vote choice consistency: the estimated average treatment effect is no higher than 0.7%, while the persuasion rate is about 3% at most. This implies that relatively few voters switch from an inconsistent to a consistent vote choice by watching debates. Our results align with those of LP, even though our methodology is quite different. We also remark that the upper bound on the persuasion rate is larger than that on the average treatment effect because, under the MTR assumption, the persuasion rate is derived by rescaling the average treatment effect. The rescaling factor specifically avoids “preaching to the converted” by conditioning on relevant events. However, even by this alternative measure, the effect of TV debates remains limited.

The rest of the paper is organized as follows. In (ref), we introduce the causal persuasion rate and related parameters, and discuss the MTR assumption, which will be used throughout. (ref) focuses on identification, highlighting that the estimands that point-identify the parameters of interest under exogenous treatments become sharp upper bounds on those parameters when the MTS assumption is imposed. In (ref), we consider pearl1999probabilities’s probabilities of causation and show that the same sharp upper bounds apply under MTR and MTS. In (ref), we apply our methodology to revisit LP’s data. In (ref), we provide concluding remarks. Finally, in the appendix, we give the proofs of our main theoretical results.

Preliminaries

For $d\in \{0,1\}$, let $Y(d)$ denote a binary potential outcome, $D$ the observed binary treatment, and therefore, the observed outcome is $Y = DY(1) + (1-D) Y(0)$. Because our identification results do not rely on covariates, we focus on the case without them. Alternatively, one can view our assumptions and identification results as “implicitly conditional on covariates.” The causal parameters we consider are:

align*[align* omitted — 137 chars of source]

The conditional probability $\theta$ is known as the average persuasion rate (APR), whereas $\theta^{(r)}$ is its reverse version (R-APR). Here, $\theta$ measures the proportion of individuals, among those who would not have taken the action of interest in the absence the treatment, who would have changed their behavior in the presence of the treatment. The reverse version $\theta^{(r)}$ focuses on individuals who would have taken the action of interest with the treatment, and it answers how many of them would not have done so without the treatment. See dellavigna2007fox, dellavigna2010persuasion, and jun2023identifying,jun2024aprt for more discussion on these and related parameters.

We start with making two assumptions.

assumption[Regularity] For all $(y,d) \in \{0,1\}^2$, $0 < \mathbb{P}\{ Y = y, D= d\} < 1$.
assumption[Monotone Treatment Response (MTR)] We have $Y(1) \geq Y(0)$ with probability one.

(ref) imposes weak regularity conditions on the observed data. (ref) has been used for partial identification manski1997,manski2000monotone,Jun2024JBES or in the context of informational treatments jun2023identifying,jun2024aprt. In the latter case, it is often interpreted as an assumption that says, “Persuasive messages are directional, and there is no backlash.”

In our application, (ref) is motivated by LP’s research design, in which respondents are surveyed twice: first to elicit pre-election voting intentions, and later to record their final post-election choices. The treatment indicator $D$ equals $1$ if the initial survey occurs after a TV debate and $0$ otherwise. The outcome variable $Y$, referred to as vote choice consistency, equals $1$ if a respondent’s pre-election intention matches their eventual vote, and $0$ otherwise; notably, $Y=0$ whenever the initial response is “undecided.” Formally, let $Y(1)$ and $Y(0)$ represent vote choice consistency with and without a debate, respectively. Panel A of (ref) illustrates the coding of these potential outcomes.

In this context, (ref) means that a debate may help clarify a voter's preference but should not confuse a previously decided voter. For instance, Scenario 1 in Panel A depicts a respondent who is undecided without the debate ($Y(0)=0$) but forms a consistent preference for Candidate A after watching it ($Y(1)=1$), a case that satisfies MTR. In contrast, Scenario 2 illustrates the reverse case, where a respondent who would have been consistent without the debate becomes undecided (and thus inconsistent) after watching it. We consider this “confusion” scenario implausible if debates serve to provide information and clarify preferences; therefore, MTR excludes this possibility. Finally, Scenarios 3 and 4 show examples where $Y(0) = Y(1) = 0$, which are consistent with MTR but yield no variation in the outcome.

It is worth noting that the plausibility of the MTR assumption depends heavily on the definitions of the treatment and outcome variables. For instance, cavgias2024media analyze the 1989 Brazilian election, where the treatment is defined as having access to a biased TV debate. In that context, the debate, which was widely regarded as biased towards the right-wing candidate, could persuade a voter to switch candidates (e.g., from Lula to Collor). Consequently, a treated voter would appear inconsistent ($Y(1)=0$) relative to their pre-debate baseline, while an untreated voter would remain consistent ($Y(0)=1$). This creates a violation of MTR. In contrast, our treatment is defined by the timing of the survey. If a debate persuades a voter to switch candidates, a respondent surveyed after the debate ($D=1$) will report their new preference, which matches their final vote ($Y(1)=1$). Meanwhile, if surveyed before the debate ($D=0$), they report their old preference, which contradicts their final vote ($Y(0)=0$). Thus, in our design, persuasion actually reinforces the MTR condition ($Y(1) \geq Y(0)$). Therefore, we view “the chance of the debate causing confusion about their own preferences” (Scenario 2) as a primary threat to our identification strategy.

Under (ref), the population consists of three (nonempty) groups: (i) never-persuadable (NP): $Y(1) = 0, Y(0) = 0$; (ii) already-persuaded (AP): $Y(1) = 1, Y(0) = 1$; and (iii) treatment-persuadable (TP): $Y(1) = 1, Y(0) = 0$. As we never observe the pair $\{Y(1), Y(0) \}$ jointly, we cannot assign a type to each observational unit; however, we can aim to identify the share of the three groups. Using the group probabilities, we can rewrite the parameters of interest as

align[align omitted — 267 chars of source]

We will now consider conditions under which we can point or partially identify the causal parameters.

Identification

We first consider the standard case of exogenous treatment as a benchmark.

assumption[Exogenous Treatment] For $d\in\{0,1\}$, $Y(d)$ is independent of $D$.

(ref) ensures that $\mathbb{P}(Y=1\mid D=1) > 0$ and $\mathbb{P}(Y=1\mid D=0) < 1$, and we define the following population quantities:

align*[align* omitted — 221 chars of source]

which are directly identified from the data. We now state the following benchmark result; see, e.g., jun2023identifying.

lemma[Benchmark] Under (ref), we have \[ \mathbb{P}\{ \mathrm{NP} \} = \mathbb{P}(Y=0\mid D=1) \quad \text{and}\quad \mathbb{P}\{ \mathrm{AP} \} = \mathbb{P}(Y=1\mid D=0). \] In particular, $\theta$ and $\theta^{(r)}$ are point-identified by $\theta_U$ and $\theta^{(r)}_U$, respectively.

(ref) is popular, but it can be restrictive, as it precludes endogenous treatment assignment. Also, we are implicit here, but (ref) is often stated after conditioning on a sufficiently large vector of covariates, which can be too demanding in practice. To go beyond (ref), we now consider an alternative assumption that allows endogenous treatment, namely monotone treatment selection.

assumption[Monotone Treatment Selection (MTS)] For all $d\in\{0,1\}$, we have \[ \mathbb{P}\{ Y(d) = 1\mid D=1 \} \geq \mathbb{P}\{ Y(d) = 1\mid D=0 \}. \]

In LP's study, the key identifying assumption is that, once potential confounders are accounted for, the date of the debate is uncorrelated with vote choice consistency. Therefore, roughly speaking, their event study design is based on (conditional) exogenous treatment assignment. However, controlling for all possible confounders in an observational study is inevitably challenging, even though LP makes a convincing case.

As a robust alternative, we employ the MTS assumption introduced by manski2000monotone. Formally, it requires that those in the treatment group be no less likely to “succeed” than those in the control group. Even if $D$ is a pure choice variable, (ref) stays valid as long as those who chose $D=1$ are not strictly inferior on average to those who opted out. In the context of LP, being surveyed later in the election cycle (i.e., \(D=1\)) naturally places respondents closer to the election day, and hence it makes vote consistency more likely, rendering MTS a plausible condition even without needing to control for covariates. See (ref) for more details.

theoremSuppose that (ref) hold. The sharp bounds on $\mathbb{P}\{ \mathrm{NP} \}$ and $\mathbb{P}\{ \mathrm{AP} \}$ are given by \begin{align*} \mathbb{P}\{ \mathrm{NP} \} \in \left[ \mathbb{P}( Y = 0 \mid D=1), \mathbb{P}( Y = 0) \right] \quad and\quad \mathbb{P}\{ \mathrm{AP} \} \in \left[ \mathbb{P}( Y = 1 \mid D=0), \mathbb{P}( Y = 1) \right], \end{align*} where $\mathbb{P}\{ \mathrm{TP}\} = 1- \mathbb{P}\{ \mathrm{NP}\} - \mathbb{P}\{ \mathrm{AP}\}$. In particular, the sharp identified intervals on $\theta, \theta^{(r)}$ are given by \begin{align*} 0 \leq \theta \leq \theta_U \quad and \quad 0 \leq \theta^{(r)} \leq \theta^{(r)}_U. \end{align*} \proof See the appendix. \qed

(ref) shows that, for both $\theta$ and $\theta^{(r)}$, the same population quantities $\theta_U$ and $\theta^{(r)}_U$ arise whether we assume (ref) or (ref), although their interpretation changes. For example, under (ref), $\theta$ is point-identified by $\theta_U$, while that same $\theta_U$ becomes the sharp upper bound on $\theta$ under (ref). Consequently, a plug-in estimator of $\theta_U$ is consistent for $\theta$ under exogeneity, and it will be consistent for the sharp upper bound on $\theta$ under the weaker assumption of MTS.

dellavigna2007fox and dellavigna2010persuasion propose the following quantity as the persuasion rate: \[ \theta_{DK} = \frac{\mathbb{P}[Y=1\mid Z=1] - \mathbb{P}[Y=1\mid Z=0]}{\mathbb{P}[D=1\mid Z=1] - \mathbb{P}[D=1\mid Z=0]} \times \frac{1}{\mathbb{P}[Y=0\mid Z=0]}, \] where $Z$ is a binary instrument. Under full compliance, that is, when $D=Z$ almost surely, this quantity coincides with the upper bound on the persuasion rate derived in (ref). More generally, under partial compliance, when $D \neq Z$, it has been pointed out that $\theta_{DK}$ can be misleading and may inflate the causal persuasion rate jun2023identifying.

In addition, (ref) establishes bounds on the shares of the NP, AP, and TP groups. In particular, $\mathbb{P}( Y = 0 \mid D=1)$ and $\mathbb{P}( Y = 1 \mid D=0)$ are now interpreted as sharp lower bounds on $\mathbb{P}\{ \mathrm{NP} \}$ and $\mathbb{P}\{ \mathrm{AP} \}$, respectively. Further, the proof of (ref) shows that (ref) yield \[ 0 \leq \mathbb{P}\{Y(1) = 1\} - \mathbb{P}\{ Y(0) = 1\} \leq \mathbb{P}( Y = 1 \mid D=1) - \mathbb{P}(Y=1\mid D=0), \] where the bounds are sharp. The sharp bounds on the share of TP follow from here because $\mathbb{P}\{\mathrm{TP}\} = \mathbb{P}\{Y(1) = 1\} - \mathbb{P}\{ Y(0) = 1\}$ equals the average treatment effect (ATE) under MTR.

Estimation and Inference

We now describe our approach to estimation and inference for the parameters of interest. Since inference for $\theta$ and $\theta^{(r)}$ follows the same principle, we will focus on the former.

Persuasion Rates

We obtain plug-in estimators for $\theta_U$ by using the coefficients from an OLS regression of $Y$ on $D$ with an intercept; e.g., the coefficient of $D$ corresponds to the numerator of $\theta_U$. If (ref) is maintained, standard two-sided confidence intervals are applicable because $\theta_U$ point-identifies $\theta$ in this case (see (ref)). Under the weaker (ref), however, $[0, \theta_U]$ is the sharp identified bounds on $\theta$, making one-sided inference appropriate.

Specifically, under (ref), the theoretical lower bound is known to be zero. Accordingly, we construct the $(1-\alpha)$ confidence interval as \[ \mathrm{CI} = \left[ 0, \ \hat{\theta}_U + c_\alpha \cdot \widehat{\mathrm{SE}}_U \right], \] where $\hat{\theta}_U$ is the estimator for the upper bound, $\widehat{\mathrm{SE}}_U$ is the corresponding standard error (computed via the delta method), and $c_\alpha = \Phi^{-1}(1-\alpha)$ is the standard one-sided critical value (e.g., $1.645$ for $\alpha=0.05$).

Notably, this confidence interval is valid for both the true parameter $\theta$ and the identified set $[0, \theta_U]$. While general partial identification methods typically require larger critical values to cover the set Imbens/Manski:04, our setting is simpler because the lower bound is fixed at zero. Consequently, the coverage probability for the identified set coincides with that of the parameter at the least favorable configuration (i.e., $\theta = \theta_U$). Thus, the standard one-sided critical value $c_\alpha$ ensures asymptotic coverage of at least $1-\alpha$ for both targets.

Furthermore, uniformity issues are simpler than in the general cases discussed in e.g., canay2017practical too. Since $\mathbb{P}(\theta \in \mathrm{CI}) \geq \mathbb{P}\{ \theta_U(P) \in \mathrm{CI} \}$ for any data-generating process $P\in \mathcal{P}$ with the corresponding $\theta_U(P) \geq \theta \geq 0$, the coverage properties of $\mathrm{CI}$ will be uniform in $P\in \mathcal{P}$, provided that the normal approximation of $\{\hat{\theta}_U - \theta_U(P)\}/\widehat{\mathrm{SE}}_U$ is uniform. See the online appendices for more details.

Specification Testing

The identification result under (ref) implies that the upper bound must be non-negative ($\theta_U \geq 0$). This restriction is testable. Following the logic of Stoye:07, we can construct a confidence set that becomes empty if the data provide strong evidence against the validity of the identification assumptions. Specifically, we define the modified confidence set $\mathrm{CI}_{\mathrm{mod}}$ as: \[ \mathrm{CI}_{\mathrm{mod}} :=

cases[0, \ \hat\theta_U + c_\alpha \widehat{\mathrm{SE}}_U] & if \hat\theta_U + c_\alpha \widehat{\mathrm{SE}}_U \geq 0, \\ \emptyset & otherwise.

\] The event $\mathrm{CI}_{\mathrm{mod}} = \emptyset$ corresponds to rejecting $H_0: \theta_U \geq 0$ in favor of $H_1: \theta_U < 0$ at significance level $\alpha$, which indicates that MTR or MTS is violated.

Shares of Already-Persuaded and Never-Persuadable Types

As before, if (ref) is maintained, standard two-sided confidence intervals are applicable because the proportions of AP and NP types are point-identified (see (ref)). Under the weaker (ref), however, the situation differs from that of persuasion rates. Here, the sharp identified sets require the bounds at both ends be estimated (see (ref)). To construct confidence intervals for these parameters, we adopt the method of Imbens/Manski:04 and Stoye:07, which adjusts the critical value to account for the estimation of the length of the identified set.

Let $[\theta_{s,L}, \theta_{s,U}]$ denote the identified set for type $s \in \{\mathrm{AP}, \mathrm{NP}\}$. The $(1-\alpha)$ confidence interval is given by \[ \mathrm{CI}_s = \left[ \hat\theta_{s,L} - c_\alpha \widehat{\mathrm{SE}}_{s,L}, \quad \hat\theta_{s,U} + c_\alpha \widehat{\mathrm{SE}}_{s,U} \right], \] where $c_\alpha$ is chosen to satisfy \[ \Phi\left( c_\alpha + \frac{\hat{\Delta}_s}{\max(\widehat{\mathrm{SE}}_{s,L}, \widehat{\mathrm{SE}}_{s,U})} \right) - \Phi(-c_\alpha) = 1-\alpha, \] with $\hat{\Delta}_s$ representing the estimated length of the interval. This approach ensures (uniformly) valid coverage for the true parameter $\theta_s$ while mitigating the conservatism inherent in intervals designed to cover the entire identified set. We refer to Stoye:07 for formal details.

Bounding the Probabilities of Causation

In this section, we discuss a few related parameters that have been discussed in the causal inference literature. Both $\theta$ and $\theta^{(r)}$ are conditioned on a counterfactual outcome, which is not directly observed. Advocating the idea that we should consider probabilities of counterfactual events given the information that is observed, pearl1999probabilities introduces the following parameters and calls them probabilities of causation:

align*[align* omitted — 187 chars of source]

where PS, PN, and PNS, respectively, denote the probability of sufficiency, the probability of necessity, and the probability of necessity and sufficiency. For further discussion, see, e.g., pearl1999probabilities,Yamamoto:2012,Dawid2014fitting,Dawid2022effects,Ding:2024:arXiv:prob_necessity. These quantities can also be expressed as

align*[align* omitted — 322 chars of source]

Under (ref), we have that $\theta = \mathrm{PS}$, $\theta^{(r)} = \mathrm{PN}$, and $\mathrm{PNS} = \mathbb{P}( Y = 1 \mid D=1) - \mathbb{P}(Y=1\mid D=0)$, as initially noted by pearl1999probabilities. More interestingly, under (ref), we have the following results.

theoremSuppose that (ref) hold. Then, the sharp identified intervals on $\mathrm{PS}$, $\mathrm{PN}$, and $\mathrm{PNS}$, respectively, are given by \begin{align*} 0 \leq \mathrm{PS} \leq \theta_U, \; 0 \leq \mathrm{PN} \leq \theta^{(r)}_U, \; and \; 0 \leq \mathrm{PNS} \leq \mathbb{P}( Y = 1 \mid D=1) - \mathbb{P}( Y = 1 \mid D=0), \end{align*} \proof See the appendix. \qed

The sharp bounds on $\textrm{PNS}$ are immediate from (ref) because $\mathbb{P}\{\mathrm{TP}\}$ equals ATE under MTR. The main point here is that the sharp identified bounds on $\mathrm{PS}$ and $\mathrm{PN}$ under (ref) coincide with those on $\theta$ and $\theta^{(r)}$, respectively.

Bounds under Alternative Assumptions

In this section, we derive the identified regions for the persuasion rate $\theta$ and the reverse persuasion rate $\theta^{(r)}$ under relaxed assumptions. We consider three scenarios: (i) no monotonicity assumptions; (ii) MTR only; and (iii) MTS only.

theoremLet (ref) hold. The sharp identified intervals for $\theta$ and $\theta^{(r)}$, respectively, are given as follows: \begin{enumerate} • With no other assumptions: $\theta \in [0, 1]$ and $\theta^{(r)} \in [0, 1]$. • With (ref) only: $\theta \in [0, 1]$ and $\theta^{(r)} \in [0, 1]$. • With (ref) only: $\theta \in [0, \overline{\theta}_{\text{MTS}}]$ and $\theta^{(r)} \in [0, \overline{\theta}^{(r)}_{\text{MTS}}]$, where \begin{align*} \overline{\theta}_{MTS} &:= \min\left\{ \frac{\mathbb{P}(Y=1\mid D=1)}{\mathbb{P}(Y=1,D=1)+\mathbb{P}(Y=0,D=0)},\ 1 \right\}, \\ \overline{\theta}^{(r)}_{MTS} &:= \min\left\{ \frac{\mathbb{P}(Y=0\mid D=0)}{\mathbb{P}(Y=1,D=1)+\mathbb{P}(Y=0,D=0)},\ 1 \right\}. \end{align*} \end{enumerate}
proofSee the online appendix. The argument closely parallels the proof of (ref).

(ref) shows that the MTR assumption alone provides no identifying power for either $\theta$ or $\theta^{(r)}$, whereas MTS alone may yield non-trivial bounds. Intuitively, the role of MTR is restricted to relating the joint probabilities of the potential outcomes to their marginal distributions, whereas MTS has identifying power for the marginals of the potential outcomes themselves.

Furthermore, note that

align*[align* omitted — 325 chars of source]

Thus, whenever MTS yields a non-trivial upper bound on $\theta$, the sharp bound on $\theta^{(r)}$ becomes trivial. In contrast, we have \[ \overline{\theta}_{\text{MTS}} \geq \theta_U \quad\text{and}\quad \overline{\theta}^{(r)}_{\text{MTS}} \geq \theta^{(r)}_U \] by construction, where both $\theta_U$ and $\theta^{(r)}_U$ can be non-trivial. We therefore adopt the joint assumption of MTR and MTS as our preferred identifying restrictions.

We now turn to the sharp identified regions for $\mathrm{PS}$, $\mathrm{PN}$, and $\mathrm{PNS}$ under the three scenarios. Define

align*[align* omitted — 451 chars of source]
theoremLet (ref) hold. The sharp identified intervals for $\mathrm{PS}$, $\mathrm{PN}$, and $\mathrm{PNS}$, respectively, are given as follows: \begin{enumerate} • With no other assumptions: $\mathrm{PS} \in [0, 1]$, $\mathrm{PN} \in [0, 1]$, and $\mathrm{PNS} \in [0, \overline{\mathrm{PNS}}_{\text{NA}}]$. • With (ref) only: $\mathrm{PS} \in [0, 1]$, $\mathrm{PN} \in [0, 1]$, and $\mathrm{PNS} \in [0, \overline{\mathrm{PNS}}_{\text{NA}}]$. • With (ref) only: $\mathrm{PS} \in [0, \overline{\mathrm{PS}}_{\text{MTS}}]$, $\mathrm{PN} \in [0, \overline{\mathrm{PN}}_{\text{MTS}}]$, and $\mathrm{PNS} \in [0, \overline{\mathrm{PNS}}_{\text{MTS}}]$. \end{enumerate}
proofCase (i) is given in tian2000probabilities, and case (ii) is given in tian2000probabilities. For case (iii), see the online appendix. The argument closely parallels the proof of (ref).

As in the case of $\theta$ and $\theta^{(r)}$, the MTR assumption alone does not have any identification power for any of the probabilities of causation. Our results under MTS only, i.e., (ref)(iii), relate to the bounds derived by tian2000probabilities under exogeneity. Comparing the two sets of the results, we find that the upper bounds yielded by MTS alone are identical to those under exogeneity. The additional identification power provided by exogeneity applies exclusively to the lower bounds.

Finally, (ref) summarizes the hierarchy of bounds for the persuasion rates and the probabilities of causation under different combinations of assumptions.

Empirical Example

In this section, we revisit the findings of LP to reevaluate the effectiveness of TV debates in persuading (undecided) voters. The debates examined in LP took place between 5 and 44 days (on average 24 days) before the election, precisely when vote choice consistency tends to increase most rapidly. Utilizing an event-study design based on 331,000 respondent-debate-election observations, LP isolate the impact of debates occurring one to three days post-event relative to the broadcast day. Their specification controls for secular time trends, sociodemographic covariates, and debate-by-election fixed effects, assuming the specific timing of the debate is exogenous conditional on these factors. See (ref) for the full econometric specification. Interestingly, LP find no significant differences in how voters' choices evolved before versus after debate periods, rejecting any effect larger than $0.5$ percentage points. This null result holds across diverse voter types, different election contexts, and even tight, uncertain races.

In our empirical analysis, we adopt an alternative strategy. Rather than controlling for the extensive set of variables used in LP, we restrict the sample to a narrow window of $[-3,+3]$ days relative to the debate. This localized approach focuses on the period immediately surrounding the event. We define the pre-debate days as the control group ($D=0$) and the debate day together with the post-debate days as the treatment group ($D=1$). Panel B of (ref) provides a visualization of the treatment definition ($D$) and the sample window.

Crucially, our method does not require conditioning on additional covariates. Unlike standard exogeneity assumptions, the MTR and MTS assumptions in this setting do not rely on covariate adjustment for their validity. MTR is imposed at the individual level, while MTS will be motivated without controlling for covariates. Therefore, our approach avoids reliance on functional form restrictions or strict exogeneity assumptions.

First, consider the plausibility of MTR, for which it is helpful to recall that the outcome variable is vote choice consistency defined by the agreement between respondents' reported voting intentions prior to the election and their realized vote choices reported after the election; if the voter is undecided before the election, then the outcome is coded as inconsistent voting whatever the final vote is. Panel A of (ref) explains how the potential outcomes would be coded in several scenarios. The idea is that because debates provide information, the likelihood of a voter remaining undecided should decrease in the presence of a debate. For example, imagine a voter who would vote for Candidate A before a debate, would become disappointed and undecided after watching the debate, and would end up just not voting at all as a result. If this voter is observed before the debate, then $Y(0) = 0$ will be observed, whereas if she is surveyed after the debate, then $Y(1) = 0$ will be observed. Therefore, MTR is still satisfied in this case. From this perspective, information revealed during a debate (including details about candidates' traits such as health or competence) helps voters refine and finalize their preferences prior to voting. Even if such information leads some voters to revise their initial intentions, this does not violate MTR, provided that debate exposure does not induce voters to become “confused” about their own preferences.

Next, regarding the plausibility of MTS, respondents surveyed later in the electoral cycle are closer to the election day and thus more likely to exhibit consistent voting behavior on average. This supports the unconditional MTS assumption. Further, our use of a $\pm$3-day estimation window around the debate adds extra support for its plausibility. While respondent composition might shift over a long campaign (e.g., urban voters being surveyed earlier than rural voters), such structural shifts are implausible within such a short period. We verify this by showing that observable covariates are balanced between the pre- and post-debate groups. Specifically, when we regress the treatment indicator on the sociodemographic controls used in the paper (gender, age, income quartiles, employment status, and education), the resulting p-value for the joint F-test of significance (with standard errors clustered at the debate level) is 0.445. This failure to reject the null hypothesis supports the validity of the unconditional MTS assumption in our local setting.

(ref) presents the estimation results divided into two panels. Panel A reports the estimates and upper confidence bounds of (the sharp upper bounds on) the Average Treatment Effect (ATE), the Average Persuasion Rate (APR, denoted as $\theta$), and the Reverse Average Persuasion Rate (R-APR, denoted as $\theta^{(r)}$). In the first row, labeled “All,” we consider the full sample ($N=42,136$) and find that the estimated impact of TV debates is small. The point estimate of the upper bound on ATE is 0.68 percentage points (standard error 0.56), with a 95% upper confidence bound (UCB) of 1.60%. These results align closely with those in LP, although we rely on a notably different methodology. Since the share of Treatment-Persuadable (TP) respondents is identical to ATE in our setting, these small values imply that very few voters' consistency relies exclusively on watching TV debates.

The middle three columns in Panel A report the upper bounds on APR, its standard error (SE), and its 95% UCB, while the last three columns present analogous estimates for R-APR. Recall that these rates rescale the share of TP relative to different denominators: \[ \textrm{APR} = \frac{\textrm{share of TP}}{\textrm{share of TP} + \textrm{share of NP}}, \quad \textrm{R-APR} = \frac{\textrm{share of TP}}{\textrm{share of TP} + \textrm{share of AP}}. \] Because both denominators are less than one, APR and R-APR are mechanically no less than ATE. Specifically, APR is estimated at 3.37% with a UCB of 7.91%, while R-APR is estimated at 0.84% with a UCB of 1.98%. The APR estimate exhibits greater amplification than that of R-APR because the share of Never-Persuadable (NP) voters is significantly smaller than that of Already-Persuaded (AP) voters. This will be confirmed in Panel B of (ref).

Economically, these measures answer different counterfactual questions. APR considers those who would not have been consistent without the debate and asks what fraction would change their behavior if they watched it. In contrast, R-APR considers those who were consistent after watching the debate and asks what fraction would have differed without it. The former effect is small, but the latter is even smaller. We note that the reported 95% confidence intervals are one-sided because the lower bounds are trivially zero, so only the upper bounds require estimation (see (ref)).

Subsequent rows of Panel A of (ref) reveal notable heterogeneity across demographic groups and countries. Younger voters (Age $< 50$) appear more susceptible to persuasion than older voters: the estimated sharp upper bound on ATE for the under-50 group is 1.06% (UCB 2.45%), compared to just 0.40% (UCB 1.36%) for the over-50 group. Similarly, APR for younger viewers is more than double that of older viewers (4.83% versus 2.13%). When disaggregated by country, the results for the U.S.\ and the U.K.\ are qualitatively similar, though the U.S.\ sample shows a slightly higher persuasion rate: APR in the U.S.\ is estimated at 5.09% (UCB 12.42%), compared to 3.69% (UCB 12.63%) in the U.K. However, the standard errors for these subsamples are larger, indicating less precision than in the pooled analysis.

Panel B of (ref) provides insight into the shares of different types: the vast majority of the population belongs to the AP or NP categories. For the full sample, the estimated proportion of AP types is bounded between 79.88% and 80.25%, while the proportion of NP types lies between 19.44% and 19.75%. The tight identification region confirms that approximately 80% of the viewing population was already committed to the outcome prior to the debate. This “ceiling effect” is more pronounced in the U.S.\ subsample, where the lower bound on the share of the AP type is 87.30%, slightly higher than the U.K.'s 83.59%.

Overall, our findings indicate that TV debates exert, at best, a modest influence on vote choice consistency, echoing LP's results. However, when focusing on APR---which excludes AP voters from the relevant population---the potential scope for debate effects is larger. Even so, this evidence remains suggestive because (i) the identified parameter is merely an upper bound on the true persuasive effect, and (ii) the standard errors associated with the APR estimates are relatively large.

Conclusions

We have revisited LP's findings on the effects of TV debates by using a new framework involving the persuasion rate. Instead of relying on exogenous treatments or other strong identifying assumptions, we have leveraged two weak monotonicity assumptions (MTR and MTS) to derive sharp bounds on various measures of persuasive effects, including the persuasion rate and its reverse version. Our analysis has shown that the impact of TV debates on vote choice consistency remains modest, which is consistent with LP's conclusion. In addition, we have highlighted that our bounds continue to serve as sharp bounds on Pearl's probabilities of causation under the same monotonicity assumptions, further underscoring the robustness of our approach. Overall, our results suggest that while TV debates may provide helpful information to voters, their influence on electoral choices is limited.

Finally, we summarize a practical roadmap for empirical implementation, as illustrated in Figure (ref). We consider a baseline setting with binary outcomes, binary treatments, and observed covariates under the MTR assumption. The first step is to assess data completeness. When outcomes or treatments are missing for some units, the sample selection bounds of possebom2022probability are applicable under treatment exogeneity. In the absence of sample selection, the appropriate estimator depends on the treatment assignment mechanism. Under exogenous treatment, persuasion rates are point identified by sample analogs, as shown in Lemma (ref). When treatment is endogenous, identification requires additional structure, such as instrumental variables jun2023identifying,yu2023binary or panel data combined with parallel trends assumptions jun2024aprt. If neither valid instruments nor panel data are available, the bounds developed in this paper remain informative under the extra assumption of MTS.

An important direction for future research is to broaden the current framework to settings beyond the reach of existing methods, including cases featuring both treatment endogeneity and sample selection, and to develop identification and inference procedures that remain informative under such combined challenges.

appendix