EconBase
← Back to paper

Learning the Effect of Persuasion via Difference-In-Differences

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

169,304 characters · 22 sections · 81 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.
center[center omitted — 519 chars of source]
center[center omitted — 348 chars of source]
center[center omitted — 63 chars of source]

Abstract. We develop a difference-in-differences framework to measure the persuasive impact of informational treatments on behavior. We introduce two causal parameters, the forward and backward average persuasion rates on the treated, which refine the average treatment effect on the treated. The forward rate excludes cases of “preaching to the converted,” while the backward rate omits “talking to a brick wall” cases. We propose both regression-based and semiparametrically efficient estimators. The framework applies to both two-period and staggered treatment settings, including event studies, and we demonstrate its usefulness with applications to a British election and a Chinese curriculum reform.

\noindentKey Words: Persuasion Rate, Treatment Effects, Media, Voting Behavior, Event Study Design. \\ \\ \noindentJEL Classification Codes: C22, C23, D72, L82

\raggedbottom

Introduction

\setcounter{footnote}{0}

Since the dawn of civilization, the meaning of persuasion has evolved from Aristotle's Rhetoric to Jane Austen's Persuasion and to curriculum reform in China cantoni2017curriculum. Persuasion is integral to both democracy and the market economy. Economists have empirically measured how persuasive efforts influence the actions of consumers, voters, donors, and investors dellavigna2010persuasion.

A widely used measure of persuasive effectiveness is the persuasion rate, introduced by dellavigna2007fox in studying biased news media and voter behavior. This measure has become a standardized metric for comparing persuasion effects across settings. For example, Table 1 in dellavigna2010persuasion and Figure 7 in Bursztyn:Yang:22\footnote{Bursztyn:Yang:22 provide a meta-analysis showing that experimental interventions correcting misperceptions can alter behavior; their Figure 7 summarizes persuasion rates based on dellavigna2007fox.} compile estimates from numerous studies. jun2023identifying formalize the persuasion rate as a causal parameter capturing the effect of a persuasive message on recipients’ behavior, identifying it under exogenous treatment or valid instrumental variables. These conditions, however, often fail in observational settings. We address this limitation by introducing two causal parameters, the forward and backward average persuasion rates on the treated, and developing a difference-in-differences (DID) framework for their identification, estimation, and inference.

To illustrate the idea, consider the following example. Let $D$ denote a binary indicator for exposure to a media platform delivering a directional message favoring a particular political party. For $d \in \{0,1\}$, let $Y(d)$ be a binary potential outcome indicating whether the agent votes for the favored party. Since the message is directional, assume $Y(1) \geq Y(0)$ almost surely; we refer to this as the no-backlash assumption. The average treatment effect on the treated (ATT) is then \[ \mathbb{P}\{ Y(1)=1 \mid D=1\} - \mathbb{P}\{ Y(0)=1 \mid D=1\} = \mathbb{P}\{ Y(1)=1,\, Y(0)=0 \mid D=1\},\footnote{This equality follows from $\mathbb{P}\{ Y(1)=1\mid D=1\} = \mathbb{P}\{ Y(1)=1,\, Y(0)=1\mid D=1\} + \mathbb{P}\{ Y(1)=1,\, Y(0)=0\mid D=1\}$ and $\mathbb{P}\{ Y(0)=1\mid D=1\} = \mathbb{P}\{ Y(0)=1,\, Y(1)=1\mid D=1\}$ under no backlash.} \] which does not address the question of what constitutes the appropriate baseline population for measuring the net persuasive effect of the treatment. Specifically,

equation[equation omitted — 232 chars of source]

and the inequality can be severely strict when $\mathbb{P}\{ Y(0)=0 \mid D=1\}$ or $\mathbb{P}(Y=1 \mid D=1)$ is small. For example, if only a small fraction of treated individuals are potential targets for persuasion, then ATT will remain small even when the message is highly effective in changing their behavior.

The conditional probabilities on the right-hand side of (ref) define our two persuasion measures: the forward and backward average persuasion rates on the treated, abbreviated as FPR and BPR, respectively.\footnote{The full abbreviations FAPRT and BAPRT sound overly long. Since all new parameters introduced in this paper are average persuasion rates on the treated, we drop “A” and “T” in the acronyms.} Compared to ATT, FPR isolates the persuasive effect by focusing on treated individuals who would not take the action without exposure, thereby excluding cases of “preaching to the converted.” This makes FPR generally larger than ATT whenever some treated individuals would act in line with the treatment even without being exposed. The term “forward” reflects that it starts from the message’s intended audience. In contrast, BPR focuses on treated individuals who actually took the action and measures how many would have behaved differently without exposure, thereby excluding “talking to a brick wall” cases, meaning individuals who would never take the action of interest regardless of treatment. The term “backward” reflects that it starts from observed takers and looks backward to assess whether treatment was necessary, which also makes BPR generally larger than ATT whenever such never-takers are present among the treated.

The two measures are complementary. FPR directly captures the influence of the message on the target audience but is based on a counterfactual subpopulation not observed in the data. BPR addresses this limitation by starting from an observed subgroup and quantifying the necessity of the informational treatment for achieving the observed behavior in the treatment group.

BPR coincides with the probability of necessity (PN) of pearl1999probabilities.\footnote{It is also referred to as the probability of causal attribution in political science Yamamoto:2012.} In contrast, FPR differs from Pearl’s probability of sufficiency (PS), which in our context corresponds to the persuasion rate on the untreated, $\mathbb{P}\{ Y(1)=0 \mid Y(0)=1,\, D=0\}$. We do not consider PS because, like the average treatment effect on the untreated, it cannot be point-identified using DID. See (ref) for further discussion.

The idea of “rescaling” to focus on a relevant event or subpopulation has a long history: for example, Bayes1763letter derived conditional probability by rescaling joint probability; imbens1994late obtained the local average treatment effect (LATE) by rescaling the intent-to-treat effect; and jun2023identifying derived the persuasion rate and local persuasion rate by rescaling the average treatment effect (ATE) and LATE, respectively. Following the same principle, we rescale ATT to define FPR and BPR under the no-backlash assumption. If the no-backlash assumption is violated, then scaling ATT leads to lower bounds for these measures. As with joint and conditional probabilities, ATT, FPR, and BPR are related but generally convey distinct information.

Beyond articulating identification issues for FPR and BPR in panel settings without instruments, we address estimation and inference in depth. The paper makes three main contributions. First, we clarify the relationships between FPR, BPR, and ATT under both backlash and no-backlash scenarios; in particular, without backlash, identification of FPR and BPR reduces to that of ATT. For ATT identification, we build on Abadie:05 and CS:2021, assuming parallel trends after conditioning on observed covariates. We analyze both the canonical two-period case and staggered treatment adoption. This contributes to the growing DID and event-study literature on heterogeneous treatment effects de2020two,Ding:2024:arXiv, causal effect estimation with panel data Ding_Li_2019,AER:event-study-design,arkhangelsky2021synthetic, and treatment effect heterogeneity under staggered adoption goodman2021difference,sun2021estimating,CS:2021,athey2022design. See also CD:22:survey,NBERw29170,roth2023s,Sun:Shapiro:22 for introductory articles and recent surveys.

Second, we propose two alternative estimation methods. Without covariates, simple regression-based estimators suffice, and we clarify their connections to standard two-way fixed effects and event-study regressions. When covariates are present, adding them separately within the regression framework may seem straightforward, but we do not recommend this approach due to contamination bias in treatment effect estimation NBERw29709,NBERw30108. Instead, we develop a class of asymptotically equivalent, efficient estimators of FPR and BPR, including doubly and locally robust estimators, derived from the efficient influence functions. This contributes to the literature on efficient and doubly robust treatment effect estimation, where ATE and ATT have been the primary focus Hahn:98,HIR:03,sant2020doubly,locally:robust:2022.

Third, we develop two complementary inference strategies that extend beyond the direct regression-based approach. The first builds on our semiparametrically efficient estimators and delivers asymptotically exact inference when the full dataset is available. The second offers a practical alternative for settings with limited information: it constructs confidence intervals for FPR and BPR using the reported ATT confidence interval and a statistic that captures the share of treated individuals who do not take the action of interest. This approach allows researchers to perform persuasion-based back-of-the-envelope inference even in the absence of full data access.

We are not aware of prior studies analyzing persuasion rates within a DID framework. Related research includes yu2023binary, who study the identification of persuasion types in the same setting as jun2023identifying, and possebom2022probability, who examine the identification of average persuasion rates under sample selection. Two companion papers, jun2025tv and jun2025rdd, analyze persuasion effects in observational data on televised debates and in regression discontinuity designs, respectively.\footnote{These companion papers were developed after the first version of the present study was posted on arXiv at \url{https://arxiv.org/abs/2410.14871}, extending the analysis of persuasion effects to other research designs.} While those studies focus on alternative identification strategies, the present paper develops a unified DID framework that accommodates both two-period and staggered adoption settings and delivers new results for identification, estimation and inference.

The causal persuasion rate is conceptually related to probabilities of causation, which have been extensively studied outside economics pearl1999probabilities,Yamamoto:2012,Dawid2014fitting,Dawid2022effects,Ding:2024:arXiv:prob_necessity. However, FPR conceptually differs from previously defined probabilities of causation, and prior research has mainly focused on deriving population bounds without connecting to DID methods or addressing inference. This line of research has been largely overlooked in economics, with possebom2022probability as a recent exception. More broadly, our analysis relates to studies of treatment effects that depend on the joint distribution of potential outcomes, including early contributions by heckman1997making and more recent work by kaji2023assessing, who extend the persuasion framework to continuous outcomes, and by ji2023model, who develop covariate-assisted bounds on persuasion rates without monotonicity assumptions.\footnote{In both kaji2023assessing and ji2023model, persuasion rates serve as illustrative examples within a broader theoretical framework that extends beyond persuasion itself; however, neither paper considers the DID setting.}

The remainder of the paper is organized as follows. (ref) study the canonical two-period case. (ref) extends the analysis to staggered treatments and emphasizes event-study-style estimation. As empirical illustrations, we revisit the impact of news media on the 1997 British election Ladd:Lenz:09,Hainmueller:2012 in (ref) and re-evaluate curriculum reforms’ effects on political attitudes in China cantoni2017curriculum in (ref). The appendices, including the online materials, contain all proofs and additional results.

The Parameters

For each observational unit $i$, let $Y_{it}$ and $D_{it}$ be binary variables observed at time $t \in \{0,1\}$, where $t=0$ denotes the pre-treatment period. Potential outcomes are denoted by $Y_{it}(d)$, so that $ Y_{it} = D_{it} Y_{it}(1) + (1 - D_{it}) Y_{it}(0). $ We also observe a vector of exogenous covariates $X_i$, with support $\mathcal{X}$ assumed common across units. As an example, $Y_{i1}(1)$ indicates whether an agent votes for a certain party in period 1 after being exposed to a news platform, while $Y_{i1}(0)$ indicates voting behavior without exposure.

Our goal is to identify and estimate the causal persuasion rates on the treated---formally defined in the following subsections---within this panel setting. We begin by stating an assumption that formalizes the setup.

assumption[No Anticipation and Non-Degenerate Probabilities] At $t=0$, no individual is treated: $Y_{i0} = Y_{i0}(0)$ and $D_{i0} = 0$ with probability one. At $t=1$, there exists $\epsilon > 0$ such that for (almost) all $x \in \mathcal{X}$, $\mathbb{P}(D_{i1} = 1 \mid X_i = x) \leq 1 - \epsilon$, and \[ \epsilon \leq \min\bigl[ \mathbb{P}\{ Y_{i1}(0) = 0, D_{i1} = 1 \mid X_i = x \},\ \mathbb{P}\{ Y_{i1}(1) = 1, D_{i1} = 1 \mid X_i = x \} \bigr]. \]

(ref) ensures that $t=0$ is a pre-treatment period and that treatment begins at $t=1$. The condition $\mathbb{P}(D_{i1}=1 \mid X_i=x) \leq 1-\epsilon$ guarantees a non-trivial untreated population given $X_i=x$. The non-degeneracy requirement ensures that both persuasion measures are well-defined: $\mathbb{P}\{ Y_{i1}(0)=0,\, D_{i1}=1 \mid X_i=x \} \geq \epsilon$ implies a subgroup of treated individuals who would not have taken the action of interest without treatment (the population relevant for FPR), while $\mathbb{P}\{ Y_{i1}(1)=1,\, D_{i1}=1 \mid X_i=x \} \geq \epsilon$ ensures that persuasion occurs (the population relevant for BPR).

In the next subsections, we define individualized persuasion rates and discuss their aggregation.

The Conditional Persuasion Rate on the Treated

We define the forward and backward conditional persuasion rates on the treated (FCPR and BCPR) as

align*[align* omitted — 200 chars of source]

which are well-defined under (ref), with subscript $c$ indicating “conditional.”

Both measures apply to treated individuals with covariates $X_i=x$ but capture different aspects of persuasion. In the voting example, $\theta^{(F)}_c(x)$ measures the fraction of treated individuals persuaded to vote for the endorsed party among those who would not have voted for it otherwise. This is the on-the-treated version of the persuasion rate in dellavigna2007fox and jun2023identifying. A drawback is that it conditions on a counterfactual subgroup that is not directly observed.

The backward rate $\theta^{(B)}_c(x)$ addresses this by focusing on treated individuals who actually voted for the party and measuring how many would have behaved differently without the treatment. As discussed in the introduction, $\theta^{(B)}_c(x)$ coincides with the conditional probability of necessity (PN) of pearl1999probabilities, commonly used in legal and policy contexts to assess causation Dawid2014fitting. In contrast, $\theta^{(F)}_c(x)$ does not correspond to Pearl’s probability of sufficiency (PS), defined as \[ \textrm{PS}(x) := \mathbb{P}\{ Y_{i1}(1)=1 \mid Y_{i1}(0)=0,\, D_{i1}=0,\, X_i=x\}, \] which concerns persuasion among untreated individuals. Although $\textrm{PS}(x)$ is based on an identifiable group, it is not point-identified under the parallel trends assumption.

To further understand FCPR and BCPR, define \[ p_{st}(x) := \mathbb{P}\{ Y_{i1}(0)=s,\, Y_{i1}(1)=t \mid D_{i1}=1,\, X_i=x \}, \qquad (s,t) \in \{0,1\}^2. \] Then

align[align omitted — 160 chars of source]

Here, $p_{00}(x)$ represents “never-persuadable” (NP) individuals, $p_{11}(x)$ “already-persuaded” (AP), and $p_{01}(x)$ “treatment-persuadable” (TP), all within the treated group with covariates $X_i=x$. Thus, $\theta^{(F)}_c(x)$ is the share of TP among NP and TP, while $\theta^{(B)}_c(x)$ is the share of TP among AP and TP.

This classification parallels the latent type framework in the LATE literature. If we reinterpret the binary outcome as a “treatment’’ and the treatment indicator as an “instrument,’’ NP, AP, and TP correspond to never-takers, always-takers, and compliers, respectively. Under this analogy, persuasion measures the extent to which treatment causally induces compliance with the targeted behavior.

Since $Y_{i1}(0)$ and $Y_{i1}(1)$ are never observed simultaneously, $\theta^{(F)}_c(x)$ and $\theta^{(B)}_c(x)$ are not point-identified without additional assumptions. Consider instead

equation[equation omitted — 186 chars of source]

Unlike $\theta^{(F)}_c(x)$ and $\theta^{(B)}_c(x)$, these lower-bound parameters depend only on the marginal potential outcome distributions, conditional on $D_{i1}=1$ and $X_i=x$, because

align*[align* omitted — 177 chars of source]

and the denominators of $\theta^{(F)}_{cL}(x)$ and $\theta^{(B)}_{cL}(x)$ equal $\mathbb{P}\{ Y_{i1}(0)=0 \mid D_{i1}=1,\, X_i=x \}$ and $\mathbb{P}\{ Y_{i1}(1)=1 \mid D_{i1}=1,\, X_i=x \}$, respectively.

Let $r \in \{F, B\}$. Since $\theta^{(r)}_{cL}(x) \leq \theta^{(r)}_c(x)$ in general, $\theta^{(r)}_{cL}(x)$ can be viewed as a conservative measure of $\theta^{(r)}_c(x)$. However, it is typically less conservative than $\textrm{CATT}(x)$, which appears as the numerator in both $\theta^{(F)}_{cL}(x)$ and $\theta^{(B)}_{cL}(x)$. In fact, $\theta^{(r)}_{cL}(x)$ represents the Fréchet–Hoeffding lower bound on $\theta^{(r)}_c(x)$ based on the marginal distributions of the potential outcomes, and this bound is sharp. See Online Appendix (ref) for further discussion. If $p_{10}(x)=0$, then $\theta^{(r)}_{cL}(x)=\theta^{(r)}_c(x)$. We now state the condition under which this holds.

assumption[No Backlash] $\mathbb{P}\{ Y_{i1}(1) \geq Y_{i1}(0) \mid D_{i1}=1,\, X_i \}=1$ almost surely.

Under (ref), $p_{10}(x)=0$, and thus $p_{00}(x)+p_{01}(x)+p_{11}(x)=1$ for (almost) all $x \in \mathcal{X}$. This assumption rules out “contrarians’’ among the treated, i.e., individuals who would act opposite to the persuasive message’s intent (analogous to “defiers’’ in the LATE literature). Consequently, the treated population, conditional on $X$, consists of only three types: never-persuadable (NP), already-persuaded (AP), and treatment-persuadable (TP).

In the voting example, (ref) states that treated voters who would have voted for the party supported by the news media even without exposure do not change their decision when exposed. Equivalently, the persuasive message is directional or biased, with no backlash among the treated. A sufficient condition for (ref) is the monotone treatment response (MTR) assumption manski1997monotone, which requires $Y_{i1}(1)\geq Y_{i1}(0)$ almost surely.\footnote{MTR is stronger, as it also requires $Y_{i1}(1)\geq Y_{i1}(0)$ for untreated individuals, whereas (ref) imposes this only for treated ones. In the voting example, our assumption allows untreated individuals who might be negatively persuaded if treated, but rules out such backlash among those actually treated.} The MTR assumption for binary outcomes is also used in pearl1999probabilities and jun2023identifying.

lemmaSuppose (ref) hold. Then for $r \in \{F,B\}$, $\theta^{(r)}_c(x)=\theta^{(r)}_{cL}(x)$ for all $x \in \mathcal{X}$. If (ref) is violated, we still have $\theta^{(r)}_c(x)\geq\theta^{(r)}_{cL}(x)$ for all $x \in \mathcal{X}$.

(ref) does not establish identification, but it shows that excluding backlash allows us to express $\theta^{(r)}_c(x)$ ($r \in \{F,B\}$) using marginal probabilities of potential outcomes. Even with possible backlash, $\theta^{(r)}_{cL}(x)$ remains a conservative lower bound for $\theta^{(r)}_c(x)$, which motivates our notation with subscript $L$. We further discuss bounds, their sharpness, and backlash effects in Online Appendix (ref).

Under (ref), point identification of FCPR and BCPR reduces to identifying CATT. Since (ref) implies $\textrm{CATT}(x)\geq 0$, both persuasion measures are no smaller than CATT: for instance,

align[align omitted — 253 chars of source]

The first inequality is strict when $\textrm{CATT}(x)>0$ and $\mathbb{P}\{ Y_{i1}(0)=1 \mid D_{i1}=1,\, X_i=x\}>0$, i.e., when some treated individuals would vote for the endorsed party without exposure. An analogous ordering holds for $\theta^{(B)}_{cL}(x)$.

Since $Y_{i1}(1)=Y_{i1}$ for $D_{i1}=1$ and $Y_{i0}(0)=Y_{i0}$, both $\theta^{(F)}_{cL}(x)$ and $\theta^{(B)}_{cL}(x)$ involve only one unidentified term:

align[align omitted — 93 chars of source]

Indeed, we have that

align*[align* omitted — 362 chars of source]

where $\mathbb{P}\{ Y_{i1}=1 \mid D_{i1}=1,\, X_i=x \}$ is directly identified from data. Before turning to the identification of $\tau_c(x)$, we next discuss aggregation.

Aggregation

The aggregated persuasion rates on the treated, called the forward and backward average persuasion rates on the treated (FPR and BPR), are defined as

equation[equation omitted — 339 chars of source]

The forward rate $\theta^{(F)}$ relates to the average persuasion rate (APR) in jun2023identifying, $\mathbb{P}\{ Y_{i1}(1)=1 \mid Y_{i1}(0)=0\}$, analogous to the ATT-ATE relationship.

Aggregating the numerators and denominators of (ref) separately, we define

equation[equation omitted — 475 chars of source]

where $\textrm{ATT} := \mathbb{E}\{ \textrm{CATT}(X_i) \mid D_{i1}=1\}$. If (ref) holds for all $x$, then $\theta^{(r)} = \theta^{(r)}_L$ for both $r \in \{F,B\}$ by Bayes’ rule.

lemmaSuppose (ref) hold. Then for $r \in \{F,B\}$, $\theta^{(r)} = \theta^{(r)}_L$. If (ref) is violated, $\theta^{(r)} \geq \theta^{(r)}_L$ in general.

Thus, averaging the numerators and denominators of $\theta^{(F)}_{cL}(\cdot)$ and $\theta^{(B)}_{cL}(\cdot)$ separately is the correct aggregation method, and identification of $\tau_c(\cdot)$ suffices to identify $\theta^{(F)}_L$ and $\theta^{(B)}_L$.

As with FCPR and BCPR, subscript $L$ indicates that $\theta^{(F)}_L$ and $\theta^{(B)}_L$ provide valid lower bounds for $\theta^{(F)}$ and $\theta^{(B)}$, respectively, without (ref). However, if (ref) fails, $\theta^{(F)}_{cL}(x)$ and $\theta^{(B)}_{cL}(x)$ can be negative for some $x \in \mathcal{X}$, whereas $\theta^{(F)}_c(x)$ and $\theta^{(B)}_c(x)$ are always nonnegative. Consequently, the aggregated parameters $\theta^{(F)}_L$ and $\theta^{(B)}_L$, derived from $\theta^{(F)}_{cL}(x)$ and $\theta^{(B)}_{cL}(x)$, may not provide the sharpest bounds based on the marginals of the potential outcomes. We discuss sharpness and interpretational issues in the absence of (ref) in Online Appendix (ref).

For intuition, assume $Y_{i1}(1) \geq Y_{i1}(0)$ almost surely so that (ref) holds, and focus on FPR. In this case, FPR is obtained by rescaling ATT, with the factor capturing the fraction of treated individuals who are not pre-converted and thus are potential persuasion targets.\footnote{In terms of the earlier decomposition of the treatment group, FPR equals $\#\mathrm{TP} / (\#\mathrm{TP} + \#\mathrm{NP}) = \#\mathrm{TP} / \#\mathrm{AP}^c$, where $\#A$ denotes the size of group $A$.} The larger the share of pre-converted individuals in the treatment group, the more important it is to adjust for “preaching to the converted’’ when measuring pure persuasive effects. If no one is pre-converted and everyone is a persuasion target, ATT and FPR coincide; otherwise, FPR typically exceeds ATT.

Identification

Identification via Parallel Trends

We focus on the two-period case and use a DID approach to identify $\tau_c(\cdot)$. Extensions to staggered treatment are discussed in (ref). The key assumption is parallel trends:

assumption[Parallel Trends] Almost surely, \begin{align*} &\mathbb{P}\{ Y_{i1}(0)=1 \mid D_{i1}=1,\, X_i \} - \mathbb{P}\{ Y_{i0}(0)=1 \mid D_{i1}=1,\, X_i \} \\ &\quad=\; \mathbb{P}\{ Y_{i1}(0)=1 \mid D_{i1}=0,\, X_i \} - \mathbb{P}\{ Y_{i0}(0)=1 \mid D_{i1}=0,\, X_i \}. \end{align*}

By Proposition 3.2 and Example 1 of roth2023parallel, (ref) is equivalent to separability: there exist functions $G$ and $H$ such that \[ \mathbb{P}\{ Y_{it}(0)=1 \mid D_{i1}=d,\, X_i=x \} = G(t,x)+H(d,x). \] This condition excludes common parametric generalized linear models such as \[ \mathbb{P}\{ Y_{it}(0)=1 \mid D_{i1}=d,\, X_i=x \} = \Lambda^{-1}\bigl( \beta_0+\beta_1 t+\beta_2 d+\beta_3^\top x \bigr), \quad \Lambda^{-1}(s)=\frac{\exp(s)}{1+\exp(s)}. \] However, it can be generalized to allow a pre-specified nonlinear link $\Lambda:[0,1]\to\mathbb{R}$, including the logistic link as a special case:

multline*[multline* omitted — 268 chars of source]

We discuss this generalization in Online Appendix (ref).

Define \[ \Psi(x) := \Pi_0(1,x)+\Pi_1(0,x)-\Pi_0(0,x), \quad \Pi_t(d,x):=\mathbb{P}(Y_{it}=1 \mid D_{i1}=d,\, X_i=x), \] for $t \in \{0,1\}$ and $(d,x^\top)^\top \in \{0,1\} \times \mathcal{X}$, all directly identified from the joint distribution of $(Y_{i0},Y_{i1},D_{i1},X_i^\top)^\top$.

theoremSuppose (ref) hold. Then, for all $x \in \mathcal{X}$, $\tau_c(x)$ in (ref) is point-identified by $\Psi(x)$.

Thus, under (ref), both $\theta^{(r)}_{cL}$ and $\theta^{(r)}_L$ for $r \in \{F,B\}$ are point-identified. For instance, for all $x \in \mathcal{X}$, \[ \theta^{(F)}_{cL}(x) = \frac{\Pi_1(1,x) - \Psi(x)}{1-\Psi(x)} = \frac{\textrm{CATT}(x)}{\textrm{CATT}(x)+1-\Pi_1(1,x)}, \] where $\textrm{CATT}(x) = \Pi_1(1,x) - \Psi(x)$ is the usual DID estimand, equivalently $\textrm{CATT}(x)=\Delta(1,x)-\Delta(0,x)$ with $\Delta(d,x):=\Pi_1(d,x)-\Pi_0(d,x)$.

Aggregated parameters are particularly relevant when covariates are multi-dimensional. Define \[ \bar\theta^{(F)}_L := \frac{\mathbb{E}\{ \Pi_1(1,X_i) - \Psi(X_i) \mid D_{i1}=1\}}{\mathbb{E}\{ 1-\Psi(X_i) \mid D_{i1}=1\}}, \quad \bar\theta^{(B)}_L := \frac{\mathbb{E}\{ \Pi_1(1,X_i) - \Psi(X_i) \mid D_{i1}=1\}}{\mathbb{E}\{ \Pi_1(1,X_i) \mid D_{i1}=1\}}. \]

corollarySuppose (ref) hold. Then for $r \in \{F,B\}$, $\theta^{(r)} = \theta^{(r)}_L = \bar\theta^{(r)}_L$. If (ref) is violated, $\theta^{(r)} \geq \theta^{(r)}_L = \bar\theta^{(r)}_L$.

Hence, when there is no backlash ((ref)), FPR and BPR are point-identified by rescaling the DID estimand. Even with potential backlash, rescaling the DID parameter yields conservative measures of FPR and BPR.

Recall that under (ref), the treated population (with or without conditioning on $X_i$) can be decomposed into three persuasion types. If (ref) also holds, then $\textrm{ATT}$ is identified, and so are the shares of each type: the share of NP among the treated is $\mathbb{E}\{ p_{00}(X) \mid D_{i1}=1\} = \mathbb{P}(Y_{i1}=0 \mid D_{i1}=1)$, the share of AP is $\mathbb{E}\{ p_{11}(X) \mid D_{i1}=1\} = \mathbb{P}(Y_{i1}=1 \mid D_{i1}=1) - \textrm{ATT}$, and the share of TP is simply $\textrm{ATT}$.

For estimation, $\Pi_t(d,x)$ can be replaced by parametric, semiparametric, or nonparametric estimates. However, for FPR and BPR, direct plug-in and aggregation are not the only approach. We discuss alternative estimation strategies in (ref).

Controlling for the Pre-Treatment Outcome and Unconfoundedness

As an alternative to parallel trends, identification can be based on unconfoundedness after conditioning on all relevant covariates. But the two approaches are related.

Specifically, consider the unconfoundedness assumption given $Z_i:= [Y_{i0}, X_i]$: i.e., $Y_{i1}(0)$ and $D_{i1}$ are independent given $Z_i$. Then, using $Z_i$ in place of $X_i$ when constructing $\bar\theta^{(F)}_L$ and $\bar\theta^{(B)}_L$ yields estimands that identify $\theta^{(F)}_L$ and $\theta^{(B)}_L$ under the unconfoundedness assumption given $Z_i$. Hence, including $Y_{i0}$ and applying the DID formulas effectively implements identification via unconfoundedness.

There is debate over conditioning on $Y_{i0}$ when estimating ATT with DID roth2023s. However, there is a testable condition under which this choice becomes irrelevant: if $Y_{i0}$ is independent of $D_{i1}$ given $X_i$, then the parallel trends and unconfoundedness approaches yield identical estimands for $\theta^{(F)}_L$ and $\theta^{(B)}_L$. In this case, it does not matter which assumption is adopted. We provide further discussion in Online Appendix (ref).

Estimation

Regression-Based Approaches

ATT is often estimated using a two-way fixed effects regression model. A similar approach applies to persuasion rates. To illustrate, we first consider the case without covariates. If (ref) hold with $X$ omitted, FPR and BPR are

align[align omitted — 510 chars of source]

where $\Delta(d):=\Pi_1(d)-\Pi_0(d)$ and $\Pi_t(d):=\mathbb{P}(Y_{it}=1\mid D_{i1}=d)$ for $(t,d)\in\{0,1\}^2$.

We treat (ref) and (ref) as the estimands of interest, assuming $\mathbb{E}(Y_{i1}\mid D_{i1}=1)>0$ and $\Psi:=\mathbb{E}(Y_{i0}\mid D_{i1}=1)+\mathbb{E}(Y_{i1}-Y_{i0}\mid D_{i1}=0)<1$. These quantities can be obtained from a simple linear regression, as shown below.

Consider the two-way fixed effects regression:

equation[equation omitted — 137 chars of source]

where $G_i=1$ for treated individuals and $0$ otherwise; $\gamma_1$ and $\gamma_2$ capture group and time fixed effects.

assumption[Two-Way Fixed Effects] \begin{enumerate} • Treatment occurs only for treated individuals in period 1: $D_{it}= \mathbb{1}(t=1) G_i$. • Time and group assignments are exogenous: $\mathbb{E}(\epsilon_{it}\mid G_i,t)=0$ for $t\in\{0,1\}$. \end{enumerate}

Under (ref), OLS consistently estimates $\gamma_0,\gamma_1,\gamma_2$, and $\gamma$. The interaction coefficient $\gamma$ corresponds to the DID estimand, the numerator of both $\bar\theta^{(F)}_L$ and $\bar\theta^{(B)}_L$ in (ref)-(ref). The following result shows that the persuasion rates can be directly obtained from this regression.

theoremSuppose $\mathbb{E}(Y_{i1}\mid D_{i1})>0$ and $\Psi<1$, ensuring $\bar\theta^{(F)}_L$ and $\bar\theta^{(B)}_L$ are well-defined. If (ref) holds, then \begin{equation} \bar\theta^{(F)}_L = \frac{\gamma}{1-\gamma_0-\gamma_1-\gamma_2}, \qquad \bar\theta^{(B)}_L = \frac{\gamma}{\gamma_0+\gamma_1+\gamma_2+\gamma}. \end{equation}

The denominators in (ref) adjust the DID estimand to obtain FPR and BPR. Inference can be conducted using standard OLS theory in combination with the delta method. One concern with this standard inference procedure is that the delta method may perform poorly in finite samples if the denominators are too close to zero. This issue can be addressed by imposing the null hypothesis for inference. For example, consider testing $H_0: \bar\theta^{(F)}_L = \theta_{F,H}$, which can be equivalently expressed as

equation[equation omitted — 97 chars of source]

Testing (ref) against its negation does not suffer from the “small denominator” problem, and inverting the test yields a robust confidence region for $\bar\theta^{(F)}_L$.

To provide additional insight, we present an equivalent approach based on instrumental variables. Define an auxiliary outcome for period 1: \[ \tilde Y_{i1} := D_{i1} + Y_{i1}(1-D_{i1}), \] which equals $Y_{i1}$ if $D_{i1}=0$, and equals $1$ if $D_{i1}=1$. Intuitively, $\tilde Y_{i1}$ represents “pseudo voters’’ who always vote for the endorsed party when treated but behave as untreated voters otherwise.

For either $A_i := \tilde Y_{i1} - Y_{i0}$ or $A_i := Y_{i1} D_{i1}$, consider the moment conditions: \[ \mathbb{E}\bigl[(Y_{i1}-Y_{i0}) - \beta_0 - \beta_1 A_i \bigr]=0, \qquad \mathbb{E}\bigl[D_{i1}\{(Y_{i1}-Y_{i0}) - \beta_0 - \beta_1 A_i\}\bigr]=0. \] Here, $\beta_1$ is the two-stage least squares (2SLS) coefficient of $Y_{i1}-Y_{i0}$ on $A_i$, using $D_{i1}$ as an instrument.\footnote{The intercept is instrumented by itself, yielding two moment conditions.}

theoremSuppose $\mathbb{E}(Y_{i1}\mid D_{i1}=1)>0$ and $\Psi<1$, so $\bar\theta^{(F)}_L$ and $\bar\theta^{(B)}_L$ are well-defined. If $A_i=\tilde Y_{i1}-Y_{i0}$, then $\beta_1=\bar\theta^{(F)}_L$; if $A_i=Y_{i1}D_{i1}$, then $\beta_1=\bar\theta^{(B)}_L$.

Thus, both persuasion rates can be estimated by 2SLS. For FPR, both the numerator and denominator of (ref) are DID estimands: the numerator is estimated by regressing $Y_{i1}-Y_{i0}$ on $D_{i1}$, and the denominator similarly uses $\tilde Y_{i1}-Y_{i0}$. Hence,

equation[equation omitted — 269 chars of source]

For BPR, the numerator matches FPR’s, while the denominator is $\mathbb{E}(Y_{i1}\mid D_{i1}=1)=\mathrm{Cov}(Y_{i1}D_{i1},D_{i1})/\mathbb{V}(D_{i1})$, using $D_{i1}^2=D_{i1}$ since $D_{i1}$ is binary.

It can be seen from (ref) that the two-way fixed effects and GMM estimators are algebraically equivalent. Inference can be performed within a generalized method of moments (GMM) framework. This formulation shows that the “small denominator” problem in the two-way fixed effect approach is exactly a weak-instrument problem in GMM. A robust inference procedure can be developed by using an Anderson-Rubin statistic that directly tests the moment conditions under the null. See Online Appendix (ref) for more details.

When covariates are present, one simple approach is to include them separately, along with their interactions and powers, directly in the regression. However, the algebraic equivalence between the fixed effects and GMM estimators no longer holds in this setting. Moreover, this approach can introduce contamination bias, meaning the estimator may not correspond to a well-defined convex average of heterogeneous treatment effects NBERw29709,NBERw30108. A more appropriate strategy is to estimate (ref) locally around $X_i=x$ and then average over $X_i$ to recover FPR or BPR. This naturally leads to a semiparametric approach, where first-step estimators are computed conditional on covariates. We discuss semiparametric estimation methods in the next subsection.

Semiparametric Approaches

We now incorporate covariates explicitly and discuss semiparametric estimation strategies. For clarity, we focus on FPR, as its formulation is slightly more involved than that of BPR; we summarize the latter at the end of the subsection. Several semiparametric estimators of $\bar\theta^{(F)}_L$ can be constructed using first-step estimates of the conditional probabilities $\Pi_t(d,x)$ or the propensity score $P(x):=\mathbb{P}(D_{i1}=1\mid X_i=x)$.\footnote{First-step estimators are typically nonparametric to reduce misspecification, though parametric choices are possible. For simplicity, we refer to all such methods as semiparametric, acknowledging that this may be a slight misnomer.}

Using the definition of $\Psi(X_i)$, we can rewrite $\bar\theta^{(F)}_L$ as

equation[equation omitted — 241 chars of source]

where $\Delta(d,x):=\Pi_1(d,x)-\Pi_0(d,x)$. Given a random sample $\{(Y_{i0},Y_{i1},D_{i1},X_i): i=1,\dots,n\}$, this suggests the following DID-based estimator: \[ \hat\theta^{(F)}_{L,DID} := \frac{\sum_{i=1}^n \{\widehat{\Delta}(1,X_i)-\widehat{\Delta}(0,X_i)\}D_{i1}}{\sum_{i=1}^n \{\widehat{\Delta}(1,X_i)-\widehat{\Delta}(0,X_i)\}D_{i1}+\sum_{i=1}^n\{1-\widehat{\Pi}_1(1,X_i)\}D_{i1}}, \] where $\widehat{\Delta}(d,X_i):=\widehat\Pi_1(d,X_i)-\widehat\Pi_0(d,X_i)$ and $\widehat\Pi_t(d,x)$ is a first-step estimator of $\Pi_t(d,x)$ chosen by the researcher.

Alternative estimators can also be derived. For instance, $\bar\theta^{(F)}_L$ can be expressed as

equation[equation omitted — 201 chars of source]

leading to the plug-in estimator

align[align omitted — 210 chars of source]

which coincides numerically with $\hat\theta^{(F)}_{L,DID}$ when no covariates are present.\footnote{In the absence of covariates, the IV estimator from (ref) is also identical.}

Applying iterated expectations and Bayes’ rule to (ref), we obtain

equation[equation omitted — 125 chars of source]

where

align[align omitted — 144 chars of source]

This leads to another estimator,

align[align omitted — 143 chars of source]

where

align*[align* omitted — 167 chars of source]

and $\widehat{P}(X_i)$ is a first-step estimator of the propensity score $P(X_i)$. This propensity-odds-weighted (POW) estimator requires estimating only $P(X_i)$ in the first step, whereas $\hat\theta^{(F)}_{L,PI}$ requires $\Delta(0,X_i)=\Pi_1(0,X_i)-\Pi_0(0,X_i)$. We revisit this difference in the next section, where we introduce a doubly (and locally) robust estimator.

Analogously, the following estimators apply to BPR:

align*[align* omitted — 397 chars of source]

These estimators are based on equivalent moment conditions and thus share the same efficient influence function for each parameter. Consequently, they are asymptotically equivalent under standard regularity conditions.

In the next section, we derive the efficient influence functions explicitly. This serves two purposes: (i) to obtain the asymptotic variance of any regular, asymptotically linear semiparametric estimator of $\bar\theta^{(F)}_L$ and $\bar\theta^{(B)}_L$, and (ii) to guide the construction of locally and doubly robust estimators.

The Efficient Influence Function

We first derive the efficient influence function for FPR $\bar\theta^{(F)}_L$, as it is more involved than that for BPR, which we discuss briefly at the end of this section. Under suitable regularity conditions, all semiparametric estimators described earlier are asymptotically linear, regular, and normal. Since semiparametric theory is well established Ackerberg2014, we do not detail all conditions for asymptotic normality. Instead, following newey1994asymptotic, we compute the semiparametric efficient influence function, whose variance equals the asymptotic variance of any regular, asymptotically linear estimator of $\bar\theta^{(F)}_L$. We assume random sampling of $(Y_{i1},Y_{i0},D_{i1},X_i^\top)^\top$.

assumptionThere exists $\epsilon>0$ such that for all $d,y_0,y_1\in\{0,1\}$ and $x\in\mathcal{X}$, \[ \epsilon \leq \mathbb{P}(Y_{i0}=y_0,\,Y_{i1}=y_1,\,D_{i1}=d \mid X_i=x) \leq 1-\epsilon. \]

This assumption ensures a well-behaved likelihood and scores. Below, we present the efficient influence function in two equivalent forms, depending on whether $P(X_i)$ or $\Pi_t(d,X_i)$ is estimated in the first step; this dual representation will motivate a doubly robust estimator.

Let $\mathcal{D}_i:=(Y_{i0},Y_{i1},D_{i1},X_i)$ and define \[ \theta^{(F)}_{L,den}:=\mathbb{E}[\{1-\Psi(X_i)\}D_{i1}], \] which appears in the denominator of the efficient influence function. For $EST \in \{\text{POW},\text{PI}\}$, define

align*[align* omitted — 270 chars of source]

where

align*[align* omitted — 596 chars of source]

We are now ready to state the main theorem of this section.

theoremSuppose (ref) holds. The semiparametric efficient influence function for $\bar\theta^{(F)}_L$ under random sampling of $(Y_1,Y_0,D_{i1},X^\top)^\top$ is \begin{align} F_{DID}(\mathcal{D}_i) := F_{POW,main}(\mathcal{D}_i)+F_{POW,adj}(\mathcal{D}_i), \end{align} which is equivalently written as \begin{align} F_{DID}(\mathcal{D}_i) = F_{PI,main}(\mathcal{D}_i)+F_{PI,adj}(\mathcal{D}_i). \end{align} For any regular, asymptotically linear estimator $\hat\theta^{(F)}_L$ based on a random sample of size $n$, \begin{align} \sqrt{n}(\hat\theta^{(F)}_L-\bar\theta^{(F)}_L) \;\overset{d}{\longrightarrow}\; N\bigl(0,\ \mathbb{E}[F_{DID}^2(\mathcal{D}_i)]\bigr). \end{align}

(ref) gives the efficient influence function in two algebraically equivalent forms. The POW representation corresponds to the moment condition \[ \mathbb{E}\bigl[H_{POW,num}(\mathcal{D}_i)-\bar\theta^{(F)}_L H_{POW,den}(\mathcal{D}_i)\bigr]=0, \] with $F_{POW,adj}(\mathcal{D}_i)$ adjusting for first-step estimation of $P(X_i)$. Similarly, the PI representation relates to $\hat\theta^{(F)}_{L,PI}$, with $F_{PI,adj}$ accounting for estimation of $\Delta(0,X_i)$.

By Theorem 2.1 of newey1994asymptotic, all regular, asymptotically linear estimators of $\bar\theta^{(F)}_L$ share this efficient influence function. Hence, $\hat\theta^{(F)}_{L,DID}$, $\hat\theta^{(F)}_{L,PI}$, and $\hat\theta^{(F)}_{L,POW}$ are asymptotically equivalent.

The efficient influence function also yields a doubly robust estimator. Using the identity

equation[equation omitted — 165 chars of source]

the estimator

equation[equation omitted — 284 chars of source]

with \[ \widehat{H}_{PI,adj}(\mathcal{D}_i) := -\frac{\widehat{P}(X_i)}{1-\widehat{P}(X_i)}(1-D_{i1})\bigl[(Y_{i1}-Y_{i0})-\widehat{\Delta}(0,X_i)\bigr], \] uses the full efficient influence function.

By construction, $\hat\theta^{(F)}_{L,DR}$ is locally robust: its asymptotic variance matches the oracle estimator that knows $P(X_i)$ and $\Delta(0,X_i)$. It is also doubly robust, remaining consistent if either $\widehat{\Delta}(0,\cdot)$ or $\widehat{P}(\cdot)$ is correctly specified. This structure naturally extends to a double/debiased machine learning (DML) estimator using machine learning first-step estimators and cross-fitting DML:EJ.

We conclude this section with the case of BPR. The only difference between FPR and BPR lies in their denominators. Specifically, BPR uses \[ \theta^{(B)}_{L,den}:=\mathbb{E}(Y_{i1}D_{i1}), \] and, under the same conditions as (ref), its efficient influence function is \[ F_{DID}^{(B)}(\mathcal{D}_i) := \frac{1}{\theta^{(B)}_{L,den}}\Bigl\{H_{EST,num}(\mathcal{D}_i)-\theta^{(B)}_L Y_{i1}D_{i1}+H_{EST,adj}(\mathcal{D}_i)\Bigr\}, \] where $EST\in\{\text{POW},\text{PI}\}$. By the same argument as for FPR, a (locally and) doubly robust estimator of $\theta^{(B)}_L$ is \[ \hat\theta^{(B)}_{L,DR} := \frac{\sum_{i=1}^n\bigl[(Y_{i1}-Y_{i0})D_{i1}-\widehat{\Delta}(0,X_i)D_{i1}+\widehat{H}_{PI,adj}(\mathcal{D}_i)\bigr]}{\sum_{i=1}^n Y_{i1}D_{i1}}. \]

Discussion

We discuss how to conduct back-of-the-envelope inference on FPR and BPR when the full dataset is unavailable but ATT is known. We assume $\textrm{ATT} \geq 0$ (as implied by (ref)). From (ref), FPR relates to ATT as \[ \textrm{FPR}(q)=\frac{\textrm{ATT}}{\textrm{ATT}+q}, \] where $q:=\mathbb{P}(Y_{i1}=0 \mid D_{i1}=1)$ such that $\textrm{ATT}+q=\mathbb{P}\{Y_{i1}(0)=0 \mid D_{i1}=1\}$. Thus, FPR exceeds ATT unless $\textrm{ATT}=0$ or $\mathbb{P}\{Y_{i1}(0)=0 \mid D_{i1}=1\}=1$, i.e., every treated individual is a persuasion target. In most settings, FPR will therefore be strictly larger than ATT.

Suppose that an interval $[\underline{q},\overline{q}] \subseteq [0,1]$ contains the true $q$ with probability approaching $1-\alpha_0$, where $0\leq\alpha_0<1$. In the empirical application in (ref), we face a situation in which the full data are unavailable, but a $(1-\alpha_0)$ confidence interval for $q$ can be used as $[\underline{q},\overline{q}]$. Since $q \mapsto \textrm{FPR}(q)$ is nonincreasing for $\textrm{ATT}\geq 0$, the implied bounds on FPR given $q \in [\underline{q},\overline{q}]$ are

align[align omitted — 192 chars of source]

These bounds provide the basis for inference on FPR using ATT estimates.

To make this concrete, suppose we have an estimate $\widehat{\textrm{ATT}}$ of $\textrm{ATT}$ with standard error $se(\widehat{\textrm{ATT}})$. Then, by the delta method, the pointwise standard error of $\widehat{\textrm{FPR}}(q):= \widehat{\textrm{ATT}}/(\widehat{\textrm{ATT}}+q)$ is $se(\widehat{\textrm{ATT}}) {q}/{(\widehat{\textrm{ATT}} + q)^2}$. Therefore, letting $z_\tau$ be the $\tau$-quantile of the standard normal distribution, an asymptotic $(1-\alpha)$ confidence interval for FPR can be obtained by $[\underline{\textrm{FPR}},\ \overline{\textrm{FPR}} ]$, where

align*[align* omitted — 430 chars of source]

because

align*[align* omitted — 635 chars of source]

For example, we may set $\alpha_0=\alpha/2$, which gives $(\alpha-\alpha_0)/2=\alpha/4$. Thus, when only ATT and its standard error are available, one can infer FPR as long as probabilistic bounds on $\mathbb{P}(Y_{i1}=0\mid D_{i1}=1)$ are known.\footnote{Unlike the inference problems in Imbens/Manski:04 and Stoye:07, our bound estimates depend on ATT and $(\underline{q},\overline{q})$ but are not jointly observed, preventing standard error calculations for the bound endpoints.}

The case of BPR is similar. Using the fact that $\mathrm{BPR}=\textrm{ATT}/(1-q)$, a $(1-\alpha)$ confidence interval is \[ \Bigg[ \widehat{\mathrm{BPR}}(\underline{q})- z_{1-(\alpha-\alpha_0)/2}\frac{se(\widehat{\textrm{ATT}})}{1-\underline{q}},\;\; \widehat{\mathrm{BPR}}(\overline{q})+ z_{1-(\alpha-\alpha_0)/2}\frac{se(\widehat{\textrm{ATT}})}{1-\overline{q}} \Bigg]. \] Finally, we remark that taking the Cartesian product of the $1-\alpha_F$ and $1-\alpha_B$ confidence intervals of FPR and BPR yields a $1-\alpha_F-\alpha_B$ Bonferroni joint confidence region of FPR and BPR.

Example I: News Media Persuasion

Setting

Ladd:Lenz:09 studied abrupt shifts in British newspaper endorsements from the Conservative to the Labour Party before the 1997 general election. Using data from the British Election Panel Study, they compared readers of newspapers that switched endorsements (treated group) with readers of newspapers that did not (control group). The binary outcome is whether a respondent voted Labour in 1997, and the binary treatment is whether a respondent read a switching newspaper. This is a two-period DID setting with no treatment in the pre-period. In addition to prior Labour voting in 1992, the dataset includes numerous predetermined covariates ($X_i$), all measured in 1992.

The data are publicly available via the Political Analysis Dataverse (\url{http://dvn.iq.harvard.edu/dvn/dv/pan}), as part of the replication materials for Hainmueller:2012. Following their approach, we use 36 predetermined covariates, excluding three prior voting variables (Labour, Conservative, and Liberal votes). The covariates cover party identification and support, ideology, parents’ voting behavior, political knowledge, television viewership, newspaper readership, authoritarianism, trade union membership, mortgage status, education, income, age, race, socioeconomic status, gender, region, and occupation.

Discussion of Assumptions

Applying our DID framework requires two key conditions: (i) no anticipation in (ref) and (ii) parallel trends, i.e., (ref). The first condition is plausible: Ladd:Lenz:09 document that newspaper endorsements of Labour were unexpected and not anticipated by readers. The parallel trends assumption implies that, in the absence of the endorsement shifts, changes in Labour voting between 1992 and 1997 would have been similar for readers of switching and nonswitching newspapers. Conditioning on covariates helps bolster the plausibility of this assumption. Indeed, imposing parallel trends directly on conditional probabilities, as in (ref), remains consistent with Figure 1 in Ladd:Lenz:09, which depicts a standard DID plot based on Labour vote shares.

We also revisit (ref). Consider a hypothetical voter who leans Labour but reads a Conservative newspaper to engage with opposing views. If the newspaper unexpectedly endorses Labour, she might infer that Tony Blair secretly supports free-market policies and consequently decide not to vote Labour, thereby violating (ref). Although such cases cannot be entirely ruled out, they are likely rare. Most voters seek confirmatory media; the endorsement shifts were perceived as credible signals of political realignment, and the reasoning required for such backlash is atypical in high-salience elections. Nevertheless, if backlash is present, our estimates of persuasion effects should be interpreted as lower bounds.

Estimation Results

(ref) reports estimates of the average persuasion rates on the treated. Panel A shows the forward measure (FPR) and Panel B the backward one (BPR). Columns FE and GMM present two-way fixed effect and GMM estimators without covariates. Although algebraically equivalent, their $t$-statistics differ slightly due to different asymptotic approximations. Columns FE-X and GMM-X include covariates additively in the regression. The two-step estimators (DID, PI, POW, DR) use first-step logistic regressions for up to five conditional probabilities: $\mathbb{P}(Y_{it}=1 \mid D_{i1}=d, X_i=x)$ for $t=0,1$ and $d=0,1$, and $\mathbb{P}(D_{i1}=1 \mid X_i=x)$.

table[table omitted — 1,145 chars of source]

To compare with the existing literature, our doubly robust ATT estimate is $0.089$, very close to the unconditional DID estimate of $0.086$ reported by Ladd:Lenz:09. Previous ATT estimates obtained using matching Ladd:Lenz:09 or entropy balancing Hainmueller:2012 are somewhat larger, ranging from $0.096$ to $0.140$. Our FPR and BPR estimates exceed all these ATT values, reflecting the substantial presence of never-persuadable (NP) and already-persuaded (AP) individuals in the treated population. Using the decomposition in (ref) and the discussion below (ref), the treated group can be partitioned into treatment-persuadable (TP) individuals with a share of $0.089$ (equal to ATT), NP with a share of $0.417$, and AP with a share of $0.494$. The predominance of NP and AP explains why FPR and BPR are considerably larger than ATT.

We also illustrate back-of-the-envelope inference. Because the full replication files of Ladd:Lenz:09 are unavailable, we rely on replication materials from Hainmueller:2012. From Table 2 of Ladd:Lenz:09, Column 3 (exact matching DID) provides an ATT estimate of $0.109$ with a standard error of $0.041$. The estimated share of non-takers in the treated group, $q = \mathbb{P}(Y_1 = 0 \mid D_1 = 1)$, is $0.583$ (with a treated group size of $211$), and its 97.5% confidence interval is $[0.507, 0.659]$. Using these values, the 95% confidence interval for FPR is $[0.039, 0.300]$ with a point estimate of $0.158$. For BPR, the corresponding interval is $[0.035, 0.589]$ with a point estimate of $0.261$.

Staggered Treatment

We extend our framework to multiple time periods with staggered treatment adoption, where units begin treatment at different times and remain treated thereafter.

Setup and Parameters of Interest

Let $t \in \{0, 1, 2, \ldots, T\}$ denote time, and $D_{it}$ the treatment status of unit $i$ at time $t$. Initially ($t = 0$), no unit is treated. If treatment starts for unit $i$ at time $s \in \{1, 2, \ldots, T\}$, then $D_{it} = 0$ for $t < s$ and $D_{it} = 1$ for $t \geq s$. Never-treated units satisfy $D_{it} = 0$ for all $t$. Exogenous covariates are $X_i$.

Define $\mathcal{S}_i := \bigl\{t\in\{1,\cdots,T\}:\ D_{it} = 1 \bigr\}$, and $S_i$ by $S_i := \min \mathcal{S}_i$ if $\mathcal{S}_i$ is non-empty, and $S_i:=\infty$ otherwise. Thus, $S_i \in \{1, 2, \ldots, T, \infty\}$ denotes the period in which treatment begins, with $S_i = \infty$ indicating that the unit is never treated. To allow for heterogeneous treatment effects, define $Y_{it}(s)$ as the potential outcome at time $t$ under treatment starting in period $s$. We impose:

assumption[Staggered Adoption with No Anticipation] At $t = 0$, no unit receives or anticipates treatment, so $Y_{i0} := Y_{i0}(\infty)$. For $t = 1,\dots, T$, the observed outcome is $Y_{it} := Y_{it}(S_i)$, and for all $s \in \{1,\dots,T\}$ and $t < s$, $Y_{it}(s) = Y_{it}(\infty)$ almost surely. Moreover, for all $s, t \in \{1,\dots,T\}$ and almost all $x \in \mathcal{X}$, there exists $\epsilon > 0$ such that $\mathbb{P}(S_i = s \mid X_i = x) \leq 1 - \epsilon$, and \[ \epsilon \leq \min \Bigl\{ \mathbb{P}\bigl( Y_{it}(\infty) = 0,\ S_i = s \mid X_i = x \bigr),\ \mathbb{P}\bigl( Y_{it} = 1,\ S_i = s \mid X_i = x \bigr) \Bigr\}. \]

The first part ensures no anticipation: untreated behavior persists until treatment begins. The overlap condition $\mathbb{P}(S_i = s \mid X_i = x) \leq 1 - \epsilon$ prevents deterministic timing given covariates. The last condition guarantees well-defined persuasion effects under staggered treatment.

The natural starting point is to define forward and backward persuasion rates at time $t$ for units treated at time $s$, conditional on covariates $X_i = x$, along with their marginal versions averaged over covariates:

align*[align* omitted — 406 chars of source]

These extend the two-period persuasion rates. For example, $\theta^{(F)}(s,t)$ measures the share of individuals treated at time $s$ who take the action of interest at time $t$, among those who would have behaved differently without the treatment. In contrast, $\theta^{(B)}(s,t)$ measures the necessity of treatment for observed actions. Both parameters reflect persuasive effects, but from complementary perspectives.

To assess cumulative dynamic persuasion effects, we aggregate across cohorts while fixing the number of periods since treatment. The $j$-period-forward persuasion rate for units treated by time $T-j$ is

align*[align* omitted — 156 chars of source]

called the forward event-study persuasion rate (FES). It captures the average persuasive effect $j$ periods after treatment across all eligible cohorts. By Bayes’ rule,

align[align omitted — 224 chars of source]

Similarly, to measure the necessity of treatment for actions observed $j$ periods later, we aggregate $\theta^{(B)}(s, s+j)$:

align*[align* omitted — 268 chars of source]

the backward event-study persuasion rate (BES). FES captures how treatment induces actions, whereas BES quantifies its necessity for observed actions.\footnote{This concise notation of FES and BES aligns with the acronyms for our baseline persuasion measures, FPR and BPR, maintaining a consistent forward/backward structure throughout the paper.}

Before moving to identification, we present an illustrative example of staggered treatment adoption in (ref). The table describes a four-period balanced panel where observations are grouped by their treatment adoption period $S_i$, and the entries show observed outcomes indexed by the event horizon $j$.

table[table omitted — 1,424 chars of source]

Identification

All parameters above depend on the joint distribution of potential outcomes. As in the two-period case, we now impose a monotonicity assumption.

assumptionFor all $t \geq s$, $\mathbb{P}\{ Y_{it}(s) \geq Y_{it}(\infty) \mid S_i = s, X_i\} = 1$ almost surely.

(ref) generalizes (ref), requiring no backlash regardless of treatment timing. It only compares $Y_{it}(s)$ and $Y_{it}(\infty)$ for each $t \geq s$ and does not restrict outcome trajectories over time.

All the persuasion rates can be expressed in terms of a single unidentified function: i.e., \[ \tau_{\mathrm{ST}}(s, t \mid x) := \mathbb{P}\bigl\{ Y_{it}(\infty) = 1 \,\big|\, S_i = s,\ X_i = x \bigr\}, \qquad t \geq s. \] We identify $\tau_{\mathrm{ST}}$ via a DID approach, for which we extend (ref).

assumptionFor all $t \geq s$, \begin{align*} &\mathbb{P}\bigl\{ Y_{it}(\infty) = 1 \,\big|\, S_i = s,\ X_i \bigr\} - \mathbb{P}\bigl\{ Y_{i,s-1}(\infty) = 1 \,\big|\, S_i = s,\ X_i \bigr\} \\ &= \mathbb{P}\bigl\{ Y_{it}(\infty) = 1 \,\big|\, S_i = \infty,\ X_i \bigr\} - \mathbb{P}\bigl\{ Y_{i,s-1}(\infty) = 1 \,\big|\, S_i = \infty,\ X_i \bigr\}. \end{align*}

(ref) extends the two-period parallel trends assumption, (ref), to staggered treatment adoption. The control group consists of “never-treated” units ($S_i = \infty$), and outcomes are compared between the period of interest $t$ and the last pre-treatment period $s-1$ for units with $S_i = s$. This assumption could be strengthened to allow any not-yet-treated group ($S_i = s' > t$) to serve as a control and to use any earlier pre-treatment period ($s-k$ for $k = 1, \ldots, s-1$) as the comparison period. In this paper, we use the simpler form of (ref), focusing on a primary control group and comparison period. For general discussions of alternative parallel trends assumptions, see, for example, CS:2021,roth2023s, and for recent advances on efficient estimation under stronger versions of this assumption, see Chen2025eff.

By (ref), for $t \geq s$, we can write

align*[align* omitted — 928 chars of source]

which is directly identified from panel data. Thus, we obtain the following identification result under the parallel trends, no-backlash, and no-anticipation assumptions.

theoremUnder (ref), $\tau_{\mathrm{ST}}(s,t \mid x)$ is point-identified by $\Psi_{\mathrm{ST}}(s,t \mid x)$. In particular, for $r \in \{F,B\}$, $\theta^{(r)}_{c}(s,t \mid x)$, $\theta^{(r)}(s,t)$, and $\theta^{(r)}_{\mathrm{ES}}(j)$ are all point-identified.

While we do not state the identification formulas for $\theta^{(r)}_{c}(s,t \mid x)$, $\theta^{(r)}(s,t)$, and $\theta^{(r)}_{\mathrm{ES}}(j)$ explicitly, their derivations are straightforward: e.g.,

align[align omitted — 800 chars of source]

The separate aggregation of the numerator and denominator in (ref) to obtain (ref) follows from Bayes' rule. The expression in (ref) is derived similarly. Since persuasion rates are conditional probabilities, they naturally appear as ratios, and when aggregated (over $X_i$ or $S_i$), the numerator and denominator are always averaged separately by Bayes' rule. Backward persuasion rates are simpler: e.g.,

align[align omitted — 320 chars of source]

Back-of-the-Envelope Inference for Event-Study Persuasion Rates

To discuss an event-study analog of (ref), define

align*[align* omitted — 173 chars of source]

Under (ref), $\mathrm{ATT}_{\mathrm{ES}}(j)$ represents the event-study average treatment effect on the treated, corresponding to $\theta_{\mathrm{ES}}(e)$ in CS:2021 in their notation. For $j = 0,1,\dots, T-1$, the following relationships hold:

equation[equation omitted — 289 chars of source]

where

align*[align* omitted — 216 chars of source]

Here, $q_{\mathrm{ES}}(j)$ is the proportion of treated units with zero outcomes $Y_{i,s+j}=0$ at horizon $j$.

As in (ref), these relationships allow back-of-the-envelope inference on $\theta^{(F)}_{\mathrm{ES}}(j)$ and $\theta^{(B)}_{\mathrm{ES}}(j)$ when estimates of $\mathrm{ATT}_{\mathrm{ES}}(j)$ and $q_{\mathrm{ES}}(j)$ are available. This is particularly appealing given recent advances in estimating $\mathrm{ATT}_{\mathrm{ES}}(j)$ CS:2021,Chen2025eff. In practice, it suffices to have a confidence interval for $q_{\mathrm{ES}}(j)$, interpreted as the share of treated units that do not take the target action, and a confidence interval for $\mathrm{ATT}_{\mathrm{ES}}(j)$. When data on $(Y_{i,s+j}, S_i)$ are available, $q_{\mathrm{ES}}(j)$ can be estimated by

align*[align* omitted — 198 chars of source]

where confidence intervals can be constructed in the usual way.

Estimation and Inference via Joint GMM

All estimators from the two-period case extend naturally to the staggered treatment setting. We focus here on the case without covariates, as it corresponds to the empirical example in (ref); semiparametric methods incorporating covariates are deferred to (ref).

Our main objects of interest are the forward and backward event-study persuasion rates, $\theta^{(F)}_{\mathrm{ES}}(j)$ and $\theta^{(B)}_{\mathrm{ES}}(j)$, which summarize dynamic persuasion effects. Throughout, we maintain the identification conditions in (ref).

For $s\in\{1,\ldots,T,\infty\}$ and calendar time $t\in\{0,\ldots,T\}$, define \[ \mu_{s,t} := \mathbb{E} \big( Y_{it} \mid S_i=s \big), \qquad \pi_s := \mathbb{P}(S_i=s). \] For each cohort $s$ and event time $j$ with $t=s+j\in\{0,\ldots,T\}$, define

align*[align* omitted — 256 chars of source]

The forward and backward persuasion rates at the cohort$\times$time level are, for $r\in \{F,B\}$,

equation[equation omitted — 125 chars of source]

As shown in (ref), aggregating across cohorts with shares $\pi_s$ gives the event-study persuasion rates, for $r\in \{F,B\}$:

equation[equation omitted — 177 chars of source]

Only cohorts $s$ for which $t=s+j$ is in-sample enter these sums.

We can express $\theta^{(r)}(s,s+j)$ as a solution to moment conditions as we did in (ref). However, that approach is not convenient to obtain event-study persuasion rates that require aggregation like (ref). Instead we construct moment conditions that identify all the component parameters below.

Let $\vartheta$ collect all cell means $\{\mu_{s,t}\}_{s\in\{1,\ldots,T,\infty\},\,t \in\{0,\ldots,T\}}$ and cohort shares $\{\pi_s\}_{s=1}^T$. Let $W_i := (Y_{i0},\ldots,Y_{iT},S_i)$, and consider the just-identified moment vector $g(W_i;\vartheta)$ defined by stacking all the elements of the following set: \[ \bigl\{ \mathbb{1}\{S_i \in \{s,t\} \} \bigl( Y_{it} - \mu_{s,t} \bigr) \bigr\}_{(s,t) \ \text{nonempty}} \ \cup\ \bigl\{ \mathbb{1}\{S_i=s\} - \pi_s \bigr\}_{s=1}^T, \] where “nonempty” denotes observed $(s,t)$ cells. The just-identified GMM estimator $\hat\vartheta$ solves the sample moment conditions; in this case, $\hat\mu_{s,t}$ coincides with the corresponding cell mean and $\hat\pi_s$ with the corresponding sample proportion. Define

align*[align* omitted — 354 chars of source]

and form $\widehat\theta^{(F)}(s,s{+}j)$, $\widehat\theta^{(B)}(s,s{+}j)$, and their event-study aggregates $\widehat\theta^{(F)}_{\mathrm{ES}}(j)$ and $\widehat\theta^{(B)}_{\mathrm{ES}}(j)$ by substituting $\widehat\mu$ and $\widehat\pi$ into the population expressions in (ref) and (ref), respectively.

The GMM formulation is convenient for estimating all $\mu_{s,t}$ and $\pi_s$ jointly and for propagating sampling uncertainty into nonlinear aggregates such as $\theta^{(F)}_{\mathrm{ES}}(j)$ and $\theta^{(B)}_{\mathrm{ES}}(j)$. We recommend cluster-robust standard errors at the level of treatment assignment (e.g., province in (ref)).\footnote{It is customary to report event-study estimates with confidence intervals for $j<0$ as a pre-trends check. For this purpose, it is preferable to use event-study ATT estimates rather than persuasion rates, both use the same parallel trend assumption for identification, while the latter are ratio statistics whose denominators may be close to zero, making them more unstable and potentially misleading in pre-treatment periods. Also, inference on the persuasion rates robust to the “small denominator” problem can be done by imposing the null hypothesis. For example, $H_0: \theta^{(F)}_{\mathrm{ES}}(j) = \theta_{F,H}$ can be equivalently formulated as \[ H_0: \sum_{s=1}^{T-j} \textrm{ATT}(s,s+j)\pi_s - \theta_{F,H} \sum_{s=1}^{T-j} \theta^{(F)}_{\mathrm{den}}(s,s+j)\pi_s = 0, \] which can be tested without relying on the delta method. }

Empirical Example II: Curriculum Persuasion

cantoni2017curriculum study the impact of a textbook reform in China, exploiting its staggered introduction across provinces. They surveyed undergraduate students at Peking University from four cohorts who entered high school between 2006 and 2009, and estimated a DID regression cantoni2017curriculum of the form:

align[align omitted — 183 chars of source]

where $i$ indexes individuals, $t$ denotes the high school entry cohort, $P_i$ is the province of high school attendance, $Y_{it}$ captures political attitudes or beliefs. The treatment variable $\text{New Curriculum}_{it}$ is an indicator equal to one if individual $i$, who entered high school as cohort $t$ in province $P_i$, was exposed to the reformed curriculum, and zero otherwise. The coefficients ${\gamma_t}$ and ${\delta_p}$ represent cohort and province fixed effects, respectively, and $\beta$ captures the effect of the curriculum reform.

From this regression, cantoni2017curriculum report persuasion rates cantoni2017curriculum. Let $\widehat{\beta}$, $\widehat{\gamma}_t$, and $\widehat{\delta}_p$ be the corresponding regression estimates. When $Y_{it}$ is not binary, it is dichotomized to one if the outcome exceeds the median response. Then, their persuasion rate is computed as:

align*[align* omitted — 153 chars of source]

where the numerator measures the treatment effect of the curriculum reform, and the denominator approximates “the share of students without the desired attitude among those taught under the old curriculum” cantoni2017curriculum.

table[table omitted — 1,432 chars of source]

If there were only two cohorts and two provinces, $\widehat{\theta}_{\textrm{CCYYZ}}$ would match our two-period regression-based estimator of FPR from (ref). However, cantoni2017curriculum feature a staggered rollout of the new curriculum across multiple cohorts and provinces. Recent work on staggered treatment de2020two,CS:2021, which postdates their analysis, shows that $\widehat{\beta}$ (the numerator of $\widehat{\theta}_{\textrm{CCYYZ}}$) may be a biased estimator when interpreted as an aggregate version of ATT. We therefore revisit their estimation. Although the survey was conducted in a single year, it covers four high school entry cohorts (2006-2009), allowing us to adapt the GMM estimation methods in (ref) to estimate persuasion rates. Since at least one pre-treatment cohort is required, the earliest treated group ($S_i = 1$) consists of students from the 2007 cohort in provinces adopting the curriculum that year. The control group ($S_i = \infty$) consists of students in provinces that adopted it in 2010, the final year of the reform. (ref) summarizes the staggered adoption pattern.

As an illustration, we consider the first outcome variable in cantoni2017curriculum, “Trust: central government.” The left panel of (ref) reports ATT, FPR, and BPR estimates for the $S_i = 2007$ group. The $x$-axis shows event time $s+j$, with 2006 as the pre-treatment baseline. Confidence intervals are clustered at the province level. As expected, both FPR and BPR exceed ATT, as they focus on more relevant subpopulations. FPR is larger than BPR, indicating that there are more already-persuaded (AP) individuals than never-persuadable (NP) individuals. Notably, for 2007, both ATT and BPR are insignificant, whereas FPR is significant, underscoring the importance of going beyond ATT. The right panel shows similar results for the $S_i = 2008$ group, with pre-treatment estimates (2006) near zero and insignificant. Overall, both panels indicate substantial effects of the textbook reform on students' trust in the central government.

figure[figure omitted — 616 chars of source]
figure[figure omitted — 718 chars of source]

We do not report separate estimates for the $S_i = 2009$ group due to imprecision and the availability of only one post-treatment period. Instead, we present FES (forward event-study persuasion rate) and BES (backward event-study persuasion rate), which summarize the overall dynamic effects. (ref) reports FES estimates for the first six outcome variables in cantoni2017curriculum. The blue dashed line shows the persuasion rate reported in the original study, which, as discussed, should be compared to FES. Interpreting FES as a more suitable aggregate than the original persuasion rates, given that we allow heterogeneous treatment effects and refine control groups by excluding provinces treated before 2007, the empirical results suggest that persuasion effects are larger in most cases (that is, the FPR estimates lie above the dashed line). As with FPR and BPR, FES estimates are generally larger than BES estimates, reflecting again that there are more already-persuaded (AP) than never-persuadable (NP) individuals. Pointwise 95% confidence intervals are computed using standard errors clustered at the province level. Pre-treatment effects are generally insignificant,\footnote{From (ref), $\theta^{(F)}_{\mathrm{ES}}(j)$ is more prone to instability than $\theta^{(B)}_{\mathrm{ES}}(j)$ in pre-treatment periods when $q_{\mathrm{ES}}(j)$ is small, as this implies a smaller denominator for $\theta^{(F)}_{\mathrm{ES}}(j)$ due to $\mathrm{ATT}_{\mathrm{ES}}(j)$ being close to zero. This pattern appears to be the case in our example.} except for “Trust: courts,” while most post-treatment effects are significantly positive. Overall, our results reaffirm the main findings of cantoni2017curriculum, providing further evidence that “studying the new curriculum led to more positive views of China's governance.”

appendix\section{Proofs of the Results in the Main Text} Proof of (ref): We generally have \begin{align*} &\mathbb{P}\{ Y_{i1}(1) = 1 \mid D_{i1} = 1, X_i = x\} - \mathbb{P}\{ Y_{i1}(0) = 1 \mid D_{i1} = 1, X_i = x\} \\ &= \mathbb{P}\{ Y_{i1}(0) = 0, Y_{i1}(1) = 1 \mid D_{i1} = 1, X_i = x\} - \mathbb{P}\{ Y_{i1}(0) = 1, Y_{i1}(1) = 0 \mid D_{i1} = 1, X_i = x\}. \end{align*} However, $\mathbb{P}\{ Y_{i1}(0) = 1, Y_{i1}(1) = 0 \mid D_{i1} = 1, X_i = x\} = 0$ under (ref). If (ref) is violated, then the claim follows from the lower bound of Fr\'{e}chet-Hoeffding inequalities: see (ref) in Online Appendix (ref) for more detail. \qed Proofs of (ref): Suppose that (ref) holds for all $x\in \mathcal{X}$. Then, $\theta^{(F)}_c(x) = \theta^{(F)}_{cL}(x)$ and $\theta^{(B)}_c(x) = \theta^{(B)}_{cL}(x)$ for all $x\in \mathcal{X}$ by (ref). Therefore, multiplying the density of $X_i$ at $x$ given $D_{i1} = 1, Y_{i1}(0) = 0$ to both sides of $\theta^{(F)}_c(x) = \theta^{(F)}_{cL}(x)$ and integrating with respect to $x$ shows $\theta^{(F)} = \theta^{(F)}_L$. The other case of $\theta^{(B)} = \theta^{(B)}_L$ is similar. \qed Proof of (ref): It follows from (ref) and the fact that $Y_{i1}(d) = Y_{i1}$ when we condition on $D_{i1}= d$. \qed Proof of (ref): It follows from (ref), and the Bayes rule. \qed (ref) in Online Appendix (ref) contains results on sharp identified bounds when (ref) can be violated. Proof of (ref): Under (ref), the coefficients in (ref) are identified by OLS, where $D_{i1} = G_i$. Therefore, the denominator of $\bar\theta^{(F)}_L$ is \begin{align*} & \bigl\{ 1 - \mathbb{E}(Y_{i0} \mid D_{i1}=1) \bigr\} - \bigl\{ \mathbb{E}( Y_{i1} \mid D_{i1} = 0) - \mathbb{E}(Y_{i0} \mid D_{i1} = 0) \bigr\} \\ &= \bigl\{ 1 - \mathbb{E}(Y_{it} \mid G_i=1, t=0) \bigr\} - \bigl\{ \mathbb{E}( Y_{it} \mid G_i = 0, t=1) - \mathbb{E}(Y_{it} \mid G_i = 0, t=0) \bigr\} \\ &= 1 - (\gamma_0+\gamma_1) - (\gamma_0 + \gamma_2) + \gamma_0 = 1-\gamma_0 - \gamma_1-\gamma_2, \end{align*} and that of $\bar\theta^{(B)}_L$ is $\mathbb{E}( Y_{i1} \mid D_{i1} = 1 ) = \mathbb{E}( Y_{it}\mid G_i = 1, t=1) = \gamma_0 + \gamma_1+\gamma_2 + \gamma$. The numerators of $\theta^{(F)}_L$ and $\theta^{(B)}_L$ are similar. \qed Proof of (ref): Let $\mathbb{E}(D_{i1}) = q_1$, and consider the denominator of the middle expression in (ref): i.e., \[ \frac{\mathrm{Cov}(\tilde Y_{i1} - Y_{i0}, D_{i1})}{\mathbb{V}(D_{i1})} = \frac{\mathbb{E}(\tilde Y_{i1} D_{i1}) - \mathbb{E}(Y_{i0}D_{i1}) - q_1 \{ \mathbb{E}(\tilde Y_{i1}) - \mathbb{E}(Y_{i0}) \}}{q_1(1-q_1)}, \] which is equal to \begin{align*} & \frac{\mathbb{E}(\tilde Y_{i1} D_{i1}) - \mathbb{E}(Y_{i0}D_{i1}) - q_1 \bigl[ \mathbb{E}(\tilde Y_{i1} D_{i1}) + \mathbb{E}\{\tilde Y_{i1}(1-D_{i1}) \} - \mathbb{E}(Y_{i0} D_{i1}) - \mathbb{E}\{ Y_{i0}(1-D_{i1}) \} \bigr]}{q_1(1-q_1)} \\ &= \frac{(1-q_1) \bigl\{ \mathbb{E}(\tilde Y_{i1} D_{i1}) - \mathbb{E}(Y_{i0}D_{i1}) \bigr\} - q_1 \bigl[ \mathbb{E}\{\tilde Y_{i1}(1-D_{i1}) \} - \mathbb{E}\{ Y_{i0}(1-D_{i1}) \} \bigr]}{q_1(1-q_1)} \\ &= \bigl\{ \mathbb{E}(\tilde Y_{i1}\mid D_{i1} = 1) - \mathbb{E}(Y_{i0}\mid D_{i1} = 1) \bigr\} - \bigl\{ \mathbb{E}(\tilde Y_{i1}\mid D_{i1} = 0) - \mathbb{E}(Y_{i0}\mid D_{i1} = 0) \bigr\} \\ &= 1 - \mathbb{E}(Y_{i0}\mid D_{i1} = 1) - \mathbb{E}(Y_{i1}\mid D_{i1} = 0) + \mathbb{E}(Y_{i0}\mid D_{i1} = 0), \end{align*} which is the denominator of $\bar\theta^{(F)}_L$. The numerator of $\bar\theta^{(F)}_L$ is similar. Finally, for the denominator of $\bar\theta^{(B)}_L$, just note that $\mathrm{Cov}( Y_{i1} D_{i1}, D_{i1}) = \mathbb{E}(Y_{i1} \mid D_{i1} = 1) \mathbb{V}(D_{i1})$. \qed \textbf{Proof of (ref): } Equivalence between the right-hand side of (ref) and that of (ref) is a simple algebraic result. The fact that $F_{DID}(Y_0,Y_1,D_1,X)$ is the semiparametrically efficient influence function follows from (ref) in Online Appendix (ref) and the fact that $F_{num}(Y_0,Y_1,D_1,X)$ and $F_{den}(Y_0,Y_1,D_1,X)$ are in the tangent space $\mathcal{T}$ described in (ref) in Online Appendix (ref): see the expressions of $F_{num}(Y_0,Y_1,D_1,X)$ and $F_{den}(Y_0,Y_1,D_1,X)$ given before the proofs of (ref) in Online Appendix (ref). Finally, asymptotic normality in (ref) follows from Theorem 2.1 of newey1994asymptotic. Specifically, the scores are given in (ref) in Online Appendix (ref), and they form a linear space. Also, any mean-zero function $s(Y_0,Y_1,D_1,X) = s_{000}(X) (1-Y_0)(1-Y_1)(1-D_1) + \cdots + s_{111}(X) Y_0Y_1D_1$ can be exactly matched with a score in the form of (ref) in Online Appendix (ref). Therefore, an asymptotically linear estimator of $\bar\theta^{(F)}_L$ must have a unique influence function, which will coincide with $F_{DID}$ by Theorem 2.1 of newey1994asymptotic. \qed \textbf{Proof of (ref): } First, identification of $\tau_{\mathrm{ST}}(s,t\mid x)$ by $\Psi_{\mathrm{ST}}(s,t\mid x)$ for $\ell< s \leq t$ immediately follows from (ref), because $\mathbb{P}\{ Y_{i,s-1}(\infty)=1\mid S_i=s, X_i\} = \mathbb{P}( Y_{i,s-1} =1\mid S_i=s, X_i )$ by (ref) that rules out anticipation. Now, we define the following objects: \begin{align*} \theta^{(F)}_{cL}(s,t\mid x) &:= \frac{\mathbb{P}\{ Y_{it}(s) = 1\mid S_i=s, X_i = x\} - \tau_{\mathrm{ST}}(s,t\mid x)}{1 - \tau_{\mathrm{ST}}(s,t\mid x)}, \\ \theta^{(F)}_{L}(s,t) &:= \frac{\mathbb{E}\bigl[ \mathbb{P}\{ Y_{it}(s) = 1\mid S_i=s, X_i\} - \tau_{\mathrm{ST}}(s,t\mid X_i) \ \big| \ S_i=s\bigr]}{1 - \mathbb{E}\bigl[ \tau_{\mathrm{ST}}(s,t\mid X_i) \ \big| \ S_i=s\bigr]}, \\ \theta^{(F)}_{\mathrm{ES},L}(j) &:= \frac{\sum_{s=1}^{T-j}\theta^{(F)}_L(s,s+j )\mathbb{P}(S_i=s)\bigl[ 1- \mathbb{E}\{ \tau_{\mathrm{ST}}(s,s+j\mid X_i) \mid S_i=s\} \bigr]}{\sum_{s=1}^{T-j}\mathbb{P}(S_i=s)\bigl[ 1- \mathbb{E}\{\tau_{\mathrm{ST}}(s,s+j\mid X_i) \mid S_i=s \} \bigr]}. \end{align*} As in the proof of (ref), we generally have \begin{multline*} \mathbb{P}\{ Y_{it}(s) = 1\mid S_i=s,X_i = x\} - \mathbb{P}\{ Y_{it}(\infty) = 1\mid S_i=s,X_i = x\} \\ = \mathbb{P}\{ Y_{it}(s) = 1, Y_{it}(\infty) = 0 \mid S_i=s,X_i = x\} - \mathbb{P}\{ Y_{it}(s) = 0, Y_{it}(\infty) = 1\mid S_i=s,X_i = x\}. \end{multline*} However, under (ref), $\mathbb{P}\{ Y_{it}(s) = 0, Y_{it}(\infty) = 1\mid S_i=s,X_i = x\} = 0$. Therefore, $\theta^{(F)}_c(s,t\mid x) = \theta^{(F)}_{cL}(s,t\mid x)$. Furthermore, we can verify that $\theta^{(F)}(s,t) = \theta^{(F)}_L(s,t)$ and $\theta^{(F)}_{\mathrm{ES}}(j) = \theta^{(F)}_{\mathrm{ES},L}(j)$ using the fact that for generic random variables $A, B$, and an event $E$, we generally have $\mathbb{E}(A\mid B\in E) = \mathbb{E}\{ \mathbb{E}(A\mid B)\mathbb{1}(B\in E) \}/\mathbb{P}(B\in E)$. Finally, replacing $\tau_{\mathrm{ST}}$ in the denominators of $\theta^{(F)}_{cL}(s,t\mid x), \theta^{(F)}_L(s,t)$, and $\theta^{(F)}_{\mathrm{ES},L}(j)$ with $\Psi_{\mathrm{ST}}$ shows the identification of $\theta^{(F)}_{c}(s,t\mid x), \theta^{(F)}(s,t)$, and $\theta^{(F)}_{\mathrm{ES}}(j)$. The backward parameters with $r=B$ are similar. \qed
appendix\pagenumbering{roman} \setcounter{page}{1} \setcounter{section}{0} \begin{center} {Online Appendices to “Learning the Effect of Persuasion via Difference-in-Differences”} \end{center} \begin{center} \begin{tabular}{ccc} Sung Jae Jun & \qquad & Sokbae Lee \\ Penn State University & \qquad & Columbia University \end{tabular} \end{center} Throughout the online appendices, to simplify notation, we drop the subscript $i$ from variables such as $Y_{i1}$, $Y_{i0}$, $D_{i1}$, $X_i$ and $S_i$ when referring to population quantities or whenever the omission does not cause confusion. \section{The Case of the Backlash} Without making (ref), the Fr\'{e}chet--Hoeffding inequality yields bounds on $\theta^{(F)}_c(x)$ and $\theta^{(B)}_c(x)$ in terms of the marginal probabilities of the potential outcomes conditional on $X=x$. Further, $\theta^{(F)}_c(x)$ and $\theta^{(B)}_c(x)$ are linearly dependent in that for all $x\in \mathcal{X}$, we have \begin{equation} \theta^{(B)}_c(x) = \frac{ \{ 1-\tau_c(x)\} \theta^{(F)}_c(x) }{\mathbb{P}\{Y_1(1) = 1\mid D_1 = 1, X=x\}}. \end{equation} Therefore, assuming that $\tau_c(x)$ is known, it suffices to have bounds on $\theta^{(F)}_c(x)$ to have bounds on $\theta^{(B)}_c(x)$. Below we first derive bounds on $\theta^{(F)}_c(x)$ without using (ref) that depend only on the marginal probabilities of the potential outcomes given $D_1 = 1$ and $X=x$. Then, bounds on $\theta^{(B)}_c(x)$ will follow from the bounds on $\theta^{(F)}_c(x)$ and (ref), or vice versa. Define \begin{align*} \mathcal{B}(x) &:= \Bigl\{ (p,q)\in \mathbb{R}^2:\ \max\{ 0, \theta^{(F)}_{cL}(x) \} \leq p \leq \min\{ \theta^{(F)}_{cU}(x), 1 \},\ q = \alpha(x) p \Bigr\} \\ &= \Bigl\{ (p,q)\in \mathbb{R}^2:\ p = q/\alpha(x),\ \max\{0, \theta^{(B)}_{cL}(x)\} \leq q \leq \min\{ \theta^{(B)}_{cU}(x), 1 \} \Bigr\}, \end{align*} where $\alpha(x) := \{ 1-\tau_c(x)\} / \mathbb{P}\{ Y_1(1) = 1\mid D_1 = 1,X=x \}$, and \begin{align*} \theta^{(F)}_{cU}(x) := \frac{\mathbb{P}\{ Y_1(1) = 1\mid D_1 = 1, X=x \}}{1 - \tau_c(x)}, \quad \theta^{(B)}_{cU}(x) := \frac{1-\tau_c(x)}{\mathbb{P}\{ Y_1(1) = 1\mid D_1 = 1, X=x \}}. \end{align*} We then have the following lemma. \begin{lemma} Suppose that (ref) holds. For all $x\in \mathcal{X}$, we have $\bigl( \theta^{(F)}_c(x), \theta^{(B)}_c(x) \bigr)\in \mathcal{B}$, which is sharp based on the information of $\mathbb{P}\{ Y_1(1) = 1 \mid D_1 = 1, X=x\}$ and $\mathbb{P}\{ Y_1(0) = 1 \mid D_1 = 1, X=x \}$. \proof It follows from (ref) and an application of Lemma I1 in Appendix I of jun2023identifying by setting $\mathbb{P}^*( A\cap B)$ in the lemma to be $\mathbb{P}\{ Y_1(0) \in A, Y_1(1)\in B\mid D_1 = 1, X=x\}$. \qed \end{lemma} The set $\mathcal{B}$ is a line that describes the bounds on $\bigl(\theta^{(F)}_c(x), \theta^{(B)}_c(x) \bigr)$ we can obtain from the marginal probabilities of the potential outcomes (given $D_1 = 1, X=x$) without using (ref). Therefore, (ref) show that rescaling CATT as in the end point $\bigl( \theta^{(F)}_{cL}(x), \theta^{(B)}_{cL}(x)\bigr)$ provides useful causal parameters to consider. Conditioning on $X=x$, $\bigl( \theta^{(F)}_{cL}(x), \theta^{(B)}_{cL}(x)\bigr)$ exactly corresponds to the persuasion rates on the treated that we are interested in if there is no backlash, while they serve as conservative measures in general, even if the backlash effect is a concern. Aggregating $X$ with appropriate conditioning shows that $\theta^{(F)}_L$ and $\theta^{(B)}_L$ are valid lower bounds on $\theta^{(F)}$ and $\theta^{(B)}$, respectively: i.e., even if there is a concern about the backlash so that (ref) may be violated, $\theta^{(F)}_L$ and $\theta^{(B)}_L$ still offer conservative measures of $\theta^{(F)}$ and $\theta^{(B)}$, respectively, although it may not be sharp in general. In order to be more precise on this issue, let \begin{align*} \mathbb{L} &:= \mathbb{E}\bigl\{ \mathbb{1}_{\mathcal{X}_L}(X)CATT(X) \ \big|\ D_1 = 1 \bigr\}, \\ \mathbb{U} &:= \mathbb{E}\bigl[ \mathbb{1}_{\mathcal{X}_U}(X) \mathbb{P}( Y_1 = 1\mid D_1 = 1,X ) + \{ 1- \mathbb{1}_{\mathcal{X}_U}(X)\}\{ 1- \tau_c(X)\} \ \big|\ D_1 = 1 \bigr], \end{align*} where $\mathcal{X}_L$ and $\mathcal{X}_U$ are defined by \begin{align*} \mathcal{X}_L &:= \{ x\in \mathcal{X}: \mathbb{P}(Y_1=1\mid D_1=1,X=x) - \tau_c(x) \geq 0 \}, \\ \mathcal{X}_U &:= \{ x\in \mathcal{X}: \tau_c(x) + \mathbb{P}(Y_1=1\mid D_1=1,X=x) \leq 1\}. \end{align*} Define \begin{align*} \mathcal{B} &:= \Bigl\{ (p,q)\in \mathbb{R}^2: \theta^{(F)*}_L \leq p \leq \theta^{(F)*}_U,\ q = \alpha p \Bigr\} \\ &= \Bigl\{ (p,q)\in \mathbb{R}^2: p = q/\alpha,\ \theta^{(B)*}_L \leq q \leq \theta^{(B)*}_U \Bigr\}, \end{align*} where $\alpha = \mathbb{E}\{ 1- \tau_c(X) \mid D_1 = 1 \} / \mathbb{P}( Y_1 = 1\mid D_1 = 1)$, and \begin{align*} \theta^{(F)*}_L := \mathbb{L} / \mathbb{E}\{ 1-\tau_c(X) \mid D_1 = 1\} &\quad and \quad \theta^{(F)*}_U := \mathbb{U} / \mathbb{E}\{ 1-\tau_c(X) \mid D_1 = 1\}, \\ \theta^{(B)*}_L := \mathbb{L} / \mathbb{P}(Y_1 = 1\mid D_1 = 1) &\quad and \quad \theta^{(B)*}_U := \mathbb{U} / \mathbb{P}(Y_1 = 1\mid D_1 = 1). \end{align*} \begin{lemma} Suppose that (ref) holds for all $x\in \mathcal{X}$. We then have $(\theta^{(F)}, \theta^{(B)})\in \mathcal{B}$, which is sharp based on the information of $\mathbb{P}\{ Y_1(1) = 1 \mid D_1 = 1, X = x \}$ and $\mathbb{P}\{ Y_1(0) = 1 \mid D_1 = 1, X = x\}$ for all $x\in \mathcal{X}$, and the distribution of $X$ given $D_1 = 1$. \proof We will only show the bounds on $\theta^{(F)}$: the set $\mathcal{B}$ will follow from them and the linear relationship between $\theta^{(F)}$ and $\theta^{(B)}$, i.e., $\theta^{(B)} = \theta^{(F)} \mathbb{E}\{ 1- \tau_c(X)\mid D_1 = 1 \}/\mathbb{P}(Y_1= 1\mid D_1 = 1)$. First, we note that \begin{align*} \mathcal{X}_L = \{x\in \mathcal{X}: \theta^{(F)}_{cL}(x) \geq 0\} \quad and\quad \mathcal{X}_U = \{ x\in \mathcal{X}: \theta^{(F)}_{cU}(x) \leq 1\} \end{align*} by definition. Therefore, we can equivalently write the sharp bounds in (ref) as \begin{equation} \mathbb{1}_{\mathcal{X}_L}(x) \theta^{(F)}_{cL}(x) \leq \theta^{(F)}_c(x) \leq \mathbb{1}_{\mathcal{X}_U}(x) \theta^{(F)}_{cU}(x) + 1- \mathbb{1}_{\mathcal{X}_U}(x). \end{equation} We then use the fact \begin{equation} f\{ x\mid Y_1(0)=0,D_1 = 1 \} = \frac{\mathbb{P}\{ Y_1(0) = 0\mid D_1 = 1,X=x\}f(x\mid D_1 = 1)}{\mathbb{P}\{ Y_1(0) = 0\mid D_1 = 1\}} \end{equation} by Bayes' rule. Specifically, multiplying the conditional density in (ref) to both sides of the inequalities in (ref) and integrating shows the claim. \qed \end{lemma} Here, $\tau_c(\cdot)$ is the only unidentified object so that $\theta^{(F)*}_L$ and $\theta^{(F)*}_U$ will be identified if $\tau_c(\cdot)$ is identified. The set $\mathcal{X}_L$ represents the values $x$ of $X$ such that $\textrm{CATT}(x)\geq 0$: i.e., it is the set of $x$ such that the lower bound on $\theta^{(F)}_c(x)$ in (ref) is nontrivial. Similarly, $\mathcal{X}_U$ is the set of the values $x \in \mathcal{X}$ such that the upper bound on $\theta^{(F)}_c(x)$ in (ref) is not trivial: or, equivalently, it is the set of values $x\in \mathcal{X}$ such that the upper bound on $\theta^{(B)}_c(x)$ is trivial. Indeed, the condition that defines $\mathcal{X}_U$ can be equivalently expressed as \begin{equation*} \mathbb{P}\{ Y_1(0) = 1, Y_1(1) = 1\mid D_1 = 1,X=x\} \leq \mathbb{P}\{Y_1(0) = 0, Y_1(1)=0\mid D_1 = 1,X=x\}. \end{equation*} If this inequality is not satisfied so that there are too many `voters' who are characterized by $X=x$ and who would vote for the party the media publicly endorses no matter what, then the upper bound on $\theta^{(F)}_c(x)$ based on the marginals of the potential outcomes will be just trivial, i.e., $\theta^{(F)}_{cU}(x) = 1$. If (ref) holds, then we have $\textrm{CATT}(X)\geq 0$ almost surely, and therefore $\theta^{(F)*}_L = \theta^{(F)}_L$ as well as $\theta^{(B)*}_L = \theta^{(B)}_L$ will follow. Therefore, (ref) show that $\theta^{(F)*}_L$ and $\theta^{(B)*}$ will be interesting parameters to consider: they are sharp lower bounds on $\theta^{(F)}$ and $\theta^{(B)}$, respectively, in general, while they are exactly equal to $\theta^{(F)}$ and $\theta^{(B)}$ under monotonicity. However, $(\theta^{(F)*}_L, \theta^{(B)*}_L)$ is a more difficult parameter than $(\theta^{(F)}_L, \theta^{(B)}_L)$ because the former contains $\mathbb{1}_{\mathcal{X}_L}(X)$ that depends on unknown objects in a nonsmooth way. In contrast, $(\theta^{(F)}_L, \theta^{(B)}_L)$ can be estimated in a more straightforward manner, as long as $\tau_c(\cdot)$ is identified. Therefore, we consider $\theta^{(F)}_L$ and $\theta^{(B)}_L$ the aggregate parameters of interest. Since $\max(0,\theta^{(F)}_L)\leq \theta^{(F)*}_L$ in general, $\theta^{(F)}_L$ is always a robust lower bound on $\theta^{(F)}$; the same comment applies to $\theta^{(B)}_L$ as well. If $\textrm{CATT}(x)$ is nonnegative for almost all $x\in\mathcal{X}$, then $\theta^{(F)}_L$ and $\theta^{(B)}_L$ are the sharp lower bounds $\theta^{(F)*}_L$ and $\theta^{(B)*}_L$, respectively. Further, if (ref) holds, then $\theta^{(F)}_L = \theta^{(F)*}_L = \theta^{(F)}$, and $\theta^{(B)}_L = \theta^{(B)*}_L = \theta^{(B)}$. Partial identification under (ref) without using (ref) is largely uneventful. For example, (ref) show the joint sharp identified set of FCPR and BCPR when (ref) is violated. For the aggregated parameters, the joint sharp identified set requires aggregation over an unknown subset of the support of $X$ in general. Below we clarify this issue, and we make a formal statement about identification of the aggregated parameters. Define \begin{align*} \bar\theta^{(F)*}_L &:= \frac{ \bar{\mathbb{L}} }{\mathbb{E}\{ 1- \Psi(X) \mid D_1 = 1 \}}, \quad \bar\theta^{(F)*}_U := \frac{\bar{\mathbb{U}}}{\mathbb{E}\{ 1- \Psi(X) \mid D_1 = 1 \}}, \\ \bar{\theta}^{(B)*}_L &:= \frac{ \bar{\mathbb{L}} }{\mathbb{E}\{ \Pi_1(1,X) \mid D_1 = 1 \}}, \quad \bar{\theta}^{(B)*}_U := \frac{\bar{\mathbb{U}}}{\mathbb{E}\{ \Pi_1(1,X) \mid D_1 = 1 \}}, \end{align*} where \begin{align*} \bar{\mathbb{L}} &:= \mathbb{E}\bigl[ \mathbb{1}_{\mathcal{X}_L^*}(X)\bigl\{ \Pi_1(1,X) - \Psi(X) \bigr\}\ \big| \ D_1 = 1 \bigr], \\ \bar{\mathbb{U}} &:= \mathbb{E}\bigl[ \mathbb{1}_{\mathcal{X}_U^*}(X)\Pi_1(1,X) + \{ 1-\mathbb{1}_{\mathcal{X}_U^*}(X)\}\{1-\Psi(X)\} \ \big| \ D_1 = 1\bigr], \end{align*} with \[ \mathcal{X}_L^* := \{ x\in \mathcal{X}:\ \Psi(x) \leq \Pi_1(1,x) \} \quad \text{and}\quad \mathcal{X}_U^* := \{ x\in \mathcal{X}:\ \Psi(x) \leq 1 - \Pi_1(1,x) \}. \] Here, $\bar\theta^{(F)}_L,\bar\theta^{(B)}_L, \bar\theta^{(F)*}_L, \bar\theta^{(B)*}_L, \bar\theta^{(F)*}_U$, and $\bar\theta^{(B)*}_U$ are all directly identified from the data. \begin{corollary} Suppose that (ref) are satisfied for all $x\in \mathcal{X}$. Then, $(\theta^{(F)}_L, \theta^{(F)*}_L, \theta^{(F)*}_U)$ and $(\theta^{(B)}_L, \theta^{(B)*}_L, \theta^{(B)*}_U)$ are identified by $(\bar\theta^{(F)}_L,\bar\theta^{(F)*}_L, \bar\theta^{(F)*}_U)$ and $(\bar\theta^{(B)}_L,\bar\theta^{(B)*}_L, \bar\theta^{(B)*}_U)$, respectively. Therefore, \begin{enumerate} • if (ref) hold for all $x\in \mathcal{X}$, then $\theta^{(F)} = \theta^{(F)}_L$ and $\theta^{(B)} = \theta^{(B)}_L$ are point-identified by $\bar\theta^{(F)}_L$ and $\bar\theta^{(B)}_L$, respectively; • if (ref) hold for all $x\in\mathcal{X}$, then the joint sharp identifiable set of $(\theta^{(F)}, \theta^{(B)})$ is given by the line connecting $[\bar\theta^{(F)*}_L,\ \bar\theta^{(B)*}_L ]$, and $[\bar\theta^{(F)*}_U,\ \bar\theta^{(B)*}_U]$. \end{enumerate} \proof Noting that $\mathcal{X}_L^*= \mathcal{X}_L$ and $\mathcal{X}_U^* = \mathcal{X}_U$ under (ref) by (ref), the claim immediately follows from (ref). \qed \end{corollary} Here, $\bar\theta^{(F)}_L$ and $\bar\theta^{(B)}_L$ are natural estimands to focus on. They are conservative measures of FPR and BPR in general, while they are exactly equal to FPR and BPR under monotonicity. Also, they are easier parameters to estimate than $\bar\theta^{(F)*}_L$ or $\bar\theta^{(B)*}_L$, which depends on unknown objects in a nonsmooth way. \section{Using a Known Link Function} As we commented in (ref), (ref) does not allow for popular parametric models such as logit or probit. However, we can modify (ref) to introduce a link function, as long as the link function is pre-specified. Below is a modification of (ref). \begin{assumption}[Parallel Trends] For some known link function $\Lambda$ on $[0,1]$ that is strictly increasing and differentiable, and for all $x\in \mathcal{X}$, $\Lambda\bigl[ \mathbb{P}\bigl\{ Y_t(0) = 1 \mid D_1 = d, X = x\bigr\} \bigr]$ is separable into the sum of a time component and a treatment component: i.e., there exist functions $G$ (of $t,x$) and $H$ (of $d,x$) such that \[ \mathbb{P}\bigl\{ Y_t(0) = 1 \mid D_1 = d, X =x \bigr\} = \Lambda^{-1}\{ G(t,x) + H(d,x) \}. \] \end{assumption} Differentiability of $\Lambda$ will be useful for calculating the efficient influence function, not strictly necessary for identification. The choice of the link function $\Lambda$ is a specification issue for the researcher: e.g., $\Lambda(s) = s$ is an obvious choice that we used throughout the main text. More generally, (ref) allows for the class of generalized linear models. For example, the logistic model with $\Lambda^{-1}(s) = \exp(s)/\{ 1+\exp(s)\}$ and \[ \mathbb{P}\{ Y_t(0) = 1 \mid D_1 = d, X=x\} = \Lambda^{-1}\bigl( \beta_0 + \beta_1 t + \beta_2 d + \beta_3 x + \beta_4^{\mathpalette\raiseT{\intercal}} t x + \beta_5^{\mathpalette\raiseT{\intercal}} dx \bigr) \] does not satisfy (ref), but it does satisfy (ref). In addition to the linear or logistic choice of $\Lambda$, it is worth considering $\Lambda^{-1}(s) = 1- \exp(-s)$ with $s\geq 0$, i.e., the distribution function of the standard exponential distribution. This choice of the link function is for the case where the time and treatment component are multiplicatively separable, and therefore there is common growth as in wooldridge2023simple. Specifically, suppose that $\mathbb{P}\{ Y_t(0) = 0\mid D_1 = d, X=x) = \tilde G(t,x) \tilde H(d,x)$ so that for all $x\in \mathcal{X}$, \[ \frac{\mathbb{P}\{ Y_1(0) = 0 \mid D_1 = 1, X = x\}}{\mathbb{P}\{ Y_0(0) = 0 \mid D_1 = 1, X= x\}} = \frac{\mathbb{P}\{ Y_1(0) = 0 \mid D_1 = 0, X=x\}}{\mathbb{P}\{ Y_0(0) = 0 \mid D_1 = 0, X=x\}}. \] In this case, the choice of $\Lambda^{-1}(s) = 1- \exp(-s)$ leads to \[ \Lambda\bigl[ \mathbb{P}\{ Y_t(0) = 1 \mid D_1 = d, X = x\} \bigr] = - \log\tilde G(t,x) - \log\tilde H(d,x). \] Under (ref), we have parallel trends with the transformation $\Lambda$: i.e., \begin{multline} \Lambda[ \mathbb{P}\bigl\{ Y_1(0) = 1 \mid D_1 = 1, X=x \bigr\} ] - \Lambda[ \mathbb{P}\bigl\{ Y_0(0) = 1\mid D_1 = 1, X=x \bigr\} ] \\ = \Lambda[ \mathbb{P}\bigl\{ Y_1(0) = 1 \mid D_1 = 0, X=x \bigr\} ] - \Lambda[ \mathbb{P}\bigl\{ Y_0(0) = 1 \mid D_1 = 0, X=x \bigr\} ]. \end{multline} Therefore, (ref) continue to hold with the modification of \begin{equation} \Psi(X) := \Lambda^{-1} \Bigl[ \Lambda\{ \Pi_0(1,X) \} + \Lambda\{ \Pi_1(0,X) \} - \Lambda\{ \Pi_0(0,X) \} \Bigr]. \end{equation} Indeed, most of our results in the main text can be extended by using (ref) instead of (ref) with some exceptions: e.g., the plug-in approach that uses $\Psi$ is valid for any choice of $\Lambda$, while the propensity-odss-weighting approach in (ref) relies on the specific choice of $\Lambda(s) = s$. \section{Robust Inference using the Anderson-Rubin Statistic} As we discussed in (ref), inference on FPR or BPR may suffer from the weak instrument problem if their denominators are too close to zero. This issue can be addressed by using an Anderson-Rubin type procedure. Below we elaborate on this topic. We will focus on the two-period case, and we consider inference on FPR first. Recall from (ref) that if $H_{F0}: \bar\theta^{(F)}_L = \theta_{F,H}$ holds, then we have the following moment conditions: \[ \mathbb{E}\bigl[ (Y_{i1}-Y_{i0}) - \beta_0 - \theta_{F,H} (\tilde Y_{i1}-Y_{i0}) \bigr]=0 \ \text{and} \ \mathbb{E}\bigl[D_{i1}\{(Y_{i1}-Y_{i0}) - \beta_0 - \theta_{F,H} (\tilde Y_{i1}-Y_{i0}) \}\bigr]=0, \] where $\tilde Y_{i1} := D_{i1} + Y_{i1}(1-D_{i1})$. Therefore, partialling out $\beta_0$ under the null $H_{F0}$ leads to the moment condition \begin{equation} \mathbb{E}\Bigl( D_{i1} \Bigl[ \bigl\{ Y_{i1} - \mathbb{E}(Y_{i1})-Y_{i0}+\mathbb{E}(Y_{i0}) \bigr\} - \theta_{F,H} \bigl\{ \tilde Y_{i1} - \mathbb{E}(\tilde Y_{i1})-Y_{i0}+\mathbb{E}(Y_{i0}) \bigr\} \Bigr] \Bigr)=0. \end{equation} Here, we note that (ref) is equivalent to \[ \mathrm{Cov}(Y_{i0} - Y_{i0}, D_{i1}) - \theta_{F,H} \mathrm{Cov}(\tilde Y_{i1} - Y_{i0}, D_{i1}) = 0, \] which is obtained by imposing the value $\theta_{F,H}$ to (ref). We can now formulate an Anderson-Rubin statistic that tests the moment condition in (ref). Specifically, having an i.i.d.\ sample $\{ (Y_{i0}, Y_{i1}, D_{i1}): i=1,2,\cdots, n\}$, define \[ \hat \xi_{F,i}(\theta_{F,H}) := D_{i1} \Bigl[ \bigl\{Y_{i1} - \bar Y_1 - Y_{i0} + \bar Y_0 \bigr\} - \theta_{F,H} \bigl\{\tilde Y_{i1} - \bar{\tilde{Y}}_1 - Y_{i0}+\bar Y_0 \bigr\} \Bigr] \] with $\bar Y_t := n^{-1}\sum_{i=1}^n Y_{it}$ for $t\in\{0,1\}$ and $\bar{\tilde{Y}}_1 := n^{-1}\sum_{i=1}^n \tilde Y_{i1}$. Then, under $H_{F0}$, the central limit theorem together with the Slutsky lemma shows that \begin{equation*} \frac{1}{\sqrt{n}}\sum_{i=1}^n \hat\xi_{F,i}(\theta_{F,H}) \stackrel{d}{\rightarrow} N\bigl( 0, V_F(\theta_{F,H}) \bigr), \end{equation*} where $V_F(\theta_{F,H})$ is the variance of \begin{equation} \{ D_{i1} - \mathbb{E}(D_{i1}) \} \Bigl[ \bigl\{ Y_{i1} - \mathbb{E}(Y_{i1})-Y_{i0}+\mathbb{E}(Y_{i0}) \bigr\} - \theta_{F,H} \bigl\{\tilde Y_{i1} - \mathbb{E}(\tilde Y_{i1})-Y_{i0}+\mathbb{E}(Y_{i0}) \bigr\} \Bigr]. \end{equation} Therefore, $V_F(\theta_{F,H})$ can be consistently estimated under $H_{F0}$ by the sample analog, say $\hat V_F(\theta_{F,H})$, of the variance of (ref). Now, the Anderson-Rubin (AR) test is to reject $H_{F0}$ whenever \begin{equation} \Bigl( \frac{1}{\sqrt{n}}\sum_{i=1}^n \hat\xi_{F,i}\theta_{F,H} \Bigr)^2 / \hat V_F(\theta_{F,H}) > c_{1-\alpha}, \end{equation} where $c_{1-\alpha}$ is the $1-\alpha$ quantile of $\chi^2(1)$. Inverting the AR test yields a robust confidence region for FPR. Testing the location of BPR is similar. Consider $H_{B0}: \bar\theta^{(B)}_L = \theta_{B,H}$, and it follows from (ref) that \begin{equation} \mathbb{E}\Bigl( D_{i1} \Bigl[ \bigl\{ Y_{i1} - \mathbb{E}(Y_{i1})-Y_{i0}+\mathbb{E}(Y_{i0}) \bigr\} - \theta_{B,H} \bigl\{ Y_{i1}D_{i1} - \mathbb{E}(Y_{i1 D_{i1}}) \bigr\} \Bigr] \Bigr)=0 \end{equation} under $H_{B0}$. In order to test (ref), define \[ \hat \xi_{B,i}(\theta_{B,H}) := D_{i1} \Bigl[ \bigl\{Y_{i1} - \bar Y_1 - Y_{i0} + \bar Y_0 \bigr\} - \theta_{B,H} \bigl\{ Y_{i1} D_{i1} - \overline{YD} \bigr\} \Bigr], \] where $\overline{YD} := n^{-1}\sum_{i=1}^n Y_{i1}D_{i1}$. Now, by the central limit theorem and the Slutsky lemma, we know that \[ \frac{1}{\sqrt{n}}\sum_{i=1}^n \hat\xi_{B,i}(\theta_{B,H}) \stackrel{d}{\rightarrow} N\bigl(0, V_B(\theta_{B,H})\bigr), \] where $V_B(\theta_{B,H})$ is the variance of \begin{equation} \{ D_{i1} - \mathbb{E}(D_{i1}) \} \Bigl[ \bigl\{ Y_{i1} - \mathbb{E}(Y_{i1})-Y_{i0}+\mathbb{E}(Y_{i0}) \bigr\} - \theta_{B,H} \bigl\{ Y_{i1}D_{i1} - \mathbb{E}(Y_{i1}D_{i1}) \bigr\} \Bigr]. \end{equation} Let $\hat V_B(\theta_{B,H})$ be the sample analog of the variance of the random variable defined in (ref), and the AR test for testing $H_{B0}$ is to reject $H_{B0}$ whenever \begin{equation} \Bigl( \frac{1}{\sqrt{n}}\sum_{i=1}^n\hat\xi_{B,i}(\theta_{F,H}) \Bigr)^2 / \hat V_B(\theta_{B,H}) > c_{1-\alpha}, \end{equation} where $c_{1-\alpha}$ is the $1-\alpha$ quantile of $\chi^2(1)$. Inverting the AR test yields a robust confidence region for BPR. Note that the tests in (ref) and (ref) are identical when $\theta_{F,H} = \theta_{B,H} = 0$. This is not surprising in view of the fact that both BPR and APR have ATT in their numerator: therefore, FPR is zero if and only if BPR is zero. If it is of interest to test $H_{J0}: \bar\theta^{(F)}_L = \theta_{F,H}, \bar\theta^{(B)}_L = \theta_{B,H}$ jointly for $(\theta_{F,H},\theta_{B,H}) \in (0,1]\times (0,1]$, then it can be done by testing the two moment conditions in (ref) and (ref) jointly. This procedure simply entails stacking $\hat\xi_{F,i}(\theta_{F,H})$ and $\hat\xi_{B,i}(\theta_{B,H})$. Specifically, under $H_{J0}$, we have \[ \frac{1}{\sqrt{n}}\sum_{i=1}^n \hat \xi_{J,i}(\theta_{F,H},\theta_{B,H}) \stackrel{d}{\rightarrow} N\bigl( 0, V_J(\theta_{F,H},\theta_{B,H}) \bigr), \] where $\hat \xi_{J,i}(\theta_{F,H},\theta_{B,H}) := \left[ \hat\xi_{F,i}(\theta_{F,H}),\ \hat\xi_{B,i}(\theta_{B,H}) \right]^{\mathpalette\raiseT{\intercal}}$, and $V_J(\theta_{F,H},\theta_{B,H})$ is the covariance matrix of the random vector defined by stacking (ref) and (ref). Therefore, the AR test in this context is to reject $H_{J0}$ whenever \[ \Bigl( \frac{1}{\sqrt{n}}\sum_{i=1}^n \hat \xi_{J,i}(\theta_{F,H},\theta_{B,H}) \Bigr)^{\mathpalette\raiseT{\intercal}} \hat V_J(\theta_{F,H},\theta_{B,H})^{-1} \Bigl( \frac{1}{\sqrt{n}}\sum_{i=1}^n \hat \xi_{J,i}(\theta_{F,H},\theta_{B,H}) \Bigr) > c_{2,1-\alpha}, \] where $\hat V_J(\theta_{F,H},\theta_{B,H})$ is the sample analog of $V_J(\theta_{F,H},\theta_{B,H})$, and $c_{2,1-\alpha}$ is the $1-\alpha$ quantile of $\chi^2(2)$. Inverting the joint AR test yields a joint confidence region for FPR and BPR. However, inverting the joint test requires searching over $(0,1]\times(0,1]$, which may not be computationally attractive in practice. In this context, it is worth noting that a Cartesian product of $1-\alpha_F$ and $1-\alpha_B$ confidence sets of FPR and BPR, respectively, yields a $1-\alpha_F-\alpha_B$ Bonferroni joint confidence region. \section{Further Discussion on Controlling for \\the Pre-Treatment Outcome and Unconfoundedness} In this part of the appendix, we return to the issues discussed in (ref) and provide a more detailed discussion. Specifically, the unconfoundedness condition is an alternative assumption that has been used in the literature especially with cross-sectional data. When the pre-treatment outcome variable is available, it is natural to include it in the conditioning variables to define the unconfoundedness assumption. Below we first compare the unconfoundedness assumption with the parallel trend assumption. Let $Z := [Y_0, X^{\mathpalette\raiseT{\intercal}}]^{\mathpalette\raiseT{\intercal}}$, and consider the following assumptions. \begin{assumption} At time $t=0$, no one is treated. At time $t=1$, there is a constant $\epsilon>0$ such that $\epsilon \leq \min\bigl[ \mathbb{P}\{ Y_1(0) = 0, D_1 = 1\mid Z\}, \mathbb{P}\{ Y_1(1) = 1, D_1 = 1\mid Z\} \bigr]$ and $\mathbb{P}( D_1 = 1\mid Z ) \leq 1-\epsilon$ with probability one. \end{assumption} \begin{assumption} $Y_1(0)$ is independent of $D_1$ conditional on $Z$. \end{assumption} (ref) is simply setting up the same environment as in the main text, but with $Z$ in lieu of $X$ (see (ref)). (ref) imposes unconfoundedness given $Z$. We remark that (ref) holds with $Z$ in lieu of $X$ if and only if (ref) holds: this can be verified by using the fact that $Z$ contains $Y_0$. Under (ref), both $\theta^{(F)}_L$ and $\theta^{(B)}_L$ are well-defined, and they are identified by \begin{equation} \begin{aligned} \tilde\theta^{(F)}_{L, Z} &:= \frac{ \mathbb{E} \bigl\{ D_1 (Y_1 - Y_0) \bigr\} - \mathbb{E}\bigl\{ D_1 \mathbb{E} ( Y_1 - Y_0 | D_1 = 0, Z) \bigr\}} { \mathbb{E} \bigl\{ D_1 (1 - Y_0) \bigr\} - \mathbb{E}\bigl\{D_1 \mathbb{E} ( Y_1 - Y_0 | D_1 = 0, Z) \bigr\}}, \\ \tilde\theta^{(B)}_{L,Z} &:= \frac{\mathbb{E} \bigl\{ D_1 (Y_1 - Y_0) \bigr\} - \mathbb{E}\bigl\{ D_1 \mathbb{E} ( Y_1 - Y_0 | D_1 = 0, Z) \bigr\}}{\mathbb{E}(D_1 Y_1)}, \end{aligned} \end{equation} where we use the fact that $Y_0$ is included in $Z$. The expressions in (ref) are reminiscent of $\bar\theta^{(F)}_L$ and $\bar\theta^{(B)}_L$ that identify $\theta^{(F)}_L$ and $\theta^{(B)}_L$ under (ref): see (ref). Indeed, \begin{equation} \begin{aligned} \bar\theta^{(F)}_L &= \frac{ \mathbb{E} \bigl\{ D_1 (Y_1 - Y_0) \bigr\} - \mathbb{E}\bigl\{ D_1 \mathbb{E} ( Y_1 - Y_0 | D_1 = 0, X) \bigr\}} { \mathbb{E} \bigl\{ D_1 (1 - Y_0) \bigr\} - \mathbb{E}\bigl\{D_1 \mathbb{E} ( Y_1 - Y_0 | D_1 = 0, X) \bigr\}}, \\ \bar\theta^{(B)}_L &= \frac{ \mathbb{E} \bigl\{ D_1 (Y_1 - Y_0) \bigr\} - \mathbb{E}\bigl\{ D_1 \mathbb{E} ( Y_1 - Y_0 | D_1 = 0, X) \bigr\}} { \mathbb{E}(D_1 Y_1)} \end{aligned} \end{equation} have the same forms as $\tilde\theta^{(F)}_{L,Z}$ and $\tilde\theta^{(B)}_{L,Z}$ except that they do not include $Y_0$ in the controls. Put differently, adding $Y_0$ to $X$ and using the DID formula is an implementation of identifying $\textrm{ATT}$ via (ref) instead of (ref). One might prefer the unconfoundedness condition after controlling for both $X$ and $Y_0$ if it is important to avoid the functional form restriction of the parallel trend assumption. Alternatively, the parallel trend assumption might be favored since it would not require fully conditioning on $Y_0$. Generally speaking, the required $X$ under the parallel trend assumption could be different from $X$ under the unconfoundedness assumption. If both are the same, which identification assumption to adopt is just reduced to the matter of whether to include $Y_0$ in the covariates or not. Further, there is a simple testable condition under which it becomes moot to distinguish the two identification strategies. If $Y_0$ is independent of $D_1$ conditional on $X$, then it can be shown that $\tilde\theta^{(F)}_{L,Z} = \bar\theta_L$ and $\tilde\theta^{(B)}_{L,Z} = \bar\theta^{(B)}_L$. Below we provide a more detailed discussion on this point. We first consider the following assumptions. \begin{assumption} $Y_1(0)$ is independent of $D_1$ conditional on $X$. \end{assumption} \begin{assumption} $Y_0$ is independent of $D_1$ conditional on $X$. \end{assumption} Unlike (ref), (ref) imposes unconfoundedness given $X$ only. (ref) are not testable, but (ref) is: $Y_0$ is observed for all observational units, but $Y_1(0)$ is not. Let $\tilde\theta^{(F)}_{L,X}$ and $\tilde\theta^{(B)}_{L,X}$ be estimands that identify $\theta^{(F)}_L$ and $\theta^{(B)}_L$, respectively, under unconfoundedness given $X$ (i.e., (ref)). That is, define \begin{align*} \tilde\theta^{(F)}_{L,X} := \frac{ \mathbb{E}( D_1 Y_1 ) - \mathbb{E}\{ D_1 \mathbb{E} ( Y_1 | D_1 = 0, X) \}} { \mathbb{E}( D_1 ) - \mathbb{E}\{D_1 \mathbb{E} ( Y_1 | D_1 = 0, X)\}}, \ \tilde\theta^{(B)}_{L,X} := \frac{ \mathbb{E}(D_1 Y_1) - \mathbb{E}\{ D_1 \mathbb{E} ( Y_1 | D_1 = 0, X)\}} { \mathbb{E}( D_1 Y_1 ) }. \end{align*} Now, in view of (ref), note that $\tilde\theta^{(F)}_{L,Z}$ and $\tilde\theta^{(B)}_{L,Z}$ can be equivalently written as $\tilde{\mathcal{N}}/\bigl[ \tilde{\mathcal{N}} +\mathbb{E}\{ D_1(1-Y_1) \} \bigr]$ and $\tilde{\mathcal{N}}/\mathbb{E}(D_1Y_1)$, respectively, where \[ \tilde{\mathcal{N}} := \mathbb{E}\Biggl\{ D_1(Y_1 - Y_0) - (1-D_1)(Y_1-Y_0) \frac{\mathbb{P}(D_1 = 1\mid Z)}{\mathbb{P}(D_1 = 0\mid Z)} \Biggr\}. \] Recall that $\bar \theta^{(F)}_L = \mathcal{N}/\bigl[ \mathcal{N} + \mathbb{E}\{ D_1(1-Y_1)\} \bigr]$ and $\bar\theta^{(B)}_L = \mathcal{N}/\mathbb{E}(D_1Y_1)$, where $\mathcal{N}$ is given in (ref). Therefore, comparing $\tilde{\mathcal{N}}$ with $\mathcal{N}$ shows that (ref) implies that $\bar \theta^{(F)}_L = \tilde\theta^{(F)}_{L,Z}$, and $\bar\theta^{(B)}_L = \tilde\theta^{(B)}_{L,Z}$. Also, it follows from \[ \mathbb{E}\bigl\{D_1 \mathbb{E} ( Y_1 | D_1 = 0, Z) \bigr\} = \mathbb{E}\Biggl\{ Y_1(1-D_1) \frac{\mathbb{P}(D_1=1\mid Z)}{\mathbb{P}(D_1 = 0\mid Z)} \Biggr\} \] that (ref) is again sufficient to ensure that $\tilde\theta^{(F)}_{L,Z} = \tilde\theta^{(F)}_{L,X}$ as well as $\tilde\theta^{(B)}_{L,Z} = \tilde\theta^{(B)}_{L,X}$. Our discussion so far can be summarized as in the following remark. \begin{remark} Suppose that (ref) holds. \begin{enumerate} • If (ref) holds, then (ref) is the same as (ref), while we have $\bar \theta^{(F)}_L = \tilde\theta^{(F)}_{L,Z} = \tilde\theta^{(F)}_{L,X}$ and $\bar\theta^{(B)}_L = \tilde\theta^{(B)}_{L,Z} = \tilde\theta^{(B)}_{L,X}$. Hence, $\theta^{(F)}_L$ and $\theta^{(B)}_L$ are identified by $\bar \theta^{(F)}_L = \tilde\theta^{(F)}_{L,Z} = \tilde\theta^{(F)}_{L,X}$ and $\bar\theta^{(B)}_L = \tilde\theta^{(B)}_{L,Z} = \tilde\theta^{(B)}_{L,X}$, respectively, under either (ref) or (ref).\footnote{In fact, it is easy to verify that (ref) implies (ref) in this case.} • If (ref) does not hold, then the researcher needs to take a stance among assumptions (ref), (ref), or (ref). \end{enumerate} \end{remark} \begin{figure}[htb] \caption{Examples of Data Generating Processes with $Z =[Y_0,X^{\mathpalette\raiseT{\intercal}}]^{\mathpalette\raiseT{\intercal}}$ } \begin{tabular}{ccc} {Panel A:\ $Y_0 \centernot{\protect\mathpalette{\protect\independenT}{\perp}} D_1\mid X$} & & {Panel B:\ $Y_0 \protect\mathpalette{\protect\independenT}{\perp} D_1 \mid X$} \\ & & \end{tabular} \parbox{6in}{ {Notes: Variables in circles are unobserved and they are independent. (ref) holds in both cases.}} \end{figure} (ref) illustrates some potential data generating processes. In Panel A, $Y_0$ is not independent of $D_1$ given $X$. Therefore, the researcher needs to take a stance. In this diagram, (ref) is satisfied, and therefore, we can identify $(\theta^{(F)}_L,\theta^{(B)}_L)$ by $(\tilde\theta^{(F)}_{L,Z},\tilde\theta^{(B)}_{L,Z})$. However, conditioning only on $X$ is not sufficient to deliver independence of $D_1$ and $Y_1(0)$, so $\tilde\theta^{(F)}_{L,X}$ and $\tilde\theta^{(B)}_{L,X}$ are not valid estimands to identify $\theta^{(F)}_L$ and $\theta^{(B)}_L$. It is not clear from this diagram whether (ref) is satisfied or not though. In Panel B, $D_1$ is independent of $Y_0$ given $X$. Hence, it does not matter what stance the researcher takes: $\bar \theta^{(F)}_L$, $\tilde\theta^{(F)}_{L,Z}$, and $\tilde\theta^{(F)}_{L,X}$ are all equal to $\theta^{(F)}_L$, while $\bar\theta^{(B)}_L, \tilde\theta^{(B)}_{L,Z}$, and $\tilde{\theta}^{(B)}_{L,X}$ are all equal to $\theta^{(B)}_L$. In the diagram, (ref) holds, and therefore, both (ref) are satisfied as well. The main takeaway from this discussion is that if $Y_0$ is independent of $D_1$ given $X$, then it is largely an unimportant question whether to control for $Y_0$ or not, or whether to rely on unconfoundedness or parallel trends. However, if $Y_0$ is not independent of $D_1$ given $X$, then the researcher needs to decide carefully which estimand to rely on to learn about $\theta^{(F)}_L$ or $\theta^{(B)}_L$. It seems feasible to establish a bracketing relationship between the two identification approaches, drawing on the findings of Ding_Li_2019 regarding ATT; however, we leave this extension for future research. \section{Derivation of the Tangent Space} We observe $(Y_0, Y_1, D_1, X^{\mathpalette\raiseT{\intercal}})^{\mathpalette\raiseT{\intercal}}$: we also observe $D_0$, but it is irrelevant for our discussion here, because $D_0 = 0$ with probability one by the setup. Let $f$ be the density of $X$, and let $P(X) := \mathbb{P}(D_1 = 1\mid X)$. For $d \in \{0,1\}$, let $Q_d(X) := \mathbb{P}(Y_0 = 1\mid D_1 = d, X)$. Further, for $d, y \in \{0,1\}$, let $R_{dy}(X) := \mathbb{P}( Y_1 = 1\mid D_1 = d, Y_0 = y, X)$. Then, the likelihood is the product of the following terms, while each line corresponds to one term: \begin{align*} & f(X) \\ & P(X)^{D_1} \{ 1- P(X)\}^{1-D_1} \\ & \bigl[ Q_1(X)^{Y_0}\{ 1- Q_1(X) \}^{1-Y_0} \bigr]^{D_1} \\ & \bigl[ Q_0(X)^{Y_0}\{ 1- Q_0(X) \}^{1-Y_0} \bigr]^{1-D_1} \\ & \Bigl( \bigl[ R_{11}(X)^{Y_1} \{ 1- R_{11}(X)\}^{1-Y_1} \bigr]^{Y_0} \bigl[ R_{10}(X)^{Y_1} \{ 1- R_{10}(X)\}^{1-Y_1} \bigr]^{1-Y_0} \Bigr)^{D_1} \\ &\Bigl( \bigl[ R_{01}(X)^{Y_1} \{ 1- R_{01}(X)\}^{1-Y_1} \bigr]^{Y_0} \bigl[ R_{00}(X)^{Y_1} \{ 1- R_{00}(X)\}^{1-Y_1} \bigr]^{1-Y_0} \Bigr)^{1-D_1}. \end{align*} We will use $\gamma$ to denote regular parametric submodels with $\gamma_0$ being the truth: e.g., $f(X;\gamma_0) = f(X)$. Here is our first lemma. \begin{lemma} Suppose that (ref) holds. Then, the tangent space has the following form: \begin{align*} \mathcal{T} := \Biggl\{ \alpha_0(X) &+ \{ D_1 - P(X) \} \alpha_1(X) + D_1\{ Y_0 - Q_1(X) \} \alpha_2(X) \\ &+ (1-D_1)\{ Y_0 - Q_0(X) \} \alpha_3(X) + D_1Y_0\{ Y_1 - R_{11}(X) \} \alpha_4(X) \\ &+ D_1(1-Y_0) \{Y_1 - R_{10}(X) \} \alpha_5(X) + (1-D_1) Y_0 \{ Y_1 - R_{01}(X) \} \alpha_6(X) \\ &+ (1-D_1) (1-Y_0) \{Y_1 - R_{00}(X) \} \alpha_7(X) \Biggr\}, \end{align*} where $\alpha_j$'s are all functions of $X$ such that $\mathbb{E}\{ \alpha_0(X) \} = 0$ and $\mathbb{E}\{ \alpha_j^2(X) \} < \infty$ for $j=0,1,\cdots, 7$. \proof The loglikelihood of regular parametric submodel is given by \[ \ell(\gamma) := \ell_0(\gamma) + \ell_{D_1}(\gamma) + \ell_{Y_0|D_1}(\gamma) + \ell_{Y_1|D_1, Y_0}(\gamma), \] where \begin{align*} \ell_0(\gamma) &:= \log f(X;\gamma), \\ \ell_{D_1}(\gamma) &:= D_1 \log P(X;\gamma) + (1-D_1)\log\{ 1- P(X;\gamma)\}, \\ \ell_{Y_0|D_1}(\gamma) &= D_1 \Bigl[ Y_0 \log Q_1(X;\gamma) + (1-Y_0) \log\{ 1- Q_1(X;\gamma) \} \Bigr] \\ &\qquad \qquad + (1-D_1)\Bigl[ Y_0 \log Q_0(X;\gamma) + (1-Y_0) \log\{ 1- Q_0(X;\gamma) \} \Bigr], \end{align*} and \begin{align*} \ell_{Y_1|D_1,Y_0}(\gamma) &:= D_1 \Bigl( Y_0 \bigl[ Y_1 \log R_{11}(X;\gamma) + (1-Y_1) \log\{ 1- R_{11}(X;\gamma)\} \bigr] \\ &\qquad \qquad +(1-Y_0) \bigl[ Y_1 \log R_{10}(X;\gamma) + (1-Y_1) \log\{ 1- R_{10}(X;\gamma)\} \bigr] \Bigr) \\ & + (1-D_1) \Bigl( Y_0 \bigl[ Y_1 \log R_{01}(X;\gamma) + (1-Y_1) \log\{ 1- R_{01}(X;\gamma)\} \bigr] \\ &\qquad \qquad +(1-Y_0) \bigl[ Y_1 \log R_{00}(X;\gamma) + (1-Y_1) \log\{ 1- R_{00}(X;\gamma)\} \bigr] \Bigr). \end{align*} Therefore, the score at the truth has the following form: \begin{align} &S(Y_0,Y_1,D_1,X) := \frac{1}{f(X)} \frac{\partial f(X; \gamma_0)}{\partial \gamma} + \frac{D_1 - P(X) }{P(X)\{ 1- P(X)\}} \frac{\partial P(X; \gamma_0)}{\partial \gamma} \\ &+ \frac{D_1\{ Y_0 - Q_1(X)\}}{Q_1(X)\{ 1- Q_1(X)\}} \frac{\partial Q_1(X; \gamma_0)}{\partial \gamma} + \frac{(1-D_1)\{ Y_0 - Q_0(X)\}}{Q_0(X)\{ 1- Q_0(X)\}} \frac{\partial Q_0(X; \gamma_0)}{\partial \gamma} \notag \\ &+ \frac{D_1Y_0 \{ Y_1 - R_{11}(X)\}}{R_{11}(X)\{ 1- R_{11}(X)\}} \frac{\partial R_{11}(X; \gamma_0)}{\partial \gamma} + \frac{D_1(1-Y_0) \{ Y_1 - R_{10}(X)\}}{R_{10}(X)\{ 1- R_{10}(X)\}} \frac{\partial R_{10}(X; \gamma_0)}{\partial \gamma} \notag \\ &+ \frac{(1-D_1)Y_0 \{ Y_1 - R_{01}(X)\}}{R_{01}(X)\{ 1- R_{01}(X)\}} \frac{\partial R_{01}(X; \gamma_0)}{\partial \gamma} + \frac{(1-D_1)(1-Y_0) \{ Y_1 - R_{00}(X)\}}{R_{00}(X)\{ 1- R_{00}(X)\}} \frac{\partial R_{00}(X; \gamma_0)}{\partial \gamma}, \notag \end{align} from which the lemma follows, because the derivatives are not restricted except for square integrability. \qed \end{lemma} \section{Derivation of the Pathwise Derivatives} Let \begin{align*} \bar\theta^{(F)}_{L,num} &:= \mathbb{E} \Bigl[ \bigl\{ \Pi_1(1,X) - \Pi_0(1,X) - \Pi_1(0,X) + \Pi_0(0,X) \bigr\} P(X) \Bigr], \\ \bar\theta^{(F)}_{L,den} &:= \mathbb{E} \Bigl[ \bigl\{ 1 - \Pi_0(1,X) - \Pi_1(0,X) + \Pi_0(0,X) \bigr\} P(X) \Bigr] \end{align*} so that $\bar \theta^{(F)}_L = \bar\theta^{(F)}_{L,num}/ \bar\theta^{(F)}_{L,den}$. Define \begin{align*} F_{num}(Y_0,Y_1,D_1,X) &:= \{ \Pi_1(1,X) - \Pi_0(1,X) - \Pi_1(0,X) + \Pi_0(0,X) \} D_1 - \bar\theta^{(F)}_{L,num} \\ & + D_1 \big[ \{ Y_1 - \Pi_1(1,X) \} - \{ Y_0 - \Pi_0(1,X) \} \big] \\ & - \frac{P(X)}{1-P(X)} (1-D_1) \big[ \{ Y_1 - \Pi_1(0,X) \} - \{ Y_0 - \Pi_0(0,X) \} \big]. \end{align*} Similarly, define \begin{align*} F_{den}&(Y_0,Y_1,D_1,X) := \{ 1 - \Pi_0(1,X) - \Pi_1(0,X) + \Pi_0(0,X) \} D_1 - \bar\theta^{(F)}_{L,den} \\ & - D_1 \{ Y_0 - \Pi_0(1,X) \} - \frac{P(X)}{1-P(X)} (1-D_1) \big[ \{ Y_1 - \Pi_1(0,X) \} - \{ Y_0 - \Pi_0(0,X) \} \big]. \end{align*} We will derive the pathwise derivatives of the numerator and denominator of $\bar\theta^{(F)}_L$ in a few steps. The following two lemmas show that $F_{num}(Y_0,Y_1,D_1,X)$ and $F_{dem}(Y_0,Y_1,D_1,X)$ are the pathwise derivatives of $\bar\theta^{(F)}_{L,num}$ and $\bar\theta^{(F)}_{L,den}$, respectively. \begin{lemma} Suppose (ref) is satisfied. Then, the pathwise derivative of $\bar\theta^{(F)}_{L,num}$ is given by $F_{num}(Y_0,Y_1,D_1,X)$. \proof Using the fact that \begin{equation} \Pi_0(d,X) = Q_d(X) \quad and \quad \Pi_1(d,X) = Q_d(X) R_{d1}(X) + \{1-Q_d(X)\} R_{d0}(X), \end{equation} we can write \begin{multline*} \bar\theta^{(F)}_{L,num} = \int \bigl[ Q_1(x) R_{11}(x) + \{1-Q_1(x)\} R_{10}(x) - Q_1(x) \\ - Q_0(x) R_{01}(x) - \{1-Q_0(x)\} R_{00}(x) + Q_0(x) \bigr] P(x) f(x) dx. \end{multline*} Therefore, the pathwise perturbation $\bar\theta^{(F)}_{L,num}(\gamma)$ of $\bar\theta^{(F)}_{L,num}$ is given by \begin{multline} \bar\theta^{(F)}_{L,num} (\gamma) = \int \bigl[ Q_1(x,\gamma) R_{11}(x,\gamma) + \{1-Q_1(x,\gamma)\} R_{10}(x,\gamma) - Q_1(x,\gamma) \\ - Q_0(x,\gamma) R_{01}(x,\gamma) - \{1-Q_0(x,\gamma)\} R_{00}(x,\gamma) + Q_0(x,\gamma) \bigr] P(x,\gamma) f(x,\gamma) dx. \end{multline} Now, by straightforward algebra, $\partial \bar\theta^{(F)}_{L,num}(\gamma_0)/\partial \gamma$ is equal to the sum of the following terms (with each line corresponding to one term): \begin{align*} & \mathbb{E}\Bigl[ \{ \Pi_1(1,X) - \Pi_0(1,X) - \Pi_1(0,X) + \Pi_0(0,X) \} P(X) \frac{\partial f(X;\gamma_0)}{\partial \gamma}\frac{1}{f(X)} \Bigr] \\ & \mathbb{E}\Bigl[ \{ \Pi_1(1,X) - \Pi_0(1,X) - \Pi_1(0,X) + \Pi_0(0,X) \} \frac{\partial P(X;\gamma_0)}{\partial \gamma} \Bigr] \\ & \mathbb{E}\Bigl[ \{ R_{11}(X) - R_{10}(X) - 1 \} P(X) \frac{\partial Q_1(X;\gamma_0)}{\partial \gamma} \Bigr] \\ & \mathbb{E}\Bigl[ \{ R_{00}(X) - R_{01}(X) + 1 \} P(X) \frac{\partial Q_0(X;\gamma_0)}{\partial \gamma} \Bigr] \\ & \mathbb{E}\Bigl[ Q_1(X) P(X) \frac{\partial R_{11}(X;\gamma_0)}{\partial \gamma} \Bigr] \\ & \mathbb{E}\Bigl[ \{ 1-Q_1(X) \} P(X) \frac{\partial R_{10}(X;\gamma_0)}{\partial \gamma} \Bigr] \\ & \mathbb{E}\Bigl[ - Q_0(X) P(X) \frac{\partial R_{01}(X;\gamma_0)}{\partial \gamma} \Bigr] \\ & \mathbb{E}\Bigl[ -\{ 1-Q_0(X) \} P(X) \frac{\partial R_{00}(X;\gamma_0)}{\partial \gamma} \Bigr] \end{align*} We are looking for some $F(Y_0,Y_1,D_1,X)$ with mean zero that satisfies \begin{equation} \frac{\partial \bar\theta^{(F)}_{L,num}(\gamma_0)}{\partial \gamma} = \mathbb{E}\bigl\{ F(Y_0,Y_1,D_1,X) S(Y_0,Y_1,D_1,X) \bigr\}, \end{equation} where $S(Y_0,Y_1,D_1,X)$ is the score described in (ref). Using the fact that the variance of a binary variable is the `'`success” probability times that of “failure,” we know by inspection that such $F(Y_0,Y_1,D_1,X)$ must be given by \begin{align*} F(Y_0,Y_1,D_1,X) & := \{ \Pi_1(1,X) - \Pi_0(1,X) - \Pi_1(0,X) + \Pi_0(0,X) \} P(X) - \bar\theta^{(F)}_{L,num} \\ & + \{ \Pi_1(1,X) - \Pi_0(1,X) - \Pi_1(0,X) + \Pi_0(0,X) \} \{D_1 - P(X) \} \\ & + \{ R_{11}(X) - R_{10}(X) - 1 \} P(X) \frac{D_1\{Y_0 - Q_1(X)\}}{P(X)} \\ & + \{ R_{00}(X) - R_{01}(X) + 1 \} P(X) \frac{(1-D_1)\{ Y_0 - Q_0(X) \}}{1-P(X)} \\ & + Q_1(X) P(X) \frac{D_1Y_0 \{ Y_1 - R_{11}(X)\} }{P(X) Q_1(X) } \\ & + \{ 1-Q_1(X) \} P(X) \frac{D_1(1-Y_0) \{ Y_1 - R_{10}(X)\} }{P(X) \{1-Q_1(X)\} } \\ & - Q_0(X) P(X) \frac{(1-D_1)Y_0 \{ Y_1 - R_{01}(X)\} }{\{1-P(X)\} Q_0(X) } \\ & - \{ 1-Q_0(X) \} P(X) \frac{(1-D_1)(1-Y_0) \{ Y_1 - R_{00}(X)\} }{\{1-P(X)\} \{1-Q_0(X)\} }. \end{align*} However, simplifying $F(Y_0,Y_1,D_1,X)$ by using (ref) yields \[ F(Y_0,Y_1,D_1,X) = F_{num}(Y_0,Y_1,D_1,X). \qedhere \] \end{lemma} \begin{lemma} Suppose (ref) is satisfied. Then, the pathwise derivative of $\bar\theta^{(F)}_{L,den}$ is given by $F_{den}(Y_0,Y_1,D_1,X)$. \proof Using (ref), we can write \begin{equation*} \bar\theta^{(F)}_{L,den} = \int \{ 1 - Q_1(x) - Q_0(x) R_{01}(x) - \{1-Q_0(x)\} R_{00}(x) + Q_0(x) \} P(x) f(x) dx. \end{equation*} Therefore, the pathwise perturbation of $\bar\theta^{(F)}_{L,den}$ is given by \begin{multline} \bar\theta^{(F)}_{L,den} (\gamma) := \int \bigl\{ 1 - Q_1(x,\gamma) - Q_0(x,\gamma) R_{01}(x,\gamma) \\ - \{1-Q_0(x,\gamma)\} R_{00}(x,\gamma) + Q_0(x,\gamma) \bigr\} P(x,\gamma) f(x,\gamma) dx. \end{multline} Hence, by simple algebra, $\partial\bar\theta^{(F)}_{L,den}(\gamma_0)/\partial\gamma$ is equal to the sum of the following terms (with each line corresponding to one term): \begin{align*} & \mathbb{E}\Bigl[ \{ 1 - \Pi_0(1,X) - \Pi_1(0,X) + \Pi_0(0,X) \} P(X) \frac{\partial f(X;\gamma_0)}{\partial \gamma}\frac{1}{f(X)} \Bigr] \\ & \mathbb{E}\Bigl[ \{ 1 - \Pi_0(1,X) - \Pi_1(0,X) + \Pi_0(0,X) \} \frac{\partial P(X;\gamma_0)}{\partial \gamma} \Bigr] \\ & \mathbb{E}\Bigl[ - P(X) \frac{\partial Q_1(X;\gamma_0)}{\partial \gamma} \Bigr] \\ & \mathbb{E}\Bigl[ \{ R_{00}(X) - R_{01}(X) + 1 \} P(X) \frac{\partial Q_0(X;\gamma_0)}{\partial \gamma} \Bigr] \\ & \mathbb{E}\Bigl[ - Q_0(X) P(X) \frac{\partial R_{01}(X;\gamma_0)}{\partial \gamma} \Bigr] \\ & \mathbb{E}\Bigl[ -\{ 1-Q_0(X) \} P(X) \frac{\partial R_{00}(X;\gamma_0)}{\partial \gamma} \Bigr]. \end{align*} Now, we are looking for some $F(Y_0,Y_1,D_1,X)$ with mean zero that satisfies \begin{equation} \frac{\partial \bar\theta^{(F)}_{L,den}(\gamma_0)}{\partial \gamma} = \mathbb{E}\bigl\{ F(Y_0,Y_1,D_1,X) S(Y_0,Y_1,D_1,X) \bigr\}, \end{equation} where $S(Y_0,Y_1,D_1,X)$ is the score described in (ref). Using the fact that the variance of a binary variable is the `'`success” probability times that of “failure,” we know by inspection that such $F(Y_0,Y_1,D_1,X)$ must be given by \begin{align*} F(Y_0,Y_1,D_1,X) &:= \{ 1 - \Pi_0(1,X) - \Pi_1(0,X) + \Pi_0(0,X) \} P(X) - \bar\theta^{(F)}_{L,den} \\ & + \{ 1 - \Pi_0(1,X) - \Pi_1(0,X) + \Pi_0(0,X) \} \{D_1 - P(X) \} \\ & - P(X) \frac{D_1\{Y_0 - Q_1(X)\}}{P(X)} \\ & + \{ R_{00}(X) - R_{01}(X) + 1 \} P(X) \frac{(1-D_1)\{ Y_0 - Q_0(X) \}}{1-P(X)} \\ & - Q_0(X) P(X) \frac{(1-D_1)Y_0 \{ Y_1 - R_{01}(X)\} }{\{1-P(X)\} Q_0(X) } \\ & - \{ 1-Q_0(X) \} P(X) \frac{(1-D_1)(1-Y_0) \{ Y_1 - R_{00}(X)\} }{\{1-P(X)\} \{1-Q_0(X)\} }. \end{align*} However, simplifying $F(Y_0,Y_1,D_1,X)$ by using (ref) yields \[ F(Y_0,Y_1,D_1,X) = F_{den}(Y_0,Y_1,D_1,X). \qedhere \] \end{lemma} In Section (ref), we have derived the scores at $\gamma_0$ and the tangent space. Also, we have shown that the pathwise derivatives of $\bar\theta^{(F)}_{L,num}$ and $\bar\theta^{(F)}_{L,den}$ are given by $F_{num}(Y_0,Y_1,D_1,X)$ and $F_{den}(Y_0,Y_1,D_1,X)$, respectively. By using these results, we now derive the pathwise derivative of $\bar\theta^{(F)}_L = \bar\theta^{(F)}_{L,num}/\bar\theta^{(F)}_{L,den}$ below. \begin{lemma} Suppose (ref) is satisfied. Then, the pathwise derivative of $\bar\theta^{(F)}_L$ is given by \begin{equation} G(Y_0,Y_1,D_1,X) := \frac{1}{\bar\theta^{(F)}_{L,den}}\Biggl( F_{num}(Y_0,Y_1,D_1,X) - \bar\theta^{(F)}_L F_{den}(Y_0,Y_1,D_1,X) \Biggr). \end{equation} \proof Since $\bar\theta^{(F)}_L = \bar\theta^{(F)}_{L,num}/\bar\theta^{(F)}_{L,den}$, we write the pathwise perturbation of $\bar\theta^{(F)}_L$ as $\bar\theta^{(F)}_L(\gamma) = \bar\theta^{(F)}_{L,num}(\gamma)/\bar\theta^{(F)}_{L,den}(\gamma)$, where the truth is denoted by $\gamma_0$: see the section on the derivation of the pathwise derivatives of $\bar\theta^{(F)}_{L,num}$ and $\bar\theta^{(F)}_{L,den}$. Now, we need to show that $G(Y_0,Y_1,D_1,X)$ satisfies \begin{equation*} \frac{\partial \bar\theta^{(F)}_L (\gamma_0)}{\partial \gamma} = \mathbb{E}\bigl\{ G(Y_0,Y_1,D_1,X) S(Y_0,Y_1,D_1,X) \bigr\}, \end{equation*} where $S(Y_0,Y_1,D_1,X)$ is the score described in (ref) in Online Appendix (ref). But, noting that \begin{equation} \frac{\partial \bar\theta^{(F)}_L (\gamma_0)}{\partial \gamma} = \frac{1}{\bar\theta^{(F)}_{L,den}} \bigg( \frac{\partial \bar\theta^{(F)}_{L,num} (\gamma_0)}{\partial \gamma} - \bar\theta^{(F)}_L \frac{\partial \bar\theta^{(F)}_{L,den} (\gamma_0)}{\partial \gamma} \bigg), \end{equation} it follows from the fact that the pathwise derivatives of $\bar\theta^{(F)}_{L,num}$ and $\bar\theta^{(F)}_{L,den}$ are given by $F_{num}(Y_0,Y_1,D_1,X)$ and $F_{den}(Y_0,Y_1,D_1,X)$, respectively: see Online Appendix (ref) on the derivations of the pathwise derivatives. \qed \end{lemma} \begin{lemma} We have $F_{DID}(Y_0,Y_1,D_1,X) = G(Y_0,Y_1,D_1,X)$, where $F_{DID}(Y_0,Y_1,D_1,X)$ is defined in (ref) in (ref). \proof This is a simple algebraic result. Specifically, if we plug $F_{num}(Y_0,Y_1,D_1,X)$ and $F_{den}(Y_0,Y_1,D_1,X)$ into (ref), then it follows that \begin{multline} G(Y_0,Y_1,D_1,X) = \frac{1}{\bar\theta^{(F)}_{L,den}}\biggl( H_1(Y_0,Y_1,D_1,X) + H_2(Y_0,Y_1,D_1,X) + H_3(Y_0,Y_1,D_1,X) \biggr), \end{multline} where \begin{align*} H_{1}(Y_0,Y_1,D_1,X) &:= D_1 \big[ \{ \Pi_1(1,X) - \Pi_0(1,X) - \Pi_1(0,X) + \Pi_0(0,X) \} \\ &\qquad \qquad - \bar\theta^{(F)}_L \{ 1 - \Pi_0(1,X) - \Pi_1(0,X) + \Pi_0(0,X) \} \big], \\ H_{2}(Y_0,Y_1,D_1,X) &:= D_1 \big[ \{ Y_1 - \Pi_1(1,X) \} - (1 - \bar\theta^{(F)}_L) \{ Y_0 - \Pi_0(1,X) \} \big] \\ H_{3}(Y_0,Y_1,D_1,X) &:= (\bar\theta^{(F)}_L - 1) \frac{P(X)}{1-P(X)} (1-D_1) \bigl[ \{ Y_1 - \Pi_1(0,X) \} - \{ Y_0 - \Pi_0(0,X) \} \bigr]. \end{align*} Further simplifications yield the form of $F_{DID}(Y_0,Y_1,D_1,X)$ defined in (ref) in (ref). In addition, the form of $F_{DID}(Y_0,Y)1,D_1,X)$ in (ref) can be obtained by manipulating the terms in (ref). \qed \end{lemma} \section{Semiparametric Estimation with Covariates in Staggered Treatment Designs} This section describes semiparametric methods that incorporate covariates in the staggered treatment setting. For the corresponding regression-based approaches without covariates, see (ref). Let $\bar S_s = \mathbb{1}(S = s) + \mathbb{1}(S = \infty)$ for each $s \in \{1, 2, \ldots, T - j\}$. We extend the semiparametric estimation framework from the two-period case in (ref) by conditioning on $\bar S_s = 1$. For example, $\theta^{(F)}(s,s+j)$ can be similarly expressed as in (ref) by \begin{align} \theta^{(F)}(s,s+j) &= \frac{\mathcal{N}_{\mathrm{ST}}(s,s+j\mid \bar S_s = 1)} { \mathcal{N}_{\mathrm{ST}}(s,s+j\mid \bar S_s = 1) + \mathbb{P}(Y_{s+j} = 0, S=s \mid \bar S_s = 1)}, \end{align} where \begin{align*} \mathcal{N}_{\mathrm{ST}}(s,s+j\mid \bar S_s = 1) &:= \mathbb{P}(Y_{s+j}=1,S=s\mid \bar S_s = 1) - \mathbb{P}(Y_{s-1}=1,S=s\mid \bar S_s = 1) \\ & - \mathbb{E}\Bigl\{ (Y_{s+j} - Y_{s-1}) \mathbb{1}(S=\infty) \frac{\mathbb{P}(S=s\mid X, \bar S_s = 1)}{1-\mathbb{P}(S=s \mid X, \bar S_s = 1)} \ \Big| \ \bar S_s = 1 \Bigr\}. \end{align*} We make “conditioning on $\bar S_s = 1$” explicit here to emphasize that our discussions on estimation and inference in the case of two periods can be applied in the current case of multiple periods, although it is in fact immaterial because $\mathbb{P}(\bar S_s = 1)$ can be multiplied to the numerator and the denominator of (ref). The other equivalent expressions based on DID or PI can be similarly obtained, and therefore, our discussions on semiparametric estimation and inference in the case of two periods can be applied to the current case of multiple periods. To be more concrete, we focus on the doubly robust estimator of the most comprehensive forward measure $\theta^{(F)}_{ES}(j)$ based on a random sample $\{(Y_{i0}, Y_{i1},\cdots, Y_{iT}, S_i, X_i^{\mathpalette\raiseT{\intercal}})^{\mathpalette\raiseT{\intercal}}:\ i=1,2,\cdots, n\}$. Also, let $\bar S_{i,s} = \mathbb{1}(S_i = s) + \mathbb{1}(S_i = \infty)$: i.e., $\bar S_{i,s}=1$ means that $S_i = s$ or $S_i = \infty$. We now start with the DR estimator of $\theta^{(F)}(s,s+j)$ for $s = 1,2,\cdots, T-j$, for which we use the same formula for the two-period DR estimator $\hat\theta^{(F)}_{L,DR}$ defined in (ref), but we only use a subsample of $(Y_{i,s-1}, Y_{i,s+j}, \mathbb{1}(S_i=s), X_i^{\mathpalette\raiseT{\intercal}})^{\mathpalette\raiseT{\intercal}}$ that satisfies $\mathbb{1}(S_i=s \text{ or } S_i = \infty)=1$. Noting that $\mathbb{1}(S_i = s)\mathbb{1}(\bar S_{i,s} = 1) = \mathbb{1}(S_i=s)$, we know that (ref) leads to the following formula: \begin{equation} \hat\theta^{(F)}_{DR}(s,s+j) := \frac{\sum_{i=1}^{n} \bigl\{ (Y_{i,s+j} - Y_{i,s-1} - \widehat\Delta_{s,s+j}(\infty,X_i) \bigr\} \mathbb{1}(S_i = s) - \widehat{C}(s,s+j)}{\sum_{i=1}^{n} \bigl\{ (1 - Y_{i,s-1} - \widehat\Delta_{s,s+j}(\infty,X_i) \bigr\} \mathbb{1}(S_i = s) - \widehat{C}(s,s+j)}, \end{equation} where \begin{align*} \widehat\Delta_{s,s+j}(\infty,X_i) &:= \widehat{\mathbb{P}}(Y_{i,s+j}=1\mid S_i = \infty, X_i) - \widehat{\mathbb{P}}(Y_{i,s-1}=1\mid S_i = \infty, X_i), \\ \widehat{C}(s,s+j) &:= \sum_{i=1}^{n} \frac{\widehat{\mathbb{P}}(S_i = s\mid X_i,\bar S_{i,s} =1)\mathbb{1}(S_i = \infty)}{1-\widehat{\mathbb{P}}(S_i = s\mid X_i, \bar S_{i,s} = 1)}\{ (Y_{i,s+j} - Y_{i,s-1}) - \widehat\Delta_{s,s+j}(\infty,X_i) \}. \end{align*} Indeed, the numerator of (ref) (divided by $n$) is a DR estimator of \[ \mathbb{P}(Y_{1,s+j}=1, S_i = s) - \mathbb{E}\Bigl\{ \Psi^{(s-1)}_{\mathrm{ST}}(s,s+j\mid X_i) \mathbb{1}(S_i = s) \Bigr\}, \] which is equal to $\mathbb{P}\{ Y_{s+j}(s) = 1, Y_{s+j}(\infty) = 0, S_i = s \}$ by (ref). Now, aggregating over $s$ in $\hat\theta_{DR}(s,s+j)$ yields a DR estimator of $\theta^{(F)}_{ES}(j)$, i.e., \begin{align} \hat\theta^{(F)}_{ES,DR}(j) := \frac{\sum_{s=1}^{T-j} \Bigl[\sum_{i=1}^{n} \bigl\{ (Y_{i,s+j} - Y_{i,s-1} - \widehat\Delta_{s,s+j}(\infty,X_i) \bigr\} \mathbb{1}(S_i = s) - \widehat{C}(s,s+j) \Bigr]}{\sum_{s=1}^{T-j} \Bigl[ \sum_{i=1}^{n} \bigl\{ (1 - Y_{i,s-1} - \widehat\Delta_{s,s+j}(\infty,X_i) \bigr\} \mathbb{1}(S_i = s) - \widehat{C}(s,s+j) \Bigr]}. \end{align} Unlike the regression approach, neither the numerator nor the denominator of (ref) is weighted by $\widehat{\mathbb{P}}(S_i = s)$. This is because the regression approach is based on objects like $\mathbb{E}(Y_{it}\mid S_i=s)$, whereas the semiparametric approach uses objects like $\mathbb{E}\{ Y_{it} \mathbb{1}(S_i=1) \}$; this distinction matters in aggregating across $s$. For the DR estimator of $\theta^{(B)}_{ES}(j)$, we simply replace the denominator of (ref) with \[ \sum_{s=1}^{T-j} \Bigl\{ \sum_{i=1}^{n} \mathbb{1}(S_i = s) Y_{i,s+j} \Bigr\}. \] We discuss the asymptotic distributions in Online Appendix (ref). \section{Inference for the Event-Study Persuasion Rate} The asymptotic influence function for regular and asymptotically linear estimators of the multi-period parameters can be derived from our previous discussion on the two-period case. We will explain this by focusing on $\hat\theta^{(F)}_{ES,DR}(j)$, which we consider the most comprehensive forward parameter: the backward version is similar, and we will remark about it later. Throughout this section, we assume that (ref) are all satisfied for identification, and we consider a random sample of size $n$ with fixed $T$, i.e., $\{(Y_{i0}, Y_{i1},\cdots, Y_{iT}, S_i, X_i): i=1,\cdots, n \}$ and $\bar S_{i,s} := \mathbb{1}(S_i = s) + \mathbb{1}(S_i = \infty)$. First, let \begin{align*} \theta_{num,DR}(s,s+j) &:= \mathbb{P}\{ Y_{i,s+j}(s) = 1, Y_{i,s+j}(\infty) = 0, S_i = s \},\\ \theta_{den,DR}(s,s+j) &:= \mathbb{P}\{ Y_{i,s+j}(\infty) = 0, S_i = s \}, \end{align*} and we have \[ \theta^{(F)}_{ES}(j) = \frac{\sum_{s=1}^{T-j}\theta^{(F)}_{num,DR}(s,s+j)}{\sum_{s=1}^{T-j}\theta_{den,DR}(s,s+j)}. \] For $k\in \{1,2,\cdots, s-1\}$, further let $\mathcal{D}_{i,s,s+j} := \bigl( Y_{i,s-1}, Y_{i,s+j}, S_i, X_i \bigr)$, and define \begin{align*} G_{num}(\mathcal{D}_{i,s,s+j}) &:= \mathbb{1}(S_i = s)(Y_{i,s+j} - Y_{i,s-1}) - \frac{\mathbb{P}(S_i = s\mid X_i,\bar S_{i,s}=1 \}}{\mathbb{P}(S_i = \infty \mid X_i,\bar S_{i,s}=1 \}}(Y_{i,s+j} - Y_{i,s-1})\mathbb{1}(S_i=\infty), \\ G_{den}(\mathcal{D}_{i,s,s+j}) &:= \mathbb{1}(S_i=s)(1-Y_{i,s-1}) - \frac{\mathbb{P}(S_i = s\mid X_i,\bar S_{i,s}=1 \}}{\mathbb{P}(S_i = \infty \mid X_i,\bar S_{i,s}=1 \}}(Y_{i,s+j} - Y_{i,s-1})\mathbb{1}(S_i=\infty), \\ G_{adj}(\mathcal{D}_{i,s,s+j}) &:= -\Bigl\{ \mathbb{1}(S_i=s) - \mathbb{1}(S_i=\infty) \frac{\mathbb{P}(S_i = s\mid X_i,\bar S_{i,s}=1 )}{\mathbb{P}(S_i = \infty \mid X_i,\bar S_{i,s}=1 ) } \Bigr\} \Delta_{s,s+j}(\infty, X_i), \\ H_{num}(\bar S_{i,s}) &:= \frac{\theta_{num,DR}(s,s+j)}{\mathbb{P}(\bar S_{i,s}=1)} \Bigl\{ \bar S_{i,s} - \mathbb{P}(\bar S_{i,s}=1) \Bigr\}, \\ H_{den}(\bar S_{i,s}) &:= \frac{\theta_{den,DR}(s,s+j)}{\mathbb{P}(\bar S_{i,s}=1)} \Bigl\{ \bar S_{i,s} - \mathbb{P}(\bar S_{i,s}=1) \Bigr\}, \end{align*} where \[ \Delta_{s,s+j}(\infty,X_i) := \mathbb{P}(Y_{i,s+j}=1 \mid S_i=\infty,X_i) - \mathbb{P}( Y_{i,s-1}=1 \mid S_i=\infty,X_i). \] Further, let \[ \mathbb{R}_i := \left[ \begin{array}{cc} \mathbb{Q}_{num,i} & \mathbb{Q}_{den,i} \end{array} \right]^{\mathpalette\raiseT{\intercal}} \] where for $r\in \{num,den\}$, \[ \mathbb{Q}_{r,i} := \sum_{s=1}^{T-j}\Bigl\{ G_r(\mathcal{D}_{i, s,s+j}) + G_{adj}(\mathcal{D}_{i, s,s+j}) + H_r(\bar S_{i,1}) \Bigr\} \] Finally, let $\hat\theta_{num,DR}(s,s+j)$ and $\hat\theta_{num,DR}(s,s+j)$ denote the numerator and the denominator (rescaled by $n$), respectively, of the right-hand expresion in (ref): \begin{align*} \hat\theta_{num,DR}(s,s+j) &:= \frac{1}{n} \sum_{i=1}^{n} \bigl\{ (Y_{i,s+j} - Y_{i,s-1} - \widehat\Delta_{s,s+j}(\infty,X_i) \bigr\} \mathbb{1}(S_i = s) - \widehat{C}(s,s+j), \\ \hat\theta_{den,DR}(s,s+j) &:= \frac{1}{n}\sum_{i=1}^{n} \bigl\{ (1 - Y_{i,s-1} - \widehat\Delta_{s,s+j}(\infty,X_i) \bigr\} \mathbb{1}(S_i = s) - \widehat{C}(s,s+j), \end{align*} which are the DR estimators of $\theta_{num,DR}(s,s+j)$ and $\theta_{den,DR}(s,s+j)$. We have the following theorem. \begin{theorem} If $\hat\theta_{num.DR}(s,s+j)$ and $\hat\theta_{den,DR}(s,s+j)$ are regular and asymptotically linear, then \[ \sqrt{n} \Biggl( \hat\theta^{(F)}_{ES,DR}(j) - \theta^{(F)}_{ES}(j) \Biggr) \stackrel{d}{\rightarrow} N\Bigl(0, J \Sigma J^{\mathpalette\raiseT{\intercal}} \Bigr), \] where \begin{align*} J := \left[ \begin{array}{cc} \frac{1}{\sum_{s=1}^{T-j} \theta_{den}(s,s+j)} & -\frac{\sum_{s=1}^{T-j} \theta_{den}(s,s+j)}{\{\sum_{s=1}^{T-j} \theta_{den}(s,s+j)\}^2} \end{array} \right] \quad and\quad \Sigma &:= \mathbb{E}\bigl( \mathbb{R}_1 \mathbb{R}_1^{\mathpalette\raiseT{\intercal}} \bigr). \end{align*} \end{theorem} \textbf{Proof: } Note that \begin{align*} \frac{\theta_{num,DR}(s,s+j)}{\mathbb{P}(\bar S_{i,s}=1)} &= \mathbb{P}\{ Y_{i,s+j}(s) = 1, Y_{i,s+j}(\infty) = 0, S_i = s \mid \bar S_{i,s}=1\} \\ \frac{\theta_{den,DR}(s,s+j)}{\mathbb{P}(\bar S_{i,s}=1)} &= \mathbb{P}\{ Y_{i,s+j}(\infty) = 0, S_i = s \mid \bar S_{i,s}=1\}. \end{align*} Further, for $r\in \{num,den\}$, the numerical value of $n \hat\theta_{r,DR}(s,s+j)$ does not change whether we use only a subsample satisfying $\bar S_{i,s} = 1$ or not. Therefore, we can repeat the same calculations as in the proof of (ref) but by using the conditional likelihood given $\bar S_{i,s} = 1$. Therefore, conditioning on $\bar S_{i,s}=1$ and letting $N_{s,\infty}$ be the size of the subsample that satisfies $\bar S_{i,s} = 1$, we know that a regular and asymptotically linear estimator $\hat\theta_{r,DR}(s,s+j)$ must have the following expansion: \begin{align*} \frac{n\hat\theta_{r,DR}(s,s+j)}{N_{s,\infty}} - \frac{\theta_{r,DR}(s,s+j)}{\mathbb{P}(\bar S_{i,s}=1)} = \frac{1}{N_{s,\infty}} \sum_{i=1}^{N_{s,\infty}} \{G_r(\mathcal{D}_{i, s,s+j}) + G_{adj}(\mathcal{D}_{i, s,s+j}) \} + o_p\left( \frac{1}{\sqrt{N_{s,\infty}}} \right), \end{align*} where we note that $G_r(\mathcal{D}_{i, s,s+j})$ and $G_{adj}(\mathcal{D}_{i, s,s+j})$ have the same forms as in the two-period case except that they use extra-conditioning on $\bar S_{i,s}=1$. But, $N_{s,\infty}/n = \mathbb{P}(\bar S_{i,s}=1) + O_p(1/\sqrt{n})$, and hence \begin{multline} \hat\theta_{r,DR}(s,s+j) - \theta_{r,DR}(s,s+j)\frac{N_{s,\infty}/n}{\mathbb{P}(\bar S_{i,s}=1)} \\ = \frac{1}{n} \sum_{i=1}^{n} \{G_r(\mathcal{D}_{i, s,s+j}) + G_{adj}(\mathcal{D}_{i, s,s+j}) \}\bar S_{i,s} + o_p\left(\frac{1}{\sqrt{n}} \right). \end{multline} Finally, we note that \[ \frac{N_{s,\infty}/n}{\mathbb{P}(\bar S_{i,s}=1)} = 1 + \frac{1}{\mathbb{P}(\bar S_{i,s}=1)} \frac{1}{n}\sum_{i=1}^n \bigl\{ \bar S_{i,s} - \mathbb{P}(\bar S_{i,s}=1) \bigr\} + o_p\left(\frac{1}{\sqrt{n}}\right), \] from which \begin{align*} \sqrt{n} &\{\hat\theta_{r,DR}(s,s+j) - \theta_{r,DR}(s,s+j) \} \\ &= \frac{1}{\sqrt{n}} \sum_{i=1}^{n} \{G_r(\mathcal{D}_{i, s,s+j}) + G_{adj}(\mathcal{D}_{i, s,s+j}) \} \bar S_{i,s} + \frac{1}{\sqrt{n}}\sum_{i=1}^n H_r(\bar S_{i,s}) + o_p(1) \\ &= \frac{1}{\sqrt{n}} \sum_{i=1}^{n} \{G_r(\mathcal{D}_{i, s,s+j}) + G_{adj}(\mathcal{D}_{i, s,s+j}) \} + \frac{1}{\sqrt{n}}\sum_{i=1}^n H_r(\bar S_{i,s}) + o_p(1), \end{align*} where the second equality is trivial by the definitions of $\bar S_{i,s}, G_r, G_{adj}$. Therefore, we have obtained \begin{equation} \sqrt{n}\sum_{s=1}^{T-j}\bigl\{ \hat\theta_{r,DR}(s,s+j) - \theta_{r,DR}(s,s+j) \bigr\} = \frac{1}{\sqrt{n}}\sum_{i=1}^n \mathbb{Q}_{r,i} + o_p(1). \end{equation} Now, use the expressions in (ref) for $r\in\{num, den\}$ together with the delta method, resulting in \begin{equation} \sqrt{n} \Biggl( \hat\theta^{(F)}_{ES,DR}(j) - \theta^{(F)}_{ES}(j) \Biggr) = \frac{1}{\sqrt{n}}\sum_{i=1}^n J \mathbb{R}_i + o_p(1). \end{equation} \qed \remark The case of the backward parameter $\theta^{(B)}_{ES}(j)$ is similar. Recall that \[ \theta^{(B)}_{ES}(j) = \frac{\sum_{s=1}^{T-j}\theta_{num,DR}(s,s+j)}{\sum_{s=1}^{T-j} \mathbb{P}(S_i = s, Y_{i,s+j}=1)}, \] and its semiparametric DR estimator can be expressed as \[ \hat\theta^{(B)}_{ES,DR}(j) = \frac{\sum_{s=1}^{T-j} \hat\theta_{num,DR}(s,s+j)}{\sum_{s=1}^{T-j} \Bigl(n^{-1}\sum_{i=1}^{n} \mathbb{1}(S_i = s) Y_{i, s+j} \Bigr)}. \] The influence function for the numerator is given in (ref). The denominator is just a simple sample average. Therefore, by the delta method, \begin{multline*} \sqrt{n}\Bigl\{ \hat\theta^{(B)}_{ES,DR}(j)- \theta^{(B)}_{ES,DR}(j) \Bigr\} \\ = \frac{1}{\sqrt{n}}\sum_{i=1}^n J \left[ \begin{array}{c} \mathbb{Q}_{num,i} \\ \sum_{s=1}^{T-j} \{ \mathbb{1}(S_i = s) Y_{i s+j} - \mathbb{P}(S_i = s, Y_{i,s+j}=1) \} \end{array} \right] + o_p(1). \end{multline*}