Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
81,035 characters · 24 sections · 38 citation commands
Factorial Difference-in-Differences
\def\spacingset#1{ {#1}} \spacingset{1}
\if11 \fi
\if01 {
} \fi
\addtocounter{page}{-1} \thispagestyle{empty}
\spacingset{1.8}
Social science research often relies on panel data to establish causality. One common approach involves exploiting cross-sectional variation in a baseline factor $G$ and temporal variation in exposure to a common event affecting all units, and applying the difference-in-differences (DID) estimator in a panel setting. As our running example, cao2022clans examine how social capital ($G$), measured by the density of genealogy books, mitigated the mortality surge during China's Great Famine from 1958 to 1961 (the event), using a county-year panel. The authors interpret the coefficient of the interaction term between $G$ and an indicator of the famine years from a two-way fixed effects (TWFE) regression as the causal effect of social capital on famine relief, and describe their approach as a DID method. Section (ref) gives details on this study and five additional examples that employ a similar approach.
Although the coefficient from a TWFE regression is numerically equal to a DID estimate and this approach is often referred to as a DID method in the empirical literature, it differs from the canonical DID popularized by card because it lacks a clean control group unexposed to the event. The corresponding causal interpretation of the DID estimator in such settings is absent from the methodological literature, and empirical applications using this approach lack clear definitions of target causal estimands.
This paper aims to bring conceptual clarity to this empirical approach, which we term {\it factorial difference-in-differences} (FDID). We define FDID as a research design, or an identification strategy, that employs the DID estimator to recover interpretable quantities of interest using observations before and after a one-time event that affects all units, provided clearly stated identifying assumptions hold. Here, we highlight the distinction between the DID estimator and the canonical DID and FDID research designs. Applied to panel data, the DID estimator, denoted by $\hat\tau_\textsc{did}$, calculates the difference in before-after differences with respect to an event between two groups, denoted by $G_i = 1$ and $G_i = 0$. In contrast, a research design encompasses not only the estimator but also the identifying assumptions and identification results. To simplify the presentation, we will refer to the FDID and canonical DID research designs as “FDID” and “canonical DID,” respectively, when no confusion is likely to arise.
We present our main theoretical results for FDID in the two-group, two-period case. The key innovation is to augment the potential outcomes framework to include both the baseline factor $G$ and the exposure indicator $Z$, motivating the term “factorial” in FDID. This formulation clarifies what the probability limit of $\hat\tau_\textsc{did}$, denoted by $\tau_{\textsc{did}}$, identifies under different identification assumptions, and how these results relate to canonical DID.
In canonical DID, $\tau_{\textsc{did}}$ identifies the average treatment effect on the treated (ATT) under the no anticipation and parallel trends assumptions angrist2009mostly. In FDID, the focus shifts to two other estimands: effect modification and causal moderation tyler, bansak2020estimating. Effect modification captures how the effect of $Z$ varies across groups defined by $G$, but does not represent $G$’s causal effect. In our running example, this corresponds to the statement that the mortality increase caused by the famine is smaller in counties with higher levels of social capital. Causal moderation, by contrast, has a direct causal interpretation as $G$’s effect on the impact of $Z$, or symmetrically, $Z$’s effect on the impact of $G$. In our example, causal moderation corresponds to two equivalent statements: (i) social capital reduced the famine’s negative impact on mortality, as emphasized in the original paper, or (ii) social capital had a stronger effect on mortality during the famine than it would have otherwise. Our key result is that under the no anticipation and canonical parallel trends assumptions, $\tau_{\textsc{did}}$ identifies effect modification, whereas recovering causal moderation requires additional assumptions. One such condition is the factorial parallel trends assumption, that is, mean independence between $G$ and potential outcomes trends. Intuitively, it holds in the absence of any unobserved confounder correlated with both $G$ and the outcome trends. For example, to identify the causal moderation of social capital on the famine’s effect on mortality, one must rule out any unobserved factor—such as income or governance quality—that is correlated with both social capital and changes in mortality between famine and non-famine years.
Moreover, we show that canonical DID can be reframed as a special case of FDID under an additional exclusion restriction requiring that exposure to the event has no effect on the outcome for units with $G_i = 0$. Under this assumption, $G$'s effect modification simplifies to the average effect of $Z$ on units with $G_i = 1$, analogous to the ATT in canonical DID. In our running example, this assumption implies that the famine had no impact on localities with low (or high) levels of social capital, an implausible claim in this setting. Alternatively, if researchers assume that $G$ has no causal impact on the outcome in the absence of the event, then causal moderation reduces to $G$'s average conditional effect given exposure to the event. With our example, this assumption implies that social capital would have no effect on mortality had the famine not occurred, which is possible but not suggested by the original paper.
We extend the framework to settings where canonical and factorial parallel trends hold only conditional on additional time-invariant covariates. We formalize identification results for the conditional DID estimator and clarify the assumptions needed to justify regression-based analysis. We also provide theoretical results for applications with repeated cross-sectional data and with a continuous baseline factor $G$.
Our contributions are twofold. First, we establish the causal interpretation of a widely used empirical approach in the social sciences, clarifying the identifying assumptions required to recover causal estimands of interest. This framework advances the discussion of causal panel analysis with the DID estimator and TWFE models---for recent reviews, see roth2023s, chiu2023, and imbens2023panel. Second, we contribute to the literature on factorial designs tyler, bansak2020estimating, han2021contrast, pashley2023causal, yu2023balancing by, to our knowledge, being the first to extend factorial designs to observational panel settings and to analyze the role of parallel trends assumptions in this context.
The rest of the paper is organized as follows. Section (ref) gives six FDID examples appearing in the empirical literature. Section (ref) formalizes FDID under the two-group, two-period panel case. Section (ref) states identification results for FDID and reconciles FDID with canonical DID. Sections (ref)--(ref) discuss extensions to conditionally valid assumptions, repeated cross-sectional data, and continuous $G$. Section (ref) illustrates our theory with the running example. Section (ref) concludes. The Supplementary Materials provide technical details, and the replication files are available on GitHub at \url{https://github.com/xuyiqing/fdid_paper}.
In this section, we present six empirical examples from economics, political science, and finance that align with the FDID research design. In each case, researchers obtain key estimates using a TWFE regression and describe the approach as a DID method. However, their intended estimands often differ.
Squicciarini2020 examines whether Catholicism, proxied by the share of refractory clergy in 1791, hindered economic growth during the Second Industrial Revolution in France. One main analysis relies on a longitudinal dataset of French departments from 1866 to 1911. The baseline factor $G$ is the 1791 share of refractory clergy, and the event is the Second Industrial Revolution in the late 19th century (post-1870), to which all departments were presumably exposed. It is not entirely clear about which estimand the study aims to identify. The author argues that the results pointed to a causal interpretation of the relationship between religiosity and economic development during the Second Industrial Revolution, which corresponds most closely to $G$’s average conditional effect. The study also describes itself as examining “the differential diffusion of technical education and industrial development” (p. 3455), which can be interpreted as targeting the effect modification of $G$.
fouka2019how studies “the effect of taste-based discrimination on the assimilation decisions of immigrant minorities” in the United States (abstract). The repeated cross-sectional data include all men born in the US from 1880 to 1930 to a German-born father, organized by state and birth year. The baseline factor $G$ is state-level measures of anti-Germanism, such as support for Woodrow Wilson in the 1916 presidential Election. The event is World War I, which started in 1917. The outcome is a foreign name index. Although not stated explicitly, the intended estimand is likely either the causal moderation or $G$'s average conditional effect post World War I.
charnysh2022explaining argues that during crises, states allocate fewer resources to less “legible” groups---those from which they cannot gather reliable information or effectively collect taxes. Using district-level panel data with yearly observations from Imperial Russia, the author studies the 1891–1892 Russian famine. The baseline factor $G$ is the district-level share of Muslims, and the event is the famine. The findings show that districts with larger Muslim populations experienced higher mortality rates during the famine. The author avoids explicit causal claims, so the analysis can be interpreted as targeting the effect modification of $G$.
chen2024powerholders argue during a major state-building episode in ancient China, the number of aristocrats from prefectures recruited into the imperial bureaucracy increased more in localities with strong military presence. The study uses prefecture-level panel data from the Northern Wei Dynasty (384–534 CE). The baseline factor $G$ is whether a prefecture had fourth-century military strongholds, and the event is a state-building reform initiated by Empress Dowager Feng (477–490 CE). The authors describe their empirical strategy as a “canonical DD strategy,” which “relies on the parallel-trend assumption to adopt a causal interpretation.” Given the use of explicit causal language, the intended estimand is likely either the causal moderation or the average conditional effect of $G$.
chen2023pledgeability examine how the ability to use corporate bonds as collateral (pledgeability) affects their prices, leveraging a policy change in Chinese bond markets. The study uses panel data on daily bond prices. The baseline factor $G$ is bond ratings, and the event is a policy change on December 8, 2014, when Chinese policymakers prohibited bonds rated below AAA from being used as collateral. The intended estimand is the ATT, the policy effect on AA and AA+ bonds. This setting reduces to canonical DID because the exclusion restriction---i.e., no policy effect on AAA and AA$-$ bonds---is plausible, given that AA$-$ bonds were already ineligible for pledging before the policy change.
Finally, cao2022clans is our running example. Recall from Section (ref) that the baseline factor $G$ is social capital, and the event is China's Great Famine from 1958 to 1961. The authors intend to estimate the causal moderation of social capital on the famine's impact on mortality. In Section (ref), we reanalyze this application using a balanced panel of 921 counties spanning the years 1954 to 1966. Figure (ref) displays the average mortality rates during this period for two types of counties in the sample: those with high social capital and those with low social capital. While the average mortality rate rose sharply during the famine years in both groups of counties, the increase was noticeably higher in counties with low social capital compared with those with high social capital.
We present our main theoretical results in the two-group, two-period panel setting. This section introduces the notation, data structure, and DID estimator, and then defines the potential outcomes that form the basis for the estimands.
Assume a standard two-group, two-period panel setting with a one-time event and a study sample of $n$ units, indexed by $i = 1, \ldots, n$. For each unit $i$, we observe a binary baseline factor $G_i \in \{0,1\}$, an outcome measured at two time points—before and after the event—denoted by $Y_{i, \textup{pre}} \in \mathbb R$ and $Y_{i, \textup{post}} \in \mathbb R$, and an exposure indicator $Z_i \in \{0,1\}$, where $Z_i = 1$ if unit $i$ is exposed when the event occurs and $Z_i = 0$ otherwise. A defining feature of FDID is that all units are exposed to the event, so $Z_i = 1$ for all $i$, as formalized in Assumption (ref) and Definition (ref) below.
In the FDID setting, the exposure indicator $Z_i$ equals one for all units and thus appears redundant. However, it is essential for defining the potential values of $Y_{i, \textup{pre}}$ and $Y_{i, \textup{post}}$ that would have been observed in the absence of the event. These potential outcomes provide the basis for defining the causal estimands and stating the identification assumptions in FDID. A similar use of $Z_i$ appears in holland1986research to clarify Lord's paradox in settings identical to FDID. The departure from canonical DID arises from the absence of a one-to-one mapping between $G$ and $Z$.
Let $\Delta Y_i = Y_{i, \textup{post}} - Y_{i, \textup{pre}}$ denote the before-after difference in the outcomes of unit $i$. The DID estimator is the difference in the average $\Delta Y_i$ between the two groups defined by the baseline factor $G$:
where $n_g$ is the number of units with $G_i = g$. We assume throughout that units are drawn from a common population distribution. Define
as the probability limit of $\hat\tau_\textsc{did}$ as $n \to \infty$, commonly referred to as the DID estimand. Our goal is to clarify the causal interpretation of $\tau_{\textsc{did}}$ in the FDID setting under various identifying assumptions.
We now define the potential outcomes under FDID. Unlike the classic DID framework, we define them with respect to both the baseline factor $G_i$ and the exposure level $Z_i$. For $t \in \{\textup{pre},\textup{post}\}$, let $Y_{it}(g,z)$ denote the potential value of $Y_{it}$ if $G$ were set at $g$ and $Z$ were set at $z$ for unit $i$. Let $1_{\{\cdot\}}$ be the indicator function. The observed outcome satisfies $Y_{it} = \sum_{g, z=0,1} 1_{\{(G_i,Z_i) = (g,z)\}} Y_{it}(g,z)= Y_{it}(G_i, Z_i)$, which reduces to $Y_{it} = Y_{it}(G_i, 1)$ under Assumption (ref). Thus, the four potential outcomes with $z=0$, $\{Y_{it}(g,0): t \in \{\textup{pre},\textup{post}\},\ g \in \{0,1\}\}$, are unobservable for all units in the FDID setting. Figure (ref) illustrates this setup.
We now formalize four estimands that $\tau_{\textsc{did}}$ can identify in the FDID setting under different assumptions. Following the literature on causal inference with factorial experiments td, bansak2020estimating, zdfact, we begin by defining three unit-level effects. Let $\tau_{i, Z \mid G = g} = Y_{i, \textup{post}}(g, 1) - Y_{i, \textup{post}}(g, 0)$ denote the effect of exposure on unit $i$ if the baseline factor $G$ were at level $g$. Let $\tau_{i, G \mid Z = z} = Y_{i, \textup{post}}(1, z) - Y_{i, \textup{post}}(0, z)$ denote the effect of the baseline factor $G$ on unit $i$ if its exposure level were at $z$. Define \\\centerline{$ \tau_{i,\textup{cm}} = \tau_{i, Z \mid G = 1} - \tau_{i, Z \mid G = 0} $} as the {\it causal moderation} of $G$ on the effect of exposure for unit $i$, capturing how the effect of exposure differs between the two levels of $G$. Note that \\\centerline{$ \tau_{i,\textup{cm}} = Y_{i, \textup{post}}(1,1) - Y_{i, \textup{post}}(1,0) - Y_{i, \textup{post}}(0,1) + Y_{i, \textup{post}}(0,0) = \tau_{i, G\mid Z = 1} - \tau_{i, G\mid Z=0}. $} Thus, $\tau_{i,\textup{cm}}$ is symmetric in $G$ and $Z$, and can also be interpreted as the causal moderation of $Z$ on the effect of $G$ for unit $i$. In the literature tyler, $\tau_{i,\textup{cm}}$ is also referred to as the {\it interaction} between $G$ and $Z$.
Definition (ref) below formalizes effect modification and causal moderation as comparisons of the exposure effect, $\tau_{i, Z\mid G=g}$, between the two levels of $G$, extending tyler and bansak2020estimating to the panel setting. To highlight the distinction between these concepts, define $\tau_{i, Z \mid G} $ as the value of $\tau_{i, Z| G =g}$ when the baseline factor $G_i$ is at its observed level, i.e., \\\centerline{$ \tau_{i, Z \mid G} = \left\{
\right. $}
Both $\tau_\textup{em}$ and $\tau_{\textup{cm}}$ capture heterogeneity in the effect of exposure, $\tau_{i, Z \mid G = g}$, across the two levels of the baseline factor $G$, but they emphasize different comparisons. On the one hand, $\tau_{\textup{cm}}$ compares $\tau_{i, Z\mid G=1}$ and $\tau_{i, Z\mid G=0}$ across all units and is conventionally regarded as a causal quantity holland1986research, frangakis2002principal, CausalImbens. It addresses the causal question: Would the effect of exposure change if one intervened on $G$? In our running example, this corresponds to whether the impact of exposure to China's Great Famine would have differed if a locality had randomly developed social capital, perhaps due to the migration of a large kinship clan. By contrast, $\tau_\textup{em}$ compares the average of $\tau_{i, Z\mid G=1}$ among units with $G_i = 1$ to the average of $\tau_{i, Z\mid G=0}$ among units with $G_i = 0$, each evaluated at the observed value of $G_i$. Accordingly, $\tau_\textup{em}$ describes differences in exposure effects between the two groups of units defined by $G$, but does not address the same causal question as $\tau_{\textup{cm}}$. See tyler for further discussion in the cross-sectional setting.
Definition (ref) below formalizes two conditional causal effects of $G$.
From Definition (ref), $\tau_{\textup{att}}$ is analogous to the ATT in canonical DID, with units satisfying $G_i = 1$ viewed as the treated group. The subscript “att” reflects this connection. The estimand $\tau_{G \mid Z = 1}$ is the average causal effect of $G$ conditional on exposure to the event. Figure (ref) summarizes the relationships among these estimands, serving as a roadmap for our identification results.
\FloatBarrier
We now present the identification results. To preview, $\tau_{\textsc{did}}$ identifies $\tau_\textup{em}$ under the canonical DID assumptions; identifying $\tau_{\textup{cm}}$, $\tau_{\textup{att}}$, and $\tau_{G \mid Z = 1}$ requires additional assumptions.
The observed pre-event outcome $Y_{i, \textup{pre}}$ and its potential values $\{Y_{i, \textup{pre}}(g,z): g,z = 0, 1\}$ all occur before the event. A common, often implicit, assumption in the DID literature is that future events do not influence past potential outcomes, known as the no anticipation assumption. We state this assumption explicitly in Assumption (ref) below. It may be violated if units anticipate the event and adjust their behavior in advance.
Recall that $\Delta Y_i = Y_{i, \textup{post}} - Y_{i, \textup{pre}}$ is the before-after difference in outcome for unit $i$. Let $\Delta Y_i (g,z) = Y_{i, \textup{post}}(g,z) - Y_{i, \textup{pre}}(g,z)$ denote the potential value of $\Delta Y_i$ if $G$ were set at $g$ and $Z$ were set at $z$ for unit $i$. Define $\Delta Y_i(G_i,0) = Y_{i, \textup{post}}(G_i,0)-Y_{i, \textup{pre}}(G_i,0)$ as the potential before-after change for unit $i$ under its observed baseline factor $G$ but assuming no exposure. Assumption (ref) below restates the canonical parallel trends assumption using the augmented potential outcomes.
Assumption (ref) states that, in the absence of the event, the average change in outcome over time, $\Delta Y_i(G_i, 0)$, would be the same across the two groups defined by the baseline factor $G$. Together, Assumptions (ref)--(ref) form the canonical identifying assumptions in the DID literature, under which $\tau_{\textsc{did}}$ identifies the ATT in the canonical DID setting. Proposition (ref) below extends this classic result to the FDID setting and shows that, under these assumptions, $\tau_{\textsc{did}}$ identifies $\tau_\textup{em}$. Figure (ref) illustrates this identification result.
Definition (ref) and Proposition (ref) underscore two key differences between FDID and canonical DID. First, FDID assumes that all units are exposed to the event, whereas canonical DID relies on a clean, unexposed control group. Second, under the no anticipation and canonical parallel trends assumptions, the DID estimator identifies the ATT, a causal quantity, in canonical DID, but identifies $\tau_\textup{em}$, a descriptive quantity, in FDID. Despite these differences, we show below that the canonical DID research design can be reframed as a special case of FDID under an additional exclusion restriction assumption that the event has no effect on a group of units defined by $G$.
{Assumption (ref)} ensures that the average post-event outcome of units with $G_i = 0$ is unaffected by exposure. Conceptually, these units are exposed but unaffected, thus resembling the clean control group in canonical DID. Together, Assumption (ref) (universal exposure) and Assumption (ref) (exclusion restriction)\ reproduce a canonical DID setting under FDID, with units satisfying $G_i = 1 $ and $G_i = 0$ serving as the treated and control groups, respectively. This justifies interpreting $\tau_{\textup{att}}$ as the ATT analog in FDID. Moreover, {Assumption (ref)} implies $\mathbb E[\tau_{i, Z \mid G = 0} \mid G_i = 0 ] = 0$, so that $\tau_\textup{em} = \mathbb E[\tau_{i, Z \mid G = 1} \mid G_i = 1 ] - \mathbb E[\tau_{i, Z \mid G = 0} \mid G_i = 0 ] =\tau_{\textup{att}}$. This provides a causal interpretation of $\tau_\textup{em}$, as formalized in Proposition (ref) below.
Recall that $\tau_{\textup{att}}$ is the ATT analog under FDID. Proposition (ref) reframes the classic identification result for the ATT in canonical DID within the FDID framework. Definition (ref) builds on this result and characterizes canonical DID as a special case of FDID. In Figure (ref), this corresponds to the post-period gaps between the solid triangle and hollow square, and between the solid circle and hollow circle, being closed.
Propositions (ref)--(ref) establish the identification of $\tau_\textup{em}$ and $\tau_{\textup{att}}$ under FDID. In many applied studies, however, the primary quantity of interest is the causal moderation $\tau_{\textup{cm}}$ or the causal effect of $G$ given exposure, $\tau_{G \mid Z=1}$; c.f. Section (ref). We discuss their identification below.
Recall that $\Delta Y_i (g,z) = Y_{i, \textup{post}}(g,z) - Y_{i, \textup{pre}}(g,z)$ denotes the potential before-after change in outcome for unit $i$. Assumption (ref) below introduces a {\it factorial parallel trends} assumption, which requires mean independence between $G_i$ and $\Delta Y_i(g,z)$.
We call Assumption (ref) the factorial parallel trends assumption because, like canonical parallel trends (Assumption (ref)), it requires equal average changes in potential outcomes across groups. However, now $\Delta Y_i(g,z)$ varies both $G$ and $Z$, implying mean independence between $G$ and all four potential outcome changes, $\Delta Y_i(g,z)$. By contrast, Assumption (ref)\ requires only mean independence between $G_i$ and $\Delta Y_i(G_i,0)$, holding $G$ fixed at its observed value. A sufficient condition for Assumption (ref) is
which states that $G$ is independent of the before-after changes in all four potential outcomes. As noted in Remark (ref), DID analysis with $(G_i,Y_{i, \textup{pre}},Y_{i, \textup{post}})$ in FDID corresponds to cross-sectional analysis with $(G_i,\Delta Y_i)$. Analogously, condition (ref) resembles the standard random assignment assumption for $(G_i,\Delta Y_i,Z_i)$ in cross-sectional settings. Condition (ref) reflects the belief that differencing the potential outcomes removes confounding with respect to $G$. As noted earlier, some applied researchers do not recognize that Assumption (ref) or condition ((ref)) is required to interpret DID estimates as causal moderation. Others, while not explicitly stating it, appear to have an intuitive understanding of this assumption and view it as more plausible than full random assignment $G_i {\perp\!\!\!\perp} Y_i(g,z)$. \footnote{For example, fouka2019how writes: “I control for the potential time-varying effect of the share of the German population in the state, which is plausibly correlated with both (lower) support for Wilson and assimilation,” and “state-level anti-Germanism is potentially endogenous to pre-existing trends in German assimilation” (p. 419), clearly acknowledging potential correlation between $G_i$ and $\Delta Y_i(g,z)$.} Nevertheless, the assumption, like unconfoundedness, is strong and untestable.
Proposition (ref) below formalizes the conditions under which $\tau_{\textup{cm}}$ and $\tau_\textup{em}$ to coincide, thereby ensuring that $\tau_{\textsc{did}}$ identifies $\tau_{\textup{cm}}$.
Assumption (ref) below introduces an alternative exclusion restriction, which requires that in the absence of the event, $G$ would not have affected the average post-period outcome. It implies $\tau_{\textup{cm}} = \tau_{G \mid Z=1}$ so that $\tau_{\textsc{did}}$ also identifies $\tau_{G \mid Z=1}$, as formalized in Proposition (ref).
Propositions (ref)--(ref) complete the roadmap in Figure (ref). Definition (ref) below formalizes the {\it FDID research design} as the combination of the FDID setting in Definition (ref) with these identification results.
\FloatBarrier
In many applications, the canonical and factorial parallel trends assumptions in Assumptions (ref) and (ref) are plausible only after conditioning on baseline covariates $X_i$ in addition to $G_i$. In this section, we present the corresponding identification and estimation results and connect them to regression methods commonly used in applied research. The unconditional setting is a special case where $X_i = \emptyset$.
The main takeaways are twofold. First, all results in Section (ref) extend to the conditional setting. Under suitable assumptions, the conditional DID estimand identifies the conditional effect modification and causal moderation, and averaging over covariates yields the marginal effects. Second, coherent with Remark (ref), standard cross-sectional methods based on unconfoundedness rosenbaum1983central, such as outcome regression and inverse propensity score weighting, carry over to the FDID setting when $(G_i, \Delta Y_i, X)_{i=1}^n$ are treated as the data. These approaches resemble the covariate-adjustment methods used in canonical DID under the conditional parallel trends assumption roth2023s. We focus here on stratification and outcome regression using linear and TWFE specifications, and provide details on inverse propensity score weighting in Section A2.6 in the Supplementary Materials.
Let $X_i$ denote the vector of covariates beyond $G_i$, taking values in $\mathcal X \subseteq \mathbb R^p$. Assumption (ref) states the overlap condition that ensures the conditional expectation $\mathbb E[\ \cdot \mid G_i = g, X_i]$ is well defined for $g = 0,1$. We maintain this assumption throughout the rest of the paper.
Under Assumption (ref), define
as the {\it conditional} DID estimand, effect modification, and causal moderation, respectively, generalizing $(\tau_{\textsc{did}}, \tau_\textup{em}, \tau_{\textup{cm}})$ in (ref) and Definition (ref). Define their marginal counterparts as
where the expectations are taken over the marginal distribution of $X_i$. Since $\tau_{\textsc{did}}(X_i)$ and $\tau_\textup{em}(X_i)$ compare $\Delta Y_i$ and $\tau_{i, Z \mid G = g}$, respectively, between the two levels of $G_i$ conditional on $X_i$, their marginal averages $(\tau_{\textup{\textsc{did}-x}}, \tau_{\textup{em-x}})$ in (ref) generally differ from $(\tau_{\textsc{did}},\tau_\textup{em})$ unless $G_i$ and $X_i$ are independent. Assumptions (ref)--(ref) below build on Assumption (ref), and extend the canonical and factorial parallel trends assumptions to the conditional setting.
Echoing the discussion below Assumption (ref), a sufficient condition for Assumption (ref) is $G_i {\perp\!\!\!\perp} \{\Delta Y_i(g,z):g, z = 0,1\} \mid X_i$, which states that $G_i$ is as-if randomly assigned with respect to changes in all four potential outcomes, conditional on covariates $X_i$. This condition is analogous to the unconfoundedness assumption for the cross-sectional data $(\Delta Y_i, G_i, Z_i, X_i)$. Like unconfoundedness, this assumption is inherently untestable; researchers can only assess its plausibility through auxiliary checks, such as placebo or sensitivity analyses imbens2024comparing.
Corollary (ref) extends Propositions (ref) and (ref), and establishes the identification of $\tau_{\textup{em-x}}$ and $\tau_{\textup{cm}}$.
Corollary (ref)(i) follows from Proposition (ref) and ensures that $\tau_{\textup{\textsc{did}-x}}$ identifies $\tau_{\textup{em-x}}$ under {Assumption (ref) (conditional canonical parallel trends)}. Corollary (ref)(ii) follows from Proposition (ref) and ensures that $\tau_{\textup{\textsc{did}-x}}$ identifies $\tau_{\textup{em-x}} = \tau_{\textup{cm}}$ under {Assumptions (ref)--(ref) (conditional canonical and factorial parallel trends)}. Moreover, Corollary (ref) shows that $\tau_{\textsc{did}}(x)$ allows us to identify the expected values of $\tau_\textup{em}(X_i)$ and $\tau_{\textup{cm}}(X_i)$ under any distribution of $X_i$, not just the marginal one. In particular, we can recover group-specific expectations, $\mathbb E[\tau_\textup{em}(X_i) \mid G_i = g]$ and $\mathbb E[\tau_{\textup{cm}}(X_i) \mid G_i = g]$, by averaging $\tau_{\textsc{did}}(X_i)$ over the conditional distribution of $X_i$ given $G_i = g$, paralleling the ATT and group causal moderation $\mathbb E[\tau_{i,\textup{cm}} \mid G_i = g]$ in Remark (ref).
To apply Corollary (ref), we need to estimate $\tau_{\textsc{did}}(x)$. From (ref), it suffices to estimate $\mathbb E[\Delta Y_i \mid G_i = g, X_i = x]$. Stratification and outcome regression are two approaches, suited to categorical and continuous $X_i$, respectively.
For categorical $X_i$ with $K$ levels indexed by $k = 1, \dots, K$, we can estimate $\mathbb E[\Delta Y_i \mid G_i=g, X_i=k]$ by the sample average of $\Delta Y_i$ among units with $(G_i,X_i)=(g,k)$, denoted by $\widehat{\Delta Y}(g,k)$. The resulting estimator of $\tau_{\textsc{did}}(k)$ is $ \hat\tau_\textsc{did}(k) = \widehat{\Delta Y}(1, k)-\widehat{\Delta Y}(0, k)$, as the {\it stratum-specific DID estimator} based on units with $X_i = k$. The marginal estimand $\tau_{\textup{\textsc{did}-x}}$ can be estimated following (ref) as $\hat\tau_{\textup{\textsc{did}-x}} = {n}^{-1}\sum_{i=1}^n \hat\tau_\textsc{did}(X_i) = \sum_{k=1}^K \pi_k \hat\tau_\textsc{did}(k)$, where $\pi_k$ is the sample proportion of units with $X_i = k$. This approach also applies when $X_i$ can be meaningfully discretized.
For continuous $X_i$, we can estimate $\mathbb E[\Delta Y_i \mid G_i=g, X_i=x]$ using regression, denoted by $\widehat{\Delta Y}(g, x)$, and obtain $\tau_{\textsc{did}}(x)$ and $\tau_{\textup{\textsc{did}-x}}$ from (ref)--(ref) as
Note that \\\centerline{$ \mathbb E[\Delta Y_i \mid G_i = g, X_i = x] = \mathbb E[Y_{i, \textup{post}} \mid G_i = g, X_i = x] - \mathbb E[Y_{i, \textup{pre}} \mid G_i = g, X_i = x], $} where the two sides correspond to two common regression-based approaches to DID analysis. The left-hand side corresponds to cross-sectional linear regression of $\Delta Y_i$ on $(G_i, X_i)$, which directly estimates $\mathbb E[\Delta Y_i \mid G_i = g, X_i = x]$. The right-hand side corresponds to TWFE regression of $Y_{it}$ on $(G_i, X_i)$ with unit and time fixed effects using {\it long-format} data, where each row corresponds to a unit-time observation. We discuss both approaches below.
Definition (ref) below presents two common specifications for estimating $\mathbb E[\Delta Y_i \mid G_i=g, X_i=x]$ using cross-sectional regression of $\Delta Y_i$.
OLS$_+$\ is a restricted version of OLS$_*$\ excluding the interactions between $G_i$ and $X_i$. Define\\ \centerline{ $\widehat{\Delta Y_*}(g,x) = \hat\beta_{1,*} + \hat\beta_{G,*} g + \hat\beta_{X,*}^\top x + \hat\beta_{GX,*}^\top gx, \quad \widehat{\Delta Y_+}(g,x) = \hat\beta_{1,+} + \hat\beta_{G,+} g + \hat\beta_{X,+}^{\top} x$} as the estimators of $\mathbb E[\Delta Y_i\mid G_i=g, X_i=x]$ under {OLS$_*$} and {OLS$_+$}, respectively. Following (ref), the DID estimators based on {OLS$_*$} and {OLS$_+$} are:
and
respectively, where $\bar X = n^{-1}\sum_{i=1}^n X_i$ is the sample mean of $X_i$. Proposition (ref) below establishes the consistency of these estimators for $\tau_{\textsc{did}}(x)$ and $\tau_{\textup{\textsc{did}-x}}$ when the corresponding models in Definition (ref) are correctly specified.
Proposition (ref) justifies the use of {OLS$_*$} and OLS$_+$\ for estimating $\tau_{\textsc{did}}(x)$ and $\tau_{\textup{\textsc{did}-x}}$ under linearity assumptions on $\mathbb E[\Delta Y_i\mid G_i, X_i]$. The identification of $\{\tau_\textup{em}(x), \tau_{\textup{em-x}}\}$ under Assumption (ref), and of $\{\tau_{\textup{cm}}(x), \tau_{\textup{cm}}\}$ under Assumptions (ref)--(ref), then follows from Corollary (ref). Together, Proposition (ref) and Corollary (ref) justify the use of OLS$_+$\ and OLS$_*$\ for DID analysis of FDID when linearity holds. In particular, Proposition (ref)(i) implies that if $\mathbb E[X_i] = 0$, then $\beta_G = \tau_{\textup{\textsc{did}-x}}$, so the coefficient of $G_i$ from {OLS$_*$}, $\hat\beta_{G,*}$, has a direct causal interpretation as a consistent estimator of $\tau_{\textup{em-x}}$ and $\tau_{\textup{cm}}$ under the corresponding assumptions. The sample version of this condition can be ensured by centering the covariates so that $\bar{X}=0$, as in hirano2001estimation and lin2013agnostic. To account for the uncertainty in $\hat\beta_{G,*}$, one can implement a unit-level cluster bootstrap procedure bertrand2004much. A subtlety is that the bootstrap samples must be generated using the original $X_i$'s, which are then recentered within each bootstrap replication. Using pre-centered covariates for resampling yields invalid inference because it ignores the sampling variability in $\bar{X}$.
By contrast, Proposition (ref)(ii) shows that the more parsimonious OLS$_+$\ may be inconsistent for $\{\tau_{\textsc{did}}(x), \tau_{\textup{\textsc{did}-x}}\}$ if the effect of $G_i$ on $\Delta Y_i$ varies with $X_i$. However, when $\mathbb E[\Delta Y_i \mid G_i, X_i]$ is truly linear in $(G_i, X_i)$, then $\tau_{\textsc{did}}(x)=\tau_{\textup{\textsc{did}-x}}$, and the coefficient $\hat\beta_{G,+}$ from {OLS$_+$} is consistent for their common value without requiring covariate centering.
Definition (ref) presents two common TWFE specifications for estimating $\mathbb E[Y_{it} \mid G_i=g, X_i=x]$, where $t = \textup{pre}, \textup{post}$. Let $1_{\{t = \textup{post}\}}$ denote an indicator for the post-period.
TWFE$_+$\ is a restricted version of {TWFE$_*$} that excludes the interactions among $G_i$, $X_i$, and $1_{\{t = \textup{post}\}}$. Standard OLS theory ensures that\\
\\ Thus, causal interpretation of {TWFE$_*$} and TWFE$_+$\ follows directly from Proposition (ref). As in OLS$_*$, when using TWFE$_*$\ for FDID analysis, it is important to center the covariates so that the coefficient on $G_i \cdot 1_{\{t = \textup{post}\}}$ has a standalone causal interpretation.
In this section, we generalize our main theoretical results to accommodate repeated cross-sectional data and to allow for a general (such as discrete or continuous) baseline factor $G$.
So far, our discussion has focused on analyses based on panel data. In many applications, however, researchers only have access to repeated cross-sectional data, where a new set of units is sampled at each time point. We extend our framework to this setting, and unify the panel and repeated cross-sections as two sampling schemes for the same population.
Previously, $i$ indexed observed units drawn from a common population. With slight abuse of notation, we now also let $i$ denote a generic random unit from this population, which may or may not belong to the study sample. Our theory ensures that the conditional DID estimand $\tau_{\textsc{did}}(x)$ identifies $\tau_\textup{em}(x)$ and $\tau_{\textup{cm}}(x)$ under suitable assumptions. Panel and repeated cross-sectional data then correspond to two sampling schemes: in panel data, the same units are re-observed over time, while in repeated cross-sectional data, new units are sampled at each time. Recall from (ref) that
where $\mathbb E[\Delta Y_i \mid G_i = g, X_i = x] = \mathbb E[Y_{i, \textup{post}} \mid G_i = g, X_i = x] - \mathbb E[Y_{i, \textup{pre}} \mid G_i = g, X_i = x]$. With panel data, $\Delta Y_i$ is observed directly, so both $\mathbb E[\Delta Y_i \mid G_i = g, X_i = x]$ and $\mathbb E[Y_{it} \mid G_i = g, X_i = x]$ can be estimated. With repeated cross-sectional data, $\Delta Y_i$ is not observed at the unit level, so estimation proceeds via $\mathbb E[Y_{it} \mid G_i = g, X_i = x]$. In either case, stratification and TWFE regression, as discussed in Section (ref), are feasible, suitable for discrete $X_i$ and continuous $X_i$, respectively. However, outcome regression based on the OLS specifications using $\Delta Y_i$ is only applicable in the panel setting.
A distinctive feature of many FDID applications fouka2019how using repeated cross-sectional data is nested sampling: the baseline factor $G$ is defined at the regional level, while outcomes are measured at the individual level. For clarity, we refer to the higher-level units as regions and the lower-level (sub)units as individuals. In practice, “individual” may also denote other sub-regional entities such as firms, schools, households, or smaller administrative units. Three sampling designs are common in multi-period studies under nested sampling:
When individual-level data are available, researchers can conduct analysis to estimate $\tau_{\textsc{did}}(x)$ either at the individual level (using individual-level data) or at the regional level (by aggregating individual level data to the regional level). When individual-level data are unavailable, methods for region-level analyses also apply.
In summary, FDID extends naturally to repeated cross-sectional data, with panel and repeated cross-sections viewed as alternative sampling schemes from the same population. In FDID with nested data structures, across the three common nested sampling schemes, stratification and TWFE regression remain applicable for both individual- and region-level analyses, with standard errors clustered by region.
The discussion so far has assumed a binary baseline factor $G$. When $G$ takes values in a general set $\mathcal G \subseteq \mathbb{R}$, all results from the binary case continue to hold for comparisons between any two levels $g, g' \in \mathcal G$.
Renew $Y_{it}(g,z)$, $\Delta Y_i(g,z)$, and $\tau_{i, Z \mid G = g} = Y_{i, \textup{post}}(g,1) - Y_{i, \textup{post}}(g,0)$ for general $g\in\mathcal G$. For $g, g'\in\mathcal G$, define $\tau_{i,\textup{cm}, g \to g'} = \tau_{i, Z\mid G = g'} - \tau_{i, Z \mid G = g}$ as the causal moderation of $G$ when its level changes from $g$ to $g'$, generalizing $\tau_{i,\textup{cm}}$. Define \\\centerline{$
$} as the conditional DID estimand, effect modification, and causal moderation when $G$ changes from $g$ to $g'$, extending $\{\tau_{did}(x), \tau_em(x), \tau_{cm}(x)\}$ in \eqref{eq:tau_X}. The marginal counterparts are $\tau_{{did-x},g \to g'} = \mathbb E[\tau_{{did}, g \to g'}(X_i)]$, $\tau_{{\textup{em-x}},g \to g'} =\mathbb E[\tau_{\textup{em},g \to g'}(X_i) ]$, and $\tau_{\textup{cm}, g \to g'} = \mathbb E[\tau_{\textup{cm}, g \to g'}(X_i)] = \mathbb E[\tau_{i,\textup{cm}, g \to g'}]$, extending $\tau_{\textup{cm}}, \tau_{\textup{\textsc{did}-x}},\tau_{\textup{em-x}}$ in Definitions (ref) and (ref).
Identification. All identification results in Section (ref) extend to general $G$. Parallel to Corollary (ref), the conditional DID estimand $\tau_{{\textsc{did}}, g \to g'}(x)$ (i) identifies $\tau_{\textup{em},g \to g'}(x)$ under the generalized conditional canonical parallel trends\ assumptions, and (ii) identifies $\tau_{\textup{cm}, g \to g'}(x) = \tau_{\textup{em},g \to g'}(x)$ under the generalized conditional canonical and factorial parallel trends\ assumptions. We provide the details in Section A2.5 in the Supplementary Materials.
Regression estimation. Renew (OLS$_*$, OLS$_+$) and (TWFE$_*$, TWFE$_+$) in Definitions (ref)--(ref) with general $G_i$. The numerical equivalence between OLS and TWFE continues to hold for general $G_i$. Parallel to Proposition (ref), OLS$_*$\ and OLS$_+$\ are consistent for estimating $\tau_{{\textsc{did}}, g \to g'}(x)$ and $\tau_{{\textup{\textsc{did}-x}},g \to g'}$ when the linear models are {correctly specified}.
With continuous $G$, we can also study incremental effects in the sense of rothenhausler2019incremental. For brevity, we relegate the details to Section A2.5 of the Supplementary Materials.
We now reanalyze the data from our running example cao2022clans to illustrate our theory. Each observation represents one county in the sample. The outcome of interest is the county-level annual mortality rate, measured as the number of deaths per thousand people. Following cao2022clans, the baseline factor takes two forms: (i) a binary indicator for high social capital, equal to one if the per capita number of genealogy books is at or above the sample median and zero otherwise; and (ii) the logarithm of the per capita number of genealogy books plus one, a continuous measure. The Great Famine began in late 1958 and, according to most historians, ended in 1961. The set of pre-famine covariates, measured at the county level, include per capita grain production, ratio of non-farming land, urbanization ratio, distance from Beijing, distance from the provincial capital, share of ethnic minorities, suitability for rice cultivation, average years of education, and log population size.
We apply three estimators, DID, {OLS$_*$} (with interactions), and {OLS$_+$} (without interactions), with the latter two incorporate covariates. As discussed earlier, in the FDID panel setting, {OLS$_*$} and {OLS$_+$} are numerically equivalent to {TWFE$_*$} and {TWFE$_+$}, respectively. We use 1957, one year before the famine, as the reference year for the pre-period, with $Y_{i, \textup{pre}}$ representing the mortality rate in county $i$ in 1957. For the post-period, we define three time windows: (i) the famine years from 1958 to 1961, (ii) the pre-famine years from 1954 to 1957, and (iii) the post-famine years from 1962 to 1966. The second time window, 1954–1957, precedes the famine and is used to construct a placebo test for the (conditional) canonical parallel trends assumption, similar to a pretrend test in canonical DID angrist2009mostly. Within each time window, we calculate the average mortality rate for each county, which serves as $Y_{i, \textup{post}}$.
cao2022clans aim to estimate “the effect of social capital on famine relief,” which we interpret as the causal moderation of social capital on the effect of exposure to famine. By Proposition (ref), if Assumptions (ref)--(ref)\ (universal exposure, no anticipation, and canonical parallel trends) hold, then $\hat\tau_\textsc{did}$ recovers effect modification, $\tau_\textup{em}$. If Assumption (ref) (factorial parallel trends) also holds, then $\hat\tau_\textsc{did}$ recovers causal moderation, $\tau_{\textup{cm}}$. Similarly, with covariates, Corollary (ref) and Proposition (ref) imply that under the conditional canonical parallel trends (Assumption (ref)), $\hat\beta_{G,*}$ and $\hat\beta_{G,+}$ recover effect modification, $\tau_{\textup{em-x}}$, under their respective model specifications. If, in addition, the conditional factorial parallel trends (Assumption (ref)) holds, then they identify causal moderation, $\tau_{\textup{cm}}$.
Table (ref) presents the results, where Panels A and B report estimates using binary and continuous measures of $G$, respectively. In each panel, Column (1) reports DID estimates without covariates, while columns (2) and (3) report estimates from {OLS$_*$} and OLS$_+$, respectively, with covariates included. Each row corresponds to a different comparison period relative to the reference year 1957. In Panel A, the first row shows that high social capital (relative to low social capital) are associated with a reduction the famine-induced increase in mortality by more than 2.9 deaths per 1,000 people per year, or nearly 12 fewer deaths per 1,000 people over the four-year period. The three estimators produce similar results and closely match those reported in the original paper using TWFE models. The estimates using the average mortality rate from 1954--1956 as $Y_{i, \textup{post}}$ (second row) are close to zero, which support the canonical parallel trends assumption. The estimates using the average mortality rate from 1962--1966 as $Y_{i, \textup{post}}$ (third row) are negative—though much smaller in magnitude than during the famine years—and statistically significant at the 5% level, suggesting a small but lasting effect of the famine on mortality rates after it ended.
To gain a better understanding on how the estimates evolves over time, we estimate it separately for each year from 1954 to 1966, excluding 1957 (the pre-period), defining $Y_{i, \textup{post}}$ as the mortality rate for each year. Figure (ref) shows the results. Coherent with findings in Panel A of Table (ref), the estimates are close to zero in the pre-famine years, negative during the famine (with particularly large negative estimates in 1959 and 1960), and remain negative but with much smaller magnitudes after the famine ended. The pre-1957 estimates are close to zero, providing suggestive evidence on the no-anticipation and (conditional) canonical parallel trends assumptions, under which the first-row estimates are effect modification. Interpreting the results as causal moderation requires the stronger (conditional) factorial parallel trends assumption, which rules out any potential confounder correlated with both social capital and the dynamics of mortality. For example, if income or governance quality are correlated with both social capital and trends in mortality, this assumption would be violated. In Section A3 in the Supplementary Materials, we present sensitivity analysis showing that a confounder with correlations to social capital and mortality changes comparable to observed covariates would have only a small impact on the estimated moderation effect. Nevertheless, we emphasize that this causal interpretation rests on a strong and inherently untestable assumption.
The patterns in Panel B mirror those in Panel A: social capital is negatively correlated with famine-induced mortality increases, but not in years leading up to the famine. For example, with $\hat\beta_{G,*}=5.15$, if linearity and constant effect assumptions hold, then a 10% increase in the per capita number of genealogy books corresponds to a reduction of 0.5 deaths per 1,000 people per year, or about 2 fewer deaths per 1,000 people over the four-year period. Overall, the evidence supports the original authors’ claim that social capital is associated with a mitigated impact of the famine. Whether this association can be interpreted as causal, however, depends on how plausible the (conditional) factorial parallel trends assumption is.
\FloatBarrier
Based on our results, we offer several recommendations for applied researchers. First, it is important to distinguish between an estimator and a research design. An estimator is merely an algorithm applied to observed data, while a research design specifies not only the estimator but also identifying assumptions that connect it to meaningful estimands. Researchers should clearly communicate each element of the research design, especially the target estimands lundberg2021your. In particular, the use of the DID estimator should not be conflated with the canonical DID research design.
Second, researchers should be cautious when interpreting DID estimates as causal in FDID. Unless the exclusion restriction that exposure $Z$ has no effect on one of the two groups is plausible, the DID estimator under the no anticipation and canonical parallel trends assumptions identifies only effect modification of $G$, not the causal effects of $Z$ or $G$. To recover causal moderation by $G$, we propose a factorial parallel trends assumption, which requires mean independence between $G_i$ and $\Delta Y_i(g,z)$ given $X_i$. Intuitively, it eliminates any unobserved confounders that are correlated with both $G$ and the outcome dynamics. Like unconfoundedness in the cross-sectional setting, this is a strong and untestable assumption. To identify $G$’s conditional effect given exposure, a quantity many studies aim to estimate, an additional exclusion restriction is required, stating that $G$ has no effect on the outcome when $Z = 0$. Table (ref) summarizes these results and illustrates the interpretation of each estimand using the running example.
Third, for estimation, in the absence of covariates, a TWFE regression including the interaction between $G$ and $1_{\{t = \textup{post}\}}$ is numerically equivalent to the DID estimator. With additional covariates $X_i$, standard cross-sectional methods based on unconfoundedness apply to $(\Delta Y_i,G_i,X_i)$, where $\Delta Y_i$ is the before-after difference of the outcome. Outcome regression, such as linear regression of $\Delta Y_i$ on $(1,G_i,X_i,G_iX_i)$, is commonly used when $X_i$ is continuous and is numerically equivalent to a TWFE regression with the interaction $G_iX_i \cdot 1_{\{t = \textup{post}\}}$. In both approaches, it is important to center $X_i$ to ensure that key regression coefficients admit a standalone causal interpretation. We discuss alternative estimation strategies in Section A2.6 of the Supplementary Materials and provide software support for their implementation through the fdid package in R.
The authors claim no conflicts of interest.
We thank Anran Liu and Rivka Lipkovitz for excellent research assistance. Peng Ding acknowledges support from the U.S. National Science Foundation (grants \# 1945136 and \# 2514234). We are grateful to Justin Grimmer, Erin Hartman, Jens Hainmueller, Laura Hatfield, Dennis Shen, Eric Tchetgen Tchetgen, Ye Wang, and seminar participants at Berkeley, MIT, Princeton, UW–Madison, UCSD, UCLA, Yale, PolMeth 2024, ACIC 2025, and the Stanford Online Causal Inference Seminar, as well as three anonymous reviewers and the Editor Hongtu Zhu, for their valuable comments. We also thank Jairui Cao and Chuanchuan Zhang for sharing the data used in cao2022clans.
{ \spacingset{1.4} }
{ \iftrue