EconBase
← Back to paper

A Short Note on Event-Study Synthetic Difference-in-Differences Estimators

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

8,217 characters · 0 sections · 11 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

A Short Note on Event-Study Synthetic Difference-in-Differences Estimators

\abstract{I propose an event study extension of Synthetic Difference-in-Differences (SDID) estimators. I show that, in simple and staggered adoption designs, estimators from arkhangelsky2021synthetic can be disaggregated into dynamic treatment effect estimators, comparing the lagged outcome differentials of treated and synthetic controls to their pre-treatment average. Estimators presented in this note can be computed using the sdid_event Stata package\footnote{The package is hosted on \href{https://github.com/DiegoCiccia/sdid/tree/main/sdid_event}{Github}.}.} \paragraph{Adapting SDID to event study analysis.}

In what follows, I use the notation from clarke2023synthetic to present the estimation procedure for event-study Synthetic Difference-in-Differences (SDID) estimators. In a setting with $N$ units observed over $T$ periods, $N_{tr} < N$ units receive treatment $D$ starting from period $a$, where $1 < a \leq T$. The treatment $D$ is binary, i.e. $D \in \lbrace 0,1\rbrace$, and it affects some outcome of interest $Y$. The outcome and the treatment are observed for all $(i,t)$ cells, meaning that the data has a balanced panel structure.

Henceforth, values of $a$ are referred to as cohorts or adoption periods. The values of $a$ are collected in $A$, i.e. the adoption date vector. For the sake of generality, we assume that $|A| > 1$, meaning that groups start receiving the treatment at different periods. The case with no differential timing can be simply retrieved by considering one cohort at a time. Time periods are indexed by $t \in \lbrace 1, ..., T\rbrace$, while units are indexed by $i \in \lbrace 1, ... N\rbrace$. Without loss of generality, the first $N_{co} = N - N_{tr}$ units are the never-treated group. As for the treated units, let $I^a$ be the subset of $\lbrace N_{co} +1, ..., N\rbrace$ containing the indices of units in cohort $a$. Lastly, we denote with $N^{a}_{tr}$ and $T^{a}_{tr}$ the number of units in cohort $a$ and the number of periods from the the onset of the treatment in the same cohort to end of the panel, respectively. These two cohort-specific quantities can be aggregated into $T_{post}$ from clarke2023synthetic, i.e. the total number of post treatment periods of all the units in every cohort. Namely,

equation[equation omitted — 62 chars of source]

is the sum of the products of $N^{a}_{tr}$ and $T^{a}_{tr}$ across all $a \in A$.

\paragraph{Disaggregating $\hat{\tau}^{sdid}_a$.}

The cohort-specific SDID estimator from arkhangelsky2021synthetic can be rearranged as follows:

equation[equation omitted — 315 chars of source]

where $\lambda_t$ and $\omega_i$ are the optimal weights chosen to best approximate the pre-treatment outcome evolution of treated and (synthetic) control units. $\tau^{sdid}_a$ compares the average outcome difference of treated in cohort $a$ and never-treated before and after the onset of the treatment. In doing so, $\tau^{sdid}_a$ encompasses all the post-treatment periods. As a result, it is possible to estimate the treatment effect $\ell$ periods after the adoption of the treatment, with $\ell \in \lbrace 1,..., T^a_{post} \rbrace$, via a simple disaggregation of $\tau^{sdid}_a$ into the following event-study estimators:

equation[equation omitted — 289 chars of source]

This estimator is very similar to those proposed by borusyak2024revisiting, liu2024practical and gardner2022two, when the design is a canonical DiD de2023difference. The only difference lies in the fact that the outcomes are weighted via unit-time specific weights. Notice that by construction

equation[equation omitted — 123 chars of source]

that is, $\hat{\tau}^{sdid}_a$ is the sample average of the cohort-specific dynamic estimators $\hat{\tau}^{sdid}_{a, \ell}$.

\paragraph{Aggregating $\hat{\tau}^{sdid}_{a,\ell}$ estimators into event-study estimates.}

Let $A_{\ell}$ be the subset of cohorts in $A$ such that $a - 1 + \ell \leq T$, i.e. such that their $\ell$-th dynamic effect can be computed, and let

equation[equation omitted — 60 chars of source]

denote the number of units in cohorts where the $\ell$-th dynamic effect can be estimated. We can use this notation to aggregate the cohort-specific dynamic effects into a single estimator. Let

equation[equation omitted — 120 chars of source]

denote the weighted sum of the cohort-specific treatment effects $\ell$ periods after the onset of the treatment, with weights corresponding to the relative number of groups participating into each cohort. This estimator aggregates the cohort-specific treatment effects, upweighting more representative cohorts in terms of units included.

As in Equation (ref), $\hat{\tau}^{sdid}_\ell$ can also be retrieved via disaggregation of another estimator from clarke2023synthetic. Let $T_{tr} = \max_{a \in A} T^{a}_{tr}$ be the maximum number of post-treatment periods across all cohorts. Equivalently, $T_{tr}$ can also be defined as the number of post-treatment periods of the earliest treated cohort. Then, one can show that

equation[equation omitted — 125 chars of source]

that is, the $\widehat{ATT}$ estimator from clarke2023synthetic is a weighted average of the event study estimators $\hat{\tau}^{sdid}_\ell$, with weights proportional to the number of units for which the $\ell$-th effect can be computed.

\paragraph{Mapping with estimates from sdid_event.} Estimators presented in this note can be computed using the sdid_event Stata package. The baseline output table of sdid_event reports the estimates of the ATT from sdid and $\hat{\tau}^{sdid}_\ell$ for $\ell \in \lbrace 1, ..., T_{tr} \rbrace$. If the command is run with the disag option, the output also includes a table with the cohort-specific treatment effects. Namely, the program returns the estimates of $\hat{\tau}^{sdid}_a$ and $\hat{\tau}^{sdid}_{a,\ell}$, whereas the former can also be retrieved from the e(tau) matrix in sdid.

\subparagraph{Proof of Equation (ref).} Let $T^{a}_{post}$ denote the total number of post-treatment periods across all units in cohort $a$.

\[

array[array omitted — 496 chars of source]

\] where the first equality comes from the definition of $\widehat{ATT}$ in clarke2023synthetic, the second equality from the definition of $T^{a}_{post}$, the third equality from Equation (ref), the fourth equality from the fact that the sets $\lbrace (a, \ell): a \in A, 1 \leq \ell \leq T^{a}_{tr} \rbrace$ and $\lbrace (a, \ell): 1 \leq \ell \leq T_{tr}, a \in A_{\ell}\rbrace$ are equal and the fifth equality from the definition of $\hat{\tau}^{sdid}_\ell$.