Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.
Difference-in-differences with as few as two cross-sectional units -- A new perspective to the democracy–growth debate
\abstract{ Pooled panel analyses often mask heterogeneity in unit-specific treatment effects. This challenge, for example, crops up in studies of the impact of democracy on economic growth, where findings vary substantially due to differences in country composition. To address this challenge, this paper introduces the Temporal Difference-in-Differences (T-DiD) estimator that leverages temporal variation in the data to estimate unit-specific average treatment effects on the treated (ATT) with as few as two cross-sectional units. Under asymptotic parallel trends, limited anticipation, and temporal dependence conditions, the proposed DiD estimator is shown to be asymptotically normal. Provided at least two control units are available, the method is further complemented with an identification test that, unlike pre-trends tests, is more powerful and can detect violations of parallel trends in post-treatment periods. Empirical results using the DiD estimator suggest Benin's economy would have been 6.4% smaller on average over the 1993-2018 period had she not democratised.}
JEL Codes: C21, C22, P16
Keywords: Treatment effects, asymptotic parallel trends, near-epoch dependence, over-identifying restrictions, identification test
\linespread{1.5}
refsection\section{Introduction}
Pooled regression analyses are commonly employed to assess the impact of interventions. However, when unit-specific effects exhibit substantial heterogeneity, estimates from pooled regressions become uninformative. Moreover, pooled analyses are often impractical when only a few cross-sectional units are available. Such analyses also heighten the risk of inadvertently violating identification conditions—such as the parallel trends assumption in the standard Difference-in-Differences (DiD) framework—by failing to carefully select appropriate control units. In certain empirical settings, including the one considered in this paper, there may be as few as a single ideal control unit. Under this two-unit baseline setting, existing methods prove inadequate. The Synthetic Control (SC) approach—the most closely related method—requires multiple candidate control units. Furthermore, unit-level placebo inference, a workhorse for SC-based inference and robustness checks, becomes infeasible with only one control unit or suffers from severe imprecision when the number of controls is small. This paper proposes a Difference-in-Differences estimator (hereafter, T-DiD) that exploits temporal variation in the data to estimate unit-specific average treatment effects on the treated (ATT) with as few as two cross-sectional units. The novelty of this paper lies not in the Difference-in-Differences estimator per se, but in the analysis of its theoretical properties—specifically, asymptotic identification and asymptotic normality—together with an identification test in a fixed-$N$, large-$T$ setting.
The DiD quantity, with as few as two units (one treated and one control) and two time periods (one pre-treatment and one post-treatment), can at best be unbiased but not consistent, since there are only two units. However, by leveraging temporal variation in the data, i.e., by taking the (weighted) average across all possible pre- and post-treatment pairs, one obtains the proposed T-DiD. This paper shows that identification holds under asymptotic parallel trends and limited anticipation conditions. The T-DiD framework allows temporal dependence in the outcome series and arbitrary cross-sectional dependence between the treated and control units without any distributional assumptions on the outcome variables themselves. Identification can hold up to an asymptotic bias, which does not interfere with inference under asymptotic parallel trends and limited anticipation assumptions. Under near-epoch dependence (NED) and standard regularity conditions, the T-DiD is asymptotically normal.\footnote{Informally, NED structures allow dependence between time series observations that decays with the time interval -- (ref) provides a formal definition.} Further, this paper proposes a valid over-identifying restrictions test of identification, which has desirable properties such as detecting a vast array of violations of identification, unlike pre-trends tests, when at least two candidate controls are available. For instance, the proposed test detects violations of identification in the post-treatment period, of which pre-tests are completely incapable. Although inference using the T-DiD is robust to cross-sectional dependence, such dependence is not sine qua non, unlike for factor and SC models; identification in the T-DiD framework does not require correlations among cross-sectional units driven by, e.g., common factors -- see hsiao-ching-wan-2012,ferman-pinto-2021-synthetic. While DiD estimators in general are inconsistent under non-trivial multiplicative factors, the SC generally fails without them. Existing robust methods---such as those by xu-2017,arkhangelsky-etal-2021,callaway-karami-2023treatment---require large $N$, a constraint our small-$N$ T-DiD approach avoids. The T-DiD does not involve regressing the outcome of a unit on that of other units; it thus avoids concerns with spurious regressions or non-standard inference in the presence of co-integration -- see, e.g., masini2022counterfactual,Li-2020-statistical for discussions.
Thanks to a linear regression-based formulation of the DiD estimator, the T-DiD accommodates observed time-varying heterogeneity through the inclusion of covariates and handles complications often encountered in time series analyses, e.g., auto-regressive processes, unit root processes, moving average processes, potentially unbounded common trends, and idiosyncratic deterministic time trends in untreated potential outcomes. However, the proposed T-DiD requires several pre- and post-treatment periods, unlike the SC, which only requires at least one post-treatment period in typical cases, and the C-DID, which needs as few as two periods. Moreover, exploring cross-sectional heterogeneity by estimating unit-specific effects using the T-DiD involves the trade-off of collapsing temporal heterogeneity through (weighted) averages, in contrast to approaches like the C-DiD and SC, which are better suited for analysing the temporal heterogeneity of treatment effects. In this sense, the T-DiD reverses the roles of time and units relative to the C-DiD: just as an extra pre-treatment period provides a testable implication (pre-test) in the C-DiD framework, an extra control unit provides a testable implication (the proposed test) in the T-DiD framework. Therefore, the T-DiD is correctly viewed as complementary to the SC and C-DiD.
The paper’s empirical application revisits the long-debated question of whether democracy fosters economic growth, a timely issue given current global democratic backsliding Coppedge2022, Foa2020. Previous studies yield conflicting evidence --- positive effects, e.g., Acemoglu2019, negative effects, e.g., Gerring2005, or insignificant effects, e.g., Murtin2014 --- due to modelling differences, data limitations, and heterogeneity in regime performance, especially among autocracies. To illustrate the T-DiD framework, the paper focuses on Benin, which has undergone significant democratisation since the 1990s, and considers neighbouring countries in West Africa as potential controls. While Burkina Faso and Niger have partially democratic histories and Nigeria’s economy diverges markedly from Benin’s, Togo emerges as the only feasible control in West Africa due to its shared colonial legacy, cultural and institutional similarities, and membership in the same monetary union.\footnote{Further discussion is available in (ref).} The empirical results show that Benin's economy would have been 6.4% smaller on average over the 1993-2018 post-treatment window had she not democratised. This effect is both statistically and economically significant.
The T-DiD estimator is not limited to the empirical example of interest in this paper; it is applicable in contexts with as few as a single treated unit and a single control unit, provided there are many pre- and post-treatment periods. Such settings are common when studying heterogeneous effects, like unit-specific effects. Relevant empirical examples include the North and South Korea experiment within the ensuing democracy-growth debate, the effect of government policy on company share prices masini2022counterfactual, the impact of anti-tax evasion laws on macroeconomic indicators carvalho-Masini-Ricardo-2018arco, Hong Kong's gross domestic product (GDP) growth after sovereignty reversion hsiao-ching-wan-2012, the economic impact of International Monetary Fund (IMF) bailouts lee-shin-2008imf, the impact of opening physical showrooms on the sales of a firm Li-2020-statistical, and the transition from federal to state management of the Clean Water Act marcus-santanna-2021role,grooms-2015-enforcing. See, e.g., carvalho-Masini-Ricardo-2018arco for additional examples.
For the remainder of the paper, (ref) defines the parameter of interest and outlines sufficient conditions for its asymptotic identification. (ref) introduces the T-DiD estimator, while (ref) develops the asymptotic theory. (ref) proposes a T-DiD-based over-identifying restrictions test of identification, and (ref) applies the method to estimate the effect of democracy on Benin's economic growth. Finally, (ref) concludes the paper. The supplementary material contains proofs of technical results, model extensions, supplementary discussions, and a simulation exercise.
\paragraph{Notation:} $ Y_t(1)$ and $Y_t(0)$ denote the treated and untreated potential outcome at period $t$, respectively, $Y_{d,t}$ denotes the outcome $Y_t$ for a unit with treatment status $D=d$. These outcomes are defined on a complete probability space $(\Omega, \mathcal{F}, P)$ and behave as stochastic processes that are Near-Epoch Dependent (NED) on a strictly stationary $\alpha$-mixing underlying process; all conditional expectations are understood to be defined almost surely. $||W||_p:= (\mathbb{E}[|W|^p])^{1/p}, \ p\geq 1 $ denotes the $L_p$-norm of the random variable $W$, and $\mathbb{V}[W]$ denotes its variance. $a_{1,n} \lesssim a_{2,n}$ means $a_{1,n} \leq c\,a_{2,n}$ for some finite $ c>0 $, and $a_{1,n} \asymp a_{2,n}$ means both $a_{1,n}\lesssim a_{2,n}$ and $a_{2,n} \lesssim a_{1,n}$ where $\{a_{1,n}:n\geq 1\}$ and $\{a_{2,n}:n\geq 1\}$ are sequences of non-negative numbers. $a\wedge b := \min\{a,b\} $, $a\vee b := \max\{a,b\} $, and $\rho_{\mathrm{min}}(\Sigma)$ denotes the minimum eigenvalue of the positive semi-definite matrix $\Sigma$. $[T]:= \{1,\ldots, T \} $, while $[-\mathcal{T}]:= \{-\mathcal{T},\ldots,-1\} $. Finally, let $\epsilon > 0$ denote a small constant, the value of which may vary across occurrences without loss of generality.
\section{The Two-unit Baseline Setting}
In the baseline case considered in this paper, the researcher has one treated unit and one control unit with large $\mathcal{T}$ periods pre-treatment and large $T$ periods post-treatment. Pre-treatment periods are labelled as $\tau \in [-\mathcal{T}]$, and post-treatment periods as $t \in [T]$. The period $t=0$ is used to denote any period in the intervening transition window between pre-treatment and post-treatment.\footnote{\url{https://freedomhouse.org}, for example, reports 1990-1991 as the transition period to democracy in Benin.} The binary variable $D$ denotes the treatment status of a unit.
\subsection{Parameter of interest}
The goal of this sub-section is to first define (in the population), the parameter of interest before establishing its asymptotic identification. The ATT at period $t$ is given by
\[
ATT(t) := \mathbb{E}[Y_t(1)-Y_t(0)\mid D=1], \ t \in [T]
\]
where potential outcomes $ \big\{ \big(Y_t(1),Y_t(0)\big), \; t \in [T] \big\} $ are random for a single treated unit. A fundamental identification problem is that $Y_t(1)$ and $Y_t(0)$ are not observed simultaneously for each $t\in [T]$. For example, unlike $Y_t(1)$, $Y_t(0)$ is not observed for the treated unit. Thus, a crucial first step in the identification of $ATT(t)$ is the identification of $\mathbb{E}[Y_t(0)\mid D=1], \, t \geq 1 $. Fix a pre-treatment period $\tau \in [-\mathcal{T}] $, then consider a standard parallel trends assumption $\mathbb{E}[Y_t(0)-Y_\tau(0)\mid D=1]=\mathbb{E}[Y_t(0)-Y_\tau(0)\mid D=0], \ t\in [T] $, e.g., roth-SantAnna-Bilinski-Poe-2023 and a no anticipation assumption $ Y_\tau(1) = Y_\tau(0) $, e.g., roth-SantAnna-Bilinski-Poe-2023. The identification of $ \big\{ ATT(t), \ t\in [T] \big\} $, namely $ ATT(t) = ATT_{\tau,t} $, where $ATT_{\tau,t} := \mathbb{E}[Y_{1,t}-Y_{1,\tau}] - \mathbb{E}[Y_{0,t}-Y_{0,\tau}]$ denotes the DiD quantity for $(\tau,t) \in [-\mathcal{T}] \times [T] $, follows from both standard identification conditions. There are four random variables in the expression of the identified $ATT(t)$: $Y_{1,t},\ Y_{1,\tau}, \ Y_{0,t}$, and $ Y_{0,\tau}$. In the two-unit baseline setting, one has at most four realisations of data thus there can be only one summand in an estimator of $ATT(t)$, namely $\widehat{ATT}_{\tau,t}:= (Y_{1,t} - Y_{1,\tau}) - (Y_{0,t}-Y_{0,\tau})$. Under the aforementioned standard C-DiD identification conditions, $\widehat{ATT}_{\tau,t}$ is unbiased for each $ATT(t), \ t\in [T] $ for a fixed $\tau \in [-\mathcal{T}] $. However, $\widehat{ATT}_{\tau,t}$ cannot be consistent, let alone lead to any meaningful inference procedure for each $t \in [T] $ since the number of cross-sectional units is fixed. Identifying variation from cross-sectional units, as in C-DiD settings with fixed $\mathcal{T},T$ and a large number of both treated and control units, cannot be exploited in the setting considered by this paper.
One can, instead, use averages of $\widehat{ATT}_{\tau,t} $ across all pairwise combinations of pre-treatment and post-treatment periods, namely
\[
\widehat{ATT}_{1,T} = \frac{1}{\mathcal{T} T}\sum_{t=1}^T\sum_{-\tau=1}^{\mathcal{T}} \widehat{ATT}_{\tau,t},
\]to target the parameter
\[
ATT_{1,T} = \frac{1}{T} \sum_{t=1}^T ATT(t).
\]Leveraging temporal variation in the data under weak dependence conditions, the above estimand can be consistently estimated, and one can obtain an asymptotically normal distribution of the estimator. Moreover, the aggregation scheme enables the parallel trends and the limited anticipation identification conditions to only hold on aggregate.
The preceding discussion focuses on the uniformly weighted estimand \( ATT_{1,T} \). In general, however, the treatment effect may vary across post-treatment periods, yielding heterogeneous \( ATT(t) \), \( t \geq 1 \), which are aggregated through averaging over \( t \geq 1 \). As previously argued, such period-specific effects \( ATT(t) \) cannot be consistently estimated individually in the present two-unit setting. To explore potential effect heterogeneity within this framework, the paper introduces a generalisation of \( ATT_{1,T} \) to a class of estimands defined as convex-weighted averages of \(\{ ATT(t), \ t \in [T] \}\), where the researcher specifies the convex weighting scheme \(\{ w_T(t), \ t \in [T] \}\):
\begin{equation}
ATT_{\omega,T} = \sum_{t=1}^T w_T(t) ATT(t),
\end{equation}
with weights satisfying \(\sum_{t=1}^T w_T(t) = 1\) and \(w_T(t) \geq 0\). $ATT(t)$ varies over time due to dynamic treatment effects, meaning that homogeneous treatment effects across time are not assumed. Although similar $ATT_{\omega,T}$-type parameters are explored in the literature (e.g., masini2022counterfactual, carvalho-Masini-Ricardo-2018arco, chernozhukov-Wuthrich-Zhu-2024t and Li-2020-statistical), identification within the DiD framework and the corresponding asymptotic theory, as developed in this paper, appear novel. Additionally, while the parameters in the cited references rely on the synthetic control (SC) framework, which requires multiple control units, the DiD approach presented here requires only a single control unit. Moreover, the SC framework typically employs sharp null hypotheses (e.g., $\mathbb{H}_o': ATT(t) = 0 \ \forall t \in [T]$), but testing a weaker hypothesis, such as $ATT_{\omega,T} = a$ for some constant $a$, may be more relevant chernozhukov-Wuthrich-Zhu-2024t. Similar to chernozhukov-Wuthrich-Zhu-2024t, this paper's approach makes a test of the latter straightforward using standard $t$-tests.
Convex weighting schemes have the form $w_T(t) = w(t)/\sum_{t=1}^{T}w(t) $ for some non-negative function $w: \mathbb{R} \mapsto \mathbb{R}_+ $. There are interesting aggregation schemes that a researcher may want to use. For example, a researcher may be interested in assigning more weight to $ATT(t), \ t\geq 0$ for post-treatment periods closer to the treatment window and less weight to those farther from it. Alternatively, the researcher might want to assign zero weight to periods immediately following treatment where treatment effects are expected to be null. As the estimand (ref) depends on the particular weighting scheme, the results in this paper, e.g., identification, asymptotic normality, and identification testing ought to hold uniformly in a suitable (sub-)class of non-stochastic convex weighting schemes. Let $w(\cdot)$ belong to a general class of non-negative functions $\mathcal{W} $, then the class of weighting schemes is given by
\begin{align}
\mathcal{W} := \bigg\{ w: \frac{w(t)}{\sum_{t'=1}^{T}w(t')} =:w_T(t) \geq 0 and \max_{t \in [T]} w_T(t)\lesssim T^{-1} \quad uniformly in T \geq 1 \bigg\}.
\end{align}
A leading example is the uniform weighting scheme: $w_T(t) = 1/T$. Another example allows linearly decreasing weighting, which puts greater weight on $ATT(t)$ closer to the period of treatment $t=0$: $w_T(t) = \frac{T - 2at}{ \sum_{t'=1}^T (T-2at')} = \frac{1}{1-a}\frac{1 - (2a/T)t}{ T-a/(1-a)}, \ a \in [0,1/2) $. Estimation of $ATT_{\omega,T}$ in this paper exploits both pre-treatment and post-treatment temporal variation for consistency and asymptotic inference. It is thus crucial that the contribution of any $ATT(t), \, t\in [T] $ to $ATT_{\omega,T}$ be asymptotically negligible otherwise inference on the sequence of $ATT_{\omega,T}$ parameters based on the T-DiD estimator cannot be valid. This paper applies the notational convention $w(0)=0$ throughout.
\subsection{Asymptotic identification}
The counterfactual component $\sum_{t=1}^T w_T(t)\mathbb{E}[Y_t(0)\mid D=1] $ of $ATT_{\omega,T}$ in (ref) is unobservable. Its identification guarantees that of $ATT_{\omega,T}$ since $\sum_{t=1}^T w_T(t)\mathbb{E}[Y_t(1)\mid D=1] = \sum_{t=1}^T w_T(t)\mathbb{E}[Y_{1,t}] $ is identified from the sampling process. Although sufficient, standard parallel trends and no-anticipation assumptions such as those mentioned in (ref) are stronger than necessary for (asymptotically) identifying $ATT_{\omega,T}$. This paper opts for an asymptotic parallel trends assumption that allows deviations away from standard parallel trends. For the baseline result presented in this section, it is maintained that potential outcomes are \emph{suitably adjusted} for, e.g., time-varying heterogeneity via the inclusion of time-varying covariates, unit roots via first differences, and deterministic trends via detrending. Details on implementation follow in (ref), (ref), and (ref). The adjustment of potential outcomes ensures the identification conditions (ref) below are robust to such complications often encountered in time series as well as ensuring the conditions hold conditional on observable non-overlapping time-varying characteristics.
\begin{assumption}[Asymptotic Parallel Trends]
\[
\frac{1}{\mathcal{T} T}\sum_{-\tau=1}^{\mathcal{T} }\sum_{t=1}^T \Big(\mathbb{E}[Y_t(0)-Y_\tau(0)\mid D=1] - \mathbb{E}[Y_t(0) - Y_\tau(0)\mid D=0]\Big) = \mathcal{O}\big((\mathcal{T} \wedge T)^{-(1/2+\gamma)}\big), \ \gamma > 0.
\]
\end{assumption}
(ref) allows some extent of mean dependence of the paths of average untreated potential outcomes on treatment status $D$. Thus, unlike standard parallel trends assumptions which impose \( \mathbb{E}[Y_t(0)-Y_\tau(0)\mid D=1] - \mathbb{E}[Y_t(0) - Y_\tau(0)\mid D=0] = 0 \) for all $(\tau,t) \in [-\mathcal{T}]\times [T] $, (ref) allows for some, albeit controlled, violations. Moreover, (ref) applies to the average over all $(\tau,t)$-pairs and \emph{not} each pair.
One could replace (ref) with an assumption of \emph{average unbiasedness}---for example, \`a la botosaru-giacomini-weidner-2023forecasted, namely,
\[
\frac{1}{\mathcal{T} T}\sum_{-\tau=1}^{\mathcal{T} }\sum_{t=1}^T \Big( \mathbb{E}[Y_t(0)-Y_\tau(0)\mid D=1] - \mathbb{E}[Y_t(0)-Y_\tau(0)\mid D=0] \Big) = 0.
\]
This condition, while weaker than standard parallel trends, can still be restrictive. For example, such a parallel trends assumption may require some form of stationarity in untreated potential outcomes for plausibility. This paper appears to be the first to use an asymptotic form of the parallel trends assumption.
To further shed light on the asymptotic parallel trends assumption, consider the following specification of expected untreated potential outcomes: $ \mathbb{E}[Y_t(0) \mid D=d] = \nu_{0}(t,d;\eta) $ where
\[
\nu_{0}(t,d;\eta) :=
\begin{cases}
\cos(t), & d=0, \\[6pt]
0.5 + \cos(t) \;+\; 0.25\,\bigl|1+0.5\sin(t)\bigr| \cdot \operatorname{sign}(t)\,|t|^{-\eta}, & d=1,\;\;|t|\geq 1.
\end{cases}
\]
The expected untreated potential outcomes (against time) for both the treated and untreated units are plotted in (ref) with $T=\mathcal{T}=5$ and $T=\mathcal{T}=25$ at $\eta=0.6$. As can be seen, the expected untreated potential outcomes are not parallel in the traditional C-DiD sense. That is, for any pairwise combination of pre- and post-treatment periods, expected untreated potential outcomes are not parallel. However, the T-DiD parallel trends condition, namely (ref) is satisfied.
\begin{figure}[!htbp]
\caption{Asymptotic Parallel Trends}
\begin{subfigure}{0.38\textwidth}
\caption
\end{subfigure}
\begin{subfigure}{0.38\textwidth}
\caption
\end{subfigure}
\begin{justify}
The plots depict the paths of untreated potential outcomes from (ref) for the control and treated units with pre- and post-treatment time horizons (a) $\mathcal{T}=T=5$ and (b) $\mathcal{T}=T=25$.
\end{justify}
\end{figure}
\begin{figure}[!htbp]
\caption{Asymptotic Parallel Trends}
\begin{subfigure}{0.3\textwidth}
\caption{ $T=\mathcal{T}=5$ }
\end{subfigure}
\begin{subfigure}{0.3\textwidth}
\caption{ $T=\mathcal{T}=25$ }
\end{subfigure}
\begin{subfigure}{0.3\textwidth}
\caption{Violation of Asymptotic Parallel Trends, $T=\mathcal{T}$ }
\end{subfigure}
\begin{justify}
Plots (a) and (b) depict the paths of pre- and post-treatment averages of untreated potential outcomes from (ref) for the control and treated units with pre- and post-treatment time horizons (a) $\mathcal{T}=T=5$ and (b) $\mathcal{T}=T=25$. Plot (c) plots violations of the parallel trends condition (Assumption (ref)) as a function of the post-treatment time horizon $T$.
\end{justify}
\end{figure}
The first and second panels of (ref) plot the paths of pre-post averages of expected untreated potential outcomes. These are the paths in untreated potential outcomes that matter for the T-DiD. At $T=\mathcal{T}=25$ in Panel (b), relative to $T=\mathcal{T}=5$ in Panel (a), the pre-post paths of average untreated potential outcomes are closer to being parallel. Panel (c) shows that the deviation from parallel paths is decreasing in $T=\mathcal{T}$.\footnote{This argument extends to $T\wedge \mathcal{T}$ in light of (ref) below. } (ref) says that although the paths of pre-post averages of expected untreated potential outcomes may not be exactly parallel in finite samples, the violation is shrinking to zero sufficiently fast in $T\wedge\mathcal{T}$.
Consider the following Data Generating Process (DGP) of untreated potential outcomes
\begin{equation}
Y_t(0) = \alpha_0 + (\alpha_1-\alpha_0)D + \varphi(t) + \nu_{0}(t,D;\eta) + e_t,
\end{equation}where $(\alpha_0,\alpha_1)$ are constant, $ \varphi(t) $ is some possibly unbounded function of $t$, and $ \mathbb{E}[e_t \mid D = d] = 0, \ d\in \{0,1\} $. The following result shows that untreated potential outcomes following (ref) satisfy (ref).
\begin{proposition}
\Copy{Key:Prop:PT_DGP}{ Let $\varphi(\cdot)$ be some possibly unbounded function $\varphi: [-\mathcal{T}]\cup [T] \rightarrow \mathbb{R}$ and $\eta > 1/2 $, then untreated potential outcomes following (ref) satisfy (ref). }
\end{proposition}
\begin{remark}
First, temporal dependence in $\{ e_t, -\mathcal{T} \leq t \leq T \}$, e.g., MA processes, is not ruled out by (ref). Second, with no bound restrictions on $\varphi(\cdot)$ in (ref), non-stationary common shocks in individual untreated potential outcomes $Y_t(0)$ conditional on $D=d$ do not violate (ref). Third, (ref) leaves unrestricted the cross-sectional dependence between treated and control unit untreated potential outcomes. Fourth, the restriction on $ \eta > 1/2 $ does not require that $\mathbb{E}[\nu_{0}(t,D;\eta)\mid D=d] =:\nu_{0}(t,d;\eta), \, d\in \{0,1\} $ or some average thereof over time be zero; thus (ref) effectively allows violations of standard parallel trends assumptions.
\end{remark}
The third point in (ref) above, in effect, allows correlated shocks to both the treated and control units. For example, neighbouring West African countries Ghana, Togo, and Benin, were all exposed to the same wave of democratisation in the early 1990s. It is thus reasonable not to rule out correlations between country outcomes over time. The T-DiD, however, does not rely on such “common factors" in pre-treatment outcomes for constructing counterfactuals unlike, e.g., the SC and factor models -- cf. abadie-gardeazabal-2003,hsiao-ching-wan-2012,ferman-pinto-2021-synthetic,Sun-Ben-Michael-Feller-2023-using.
\begin{example}[Economic Interpretation of (ref)]
With the outcome defined as $Y_t = \log(GDP_t)$, (ref) says that, had Benin not democratised, her average pre- to post-democratisation economic growth rate would have differed “negligibly" from that of Togo. The $\mathcal{O}\big((\mathcal{T} \wedge T)^{-(1/2+\gamma)}\big)$ term in (ref) translates the magnitude of the “negligibility" required of the violations of the canonical parallel trends assumption.
\end{example}
To conclude the discussion of (ref), it is worth noting that, although one might argue that requiring a large number of post-treatment periods \( T \) constitutes a restrictive condition, this concern is less pertinent in the context considered here. Crucially, this paper focuses on settings in which the condition of a large \( T \) is credible. For instance, regime changes—such as transitions between autocracy and democracy—often entail significant costs in terms of time and human lives, thereby lending plausibility to the large-\( \mathcal{T} \wedge T \) asymptotics adopted in this paper.
Rational expectations of economic agents, e.g., foreign investors, can increase the uncertainty around the timing of treatment, i.e., when treatment \emph{actually} begins. Per the running example, the anticipation of democratisation by investors can already induce a change in foreign direct investment and private domestic investment before the documented period of transition to democratisation. While a standard no-anticipation assumption, e.g., $Y_\tau(1)=Y_\tau(0)$ or the weaker $\mathbb{E}[Y_\tau(1) - Y_\tau(0)\mid D=1] = 0$ for $\tau \in [-\mathcal{T}] $, together with (ref), should suffice for the asymptotic identification of $ATT_{\omega,T}$, it is not necessary. The assumption below allows for deviations away from standard no-anticipation assumptions.
\begin{assumption}[Asymptotically Limited Anticipation]
\[
\frac{1}{\mathcal{T} }\sum_{-\tau=1}^{\mathcal{T} } \mathbb{E}\big[Y_{\tau}(1) - Y_{\tau}(0)\mid D=1 \big] = \mathcal{O}(\mathcal{T} ^{-(1/2+\delta)}), \ \delta > 0.
\]
\end{assumption}
By definition, \( \mathbb{E}\big[Y_{\tau}(1) - Y_{\tau}(0)\mid D=1 \big] = ATT(\tau), \ \tau \leq -1 \). (ref) requires that the average anticipatory effect of treatment on the treated unit is asymptotically negligible. The following illustration draws on the running example to provide further intuition.
\begin{example}[Economic interpretation of (ref)]
In general, economic analysts and prospective investors use past democratisation experiences, among other indicators, to form expectations of a country's future political stability. Thus, the anticipation that the wave of democratisation sweeping across African countries may eventually materialise in Benin could bolster investor confidence. Such a boost in confidence can subsequently lead to a cautious increase in investment before the actual transition to democracy begins. (ref) says that a boost in investment (and GDP for that matter) owing to anticipations of democratisation is \emph{negligible} (up to some $\mathcal{O}(\mathcal{T} ^{-(1/2+\delta)})$ term).
\end{example}
Like callaway-santanna-2021, (ref) has, as a special case, limited treatment anticipation during a fixed number of periods before treatment. This special case simply involves dropping those periods out of the sample. This is further supplemented by (ref), thus providing a buffer against residual anticipatory effects.
The following remark is useful in explaining the gains from choosing suitable control units for a treated unit.
\begin{remark}
Consider a control unit, which shares a set of characteristics $Z_t$ at time $t$. For example, Togo shares several characteristics with Benin: a long border, a similar colonial history, the same monetary policy, the same currency, very similar climate and weather patterns, very similar baskets of export and import goods, and natural resources. Thus, imposing (ref) in the running example implies they hold conditional on shared characteristics.
\end{remark}
Choosing controls that share characteristics $Z_t$ (both observable and unobservable) with the treated unit serves to control for the shared characteristics and therefore weakens both parallel trends and limited anticipation conditions ((ref)) to conditional ones (on $Z_t$) -- cf. callaway-santanna-2021 and roth-SantAnna-Bilinski-Poe-2023. A main takeaway from (ref) is that the greater the overlap in characteristics between the treated and control unit, the weaker and more credible are (ref). Moreover, the rates of decay in (ref) allow controlled degrees of violation of standard parallel trends and no anticipation conditions conditional on shared characteristics.
Fixing a pre-treatment “base" period $\tau \leq -1 $,
\[
ATT_{\tau,t} = \mathbb{E}[Y_t(1)-Y_\tau(1)\mid D=1] - \mathbb{E}[Y_t(0) - Y_\tau(0)\mid D=0] = \mathbb{E}\big[Y_t - Y_\tau \mid D = 1 \big] - \mathbb{E}\big[Y_t - Y_\tau \mid D = 0 \big]
\]
for a post-treatment period $t\geq 1$ where the second equality holds because treated (resp. untreated) potential outcomes are observed for the treated (resp. untreated) unit. From the proof of (ref) below,
\begin{align*}
ATT(t) - ATT_{\tau,t} =&: -R_{\tau,t}\\
=& -\Big\{\underbrace{ \mathbb{E}[Y_t(0)-Y_\tau(0)\mid D=1] - \mathbb{E}[Y_t(0)-Y_\tau(0)\mid D=0] }_{\text{Trend Bias}(\tau,t)} - \underbrace{\mathbb{E}[Y_\tau(1)-Y_\tau(0)\mid D=1]}_{\text{Anticipation Bias}(\tau)}\Big\}.
\end{align*}
Although $ATT_{\tau,t}$ is identified, $ATT(t)$ is not identified because the difference between the trend and anticipation bias terms, i.e., $R_{\tau,t}$, is not identified from the data sampling process. To exploit temporal variation from pre-treatment periods, another convex weighting scheme on the same class $ \mathcal{W} $ is introduced, namely, $\psi \in \mathcal{W} $ with $\psi_{\mathcal{T}}(-\tau) \geq 0$ for all $\tau\leq -1$. Under the uniform weighting scheme, for example, where each period is equally weighted post- and pre-treatment, $ w_T(t) = 1/T $ and $ \psi_{\mathcal{T}}(-\tau) = 1/\mathcal{T} $. Define the effective estimand
\[
\widetilde{ATT}_n^{w,\psi}:= \sum_{t=1}^{T}\sum_{-\tau=1}^{\mathcal{T} } \psi_{\mathcal{T} }(-\tau) w_T(t)ATT_{\tau,t}.
\]
Let $n: = \mathcal{T} + T$ and $\lambda_n = T/n$. The following regularity condition controls the ratio $\lambda_n$.
\begin{assumption}
Uniformly in $n$, $ \lambda_n \in [\epsilon, \, (1-\epsilon) ] $ for some constant $ \epsilon \in (0,1/2] $.
\end{assumption}
(ref) is essential for theoretical results under large $T,\mathcal{T} $ asymptotics -- cf. chan2021pcdid and carvalho-Masini-Ricardo-2018arco. It ensures that the ratio of pre- and post-treatment periods to the total number of periods does not vanish even in the limit while also accommodating the possibility that the limit does not exist.\footnote{For example, $ \lambda_n = \epsilon + (1 - 2\epsilon) \cdot \frac{1 + \sin(n)}{2} \in [\epsilon, \, (1-\epsilon)], \quad \text{for all } n \geq 1 $.}
The following provides the (asymptotic) identification result.
\begin{theorem}[Asymptotic Identification]
\Copy{Key:Theorem:Identification}{Let (ref) hold, then
(a) $\displaystyle \sup_{\{w,\psi\} \in \mathcal{W}^2 } |ATT_{\omega,T} - \widetilde{ATT}_n^{w,\psi}| = \mathcal{O}\big(n^{-(1/2+\gamma\wedge \delta)}\big) $ and
(b) $\displaystyle \sup_{ \{\psi,\phi\} \in \mathcal{W}^2 } |\widetilde{ATT}_n^{w,\psi} - \widetilde{ATT}_n^{w,\phi}| = \mathcal{O}\big(n^{-(1/2+\gamma\wedge \delta)}\big) $ uniformly in $w\in\mathcal{W}$, where the constants $\gamma>0$ and $\delta>0$ are defined in (ref), respectively.}
\end{theorem}
Part (a) of (ref) shows that although identification may not hold exactly in finite samples, it does asymptotically as $n \rightarrow \infty $ subject to (ref). Precisely, it says $ATT_{\omega,T}$ is identified up to a $o\big(n^{-1/2}\big)$ term uniformly in $\mathcal{W}^2 $. Part (b) is particularly important as it shows that identification of $ATT_{\omega,T}$ is asymptotically invariant to the weighting scheme $\psi\in\mathcal{W} $ applied to pre-treatment outcomes. This is reasonable because pre-treatment weighting is simply a methodological device and not an intrinsic part of the $ATT_{\omega,T}$ parameter. Thus, the \emph{effective estimand} $\widetilde{ATT}_n^{w,\psi}$ is asymptotically invariant to the pre-treatment weighting scheme $\psi \in \mathcal{W} $.\footnote{Recall that the post-treatment weighting scheme $w$ is an effective part of the definition of $ATT_{\omega,T}$.}
The following decomposition
\begin{equation}
ATT_{\omega,T} = \widetilde{ATT}_n^{w,\psi} - R_n^{w,\psi}
\end{equation}
is used in the proof of (ref) where $R_n^{w,\psi}$ is the difference between the convex-weighted averages of the trend and anticipation biases, i.e., $\displaystyle R_n^{w,\psi}:=\sum_{-\tau=1}^{\mathcal{T} }\sum_{t=1}^T \psi_{\mathcal{T} }(-\tau) w_T(t)R_{\tau,t}$.
\begin{remark}
The identification result in (ref) is equivalent to $R_n^{w,\psi}=o(n^{-1/2})$ uniformly in $\mathcal{W}^2$. (ref) are imposed separately following conventional DiD identification arguments, mostly for economic interpretability and clarity. They are thus jointly sufficient but not necessary for (ref) as it suffices that $R_n^{w,\psi}$ be $o(n^{-1/2})$ uniformly in $\mathcal{W}^2$. The latter condition, which is both sufficient and necessary for asymptotic identification while allowing for an asymptotically normal estimator, comes at the cost of less transparency and intuition. Although the rates in (ref) are sufficient for identification, weaker rates that deliver $R_n^{w,\psi}=o(1)$ are also sufficient for asymptotic identification.
\end{remark}
\section{Estimation}
\subsection{The T-DiD estimator}
In view of (ref), the DiD estimator of $ATT_{\omega,T}$, which is the T-DiD, can be stated as $\widehat{ATT}_{\omega,T} = \sum_{-\tau=1}^{\mathcal{T} }\sum_{t=1}^T \psi_{\mathcal{T} }(-\tau) w_T(t)\widehat{ATT}_{\tau,t}$ where $\widehat{ATT}_{\tau,t}:= (Y_{1,t} - Y_{1,\tau}) - (Y_{0,t}-Y_{0,\tau}) $. This suggests the T-DiD can be alternatively expressed as
\begin{equation}
\widehat{ATT}_{\omega,T} = \sum_{t=1}^{T}w_T(t)(Y_{1,t} - Y_{0,t}) - \sum_{-\tau=1}^{\mathcal{T} }\psi_{\mathcal{T} }(-\tau)(Y_{1,\tau} - Y_{0,\tau}).
\end{equation}
The above expression clearly shows that $\widehat{ATT}_{\omega,T}$ is a difference of post-treatment and pre-treatment differences in the outcomes of the treated and control units; it is thus a \emph{bona fide} difference-in-differences estimator -- cf. arkhangelsky-etal-2021. While the above estimator may seem obvious, a formal treatment under identification and sampling conditions (as done in this paper) does not appear to have been done.
\begin{remark}
The main difference between (ref) and a conventional DiD estimator lies in the variation exploited for identification, estimation, and inference. In a conventional DiD setting with short panels or repeated cross-sections, the number of cross-sectional units (or clusters)—typically independently sampled—in both treated and control groups is large, while the number of time periods remains fixed. By contrast, the framework considered in this paper requires a large number of pre- and post-treatment periods, but a fixed number of cross-sectional units in the treated and control groups.
\end{remark}
From a practical point of view, a regression-based estimator is desirable as existing routines for estimation and heteroskedasticity- and auto-correlation-robust standard errors are readily applicable without modification. Define the random variable $X_t:= Y_{1,t} - Y_{0,t}$. As arbitrary cross-sectional dependence between the treated and the untreated unit is possible, one can cast the sequence of identified parameters $\widetilde{ATT}_n^{w,\psi}$ in the following (weighted) linear model:
\begin{equation}
X_t = \beta_0 + \widetilde{ATT}_n^{w,\psi} \mathbbm{1}\{t\geq 1\} + U_t
\end{equation}
using the set of \emph{non-negative weights} $\{\widetilde{w}_n(t), t \in [-\mathcal{T}] \cup [T] \} $ where $\widetilde{w}_n(t):= w_T(t)\mathbbm{1}\{t \geq 1\} + \psi_{\mathcal{T}}(t)\mathbbm{1}\{t\leq -1\} $. (ref) in the supplement shows the numerical equivalence of (ref) to a regression-based one using (ref), namely $ \widehat{ATT}_{\omega,T} = \widehat{B} $ where $\displaystyle (\widehat{\beta}_0,\widehat{B})' = \operatorname*{argmin}_{(\beta_0,B)'} S_{\widetilde{w},n}(\beta_0,B)$ and $S_{\widetilde{w},n}(\beta_0,B): = \sum_{t=-\mathcal{T}}^T \widetilde{w}_n(t)(X_t - \beta_0 - B \mathbbm{1}\{t\geq 1\})^2.$ The differencing in $X_t$ eliminates \emph{arbitrarily unbounded} common trends or common shocks in $Y_{1,t}$ and $Y_{0,t}$.
(ref) hold for suitably transformed potential outcomes. The regression-based form in (ref) renders such adjustments feasible in a straightforward manner by (1) accounting for observed uncommon time-varying heterogeneity: $X_t = \beta_0 + \widetilde{ATT}_n^{w,\psi} \mathbbm{1}\{t\geq 1\} + (\dot{Z}_{1,t}-\dot{Z}_{0,t})\bm{\beta} + U_t$ where $\dot{Z}_{d,t}$ comprises flexible transformations, e.g., polynomials, of observed uncommon time-varying covariates of unit $d \in \{0,1\} $, (2) mitigating persistence in the errors $U_t$ in order to improve efficiency, e.g., $X_t = \beta_0 + \widetilde{ATT}_n^{w,\psi} \mathbbm{1}\{t\geq 1\} + \sum_{l=1}^p X_{t-l}\beta_l + U_t$, (3) taking first differences of $X_t$ to remove unit roots: $(X_t - X_{t-1}) = \beta_0 + \widetilde{ATT}_n^{w,\psi} \mathbbm{1}\{t\geq 1\} + U_t$, (4) fitting a moving average process in (ref) to mitigate serial correlation in $U_t$, and (5) controlling for idiosyncratic time trends in $X_t$ to remove non-stationarity that can otherwise hurt asymptotic identification and inference -- details follow in (ref). These adjustments are useful. For example, adjusting for time-varying heterogeneity reduces the risk of omitted variable bias in the estimation of $\widetilde{ATT}_n^{w,\psi}$, thereby broadening the applicability of the T-DiD.\footnote{Maintaining the $\widetilde{ATT}_n^{w,\psi}$ notation in the above extensions is to avoid notational clutter.}
\subsection{Relation to alternative estimators}
The Synthetic Control (SC) method is, in many respects, the most closely related to the T-DiD. Other comparable approaches include the Synthetic DiD and Before--After (BA) estimators. This section provides a brief overview of these methods and discusses their connection to the T-DiD.
\subsubsection{Synthetic Controls}
With a single control unit, the convex pre-treatment SC weight is trivially 1. Thus, the SC estimator adapted to the baseline case considered in this paper is simply $ \widehat{ATT}^{SC} = \sum_{t=1}^{T} w_T(t) X_t = \sum_{t=1}^{T} \omega_n(t) X_t $ where $\omega_n(t):= w_T(t)\mathbbm{1}\{t\geq 1\} - \psi_{\mathcal{T} }(t)\mathbbm{1}\{t\leq -1\} $. Observe for instance that, unlike DiD estimators in general, the $\widehat{ATT}^{SC}$ imposes a stronger condition of a zero mean-difference on the post-treatment untreated potential outcomes, i.e., $\displaystyle \sum_{t=1}^{T} w_T(t)\mathbb{E}[Y_t(0)\mid D=1] = \sum_{t=1}^{T} w_T(t) \mathbb{E}[Y_{0,t}] $. One observes from $\widehat{ATT}^{SC}$ above that no use is made of pre-treatment observations; this suggests that under this strong assumption, no pre-treatment observations are needed.
A clear advantage of the T-DiD estimator emerges. First, the convex-weighted SC is valid under a strong assumption that post-treatment untreated potential outcomes have equal means (uniformly in $\mathcal{W}$). This underlying assumption can fail, and the SC suffers from identification failure as a result. Granted that (ref) are satisfied, the identification of $ATT_{\omega,T}$ holds asymptotically and inference thereon using the T-DiD framework is valid under standard regularity conditions outlined in the next section.
ferman-pinto-2021-synthetic,tian-lee-panchenko-2024-synthetic advocate demeaning outcomes before pre-treatment fitting.\footnote{A similar approach is discussed in doudchenko-imbens-2016, where an intercept term is included in the pre-treatment fit.} The demeaned convex-weighted SC with a single control unit and convex-weighted ATTs leads to a form of the T-DiD -- see (ref) for the derivation. However, the \( T \)-DiD framework, taken \emph{en bloc} as a unified approach to identification, estimation, inference, and identification testing, is not a special case of the demeaned synthetic controls in tian-lee-panchenko-2024-synthetic,ferman-pinto-2021-synthetic, as those methods require multiple control units and/or multiple outcomes for feasible SC inference.
The synthetic DiD (S-DiD) estimator introduced by arkhangelsky-etal-2021 seeks to merge desirable features of the SC and the C-DiD using panel data. The estimator is cast in a weighted two-way fixed effects regression framework:
\begin{equation*}
(\widehat{ATT}^{sdid},\widehat{\mu}_Y,\widehat{a},\widehat{b}) = \operatorname*{argmin}_{ATT,\mu_Y,a,b} \Big\{ \sum_{i=1}^N \sum_{t=-\mathcal{T}}^T \big(Y_{it} - \mu_Y - a_i - b_t - D_i\mathbbm{1}\{t\geq 1\} ATT \big)^2\hat{\omega}_i^{sdid}\lambda_t^{sdid} \Big\}
\end{equation*}
where the weights are data-driven. Unit-specific weights $\{\hat{\omega}_i^{sdid}, \ i\in [N] \}$ ensure that the average outcome for treated units aligns closely with the weighted average for control units in the pre-treatment periods, while period-specific weights $\{\lambda_t^{sdid}, \ t \in [-\mathcal{T}] \cup [T] \}$ ensure that the average post-treatment outcome for control units differs from the weighted pre-treatment average of the same control units by a constant. In this sense, the S-DiD borrows from the SC the idea of re-weighting control units to approximate pre-treatment trends, and from the DiD the robustness to additive unit-level shifts --- the unit fixed effects $a_i$ absorb permanent level differences that the SC weights need not eliminate exactly. The S-DiD is not suitable for the two-unit baseline setting of primary interest in this paper as it requires several control units; see arkhangelsky-etal-2021. More fundamentally, its asymptotic theory is driven by a large cross-section: both the number control units and pre-treatment periods ($N_{co}$ and $T_{pre}$, respectively) must diverge at comparable rates, so that identification and inference lean on cross-sectional variation rather than the temporal variation exploited in this paper.
\subsubsection{Before-After (BA)}
Another interesting estimator that emerges in the present context is the before-after estimator carvalho-Masini-Ricardo-2018arco. Unlike the SC above which uses post-treatment outcomes of the untreated unit to construct the counterfactual, the BA uses pre-treatment outcomes of the treated unit to construct the counterfactual. The expression of the BA estimator is given by $\widehat{ATT}^{BA}:= \sum_{t=1}^T w_T(t) Y_{1,t} - \sum_{-\tau=1}^{\mathcal{T} } \psi_{\mathcal{T}}(\tau) Y_{1,\tau} = \sum_{t=-\mathcal{T}}^T \omega_n(t) Y_{1,t}$. The BA estimator effectively uses the pre-treatment mean of the outcome of the treated unit to impute the mean post-treatment untreated potential outcome of the treated unit.
Denote
$ \widehat{ATT}_d^{BA}:= \sum_{t=1}^T w_T(t) Y_{d,t} - \sum_{-\tau=1}^{\mathcal{T} } \psi_{\mathcal{T}}(\tau) Y_{d,\tau} $ for $d\in \{0,1\}$ and note that $\widehat{ATT}_{\omega,T} = \widehat{ATT}_1^{BA} - \widehat{ATT}_0^{BA} = \widehat{ATT}^{BA} - \widehat{ATT}_0^{BA} $. Thus, the T-DiD is the difference between the BA estimators of the treated and control units, while the BA estimator ignores the second term. The BA is therefore not robust to non-negligible correlated or common shocks to which both treated and control units may be exposed. Unlike the T-DiD, the BA is not invariant to the pre-treatment weighting scheme. Thus, using non-uniform weights for the BA estimator may not be meaningful.
\section{Asymptotic Theory}
\subsection{Baseline - Two units}
For the treatment of the asymptotic theory of the T-DiD estimator (ref), it is notationally convenient to opt for an alternative formulation:
\begin{equation}
\widehat{ATT}_{\omega,T} = \sum_{t=-\mathcal{T} }^{T} \omega_n(t)X_t,
\end{equation}where $\omega_n(t):= w_T(t)\mathbbm{1}\{t\geq 1\} - \psi_{\mathcal{T} }(t)\mathbbm{1}\{t\leq -1\} $ and $\omega_n(0)=0$ by notational convention. This paper allows heterogeneity and weak dependence of quite general forms in the sampling process of the data $\{(Y_{1,t},Y_{0,t}), t \in [-\mathcal{T}] \cup [T] \}$. For example, $ \big \{X_t:= Y_{1,t}-Y_{0,t}, \ t \in [-\mathcal{T}] \cup [T] \big \}$ are allowed to be auto-correlated and heterogeneous, thereby accommodating various interesting forms of dependence. To this end, the following statement supplies the definition of near-epoch dependence (NED), which is flexible enough to accommodate several empirically relevant forms of temporal dependence and heterogeneity.
\begin{definition}[Near-Epoch Dependence -- davidson2021stochastic Definition 18.2]
Let $ \{\{V_{nt}\}_{t=-\infty}^{\infty}\}_{n=1}^{\infty} $ be a possibly vector-valued stochastic array defined on the probability space $ (\Omega,\mathcal{F},P) $ and $\mathcal{F}_{n,t-m}^{t+m}:= \sigma(V_{n,t-m},\ldots,V_{n,t+m}) $ where $\sigma(\cdot)$ denotes a sigma-algebra. An integrable array $ \{\{U_{nt}\}_{t=-\infty}^{+\infty}\}_{n=1}^{\infty} $ is $L_p$-NED on $ \{ V_{nt} \} $ if it satisfies $||U_{nt} - \mathbb{E}[U_{nt}|\mathcal{F}_{n,t-m}^{t+m}]||_p \leq d_{nt}\nu_m $ where $\nu_m \rightarrow 0 $ and $ \{d_{nt}\} $ is an array of positive constants.
\end{definition}
davidson2021stochastic provides some examples of NED processes: (1) linear processes with absolutely summable coefficients davidson2021stochastic, (2) bi-linear auto-regressive moving average processes davidson2021stochastic, (3) a wide class of lag functions subject to dynamic stability conditions davidson2021stochastic, and GARCH(1,1) processes under given conditions hansen-1991-garch.
Define
\begin{align*}
s_{\omega,n}^2:= \mathbb{V}\Big[\sum_{t=-\mathcal{T} }^{T} \omega_n(t)X_t \Big] = \sum_{t=-\mathcal{T} }^{T} \omega_n(t)^2\mathbb{V}[X_t] + 2 \sum_{t=-\mathcal{T} }^{T}\sum_{t'=-\mathcal{T} }^{t-1} \omega_n(t)\omega_n(t')\mathrm{cov}[X_t,X_{t'}]
\end{align*}
and $ X_{nt}:= \omega_n(t)(X_t - \mathbb{E}[X_t])/s_{\omega,n} $, then $\big\{X_{nt}, \ t\in [-\mathcal{T}] \cup [T], n \in \mathbb{N} \big\}$ constitutes a triangular array of zero-mean random variables. Also, define $\sigma_{nt}:= \sqrt{\mathbb{E}[X_{nt}^2]} $, $c_{nt}:= \max\{\sigma_{nt}, s_{\omega,n}\}$, and $d_{nt} \leq \bar{d}||X_{nt}||_2$ uniformly in $n$ and $t$ for some constant $\bar{d}>0$. The following set of assumptions is important for establishing the asymptotic normality of the T-DiD estimator. Let $r$ and $C$ be constants such that $r>2$ and $0<C<\infty$.
\begin{assumption}
$ \displaystyle \sup_{t \in [-\mathcal{T}] \cup [T]} ||(Y_{1,t} - Y_{0,t}) - \mathbb{E}[(Y_{1,t} - Y_{0,t})]||_r \leq C $ uniformly in $n$.
\end{assumption}
(ref) is weak in two dimensions: (1) it is imposed on the difference $X_t:=Y_{1,t} - Y_{0,t}$ and \emph{not} on the levels $Y_{d,t}$, $(d,t) \in \{0,1\} \times [-\mathcal{T}] \cup [T] $ and (2) it is imposed on deviations of $X_t$ from its mean $\mathbb{E}[X_t]$. (ref) thus does not rule out non-stationary common shocks. (ref) also does not rule out possibly unbounded deterministic trends in $X_t$. It is weaker than sub-Gaussianity and variance-homogeneity conditions needed for some SC methods -- cf. arkhangelsky-etal-2021 and Sun-Ben-Michael-Feller-2023-using.
The following is a standard regularity condition on $s_{\omega,n}$.
\begin{assumption}
$ \sqrt{n}s_{\omega,n} \geq \epsilon $ for some small constant $\epsilon>0$ uniformly in $\mathcal{W}^2$.
\end{assumption}
(ref) is a flexible technical condition on $s_{\omega,n}$ that accommodates, e.g., covariance stationarity, i,e., $s_{\omega,n}=\mathcal{O}(n^{-1/2})$. It, however, rules out degeneracy in the limit: $\sqrt{n}s_{\omega,n} \rightarrow 0 $ as $n\rightarrow \infty$. As (ref) restricts only $s_{\omega,n}$, standard deviations $\sigma_{nt}$ for some $t \in [-\mathcal{T}] \cup [T] $ are allowed to be zero (in the limit). (ref), however, restricts this form of degeneracy.
This paper uses the general representation $X_{nt} = g_t(\ldots,V_{n,t-1},V_{nt},V_{n,t+1},\ldots)$ for some mixing array $V_{nt}$ and time-dependent function $g_t(\cdot)$. The following assumption ensures that this representation satisfies the required near epoch dependence and mixing properties.
\begin{assumption}
(a) $\{X_{nt}\} $ is $L_2$-near epoch dependent of size $-1/2$ with respect to a constant array $\{d_{nt}\}$ on $ \{V_{nt}\} $.
(b) $ \{V_{nt}\} $ is an $\alpha$-mixing array of size $-r/(r-1)$.
\end{assumption}
(ref) allows the observed (weighted) data to be heterogeneous and weakly dependent. Besides, $\{ATT(t), \ t \in [T]\} $ and the user-specified weighting scheme $ \{\omega_n(t), \ t\in [-\mathcal{T}] \cup [T] \} $ are also allowed to be heterogeneous in $t$. (ref) is a more general form of weak dependence than commonly imposed mixing conditions in the literature -- cf. chan2021pcdid,Sun-Ben-Michael-Feller-2023-using,fry-2024-method.
The next assumption is essential in establishing central limit theorem results under NED.
\begin{assumption}
\[\max_{1\leq j\leq r_n+1}M_{nj}=o(b_n^{-1/2}) \text{ and }\sum_{j=1}^{r_n} M_{nj}^2 = \mathcal{O}(b_n^{-1}) \]
where $\displaystyle M_{nj}=\max_{(j-1)b_n+1\leq t \leq jb_n} c_{nt}$, for $j \in [r_n]$, $\displaystyle M_{n,r_n+1} = \max_{r_nb_n+1\leq t \leq n} c_{nt}$, $b_n= \lfloor n^{1-\alpha} \rfloor $, $\alpha\in(0,1]$, and $r_n= \lfloor n/b_n \rfloor $.
\end{assumption}
While (ref) allows heterogeneity in observed data, (ref) serves to restrict the degree of heterogeneity allowed -- see deJong-1997central. Also, while (ref) characterises a lower bound on the growth rate of $s_{\omega,n}$ allowed, (ref) characterises an upper bound on its growth rate, i.e., $n^{(1-\alpha)/2}c_{nt}= n^{(1-\alpha)/2}\max\{\sigma_{nt}, s_{\omega,n}\} = o(1), \ \alpha \in (0,1] $. The absence of this upper bound on the growth rate of $s_{\omega,n}$ can result in long-range dependence, and asymptotic normality fails.
With the foregoing assumptions in hand, the asymptotic normality of the T-DiD follows.
\begin{theorem}[Asymptotic Normality]
\Copy{Key:Theorem:AsympN}{Under (ref), $
s_{\omega,n}^{-1}(\widehat{ATT}_{\omega,T} - ATT_{\omega,T}) \xrightarrow{d} \mathcal{N}(0,1) $ uniformly in $\mathcal{W}^2$.}
\end{theorem}
In practice, $s_{\omega,n}$ is estimated using a heteroskedasticity and auto-correlation robust procedure, e.g., the newey-west-1987-simple procedure. Using the self-normalised $t$-test statistic of chernozhukov-Wuthrich-Zhu-2024t is an alternative inference procedure. Exploring this path to inference in the current framework is, however, left for future work due to considerations of scope and space. The asymptotic theory for the regression-based estimator of $ATT_{\omega,T}$ in (ref) and its variants is analogous to the preceding analysis; formal details are omitted. (ref) formally considers deterministic time trends.
\subsection{Extension -- Multiple control units}
Beyond the two-unit baseline case, practical cases often involve a fixed number of units with large $\mathcal{T}, T$ (e.g., marcus-santanna-2021role,grooms-2015-enforcing), where both SC and T-DiD might be suitable. This section focuses on scenarios with multiple control units, say two or three, and large $\mathcal{T}$ and $T$, which still fall within the T-DiD framework. In particular, the availability of multiple control units is exploited in this paper to propose an over-identifying restrictions test of identification. Two cases of interest arise when multiple units are available: (1) \( J > 1 \) control units for a given treated unit, and (2) \( I > 1 \) treated units, each of which has at least one corresponding control unit. To maintain focus on the primary case of a single treated unit, the generalisation to multiple treated units is deferred to (ref). It is assumed that the researcher has control units $j\in [J]$ for the treated unit. Whenever there are $J>1$ valid control units for the treated unit, one can, using the theory elaborated in (ref) above, obtain $J$ estimates of $ATT_{\omega,T}$. In what follows, let $\widehat{ATT}_{\omega,jT}$ denote the estimator (ref) using control unit $j \in [J] $ and the treated unit.
Let $\widehat{ATT}_{\omega,\cdot T}:=(\widehat{ATT}_{\omega, 1 T},\ldots,\widehat{ATT}_{\omega, J T})'$ denote a $J\times 1$ vector of T-DiD estimators of $ATT_{\omega,T}$. Define the $J\times J$ matrix $S_{\omega,n}:= \mathbb{E}[(\widehat{ATT}_{\omega,\cdot T} - \mathbbm{1}_JATT_{\omega,T})(\widehat{ATT}_{\omega,\cdot T} - \mathbbm{1}_J ATT_{\omega,T})']$ where $\mathbbm{1}_J$ denotes a $J \times 1$ vector of ones. The following presents an extension of (ref) to the multiple-control setting. It requires that the minimum eigenvalue of \( nS_{\omega,n} \) be uniformly bounded below by \( \epsilon^2 > 0 \).
\begin{namedassumption}{(ref)-Ext}
Uniformly in $n$, $\rho_{\mathrm{min}}(nS_{\omega,n})\geq\epsilon^2$.
\end{namedassumption}
The following provides the (ref)-analogue of the multiple-control case. Define $\mathbb{S}^J:= \{ \uptau \in \mathbb{R}^J: ||\uptau|| = 1 \} $ where $J \in \mathbb{N} $.
\begin{corollary}
\Copy{Key:Theorem:AsympN_ext}{Suppose (ref) hold for each of the $J$ unique treated-control pairs. Suppose further that (ref) holds, then $(\uptau_J'S_{\omega,n}\uptau_J)^{-1/2}\uptau_J'(\widehat{ATT}_{\omega,\cdot T} -\mathbbm{1}_J ATT_{\omega,T} ) \xrightarrow{d} \mathcal{N}(0,1)$ uniformly in $\mathcal{W}^2 $ for all $\uptau_J \in \mathbb{S}^J$.}
\end{corollary}
A useful deduction from (ref) using the Cram\'er-Wold device is that $S_{\omega,n}^{-1/2}(\widehat{ATT}_{\omega,\cdot T} - \mathbbm{1}_J ATT_{\omega,T}) \xrightarrow{d} \mathcal{N}(0,\mathrm{I}_J)$ under the imposed conditions. Although theoretically useful in its own right, (ref) is not interesting from a practical point of view as no direction vector $\uptau_J \in \mathbb{S}^J $ is specified. With asymptotically “efficient" linear combinations of $\widehat{ATT}_{\omega,\cdot T}$ and over-identifying restrictions tests of identification in mind, it is useful to cast the problem of finding an optimal linear combination as follows:
\begin{align}
\widehat{ATT}_{\omega,T}^{h_J}: = \operatorname*{argmin}_{a} \big\{ (\widehat{ATT}_{\omega,\cdot T} - \mathbbm{1}_Ja)'H_n(\widehat{ATT}_{\omega,\cdot T} - \mathbbm{1}_Ja) \big\} = h_J'\widehat{ATT}_{\omega,\cdot T}
\end{align}
where $h_J:= H_n'\mathbbm{1}_J(\mathbbm{1}_J'H_n\mathbbm{1}_J)^{-1}$, $H_n$ is a sequence of $J\times J$ positive definite matrices, and $\mathbb{V}[\widehat{ATT}_{\omega,T}^{h_J}] = h_J'S_{\omega,n}h_J = (\mathbbm{1}_J'H_n\mathbbm{1}_J)^{-2}\mathbbm{1}_J'H_nS_{\omega,n}H_n\mathbbm{1}_J $. Thus $\widehat{ATT}_{\omega,T}^{h_J}$ entails a specific linear combination of $\widehat{ATT}_{\omega, \cdot T}$, namely $h_J$ which in turn is a function of the positive definite $H_n$. Using standard Generalised Method of Moments (GMM) and minimum distance estimation (MD) arguments, e.g., wooldridge-2010, the efficient choice of $H_n$ within the MD class of $ATT_{\omega,T}$ estimators $\widehat{ATT}_{\omega,T}^{h_J}$ is $H_n=S_{\omega,n}^{-1}$ whence
\[
\widehat{ATT}_{\omega,T}^*:= \widehat{ATT}_{\omega,T}^{h_J^*} \text{ with } (s_{\omega,n}^*)^2: = \mathbb{V}[\widehat{ATT}_{\omega,T}^*]=1/(\mathbbm{1}_J'S_{\omega,n}^{-1}\mathbbm{1}_J).
\]
For completeness, the efficiency result is provided in (ref). Let $\mathbb{H}^J:= \{h\in \mathbb{R}^J: h'\mathbbm{1}_J = 1\} $ denote the space of vectors in $\mathbb{R}^J$ whose elements sum to 1, and observe that $h_J\in \mathbb{H}^J$.
\begin{proposition}
\Copy{Key:Prop:Eff_ATT}{Under the assumptions of (ref), $h_{J}^*:= S_{\omega,n}^{-1}\mathbbm{1}_J(\mathbbm{1}_J'S_{\omega,n}^{-1}\mathbbm{1}_J)^{-1} $ delivers the most efficient estimator of $ATT_{\omega,T}$ among all $h_{J} \in \mathbb{H}^J$.}
\end{proposition}
\section{Over-identifying Restrictions Test}
Whenever all available candidate control units of the treated unit are valid, each element in $\widehat{ATT}_{\omega,\cdot T}$ is a consistent estimator for the same $ATT_{\omega,T}$. This provides the basis for a test of over-identifying restrictions. Thus, the proposed test is similar in spirit to that of marcus-santanna-2021role, exploiting over-identification for efficiency and identification testing.\footnote{Such a test is not feasible in the SC framework, when a pre-treatment fit is taken of \emph{all} control outcomes.}
\subsection{Two-candidate controls}
To drive intuition, consider the simple case with two candidate control units. One would expect the same $ATT_{\omega,T}$ up to a negligible $o(n^{-1/2})$ bias term from both $\widetilde{ATT}_{1n}^{w,\psi}$ and $\widetilde{ATT}_{2n}^{w,\psi}$, where $\widetilde{ATT}_{jn}^{w,\psi}$ is the effective estimand $\widetilde{ATT}_n^{w,\psi}$ using control unit $j\in [2]$. Thus from the decomposition (ref),
\begin{equation*}
\widetilde{ATT}_{1n}^{w,\psi} - \widetilde{ATT}_{2n}^{w,\psi} = R_{1n}^{w,\psi} - R_{2n}^{w,\psi}
\end{equation*} where $R_{jn}^{w,\psi}$ denotes $R_n^{w,\psi}$ using control unit $j$. Let $Y_{0,t}^{(j)}$ denote the outcome of the $j$'th control unit at period $t$, then
$$
\widehat{ATT}_{\omega,1T} - \widehat{ATT}_{\omega,2T} = \sum_{t=1}^{T}w_T(t)(Y_{0,t}^{(1)} - Y_{0,t}^{(2)}) - \sum_{-\tau=1}^{\mathcal{T} }\psi_{\mathcal{T} }(-\tau)(Y_{0,\tau}^{(1)} - Y_{0,\tau}^{(2)}).
$$
The expression is a difference-in-differences estimator akin to (ref) using two control units since the outcome of the treated unit cancels out. This baseline case is particularly appealing because it involves a straightforward implementation: estimating $ATT_{\omega,T}$ with just two controls and then testing for a null effect using a $t$-test. Suppose (ref) holds for each candidate control unit --- an implication of (ref), with (ref) maintained --- then it follows that $R_{jn}^{w,\psi} = o(n^{-1/2})$ for each $j\in [2]$. This further implies $\widetilde{ATT}_{1n}^{w,\psi} - \widetilde{ATT}_{2n}^{w,\psi} = o(n^{-1/2})$. $\widehat{ATT}_{\omega,1T} - \widehat{ATT}_{\omega,2T}$ would not be statistically significant if both controls were valid, i.e., each control guarantees that (ref) hold for a given treated unit.
\subsection{Multiple candidate controls}
Generally, when more than two candidate control units are available, it follows from the identification conditions in (ref)—see (ref)—that \( \| R_{\cdot n}^{w,\psi} \| = o(n^{-1/2}) \) uniformly over \( \mathcal{W}^2 \), where \( R_{\cdot n}^{w,\psi} := (R_{1n}^{w,\psi}, \ldots, R_{Jn}^{w,\psi})' \) denotes the \( J \times 1 \) vector of bias indexed by the \( J \) control units for a given treated unit. In view of (ref), (ref), and (ref), one can characterise the test hypotheses using the representation
$$
||\sqrt{n}R_{\cdot n}^{w,\psi}|| = C_{\omega,n} n^{1/2 - \tilde{\delta}}
$$ where $ \{C_{\omega,n}: n \geq 1 \}$ is a sequence of bounded positive constants. $\tilde{\delta} > 1/2$ if asymptotic identification ((ref)) holds for each $j \in [J] $, and $\tilde{\delta}\leq 1/2$ otherwise for at least one control unit $j$. The following correspond to the null, local alternative, and fixed alternative hypotheses:
\[
\mathbb{H}_o: \tilde{\delta} > 1/2; \quad \mathbb{H}_{an}: \tilde{\delta} = 1/2; \quad \text{and} \quad \mathbb{H}_{a}: \tilde{\delta} < 1/2.
\]
If (ref) hold for each control unit \( j \in [J] \), then the null hypothesis \( \mathbb{H}_o \) is satisfied. Exact identification in the C-DiD set-up holds under $\mathbb{H}_o$ with $\tilde{\delta} = \infty$. Thus, the current framework only requires $\tilde{\delta} > 1/2$. Under the local alternative $\mathbb{H}_{an}$ or fixed alternative $\mathbb{H}_a$, either (ref) or (ref) is violated for at least one control unit.
Define $\widehat{S}_{\omega,n}$ as the estimator of the covariance matrix $\widetilde{S}_{\omega,n}:=\mathbb{E}[(\widehat{ATT}_{\omega,\cdot T} - \mathbb{E}[\widehat{ATT}_{\omega,\cdot T}])(\widehat{ATT}_{\omega,\cdot T} - \mathbb{E}[\widehat{ATT}_{\omega,\cdot T}])']$.\footnote{The separate notation of the covariance matrix is introduced since $\widetilde{S}_{\omega,n}=S_{\omega,n}$ only holds under $\mathbb{H}_o$. $\widetilde{S}_{\omega,n}$, in contrast, remains a valid covariance matrix under $\mathbb{H}_o$, $\mathbb{H}_{an}$, and $\mathbb{H}_a$.} The following consistency condition is imposed on $\widehat{S}_{\omega,n}$.
\begin{assumption}
$ \displaystyle \frac{\uptau_J'\widehat{S}_{\omega,n}\uptau_J}{\uptau_J'\widetilde{S}_{\omega,n}\uptau_J} \xrightarrow{p} 1 $ for all $\uptau_J \in \mathbb{S}^J $.
\end{assumption}
The over-identifying restrictions test statistic for testing $\mathbb{H}_o$ above is
\begin{align*}
\widehat{Q}_{\omega,n}: = (\widehat{ATT}_{\omega,\cdot T} - \mathbbm{1}_J\widehat{ATT}_{\omega,T}^*)'\widehat{S}_{\omega,n}^{-1}(\widehat{ATT}_{\omega,\cdot T} - \mathbbm{1}_J\widehat{ATT}_{\omega,T}^*).
\end{align*}
While (ref) imply $\mathbb{H}_o$, the converse generally does not hold. There exist violations of (ref), (ref), or both that imply neither $\mathbb{H}_{an}$ nor $\mathbb{H}_{a}$. Technically, the proposed test derives power under the condition that $\uptau_J^R:= ||\sqrt{n}R_{\cdot n}||^{-1} \sqrt{n}R_{\cdot n} $ does not lie in the null space of $(n\widetilde{S}_{\omega,n})^{-1/2}[\mathrm{I}_J - \widetilde{P}_{n}](n\widetilde{S}_{\omega,n})^{-1/2}$ as $n\rightarrow \infty$ where $\widetilde{P}_{n}:= \widetilde{S}_{\omega,n}^{-1/2}\mathbbm{1}_J(\mathbbm{1}_J'\widetilde{S}_{\omega,n}^{-1}\mathbbm{1}_J)^{-1} \mathbbm{1}_J'\widetilde{S}_{\omega,n}^{-1/2} $. Thus, the test has trivial power under $\mathbb{H}_{an}$ and $\mathbb{H}_a$ if the above condition is violated. This constitutes the source of the inconsistency of over-identifying restrictions tests, including the one proposed above.\footnote{guggenberger2012note provides a detailed discussion.}
As to whether ${\uptau_J^R}'(n\widetilde{S}_{\omega,n})^{-1/2}[\mathrm{I}_J - \widetilde{P}_{n}](n\widetilde{S}_{\omega,n})^{-1/2}\uptau_J^R=o(n^{-1/2})$ can hold under $\mathbb{H}_{an}$ or $\mathbb{H}_{a}$ is less of a statistical problem and more of the economic question at hand. Thus, from a statistical point of view, the above test is not consistent, i.e., it cannot detect \emph{all} violations of $\mathbb{H}_o$. Thus, a good understanding of the economics of such a violation would be useful in complementing the over-identifying restrictions test of $\mathbb{H}_o$. To illustrate in the context of the empirical example, suppose the asymptotic parallel trends assumption ((ref)), with Togo as a control for Benin, were violated. If the path of pre-post log GDP per capita averages of Cameroon, a second control unit, coincided with that of Togo, the proposed test would be incapable of detecting this violation. In other words, the test fails to reject the null if the trend biases across control units are identical up to a $o(n^{-1/2})$ term.
Let $F_n^\omega$ be the distribution from which observed data are drawn, and let $\mathcal{F}_o^\omega \subset \mathcal{F}^\omega$ characterise the set of distributions over which the test is not consistent, i.e., the set of distributions of the data for which $\mathbb{H}_o$ is violated but the test lacks power, i.e., ${\uptau_J^R}'(n\widetilde{S}_{\omega,n})^{-1/2}[\mathrm{I}_J - \widetilde{P}_{n}](n\widetilde{S}_{\omega,n})^{-1/2}\uptau_J^R=o(n^{-1/2})$.\footnote{The superscript $\omega $ in $F_o^\omega $ and $ \mathcal{F}^\omega$ emphasises the dependence on the weighting scheme in $\mathcal{W}^2$.} Thus, $\displaystyle 0 < \theta:= \lim_{n\rightarrow \infty} ||[\mathrm{I}_J - \widetilde{P}_{n}](n\widetilde{S}_{\omega,n})^{-1/2}\uptau_J^R||^2\cdot ||\sqrt{n}R_{\cdot n}||^2$ if $F_n^\omega \not\in \mathcal{F}_o^\omega $. Observe that distributions for which $\mathbb{H}_o$ is true, namely $\mathcal{F}_{\mathbb{H}_o}^\omega$ belong to $\mathcal{F}_o^\omega$. The following shows the validity and non-trivial power of the test over the set of distributions $\mathcal{F}^\omega\setminus\mathcal{F}_o^\omega$.
\begin{theorem}
\Copy{Key:Theorem:Test_Implication}{Let (ref) hold, then (a) under $\mathbb{H}_o$, $\widehat{Q}_{\omega,n} \xrightarrow{d} \chi_{J-1}^2$; (b) under $\mathbb{H}_{an}$ and if $F_n^\omega \not\in \mathcal{F}_o^\omega $, $\widehat{Q}_{\omega,n} \xrightarrow{d} \chi_{J-1}^2(\theta)$; and (c) under $\mathbb{H}_{a}$ and if $F_n^\omega \not\in \mathcal{F}_o^\omega $, $\widehat{Q}_{\omega,n} \rightarrow \infty$ uniformly in $\mathcal{W}^2$.}
\end{theorem}
Simulations in (ref) support the theory developed in this paper by demonstrating the strong performance of the DiD estimator and the identification test in terms of low bias, good empirical size control, and substantial power under alternatives.
\section{Empirical Analysis}
This section concerns the empirical results. The first part briefly explains the economic model of the democracy-growth nexus, the second describes the data, and the third applies the T-DiD to estimate the effect of democracy on Benin's GDP per capita.
\subsection{The economic model}
On the one hand, democracy is a useful indicator of stability in a country. Thus, democracy can be an indicator of low country risk from the perspective of investors. Investment is useful not just in maintaining depreciating capital but also in growing the capital stock. From standard macroeconomic models, e.g., the Solow model, output per worker increases in capital per worker. In this sense, one can expect democracy to drive economic growth. On the other hand, democracy involves elections at sometimes unpredictable frequencies, which can be a source of uncertainty since a change of government can lead to radical changes in policy direction. To the extent that such uncertainty undermines investor confidence, one can expect democracy to negatively impact economic growth. Thus, it may not be \textit{a priori} clear if and in which direction democracy generally impacts economic growth. The answer may therefore be country-specific.
\subsection{Data}
The democracy measure is sourced from the \textit{Varieties of Democracy} (V-Dem) project. The data comprise annual observations from 1960 to 2018. The measure of democracy is defined as the average of the five high-level democracy indices of the V-Dem project, namely the electoral democracy index, liberal democracy index, participatory democracy index, deliberative democracy index, and egalitarian democracy index. Each index is an aggregate index capturing specific dimensions of the concept of democracy and ranges from 0 (zero democracy) to 1 (full democracy).
\pgfplotsset{my personal style/.style=
{font=},width=12cm,height=7.0 cm}
\begin{figure}[h!]
\begin{center}
\caption{Democracy index}
\begin{tikzpicture}[]
\begin{axis}[my personal style,minor x tick num=1,
xlabel=Period,
ymin=,
ymax=1,
xticklabels={,,1960,1970,1980,1990,2000,2010,2020},
ylabel=,
title=,legend style={
at={(0.50,0.9)},
anchor=north east}]
\addplot[no markers,color=black,line width=1pt,solid,mark size=0.5pt,smooth] table[col sep=tab,x=Year,y=Benin] {figures/DemocracyIndexAll.txt};
\addplot[no markers,color=black,line width=1pt,densely dotted,mark size=0.5pt,smooth] table[col sep=tab,x=Year,y=Togo] {figures/DemocracyIndexAll.txt};
\addplot[draw=black] coordinates {(1990,0) (1990,0.8) };
\addplot[draw=gray] coordinates {(1960,0.5) (2016,0.5) };
\legend{Benin, Togo, Threshold (0.5)}
\end{axis}
\end{tikzpicture}
\end{center}
{
\textit{Notes:} The plots provide the aggregate democracy indices for Benin and Togo from 1960 through 2018. The thin vertical line at 1990 marks the beginning of Benin's democratisation process, whereas the horizontal line marks the 0.5 threshold for democracy.
}
\end{figure}
(ref) plots the democracy index for Benin and Togo. As observed, Benin and Togo have similar levels of democracy until 1990; the average gap in democracy between the two countries is around 0.039, with Benin being the more democratic country. Starting in 1990, when the democratisation process begins, the two countries diverge; the gap now stands at approximately 0.269, seven times the pre-1990 gap, with Benin becoming the more democratic country. Overall, (ref) indicates a 600-percent gain in Benin's democracy relative to Togo during the 1990s, 2000s, and 2010s. Although democracy in Togo improves since 1990, it does not reach a level that qualifies as democratic. This is why Togo serves as a control unit for Benin. For the empirical analyses, the main transition window considered is 1990-1992, and treatment remains an absorbing state up to 2018.
The outcome variable is GDP per capita of Benin and Togo \emph{in constant 2015 US dollars} from 1960 to 2018. The data are sourced from the World Bank Development Indicators.\footnote{The preferred measure of the outcome, PPP-adjusted GDP per capita, is not available for the 1960-1989 pre-treatment period.} (ref) shows the GDP per capita of Benin and Togo from 1960 to 2018. As can be observed, neither country's GDP per capita dominates until 1990 when Benin begins the democratisation process. After that, Benin's economic performance markedly surpasses Togo's.
\begin{figure}[!htbp]
\caption{GDP per Capita}
\begin{subfigure}{0.49\textwidth}
\caption{Level}
\end{subfigure}
\begin{subfigure}{0.49\textwidth}
\caption{Log}
\end{subfigure}
\begin{justify}
{
\textit{Notes:} The above plots show level and log GDP per capita, measured in constant 2015 US dollars. The vertical dotted lines at 1990 mark the beginning of Benin's democratisation process.
}
\end{justify}
\end{figure}
\subsection{The effect of democracy on growth}
(ref) presents the empirical results. Column-wise, results in Panel A are based on log GDP per capita, while results in Panel B are based on GDP per capita in levels (in constant 2015 US dollars). For each panel, there are four specifications by the transition window taken as the transition period of Benin from an autocracy to a democracy: (1) 1990-1992, (2) 1990, (3) 1991, and (4) 1992. Row-wise, the first part of (ref) corresponds to estimates of $ATT_{\omega,T}$ and the parameter $\beta_1$ on $X_{t-1}$, which captures persistence in the series $X_t$. The estimates are accompanied by heteroskedasticity- and auto-correlation-robust standard errors (in parentheses). The second part comprises typical diagnostic tests: (1) the Augmented Dickey-Fuller (ADF) test, (2) the Kwiatkowski-Phillips-Schmidt-Shin (KPSS) test, and (3) the Durbin-Watson (DW) test with respective null hypotheses (1) unit root, (2) trend stationarity, and (3) zero auto-correlation against a two-sided alternative. The uniform weighting scheme $ \omega_n(t) = T^{-1}\mathbbm{1}\{t\geq 1\} - \mathcal{T}^{-1}\mathbbm{1}\{t\leq -1\} $ is applied throughout. Estimates of \( ATT_{\omega,T} \) are interpreted as the loss in GDP per capita that would have occurred had Benin not democratised.\footnote{For results in Panel A, note that $ \displaystyle \log(Y_t(1)) - \log(Y_t(0)) = -\log\Big(\frac{Y_t(0)}{Y_t(1)}\Big) = -\log\Big(\frac{Y_t(0)}{Y_t(1)} - 1 + 1\Big) = -\log\Big(\frac{Y_t(0) - Y_t(1)}{Y_t(1)} + 1\Big) \approx - \frac{Y_t(0) - Y_t(1)}{Y_t(1)} $ for small changes.}
\begin{table}[!htbp]
\caption{The Effect of Democracy on Growth - Empirical Estimates}
{2pt}
\begin{tabular}{@rccccccccc@}
\toprule
\multicolumn{1}{l} & \multicolumn{4}{c}{Panel A} & & \multicolumn{4}{c}{Panel B}\\
\multicolumn{1}{l} & \multicolumn{4}{c}{Log GDP per capita} & & \multicolumn{4}{c}{GDP per capita}\\ \cmidrule(l){2-5} \cmidrule(r){7-10}
\multicolumn{1}{l|}{Transition Window} & \multicolumn{1}{c} & \multicolumn{1}{c} & \multicolumn{1}{c} & \multicolumn{1}{c} & $\quad$ & \multicolumn{1}{c} & \multicolumn{1}{c} & \multicolumn{1}{c} & \multicolumn{1}{c} \\
\multicolumn{1}{l|}{$(t=0)$} & \multicolumn{1}{c}{1990-1992} & \multicolumn{1}{c}{1990} & \multicolumn{1}{c}{1991} & \multicolumn{1}{c}{1992} & $\quad$ & \multicolumn{1}{c}{1990-1992} & \multicolumn{1}{c}{1990} & \multicolumn{1}{c}{1991} & \multicolumn{1}{c}{1992} \\ \midrule
\multicolumn{1}{l|}{Estimate} & & & & & & & & & \\ \cmidrule(r){1-1}
\multicolumn{1}{r|}{$ATT_{\omega,T}$}
&0.064 &0.078 &0.074 &0.059 & &46.586 &56.536 &54.101 &42.786 \\
\multicolumn{1}{r|}
&(0.017) &(0.018) &(0.019) &(0.017) & &(16.178) &(14.886) &(16.040) &(16.076) \\
\multicolumn{1}{r|}{p-val}
&0.000 &0.000 &0.000 &0.000 & &0.004 &0.000 &0.001 &0.008 \\
\multicolumn{1}{r|}{2015 BMK}
& & & & & & 4.5% &5.4% &5.2% &4.1% \\
\multicolumn{1}{r|}
& & & & & & & & & \\
\multicolumn{1}{r|}{ $X_{t-1}$}
&0.803 &0.775 &0.775 &0.802 & &0.835 &0.806 &0.805 &0.835 \\
\multicolumn{1}{r|}
&(0.050) &(0.060) &(0.060) &(0.049) & &(0.060) &(0.060) &(0.064) &(0.059) \\
\midrule
\multicolumn{1}{l|}{Diagnostic Tests} & & & & & & & & & \\ \cmidrule(r){1-1}
\multicolumn{1}{r|}{ADF p-val} &0.028 &0.025 &0.035 &0.043 & &0.023 &0.034 &0.033 &0.024 \\
\multicolumn{1}{r|}{KPSS p-val} &0.100 &0.100 &0.100 &0.100 & &0.100 &0.100 &0.100 &0.100 \\
\multicolumn{1}{r|}{DW p-val} &0.159 &0.977 &0.798 &0.153 & &0.432 &0.822 &0.906 &0.441 \\
\bottomrule
\end{tabular}
\begin{justify}
{
\textit{Notes:} GDP per capita is measured in constant 2015 US dollars. All estimates are based on quasi-differenced $X_t$ following (ref) in the supplement. All standard errors (in parentheses) are heteroskedasticity and auto-correlation robust using the newey-west-1987-simple procedure. ADF is the Augmented Dickey-Fuller test of a unit root null hypothesis. KPSS denotes the Kwiatkowski-Phillips-Schmidt-Shin (KPSS) test of a trend stationarity null hypothesis. DW is the Durbin-Watson test of a zero auto-correlation null hypothesis against the two-sided alternative.}
\end{justify}
\end{table}
\begin{table}[!htbp]
\caption{Empirical Estimates with Transition Window: 1990-1992 }
{2pt}
\begin{tabular}{@rccccccccc@}
\toprule
\multicolumn{1}{l} & \multicolumn{4}{c}{Panel A} & & \multicolumn{4}{c}{Panel B}\\
\multicolumn{1}{l} & \multicolumn{4}{c}{Log GDP per capita} & & \multicolumn{4}{c}{GDP per capita}\\ \cmidrule(l){2-5} \cmidrule(r){7-10}
\multicolumn{1}{l|}{Transition Window} & \multicolumn{1}{c} & \multicolumn{1}{c} & \multicolumn{1}{c} & \multicolumn{1}{c} & $\quad$ & \multicolumn{1}{c} & \multicolumn{1}{c} & \multicolumn{1}{c} & \multicolumn{1}{c} \\
\multicolumn{1}{l|} & \multicolumn{1}{c}{ T-DiD$^*$ } & \multicolumn{1}{c}{S-DiD} & \multicolumn{1}{c}{SC} & \multicolumn{1}{c}{ASC} & $\quad$ & \multicolumn{1}{c}{ T-DiD$^*$ } & \multicolumn{1}{c}{S-DiD} & \multicolumn{1}{c}{SC} & \multicolumn{1}{c}{ASC} \\ \midrule
\multicolumn{1}{l|}{Estimate} & & & & & & & & & \\ \cmidrule(r){1-1}
\multicolumn{1}{r|}{$ATT_{\omega,T}$}
&0.042 &0.249 &0.090 &0.405 & &32.493 &101.582 &-43.968 &254.327 \\
\multicolumn{1}{r|}
&(0.011) &(0.282) & & & &(10.223) &(1519.637) & & \\
\multicolumn{1}{r|}{p-val}
&0.000 & &0.700 &0.417 & &0.001 & &0.800 &0.316 \\
\multicolumn{1}{r|}{2015 BMK}
& & & & & &3.1% &9.8% &-4.2% &24.4% \\ \midrule
\multicolumn{1}{l|}{Identification test}
& & & & & & & & & \\ \cmidrule(r){1-1}
\multicolumn{1}{r|}{p-val}&0.882 & & & & &0.938 & & & \\
\bottomrule
\end{tabular}
\begin{justify}
{
\textit{Notes:} Methods compared to the T-DiD include the Synthetic difference-in-differences (S-DiD) of arkhangelsky-etal-2021, the Synthetic Control (SC) of abadie-gardeazabal-2003, and the Augmented Synthetic Control (ASC) of ben-feller-rothstein-2021. The S-DiD, SC, and ASC p-values are computed from permutations. The transition window of 1990-1992 is maintained for all results above. T-DiD uses Togo and Cameroon as controls following (ref). }
\end{justify}
\end{table}
\begin{table}[htbp]
\caption{ Synthetic Control Weights with Transition Window: 1990-1992 }
\begin{tabular}{l|ccc|ccc}
\toprule
& \multicolumn{3}{c|}{Log GDP per capita} & \multicolumn{3}{c}{Level GDP per capita} \\ \cmidrule(l){2-4} \cmidrule(r){5-7}
Country & S-DiD & SC & ASC & S-DiD & SC & ASC \\ \midrule
Algeria & 0.056 & 0.035 & 0.039 & 0.038 & 0.020 & 0.014 \\
Burundi & 0.068 & 0.068 & 0.222 & 0.060 & 0.039 & 0.304 \\
Cameroon & 0.115 & 0.048 & 0.180 & 0.057 & 0.029 & 0.127 \\
Central African Republic & 0.060 & 0.059 & 0.032 & 0.061 & 0.034 & -0.001 \\
Chad & 0.057 & 0.057 & -0.009 & 0.064 & 0.034 & -0.005 \\
Gabon & 0.000 & 0.023 & -0.033 & 0.000 & 0.011 & -0.005 \\
Kenya & 0.044 & 0.050 & -0.021 & 0.054 & 0.030 & 0.000 \\
Libya & 0.000 & 0.016 & -0.016 & 0.000 & 0.007 & 0.001 \\
Madagascar & 0.054 & 0.056 & 0.006 & 0.063 & 0.033 & 0.030 \\
Malawi & 0.020 & 0.067 & -0.022 & 0.059 & 0.037 & -0.002 \\
Nigeria & 0.000 & 0.043 & -0.010 & 0.048 & 0.026 & -0.001 \\
Rwanda & 0.075 & 0.067 & 0.051 & 0.060 & 0.038 & 0.129 \\
Seychelles & 0.034 & 0.027 & -0.039 & 0.013 & 0.015 & -0.005 \\
Sierra Leone & 0.079 & 0.050 & 0.270 & 0.058 & 0.030 & 0.065 \\
Somalia & 0.020 & 0.071 & -0.041 & 0.061 & 0.458 & -0.003 \\
Sudan & 0.012 & 0.055 & -0.101 & 0.060 & 0.033 & -0.010 \\
Tanzania & 0.108 & 0.057 & 0.179 & 0.068 & 0.033 & 0.148 \\
Togo & 0.060 & 0.056 & 0.048 & 0.058 & 0.033 & 0.026 \\
Zambia & 0.067 & 0.049 & 0.082 & 0.063 & 0.030 & 0.083 \\
Zimbabwe & 0.071 & 0.045 & 0.183 & 0.055 & 0.028 & 0.106 \\
\midrule
\# Controls & 17 & 20 & 11 & 18 & 20 & 11 \\
\bottomrule
\end{tabular}
\begin{justify}
{Synthetic Control Weights: S-DiD, SC, and ASC estimators for Log and Level GDP per capita. Entries are reported to three decimal places. The last row counts the number of countries with non-zero weights. Estimated non-zero S-DiD pre-treatment period weights are highly concentrated in a small subset of years. For log GDP, the estimated lambda weights are 0.045 in 1960, 0.029 in 1963, 0.012 in 1978, 0.011 in 1979, and 0.903 in 1989. For GDP in levels, the only non-zero pre-treatment weight is 1989, with estimated weight 1.0. In both specifications, all other pre-treatment years receive zero estimated weight.}
\end{justify}
\end{table}
From Panel A, one observes a fairly narrow range of the $ATT_{\omega,T}$ estimates by transition window. One observes that had Benin not democratised, her economy would have been $5.9\% - 7.8\% $ smaller on average over the period spanning the early 1990s up to 2018. To have a sense of the economic significance of these effects, consider comparisons with the average World GDP per capita growth rate over the 1991-2018 period.\footnote{The average World GDP per capita growth rate over the 1991-2018 period is 3.597% -- source: authors' calculation using World Bank data sourced from \url{www.macrotrends.net}.} The $ATT_{\omega,T}$ estimates are economically significant and interesting -- Benin's economy would have, on average over the period spanning the early 1990s up to 2018, “shrunk" by a factor of $1.6 - 2.2$ times the average World GDP per capita growth rate over the 1991-2018 period had she not democratised. Moreover, the estimates are statistically significant at all conventional levels.
Panel B uses Benin's GDP per capita (in constant 2015 US dollars) of $ 1041.653 $ as a benchmark for interpreting results on the Level GDP per capita.\footnote{The corresponding row is labelled “2015 BMK".} Thus, one observes that Benin's annual average GDP per capita over the period spanning the early 1990s through 2018 as a percentage of her $2015$ GDP per capita (in 2015 constant USD) would have been $4.1\% - 5.4\% $ smaller had she not democratised. These values are economically significant and interesting as the margin is economically non-trivial.\footnote{Level $ATT_{\omega,T}$ estimates are expressed in 2015 USD in Benin and hence are generally dependent on 2015 price levels and exchange rates in Benin. To further aid interpretation, taking the 2015 ratio of Benin's Purchasing Power Parity (PPP) adjusted GDP per capita in constant 2021 USD to the GDP per capita in constant 2015 USD of $2.9$ (source: World Bank Indicators), assuming this ratio is stable over the post-treatment period, the 2015 constant USD PPP-adjusted $ATT_{\omega,T}$ estimates of Panel B read: $ 131.9819, \ 163.4295, \ 155.8257, \text{ and } 120.6922 $.} Moreover, all $ATT_{\omega,T}$ level estimates are statistically significant at all conventional levels. Thus, across different specifications of the transition window, the effects of democratisation on economic growth are positive and statistically significant. That the range of estimates $4.1\% - 5.4\% $ is narrow indicates the robustness of the results to the specification of the transition window.
The estimates on $X_{t-1}$ are indicative of a high level of persistence in the series $X_t$ which is the difference between the annual GDP per capita of Benin and Togo. The Augmented Dickey-Fuller (ADF) test suggests the unit-root null hypothesis is rejected at the $ 5\%$ level across all specifications in both panels A and B. The Kwiatkowski-Phillips-Schmidt-Shin (KPSS) test of trend stationarity suggests a failure to detect trend non-stationarity at the $5\%$ level, while the Durbin-Watson test indicates a failure to reject the null of zero auto-correlation (against the two-sided alternative) at any conventional nominal level.\footnote{The DW test also serves as a specification test in this setting. Suppose the true model had an MA(1) error instead, i.e., $ X_t = \beta_0 + \widetilde{ATT}_n^{w,\psi} \mathbbm{1}\{t \geq 1\} + \mathcal{E}_t + \beta_1 \mathcal{E}_{t-1} $, then the resulting errors from the “mis-specified" model (with the lagged $X_t$ above ) would be auto-correlated. This is the violation the DW test is apt to detect.} Taken together, the diagnostic tests suggest the reliability of the T-DiD results in (ref).
To compare the proposed \( T \)-DiD estimator with alternative methods in this empirical setting, the S-DID, SC, and ASC results reported in (ref) use a donor pool comprising \emph{all} African countries \emph{without} missing constant 2015 USD GDP per capita data, whose average democracy index remained below the 0.5 threshold throughout the 1960--2018 period. The resulting donor pool consists of 20 countries, along with the corresponding SC weights, as detailed in (ref). 1990-1992 is maintained as the transition window. The efficient T-DiD following (ref) with Cameroon and Togo as controls is compared with the Synthetic DiD of arkhangelsky-etal-2021, the Synthetic Control method of abadie-gardeazabal-2003, and the Augmented Synthetic Control method of ben-feller-rothstein-2021. Relative to the T-DiD, these estimates are not precisely estimated. This is not surprising as the minimum attainable p-value (or test size under the null) in a permutation test applied to the SC and ASC is $1/21 \approx 0.0476 $.\footnote{Besides, the discussion in (ref) suggests that the SC estimators without adjusting for auto-correlation in the outcome and its fitted counterfactual may be targeting the long-run (multiplier) effect instead.}
\section{Conclusion}
While being instrumental in answering the empirical question at hand, the T-DiD estimator equally contributes to the larger applied and theoretical literature on causal inference in large-$\mathcal{T},T$ settings with a fixed number of cross-sectional units. The $ATT_{\omega,T}$ parameter of interest indexed over a space of convex weighting schemes is asymptotically identified under identification conditions where even both asymptotic parallel trends and limited anticipation identification conditions may be individually violated. Under NED—a fairly general form of temporal dependence—and weak dominance regularity conditions, the T-DiD estimator is asymptotically normal, thereby paving the way for standard inference using, e.g., $t$-tests.
This paper further extends the baseline theory to accommodate possible intricacies underpinned by stochastic and deterministic trends in $X_t:= Y_{1,t} - Y_{0,t} $ and multiple control units. This paper exploits over-identification in the multiple-control setting to propose an identification test with more desirable statistical properties relative to the pre-trends test. For example, the proposed identification test can detect violations of identification in post-treatment periods, unlike pre-trends tests.
Sustained interest in the economic gains of democracy spawns a vast array of interesting contributions that fail to reach a clear consensus. This issue is largely due to the immense heterogeneity in effects, which renders the generalisability of results quite difficult. This paper departs from the extant literature by adopting a novel perspective -- estimating country-specific effects. This paper equally meets the methodological challenge that this new approach engenders --- proposing the T-DiD, which is adaptable to as few as a single treated unit and a single control unit with large pre- and post-treatment periods. The estimated effect of democracy on economic growth in this paper is not only statistically significant at all conventional levels; it is also economically interesting. Benin's economy would have been $5.9\%-7.8\% $ smaller on average had she not democratised.
\section*{Acknowledgements}
This paper benefited from the invaluable feedback of Al-mouksit Akim, Brantly Callaway, Alecia Cassidy, Traviss Cassidy, Daniel Henderson, Jonathan Hall, Feiyu Jiang, Eric Kadio, D\'esir\'e K\'edagni, Junsoo Lee, Xiaochun Liu, Tigran Melkonyan, and Julius Owusu. This manuscript benefited from the use of ChatGPT (OpenAI, GPT-5.3) for language editing and presentation.
\section*{Declaration of Interests}
The authors have no interests to declare.
\section*{Data Availability Statement}
Data used in the study are publicly available. The replication package, including the dataset, is available on author E.S.T's website.
\begingroup
\setstretch{1.25}
\setlength\bibitemsep{0.5 pt}
\printbibliography
\endgroup
\setcounter{page}{1}