Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
71,180 characters · 17 sections · 65 citation commands
Testing for equivalence of pre-trends in Difference-in-Differences estimation
\def\spacingset#1{ {#1}} \spacingset{1}
\if00 \fi
\if10 {
} \fi
\spacingset{1.45}
In the classic case, the Difference-in-Differences (DiD) framework consists of two groups observed over two periods of time, where the “treatment group” is untreated in the initial period and has received a treatment in the second period whereas the “control group” is untreated in both periods. The key condition under which the DiD estimator yields sensible point estimates of the average treatment effect on the treated is known as the “parallel paths” or “parallel trends assumption”, henceforth referred to as PTA, which states that in the absence of treatment both groups would have experienced the same temporal trends in the outcome variable on average. If pre-treatment observations are available for both groups, the plausibility of this assumption is typically assessed by plots accompanied by a formal testing procedure showing that there is no evidence of differences in trends over time between the treatment and the control group. However, traditional pre-tests can suffer from low power to detect violations of the PTA kahn2020promise. Thus, finding no evidence of differences in trends in finite samples does not imply that there are no differences in trends in the population. More concerningly, roth2018pre points out that if differences in trends exist, conditional on not detecting violations of parallel trends at the pre-testing stage, the bias of DiD-estimators may be greatly amplified.
Given the severe consequences of falsely accepting the PTA, we propose that instead of testing the null hypothesis of "no differences in trends" between the treatment and the control group in the pre-treatments periods, one should apply a test for statistical equivalence. We provide three distinct types of equivalence that impose bounds on the maximum, the average and the root mean square change over time in the group mean difference between treatment and control in the pre-treatment periods. Given a threshold below which deviations from the PTA can be considered negligible, these tests allow the researcher to provide statistical evidence in favor of the PTA, thus increasing its credibility. If no sensible equivalence threshold can be determined before analyzing the data, we propose to report the smallest equivalence threshold for which the null hypothesis of "non-negligible trend differences" can still be rejected at a given level of significance. Conceptually, this idea is similar to the “equivalence confidence interval” in hartman2018 applied to a DiD setting. Our procedure reverses the burden of proof since the data has to provide evidence in favor of similar trends in the treatment and the control group, which is arguably more appropriate for an assumption as crucial to the DiD-framework as the PTA. Furthermore, the power to reject the null hypothesis of “non-negligible differences” is increasing with the sample size (also see hartman2018). This improves upon the current practice of testing the null hypothesis of “no difference”, since large samples increase the chances of rejecting this null hypothesis (and thus seemingly making the DiD framework inapplicable), even if the true difference between treatment and control may be negligible in the given context. Finally, our equivalence test statistics can easily be implemented in practice. While we motivate our tests in the standard two-way fixed effects model with panel data, we discuss how our tests can be applied in situations where treatment timing differs across groups (i.e.\ “staggered treatment assignment”) or where average treatment effects depend on some observable characteristics.
As we use equivalence tests, our paper is closely related to bilinski2020nothing, who provide a discussion on the benefits of using equivalence (or "non-inferiority") tests when testing for violations of modeling assumptions. Their “one-step-up” approach is based on a non-inferiority test of treatment effect estimates obtained from a standard DiD model and from a model augmented with a particular violation of the parallel trends assumption (e.g.\ a linear trend). While both approaches stress the potential benefits of equivalence testing in DiD setups, a distinctive feature is that we do not necessarily focus on a particular violation of the PTA. As pointed out in kahn2020promise, including for instance group-specific linear time trends can lead to a loss in degrees of freedom and thus to a substantial loss in power. In contrast, our approach focuses on testing for any non-negligible differences between treatment and control in the pre-treatment periods. Our paper is also related to other approaches that allow for certain deviations from exactly parallel trends. roth2020honest partially identify the ATT under restrictions that impose that the post-treatment violation of parallel trends is not too large relative to the pre-treatment violation. By contrast, we propose tests for the null hypothesis that the pre-treatment violation of parallel trends is large. If the null is rejected, so that the pre-treatment violation is determined to be small, then the researcher may decide that violations of parallel trends can safely be ignored with minimal bias. Alternatively, the researcher could use the upper bound on the pre-treatment violation given by our approach as a way of determining reasonable bounds on the post-treatment violations of parallel trends, which could then be used as input to the partial identification frameworks of roth2020honest or manski2018right.
The rest of the paper is organized as follows. Section 2 introduces the main two-way fixed effects model and discusses the widely used practice of testing for violations of the PTA. Equivalence tests and our hypotheses are discussed in Section 3. Section 4 introduces our assumptions and presents the test statistics for our hypotheses. In Section 5, we present extensions of the main model that allow for heterogeneous treatment effects due to differences in treatment timing or observable characteristics. Simulation evidence on the performance of our test procedures is provided in Section 6, while Section 7 contains an empirical illustration of our approach. Section 8 concludes. Finally, mathematical details and tables are collected in the Appendix.
We initially contrast our test procedures with the usual pre-test in the canonical DiD case with only two groups and common treatment timing. The researcher observes a balanced panel of $n$ individuals recorded over $T+2$ periods of time. We assume that in the first $T+1$ periods none of the individuals is treated whereas in period $T+2$ a subset has received treatment. Individual $i$ is a member of the treatment group if $G_i=1$ and a member of the control group if $G_i=0$. Following kahn2020promise we refer to period $T+1$ as the “base period” as treatment effects are usually assessed by comparing post-treatment outcomes with the outcomes in the last pre-treatment period. The potential outcomes of unit $i$ in period $t$ when treated and in the absence of treatment are denoted as $Y_{i,t}(1)$ and $Y_{i,t}(0)$, respectively. The object of interest is the average treatment effect on the treated (ATT), given as
For identification of the ATT, we need to make several assumptions. First, we require that $\Pr(G_i=1)=p\in (0,1)$, which is an “overlap” condition that ensures that both treatment and control are non-empty (santanna-zhao-2020). Next, we assume "no-anticipation", which rules out treatment effects before the actual treatment date. Adapting borusyak, we assume
with $D_{l}(t) =\mathbbm{1}\{l=t\}$. This implies that the expected observed outcome coincides with the expected potential outcome corresponding to the actual treatment status both for the control units and the eventually treated units. Next, let
where $\alpha_i, \lambda_t$ and $\gamma_t$ are some non-stochastic constants. Combining (ref) with (ref), we can write
with $u_{i,t}=Y_{i,t}-\mathbb{E}[Y_{i,t}|G_i]$. It is clear from (ref) that $\pi_{ATT}$ is not identified without further restrictions on $\gamma_t$. The fundamental assumption that leads to the DiD estimator is the (augmented) PTA
which implies that $\gamma_t-\gamma_{T+1}=0$ for all $t=1,...,T+2$. Notice that this assumption over-identifies $\pi_{ATT}$, as identification only requires parallel trends between the post-treatment and the base period, i.e.\ $\gamma_{T+2}-\gamma_{T+1}=0$. However, as it is often difficult to imagine circumstances under which the latter condition is satisfied whereas (ref) is not, the augmented PTA can be useful for a pre-testing procedure that allows researchers to assess the plausibility of the PTA post-treatment. To do so, one uses the data from periods $1,...,T+1$ to estimate the two-way fixed effects (TWFE) model
where $\beta_l=\gamma_l-\gamma_{T+1}$. Here, $\beta=(\beta_1,...,\beta_T)'$ collects all “placebo” treatment effects so that under (ref) and (ref), $\beta=0$ if and only if the augmented PTA holds. As shown, for instance, in Baltagi2021 or wooldridge2021two, $\beta$ can be estimated by pooled OLS on “double-demeaned” data. Let $W_{i,t,l}=G_i D_l(t)$ and $W_{i,t}=(W_{i,t,1},...,W_{i,t,T})'$. Double-demeaning (ref) then yields
for $ i=1,...,n$ and $t=1,...,T+1$, where $\ddot{Y}_{i,t}=Y_{i,t}-\bar{Y}_{i,\cdot}-\bar{Y}_{\cdot, t}+\bar{Y}_{\cdot\cdot}$, \[ \bar{Y}_{i,\cdot}=\frac{1}{T+1}\sum_{t=1}^{T+1}Y_{i,t},~ \bar{Y}_{\cdot, t}=\frac{1}{n}\sum_{i=1}^nY_{i,t},~ \bar{Y}_{\cdot\cdot}=\frac{1}{n(T+1)}\sum_{i=1}^n\sum_{t=1}^{T+1}Y_{i,t} \] and $\ddot{W}_{i,t}$ and $\ddot{u}_{i,t}$ are defined analogously. Since $\mathbb{E}[u_{i,t}|G_i]=0$ by construction, consistency and asymptotic normality of the pooled OLS estimator in (ref), denoted as $\hat{\beta}$, follow under mild conditions. To find evidence against the plausibility of parallel trends, it is therefore common in applied economic research to test for individual significance (see roth2018pre), i.e.\ for every $l\in\{1,...,T\}$ we test
If the null hypothesis is rejected in a pre-treatment period, the PTA is deemed unreasonable, and consequently the DiD framework is often regarded as unsuitable in the corresponding context. This procedure has several shortcomings. For instance, a problematic common practice is to treat failure to reject the null hypothesis in (ref) as evidence in favor of $\mathrm{H}_0$, i.e.\ one proceeds as if the null hypothesis was true and as if the PTA held. From a statistical point of view, this practice is incorrect as it neglects the error of type II. In some cases, there may be differences in trends between both groups in the population that cannot be detected with traditional test of (ref) due to a lack of statistical power. roth2018pre points out that ignoring these differences can amplify the bias and thus raise concerns of a “publication bias”, since articles using a DiD identification argument are more likely to be deemed publishable when a test of (ref) could not detect evidence against the PTA. Moreover, the DiD framework is sometimes used even when $\mathrm{H}_0$ in (ref) is rejected, as some statistically significant differences are deemed negligible in a given context. However, a potential threshold $\mathcal{U}>0$ that quantifies what constitutes a negligible difference is often not adequately discussed. On the other hand, if the DiD framework is not applied when $\mathrm{H}_0$ in (ref) is rejected in at least one pre-treatment period, useful information may be lost if $\mathcal{U}$ can be interpreted as a plausible “upper bound” for trend differences. For these reasons, the plausibility of the PTA as the fundamental modeling assumption of the DiD framework can be more convincingly assessed using statistical equivalence tests as these tests address all of the above shortcomings of the current standard testing procedure. For instance, to rewrite (ref) in terms of statistical equivalence, for some $l\in\{1,...,T\}$ one would define the equivalence threshold $\mathcal{U}>0$ and test
so that rejecting $\mathrm{H}_0$ yields evidence in favor of negligible trend differences between periods $l$ and $T+1$. While the hypothesis (ref) can easily be tested with the "two one-sided tests" procedure of schuirmann1987, we do not recommend this approach as it would lead to an accumulation of Type I error due to multiple testing. In the following, we elaborate on the benefits of equivalence tests and provide ways of summarizing the statistical evidence in favor of the PTA in the pre-treatment periods by formulating joint hypotheses.
Equivalence testing is well known in biostatistics (see berger1996bioequivalence or wellek2010). While it has recently been considered in the statistical literature for the analysis of structural breaks (e.g.\ Dette2014b, dettewu2019, detkokaue2020 or DKV2018), it is less frequently used in econometrics. Instead of assuming that treatment and control are perfectly comparable unless there is strong evidence against this assumption, we suggest several testing procedures that explicitly require finding evidence in favor of the comparability of both groups. Each of the tests is based on an upper bound $\mathcal{U}> 0$ for changes in the group mean differences in the pre-treatment periods relative to the base period. There are two ways in which one can make use of the upper bound $\mathcal{U}$. First, as in the “classic” use of equivalence tests, one can specify a threshold $\mathcal{U}$ below which changes in the group mean differences over time are deemed negligible. Rejecting the null hypothesis that trend differences are larger than $\mathcal{U}$ at level of significance $\alpha$ then implies that deviations from parallel trends in the pre-treatment periods are negligible at confidence level $1-\alpha$. The researcher may then interpret this as support for negligible differences in trends post-treatment and hence assume that the PTA holds. This procedure improves upon the current pre-test as it requires an explicit rationalization of the threshold $\mathcal{U}$ and sufficient data to support the assumption of negligible violations of the PTA. The choice of the threshold $\mathcal{U}$ should thus reflect the specific scientific background of the application. In bio-statistics, the popularity of equivalence tests has led to a consensus on sensible choices for $\mathcal{U}$, and regulators frequently specify the equivalence thresholds that should be employed (see wellek2010 for a recent review). We expect that with a more frequent adoption of equivalence testing in applied economics a similar consensus will be reached. However, in some applications, it may still be difficult to objectively argue that a certain extent of violations of the PTA can be ignored in practice. It may then be sensible to report $\mathcal{U}^*$ as the smallest value at which $\mathrm{H}_0$ can be rejected at a given level of significance (i.e.\ for which “equivalence of pre-trends" can be concluded). Small values of $\mathcal{U}^*$ relative to the estimated treatment effect may then be regarded as reassuring as it is unlikely that the treatment effect is merely an artifact of differences in trends. On the contrary, if $\mathcal{U}^*$ is relatively large, the credibility of the estimated effect is in serious doubt. A similar idea has been proposed in hartman2018. It can further be related to the "breakdown frontier" proposed by Masten2020, as treatment effects can be considered non-robust to violations of the PTA when $\mathcal{U}^*$ exceeds the estimated treatment effect. Finally, in cases where the choice of the threshold is difficult, the methodology presented here can also be used to provide (asymptotic) confidence intervals for violations of the PTA.
We assume that there exists a vector $\beta\in\mathbb{R}^T$ that can be used to assess the plausibility of the augmented PTA. For instance, in (ref), each $\beta_l$ corresponds to a "placebo" treatment effect in period $l\in\{1,...,T\}$. To keep the notation tractable, the dimension of $\beta$ corresponds to the number of available pre-treatment periods $T$. In practice, it is of course possible that the dimension of $\beta$ is smaller than $T$ (e.g.\ when only a subset of all pre-treatment period is used for a pre-test) or exceeds $T$ (e.g.\ when a conditional PTA is tested; see Section (ref)). Our tests can then be applied with minor adjustments.
Overall, we consider three distinct hypotheses to test for equivalence of pre-trends in treatment and control. We start with a discussion of the maximum placebo treatment effect. For $\beta\in\mathbb{R}^T$, level of significance $\alpha$ and the equivalence threshold $\delta>0$, we test
where $\|\beta\|_{\infty}\vcentcolon= \max_{l\in\{1,...,T\}}|\beta_l|$. Since we are now controlling the type I error, this implies that with confidence level of at least $1-\alpha$, $\delta$ is an upper bound for the maximum placebo treatment effect.
In many applications, pre- and post-treatment periods are pooled, for instance to increase statistical power. Similarly, it may be sensible in some applications to consider a pooled or average measure of the pre-treatment deviations from parallel trends. Thus, defining $\bar{\beta}\vcentcolon= \frac{1}{T}\sum_{l=1}^{T}\beta_l$, one can find bounds on the average placebo effect by testing
One disadvantage of (ref) is that there may be cancellation effects in situations where the components of $\beta$ are large in absolute terms but have opposing signs. Therefore, (ref) should be used when differences in pre-trends can safely assumed to be of the same sign. As pointed out in roth2020honest, monotone violations of the PTA are frequently discussed in the applied literature. For instance, treatment effect estimates are often considered robust if potential violations of the PTA are of the opposing sign and can thus be ruled out as an explanation for the estimated effects. As an alternative to (ref) that does not suffer from potential cancellation effects, we further consider the root mean square (RMS) of $\beta$, i.e. $\beta_{RMS}\vcentcolon= \|\beta\|/\sqrt{T}= \sqrt{\frac{1}{T}\sum_{l=1}^{T}\beta_l^2}$, where $\|\cdot\|$ denotes the euclidean norm on $\mathbb{R}^{T}$. The RMS of $\beta$ can thus be interpreted as the euclidean distance between treatment and control in the pre-treatment periods relative to the distance in the base period scaled by the number of pre-treatment periods. The scaling is induced to ensure that this distance between treatment and control does not increase with the number of pre-treatment periods available. The hypothesis is then formulated as
which can equivalently be written as
In Section (ref) below we develop a test statistic for (ref) and recover $\zeta$ as $\sqrt{\zeta^2}$.
We begin this section by introducing the assumptions underlying our tests. We then proceed by discussing the implementation and the statistical properties of our tests.
For our statistical theory, we mainly need that a “sequential” version of asymptotic normality holds:
Notice that (ref) comprises “conventional" asymptotic normality of $\hat{\beta}$ when $\lambda=1$, i.e.
where $\Sigma$ denotes a $T\times T$-dimensional covariance matrix. A further assumption required in some of our results is that this matrix can be consistently estimated:
The proofs of the validity of the first test for the hypothesis (ref) and of the test for the hypothesis (ref) do not require Assumption (ref) but only the weaker condition (ref) together with Assumption (ref) (see the first part of Section (ref) and Section (ref) below for the exact definition of these two tests). A second more powerful test for the hypothesis (ref) is introduced in the second part of Section (ref), and its validity will be established in the TWFE model (ref), under conditions that imply (ref). In contrast, our test for the hypothesis (ref) requires (ref), but Assumption (ref) is not needed, as this test is based on “self-normalization”. We emphasize that without any additional assumptions the asymptotic normality in (ref) does not imply the process convergence in (ref). In fact, this a very delicate probabilistic question, which has, to our best knowledge, only been solved for sums of random variables. For example, we refer to kuelbs1973 for the independent case and to samur1987 for the dependent case (note that these authors consider Banach space valued random variables, which contains the case of finite dimensional vectors).
Sequential asymptotic normality as in Assumption (ref) can be shown to hold for stationary processes under various forms of dependence such as mixing or physical dependence (see merlevede2006 and the references therein). Moreover, our tests can be flexibly adapted to different types of data (e.g.\ panel data or repeated cross-sections) and other features of the design such as clustering or staggered treatment assignment. As an illustration, in Appendix (ref), we verify that Assumption (ref) holds under mild conditions in the canonical DiD model of Section (ref).
Notice that our approach is a test of the PTA conditional on additional parametric restrictions (for instance, (ref) and (ref), or parametric restriction on the effect of observed covariates). While arguably the vast majority of empirical studies are willing to impose similar assumptions, one might want to test the PTA without these restrictions. In this case, as an alternative to our method, one could consider tests based on conditional moment inequalities (see, among others, andrews-soares-2010, romano-shaikh-wolf-2014 or canay-shaikh-2017).
Based on Assumptions (ref) and (ref), we now derive test statistics for our equivalence hypotheses. We further analyze the statistical properties of the resulting test procedures.
To describe the first test for the hypothesis (ref) we initially consider the case $T=1$ so that our objective is to test whether a single parameter $\beta_1$ exceeds a certain threshold. As $\hat{\beta}_1$ is approximately distributed as $\mathrm{N}(\beta_1,\Sigma_{11}/n)$, the test statistic $|\hat{\beta}_1|$ approximately follows a folded normal distribution. We therefore propose to reject the null hypothesis in (ref) whenever
where $\mathcal{Q}_{\mathrm{N}_F(\delta,\sigma^{2})}(\alpha)$ denotes the $\alpha$ quantile of the folded normal distribution with mean $\delta$ and variance $\sigma^{2}$. It is shown in Appendix (ref) that this test is consistent, has asymptotic level $\alpha$ and is (asymptotically) uniformly most powerful for testing the hypothesis (ref) in the case $T=1$. In particular this test is more powerful than the two-sided $t$-test (TOST), which could be developed following the arguments in hartman2018. For $T>1$, we apply the idea of intersection-union (IU) tests outlined in berger1996bioequivalence and reject the null hypothesis in (ref), whenever
While this test is computationally attractive, a well-known disadvantage of testing procedures based on the IU principle is that they tend to be rather conservative berger1996bioequivalence, which is confirmed by our simulation study (see Table (ref)).
Specifically for the TWFE model (ref), it is possible to obtain a more powerful test for the hypothesis (ref) as follows: In the first step, estimate $\beta$ in (ref) to obtain the unconstrained TWFE estimator $\hat{\beta}_u$. In the second step, re-estimate (ref) by minimizing the sum of squared residuals under the constraint $ \| \beta \|_\infty = \max_{l=1,...,T}|\beta_l| = \delta$ to obtain a constrained estimator, say $\hat{\beta}_c$. We then define new estimators of the parameters as
and
Note that the vector ${\hat{\hat{\beta}}_c}$ satisfies the null hypothesis in (ref). In the third step, for $b=1,...,B\in\mathbb{N}$, we generate bootstrap samples with $\ddot{u}_{i,t}^{(b)}\overset{\mathrm{i.i.d.}}{\sim} \mathrm{N} (0,\hat{\hat{\sigma}}_c)$ and $\ddot{Y}_{i,t}^{(b)}=\ddot{W}_{i,t}'\hat{\hat \beta}_c+\ddot{u}_{i,t}^{(b)}$ for $i=1,...,n$ and $t=1,...,T+1$. For each bootstrap sample, estimate $\hat{\beta}^{(b)}$ and compute $\mathcal{Q}_\alpha^*$ as the empirical $\alpha$-quantile of the bootstrap sample $\{\max_{l=1,...,T}|\hat{\beta}_l^{(b)}| : b=1,...,B\}$. Finally, reject the null hypothesis $\mathrm{H}_0$ in (ref) if
Notice that the condition $ \| \beta \|_\infty = \delta$ in the calculation of the constrained estimator $\hat \beta_c$ is made for technical reasons in the proof of this statement. Numerical results show that the test has a very similar behavior if $\hat \beta_c$ is calculated under the condition that $ \| \beta \|_\infty \geq \delta$. The following result shows that this test is consistent and has asymptotic level $\alpha$.
For some fixed $\tau>0$, a test can be constructed by first computing the statistic \[ \bar{\hat{\beta}}\vcentcolon=\frac{1}{T}\sum_{t=1}^{T}\hat{\beta}_t = \mathds{1}'\hat{\beta}/T, \] where $\mathds{1} = (1,\ldots , 1 )' \in \mathbb{R}^{T}$. Note that it follows from Assumptions (ref) and (ref) that $$\sqrt{n} \mathds{1}'({\hat \beta} - \beta) \to \mathrm{N}(0, \mathds{1}' \Sigma \mathds{1})).$$ Consequently, based on the discussion in the first part of (ref), we propose to reject the null hypothesis in (ref), whenever
where $\hat{\sigma}^{2}=\mathds{1}'{\hat{\Sigma}} \mathds{1}$/n, and $\hat{\Sigma}$ is a consistent estimator of $\Sigma$ which is known by Assumption (ref).
For $\lambda\in [\varepsilon,1]$, let $\hat{\beta}_{RMS}(\lambda)$ denote the RMS of $\hat{\beta}$ based on a random subsample of size $\lfloor \lambda n\rfloor$ of the original sample of $n$ individuals. In order to construct a pivot test for the hypothesis (ref), define
where
and $\nu$ denotes a measure on the interval $[\varepsilon,1]$. The following result is proved in the Appendix.
It follows from the proof of Theorem (ref) that the statistic $\hat{\beta}_{RMS}^2$ is a consistent estimator of $\beta_{RMS}^2$. Therefore, we propose to reject the null hypothesis $\mathrm{H}_0$ in (ref) (and consequently $\mathrm{H}_0$ in (ref)), whenever
where $\mathcal{Q}_{\mathbb{W}}(\alpha)$ is the $\alpha$-quantile of the limiting distribution of the random variable $\mathbb{W}$ on the right-hand side of (ref). Note that these quantiles can be easily obtained by simulation because the distribution of $\mathbb{W}$ is completely known. For instance, $\mathcal{Q}_{\mathbb{W}}(0.05)\approx -2.1$. The following result shows that this decision rule defines a valid test for the hypothesis (ref).
The use of the simple TWFE model has recently experienced substantial criticism in the presence of multiple groups, heterogeneous treatment effects and differences in treatment timing. In this situation, the TWFE estimator often does not correspond to a reasonable estimate of the ATT, and alternative estimators have been proposed by several authors (see, for instance, goodman2018difference, callaway2020difference, abraham2018estimating, borusyak or chaisemartin2020. Excellent reviews of this fast-growing literature are provided by roth-survey and chaisemartinDiD.). Specifically, wooldridge2021two shows that the deficiency of the TWFE estimator can be regarded as a model misspecification problem. He then proposes model adjustments that allow for treatment effect heterogeneity due to differences in treatment timing and observed characteristics (which are assumed to be unaffected by the treatment). While we conjecture that our equivalence tests can be used with most (if not all) of the mentioned estimators, we focus on the regression-based approach of wooldridge2021two in the following, as it is straightforward to show that our assumptions hold with minor adjustments to the arguments in Appendix (ref). We now consider the case of the staggered adoption (i.e.\ the initial treatment period varies across groups) of an absorbing treatment (i.e.\ the treatment status does not change after the initial treatment) in the presence of a never-treated group. Following wooldridge2021two, we assume that the time since the initial treatment adoption produces different levels of exposure to the treatment, resulting in treatment effect heterogeneity across time. As in Section (ref), we consider a balanced panel of $n$ individuals that are observed in $T+1$ pre-treatment periods. In periods $T+2,...,\overline{T}$, a subset of individuals adopts treatment, leading to “treatment cohorts”. As before, period $T+1$ is used as the base period. To define a treatment cohort dummy, let $G_i^r=1$ if individual $i$ has first adopted treatment in period $r\in\mathcal{R}\vcentcolon=\{T+2,...,\overline{T},\infty\}$ and zero otherwise, where $G_i^\infty$ is a dummy indicating that individual $i$ is a member of the never treated group. The potential outcome of unit $i$ in treatment cohort $r\in\mathcal{R}$ observed in time period $t\in\{1,...,\overline{T}\}$ is denoted by $Y_{i,t}(r)$, where the “baseline” potential outcome in period $t$ if unit $i$ is untreated is given by $Y_{i,t}(\infty)$. The objects of interest are the cohort and time-specific treatment effects that may depend on a vector of observed and time-invariant covariates $X_i$, i.e.
In many applications, it is assumed that the treatment effect is a linear function of the observed covariates (which may contain polynomials of $X_i$), so that
where $\dot{X}_i=X_i-\mathbb{E}[X_i|G_i^r=1]$ is the covariate centered around the cohort mean so that $\pi_{r,t}$ is the treatment effect averaged across the distribution of $X_i$ conditional on $G_i^r=1$. Adapting (ref), we assume
where $\mathbb{G}_i=(G_i^{T+2},...,G_i^{\overline{T}})$, so that deviation from the designated treatment path (e.g.\ through anticipation) are ruled out. Further extending (ref),
where $\gamma_{r,s}$ is a non-stochastic constant that differs across cohorts and time. Notice that the covariates are again assumed to enter the baseline outcome linearly. As in Section (ref), the cohort and time-specific ATTs cannot be identified without further restrictions. In order to ensure that $\gamma_{r,s}=0$, we adapt (ref) to impose a conditional staggered parallel trends assumption (CSPTA), using the never-treated group as the control:
Combining the above assumptions, we obtain
Simple algebra sows that the placebo conditional cohort and time specific treatment effects $\tilde{\pi}_{m,k}+\tilde{\rho}_{m,k}'\dot{X}_i$ are identified by taking differences-in-differences, i.e.\ by comparing the evolution in the average outcome between periods $k$ and the base $T+1$ between treatment cohort $m$ and the never treated group. Given (ref)--(ref), the CSPTA implies that the placebo treatments are zero. We therefore avoid any “contamination” by treatment effects at time $m'>m$, which, as noted by abraham2018estimating, can lead to a rejection of the CSPTA in the pre-treatment periods even in cases where it actually holds. Since the assumptions in Section (ref) can easily be shown to hold under mild conditions by adapting the arguments in Section (ref), we can directly apply our equivalence tests to the placebo treatment effects. Moreover, the model in (ref) can be flexibly adjusted to the problem at hand. For instance, one may be willing to exclude a subset of the placebo treatment effects from the model in order to allow for some pooling across cohorts and time. As noted by wooldridge2021two, in this case, the pooled OLS estimator of $\pi_{r,s}$ is an averaged “rolling DiD” where, on top of the never-treated group and the base period, any cohort and period that corresponds to an omitted placebo treatment effect is used as a control. In this case, the CSPTA needs to be adjusted accordingly (e.g.\ as in roth-survey), as parallel trends need to be plausible between multiple groups and periods. Finally, notice that in practice $\mathbb{E}[X_i|G_i^r=1]$ needs to be replaced by the sample average of $X_i$ in cohort $r$. As suggested in wooldridge2021two, one should adjust the standard errors to account for the additional sampling variation.
As a further note of caution, practitioners should be aware that our tests only consider an implication of the CSPTA, as we are relying on assumptions that impose a particular form of treatment effect heterogeneity (e.g.\ linear dependence on observed covariates). In fact, we are thus testing a joint test of a parametric restriction on treatment effect heterogeneity and conditional parallel trends. While parametric restrictions are popular in the applied literature, one might still be concerned about their validity when testing for parallel trends. As an alternative, one may then consider tests based on conditional moment inequalities (see, for instance, andrews-shi).
In order to investigate the small sample properties of our tests, we conduct a simulation study in R. For that, we create a panel data set with $T\in\{1,4,8,12\}$ and $n\in\{100,1000\}$. For $i=1,...,n$ and $t=1,...,T+2$, we generate the data from
with $\alpha_i$ and $\lambda_t$ standard normal, $\Pr(G_i=1)=\frac{1}{2}$ and $\pi_{ATT}=0$.
In an initial step, we investigate the level of the proposed tests. To do so, we set the level of significance to $\alpha=5\%$ and the equivalence threshold for all hypotheses to $1$. We then choose the parameters $\beta_l$ in the pre-treatment periods such that we are on the ”boundary” of the hypotheses, that is $\beta_1=1$ and $\beta_l=0$ for $l>1$ or $\beta_l=1$ for all $l=1,...,T$. Moreover, we also investigate the power of the test procedures by choosing $\beta_l\in\{0.8,0.9\}$ for all $l=1,...,T$. The bootstrap based tests for (ref) are computed using $1000$ bootstrap draws. Finally, the generated error terms are both serially correlated and heteroskedastic, as they are drawn from a stationary AR(3) process with autoregressive parameters $(\phi_1,\phi_2,\phi_3)=(0.5,0.3,0.1)$ and with standard deviation $1+G_i$. Consequently, the tests that require an estimate of $\Sigma$ are based on standard errors that are clustered on the individual level. The results for all tests based on $20000$ simulations are presented in Tables (ref)--(ref).
In the following scenarios, we choose the level of significance $\alpha=5\%$ and compute $\delta^*_{IU}$, $\delta^*_{Boot}$ and $\delta^*_{c.Boot}$ as the smallest equivalence thresholds for the IU, the bootstrap and the cluster bootstrap tests such that the null hypothesis in (ref) can still be rejected (i.e.\ for which equivalence of pre-trends can be concluded). Similarly, we compute the smallest equivalence thresholds $\tau^*$ and $\zeta^*$ for the corresponding null hypotheses in (ref) and (ref) using the tests in (ref) and (ref), respectively. The reported numbers correspond to the average over $M=2500$ simulations and can be used to assess at what value of the equivalence threshold a particular test can be expected to reject the null hypothesis. Finally, we report the usual $95\%$ confidence interval $\mathrm{CI}^{\hat{\pi}_{ATT}}$ and the number of simulations in which each $\beta_l$ for $l=1,...,T$ was found to be statistically insignificant. We then investigate how violations of the PTA affect the chance of falsely detecting a treatment effect and how these violations affect the smallest equivalence thresholds for which equivalence can be concluded. For simplicity, the model errors are drawn independently from a standard normal distribution.
Table (ref) shows the results under the PTA. We further simulate scenarios in which the PTA is violated due to the presence of unobserved effects that affect the treatment group but not the control group. First, we simulate a pre-program shock, also known as “Ashenfelter's dip” by replacing $u_{i,t}$ in (ref) by $\tilde{u}_{i,t}=u_{i,t}+ G_i D_{T+1}(l) V_i$, where $V_i\sim \mathrm{N} (\mu,1)$ with mean $\mu\in\{\frac{1}{4},\frac{1}{2}\}$. This implies that $\mathbb{E}[\hat{\pi}_{ATT}]=\mathbb{E}[\hat{\beta}_l]=-\mu$ for $l=1,...,T$. The results are presented in Tables (ref) and (ref). In our second scenario, $\tilde{u}_{i,t}=\psi\times t\times G_i$ with $\psi=0.025$, modeling an unobserved linear difference in time trends, starting in $t=1$. Consequently, $\mathbb{E}[\hat{\beta}_l]=\psi(l-T-1)$ and $\mathbb{E}[\hat{\pi}_{ATT}]=\psi$. The results are presented in Table (ref).
Table (ref) shows that the test in (ref) approximately keeps the desired level for every $T$ even in small samples. The test in (ref) appears to be slightly over-rejecting when $n =100$ but keeps its nominal level in larger samples. Notice that in Table (ref) the tests in (ref) and (ref) rightfully reject the null hypothesis in an increasing number of cases as $n $ and $T$ increase as $\bar{\beta}$ and $\beta_{RMS}$ are further away from the boundary of the null, resulting in an increase in statistical power. Regarding the tests for (ref) Tables (ref) and (ref) illustrate that the IU and cluster bootstrap based tests maintain their nominal level for $T=2$. When only one parameter is at the boundary of the null hypothesis, both tests also perform well in the sense that the empirical rejection frequency is close to the nominal level for sufficiently large $n$. In comparison, the test in (ref) over-rejects even for large $n$, showing that it is not robust to high levels of serial correlation. If $\beta_l=1$ for all pre-treatment periods, all three tests become conservative for larger values of $T$. This phenomenon appears to be much more pronounced for the test based on the IU principle, for which it is well-documented berger1996bioequivalence. For instance, the empirical level of the IU based test is more than 7 times smaller than the corresponding level of the bootstrap based tests for $T=12$ (see Table (ref)). As shown in Table (ref), this has important consequences for the power of both tests as the cluster bootstrap based test procedure outperforms the IU based test for $T>2$. On the other hand, the IU based test may still be attractive for practical applications as it is numerically much less demanding. As compared to the tests for (ref), the power of our test in (ref) is substantially larger, only surpassed by the power of the test in (ref). All tests have in common that the power decreases with $T$. This is true even for the test in (ref), which is the (asymptotically) uniformly most powerful test for (ref) for any $T$. Thus, concluding equivalence of pre-trends becomes more demanding with an increase in the number of pre-treatment periods. This makes intuitive sense in the DiD setup, where equivalence of pre-trends in a larger number of periods is often regarded as stronger evidence for the plausibility of the PTA.
For $T=1$, we find that $\zeta^*>\delta^*_{Boot}\approx \delta^*_{c.Boot}>\delta^*_{IU}=\tau^*$. This is not surprising, since the IU based test is asymptotically uniformly most powerful for the hypothesis (ref), where (ref) coincides with (ref) and (ref) when a single pre-treatment parameter is tested. For $T>1$, we roughly observe that $\tau^*<\zeta^*\leq\delta^*_{Boot}\leq \delta^*_{c.Boot} <\delta^*_{IU}$. This relationship between the smallest equivalence thresholds of the maximum, average and RMS tests is expected, as $|\bar{\beta}|\leq \beta_{RMS}\leq \|\beta \|_{\infty}$. The better performance of the bootstrap based tests for (ref) as compared to the IU based test may attributed to their higher power as evidenced in Tables (ref) and (ref). We also observe that $\delta^*_{Boot}$ performs only slightly better than $\delta^*_{c.Boot}$ under the PTA when errors are spherical. Under violations of the PTA, $\delta^*_{Boot}<\delta^*_{c.Boot}$ for $T>1$. However, as $\tilde{u}_{i,t}$ becomes non-spherical, only the cluster bootstrap based test maintains its nominal level (see Table (ref)). We therefore recommend to use the cluster-robust version of the bootstrap based test in practice. Further notice that even when the PTA holds, the practice of rejecting the DiD framework when $\hat{\beta}_l$ is statistically insignificant for at least one $l\in\{1,...,T\}$ is clearly inefficient as is shown by the first row of Table (ref), as an increase in available pre-treatment periods increases the chance of incorrectly rejecting the DiD framework under the PTA. A similar observation can be made in the presence of a linear time trend as shown in Table (ref). Here, even when the empirical coverage rate of the usual confidence interval is only slightly lower than the nominal level, the DiD framework is rejected in a large number of cases.
When the PTA is violated due to a small temporary shock as in Table (ref), the usual practice of adopting the PTA when no significant differences in pre-trends could be found can lead to a false discovery of a non-zero treatment effect in a substantial number of cases, in particular when the sample size is small. If the temporal shock is larger as in Table (ref), a non-existing treatment effect will be found to be significantly different from zero in almost all cases. All our test procedures require an unrealistically large equivalence threshold in order to be able to conclude equivalence of pre-trends. In particular, any equivalence threshold for which equivalence could be concluded would have to be larger than the estimated treatment effect, therefore casting serious doubt on the validity of the latter. Similarly, when the PTA is violated due to a linear difference in trends (see Table (ref)), the equivalence thresholds would have to be chosen larger than the estimated treatment effect, thus suggesting that the estimated ATT may contain bias due to insufficient support for the PTA. Moreover, our methodology can be useful in identifying the presence of a linear time trend, as $\tau^*$ and $\zeta^*$ tend do remain stable in $T$ under the PTA or when the violation of the PTA is only temporary, whereas under the presence of a linear trend, they increase with $T$.
We illustrate our approach by re-considering the influential Difference-in-Differences analysis in ditellacrime2004. They use a shock to the allocation of police forces as a consequence of a terrorist attack on a Jewish institution as a natural experiment to study the the effect of police on crime. The original authors conduct the usual pre-test in (ref) and find no evidence for violations of the PTA. However, donohue2013police point out several shortcomings of the original paper (e.g.\ spillover effects from the treated to the untreated group). In particular, they find that the PTA is not plausible if the pre-treatment data is inspected on a more granular level, thus casting doubt on the validity of the estimated treatment effects. While the traditional test failed to detect evidence against the PTA, we will apply our test procedures to analyze how much evidence in favor of the PTA can be extracted from the original specification in ditellacrime2004.
The data consists of monthly averages of the number of car thefts between April and December 1994 in each out of 876 Buenos Aires city blocks out of which 37 blocks received additional protection after the attack. The main specification in ditellacrime2004 is given by $Y_{it}=\alpha_i+\lambda_t+\beta D_{it}$, where $Y_{it}$ denotes the number of car thefts in block $i$ and month $t$ and $D_{it}$ is a dummy variable taking the value 1 if block $i$ is treated in period $t$. Finally, $\alpha_i$ and $\lambda_t$ are block- and time-specific fixed effects. By using this specification, the pre- and post-treatment periods are pooled together so that the estimated treatment effect compares the post-treatment difference in car thefts between treated and non-treated blocks to the corresponding pre-treatment difference. To analyze group mean differences in the pre-treatment periods, we fit (ref) to subsets of the data that include one, two or three pre-treatment periods, corresponding to June, May and June and April--June. As in the original paper, we include block- and time-specific effects and cluster on the block level. We find that using heteroskedasticity-robust standard errors instead of clustering has no substantial effect. Finally, we compute $\delta^*_{IU}$, $\delta_{Boot}^*$, $\delta_{c.Boot}^*$, $\tau^*$ and $\zeta^*$. The results are summarized in Table (ref) below.
Notice that for all tests the smallest equivalence threshold that still allows us to conclude equivalence of pre-trends are the largest when only the pre-treatment period June is used. This hints towards a temporary shock to treatment or control in June which may bias the pooled estimates in Table 3 of ditellacrime2004. The latter are significant and range between $-0.058$ and $-0.081$. One important outcome of our equivalence test based analysis is that, even without the granular data inspection of donohue2013police, the equivalence thresholds have to be chosen unrealistically large in order to conclude equivalence of pre-trends. In fact, the smallest equivalence thresholds for which the null hypotheses can be rejected are larger than the estimated effect size of police on crime. Therefore, it is questionable whether there is any effect at all, since the estimated effect may be an artifact of the violated PTA only.
We have derived several tests that allow researchers to provide statistical evidence in support of the parallel trends assumption in difference-in-differences estimation by testing for negligible differences in pre-trends in treatment and control. The tests can easily be implemented in popular models such as the two-way fixed effects model, and they can be flexibly adjusted to accommodate more complex settings. Our simulation analysis yields support for our theoretical results as our tests maintain their nominal level and exhibit high statistical power in sufficiently large samples. Finally, we apply our methodology to the data provided by ditellacrime2004. Even without a granular inspection of the data as in donohue2013police, our methodology casts doubt on the estimated effects, as they may simply be the result of trend differences that could not be detected using traditional pre-tests.
We thank the editor Ivan Canay, an associate editor, and two anonymous referees for comments that greatly improved this paper. We further thank participants of the International Panel Data Conference 2023, the European Summer Meeting of the Econometric Society 2023 and the European Winter Meeting of the Econometric Society 2023.
The authors report there are no competing interests to declare.