EconBase
← Back to paper

MSE-Optimal Difference-in-Differences Estimator

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

38,837 characters · 14 sections · 5 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

MSE-Optimal Difference-in-Differences Estimator

\selectlanguage{english}

abstractThis paper develops a difference-in-differences (DiD) estimation method that selects the optimal length of pre-trends by minimizing the mean squared error (MSE). Conventional DiD regression models, such as the two-way fixed effects model or the event study model, may suffer from accuracy and validity concerns. If the sample size is small, the estimator may have a larger variance. Also, pre-tests often lack power to detect violations of the parallel trends assumption as \citeA{Roth:2022} highlights. By focusing on the bias and variance tradeoff, the proposed method derives the MSE-optimal estimator from the optimal length of pre-trends. Simulation results and an empirical application demonstrate the practical applicability of the proposed method.

Keywords: Difference-in-Differences, Pre-trends, Mean Squared Error

Introduction

Difference-in-differences (DiD) is a widely used method in economics for estimating causal effects. One of the well-known studies that implemented DiD is \citeA{Card_and_Krueger:2022}, which studies the effect of the minimum wage increase in New Jersey. Recently, there have been many theoretical advances in DiD methods enabling researchers to estimate treatment effect in a wide range of settings.

In practice, DiD is usually implemented via two regression models: the two-way fixed effects (TWFE) model and the event study model. Aside from the problem of staggered treatment timing, those models have several problems. One of the problems is that precise estimation is difficult when the sample size is small. Event study plots, which are commonly used to present event study estimates, tend to have wider confidence intervals particularly when the number of observations is small. This can be viewed as a variance problem. Another problem is that the pre-trends test is not always effective Roth:2022. Even though there may be a potential bias between the treatment and control groups, the pre-trends test cannot detect it. This can be viewed as a bias problem.

To address these problems, this paper proposes an MSE-optimal approach to difference-in-differences estimation. The framework is the following. When longer pre-trends are included, the estimator tends to have a smaller variance. At the same time, the estimator may be biased if there is a potential bias between the treatment and control groups. On the other hand, the estimator tends to have a larger variance when only a shorter pre-trends are included. Also, the estimator is likely to be less biased. Using the bias and variance tradeoff, this method derives the MSE-optimal estimator by choosing the optimal length of the pre-trends. I illustrate this method through simulations and an empirical application. The results show that the proposed method performs well in terms of MSE relative to conventional methods.

Several studies examine pre-trends in the DiD setting. \citeA{Roth:2022} shows that the conventional pre-trends tests often have low power and that depending on such tests can distort estimation and inference. \citeA{Rambachan_and_Roth:2023} propose robust inference procedures that allow for possible violations of the parallel trends assumption. My paper takes the possibility of such bias seriously and proposes a procedure that minimizes the MSE of the estimator by explicitly balancing the bias and variance. \citeA{Egami_and_Yamauchi:2023} show how multiple pre-treatment periods can improve DID estimation and propose double DID, which can be more efficient and rely on weaker assumptions than the conventional DID. My paper also focuses on the role of pre-treatment periods in DiD estimation, by viewing the length of the pre-trends as a choice variable in estimation. \citeA{Ishimaru:2026} studies treatment effect estimation in panel data without relying on the parallel trends assumption and proposes alternative conditions for identifying the average treatment effect on the treated. By contrast, my paper works within a DiD framework in which deviations from parallel trends are assumed to be limited, so that the resulting distortion can be interpreted as a bias and incorporated into the MSE criterion.

The bias-variance tradeoff is also fundamental in the selection of the optimal bandwidth in regression discontinuity design (RDD) Imbens_and_Kalyanaraman:2012, Calonico_and_Cattaneo_and_Titiunik:2014. A wider bandwidth includes more observations in estimation, which generally lowers variance. At the same time, however, observations farther from the cutoff may be less informative about the local treatment effect at the threshold, so a wider bandwidth may induce greater bias. Conversely, a narrower bandwidth focuses on observations closer to the cutoff, thereby reducing potential bias because the identifying assumptions are more plausible. However, fewer observations are available; therefore the estimator typically has higher variance. Thus, the bandwidth choice in RDD is naturally characterized as a bias-variance tradeoff. This paper applies a similar idea to the DiD setting: extending the pre-treatment period may improve precision, but it may also increase bias because the conventional pre-trends test may fail to detect the bias.

The rest of the paper is organized as follows: In the next section, I describe the canonical DiD models as baseline DiD models. In Section $3$, I show the proposed model and explain the implementations in the static and dynamic treatment effect cases. Then, I illustrate the model with simulations in Section $4$. I also apply the model to the empirical example in Section $5$, and I conclude in Section $6$.

Baseline

This section describes the standard DiD framework as compared with the model this paper proposes. Section $2.1$ is about the basic $2 \times 2$ design and some key assumptions. Section $2.2$ explains the two-way fixed effects model which is a commonly used regression model for static treatment effect. Section $2.3$ illustrates the event study model which is a commonly used regression model for dynamic treatment effect.

Canonical Difference-in-differences

The simplest DiD setting is a $2 \times 2$ (2 groups (treatment group vs control group) and 2 periods (pre-treatment vs post-treatment)) design. There have been many advances related to DiD methods even in recent years (see, e.g. \citeA{Baker_etal:2022}, \citeA{Baker_etal:2026} and \citeA{Roth_etal:2023} for a review) and those newer methods differ technically from one another, but they fundamentally depend on the same idea. Therefore, I begin with the simplest $2 \times 2$ design and specify some important assumptions.

The notation used in this paper is as follows. $Y_{i,t}$ denotes an outcome of individual $i$ and time $t$. $D_{i,t}$ indicates whether or not the individual $i$ receives treatment in period $t$. $Y_{i,t}(D_{i,t} = 1) = Y_{i,t}(1)$ represents the potential outcome of $Y$ if the individual $i$ receives a treatment at period $t$, and $Y_{i,t}(D_{i,t} = 0) = Y_{i,t}(0)$ represents the potential outcome of $Y$ if the individual $i$ does not receive a treatment at period $t$. Lastly, let $t^*_i$ be the time just before the individual $i$ receives a treatment. Hereafter, $t^*_i$ is set to $-1$ unless specified otherwise.

One common target parameter of DiD is the average treatment effect on the treated (ATT). In the $2 \times 2$ design, the ATT can be expressed as

align*[align* omitted — 73 chars of source]

To obtain this ATT, the following three assumptions - Assumption (ref): Parallel Trends, Assumption (ref): No Anticipation and Assumption (ref): No Spillover - are crucial.

assumption{\normalfont Parallel Trends}\\ \begin{equation*} \mathbb{E}[Y_{i,t}(0) - Y_{i,t-1}(0)|D_i = 1] = \mathbb{E}[Y_{i,t}(0) - Y_{i,t-1}(0)|D_i = 0] \end{equation*}

Assumption (ref) states the average changes in outcomes are the same for the treatment group and the control group if both groups do not receive the treatment.

assumption{\normalfont No Anticipation}\\ \begin{equation*} Y_{i,t} = Y_{i,t}(0) \ for all \ t \leq t^{*} \ and \ i \ with \ D_i = 1 \end{equation*}

Assumption (ref) is an assumption that implies the treatment does not affect any actions before the treatment timing. This assumption can be interpreted as a situation that no one could know whether or not they would receive a treatment.

assumption{\normalfont No Spillover\footnote{Although this paper maintains the No Spillover assumption, recent work has developed DiD methods that explicitly account for spillover effects; see \citeA{Butts:2023} and \citeA{Fiorini_Lee_and_Pfeifer:2024}.}}\\ \begin{equation*} \forall j \neq i, \ Y_{i,t} (D_i, D_j) = Y_{i,t} (D_i, D_j') \ for all \ D_j\ and \ D_j' \end{equation*} {\normalfont }

Assumption (ref) states an individual's outcome is not affected by the treatment status of others.

Under these assumptions, ATT in the $2 \times 2$ setting can be transformed into the following equation.

align[align omitted — 297 chars of source]

The above discussion presents the $2 \times 2$ DiD framework. I extend the $2 \times 2$ DiD framework to a more general setting with multiple individuals observed over multiple time periods, focusing in particular on the extension from two periods to many periods. Throughout the analysis, I maintain Assumptions 2 and 3.

Two-Way Fixed Effects

In practice, researchers typically estimate the ATT in a DiD framework using regression-based methods. One of the most commonly used specifications is the two-way fixed effects (TWFE) model. The TWFE specification without covariates is given by:

equation[equation omitted — 90 chars of source]

where $\mu_i$ is an individual fixed effect and $\theta_t$ is a time fixed effect. Then, the coefficient $\beta$ represents the ATT. This model is suitable when the treatment effect is “static” (e.g. Figure(ref) A). If the treatment effect is constant across post-treatment periods, a time-invariant $\beta$ is appropriate. Although the TWFE model had been used widely in empirical research, this model has a problem with staggered treatment timing Goodman-Bacon:2021,Callaway_and_Sant'Anna:2022,de_Chaisemartin_and_d'Haultfoeuille:2020,de_Chaisemartin_and_d'Haultfoeuille:2023. This paper assumes a treatment occurs at the same time for all individuals, but it might be possible to extend to staggered treatment timing.

Event Study

Another model to derive a DiD estimator is the event study model. This model is suitable when the treatment effect is “dynamic” (e.g. Figure(ref) B). The event study model without covariates can be written as:

equation[equation omitted — 135 chars of source]

where $\mu_i$ and $\theta_t$ are the same as the TWFE model. The difference is the third term, which makes it possible to estimate the dynamic treatment effect. When the treatment effect is dynamic, $\beta_l$ must be time-variant to capture the change in the treatment effect. This paper does not address this issue, but staggered treatment causes a problem in the event study model similarly to the TWFE model Sun_and_Abraham:2021.

Model

This section describes the model this paper proposes. I first clarify the idea of the proposed model in Section $3.1$. Then, I illustrate two cases depending on the type of treatment effect (static and dynamic) in Section $3.2$ and $3.3$.

Model Overview

The main idea of this paper is to construct the MSE-optimal DiD estimator by changing the length of pre-trends. I first explain why it is appropriate to consider the MSE-minimizing approach in this context. In general, estimators including DiD estimators have smaller standard errors when the sample size is large. This means a DiD estimator would be more precise if the pre-trends are longer. At first glance, longer pre-trends may seem preferable to shorter ones. However, longer pre-trends might bias the estimator unintentionally. In DiD studies, it is common to check the parallel trends assumption by examining whether or not pre-treatment estimates are statistically close to zero using $95\%$ confidence intervals (CI), but such a pre-trends test may fail to detect bias Roth:2022. Therefore, it is possible that the estimator is biased. In particular, if there are longer potentially nonparallel pre-trends, the bias would be larger. Hence, longer pre-trends might have a smaller variance and a larger bias, and shorter pre-trends might have a larger variance and a smaller bias. It can be said that there is a bias and variance tradeoff. Note that MSE is expressed as the following:

equation[equation omitted — 73 chars of source]

I define the optimal choice as the one that minimizes the MSE.

equation[equation omitted — 98 chars of source]

This is why it is possible to incorporate the MSE-minimizing approach into the DiD framework by changing the length of pre-trends.

To analyze the proposed procedure more precisely, I consider an oracle benchmark. Let $\mathcal{L}=\{0,1,\ldots,L\}$ denote the set of candidate length of pre-trends. Also, let $\ell \in \mathcal{L}$ denote the length of pre-trends and $\ell_{\text{post}} (\geq 1)$ denote the length of post-treatment periods, which is fixed in this setting. In this benchmark, I consider the TWFE specification in ((ref)), and the target parameter is $\beta$. For each $\ell \in \mathcal{L}$, let $\hat{\beta}_{\ell}$ denote the DiD estimator constructed with pre-treatment length $\ell$. I impose the following two assumptions.

assumptionAssume that deviations from parallel trends in the pre-treatment periods follow a linear form, $\gamma \ell$.
assumptionAssume that the variance of the candidate estimators is finite.

Then, the following proposition holds.

propositionSuppose that the deviations from parallel trends in the pre-treatment periods satisfy Assumption (ref) and the variance of the candidate estimators satisfies Assumption (ref). Also, suppose that the candidate set $\mathcal{L}=\{0,1,\ldots,L\}$ is finite. Then, there exists an MSE-optimal DiD estimator $\hat{\beta}^*$.
proofFor each $\ell \in \mathcal{L}$, the MSE of $\hat{\beta}_{\ell}$ is \begin{align*} MSE\big(\hat{\beta}_{\ell}\big) &= Bias\big(\hat{\beta}_{\ell}\big)^2 + Var\big(\hat{\beta}_{\ell}\big) \\ &= \Big(\frac{\gamma \ell}{\ell + \ell_{post}}\Big)^2 + Var\big(\hat{\beta}_{\ell}\big) \ ( < \infty ) \end{align*} Then, the optimal length of pre-trends $\ell^*$ is \begin{align*} \ell^* = \operatorname*{arg\,min}_{\ell \in \mathcal{L}} MSE\big(\hat{\beta}_{\ell}\big) = \operatorname*{arg\,min}_{\ell \in \mathcal{L}} \Bigg( \Big(\frac{\gamma \ell}{\ell + \ell_{\text{post}}}\Big)^2 + \text{Var}\big(\hat{\beta}_{\ell}\big) \Bigg) \end{align*} and the MSE-optimal DiD estimator $\hat{\beta}^*$ is \begin{align*} \hat{\beta}^* = \hat{\beta}_{\ell^*} \end{align*}

This oracle rule is not feasible in practice because the true bias and variance are generally unknown. However, it serves as a theoretical benchmark for the feasible selection rule developed below.

In the feasible implementation, the variance is approximated using the squared standard error of the estimate. On the other hand, the true bias is unobserved in practice. I therefore use the estimator with pre-treatment length $0$ as a benchmark. Specifically, I define the approximated bias for the estimator with pre-treatment length $\ell$ as

equation*[equation* omitted — 170 chars of source]
figure[figure omitted — 703 chars of source]

Static Treatment Effect Estimation

I consider the MSE-minimizing approach in the case of a static treatment effect. In this case, the TWFE model is often used as I mentioned in the previous section. Therefore, I incorporate the MSE-minimizing approach into the TWFE model. Figure(ref) shows the changes of the outcome with pre-trends length $9, 4,$ and $ 0$. When the length of pre-trends is $9$, more observations are available to estimate $\beta$ in ((ref)). Hence, the variance of the estimate would be smaller. However, pre-trends are not completely parallel and this would bias the estimate. When the length of pre-trends is $0$, there are fewer observations to estimate $\beta$ in ((ref)). Thus, the variance of the estimate would be larger. On the other hand, there are no pre-trends that would induce bias in the estimate. When the length of pre-trends is $4$, there is a moderate number of observations to estimate $\beta$ in ((ref)). Namely, the variance of the estimate would be moderate. Also, there is an intermediate length of pre-trends that would cause the estimate to be biased. By calculating the variance and the bias of each length of pre-trends, the optimal length of pre-trends and the MSE-optimal TWFE estimator can be obtained.

figure[figure omitted — 812 chars of source]

Dynamic Treatment Effect Estimation

I next consider the MSE-minimizing approach in the case of dynamic treatment effect. Let $\beta_{\ell}$ denote a coefficient for the treatment effect in the $\ell$th period after the treatment, and suppose that this coefficient is the parameter of interest. The goal is to minimize the MSE of the estimator for $\beta_{\ell}$. Although the basic idea is the same as in the static treatment effect case, an additional step is required in the dynamic setting before changing the length of pre-trends.

The estimator of the conventional event study model ((ref)) for the treatment effect in the $\ell$th period after the treatment can be written as follows:

align[align omitted — 244 chars of source]

This means the estimator $\hat{\beta}_{\ell}$ is not affected by the values of $Y_{i, t}$ other than $Y_{i, t^*_i + \ell}, Y_{i, t^*_i}$. Hence, changing the length of pre-trends does not affect the bias and variance of the estimator $\hat{\beta}_{\ell}$.

Therefore, I propose the following modified model ((ref)) when the treatment effect is dynamic.

equation[equation omitted — 147 chars of source]

Then, the estimator of this modified model ((ref)) for the treatment effect in the $\ell$th period after the treatment can be written as follows:

align[align omitted — 273 chars of source]

Compared with the conventional event-study model, this modified specification does not include the pre-treatment coefficients $\beta_{\ell}$ for $\ell \leq -2$. As a result, more observations can be used to estimate the post-treatment coefficients $\beta_{\ell}$, which may reduce the variance relative to the conventional event-study model ((ref)). However, because the modified specification does not estimate coefficients for pre-treatment periods, the conventional pre-trends test cannot be conducted within this model itself. For this reason, the use of the modified specification is justified when the data have already passed the conventional pre-trends test under the conventional event-study model ((ref)).

A similar approach can then be applied as in the TWFE case. When the pre-treatment period is longer, the estimator tends to have a smaller variance because more observations are used in the estimation. However, the estimator may also be more biased, since the model imposes the parallel trends assumption between the treatment and control groups over a longer pre-treatment window. By contrast, when the pre-treatment period is shorter, the estimator tends to have a larger variance because fewer observations are available. At the same time, the potential bias may be smaller, because the model relies on the parallel trends assumption over a narrower pre-treatment window. By minimizing the MSE of the estimator, the optimal length can be selected, yielding the MSE-optimal event-study estimator.

Simulation

In this section, I conduct two simulations to illustrate the proposed model. I first illustrate the case of the static treatment effect in Section $4.1$. I then explain the dynamic treatment effect case in Section $4.2$.

Static Treatment Effect Estimation

This simulation applies the method in Section 3.2 to simulated data with fundamentally non-parallel trends between the treatment and the control groups. The simulated data are generated according to the following setup:

align*[align* omitted — 392 chars of source]

Both the treatment group and the control group are set to have $50$ individuals and $t \in [-10,10]$. The trends of each group are shown in Figure (ref).

The results are reported in Table (ref). The bias tends to increase when the pre-trends extend over a longer period. At the same time, the variance tends to get smaller. The MSE is minimized when the pre-trends length is $3$, which is the optimal pre-trends length. The MSE-optimal TWFE estimate is $95.449 \ (6.076)$. The conventional TWFE model does not change the pre-trends length; therefore, the conventional TWFE model estimate is $91.570 \ (4.577)$. Compared with this estimate, the MSE-optimal estimate is less biased but it has a larger variance.

table[table omitted — 2,305 chars of source]

Dynamic Treatment Effect Estimation

The next simulation considers a dynamic treatment effect. This simulation applies the method in Section 3.3 to different simulated data which also have fundamentally non-parallel trends between the treatment and the control groups. The simulated data are generated according to the following setup:

align*[align* omitted — 423 chars of source]

In the same way as the simulation in Section $4.1$, both the treatment group and the control group are set to have $50$ individuals and $t \in [-10,10]$. The trends of each group are shown in Figure (ref).

Suppose that the target parameter is $\beta_5$. Note that the model to estimate is model ((ref)). The results are reported in Table (ref). The bias of the estimates changes nonmonotonically but it seems to get larger when the pre-trends extend over a longer period. The variance tends to get smaller as the pre-trends become longer. The MSE is minimized when the pre-trends length is $6$, which is the optimal pre-trends length. The MSE-optimal event study estimate is $50.943 \ (22.021)$.

table[table omitted — 2,373 chars of source]
landscape\begin{figure} \caption{Event study plots with different models} \begin{minipage}{14cm} Notes: These figures represent the event study plots with three models: (A) the traditional event study model ((ref)), (B) the modified event study model ((ref)) with the full length of pre-trends, (C) the MSE-optimal event study model. The black points are the point estimates and the vertical lines are the $95\%$ confidence intervals. The triangle points are the true values of the simulation. \end{minipage} \end{figure}

I compare the results with the conventional event study model. Figure (ref) (A) represents the event study plot of the conventional event study model ((ref)). It barely passes the pre-trends test, but the estimates have larger variances. Next, I apply the modified model ((ref)). Figure (ref) (B) shows the event study plot of the modified event study model ((ref)) with the full length of pre-trends. Owing to the increase of the sample size for the estimation, the estimates have smaller variances. However, imposing the parallel trends assumption on pre-treatment periods for the two groups introduces bias into the estimates. Lastly, Figure (ref) (C) shows the event study plot of the MSE-optimal event study model. The estimates have smaller variances compared with the traditional event study model ((ref)) and are less biased compared with the modified event study model ((ref)) with the full length of pre-trends.

Empirical Application

I apply the MSE-optimal DiD model to the setting of \citeA{Fitzpatrick_and_Lovenheim:2014}. This research studies the effect of the Early Retirement Incentives (ERI) program in Illinois in 1993 on student achievement. They estimate regressions of the following form:

equation[equation omitted — 278 chars of source]

where $Y_{igt}^s$ is the standardized test score in grade $g$ for subject $s$ in school $i$ and year $t$. The variable $Teachers \geq 15$ is the average number of teachers in a given grade with at least $15$ years of experience pre-$1994$. $Teachers$ is the average total number of teachers in a grade and school in the pre-ERI period, and $Post$ is an indicator variable equal to $1$ for school years after $1993$. The vector $\mathbb{X}$ contains the set of school-by-year demographic variables. $\delta_{ig}$ is school-by-grade fixed effects, and $\phi_{tg}$ is grade-by-year fixed effects.

In one analysis, they estimate the effect of the ERI program on the Math test score using a conventional event study model based on the regression ((ref)). The results are shown in the left panel of Figure (ref). It passes the pre-trends test but the confidence intervals are relatively wide. In other words, the variances of the estimates are large.

Suppose that the target parameter is the $1997$ coefficient because the effect of the ERI program on student achievement may arise with a delay. Then, the results of the MSE-optimal event study model are reported in Table (ref). Based on the calculation, MSE is minimized when the length of the pre-trends is $3$. Therefore, the MSE-optimal event study plot is shown in the right panel of the Figure (ref).

figure[figure omitted — 741 chars of source]
table[table omitted — 1,453 chars of source]

Conclusion

This paper proposes an MSE-optimal approach to difference-in-differences estimation. Conventional DiD designs (TWFE model, event study model) may suffer from either high variance or substantial bias. The proposed model changes the length of pre-trends and finds the optimal one that minimizes the MSE of the estimator. I illustrate the method using simulations and an empirical application. As future work, it might be possible to apply this method to a staggered treatment setting.

\nocite{*}