Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
63,521 characters · 13 sections · 54 citation commands
Misclassification in Difference-in-Differences Models
The difference-in-differences (DID) method is a popular quasi-experimental technique used to identify causal effects of a treatment (e.g., policy intervention) when data is available on the pre- and post-treatment periods. As of 2018, currie2020technology reports that 25 percent of National Bureau of Economic Research working papers in applied microeconomics and 15 percent of papers in “top five” economics journals mention DID.\footnote{Recently, the DID method has drawn significant attention from methodological researchers working to clarify several identification and estimation issues; see reviews in deChaisemartin2022 and roth2022s.} However, when the treatment variable is observed with errors (in which case we say the treatment is misclassified), the DID estimand may not have a clear causal interpretation even when the identifying parallel trends (PT) assumption holds.
Misclassification can arise in DID designs from many sources. One common scenario occurs when researchers estimate or infer the treatment variable from auxiliary data, thereby potentially introducing misclassification (cortes2013achieving, deChaisemartin2018, and fortson2009hiv). This scenario frequently occurs when researchers use a proxy variable to classify units into treated and control groups because the treatment variable is missing for either the pre- or post-intervention period. For instance, groen2008effect studies the impact of Hurricane Katrina on labor market outcomes of evacuees. The treatment variable in their study is a binary variable for being an evacuee, but the Current Population Survey only collected information on evacuee status after Katrina. As such, the authors defined treated units (evacuees) in the pre-treatment period as those living in Katrina-affected areas, potentially introducing misclassification because not everyone living in those areas evacuated after the storm.
In other cases, researchers use a mismeasured continuous treatment variable to define a binary treatment variable for DID estimation. This is often the case when units are classified as treated when an estimated index or rate exceeds a specified threshold chosen by the researcher (Kessler_al2022 and Miller2012).\footnote{Another likely reason for converting continuous treatments to binary variables is the lack of theoretical work on the identification of treatment effects with continuous treatments in DID designs. However, a recent notable exception is callaway2024difference.} The resulting binary variable from the mismeasured continuous treatment is necessarily misclassified. Even when the underlying continuous variable is correctly measured, researchers often resort to creating binary variables for DID estimation with the aim of capturing the intensity of exposure to some policy intervention (draca2011minimum and galasso2022does).
In other cases, there is ambiguity regarding the exact timing of a reform's passage or implementation. For instance, when a significant amount of time elapses between a legislation's proposed date and its eventual enactment, the researcher may opt to use the former to define the timing of treatment (kresch2020buck). Researchers might also lack information about the actual implementation of a policy when there is a lag between the passage and its effective implementation or when they use the lag to define the post-intervention period with the aim of allowing sufficient time for the policy's impact to kick in (bindler2018punishment and murray2016mice). Even when researchers know the true treatment date, data unavailability might compel them to define overlapping pre- and post-intervention periods that introduce misclassification (buchmueller2011effect). Most of the above studies admit the misclassification problems confronting them, but there is a conspicuous lack of methodological work addressing it.
In this paper, we study the identification of the average treatment effect on the treated (ATT) in the DID framework when the treatment is subject to misclassification. We characterize the resulting bias and propose a partial identification approach that researchers can use to investigate the sensitivity of their DID estimates. Our framework distinguishes the latent treatment from its observed (misclassified) counterpart and characterizes how the DID estimand aggregates causal effects across correctly classified and misclassified subpopulations. This characterization clarifies when misclassification leads to attenuation, when it can generate sign reversals, and which additional assumptions are needed to recover interpretable causal estimates.
Specifically, we make several contributions to the literature. First, this paper is one of the first to study the identification of causal effects in the DID setting when the treatment is misclassified. In work concurrent to ours, negi2025difference apply a one-sided misreporting model (nguimkeu2019estimation) and instrumental variables for both treatment and misreporting to achieve point identification in a parametric DID regression. Their approach is useful when researchers are willing to impose a one-sided misreporting structure and have credible instruments for both treatment and misreporting. In contrast, our contribution targets scenarios where misclassification may be bidirectional, instruments are unavailable, or researchers prefer to avoid strong parametric assumptions. In our framework, we show that under the standard PT assumption, the DID estimand recovers a weighted average of the ATT for the correctly classified and misclassified subpopulations with a non-positive weight for the misclassified units. This finding mirrors the recent “negative weighting” criticism of standard two-way fixed effects (TWFE) estimators (deChaisemartin2022 and goodman2021difference). However, while the negative weights in those TWFE contexts arise from treatment effect heterogeneity in staggered adoption designs, we demonstrate that misclassification alone can generate non-positive weighting, potentially biasing estimates even in simple canonical DID settings. Furthermore, the weights do not necessarily sum up to one. In general, the direction of the bias is unknown, implying that the DID estimand may induce a sign-reversal phenomenon, where its sign could differ from the true causal effect.
Second, we establish a linear relationship between the ATT and the DID estimand where the coefficients are unidentified. This relationship allows us to discuss conditions under which various types of biases may occur. Using this relationship, we then provide a sufficient condition for the existence of a fixed point where the DID estimand still identifies the ATT even in the presence of misclassification.
Third, we show that under additional assumptions, the DID method identifies the sign of the ATT but remains biased in magnitude. In particular, if the misclassification error is nondifferential and a monotonicity condition holds, then the DID estimand only suffers from attenuation bias, producing smaller estimates in magnitude. When the extent of the misclassification is bounded by a sensitivity parameter, we derive bounds on the ATT under the aforementioned assumptions. The choice of the sensitivity parameter is specific to each application and we suggest that researchers use institutional or contextual knowledge and the structure of their data to inform their choice.
Finally, we develop a partial identification approach to overcome arbitrary misclassification in settings where the researcher has access to multiple data sources drawn from the same underlying population but potentially subject to different rates of misclassification. This scenario arises naturally when, for example, treatment status is recorded in both administrative records and survey data, or when multiple proxy variables are available to classify units into treatment arms. Under the assumption that the data source affects misclassification rates but not potential outcomes or true treatment status, coupled with a relevance condition that misclassification rates vary across sources, we show that the observed outcome distributions within each source decompose into finite mixtures whose components are source-invariant but whose weights vary with the source. We exploit this mixture structure to partially identify stratum-specific ATTs for correctly classified and misclassified subpopulations. We then derive bounds on the unconditional ATT as a convex weighted average of the two stratum-specific ATTs where the weights are partially identified. Unlike the previous results above, these multi-data source bounds are data-driven and do not require the researcher to specify a misclassification rate. We conduct inference using the intersection bounds framework of CLR2013.
We illustrate our theoretical results through simulation and empirical exercises. In the simulations, we consider various designs---differential vs nondifferential misclassification, and symmetric vs asymmetric misclassification. The simulation results display sign reversal in the DID estimates when the measurement error is differential regardless of whether the misclassification is symmetric or not. In the case of nondifferential misclassification, we find that the attenuation bias can be substantial depending on the design. For the multi-data source bounds, our simulations show that the method yields informative confidence sets that correctly identify the sign of the ATT.
For applied work, we recommend a simple reporting template that includes the conventional DID estimate under PT and bound sets for the ATT over a transparent range of misclassification rates informed by institutional knowledge. We illustrate these results using two empirical studies. In the main paper, we analyze federalism in the Brazilian water and sanitation sector (kresch2020buck). In Appendix (ref), we revisit the impact of abolishing capital punishment in England between 1772 and 1871 (bindler2018punishment). We discuss the possibility that the treatment is misclassified in these studies and illustrate how we can use the data and contextual information to implement our sensitivity bounding analysis with reasonable choices of the sensitivity parameter.
Our work unites the vast literature on difference-in-differences designs and measurement error. Our work is directly related to the longstanding literature on the identification of causal parameters when a binary treatment variable is misclassified. One set of studies in this literature leverages instrumental variables, auxiliary data, or parametric assumptions to achieve point identification (e.g., bollinger2017bayesian; nguimkeu2019estimation). Other studies focus on partial identification and bounding strategies in settings such as heterogeneous instrumental variable models (e.g., Acerenza_al2021; Chalak2017; Kreideral2012; Ura2018; possebom2025crime; yanagi2019inference). Although these studies and the additional papers cited therein cover many quasi-experimental designs, they do not consider misclassification in the DID framework. Our paper also connects with the recent literature exploring a related but different problem of missing data in DID analysis. These studies provide point and partial identification results in the DID framework when the treatment variable is missing for either the pre- or post-treatment period (botosaru2018difference and fan2017partial).
The remainder of the paper is organized as follows. Section (ref) presents the model, the assumptions, and the main identification results. Section (ref) presents our theoretical results for bounding the ATT. Section (ref) presents simulation results, and Section (ref) presents an empirical illustration. Section (ref) concludes. Proofs and additional results are in the Appendix.
Our framework is the canonical DID design comprising two groups and two periods. Consider the following model:\footnote{The specification $D = D^*(1-\varepsilon) + (1-D^*) \varepsilon$ is shown in Acerenza_al2021 to be without loss of generality.}
where the vector $(Y_0, Y_1,D)$ represents the observed data, while the vector $(Y_t(0), Y_t(1), D^*, \varepsilon)$ is latent. In this model, the variables $Y_t(0)$ and $Y_t(1)$ are the potential outcomes that would have been observed in period $t\in\{0,1\}$ had the treatment been externally set to 0 and 1, respectively. The variables $Y_0, Y_1 \in \mathcal Y$ are the observed outcomes in the baseline $(t=0)$ and the post-intervention $(t=1)$ periods, respectively. The variable $D^*\in \left\{0,1\right\}$ is the true treatment occurring between periods 0 and 1, while $D$ is a potentially misclassified version of $D^*$. The latent variable $\varepsilon$ is the indicator for misclassification. When $\varepsilon=0$, there is no misclassification, but the observed treatment is misclassified whenever $\varepsilon=1$.
As is customary in the DID literature, we assume away any anticipatory effects of the treatment, so that $Y_0(1)=Y_0(0)$. We also assume $0< \mathbb P(D^*=1) < 1$, and $0< \mathbb P(D=1) < 1$, implying that a fraction of the population is treated whether or not the observed treatment is misclassified. In this paper, we are interested in identifying the ATT defined as
In general, the presence of misclassification complicates the identification of the ATT in model (ref). Such misclassification may induce a violation of the parallel trends assumption in the observed treatment variable $D$, in which case the DID estimand will not have a clear causal interpretation. Even when the parallel trends assumption holds in the misclassified $D$, the causal interpretation of the DID estimand remains unclear. Below, we study the consequences of misclassification on the DID estimand when the researcher assumes parallel trends in $D$. We discuss why the researcher may assume parallel trends in $D$ and provide sufficient conditions for this assumption to hold despite the misclassification. We state the assumptions and present our findings.
Assumption (ref) is the standard parallel trends assumption commonly used in the literature on difference-in-differences. It states that in the absence of treatment, the control and treatment groups would have followed the same trend on average. It is equivalent to $$\mathbb E[Y_1(0) \vert D=1]-\mathbb E[Y_1(0) \vert D=0]=\mathbb E[Y_0\vert D=1] - \mathbb E[Y_0 \vert D=0].$$ In the observed model, the standard DID estimand can be defined as $$\theta_{DID} \equiv \theta^1_{OLS}-\theta^0_{OLS},$$ where $\theta^0_{OLS}$ and $\theta^1_{OLS}$ are the ordinary least squares (OLS) (or difference-in-means) estimands at periods 0 and 1, respectively given by
As discussed above, Assumption (ref) is stated in terms of the observed treatment variable (instead of the true, unobserved treatment). This is because when the researcher suspects that the treatment variable is misclassified, they may naturally make Assumption (ref) to proceed with the DID identification strategy for several reasons. They might do so because they want to ignore the problem, (incorrectly) assert that misclassification has minimal consequences or for convenience due to lack of a viable solution (i.e., alternative estimation method). As a result, we study the causal interpretation of the DID estimand under Assumption (ref) for practical considerations. However, before we present the results, we provide sufficient conditions (assumptions) on $(D^*, \varepsilon, Y_1(0), Y_0(0))$ for Assumption (ref) to hold.
The first part of Assumption (ref) is the standard parallel trends assumption stated in terms of the true treatment variable. The second part states that conditional on the true treatment variable, the average outcomes for the correctly classified and misclassified groups would have followed the same trend. This parallel trends assumption permits different average outcome trends over time across the misclassification groups (defined by $\varepsilon$) in each treatment arm. The following lemma shows that these two assumptions imply Assumption (ref).
This section provides our main results for the consequences of misclassification in the DID framework. We first allow misclassification to be arbitrary (potentially differential) with no structure imposed on it. The following proposition provides an expression of the DID estimand when the misclassified treatment variable is used for identification.
Proposition (ref) shows that the DID estimand does not recover the true ATT in equation ((ref)), but rather a weighted average of the ATT for two subpopulations---the correctly classified and misclassified treated groups. Importantly, the weights are positive for the correctly observed units and non-positive for the misclassified observations, and do not necessarily sum up to one. It follows that the DID estimand could yield an opposite sign for the treatment effect in some circumstances. Proposition (ref) shows that the standard DID estimand could be negative even if the true ATT is positive under misclassification and vice versa. Indeed, we note from equation ((ref)) and the law of iterated expectations that
If the quantities $\mathbb E\left[Y_1(1)-Y_1(0) \vert D^*=1, \varepsilon=0\right]$ and $\mathbb E\left[Y_1(1)-Y_1(0) \vert D^*=1, \varepsilon=1\right]$ are positive, then from equation ((ref)), we know that the ATT is positive. It is then straightforward to check that Proposition (ref) implies the DID estimand could be negative in this scenario. In other words, the DID estimand using a misclassified treatment fails to identify an interesting causal parameter and could potentially yield misleading conclusions. Our pessimistic results suggest that the consequences of misclassification in DID analysis are more severe than previously understood. This finding is contrary to observations in some previous empirical papers where the researchers concerned about possible misclassification suggest that their estimates are potentially only attenuated (draca2011minimum, galasso2022does, and Miller2012).\footnote{For instance, in reference to a binary treatment variable created from a potentially mismeasured continuous variable representing the 2005 county-level uninsurance rate in Massachusetts (Uninsured2005c), Miller2012 observes that “One advantage of using a binary indicator, rather than a continuous measure, is that it is not reliant on the assumption of a linear relationship between insurance coverage and emergency room usage and is more robust to measurement error in the variable Uninsured2005c.” Also, galasso2022does remarks that “Second, because of the threshold approach that we use to define the treatment and control groups, the control subclasses also include implant patents. In principle, this will cause attenuation bias and lead to an underestimation of the impact of the increase in liability.” }
To elaborate on the nature of the resulting bias due to misclassification, we combine equations ((ref)) and ((ref)) to obtain the following relationship between the $ATT$ and $\theta_{DID}$.
Proposition (ref) shows that the ATT is linearly related to the DID estimand under misclassification, albeit the coefficients are not identified from the observed data. Nonetheless, it is instructive to use this relationship to study the resulting bias.
Let the slope coefficient be denoted by $A=\frac{\mathbb P(D=1)}{\mathbb P(D^*=1)}.$ Also, denote the numerator and denominator of the intercept term by $B=\mathbb E[(Y_1(1)-Y_1(0))\varepsilon \vert D^*=1],$ and $C=P(D=0)$, respectively. Furthermore, suppose that the probability of having a false positive is less than that of having a false negative, i.e., $\mathbb P(D^*=0,\varepsilon=1) < \mathbb P(D^*=1,\varepsilon=1)$. This implies $A<1$. Under these assumptions, Figure (ref) provides a graphical illustration of the bias in the DID estimand under misclassification based on Proposition (ref). The blue line is drawn assuming that $B>0$ and the red line represents the 45-degree line.\footnote{For example, $B>0$ could occur when the average treatment effect is positive for the misclassified treated group.} In this case, the sign of the treatment effect is not identified for certain positive values. That is, when the ATT is positive, the DID takes on the wrong (negative) sign whenever $0<ATT<\frac{B}{C}$. The sign-reversal region is indicated by the gray shaded area in Figure (ref). Further, when the ATT lies between $\frac{B}{C}$ and $\frac{B}{C(1-A)}$, the DID estimand is biased downwards (attenuation bias). The DID estimand yields an expansion bias whenever $ATT>\frac{B}{C(1-A)}>0$.
To summarize, when $\theta_{DID} < \frac{B}{C(1-A)}$, the DID estimand is biased downwards and it is biased upwards when $\theta_{DID} > \frac{B}{C(1-A)}$. When $\theta_{DID}=\frac{B}{C(1-A)}$, the DID estimand is equal to the ATT, even in the presence of misclassification. This latter scenario corresponds to the fixed point in Figure (ref). In Corollary (ref), we provide a sufficient condition under which the fixed point result may occur.
We conclude this section with two special cases under arbitrary forms of misclassification. These special cases do not exhaust the possible misclassification mechanisms researchers may be interested in, but they provide results for important practical scenarios. In the first case, we examine one-sided misclassification and assume that only false positives are present. Unidirectional measurement error in a treatment variable has been studied in previous works nguimkeu2019estimation. An example of one-sided misclassification in DID settings is one discussed in an analysis of the impact of Hurricane Katrina on labor market outcomes of evacuees (groen2008effect). In that study, the treated group was correctly measured in the post-treatment period as those who evacuated due to the storm. However, before the storm, the treated units are subject to misclassification (false positives) because they were determined based on their residence in affected areas. False positives arise because some of those who lived in affected areas before the storm may not have evacuated in the aftermath of the storm, hence not truly treated.
Corollary (ref) shows that even when misclassification is arbitrary, its consequence for DID estimation is less severe when the misclassification is one-sided. In the next section, we provide alternative conditions under which attenuation bias can occur.
In the second special case, we consider what happens when misclassification arises following a Roy selection mechanism such that units are misclassified when their treatment effect exceeds some threshold.
Corollary (ref) implies that if $q$ is positive and the DID estimand $\theta_{DID}$ is also positive, the sign of the ATT is identified as positive. But, it is also possible that $\theta_{DID}$ be negative while the $ATT$ is positive. When the threshold $q$ is “sufficiently” large, we can ensure that the $ATT$ is positive regardless of the value of $\theta_{DID}$. The proof of this corollary follows from Proposition (ref) and Assumption (ref).
We now show how the DID estimand fares under additional assumptions on the nature of misclassification. The literature on measurement error often invokes additional assumptions when the researcher believes that the nature of misclassification is uncorrelated with potential outcomes.\footnote{See bound2001measurement for additional discussions on various types of measurement errors.} We show that if the measurement error is nondifferential and a monotonicity condition holds, then the DID estimand yields an attenuation bias.
Assumption (ref) states that conditional on the true (unobserved) treatment, misclassification is independent of the potential outcomes. This type of measurement error is likely in some empirical contexts, especially when the misclassification stems from uncertainty regarding the timing of treatment (e.g., when the researcher uses a reform's proposed date instead of its date of passage).
Assumption (ref) is equivalent to the well-known monotonicity condition in Hausman_al1998, which states that the sum of false positive and false negative rates of misclassification may not exceed one, i.e., $\mathbb P(\varepsilon=1\vert D^*=1)+\mathbb P(\varepsilon=1\vert D^*=0) < 1$. In Appendix (ref), we prove the equivalence between Assumption (ref) and the monotonicity condition in Hausman_al1998. Assumption (ref) is a minimal requirement to ensure that the misclassification problem is not too severe to render the research project infeasible.
Under Assumption (ref), we have $\mathbb E\left[Y_1(1)-Y_1(0) \vert D^*, \varepsilon \right]=\mathbb E\left[Y_1(1)-Y_1(0) \vert D^*\right]$. Hence, the following corollary holds.
Corollary (ref) shows that the standard DID method under Assumptions (ref), (ref) and (ref) produces smaller effects in magnitude when the treatment is misclassified (attenuation bias).
As mentioned earlier, we now provide a sufficient condition that yields a fixed point result under nondifferential misclassification. Here, the DID method is theoretically unaffected by the presence of misclassification.
Corollary (ref) provides a sufficient condition under which there exists a fixed point in the relationship between the $ATT$ and the DID estimand. At the fixed point, the DID estimand is robust to any misclassification in the treatment variable.
Given the above results, how can researchers overcome the bias from misclassification? In the measurement error literature, there are various approaches to addressing misclassification, including partial and point identification methods. We believe that the appropriate approach for overcoming measurement error in the DID framework is context-dependent, especially given the various ways misclassification may arise. If the researcher is willing to impose a one-sided measurement error structure and maintain parametric assumptions about the error structure, negi2025difference propose a two-step estimator. In the next two subsections, we propose complementary partial identification methods for addressing misclassification under nondifferential and arbitrary misclassification.
If the researcher is willing to invoke the assumption of nondifferential misclassification (Assumption (ref)), we develop bounds on the ATT that allow them to investigate the sensitivity of their results to misclassification using contextual knowledge to infer the extent of misclassification. Specifically, we use the extent of misclassification to bound the ATT by scaling the DID estimate accordingly. We, therefore, introduce the following assumption and derive the bounds in the subsequent corollary.
Corollary (ref) provides bounds on the ATT based on the DID estimand and the upper bound on the extent of misclassification, $\lambda$, which plays the role of a sensitivity parameter. Larger values of $\lambda$ imply wider bounds for the ATT, while smaller values of $\lambda$ yield tighter bounds. In the special case where $\lambda=0$ (no misclassification), the bounds collapse to a point, and the ATT is point-identified as the standard DID estimand. In the next subsection, we develop an alternative partial identification approach to bounding the ATT if the researcher departs from Assumption (ref) and has access to multiple data sources drawn from the same population.
Under arbitrary forms of misclassification, we show how to combine multiple exogenous data sources to partially identify the treatment effect. We introduce the following identifying assumptions.
Assumption (ref) states that $(i)$ the data source is excluded from the outcome and the actual treatment status, and $(ii)$ the potential outcomes and the actual treatment status are independent of the data source conditional on misclassification. However, note that data source $S$ is allowed to influence misclassification $\varepsilon$. This assumption is intuitive and is likely to hold in many settings because survey administrators who collect the data and policymakers are typically separate.
Assumption (ref) states that the misclassification probability for each observed treatment status is changing with the data source. This assumption combined with Assumption (ref) imply that the data source induces some variations in the observed outcome and treatment distributions without changing the potential outcome and potential treatment distributions.
Under Assumption (ref), we show in Appendix (ref) that the following equalities hold.
We have a finite mixture model where mixture weights vary across data sources while mixture distributions do not. We follow HKS2014 and kedagni2023identifying to partially identify the mixture components. See details in Appendix (ref).
To partially identify the ATT, we rely on the following standard parallel trends assumption (stated in terms of the true treatment) for the correctly classified and misclassified groups.
In the following proposition, we use the notation $\alpha^d(s)\equiv \mathbb P(\varepsilon=1\vert D=d, S=s)$.
In this section, we present Monte Carlo simulations to illustrate the consequences of misclassification in DID designs and investigate the usefulness of our proposed bounds. Our simulations cover several data generation processes (DGPs) corresponding to various types (differential and nondifferential) and degrees of misclassification.
All our simulation designs have the same basic structure but differ in terms of the nature of misclassification. We draw potential outcomes based on the two-period model in ((ref)) as follows:
where $Y_{0}(0)$ and $Y_{1}(d), d\in \{0,1\},$ are the pre-treatment and post-intervention period potential outcomes, respectively; $\mathcal U_{[a,b]}$ are continuous uniform random variables on the interval $[a,b]$; $u_j, j\in \{0,1\},$ are standard normal random variables; and $D^*$ is the true, unobserved treatment. To focus on the misclassification problem, all our designs maintain the parallel trends assumption by setting $Y_1(0)=Y_0(0)$.
We generate the true treatment indicators as Bernoulli random variables with a success probability of 0.5. The researcher observes a possibly misclassified treatment variable depending on the misclassification dummy, $\varepsilon$. When misclassification is differential, we generate the misclassification dummy as $\varepsilon = \mathbbm{1}\{Y_1(1)-Y_1(0) >q\}$, where $q$ denotes an appropriately chosen quantile of the distribution of $Y_1(1)-Y_1(0)$ to obtain the desired level of misclassification. For nondifferential misclassification, $ \varepsilon = D^* \mathcal Bernoulli(p)+ (1-D^*)(1-\mathcal Bernoulli(p)),$ where the success probability, $p,$ determines the degree of misclassification.
In both cases (differential and nondifferential), we consider various types of misclassification. We examine the scenario where the misclassification is symmetric or asymmetric across the true treatment states. We refer to misclassification as symmetric when we have equal rates of false negatives and false positives; the errors are asymmetric when those error rates are different. We also study the special case where the misclassification is one-sided, with only false positives or false negatives being present.
For the multi-data source simulations, the nondifferential case uses the same Bernoulli mechanism described above, while the differential case uses an alternative DGP in which the misclassification probability depends on how extreme the individual treatment effect (ITE) gains are. Specifically, let $ITE = Y_{1}(1) - Y_{1}(0)$ denote the individual treatment effect, and define the standardized score $Z = |ITE - ATE|/\sigma_{ITE}$, where $ATE=\mathbb E[ITE]$ and $\sigma_{ITE}$ are the cross-sectional expectation and standard deviation of $ITE$. Individuals with the highest scores are deterministically misclassified: the misclassification indicator is $\varepsilon = \mathbbm{1}\{Z \geq z^*\}$, where $z^*$ is the threshold that yields the desired misclassification rate. We consider the case where each data source $s$ has a different misclassification rate (i.e., 5% vs. 10%).
The observed data is given by $\{Y_{0}, Y_{1}, D \}$, where the outcomes (pre- and post-treatment) and treatment variable are respectively governed by the following observation mechanism. For outcomes, $Y_0 = Y_{0}(0)$ and $Y_1=Y_1(1)D^*+Y_1(0)(1-D^*)$, and for treatment, $D = D^*(1-\varepsilon) + (1-D^*) \varepsilon$.
This section presents our Monte Carlo study results based on samples of size 10,000, which are aggregated across 10,000 iterations. The true treatment effect (ATT) equals 3 in all the experiments. Overall, the simulation results show that DID estimator is severely biased downwards and sometimes yields a sign opposite of the true treatment effect. This general finding aligns with our main theoretical result in Proposition (ref). Even when the parallel trends assumption holds, we cannot generally sign the ensuing bias in DID estimation when misclassification is allowed to be differential.\footnote{Appendix Table (ref) presents simulation results for nondifferential misclassification. Those results show that the DID estimates exhibit an attenuation bias as shown in Corollary (ref) for all levels of misclassification. In addition, the proposed bounds in Corollary (ref) are reported in Column 5 of Appendix Table (ref).}
For differential misclassification (Table (ref)), the DID estimates have the wrong (negative) signs in all instances (Panels A through C) except for the case where only false positives are present. We find that the sign switching of the DID estimate occurs with at least a 20% unconditional misclassification rate in Panels B and C but with at least 30% in Panel A. All else equal, one likely reason why the sign reversal occurs at higher levels of misclassification in Panel A is that the errors are symmetric and the biases (from false negatives and false positives) might cancel out more evenly, leading to a less severe overall bias. The sign-reversal result we obtain in the DID context when the parallel trends assumption holds bears resemblance to the consequences of measurement error in linear treatment effect models under strict exogeneity. For instance, nguimkeu2019estimation shows that, even when treatment is exogenous, the sign of the OLS estimator is not generally identified with endogenous (differential) misreporting.
In the case where the misclassification is solely false positives (Panel D), we find that the DID estimates are attenuated but take on the correct (positive) sign of the true ATT. This finding is consistent with Proposition (ref) by observing that the second term on the right-hand side of equation (ref) is zero when there are no false negatives. The simulation results illustrate the severe consequences of misclassification, but do not exhaust the set of possible outcomes thereof. As shown in Proposition (ref) and the discussion immediately following it, the resulting bias in the DID estimator can take any form---attenuation bias, expansion bias, or sign reversal---depending on various parameters in specific contexts and data generating processes. In summary, the simulations illustrate our theoretical results showing that the consequences of misclassification in DID designs can be severe, sometimes going beyond an attenuation bias to producing the incorrect sign of the treatment effect.
Table (ref) presents simulation-based confidence intervals for the ATT bounds derived from combining two data sources with different misclassification rates. For both the nondifferential case (Panel A) and the differential case (Panel B), the 95% confidence intervals based on Propositions (ref) and (ref) correctly identify the true sign of the ATT, demonstrating that exploiting variation in misclassification rates across data sources yields informative bounds. The multi-data source theoretical results and simulations suggest this partial identification approach may be a viable solution to measurement error if the researcher can access multiple data sources applicable to their research question.
In this section, we illustrate our theoretical results using an empirical application highlighting how misclassification may arise in DID designs. Our objective is to demonstrate the usefulness of our proposed methods and provide guidance on how researchers may use them. The application considers the effect of a 2007 legal reform in Brazil that legislated municipal governments (as opposed to state governments) as the ultimate authority to provide public services in the water and sanitation sector kresch2020buck. A second empirical illustration revisiting how punishment severity affects jury decision-making bindler2018punishment is provided in Appendix (ref).
In January 2007, the Brazilian Congress passed a law (National Water Law 11.447) that gave municipalities the ultimate authority to provide water and sanitation services, eliminating the risk of state governments overtaking municipal-run companies. kresch2020buck uses a DID design comparing self-run municipalities (treatment) to state-run municipalities (control) and finds that this reform led to significant increases in total system investment, access, and decreases in child mortality.\footnote{As in the original study, we maintain that the parallel trends assumption plausibly holds given that the decision to be a self-run or state-run municipality was made in the 1970s, with no companies switching their status during the study period.}
Crucially, kresch2020buck used the date of the proposed legislation (2005) rather than its passage date (2007) to define the post-intervention period, arguing that the bill's likely passage may have shifted investment decisions earlier. We view this choice as introducing potential misclassification in the DID framework.\footnote{One alternative interpretation is that using the proposed date identifies anticipatory effects of the reform rather than introducing misclassification. However, we think this alternate view moves the goalpost and changes the interpretation of the policy-relevant parameter identified by the DID estimand.} Unlike most empirical settings, kresch2020buck can directly compare estimates using both dates. Our proposed bounds provide a complementary sensitivity tool for researchers who lack this advantage.
Table (ref) presents our findings on the effect of the legal reform on investment decisions. We replicate the original analysis without covariates.\footnote{Our results replicate the estimates in Table 3 of the original study, albeit without covariates. Some of the covariates used in the original study include municipality characteristics such as population size, municipality finance measures, and temperature and rainfall variables.} Columns 1 and 2 present the DID estimates using 2005 and 2007 to mark the post-intervention periods, respectively. Both sets of DID estimates show that the legal reform eliminating take-over risk led to a statistically significant increase in overall investments and all types of investments except for government grants and investment in water. While the results across the two legislation dates are qualitatively the same, the estimates using the earlier proposed date are slightly smaller.\footnote{However, we cannot reject the null hypothesis that both sets of estimates are equal at the 5% level of significance.}
Column 3 presents bounds on the ATT based on Corollary (ref) under nondifferential misclassification. We estimate the sensitivity parameter as $\lambda=0.29$ (i.e., $2/7$), reflecting the two-year discrepancy between the proposed and passage dates relative to the seven-year post-intervention period.\footnote{Alternatively, one can report alternative bounds using various rates of misclassification based on reasonable beliefs regarding the extent of misclassification.} The upper limits of the bounds are roughly 40% higher than the DID estimates (based on the proposal date) across most investment types, and in all but one instance the bounds include the DID estimate using the passage date. This finding is reassuring: our proposed bounds provide an informative and complementary way to assess the sensitivity of DID estimates to misclassification when the researcher cannot directly verify robustness to alternative treatment definitions.
The difference-in-differences design remains one of the most widely used quasi-experimental designs in economics and related disciplines. Under a parallel trends assumption, this method identifies the average treated effect on the treated. Part of the DID's appeal is its simplicity and ease of use. The canonical DID design provides a nonparametric estimate of the ATT by comparing the differences in outcomes across two groups and periods. The DID method continues to attract a great deal of attention from methodological researchers along several dimensions. One aspect of the DID design that has received little to no attention is the misclassification of treatment status. When empirical researchers encounter misclassification in their DID analysis, they often assert that its consequences are likely minimal, resulting in an attenuation bias.
This paper investigates the identification of the DID estimand under arbitrary misclassification. We show that the consequences of misclassification are more severe than previously articulated in the literature, with the sign of the DID estimand possibly being different from the true treatment effect. In particular, the DID estimand using a possibly misclassified treatment variable identifies a weighted average of the ATT for the correctly classified and misclassified subgroups, with the weights being negative for the latter.
Under nondifferential misclassification, we find that the DID estimand is attenuated but recovers the correct sign. We propose bounds on the ATT under reasonable assumptions on the extent of misclassification. Researchers may use these bounds for sensitivity analysis when they suspect misclassification of the treatment variable in their DID framework. We also propose a partial identification approach under arbitrary misclassification when the researcher has access to multiple data sources drawn from the same underlying population but subject to different misclassification rates.
We provide simulation evidence to demonstrate our theoretical results and conclude with empirical applications that provide guidance on how to estimate the sensitivity parameter used for our bounding exercise. This paper contributes to the literature by providing new insights on the implications of measurement error in DID designs. Future work can consider extensions of our work to staggered DID designs and consider the inclusion of covariates.