Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
111,118 characters · 23 sections · 171 citation commands
Sensitivity Analysis for Treatment Effects in Difference-in-Differences Models using Riesz Representation
\def\spacingset#1{ {#1}} \spacingset{1}
\if11 { \singlespacing }
\affil[a]{\it School of Business & Economics, Freie Universität Berlin, Boltzmannstr. 20, 14195 Berlin, Germany} \affil[b]{\it Chair of Statistics with Application in Business Analytics, University of Hamburg Business School, Moorweidenstr. 18, 20148 Hamburg, Hamburg, Germany} \affil[c]{\it Economic AI, Nürnberger Str. 262 A, 93059 Regensburg, Bayern, Germany} \affil[d]{\it Düsseldorf Institute for Competition Economics, Heinrich Heine University Düsseldorf, Universitätstr. 1, 20225 Düsseldorf, Nordrhein-Westphalen, Germany }
\noindentKeywords: Sensitivity Analysis, Difference-in-differences, Double Machine Learning, Riesz Representation, Causal Inference
Identification of causal effects in difference-in-differences (DiD) models fundamentally relies on the parallel trends assumption. For example, in the canonical $2\times 2$ design with two periods and two groups, it is assumed that, in the absence of treatment, the expected potential outcomes of both groups would have followed parallel trends over time. In practice, however, this assumption is often only plausible after conditioning on observed pre-treatment confounders. Empirical researchers typically select these covariates based on domain knowledge or economic reasoning relevant to the context of the study. The treatment assignment in empirical DiD studies often arises from the decisions of individual units or groups, such as states or countries choosing to adopt certain policies, in response to economic considerations and other factors, some of which may be unobserved. Empirical researchers try to model these decision processes by accounting for pre-treatment covariates in order to justify parallel trends conditionally on these characteristics. However, identification of the Average Treatment Effect on the Treated (ATT) fails if these variables do not adequately account for all relevant confounding information: The reported ATT estimate will be contaminated by the bias from unobserved self-selection into the treatment groups.
In such settings, it seems natural to question the validity of the conditional parallel trend assumption: Is the researcher really able to account for all systematic selection mechanisms that make certain types of individuals more likely to receive (or rather choose) the treatment than others? Quantifying the implications of such violations can help to assess the robustness of causal findings: If the parallel trend assumption were violated, what would be the resulting bias for the causal estimate? Would this bias be sufficient to substantially change the conclusions of the causal analysis, for example changing the significance of an effect estimate? Such sensitivity considerations are useful for an appropriate interpretation and transparent communication of causal results according to the plausible strength of the parallel trend violation. Building on previous work by chernozhukov2022long and cinelli2020making, we develop a new approach for sensitivity analysis in DiD models with panel data exploiting the Riesz representation for the ATT in the canonical $2\times 2$ DiD setting and group-time specific average treatment effects, $ATT(\mathrm{g},t)$ in multi-period settings with staggered adoption. Our sensitivity approach helps to quantify the implications from omitting one or several pre-treatment confounders in terms of the corresponding explanatory power for the treatment assignment and the observed difference in outcomes over time. Consequently, our framework makes it possible to bound the bias from omitting pre-treatment confounders, adjust the ATT estimators accordingly and to compute critical values, which are also known as “robustness values” cinelli2020making, chernozhukov2022long or “breakdown” values rambachan2023.
Our approach builds on the doubly robust ATT estimator introduced by zimmert2018efficient, chang2020double and sant2020doubly, which is compatible with Machine Learning (ML) nuisance estimators in the Double/Debiased Machine Learning (DML) framework Chernozhukov2018. We derive the Riesz representation for the ATT in the canonical $2\times2$ DiD setting to establish the bias bounds from parallel trends violation and, thus, extend prior work on cross-sectional data by chernozhukov2022long and cinelli2020making. To relate our sensitivity approach to the common practice of pre-testing in event studies, we also present results for group-time specific average treatment effects in multi-period settings with staggered treatment adoption callaway2021difference.
We consider a canonical $2\times 2$ DiD setting with two treatment groups and two periods, in which the conditional parallel trend assumption would be satisfied only if we had access to observed pre-treatment confounders $X_i$ and one or several unobserved confounding variables $A_i$. In general, not taking into account the confounding through $A_i$ will lead to a systematic bias in the estimation of the ATT, which is our target causal parameter $\theta_0$. This corresponds to a situation in which researchers worry about the comparability of the treated and control group prior to treatment onset due to differences in pre-treatment characteristics. These would translate into non-parallel trends of the potential outcomes without treatment over the considered evaluation period. A prominent example where individual characteristics play an important role to establish parallel trends is the famous Lalonde data as re-analyzed in smith2005. In this example, the individual characteristics of participants in the National Supported Work (NSW) experiment are not only important to predict whether individuals enter the sample of the treated or control group, but also for the change of the outcome variable (earnings). smith2005 state that accounting for these time-invariant confounders in a DiD model is crucial to reduce the bias from selection into the treatment groups. In our first empirical example in Section (ref), we apply the suggested sensitivity approach to obtain bias bounds on the ATT in a reevaluation of the data used by smith2005 and lalonde1986. The second empirical application is a replication of draca2011minimum, demonstrating the use of pre-testing information for scenarios of parallel trend violations.
Our theoretical results show that the asymptotic bias resulting from violations of parallel trends is closely related to the explanatory power of the unobserved confounding variables $A_i$ for the treatment assignment probability and the outcome difference over time in addition to $X_i$. We provide asymptotic bounds on the causal parameter $\theta_0$ and corresponding $(1-a)$-confidence limits if we have access to the observed data only, i.e., $Y_{i,t},D_{i,t},X_{i}$. For example, this reflects a setting where measurements for pre-treatment confounding are imperfect proxies, such as variables for educational achievement and income capturing confounding from socio-economic status.
A critical step in sensitivity analysis is to define plausible and realistic scenarios of identification violations. A scenario specifies explicit numeric values for the sensitivity parameters, which enter the asymptotic bias formula that is obtained using the Riesz representation for the causal parameters. In our approach, the sensitivity parameters quantify the strength of the parallel trend violation in terms of the explanatory power of $A_i$ for the treatment assignment and the difference in outcomes. We propose three ways to formulate parallel trends violation scenarios: (1) Include information from pre-testing, which is a common practice in event studies with access to pre-treatment periods, (2) exploiting knowledge from leaving-out known and observed pre-treatment confounders (so-called benchmarking) and (3) standard reporting measures and visualizations such as contour plots. In addition to two empirical applications, we also provide evidence on the validity of our sensitivity framework in systematic simulation studies. We would like to highlight that, to the best of our knowledge, the simulation results are the first thorough numerical experiments underscoring the validity of sensitivity analysis based on Riesz representation chernozhukov2022long. Our results demonstrate that bias bounds for the confidence intervals achieve near-to-nominal empirical coverage. Moreover, in our simulation experiments, we compare the performance of the sensitivity bounds for the ATT to that of the (infeasible) oracle estimator for the ATT, which would be obtained if one had access to all observed and unobserved confounders. Our results show that having knowledge on the explanatory power of the unobserved pre-treatment confounder $A_i$ with regard to the treatment status and outcome difference is equivalent to directly having access to the unobserved confounders, already with moderate sample size. We share an open source implementation of our DiD sensitivity analysis with additional documentation and examples through the DoubleML package for Python DoubleML2022Python, DoubleML.\footnote{More information on DoubleML available at \url{https://docs.doubleml.org}.}
The remainder of this paper is structured as follows. In Section (ref), we briefly review the existing literature on difference-in-differences with a focus on violations of the parallel trend assumption, doubly robust estimation, and sensitivity analysis. Section (ref) introduces the difference-in-differences setting and the major idea of our sensitivity approach for DiD. We do so by considering a motivation example, reviewing the underlying ideas of chernozhukov2022long and then presenting our new approach in the canonical $2\times 2$ DiD model. Section (ref) extends this framework for sensitivity analysis to the multi-period DiD model with staggered adoption. Section (ref) provides simulation studies and Section (ref) provides two empirical applications of our proposed framework. Finally, Section (ref) concludes and provides an outlook on future research.
Difference-in-differences is probably the most frequently used approach to causal inference with observational data, see for example recent textbooks by chaisemartin2023credible, huber2023causal and cunningham2021causal. The recent econometric literature has vastly been impacted by innovative developments in the difference-in-differences literature: New approaches have addressed limitations of classical estimation procedures such as two-way fixed effects (TWFE) in settings with multiple treatment groups and heterogeneous effects, for example including work by goodman2021difference, borusyak2024revisiting, sun2021estimating, de2020two, athey2022design and callaway2021difference. Recent surveys are available in roth2023, chaisemartin2023, and callaway2023difference. A practice-oriented guide to difference-in-differences is provided by baker2025difference. Our study builds on the doubly robust estimator for the ATT in the canonical $2\times 2$ setting as suggested by sant2020doubly and on group-time specific average treatment effects, $ATT(\mathrm{g},t)$, in multi-period designs with staggered adoption as considered by callaway2021difference. The latter has been a seminal study to overcome typical limitations of traditional TWFE estimators by recognizing that causal parameters in complex DiD designs can be modeled as aggregations of possibly many $2$-by-$2$ comparisons. Moreover, new estimation approaches have been suggested building on the property of double robustness or Neyman orthogonality. Orthogonality makes it possible to utilize machine learning (ML) for estimation in the double machine learning framework Chernozhukov2018, sant2020doubly, chang2020double, zimmert2018efficient.
For a long time, empirical economists and econometricians have recognized the importance of the parallel trend assumption for the causal interpretation of empirical DiD results, which is assumed to hold conditionally or unconditionally on observable characteristics. In the canonical setting with one pre- and one post-treatment period, this assumption states that, in expectation, the potential outcomes under no treatment for the treatment and control group develop in parallel over time. However, the parallel trend assumption is untestable. A common practice that is suitable if observations from pre-treatment periods are available, is so-called pre-testing. For example, pre-testing is recommended as part of Pedro Sant’Anna's DiD checklist,\footnote{Available at \url{https://psantanna.com/DiD/checklist.png}.} roth2023 and in the conclusion of baker2025difference. It is worth noting that the parallel trend assumption makes a statement on the development of the expected potential outcome of the treated group in the post-treatment period if the treatment had not occurred, which is a fundamentally unobservable or counterfactual quantity. The idea of pre-testing is to collect evidence on observable counterparts of this unknown average potential outcome from pre-treatment periods: If researchers have access to observations prior to the treatment, it is possible to assess whether the observed average difference in the outcome of the treated and control group are the same in these periods. As there is no treatment effective prior to the actual assignment (by ruling out anticipatory effects), the observed outcome difference should be zero. Otherwise, a significant effect would indicate a violation of the parallel trends assumption in the pre-treatment period under consideration. Pre-testing serves as a plausibility check rather than a proper statistical test of the (untestable) parallel trends assumption. There is no guarantee that evidence suggesting parallel trends prior to the treatment actually correspond to parallel trends over the considered post-treatment period. The same is true vice versa: Significant pre-trends do not necessarily mean that parallel trends are actually violated after the treatment occurred. In the end, the conclusions from pre-testing exercises have to be interpreted in the specific context of an empirical analysis baker2025difference. Recent work by freyaldenhoven2019pre, bilinski2018nothing, roth2022pretest, kahn2020promise addresses limitations of pre-testing due to low power. Moreover, roth2022pretest point to a risk of selection bias, which arises if researchers only evaluate data for which no pre-treatment violations can be rejected.
Sensitivity analysis with regard to parallel trend violations is listed as a recommended step in DiD analysis in baker2025difference. Despite their relevance, sensitivity approaches to parallel trend violations have only recently been developed. A frequently used approach with a focus on pre-testing has been introduced by rambachan2023 who extend prior work by manski2018right.\footnote{For example, the approach of rambachan2023 has been used in callaway2023difference, baker2025difference, chiu2025.} rambachan2023 develop finite-sample and uniformly valid asymptotic bounds when the parallel trends assumption is relaxed. Unlike the point identification of causal parameters under the exact parallel trends assumption, partial identification of a set of causal parameters is obtained under a “bounded differential trends” assumption chaisemartin2023credible. The bounds can be based on user-provided specifications on the shape of the parallel trends violation resulting in a set of different post-treatment trends $\Delta$. The choice of $\Delta$ can be motivated from smoothness assumptions on the differences in trends, pre-testing and their combinations. From a practical perspective, it is appealing that the knowledge gathered from pre-testing can be exploited, for example relative to a multiple of the strongest pre-treatment difference in the parallel-trends in two consecutive periods. By default, this implements a linear extrapolation of pre-treatment violations over the treatment evaluation periods, which for example can be used to point down so-called breakdown values that correspond to a reduction of the reported effects to zero.
An earlier approach to sensitivity analysis with regard to violations of parallel trends is provided by keele2019patterns, who adapt previous sensitivity analysis by rosenbaum2009amplification and rosenbaum2002 to DiD designs. In addition, freyaldenhoven2019pre develop an approach for identification of the causal parameter in a linear panel setting under violations of the parallel trends assumption. They require an additional identification assumption on the confounding relationships between the unobserved confounders, an additional covariate and the outcome variable: The confounding variables are assumed to affect the additional covariate in a similar way as the outcome, but the treatment variable is not allowed to have an impact on the auxiliary variable. huber2024joint develop a joint test for unconfoundedness and conditional parallel trend in a DiD setting based on a testing idea introduced in huber2023testing.
In contrast, our approach builds on sensitivity analysis based on Riesz representation as established in chernozhukov2022long. A detailed comparison of the framework of chernozhukov2022long to Rosenbaum's approach in cross-sectional settings is provided in Appendix E of chernozhukov2022long, which similarly applies for the difference-in-differences setting considered here and in keele2019patterns. In terms of estimation, the methodology in keele2019patterns focuses on a matching approach, whereas we build on the doubly robust estimators of sant2020doubly and callaway2021difference.
Our approach extends the current literature on difference-in-differences and sensitivity to parallel trends violations in various regards. It provides a new set of tools for analyzing the robustness of DiD estimation results to the existence of unobserved pre-treatment confounders in settings with and without pre-treatment periods. rambachan2023 establish bounds based on a user-provided description of parallel trends violations through a specification of smoothness conditions or relative magnitudes to pre-testing violations, leading to a possibly very flexible set $\Delta$. Our bias bounds are motivated by the additional explanatory power of omitted pre-treatment confounding variables relative to the observed variables $X_i$. In analogy to the framework of chernozhukov2022long, we distinguish two different models: A long and a short model, with corresponding values for the identified (long and short) parameters. Accordingly, the model that contains both observed confounders $X_i$ and unobserved confounders $A_i$ is referred to as the long model. This model would have access to all confounding variables that are sufficient to establish conditional parallel trends and, hence, correspond to what is often called an oracle model. In contrast, the short model only has access to the observable data, i.e., $X_i$, $D_{i,t}$ and $Y_{i,t}$ and thus omits $A_i$. A difference to the sensitivity approach by rambachan2023 is that their framework does not assume the existence of such an oracle model. Intuitively, our bias bounds are obtained from a systematic comparison of the long and short parameters as identified by the corresponding models using their Riesz representation, whereas rambachan2023 base their bias bounds on user-provided specification of possible parallel trend violations through $\Delta$, which can be challenging for applied researchers. From our point of view, we consider the existence of a long model as plausible and intuitive to applied researchers in many cases, as it often serves as the starting point for identification in DiD models in empirical studies. For example, a common reason for violations to causal identification in economic applications is the existence of socio-economic status (SES). Usually, empirical researchers employ possibly imperfect proxy measures for SES such as educational attainment, occupation, and income. The sensitivity parameters that we will use to derive the bias from parallel trend violations are defined as (nonparametric) partial $R^2$ values, for example quantifying the share of the residual variation in the outcome difference over time that could be explained by SES in addition to the included and imperfect proxy variables. When applied researchers postulate identification under conditional parallel trends, we believe that this modeling approach often explicitly or implicitly assumes the existence of such an oracle model.
Another difference in our approach compared to rambachan2023 is the inferential framework, which is built on previous work by chernozhukov2022long and cinelli2020making. chernozhukov2022long provide a general framework for analyzing the omitted variable bias in cross-sectional settings for a wide class of target parameters which can be characterized by a so-called Riesz representation chernozhukov2022automatic, chernozhukov2021automatic, riesznet. Accordingly, the estimator of interest is obtained as a solution to a moment equation which can be represented by a linear functional containing two terms: The conditional expectation of the outcome variable and the so-called Riesz representer. The latter models the relationship between the covariates and the treatment variable, such as the Horvitz-Thompson transform in augmented inverse probability weighting or Frisch-Waugh-Lovell partialling out of the covariates from the treatment variable in a partially linear regression model. We review the general sensitivity framework of Chernozhukov2018 in Section (ref). Building on the initial work on linear regression by cinelli2020making, chernozhukov2022long establish asymptotic bounds on the omitted variable bias and coverage guarantees for confidence bounds in a variety of causal models.
As described before, the bias bounds for the causal parameter and confidence intervals are a function of the sensitivity parameters, which characterize the strengths of the unobserved confounding relationships. An appealing feature based on this modeling approach is that it makes it possible to obtain bounds on the ATT parameter even if no pre-treatment periods are available for pre-testing. In these settings, it would be very difficult to plausibly specify the set of possible parallel trend violations $\Delta$ in the approach by rambachan2023. In contrast, our approach allows to leverage the explanatory power of one, several or all observed pre-treatment confounders in our framework to inform violation scenarios. In so-called benchmarking, these variables are left out from the model, which is used to compute empirically grounded values for the sensitivity parameter in the bias formula. Finally, in analogy to the breakdown values in rambachan2023, our approach is compatible with the standard reporting statistics of cinelli2020making and chernozhukov2022long, which inform researchers of how strong unobserved confounding would have to be to cause a reduction of the reported estimate to zero.
Our paper contributes to the existing literature on sensitivity analysis in difference-in-difference models in various regards. First, we establish new results on the asymptotic bias from parallel trends violation for the $ATT$ in the canonical $2\times2$ DiD setting as well as on $ATT(\mathrm{g},t)$ parameters in multi-period settings with staggered adoption. Our bounds quantify the bias as a function of the explanatory power of the unobserved pre-treatment confounders in terms ot the treatment status and the outcome difference. Second, we propose practical approaches to parameterize the parallel trend violation scenarios. Our approach is compatible with the common practice of pre-testing, but remains applicable even if no-pretreatment periods are available. To the best of our knowledge, our third contribution is the first systematic evaluation of sensitivity analysis based on Riesz representation, providing supportive evidence and facilitating the interpretation of its properties, such as the empirical distribution of bias bounds, their comparison to oracle (long) estimates and the empirical coverage of confidence sensitivity bounds. Finally, we demonstrate the use of our sensitivity approach in a $2\times2$ DiD setting based on lalonde1986 and smith2005 and a multi-period setting in a reassessment of draca2011minimum.
In this section, we motivate and introduce our sensitivity approach in the canonical $2\times 2$ DiD setting with two periods and two treatment groups (treated and control). In Section (ref), we will extend the sensitivity analysis to group-time average treatment effects in multi-period settings with staggered adoption.
Before introducing the formal sensitivity framework in the following sections, we briefly illustrate the main ideas in a motivation example.
We consider the common situation that researchers face in empirical applications of difference-in-differences: Given a set of observed pre-treatment confounders $X_i$, the researcher might worry about unobserved confounding, which is not accounted for by $X_i$. As a consequence, the conditional parallel trend violation would be violated leading to a possibly substantial bias of the causal estimate. We can imagine two different models: A feasible model (denoted as the “short” model in chernozhukov2022long) with access only to the observed pre-treatment confounders $X_i$ and an infeasible (or “long”) model, which would account for $X_i$ and $A_i$. The short and the long model identify the short parameter $\theta_s$ and the long parameter $\theta_0$, which is the true target parameter, respectively. The target parameter in the $2\times 2$ DiD setting is the ATT defined as
Here, the treatment variable $D_{i,t}=1$ indicates that unit $i$ is treated before time $t$ (otherwise $D_{i,t}=0$). Since $D_{i,0}=0$ for all $i$, we define $D_{i}:=D_{i,1}$. Furthermore, $Y_{i,t}(0)$ denotes the potential outcome of unit $i$ at time $t$ if the unit did not receive treatment up to time $t$ and analogously $Y_{i,t}(1)$ denotes the potential outcome of unit $i$ at time $t$ if the unit did receive treatment. Our goal is therefore to make statements on the expected bias of the short parameter as compared to the true ATT,
As we will show later, the magnitude of this bias will depend on the explanatory power of the unobserved pre-treatment covariates additionally to the observed $X_i$. Figure (ref) illustrates the consequences of a parallel trend violation due to omitting the unobserved pre-treatment confounder $A_i$ in the short model. The results show the DML point estimates for the ATT and two-sided $90\%$-confidence intervals obtained from the long (colored orange) and short model (blue) in a simulated data example. The example is based on a modified version of a data generating process (DGP) from sant2020doubly, which is further explained in Section (ref). In practice, researchers would not be able to know whether the estimated ATT, $\widehat{\theta}_s \approx 5.338$, is close to its true value or not, which is $\theta_0=5.0$ in the example. Applying our suggested approach for sensitivity analysis would result in a lower bound of the ATT of around $\hat{\theta}_{-}=5.059$, which is substantially closer to the true value. Moreover, the standard $90\%$-confidence intervals for the point estimate from the short model do not cover the true ATT. In contrast, we can observe that the one-sided $90\%$-confidence sensitivity bounds derived from our sensitivity approach exhibit coverage of the true effect. The example illustrates that effectively using sensitivity analysis helps to improve the quality of causal statements in DiD studies when researchers are worried about the validity of the parallel trends assumption. As we will later emphasize in our empirical examples, a key ingredient to sensitivity analysis is the proper definition of plausible scenarios of parallel trends violations. In the motivating example, we employed the population values for the underlying sensitivity parameters, which we calibrated in our data generating process.
We will rely on a notation similar to sant2020doubly. Let $Y_{i,t}$ be the outcome of interest for unit $i$ observed at time $t\in\{0,1\}$. The observed outcome for unit $i$ at time $t$ is determined by the treatment status,\footnote{Note that the “switching” Equation (ref) for the observed outcome incorporates a no-anticipation assumption stating that $Y_{i,t-1}(1)=Y_{i,t-1}(0)$ which excludes an effect of the treatment variable prior to the realization of $D_i=1$ in period $t$, i.e. $D_{i,t}=1$, see for example callaway2023difference.}
Further, let $X_i$ be a vector of observed pre-treatment covariates for unit $i$. Moreover, one or multiple pre-treatment confounders, $A_i$, are unobserved. Throughout, we work in a panel setting and assume that the data $W_i = (Y_{i,0}, Y_{i,1}, D_i, X_i, A_i)$ are i.i.d. across units $i \in \{1,\dots,n\}$. Again, the target parameter of interest is the average treatment effect on the treated (ATT)
It is useful to define the first difference in observed outcomes over time, $\Delta Y_i := Y_{i,1} - Y_{i,0}$. In the following, we may occasionally abstract from the index $i$ to keep the notation simple. Estimation of the ATT parameter can be based on different approaches, including inverse probability weighting, outcome regression or doubly robust approaches sant2020doubly, abadie2005semiparametric. Our sensitivity approach builds on double machine learning (DML). We define the nuisance parameters $m(x,a)$, denoting the propensity score, and $g(d,x,a)$, denoting the outcome difference conditional on treatment status $d$ and covariates $(x,a)$ as
Identification of the ATT is based on the following standard assumptions.
Assumption (ref) states that, in expectation, the change in the potential outcomes without treatment would be the same in the treated and untreated group, conditional on all observed and unobserved pre-treatment covariates $X$ and $A$. Identification is therefore compatible with time trends that are specific to the values of $X$ and $A$. Assumption (ref) is a commonly imposed overlap assumption that requires a positive fraction of treated individuals and that the propensity score is restricted to values strictly smaller than $1$.
We focus on estimation using Double/Debiased Machine Learning (DML) as generally established in Chernozhukov2018 and adapted to estimation of the ATT in the $2 \times 2$ DiD setting by chang2020double, zimmert2018efficient and sant2020doubly. DML relies on three key ingredients: (1) Neyman orthogonality, (2) high-quality machine learners, and (3) sample splitting. Under these three requirements, the DML estimator is consistent and asymptotically normal. Asymptotically valid confidence bounds are provided by Chernozhukov2018. The DML estimator for the ATT is the solution to the orthogonal moment condition and corresponds to the doubly robust estimator in sant2020doubly
with nuisance components $\eta=(m, g)$ being estimated by machine learning. For the sake of notational simplicity, we abstract from in-sample normalization, which is commonly implemented as a finite-sample adjustment. The score function and sensitivity results for the case with in-sample normalization are presented in Appendix (ref). Similarly, we abstract from a dedicated notation to highlight out-of-sample predictions, but rather assume that all nuisance predictions are obtained from hold-out partitions from the data to safeguard against overfitting-induced bias Chernozhukov2018.
Before we establish new bias bounds for the point estimator of the ATT and confidence intervals in the canonical DiD setting in Section (ref), we first give a brief review of the general approach for omitted variable bias under violations of the unconfoundedness assumption in cross-sectional settings as established by chernozhukov2022long.\footnote{Unconfoundedness is also referred to as selection-on-observables, conditional ignorability or conditional exogeneity.} chernozhukov2022long extend previous results on sensitivity analysis for the average treatment effect (ATE) in a linear regression model in cinelli2020making to a class of estimators that can be characterized as a linear functional of the conditional expectation of the outcome variable and a so-called Riesz representer,
In the cross-sectional setting, $g(W)$ generally refers to the conditional expectation of the outcome variable. $W$ denotes the data. Furthermore, $\alpha(W)$ is the so-called Riesz representer (RR). The Riesz representer plays a key role for debiasing and implements Neyman orthogonality either through a known analytical expression or through an approximation by an automatic estimation algorithm chernozhukov2021automatic, chernozhukov2022automatic.
According to chernozhukov2022long, $\theta_0$ can be identified by exploiting the Riesz representation $\mathbb{E}[g(W) \alpha(W)]$.
Identification of $\theta_0$ is feasible only if $X$ and $A$ were observed, which, however, is infeasible in an empirical analysis. Hence, chernozhukov2022long establish asymptotic bias bounds on the resulting omitted variable bias quantifying the deviation of the short model from the long model. The corresponding quantities for the short model are defined as $g_s(W^s)$ and $\alpha_s(W^s)$ with $W^s=(Y,D,X)$. Accordingly, the bias can be expressed as
Importantly, $\alpha(W^s)$ is the projection of $\alpha$ given the short data,
chernozhukov2022long show that the magnitude of the bias depends on the values of the sensitivity parameters, which quantify the difference of the short from the long quantities,
with $S^2 = \mathbb{E}[(Y-g_s)^2] \mathbb{E} [\alpha_s^2]$ being a scaling factor that can be estimated empirically and $\rho^2=\text{Cor}^2(g-g_s, \alpha-\alpha_s)$ quantifying the correlation between the residual confounding in the outcome regression and the Riesz representer. For conservative bounds, $\rho$ is set to a value of $\rho=1$. Importantly, the sensitivity parameters $C_Y^2$ and $C_D^2$ quantify the proportion of residual variance in $Y$ and $\alpha$, respectively, that can be attributed to $A$,
with $R^2_{V\sim W}=\frac{\mathrm{Var} (\mathbb{E}[V|W])}{\mathrm{Var} (V)}=\frac{\mathrm{Var}(V) - \mathbb{E}[\mathrm{Var}(V|W)]}{\mathrm{Var}(V)}$ denoting the nonparametric $R^2$ chernozhukov2022long. The interpretation of the sensitivity parameter $C_Y^2$ is often similar for different causal models and parameters. However, since $C_D^2$ refers to the interpretation of the Riesz representer $\alpha$, which differs across causal models and parameters, the direct interpretation of $C_D^2$ has to be adapted accordingly.
According to Theorem 4 in chernozhukov2022long, the DML plug-in estimators for the lower and upper bounds of the target parameter, $\hat{\theta}_{\pm}=\hat{\theta}_s \pm |\rho| C_Y C_D \hat{S}$ are asymptotically normally distributed and centered around their population counterparts, $\theta_{\pm}$. Moreover, the result gives rise to a coverage property of the corresponding asymptotic one-sided confidence bounds, $\ell_-$ and $u_+$, such that $P(\theta_- \ge \ell_-)\rightarrow 1-a$ and $P(\theta_+ \le u_+) \rightarrow 1-a$ given a significance level $a$. Evaluating the bias formula in Equation (ref) with the oracle values for the sensitivity parameters $\rho$, $C_Y^2$, and $C_D^2$, it is possible to identify the absolute value of the bias. However, these values are generally unknown and researchers have to find plausible, ideally empirically grounded choices to obtain bounds on the bias. These can be obtained from domain expertise and/or benchmarking exercises that take advantage of the explanatory relationships between $Y$, $D$ and the observed covariates $X$ in the data.
The results in chernozhukov2022long refer to identification under the unconfoundedness assumption in cross-sectional setting. In this section, we extend their work to sensitivity analysis with regard to violations of the conditional parallel trend assumption and focus on the case of panel data. Extending our sensitivity approach to repeated cross-sectional data as considered in sant2020doubly would be straightforward. This case is covered in the user guide of the DoubleML library for Python with an implementation also provided by DoubleML.\footnote{More information available at \url{https://docs.doubleml.org/stable/guide/sensitivity.html}.} Appendix (ref) also presents the Riesz representer in the case of in-sample normalization, which is commonly implemented in practice. We first derive the Riesz representation for the ATT in the canonical $2\times2$ DiD setting. Importantly, other than in the cross-sectional data settings, the conditional expectation in the DiD model with panel data refers to the difference in the outcome over time, $\Delta Y$,
where $g(W)=\mathbb{E}[\Delta Y|D,X,A]$ and $W=(\Delta Y,X,D,A)$. The Riesz representer for the ATT is given by
with $m(X,A)= E[D|X,A]$ and $p=P(D=1)$. Hence, the ATT corresponds to
with $\mathcal{M}(W,g):=D/p (g(1,X,A) - g(0,X,A))$. Details are provided in Appendix (ref). It is worth noting that $\theta_0$ is only identified if we would observe $A$ because Assumption (ref) is assumed to only hold conditionally on $X$ and $A$. In practice, we can only work with the observed data $W^s=(\Delta Y,X,D)$. Hence, the short parameter $\theta_s$ is given by
with $g_s(W^s):=\mathbb{E}[\Delta Y|D,X]$. Consequently, we can apply the sensitivity methodology from chernozhukov2022long and obtain the resulting bounds for the bias
with $S^2:=\mathbb{E}[(\Delta Y-g_s(W^s))^2]\mathbb{E}[\alpha_s^2(W^s)]$. The interpretation of the corresponding sensitivity parameters
leads to interesting results. The strength of confounding generated in the outcome regression $C^2_{\Delta Y}$ directly takes the form
which measures the proportion of residual variation in the differenced outcomes $\Delta Y$ that can be explained by the unobserved pre-treatment confounders $A$ in addition to the observed covariates $X$. This directly refers to the effect of violating the conditional parallel trend assumption. Furthermore, the strength of confounding generated in the treatment $C^2_D$ can be rewritten as
This shows that $C_D^2$ can be interpreted as the relative increase in the average odds of entering the treatment group due to $A$ after taking into account $X$. In the implementation and application of our sensitivity framework, we use a modified version of $C^2_D$ that is bounded by $0$ and $1$ Bach2024Sensitivity,
Here, $\widetilde{C}^2_D$ can be interpreted as the relative decrease in the average odds of being in the treatment group due to only observing $X$ but not $A$. Again, technical details are provided in Appendix (ref).
The sensitivity parameters emphasize the role of the pre-treatment confounders for identification in difference-in-differences. In order to identify the ATT, it is crucial to account for pre-treatment confounding such that the parallel evolution of the expected potential outcome under no treatment can be justified. The bias from missing information in this regard by not observing $A$ is proportional to the share of the variation in the outcome difference $\Delta Y$ that can be attributed to the unobserved pre-treatment confounding through $A$. For the treatment variable, the confounding variables help to better separate between treated and untreated individuals, as reflected by the odds ratio in the definition of $\widetilde{C}_D^2$. The larger the explanatory power of $A$ for selection into the treatment group, the larger will be the corresponding bias for the ATT.
In this section, we establish sensitivity analysis for causal parameters in the multi-period DiD setting with staggered adoption as, for example, considered by callaway2021difference.
In the following, we use a notation and identifcation assumptions that are based on callaway2021difference. In their paper, callaway2021difference show that group-time specific causal effect parameters in a multi-period setting with staggered adoption can be expressed as binary comparisons of the considered treatment and control group. In our presentation, we slightly adjust the notation to better fit into the common naming conventions in the Double/Debiased Machine Learning literature, sometimes slightly abusing notation. The framework is a generalization of the previously presented $2\times2$ DiD setting. As before, we focus on panel data to abstract from complex notation in the case of repeated cross-sectional data.
We consider $n$ observational units at time $t = 1, \ldots, \mathcal{T}$. The treatment status for unit $i$ in period $t$ is indicated by the binary variable $D_{i,t}$. In settings with staggered treatment adoption, it is common to define treatment groups according to the first period after treatment receipt, which requires additional notation. We focus on the case of an absorbing treatment status, i.e., if individual $i$ is treated first in period $\mathrm{g}$, the individual will remain treated until the final period $\mathcal{T}$. The variable $G_i^\mathrm{g}$ is an indicator variable that takes value one if $i$ is treated for the first time in period $\mathrm{g}$, $G_i^\mathrm{g}=\mathds{1}\{G_i=\mathrm{g}\}$ with $G_i$ referring to the first post-treatment period. In the setting with absorbing treatment, we have $D_{i,t} = 1,$ $\forall t \ge G_i$ almost surely (cf. Assumption 1 in callaway2021difference). If individuals are never exposed to the treatment, we define $G_i = \infty$ and $G_i^\mathrm{g}=0, \forall t = 1, \ldots, \mathcal{T}$. We define $Y_{i,t}(0)$ as the potential outcome of individual $i$ in period $t$ if no treatment has been assigned until period $\mathcal{T}$. We summarize the assumptions as follows.
The target causal parameters are defined in terms of differences in potential outcomes. The potential outcome of an individual $i$ that has been treated first in period $\mathrm{g}$ can be evaluated in period $t$ by
As a measure of the average causal effect of the treatment, it is common to define a group-time average treatment effect parameter, $ATT(\mathrm{g},t)$. This target parameter quantifies the average change of the potential outcomes due to being treated first in period $\mathrm{g}$ as evaluated in period $t$,
In the $2\times 2$ DiD setting, the counterfactual average outcome for the treated group (which would have realized had the group not received the treatment) is estimated based on the information from the untreated group. This is valid under the conditional parallel trend assumption, which ensures that conditional on the pre-treatment covariates, the expected average outcome difference over time is the same for the treated and untreated. However, in multi-period DiD settings with staggered adoption, there is no longer one unique definition of the control group, whose information can be exploited for identification of the causal parameters. To characterize the control groups in line with the literature, we define an indicator variable $C_{i,t}$, which depends on whether never treated or not yet treated units are used for comparison.
As a consequence, the parallel trend assumption will be adapted to the control group under consideration. To account for anticipation effects, we will first introduce the limited anticipation assumption of callaway2021difference using an anticipation parameter $\delta$.
Assumption (ref) a. assumes that, in the absence of treatment, the expected outcome for the group treated first in period $\mathrm{g}$ would have evolved in parallel over the time periods considered as compared to the group that never received the treatment. This assumption is qualitatively different from Assumption (ref) b., which imposes parallel trends of the group treated first in period $\mathrm{g}$ as compared to groups that have not yet been exposed to the treatment at time $t+\delta$ callaway2021difference. Whether never treated or not yet treated are used as a control group depends on the empirical context of an analysis and the overall (causal) research question. Often, only one of these groups is available or relevant for causal evaluations. In addition to the previous assumptions, an overlap assumption has to be made in the multi-period DiD setting to achieve identification of the group-time specific causal parameters.
Identification of the long parameters, i.e., the $ATT(\mathrm{g},t)$ in Equation (ref), is granted by Assumptions (ref) to (ref) callaway2021difference. Again, we note that identification of these target parameters require accounting for $X$ and $A$ in order to justify that the trends in the average outcomes between the defined control and treatment groups develop in parallel over the evaluation period under consideration.
It is possible to extend the machine-learning based estimation of the $ATT$ parameter presented in Section (ref) to the muli-period DiD model with staggered adoption as presented by callaway2021difference. Note that the corresponding nuisance functions depend on the control group used for the estimation of the target parameter. By slight abuse of notation we use the same notation for both control groups $C_{i,t}^{(nev)}$ and $C_{i,t}^{(nyt)}$. More specifically, the control group only depends on $\delta$ for not yet treated units.
with nuisance elements $\eta_0=\big(g_{0, \mathrm{g}, t_\text{pre}, t_\text{eval}}, m_{0, \mathrm{g}, t_\text{eval}}\big)$. Here, $g_{0, \mathrm{g}, t_\text{pre}, t_\text{eval},\delta}(\cdot)$ denotes the population outcome change regression function for the control group that is specified to evaluate the causal effect for the group treated first in $\mathrm{g}$ over the pre-period $t_\text{pre}$ and evaluation period $t_\text{eval}$. $t_\text{pre}$ is also denoted as the base period and often specified as the last period before a group received the treatment for the first time, $t_{\text{pre}}= g -\delta - 1$. Furthermore, $m_{0, \mathrm{g}, t_\text{eval} + \delta}(X_i, A_i)$ is the generalized propensity score. For notational purposes, we will omit the subscripts $\mathrm{g}, t_{\text{pre}}, t_{\text{eval}},\delta$ and refer to the corresponding functions by the simplified versions
Note that for estimation of the causal parameters in DiD models, it suffices to estimate the conditional expectation of the outcome differences for the specified control group as indicated by conditioning on $C_{i,t_{\text{eval}+\delta}}=1$ in Equation (ref). However, for sensitivity analysis, we require additional estimation of the conditional outcome difference for the group treated first in period $\mathrm{g}$, which we define as
machine-learning based estimation of the $ATT(\mathrm{g},t)$ parameters can be based on the multi-period analog of the Neyman-orthogonal score for the $ATT$ in the canonical DiD setting in Equation (ref), \footnote{Again, we abstract from in-sample normalization in our presentation. The score function and Riesz representation with in-sample normalization are presented in Appendix (ref).}
To extend the sensitivity framework from Section (ref), we derive the Riesz representer for the multi-period DiD setting with staggered adoption. In this setting, the moment equation is
where $\max(G^{\mathrm{g}}, C^{(\cdot)})$ is an indicator that takes value one if an individual belongs to the treated or the specified control group. Including this indicator ensures that only observations from the relevant treatment and control groups are used for identification. When we later present the sensitivity parameters, it is important to account only for variation in the data that is relevant for the target causal parameter. This gives rise to the Riesz representer for the $ATT(\mathrm{g},t)$
To have a compact presentation of the sensitivity parameters we refer to the outcome difference that is estimated for evaluation of the $ATT(g,t)$ parameter by $\Delta_{t_{\text{pre}},t_{\text{eval}}}Y = Y_{t_{\text{eval}}} - Y_{t_{\text{pre}}}$.
Accordingly, the sensitivity parameters are given by
and, for the selection into the treatment groups,
As before, it is useful to use the transformed version $\widetilde{C}_D^2$.
The sensitivity parameters quantify the explanatory power of the unobserved confounders for the considered outcome difference and the treatment assignment, which is analguous to their interpretation in the $2\times 2$ case. However, we have to restrict attention to only those observations in the data that are relevant for the binary comparison in the multi-period DiD setting. This is done by conditioning on the indicator $\max(G^\mathrm{g}, C^{(\cdot)})$.
$C^2_{\Delta Y}$ is then measuring the share of the residual variation in the difference of the outcome variable over the period $t_{\text{pre}}$ to $t_{\text{eval}}$ for the group that received treatment first in period $\mathrm{g}$, $\Delta_{t_{\text{pre}},t_{\text{eval}}}Y$, which can be explained by accounting for the unobserved pre-treatment confounders $A$. $\widetilde C^2_D$ measures the relative decrease in the odds ratio to receive the treatment for the first time in period $\mathrm{g}$ due to not accounting for the explanatory power of $A$.
The previous sections provided a bias formula for causal parameters in difference-in-differences models to quantify violations of the conditional parallel trend assumption. Once we plug in values for $C_D^2$ and $C_{\Delta Y}^2$, we obtain an upper and lower bound for the ATT and the $ATT(g,t)$ parameters, respectively. Moreover, we can exploit Theorem 4 of chernozhukov2022long, which makes it possible to establish coverage guarantees of the confidence bounds. Heuristically speaking, plugging in the population values for the sensitivity parameters in the bias formula in Equation (ref) is asymptotically equivalent to having direct access to the $A$ in the oracle model.
In empirical applications, researchers cannot use the long data $W=(Y,D,X,A)$, nor can they know the population values for the sensitivity parameters. As they are concerned with the presence of unobserved pre-treatment confounders $A$, they are faced with choosing values for $C_D^2$, $C_{\Delta Y}^2$, and $\rho$ that are plausible for their specific empirical data setting and, ideally, close to the population values of the sensitivity parameters.\footnote{It is generally recommended to start with the most conservative value $\rho=1$.} We denote the choice of specific values for $C_D^2$ and $C_{\Delta Y}^2$ as parametrizing a violation scenario of the conditional parallel trend assumption (or shorter, simply violation scenarios), for which the asymptotic bias is estimated. We suggest four practical ways, in which empirical researchers can obtain plausible values for $C_D^2$ and $C_{\Delta Y}^2$: First, domain expertise can help to quantify the plausible explanatory power of unobserved pre-treatment confounders for $\Delta_{t_{\text{pre}},t_{\text{eval}}}Y$ and $G^\mathrm{g}$. Second, standard reporting and visualizations can serve as informative measures, which might be complementary to domain expertise. Third, information from pre-treatment periods might be exploited, following the rationale of pre-testing in event studies. Fourth, so-called benchmarking, i.e., leaving out observed pre-treatment confounding variables, helps to obtain empirically grounded choices for $C_D^2$ and $C_{\Delta Y}^2$. In the end, the definition of one or more violation scenarios can be the result of combining these four procedures.
cinelli2020making and chernozhukov2022long develop a suite of standard reporting measures in their sensitivity framework, which can also be applied in difference-in-differences models. The so-called robustness values $RV$ and $RV_a$ indicate critical scenarios of violations of the identifying assumptions: In the difference-in-differences setting, $RV$ corresponds to the minimum strength of violating the conditional parallel trends assumption that suffices to set the ATT or, respectively, the $ATT(\mathrm{g},t)$ estimate equal to zero or an alternatively chosen null hypothesis. $RV_a$ accounts for estimation uncertainty and reports the minimum violation scenario that is sufficient to make the causal parameter non-significant, hence $RV \ge RV_a$ in general. Note that the robustness values always refer to symmetric violation scenarios, $C_{\Delta Y}^2 = \widetilde C_D^2$. The bias bounds for the point estimate or the confidence limits can be illustrated in contour plots, as for example presented in Figure (ref) in Section (ref). A contour line indicates all combinations of $C_{\Delta Y}^2$ and $\widetilde C_D^2$ that correspond to the same value of the bias-adjusted estimate.
Pre-testing is a common practical procedure to assess violations of parallel trends in difference-in-differences models with multiple periods. Whereas no formal statistical test exists for testing a violation of parallel trends over the relevant time periods for estimation of the causal parameter, the underlying idea is to provide evidence from pre-treatment periods. In absence of an actual treatment, a significant causal effect for one or several pre-treatment comparisons might cast doubt on the validity of conditional parallel trends in the relevant treatment periods. However, not rejecting the null hypothesis of a zero effect in the pre-treatment period does not serve as evidence for the validity of the identification assumptions. In the end, the conclusions from pre-testing evidence are dependent on the context of the empirical application.
It is possible to use pre-testing information to empirically support the choice of sensitivity parameters $C_D^2$ and $C_{\Delta Y}^2$. Here, we focus on sensitivity analysis for the average treatment effect for one treatment group that has received the treatment first in period $\mathrm{g}$ being evaluated in the first post-treatment period, i.e., $ATT(\mathrm{g}, \mathrm{g})$ with pre-treatment period $\mathrm{g}-1$. Suppose, we have access to $\mathcal{T}_{\text{pre}}$ pre-treatment periods, i.e, we have periods $1,\ldots, \mathcal{T}_{\text{pre}}, \mathrm{g}, \mathrm{g}+1, \ldots, \mathcal{T}$. Pre-testing makes it possible to obtain $\mathcal{T}_{\text{pre}}-1$ placebo effects $\hat{\theta}_{2}, \ldots, \hat{\theta}_{\mathcal{T}_{\text{pre}}}$. Because there is no treatment in the placebo periods and anticipatory effects are ruled out by assumption, any measured pre-testing effect is the result of a parallel trends violation roth2022pretest, rambachan2023. Hence, it is possible to bound the bias in the pre-treatment periods by
The robustness value for each pre-treatment period $t$, $RV_t$, indicates the values $C_{\Delta Y}^2$ and $\widetilde C_D^2$ that would suffice to reduce (or rather correct) the bias from a conditional parallel trend violation to zero. In our application in Section (ref), we employ a conservative rule that takes the maximum robustness values from all pre-treatment periods (or a $k$-fold multiple thereof, $k>0$):\footnote{It might be that the sign of the pre-treatment estimates is different than that of the suspected bias. In that case, the rule might be adjusted to select the maximum among all pre-treatment violations in the suspected direction of the bias.}
An additional way to obtain specific values for the sensitivity parameters, which is applicable also if no pre-treatment periods are available, is so-called benchmarking. Leaving out observed confounding variables has previously suggested for example in cinelli2020making and chernozhukov2022long who build on prior work by imbens2003sensitivity, altonji2005selection, and oster2019unobservable. Adapting the idea of leaving-out-confounders from the cross-sectional setting, we can mimic the omission of unobserved pre-treatment confounders by leaving out one or multiple variables from $X$. Whereas there is no formal guarantee that the actual confounding variable $A$ shares the same explanatory power with the benchmarking variable $X_j$, these modeling exercises are very useful for empirical researchers: In many cases, domain experts and empirical researchers know about the role of the most important pre-treatment covariates, which are crucial to justify the conditional parallel trend assumption. Leaving these variables out can be helpful to judge the plausibility of violation scenarios based on the empirical evidence.
Benchmarking works as follows: Let $X_j$ denote the benchmarking pre-treatment covariates, leaving only $X_{-j}=X \setminus \{X_j\}$ for estimation of the nuisance functions $g(\cdot)$ and $m(\cdot)$. Then it is possible to recompute the sensitivity parameters that correspond to the change in the estimate of the ATT or, respectively, $ATT(\mathrm{g},t)$ that is caused by removing $X_j$ from the proxy model. We denote these calibrated sensitivity parameters as $\widetilde C_D^{2,\text{bench}}$ and $C_{\Delta Y}^{2,\text{bench}}$. It is necessary to incorporate a correction factor $\kappa$ that adjusts for the change in the residual variation that is left after removing $X_j$ from $X$ in the proxy model, see also Appendix F of chernozhukov2022long.
In this section, we report the results of a simulation study to assess the finite-sample performance of our sensitivity framework. To the best of our knowledge, these are the first systematic simulation results for sensitivity analysis based on Riesz representation. We extend a data generating process (DGP) from sant2020doubly for the canonical $2\times 2$ DiD setting in terms of an unknown confounder $A$. As we will clarify in the following, we implement a confounding setting with known population values for the sensitivity parameters $\rho, C_{\Delta Y}^2$, and $\widetilde C_D^2$. This makes it possible to assess the empirical performance of our sensitivity approach in four regards: First, given the oracle values for the sensitivity parameters, we can estimate the lower and upper bias bounds for the ATT as obtained from Equation (ref), $\hat{\theta}_-$ and $\hat{\theta}_+$, and compare them to the true parameter value, $\theta_0$. In the scenario considered, the parallel trend violation implements an upward bias of the short ATT parameter, such that we focus on the evaluation of the lower bias bound. Second, we can evaluate the empirical performance of the robustness values $RV$ and $RV_a$, which are expected to be close to the oracle values of $C_{\Delta Y}^2$, and $\widetilde C_D^2$.\footnote{We implemented a symmetric parallel trend violation scenario, such that the population values are calibrated to $C_{\Delta Y}^2=\widetilde C_D^2=0.1$, which implements a nominal level for the robustness values $RV$ and $RV_a$ of $0.1$ .} Third, we can evaluate the performance of the lower and upper bias bounds for the confidence limits, $\hat{\ell}_-$ and $\hat{u}_+$. Evaluating the sensitivity bounds at the population values for $C_{\Delta Y}^2$ and $\widetilde C_D^2$, the one-sided bounds are expected to cover the true ATT in $1-a$ percent of the data realizations. Lastly, we can compare the bias bounds to the long estimate of the ATT, $\hat{\theta}_{\text{long}}$, which is only feasible in a simulation setting. Unlike in real data analysis, we can use the simulated pre-treatment confounder $A$ to compute the DML estimate that would be obtained from the long data. This is informative to assess the usability of the sensitivity bounds according to the sensitivity parameters as compared to having access to the unobserved pre-treatment confounder $A$.
We consider an adapted version of the Monte Carlo simulation considered in sant2020doubly. Let $X=(X_1, X_2, X_3, X_4, X_5)^T \sim \mathcal{N}(0,\Sigma),$ where $\Sigma$ corresponds to the identity matrix. For $j=1,2,3,4,5$, define $Z_j=(\widetilde{Z}_j-\mathbb{E}_n[\widetilde{Z}_j])/\mathrm{Var}(\widetilde{Z}_j)$, where
For generic $V=(V_1, V_2, V_3, V_4, V_5)^T$, define
Using only the observed pre-treatment confounders $X$ or, respectively, $Z$ in the population model would basically implement the DGP from sant2020doubly. However, we extend the simulation design in terms of an unobserved pre-treatment confounder $A$, which is uniformly distributed over an interval $(-1,1)$, i.e., $A\sim \mathcal{U}(-1,1)$. $A$ enters the equation of the propensity score and the outcome difference regression in an additive way\footnote{Additivity in the propensity score helps to compute the population values for the Riesz representer. To ensure that $p(Z,A)\in (0,1)$, we impose an additional clipping such that $0.1 \le p(Z,A) \le 0.9$.}
with $U \sim\mathcal{U}[0,1]$. The outcome $Y$ is generated as
where $\varepsilon_0,\varepsilon_1(D)\sim\mathcal{N}(0,\sigma_\varepsilon^2)$ and $\theta\in \mathbb{R}$.
We parametrize the causal model such that for the ATT we have $\theta_0 = 5$. Simulation studies for sensitivity analysis are characterized by a specific challenge. We would like to implement a specific confounding scenario, which is in line with the previously presented theoretic framework. To do so, we calibrate the values for $\gamma_A$ and $\beta_A$ based on a super-population model with $1,000,000$ observations, for which we can compute the long and short model. Accordingly, we can compute the population-level values of the sensitivity parameters $\widetilde C_D^2$, $C_{\Delta Y}^2$, and $\rho$. These values are then used as the evaluated parallel trend violation scenario in the empirical application of the sensitivity framework, such that we can measure the performance of the corresponding bias bounds at these population values. We consider settings with sample size $n \in \{500, 100, 5000, 50000\}$ and report results from $R=10,000$ simulation repetitions. We consider specification for the outcome regression and propensity scores that are linear in the covariates, i.e., $Z=\widetilde{Z}$, which corresponds to DGP 1 in sant2020doubly. For estimation, we use unpenalized linear and logistic regression learners. Hence, the resulting bias of the ATT estimate will be only the consequence of the parallel trends violation and not reflect misspecification bias.
Table (ref) shows the average results for the point estimation and bias bounds for the ATT as obtained from $R=10,000$ simulation repetitions for different sample sizes $n$. In all settings, the short estimate of the ATT, $\hat{\theta}_s$, exhibits an upward bias that results from omitting $A$ from the model. Applying the bias formula in Equation (ref) according to the definitions of the sensitivity parameters presented in the previous sections leads to a lower bound, $\hat{\theta}_-$, that is very close to the true value $\theta_0=5.0$. In all settings, the robustness value is close to its nominal level of $RV_{\theta=5}=0.1$. In the setting with the smallest sample size, $n=500$, the robustness value is slightly too optimistic suggesting to use conservative settings in small-sample settings. This result can be explained by some numerical instabilities in small samples, which is partly reflected by the higher standard variation. We recommend considering estimation uncertainty when interpreting robustness values, as also reflected by low values for the robustness values $RV_{\theta=5,a=0.1}$ in these settings, cf. Table (ref). Interestingly, comparing the lower bound, $\hat{\theta}_-$, to the oracle estimator $\hat{\theta}_{\text{long}}$ reveals that using the oracle confounding scenario is almost equivalent to directly using the unobserved pre-treatment confounder $A$. In small samples, the bias bounds are slightly more variable than the oracle estimate. However, with increasing sample size, this difference becomes negligible.
Table (ref) shows the average results for estimation and sensitivity bounds of the lower confidence limit. The results illustrate that, in line with the expectations, the empirical coverage of the one-sided confidence interval $[\hat{\ell}_s,\infty)$ for the ATT is below the nominal level of $90\%$ and approaches $0.00$ as estimation uncertainty diminishes in larger samples. In contrast, the lower confidence bias bounds, $\hat{\ell}_-$, cover the true value of the ATT in $91.9\%$ to $92.7\%$ of the cases, approaching a nominal coverage of $90\%$. In larger sample settings, the lower bound of the confidence limit gets closer to the true value for the ATT.\footnote{Note that the population lower bound, $\theta_-$, evaluated in the population parallel trend violation scenario is approximating the true ATT, $\theta_0$ (subject to numerical differences).} The robustness value $RV_{\theta=5,a=0.1}$ appears to be conservative in small samples but approaches the nominal value of $10\%$ in larger samples. Comparing the performance of the lower confidence bias bound to that of the lower confidence corresponding to the long point estimator reveals that their performance is similar in terms of empirical coverage and variability. The average value for $\hat{\ell}_-$ is very close to $\hat{\ell}_{\text{long}}$, with differences becoming smaller in larger samples. Moreover, increasing the sample size leads to smaller variability of the bias bounds, such that the corresponding standard deviations approach those of the oracle confidence bound in moderate and large samples.
To get more insight on the distribution of the lower bounds for the point estimate and the lower $90\%$ confidence limit, we provide histograms of their standardized versions in Figure (ref). The histograms show that the lower bounds $\hat{\theta}_-$ (Panel ($i$)) and $\hat{\ell}_-$ (Panel ($ii$)) are normally distributed with $\hat{\theta}_-$ being close to the true ATT, on average. The distribution of the lower bias bounds is rather dispersed in small data settings, with the approximation of a standard normal distribution becoming better with larger sample sizes. As indicated by the solid green vertical line, the $90th$ percentile of the empirical distribution of $\hat{\ell}_-$ is close to the true ATT (red dashed line), which provides evidence on the close-to-nominal level empirical coverage of the sensitivity confidence bounds with decreasing estimation uncertainty. Figure (ref) illustrates the empirical distribution of the DML estimate (short model), $\hat{\theta}_s$, and the lower bound, $\hat{\theta}_-$, for settings with increasing sample size in a density plot. The figure illustrates that with larger sample size, the lower bias bound (right panel) concentrates around the true value, whereas the DML estimate exhibits a bias irrespective of the reduced estimation uncertainty.
Figure (ref) shows the histogram of the standardized robustness values. The results for the small sample settings with $n=500$ and $n=1000$ show that the estimation of the $RV$ might be complicated by numerical instabilities, as it is often estimated to be $0$. With moderate and larger samples, the $RV$ approaches its nominal value and exhibits an empirical distribution that is similar to the standard normal distribution. More results can be found in Appendix (ref).
We apply the framework for sensitivity analysis to the famous LaLonde data. In his influential study, lalonde1986 evaluated the causal effect of participation in the National Supported Work Demonstration (NSW) program on earnings combining two different data sources. First, he used data from a field experiment, where individuals were randomly assigned to participate in a job training. Second, he constructs additional data sets from the Current Population Survey (CPS)\footnote{We follow smith2005 and focus on using the CPS samples.} and the Panel Study of Income Dynamics (PSID). The idea of using these additional data sets was to mimic an observational study, for which the results could be compared to those of an experimental evaluation. The so-called Lalonde critique casts doubts on the validity of (at that time state-of-the-art) observational causal inference techniques and sparked an intense debate on the use of observational data in the economics and econometrics literature. Important studies include dehejia1999causal, dehejia2002propensity (henceforth DW), heckman1997matching, and smith2005, among others. A comprehensive survey and reevaluation has recently been provided by imbens2024lalonde, which includes state-of-the-art causal estimators and new data sets. imbens2024lalonde also perform sensitivity analysis using the framework of cinelli2020making for linear regression under the unconfoundedness assumption.
We perform sensitivity analysis on the ATT in the $2\times 2$ DiD setting building on the previous work by smith2005, which was also evaluated using the doubly robust ATT estimator in sant2020doubly. Also huber2024joint use the Lalonde-PSID sample for their joint test for unconfoundedness and parallel trends.
smith2005 employ a matching DiD estimator to analyze three different data sets: (1) The original Lalonde-CPS sample, (2) a modified version of the Lalonde-CPS data used in DW, and (3) a refined version of the DW data that focuses on individuals from an early phase of the field experiment. The appealing feature of the LaLonde data is the availability of an experimental benchmark for observational causal estimates. This makes it possible to run two different types of a quasi-observational causal evaluation: First, following lalonde1986 and DW, the individuals that were actually assigned to participate in the job training are considerd as a treatment group and the control group is composed from the CPS data. Second, it is possible to evaluate a placebo effect: Those individuals who were not assigned to the treatment in the experiment are considered as a pseudo-treated group, for which the ATT is evaluated against the control group from the CPS data. The placebo analysis is useful to quantify the so-called “evaluation bias” smith2005 that originates from systematic differences in the experimental and observational samples. Because the pseudo-treated group did not actually receive the treatment, a nonzero and possibly significant ATT estimate only reflects selection into the experimental sample ($\theta_0=0$ in this case). In their results, smith2005 and sant2020doubly report the evaluation bias as $\frac{\hat{\theta}}{\hat{\theta}_\text{Exp}}$, where $\hat{\theta}_\text{Exp}$ is the ATT estimate from the original experimental evaluation in lalonde1986.\footnote{A bias of $1$ would indicate that a reported ATT estimate is to $100\%$ reflecting a bias due to selection into the experimental sample.}
The rationale for the placebo analysis in smith2005 is that the previously reported observational estimates in DW are not only reflecting the causal effect, which is estimated according to their flexibly specified propensity score matching estimators, but also affected by such sample selection bias. To account for systematic differences in the experimental and observational samples related to the measurement of the outcome variable and accounting for local labor market characteristics, smith2005 suggest to use matching difference-in-differences estimators. Accordingly, smith2005 conclude that DiD estimators are better able to adjust for these time-invariant factors.
Moreover, smith2005 emphasize that the functional form specification for the propensity score plays an important role to replicate results in DW, which is in line with the results in sant2020doubly. sant2020doubly use linear and logistic regression based on manually constructed specifications including polynomial and interaction terms as motivated by DW. Table (ref) show the point estimates for the DML DiD ATT in the placebo analysis according to the preferred choices for modeling the nuisance components. We chose the learner specification that performed best in terms of the predictive performance for the nuisance functions $g()$ and $m()$. To address overlap issues, which are known to be a major challenge in the LaLonde data evaluation, we calibrated the learners using isotonic regression van2024doubly, van2025automatic, klaassen2025calibration, ballinari2024calibrating. More results, including those from other learner choices, are available in Appendix (ref). The DML DiD estimators and the evaluation bias in Panel (1) of Table (ref) are all in the range of the results reported in sant2020doubly (Table 3), with slightly reduced standard errors. Overall, the estimation of the ATT is very variable leading to non-significant estimates, due the number of pseudo-treated individuals being very small as compared to the control group. As a consequence the robustness values $RV$ displayed in Panel (2) of Table (ref) are very close to zero pointing at non-robust effects. However, we estimate the bias bounds for the ATT in the sample with the actual treated based on the placebo settings, i.e., we set $C_{\Delta Y}^2$ and $\widetilde{C}_D^2$ to the corresponding robustness values from the placebo analysis. The upper bias bounds for the point estimates and confidence bounds are shown in Figure (ref). It is possible to see that in the LaLonde CPS and early randomized samples, the upper sensitivity bounds for the ATT are now closer to the experimental benchmark. In all cases, the confidence sensitivity bounds cover the experimental benchmarks.
As a second empirical example, we apply sensitivity analysis to a study by draca2011minimum who evaluate the causal effect of the introduction of the national minimum wage (NWM) in the UK on firm profitability in 1999. Unlike the original study, we only consider the case of a balanced panel, i.e., we drop observations that are not observed during the entire study period from 1994 to 2002. The unit of observation $i$ at time $t$ is a firm. In contrast to the original study with data of up to $771$ firms, our balanced panel data set contains $337$ with $57$ treated firms. draca2011minimum define the treatment based on pre-treatment wages: Firms that are expected not to be affected or less affected according to average wages in the time before the NWM introduction are categorized as untreated. Firms with lower average wages prior to the introduction of the NWM are considered as treated. In line with the original study, we consider the effect of the NWM introduction on two outcome variables, the log average wage at the firm (ln_avwage) and the firm's profit margin defined as the profit to sales ratio (net_pcm). In this section, we focus on the analysis with regard to net_pcm and provide the results on average wages in Appendix (ref). The original study employs two-way fixed effects regression. In our analysis, we focus on the previously presented DML DiD estimator in the multi-period setting. Due to the different sample compositions, our results deviate slightly from the original study.
In our replication of draca2011minimum, we base identification on the conditional parallel trends assumption, which we assume to hold after accounting for pre-treatment information on a firm's industry (2-digit industry classification, sic2), the government office region of workplace (gorwk), the share of part-time workers (ptwk), the share of female workers (female), and the share of union members (unionmem) by the three-digit industry classification. The latter three of these variables are time variant whereas there is no variation in industry and region information over time. For the time-varying variables, we condition on the pre-treatment level. Figure (ref) illustrate the $ATT(\mathrm{g},t)$ estimates as obtained from double machine learning using a random forest learner for the propensity score and the outcome difference regression. We estimate an effect of the NWM introduction on the net profit margin in the first post-treatment period of $-0.021$ ($95\%$ confidence interval $[-0.041,-0.001]$), which is in line with the results in draca2011minimum.
We apply our sensitivity approach to the $ATT(\mathrm{g},t)$ parameters and show the resulting robustness values in Figure (ref). The $RV$s for the post-treatment $ATT(\mathrm{g},t)$ estimates range between $6.5\%$ for the third post-treatment period to $9.6\%$ in the second post-treatment period. To set these values into context, we can exploit the information from pre-testing, for which we expect non-significant effects and, thus, small $RV$ values under valid conditional parallel trends (cf. Section (ref)). The effect estimates for the pre-treatment periods are relatively close to zero and not significant. The maximum robustness value from pre-testing is $2.56\%$ as obtained for the last pre-treatment period. Compared to the post-treatment $RV$ values , the pre-testing $RV$ is relatively small. Note that the pre-testing coefficient also has a different sign than the post-treatment $ATT(\mathrm{g},t)$ estimate, which might, hence, result in a rather conservative parallel trend violation scenario.
An appealing feature of the study by draca2011minimum is the availability of a rich set of pre-treatment confounding variables, which can be exploited for benchmarking exercises and, thus, inform plausible parallel trend violation scenarios. Following the general benchmarking procedure described in Appendix F of chernozhukov2022long, we can compute gain statistics from leaving out covariates. The corresponding values for the sensitivity parameters are presented in Table (ref). In many cases, we find that omitting the covariates, as for example the information on industry classification, has considerably more explanatory power for the treatment status than for the difference in outcomes. This is in line with the general intuition that information on a firm's industry is likely related to the average wage level (prior to the NMW introduction), which is used to define the treatment status in draca2011minimum.
The previous sensitivity exercises have resulted in a set of parallel trend violation scenarios. As a next step, we consider the critical parallel trend violation that would lead to a substantial change in the causal results, i.e., reduce the (negative) effect estimate of the NWM introduction to zero. In the following, we focus on the sensitivity bounds for the $ATT(\mathrm{g},t)$ in the first post-treatment period. For this parameter, we obtain robustness values $RV=8.76\%$ and $RV_{a=0.1}=1.81$. Hence, an unobserved pre-treatment confounder $A$, which could explain $8.76\%$ of the residual variation in the conditional expected outcome difference and lower the odds ratio for the treatment status according as explained by the oracle model by $8.76\%$ would be required in order to set the effect to zero.
As a next step, we can estimate the upper sensitivity bounds for the $ATT(\mathrm{g},t)$ parameter in the first post-treatment period. To do so, we use an additional visualization of the upper bound for the causal parameter in the different parallel trend violation scenarios through a contour plot, presented in Figure (ref). Note that we enforced $\lvert \rho \lvert=1$ for the indicated scenarios in the contour plot to maintain the comparability of the different scenarios.\footnote{$\rho$ operates as a scaling factor in the bias formula in Equation (ref).} As a consequence, the evaluated settings are worst-case scenarios with a maximum correlation of the confounding variation in the treatment assignment and outcome difference. The contour plot shows that the upper bound for the $ATT(\mathrm{g},t)$ parameter in the first post-treatment period is negative in all scenarios considered. A parallel trend violation corresponding to the strongest pre-treatment violation would result in a reduction of the causal effect (in absolute values) to $-0.015$. The strongest (conservative) benchmarking scenario corresponds to leaving out the share of part-time workers by the three-digit industry classification. A parallel trend violation that is comparably strong in terms of the explanatory power for the treatment assignment and the considered outcome difference would induce an adjustment of the $ATT(\mathrm{g},t)$ parameter to $-0.014$.
Moreover, it is possible to vary the strength in terms of a $k$-fold multiple of the pretesting scenario sensitivity parameters, which would be similar to the type of results commonly reported in applications of the approach by rambachan2023. Figure (ref) presents the upper bound of the $ATT(\mathrm{g},t)$ parameter in the first post-treatment period and the one-sided upper confidence bound at a nominal level of $90\%$. The figure illustrates that the reported causal effect in the first post-treatment would become non-significant if the conditional parallel trend violation was comparable to the pre-testing scenario.
The previous sensitivity results point at a rather robust effect in the considered scenarios. Of course, the overall conclusion from our empirical analyses have to be set into the context of the study. draca2011minimum study a country-wide introduction of a minimum wage, which is a policy affecting all companies operating in the UK. Differences in the probability to be affected by the NMW introduction and the change in the profitability over the considered time horizon might be related to firm characteristics, such as industry and characteristics of the workers. We find that benchmarking against observed variables in the data point at a rather robust effect. However, considering estimation uncertainty, the results in Figure (ref) show that the effect quickly becomes non-significant in our pre-testing sensitivity exercises.
Our study contributes to the existing and quickly growing literature on difference-in-differences, causal machine learning and sensitivity analysis. In the previous section, we presented a new approach to sensitivity analysis with regard to violations of the conditional parallel trend assumption in common difference-in-differences models. In many empirical applications, researchers are concerned with potential violations of the parallel trend assumptions and a corresponding bias of the causal parameter estimate. To assess the robustness of causal estimation results due to such violations, researchers can use our suggested sensitivity approach and obtain asymptotic bounds for the point estimates and confidence intervals for the target parameters. Specific violation scenarios in terms of a set of values for the sensitivity parameters in the asymptotic bias formula can be based on pre-testing evidence, benchmarking analyses, standard reporting statistics and domain expertise. In addition to the theoretical results on Riesz representation in the canonical and multi-period DiD setting with staggered adoption, we provide new evidence on the validity of Riesz-representation-based sensitivity analysis in a simulation study. Moreover, we provide an open source implementation of our approach in DoubleML for Python and we demonstrate the application of DiD sensitivity analysis in two empirical examples.
There are various extensions worth to explore in future research. A natural extension of the current approach is to aggregations of the $ATT(\mathrm{g},t)$ parameters in the multi-period difference-in-differences setting as considered in callaway2021difference, For example, often the treatment effect is evaluated relative to the treatment receipt in event studies. Riesz-representation based sensitivity analysis for aggregated effects is non-trivial and, hence, left for future research. Further extensions might be interesting for recently considered difference-in-differences models such as models with continuous treatments. Finally, empirical application of Riesz-representation-based sensitivity analysis is important to address empirical challenges of these new techniques for causal analysis.