The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.
80,181 characters
Omitted Variable Bias in Difference-in-Differences Designs
\def\spacingset#1{\small\normalsize}
\begin{bibunit}
\spacingset{1}
\if11
{
\title{\bf Omitted Variable Bias in Difference-in-Differences Designs}
\author{
Juejue Wang \\
Department of Statistics, University of Washington \\
and \\
Pedro H.C. Sant'Anna \\
Department of Economics, Emory University \\
and \\
Victor Chernozhukov \\
Department of Economics, MIT \\
and \\
Carlos Cinelli \\
Department of Statistics, University of Washington}
\date{}
\maketitle
} \fi
\if01
{
\title{\bf Omitted Variable Bias in Difference-in-Differences Designs}
\author{
}
\date{}
\maketitle
} \fi
\begin{abstract}
We study the omitted variable bias (OVB) problem in canonical difference-in-differences (DiD) designs when unobserved confounding induces departures from the parallel trends assumption. Our results provide a novel characterization of the OVB formula for the average treatment effect on the treated (ATT), which is of independent interest. We show how the ATT bias is mainly governed by the strength of confounding in the treatment assignment mechanism and provide alternative ways of quantifying this strength, such as (i) changes in the average odds of treatment among the treated, (ii) confounding imbalance between treated and control units, or (iii) variation explained in treatment odds among the untreated.
Building on these results, we offer sensitivity statistics for routine reporting, describing the minimum strength of confounding required to overturn the conclusions of a DiD study, as well as formal bounds on the strength of confounders based on comparisons to observed covariates or pre-trends. Finally, we provide flexible and efficient statistical inference methods for the bounds on ATT, which can leverage modern machine learning algorithms for estimation. We demonstrate the utility of our approach in an empirical example that estimates the effects of minimum wage on teen employment.
\end{abstract}
\noindent
{\it Keywords:} Sensitivity Analysis; Difference-in-Differences; Parallel Trends; Unobserved Confounding; Double Machine Learning; Robustness Value.
\vfill
\newpage
\spacingset{1.73}
\section{Introduction}
Difference-in-differences (DiD) has become the most widely used research design for causal inference with observational data in economics and other empirical sciences \citep{goldsmith2024tracking}. Its appeal is easy to understand. The method is simple to implement, makes modest demands on the data (it requires only a treated and control group, observed before and after treatment takes place), and it even allows for some types of treatment selection. In its canonical form, DiD identifies the average treatment effect on the treated (ATT) under a parallel trends assumption (PTA), which states that, in the absence of treatment, the average outcomes of treated and control units would have evolved in parallel over time---possibly only after conditioning on a set of observed pre-treatment covariates. As with any method for causal inference, however, the reliability of DiD rests on the plausibility of its identifying assumption. Because the PTA concerns counterfactual trajectories that are never observed for treated units after treatment, it is fundamentally untestable, and investigators marshal what evidence they can to defend its plausibility in the context of their investigations.
By far the most common way to gather such evidence is to perform placebo tests using pre-treatment information. Researchers typically check whether the outcomes of treated and control groups evolve in parallel before treatment, and read the absence of differential pre-treatment trends (often known as ``pre-trends'') as evidence for post-treatment parallel trends. Another common practice is to check whether treated and control groups are balanced on observed pre-treatment covariates that are also thought to be determinants of changes in the untreated potential outcome. While null findings on either of these tests are consistent with the parallel trends assumption, they are not dispositive. Pre-trends speak only to what happened before treatment, and parallel trends may well hold ex-ante and fail ex-post. Similarly, balance on observed covariates cannot rule out \emph{unobserved} confounders that induce violations of parallel trends. And in many settings, pre-trends do differ across groups, and observed covariates are imbalanced. Yet researchers would still like to learn something about the ATT, even if there is evidence that the parallel trends assumption does not hold exactly.
To this end, sensitivity analyses have been proposed that use observed pre-treatment deviations from parallel trends to bound the plausible magnitude of post-treatment violations (e.g., \citealp{rambachan2023more}). This framework has been widely adopted, and is now recommended as a standard sensitivity analysis tool in the DiD literature \citep{roth2023s, baker2025difference}. While a step forward from naively assuming that parallel trends holds exactly, this approach still suffers from shortcomings. In applied work, violations of parallel trends are typically attributed to \emph{unobserved confounders} that \emph{jointly} shape the \emph{selection} of units \emph{into treatment} and their untreated \emph{outcome trends} \citep{callaway2021difference, roth2023s, baker2025difference}. Yet pre-trend extrapolation offers no framework for reasoning about these forces, and \emph{why} parallel trends might fail in the first place. It is informative only insofar as pre-treatment dynamics are a reliable proxy for the post-treatment counterfactual, an assumption that it cannot itself adjudicate. One immediate symptom of this deficiency is that the framework cannot be applied to the simplest, canonical two-period DiD design, where pre-trends are not available. How can we translate such confounding into formal statements about deviations from parallel trends and the ATT bias? How strong would it have to be to overturn a given conclusion in a DiD study?
In this paper, we develop a suite of sensitivity analysis tools for DiD that allows one to easily answer such questions.
We first derive a novel characterization of the omitted variable bias (OVB) formula for the ATT. While motivated by DiD, this result is important in its own right and of independent interest---we elaborate on this point in the related literature below. We show how the bias due to violations of parallel trends admits a formal decomposition into the product of a scaling factor estimable from the data, and three interpretable bias factors that researchers already informally invoke: (i)~the strength of unobserved confounders in shaping selection into treatment, (ii)~their strength in shaping the evolution of untreated potential outcomes, and (iii)~the alignment between these two channels. Among these, we further discuss how selection into treatment is the most important factor for bounding this bias (since the other two factors are each upper bounded by one), and we offer a rich set of equivalent ways that practitioners can use to reason about it, such as how unmeasured confounding increases the average odds of treatment among the treated, widens the covariate imbalance between treated and control units, or explains variation in treatment odds among the untreated.
Next, we derive (extreme) robustness values for DiD. These are sensitivity statistics that characterize the minimum strength of unobserved confounding needed to overturn a given conclusion about the ATT \citep{cinelli2020making,cinelli2025omitted,chernozhukov2022long}. Routine reporting of these quantities, alongside point estimates and standard errors, provides a quick and simple way to communicate how robust DiD findings are to violations of the parallel trends assumption. We also offer formal bounds on the ATT based on plausibility judgments about how the strength of unobserved confounders compares against the explanatory power of observed covariates. These bounds connect naturally to balance checks already performed by applied researchers, who compare observed covariates across treated and control groups as informal evidence for parallel trends. Our framework gives these comparisons a direct role in sensitivity analysis, allowing the observed imbalance on measured covariates to discipline judgments about the unobserved imbalance one is willing to entertain. When multiple pre-treatment periods are available, the same logic applies to pre-treatment dynamics, allowing our framework to be used in tandem with current pre-trend extrapolation approaches.
Finally, we provide flexible and efficient methods for statistical inference using debiased machine learning (DML), which allows the use of modern machine learning algorithms for estimation---though we note our approach can also be used with standard parametric regressions. We demonstrate the utility of our approach in an empirical application revisiting the analysis of the effect of minimum wage on teen employment in \citet{callaway2021difference}. Open-source software for R implements the methods discussed in this paper.
\paragraph{Related literature.}
Our work contributes to the growing literature on sensitivity analysis for DiD, where the dominant approach is the pre-trend extrapolation framework of \cite{rambachan2023more}. Pre-trend extrapolation is best understood as a way of benchmarking the plausible magnitude of post-treatment violations against observed pre-treatment dynamics. However, it cannot explain how such violations arise in the first place. Our framework fills this gap by decomposing the bias into interpretable factors tied to selection into treatment and untreated outcome evolution, and in doing so, it opens the door to a broader set of benchmarking strategies. Observed covariates, for instance, can also be used to benchmark the plausibility of unobserved confounding, further connecting sensitivity analysis to the balance tests that applied researchers already perform as a matter of course. See Section~\ref{sec:benchmarking} for further discussion.
Our work is also related to the broader literature on sensitivity analysis for the ATT. In particular, our results are most closely related to \citet{chernozhukov2022long}, who develop a general OVB framework for linear functionals of the conditional expectation of the outcome, encompassing the ATT as a special case; see \cite{bach2025sensitivity} for a direct application of \citet{chernozhukov2022long} to DiD designs using this unconditional parameterization. Though we build on \citet{chernozhukov2022long}, our OVB formula for the ATT differs from theirs because we exploit a structural feature of the ATT---namely, that the bias admits a parameterization conditional solely on the untreated units. This parameterization has several desirable properties, such as requiring plausibility judgments only on the local parameters that directly drive the bias. We further extend \citet{chernozhukov2022long} by providing alternative ways of quantifying selection strength in terms of treatment odds and covariate imbalance, both for our preferred conditional parameterization and for the original unconditional parameterization.
Our results are also closely related to the variance-based sensitivity analysis proposed in \citet{huang2025variance} for weighting estimators, which bounds the bias of the ATT by restricting an $R^2$ measure of the balancing weights. Our approach differs from theirs in four main ways. First, we use a non-centered parameterization of selection strength, which naturally avoids a zero-denominator issue acknowledged in \citet{huang2025variance}. Second, we further decompose the bias into three components---the strength of confounding in treatment selection, its strength in untreated outcome trends, and the alignment between the two. We show that \citet{huang2025variance}'s parameterization bundles the trend and alignment components together, so that their upper bound on the bias implies setting these two components to one. Third, their benchmarking analysis uses observed covariates in a way that is less conservative than ours. Fourth, we use debiased machine learning for statistical inference.
We further extend the derivations in \citet{huang2025variance} to obtain results analogous to ours in their centered parameterization.
Thus, a distinct contribution of this paper is not only a novel OVB formula for the ATT, but a complete account of the different parameterization choices available for this bias, along with the theoretical and practical consequences of these choices. See Section~\ref{sec:parameterization} and Appendix~\ref{app:all-comparison} for further details.
\paragraph{Notation.} $R^2_{Y\sim X}$ denotes the (possibly uncentered) $R^2$
from the orthogonal linear projection of a scalar $Y$ onto a random vector
$X$, and $\eta^2_{Y\sim X}$ its non-parametric analogue, with $E[Y\mid X]$ in
place of the best linear predictor. The partial $R^2$ of $Y$ with $U$ given
$X$ is defined as $R^2_{Y\sim U\mid X}:=(R^2_{Y\sim U+X}-R^2_{Y\sim X})/(1-R^2_{Y\sim
X})$, with an analogous definition for $\eta^2_{Y\sim U\mid X}$. We use $\mathop{}\!\mathrm d L/\mathop{}\!\mathrm d P$ to denote the Radon--Nikodym
derivative of $L$ with respect to $P$; for distributions $P \ll Q$,
$\chi^2(P\|Q):=\int(\mathop{}\!\mathrm d P/\mathop{}\!\mathrm d Q-1)^2\,\mathop{}\!\mathrm d Q$ denotes the chi-squared divergence of
$P$ from $Q$. Finally, $P_{V\mid W=w}$ denotes the distribution of $V$ given
$W=w$, and we use $E_n[f(X_i)]:=\frac{1}{n}\sum_{i=1}^n f(X_i)$ for empirical
expectations.
\section{Background and running example}
\label{sec:background}
In this section we introduce the running example used throughout the paper. This example serves three main purposes: (i)~it illustrates the standard use of pre-trends and covariates as a tool for assessing the plausibility of parallel trends; (ii)~it shows how our framework already refines this standard analysis through a richer characterization of \emph{observed} biases; and (iii)~it sets the stage for the formal sensitivity analysis to \emph{unobserved} confounders that we develop in the rest of the paper.
\subsection{The effect of minimum wage on teen employment}
Whether increases in minimum wages reduce teen employment has been actively debated in labor economics for decades, with credible evidence on both sides \citep{neumark2022myth, dube2024minimum}. Here we revisit the analysis of \citet{callaway2021difference}, who study this question using a difference-in-differences design.
Between 2001 and 2007, the federal minimum wage remained at \$5.15, but several states raised their own minimums above this floor at different points in time. \citet{callaway2021difference} exploit this variation in treatment timing to estimate group-time average treatment effects for each cohort of states, defined by when they first raised their minimum wage above the federal level. To fix ideas, here we focus on the last cohort, the states that raised their minimum wage in 2007. The treated group consists of the 584 counties in those states, and the control group consists of the 1,377 counties in states that never raised their minimum wage above the federal floor before the end of 2007.
An investigator interested in estimating the ATT might start the analysis by positing that, absent the policy change, teen employment in treated and control counties would have evolved in parallel over time. Under this parallel trends assumption, the ATT is identified by the usual canonical two-by-two DiD estimand. The results are shown in panel (a) of Figure~\ref{fig:nt_ATT_multi_main}.
The treatment effect of interest is the estimate for 2007, and we find that increases in minimum wages resulted in a 2.77\% reduction in teen employment for that group. This is a sizable, policy-relevant negative effect, but how credible is this estimate?
\begin{figure}[t]
\centering
\begin{subfigure}{0.49\textwidth}
\centering
\includegraphics[width=\textwidth]{plots/uncond_nt_ATT_multi.pdf}
\caption{{\scriptsize Unconditional Parallel Trends, never-treated as control}}
\end{subfigure}
\hspace{0.015\textwidth}
\begin{subfigure}{0.49\textwidth}
\centering
\includegraphics[width=\textwidth]{plots/cond_nt_ATT_multi.pdf}
\caption{{\scriptsize Conditional Parallel Trends, never-treated as control}}
\end{subfigure}
\caption{{\footnotesize
Effect of the minimum wage on teen employment. Estimates use DML with a canonical DiD design and the preceding year as base period. Red and blue lines give point estimates with uniform 95\% confidence bands for pre-treatment periods and treatment effects, respectively.}}
\label{fig:nt_ATT_multi_main}
\end{figure}
Beyond the effect of interest, Figure~\ref{fig:nt_ATT_multi_main}(a) also shows pre-treatment estimates (``pre-trends'') from 2002 to 2006, which are now commonly used to assess the plausibility of the parallel trends assumption. If the untreated trends of both groups were indeed evolving in parallel over time, these pre-trends should be equal to zero, barring sampling variation.
Yet, the 2006 pre-trend is -3.9\%, which is not only large and statistically significant, but also of the same order of magnitude as the treatment effect estimate itself. This suggests that treated and control counties were already on diverging trends before any policy changes.
What might explain such deviations from parallel trends? A natural concern is that employment in counties with \emph{different characteristics} may not evolve in parallel over time. To probe this, another common approach is to check whether the treated and control groups are \emph{balanced} in terms of covariates that could plausibly affect employment trajectories. Table~\ref{tab:balance_main} in Appendix~\ref{app:additional-results} reports standardized differences in means for several such characteristics. We find that treated counties are more populous, whiter, more educated and have lower poverty rates than control counties; they are also concentrated in the Midwest and West, with the South substantially underrepresented. Most differences exceed the 0.25 threshold, conventionally regarded as problematic \citep{imbens2015causal,baker2025difference}. These are the kind of imbalances that could account for parallel trends violations, and they motivate adjusting for observed covariates before drawing any conclusions about the policy effect.
To alleviate such concerns, panel (b) of Figure~\ref{fig:nt_ATT_multi_main} reports DiD estimates under a \emph{conditional} parallel trends assumption, which requires only that counties with the \emph{same observed characteristics} would have evolved in parallel over time. These estimates account for observed covariates nonparametrically, using random forests and debiased machine learning as we will discuss in Section~\ref{sec:est-and-inf}. The treatment effect estimate is now -3.66\%, somewhat larger in magnitude than the unconditional estimate, and points to the same conclusions. However, the 2006 pre-trend estimate also remains essentially unchanged. In other words, though adjusting for observed covariates has not changed our treatment effect estimate, it has not resolved the pre-trend problem either. This persistent pre-trend deviation suggests confounding is acting through factors not captured by the observed data.
At this point, the standard recommendation in the DiD literature is to proceed with a sensitivity analysis of the type of \citet{rambachan2023more}. It proceeds as follows: take the conditional 2006 deviation of -3.8\%,
assume the same deviation persists in 2007, and subtract it from the conditional estimate of -3.66\%. Since 3.8\% is larger than 3.66\%, this is evidently enough bias to overturn our original conclusion. This approach is straightforward and an improvement over simply assuming parallel trends and ignoring the problem. But it is also somewhat unsatisfactory as it leaves several questions unanswered. Why have the observed covariates barely changed these estimates, despite their substantial imbalance? Should we be looking at standardized mean differences or something else? What kind of unobserved confounding would have to be at work to produce these biases? And is pre-trend extrapolation the only way to gauge the plausible magnitude of such confounders?
As a preview of how the results of this paper can help analysts answer these questions,
we show that the difference between the conditional and unconditional estimates in fact admits the following exact decomposition:
\[
\underbrace{(-0.0277)}_{\text{Unconditional Estimate}} - \underbrace{(-0.0366)}_{\text{Conditional Estimate}} =~ \underbrace{0.0089}_{\text{Bias}} ~= \underbrace{0.194}_{\text{Alignment}} \times \underbrace{0.117}_{\text{Trend}} \times \underbrace{2.548}_{\text{Selection}} \times \underbrace{0.154}_{\text{Scale}}.
\vspace{.2cm}
\]
Table~\ref{tab:obs_strength} in Appendix~\ref{app:additional-results} reproduces the above decomposition using all covariates, and also presents analogous decompositions obtained by adjusting for each observed covariate separately. Almost every applied DiD paper conditions on covariates, yet the standard narrative about why they matter rarely extends beyond covariate imbalance. Decompositions such as the one above provide an exact post-mortem of how covariates affect (or not) the results.
First, the ``selection'' component confirms that indeed there is a sizable distributional imbalance between treated and control groups, corresponding to a chi-squared divergence of $2.548^2 \approx 6.5$---note the relevant metric of distributional imbalance here is not the commonly used standardized mean difference, though in some cases the two may be close to each other, as we will discuss. Second, one of the reasons why this imbalance does not translate into a substantial difference in estimates in our running example is because these covariates explain only $0.117^2\approx 1.4\%$ of the variation in trends among the control group (the ``trend'' component). Third, for bias to arise, covariates must affect both channels jointly---their effect on the outcome trend must correlate with their effect on selection. This correlation, the ``alignment'' component, equals only 0.194 in our example, further attenuating the bias. As with the conditional estimate, all these components can also be estimated non-parametrically using debiased machine learning with the tools we provide in this paper. We argue that exercises such as the one above should be a routine part of DiD studies.
But more importantly, this same decomposition used to explain how observed covariates \emph{did} change our estimates can also be used to contemplate how unobserved confounders \emph{would} have changed our estimates. For example, it explains what unobserved confounders need to look like, in terms of their strength in predicting outcome trends and treatment selection, in order to induce any certain amount of bias---including biases derived from pre-trend extrapolation.
Crucially, it also opens up different ways of gauging the plausible magnitude of confounding. For instance, if an investigator has grounds to claim that no plausible unobserved confounder is more strongly imbalanced than the observed covariate region---or, equivalently, in terms of how it changes the odds of belonging to the treated group---this is also sufficient to bound the bias. We develop these tools formally in the rest of the paper, beginning with the bias decomposition itself in the next section.
\section{Omitted variable bias in difference-in-differences designs}
\label{sec:ovb}
\subsection{Problem setup}
\label{sec:setup}
We begin by introducing the omitted variable bias problem in the canonical two-by-two DiD setup. Let $t \in \{1,2\}$ index the pre- and post-treatment periods, respectively. Let $D\in \{0,1\}$ denote treatment group assignment, with $D=1$ for treated and $D=0$ for control. For each period $t$, $Y_t(1)$ and $Y_t(0)$ denote the treated and untreated potential outcomes at time $t$, respectively. We assume the observed outcome at time $t$ satisfies
\begin{equation}
\label{eq:consistency}
Y_t = D Y_t(1) + (1 - D) Y_t(0).
\end{equation}
Equation~(\ref{eq:consistency}) is often referred to as the consistency assumption, or the stable unit treatment value assumption (SUTVA).
The causal parameter of interest is the average treatment effect on the treated (ATT) at the post-treatment period, defined~as
\begin{equation}
\label{eq:att}
\text{ATT} := E[Y_2(1) \mid D=1] - E[Y_2(0) \mid D=1].
\end{equation}
Now let $\Delta Y:= Y_2 - Y_1$ denote the outcome evolution, $\Delta Y(0) := Y_2(0) - Y_1(0)$ the untreated potential outcome evolution, $X$ denote a vector of observed covariates and $U$ denote a vector of \emph{unobserved} covariates. Suppose that, conditionally on both $X$ and $U$, the outcomes of treated and control units would have evolved in parallel over time in the absence of treatment,~i.e.,
\begin{equation}
\label{eq:cpta}
E[\Delta Y(0) \mid D=1, X, U] = E[\Delta Y(0) \mid D=0, X, U].
\end{equation}
Under this conditional parallel trends assumption---along with the assumptions of no anticipation and other regularity conditions\footnote{See Appendix~\ref{app:did-assump} to recall the usual standard DiD assumptions.}, such as weak overlap below---the ATT is identified by the following well-known \emph{difference-in-differences estimand},
\begin{equation}
\label{eq:long}
\theta:= E[~E[\Delta Y \mid D=1, X,U] - E[\Delta Y \mid D=0,X,U] ~\mid~ D=1~].
\end{equation}
We refer to $\theta$ as the ``long'' DiD estimand because it depends both on the observed covariates $X$ and the unobserved confounders $U$. The starting point of our analysis is that the investigator has determined that she is interested in the DiD estimand $\theta$ as in (\ref{eq:long}).
\begin{remark}
Different formulations of the DiD assumptions can lead to the same statistical estimand (\ref{eq:long}), after appropriately redefining assumptions and variables. See, for example, \cite{callaway2021difference} for identification results involving group-time average treatment effects with multiple time periods, which invoke various different choices of control groups, parallel trends and no-anticipation assumptions.
Another example is the identification of the ATT under an unconfoundedness assumption, instead of parallel trends---for that case, one simply replaces $\Delta Y$ with $Y$.
Our results apply to any such estimand that can be expressed in the generic form (\ref{eq:long}).
\end{remark}
We further impose the following regularity conditions.
\begin{assumption}[Regularity Conditions]
\label{asmp:regularity}
We assume the outcome evolution is square-integrable $E[\Delta Y^2] < \infty$ and weak overlap $E\left[P(D=1)^{-1}P(D=0|X,U)^{-1} \mid D=1 \right] < \infty$.
\end{assumption}
In practice, however, the unobserved confounders $U$ are not available to the researcher, so we cannot estimate the ``long" DiD parameter $\theta$. Instead, we are forced to estimate its ``short'' version, omitting $U$ from the analysis,
\begin{equation}
\label{eq:short}
\theta_s:= E[~E[\Delta Y \mid D=1, X] - E[\Delta Y \mid D=0,X] ~\mid~ D=1~].
\end{equation}
Note that the parallel trends assumption may not hold if we condition on $X$ alone, and our estimates may thus suffer from omitted variable bias. In fact, under standard DiD assumptions, the omitted variable bias due to the omission of $U$ exactly equals the average deviation from parallel trends, among the treated, from conditioning on $X$ alone, i.e.,
\[\theta - \theta_{s}
= -
E\left[~E[\Delta Y(0) \mid X, D=1] - E[\Delta Y(0) \mid X, D=0] ~\mid~ \, D=1 ~\right].\]
Consequently, judgments about deviations from parallel trends can be translated into judgments about omitted confounding. Conversely, the presence of omitted confounders provides a natural way to explain deviations from parallel trends.
Our goal is to characterize the omitted variable bias---or, equivalently, the deviations from parallel trends---in terms of the strength of the omitted variable $U$, capturing the distinct channels through which $U$ confounds our estimand.
\subsection{The omitted variable bias formula}
Let $g_0$ and $g_{0s}$ be the long and short regression functions of outcome trend for the untreated,
\[g_0:= E[\Delta Y|D=0, X,U], \qquad g_{0s}:= E[\Delta Y|D=0, X].\]
The long and short DiD estimands can thus be compactly written as
\[\theta = E[\Delta Y -g_0\mid D=1], \qquad \theta_s = E[\Delta Y - g_{0s}\mid D=1].\]
The average treated trend $E[\Delta Y|D=1 ]$ is the same in both the long and short parameters, and it cancels out from the bias. The OVB is therefore fully governed by errors in the extrapolation of the untreated trend from the control group to the treated group. That is,
\[\theta - \theta_s = -(\theta_0 - \theta_{0s}),\]
where,
\begin{equation}
\label{eq:theta0}
\theta_0 :=E[g_0|D=1], \qquad \text{and}, \qquad \theta_{0s}:=E[g_{0s}|D=1].
\end{equation}
In order to better understand the nature of this extrapolation, consider the marginal odds, as well as the short and long conditional odds of treatment,
\[
O:=\frac{P(D=1)}{P(D=0)},
\qquad
O_{XU}:=\frac{P(D=1\mid X,U)}{P(D=0\mid X,U)},
\qquad
O_X:=\frac{P(D=1\mid X)}{P(D=0\mid X)}.
\]
A key step for our approach is the following lemma, which rewrites $\theta_0$ and $\theta_{0s}$ as inner products of the conditional regression functions and rebalancing weights given by the conditional to marginal odds ratio, under the control population.
\begin{Lemma}
\label{lemma:cond_RR_main}
Under Assumption~\ref{asmp:regularity}, $\theta_0$ and $\theta_{0s}$ in (\ref{eq:theta0}) can be written as
\[\theta_0 = E
\left[g_0 \left(\frac{O_{XU}}{O}\right)\mid D=0 \right], \quad \theta_{0s} = E\left[g_{0s} \left(\frac{O_{X}}{O}\right)\mid D=0\right].\]
Moreover, $g_{0s} = E[g_0 \mid D=0,X]$ and $O_{X} = E[O_{XU}\mid D=0,X]$.
\end{Lemma}
The fact that $g_{0s} = E[g_0 \mid D=0,X]$ follows directly from the law of iterated expectations. The property $O_{X} = E[O_{XU}\mid D=0,X]$ can be verified by applying Bayes' rule. Using this lemma, we can derive the following characterization of the OVB for the ATT. It shows that deviations from parallel trends can be exactly characterized by the ability of omitted confounders to explain treatment selection---as parameterized by the odds of treatment---and trend variation, jointly.
\begin{Theorem} [OVB for the ATT]
\label{thm:main_npm}
Consider the long and short estimands $\theta$ and $\theta_{s}$ in (\ref{eq:long}) and (\ref{eq:short}). Under the regularity conditions of Assumption~\ref{asmp:regularity}, the OVB is given by the (negative of the) covariance of errors induced~by~the omission of $U$ both in the regression function and in the odds ratio, among the untreated units,
\begin{align*}
\theta - \theta_s = - \operatorname{Cov}\left(g_0 - g_{0s},~ \frac{O_{XU}}{O}- \frac{O_{X}}{O} ~\middle|~ D=0 \right).
\end{align*}
Moreover, this bias can be further expressed in terms of $R^2$ measures:
\[\theta - \theta_s = -\underbrace{\rho_0}_{alignment} \underbrace{C_{0\Delta Y}}_{trend} \underbrace{C_{0D}}_{selection}
\underbrace{S_0}_{scale}, \]
where,
\[\rho_0 := \operatorname{Cor}\left(g_0 - g_{0s}, O_{XU} - O_X \mid D=0\right),\quad
C^2_{0\Delta Y} := \eta^2_{\Delta Y\sim U\mid X,D=0},\quad
C^2_{0D}:= \frac{1-R^2_{O_{XU}\sim O_{X}|D=0}}{R^2_{O_{XU}\sim O_{X}|D=0}},\]
and
\[S^2_0 := \sigma^2_{0s} \nu^2_{0s}, \qquad
\sigma^2_{0s}:= E[\operatorname{Var}(\Delta Y \mid X,D=0)\mid D=0], \qquad
\nu^2_{0s}:= E\left[\left(\frac{O_X}{O}\right)^2 \middle| D=0\right].\]
\end{Theorem}
Theorem~\ref{thm:main_npm} rewrites the average deviation from parallel trends in terms that more conveniently rely on scale-free partial $R^2$ measures of association characterizing the strength of unobserved confounding. The bias decomposes into the product of three non-identifiable bias factors, $\rho_0$, $C_{0\Delta Y}$, $C_{0D}$---which need to be restricted by plausibility judgments on the strength of confounders---and an identifiable scale factor, $S_0$, which is estimable from the observed data. Note that bias arises only when all bias factors are non-zero.
The ``trend'' component $C^2_{0\Delta Y}$ measures the strength of confounding in explaining trend variation. Formally, this is measured by $\eta^2_{\Delta Y \sim U|X,D=0}$, which is the \emph{non-parametric} partial $R^2$ of $\Delta Y$ with $U$, given $X$, among the untreated. This quantifies how much residual variance of the trend in the control group is explained by the omitted confounders, after taking into account what is already explained by~$X$. The less confounders can explain trend variation in the control group, the less the potential of such confounders to induce deviations from parallel trends. Note this component can always be left unconstrained by the investigator, as it is naturally upper bounded by 1.
The ``selection'' component $C^2_{0D}$ measures the strength of confounding in explaining selection into treatment. Formally, this is driven by $1-R^2_{O_{XU}\sim O_{X}|D=0}$, i.e., the share of variation in the long treatment odds of the control group which cannot be explained by observed covariates alone. While this latter quantity is also an $R^2$ which can be upper bounded by 1, note the selection term $C^2_{0D}$ is given by the ratio, $(1-R^2_{O_{XU}\sim O_{X}|D=0})/R^2_{O_{XU}\sim O_{X}|D=0}$, which can be arbitrarily large. Therefore, the treatment selection component must always be restricted. To aid plausibility judgments regarding $C^2_{0D}$, we provide alternative characterizations of treatment selection in the next section below.
Together, the two sensitivity parameters, $\eta^2_{\Delta Y \sim U|X,D=0}$ and $1-R^2_{O_{XU}\sim O_{X}|D=0}$, characterize the strength of confounding in explaining trends and treatment odds. However, for bias to arise, it is not sufficient for confounders to create errors in the trend and selection components; these errors must be systematically aligned. This is captured by the ``alignment'' component $\rho_0$, defined as the correlation between these two errors in the control arm. Similarly to the trend component, in the worst case, the alignment component can also be left unconstrained by the investigator, as its magnitude is upper bounded by 1.
Finally, the identifiable scale factor $S_0$ translates these three scale-free measures of the strength of confounding back into the original units of the ATT. It is characterized by the strength of \emph{observed covariates} in explaining variation in outcome trend and selection into treatment. The first component, $\sigma_{0s}^2$, measures the leftover variation of the trend in the control group after taking into account the part explained by $X$. Note the more variation $X$ explains of the outcome trend, the less room there is for unobserved confounding to create bias. Moving to $\nu_{0s}^2$, we see that the opposite relationship holds with respect to the strength of covariates in explaining treatment assignment. This term measures how well observed covariates separate treated from control units. The stronger this separation, the less identifying variation is being used to estimate the ATT, and the more leverage unobserved confounding has to drive the bias; $\nu_{0s}^2$ is also directly interpretable as the imbalance of observed covariates between treated and control units, as measured by the chi-squared divergence, as we show next.
\subsection{Making sense of treatment selection strength}
\label{sec:treatment-selection}
As we have seen, the OVB formula for the ATT is particularly sensitive to selection strength, both unobserved (via $C_{0D}^2$) and observed (via $\nu^2_{0s}$). It is therefore useful to develop alternative ways of understanding and reasoning about these quantities. In this section, we provide such characterizations in terms of increases in average treatment odds and covariate imbalance, helping researchers specify plausible upper bounds on the unobserved selection strength $C^2_{0D}$ using the scale most familiar to them in their own applications.
For example, in genetics, psychology, economics and other quantitative social sciences, researchers routinely think in terms of variance explained and related $R^2$ measures \citep{cohen2013statistical,Imbens03,oster_unobservable_2019,cinelli2020making,cinelli2025omitted,chernozhukov2022long}. In medicine and epidemiology, odds and odds ratios are commonly used to characterize selection into treatment \citep{DingVanderweele16}. Covariate imbalance---often summarized by the standardized mean difference, itself a common effect size, \citep{cohen2013statistical}---is a widely used balance diagnostic in observational studies and randomized trials \citep{stuart2010matching}. Each community has developed intuition for what counts as ``large'' or ``small'' on its own scale. Our goal is to formally connect these measures, allowing researchers to approach the OVB problem from different perspectives. The following lemma provides the basic identities linking treatment odds, changes of measure, and covariate imbalance.
\begin{Lemma}
\label{thm:main_connection}
Under the weak overlap condition of Assumption~\ref{asmp:regularity}, the following holds:
\begin{enumerate}[label=(\roman*), leftmargin=*, itemsep=2pt, topsep=3pt]
\item \textbf{Odds ratio as change of measure:}
\(\displaystyle O_{XU}/O
= \mathop{}\!\mathrm d P_{X,U|D=1}/\mathop{}\!\mathrm d P_{X,U|D=0}\).
\item \textbf{Control-to-treated odds identity:}
\(\displaystyle E[O_{XU}^2\mid X,D=0]
=O_XE[O_{XU}\mid X,D=1]\).
\item \textbf{Imbalance as average odds:}
\(\displaystyle \chi^2(P_{X,U|D=1}\|P_{X,U|D=0})
=E[O_{XU}/O\mid D=1]-1\).
\end{enumerate}
\end{Lemma}
Applying these identities yields equivalent expressions for the selection components $C^2_{0D}$ and $\nu^2_{0s}$ in terms of increase in average treatment odds and $\chi^2$-divergence.
\begin{Corollary}[Alternative characterizations of $C_{0D}^2$ and $\nu^2_{0s}$]
\label{thm:alt-selection}
Under the weak overlap condition of Assumption~\ref{asmp:regularity}, $C^2_{0D}$ and $\nu^2_{0s}$ of Theorem~\ref{thm:main_npm} admit the additional interpretations:
\begin{enumerate}
\item Increase in average treatment odds among the treated:
{
\begin{align*}
C_{0D}^2 = \frac{E[O_{XU}\mid D=1] - E[O_X\mid D=1]}{E[O_X\mid D=1]}, \qquad
\nu^2_{0s} = \frac{E[O_X|D=1]}{O}.
\end{align*}
}
\item Covariate imbalance between treated and control units:
{
\begin{align*}
C_{0D}^2 = \frac{\chi^2(P_{X,U|D=1}\|P_{X,U|D=0}) - \chi^2(P_{X|D=1}\|P_{X|D=0})}{\chi^2(P_{X|D=1}\|P_{X|D=0}) + 1}, \qquad
\nu^2_{0s}=\chi^2(P_{X|D=1}\|P_{X|D=0})+1.
\end{align*}
}
\end{enumerate}
\end{Corollary}
The odds characterization gives a direct way to translate substantive restrictions on the propensity score into bounds on $C_{0D}^2$. For example, if one believes that the inclusion of $U$ at most doubles the average odds of treatment, among treated units, then Corollary~\ref{thm:alt-selection} implies $C_{0D}^2 \le 1$. This relationship also clarifies the connection between the selection component $C^2_{0D}$ of the OVB formula and sensitivity models based on worst-case bounds on odds ratios. The marginal sensitivity model of \citet{tan2006distributional}, for instance, assumes
\[
\Lambda^{-1} \leq O_{XU}/O_X \leq \Lambda
\]
almost surely, for some $\Lambda >1$. Such a restriction implies the sharp bound (see Corollary~\ref{cor:msm-parameterization}),
\[
C_{0D}^2 \leq (\Lambda-1)^2/\Lambda.
\]
Thus, if one believes $U$ cannot double (or halve) the odds of treatment not only on average, but uniformly, this implies the tighter restriction $C^2_{0D} \leq 1/2$. This also makes it clear that worst-case restrictions on odds ratios are sufficient, but not necessary, to bound the bias.
Beyond treatment odds, Corollary~\ref{thm:alt-selection} also provides a covariate balance interpretation of $C_{0D}^2$. It shows that $C_{0D}^2$ captures the additional imbalance introduced by $U$, relative to the imbalance already captured by $X$. Checking covariate balance is standard practice in applied research, often through standardized mean differences (SMD) \citep{imbens2015causal,baker2025difference}. As the OVB formula shows, however, the imbalance measure directly relevant for quantifying the bias of the ATT is the $\chi^2$-divergence, which summarizes the full distributional difference between treated and control groups.
\subsection{On the choice of parameterization of the ATT bias}
\label{sec:parameterization}
Theorem~\ref{thm:main_npm} provides a novel parameterization of the omitted variable bias of the ATT. This choice of parameterization is not innocuous. It determines the quantities about which researchers must make plausibility judgments, as well as the allowable values that such sensitivity parameters can take. In this section, we present some key differences of our result from prior OVB formulas for the ATT derived in \citet{chernozhukov2022long} and \citet{huang2025variance}. A complete analysis is deferred to Appendix~\ref{app:all-comparison}.
The result of \citet{chernozhukov2022long} applied to the ATT can be written as
\begin{align}
\theta-\theta_s
&=\rho\,C_{\Delta Y}\,C_D\,S,
\label{eq:lls}
\end{align}
where
\(
C_{\Delta Y}^2:=\eta^2_{\Delta Y\sim U\mid X,D},
C_D^2:=\frac{E[O_{XU}]-E[O_X]}{E[O_X]},
S^2:=E\!\left[\operatorname{Var}(\Delta Y\mid X,D)\right]\frac{E[O_X]}{P(D=1)^2},
\) and
\(
\rho:=
\operatorname{Cor}\left(
E[\Delta Y\mid D,X,U]-E[\Delta Y\mid D,X],\;
-(1-D)(O_{XU}-O_X)
\right).
\)
Note that the sensitivity parameters in (\ref{eq:lls}) are defined in terms of the full population, even though Theorem~\ref{thm:main_npm} shows that the bias for the ATT can be localized solely on the untreated units. Therefore, the parameterization in (\ref{eq:lls}) asks users to make plausibility judgments on quantities that are not directly relevant for the bias, such as $\eta^2_{\Delta Y\sim U\mid X,D}$ instead of $\eta^2_{\Delta Y\sim U\mid X,D=0}$.
This fact has consequences beyond interpretation. In particular, these unconditional trend and alignment bias factors obey the following restriction.
\begin{Proposition}
\label{prop:lls-restriction}
The trend and alignment components in equation~\eqref{eq:lls} satisfy
\[
\rho^2 C_{\Delta Y}^2
=
\frac{P(D=0)\,\sigma_{0s}^2}
{E[\operatorname{Var}(\Delta Y\mid X,D)]}\,
\rho_0^2\,C_{0\Delta Y}^2
\;\leq\;
\frac{P(D=0)\,\sigma_{0s}^2}
{E[\operatorname{Var}(\Delta Y\mid X,D)]}.
\]
\end{Proposition}
Thus, whenever the residual outcome variance among treated units is positive, the upper bound in Proposition~\ref{prop:lls-restriction} is strictly smaller than one, meaning that the trend and alignment components cannot vary freely. For example, applying the bounds $|\rho|\leq1$ and $C_{\Delta Y}\leq1$ directly to (\ref{eq:lls}) gives
\(
|\theta-\theta_s|\leq C_D\,S,
\)
whereas Proposition~\ref{prop:lls-restriction} implies the smaller bound
\(
|\theta-\theta_s|
\leq
\sqrt{
\frac{P(D=0)\,\sigma_{0s}^2}
{E[\operatorname{Var}(\Delta Y\mid X,D)]}
}
\,C_D\,S
=
C_{0D}\,S_0,
\)
which is exactly the bound implied by our OVB decomposition in Theorem~\ref{thm:main_npm}.
As for the characterization of \citet{huang2025variance}, in Appendix~\ref{app:hp-comparison} we show that it admits the following representation:
\begin{align}
\theta-\theta_s
&=-\rho_{w0}\,C_{w0D}\,S_{w0},
\label{eq:hp}
\end{align}
where
\(
\rho_{w0}:=\operatorname{Cor}(\Delta Y,O_{XU}-O_X\mid D=0),
C_{w0D}^2:=\frac{E[O_{XU}]-E[O_X]}{E[O_X]-O},
\)
and
\(
S_{w0}^2
:=
\operatorname{Var}(\Delta Y\mid D=0)\,
\frac{P(D=0)\{E[O_X]-O\}}{P(D=1)^2}.
\) For sensitivity analysis, \citet{huang2025variance} propose bounding $|\rho_{w0}|$ by the estimable quantity $\bar\rho_{w0}:= \sqrt{
1-\operatorname{Cor}^2(O_X,\Delta Y\mid D=0)
}.$
The first noticeable difference with respect to our result in Theorem~\ref{thm:main_npm} concerns the selection components. If $X$ does not predict treatment, which includes the canonical 2x2 DiD setup without covariates, $C_{w0D}$ has a zero denominator, and the decomposition (\ref{eq:hp}) is not well defined. If $X$ only weakly predicts treatment, (\ref{eq:hp}) is well defined, but $C_{w0D}$ can grow arbitrarily large, while the estimation of $S_{w0}^2$ becomes difficult.
The second difference concerns the trend and alignment components. The parameterization in equation~\eqref{eq:hp} does not contain a separate sensitivity parameter measuring the strength of confounding with the untreated outcome trend. Instead, trend and alignment components are bundled into a single correlation $\rho_{w0}$, and separate plausibility judgments about these parameters are not possible. Furthermore, the parameterization obeys the following restriction.
\begin{Proposition}
\label{prop:hp-correlation}
When well defined, $\rho_{w0}$ in equation~\eqref{eq:hp} satisfies
\[
|\rho_{w0}|
\leq
\sqrt{
\frac{\sigma_{0s}^2}
{\operatorname{Var}(\Delta Y\mid D=0)}
}
\leq
\bar\rho_{w0}.
\]
The first equality holds iff $\rho_0^2C_{0\Delta Y}^2=1$, and the second iff $g_{0s}$ is an affine function of~$O_X$ among the untreated.
\end{Proposition}
Consequently, their reported upper bound is larger than necessary. When expressed in our parameterization, it can be attained only when the trend and alignment components reach their worst-case values and, among untreated units, the conditional outcome trend can be expressed as an affine function of the treatment odds.
The consequences of these different parameterizations are especially transparent under the marginal sensitivity model (MSM) of \citet{tan2006distributional} introduced in Section~\ref{sec:treatment-selection}. The same restriction on treatment odds yields the following sharp bounds.
\begin{Corollary}
\label{cor:msm-parameterization}
Under the MSM, the selection parameters satisfy
\[
C_D^2\leq\left(\frac{E[O_X]-p}{E[O_X]}\right)\frac{(\Lambda-1)^2}{\Lambda},\qquad C_{0D}^2\leq\frac{(\Lambda-1)^2}{\Lambda},\qquad C_{w0D}^2\leq\left(\frac{E[O_X]-p}{E[O_X]-O}\right)\frac{(\Lambda-1)^2}{\Lambda}.
\]
All three bounds are sharp, with the last applying when $C_{w0D}^2$ is well defined.
\end{Corollary}
Thus, under the same MSM, the sharp bound for $C_{0D}^2$ is invariant to the observed treatment assignment, whereas the sharp bounds for the alternative parameterizations depend on it, with that for $C_{w0D}^2$ diverging and becoming uninformative as observed imbalance vanishes.
\section{Benchmarking the strength of unobserved confounding}
\label{sec:benchmarking}
The main difficulty in sensitivity analysis is specifying plausible strengths for unobserved confounding. While in some cases researchers may not be able to make absolute judgments about the magnitude of an omitted confounder, they may still have grounds to make judgments of its relative importance. For example, a researcher may be willing to argue that any unobserved confounder is no stronger than some observed covariate, or that post-treatment confounding is no stronger than the confounding revealed by pre-treatment periods. In this section, we develop two benchmarking procedures that formalize these comparisons.
\subsection{Comparing unobserved confounders against observed covariates}
\label{sec:bench-covariates}
Our first analysis compares the gains in explanatory power from unobserved confounders with those from key observed covariates. Let $X_j$ denote the benchmark covariates against which we wish to calibrate the strength of $U$, and let $X_{-j}$ denote the remaining covariates, so that $X = (X_j, X_{-j})$. We define the following relative strength parameters:
\begin{align*}
k_{0\Delta Y,j} := \frac{ \eta^2_{\Delta Y \sim U,X_j \mid X_{-j},D=0} - \eta^2_{\Delta Y \sim X_j \mid X_{-j},D=0}}{\eta^2_{\Delta Y \sim X_j \mid X_{-j},D=0}},
\quad
k_{0D,j} := \frac{R^2_{O_{X} \sim O_{X_{-j}}\mid D=0} - R^2_{O_{XU} \sim O_{X_{-j}}\mid D=0}}{1 - R^2_{O_X \sim O_{X_{-j}}\mid D=0}}.
\end{align*}
The parameter $k_{0\Delta Y, j}$ measures how much additional variation in the untreated outcome evolution is explained by $U$, relative to the observed gains in variation explained by $X_j$. For example, $k_{0\Delta Y, j} =1$ implies that further adding $U$ to the trend regression would lead to similar additive gains in explanatory power to those observed by adding $X_j$. The parameter $k_{0D, j}$ uses the same logic to measure the relative gain in explanatory power with the treatment odds. As discussed in Section~\ref{sec:treatment-selection}, $k_{0D, j}$ admits alternative interpretations in terms of the relative increase in the \textit{average treatment odds} or the additional \textit{imbalance between treated and control groups} attributable to $U$, relative to that attributable to $X_j$.
These measures allow us to re-express the absolute strength of $U$ in the OVB formula of Theorem~\ref{thm:main_npm} in terms of its relative strength as compared to that of $X_j$:
\begin{align*}
\eta^2_{\Delta Y \sim U \mid X,D=0} = k_{0\Delta Y,j}
\left(\frac{\eta^2_{\Delta Y \sim X_j \mid X_{-j},D=0}}{1 - \eta^2_{\Delta Y \sim X_j \mid X_{-j},D=0}}\right),
\quad
1 - R^2_{O_{XU} \sim O_X\mid D=0} &= k_{0D,j}
\left(\frac{1 - R^2_{O_X \sim O_{X_{-j}}\mid D=0}}{R^2_{O_X \sim O_{X_{-j}}\mid D=0}}\right).
\end{align*}
Therefore, plausibility judgments on the \emph{relative importance} of $U$ compared to $X_j$, both in explaining outcome evolution, and treatment odds, can be leveraged to bound the bias.
Finally, one may benchmark plausible values for the correlation of errors $\rho_0$. Notice that $\rho_0$ is not a measure of explanatory power of the confounders. Rather, it measures how systematically $U$ shifts the odds of treatment and outcome evolution in the same direction. This is partly connected to the functional form of confounding. For example, if $U$ enters the treatment and trend equations with different functional forms, this will generally attenuate the bias. To calibrate this empirically, we may use as a reference the observed correlation between the errors in the outcome regression and treatment odds induced by $X_j$, as measured by
\[
\rho_{0,j}:=\operatorname{Cor}\!\left(g_{0s} - g_{0s,-j},\; O_X - O_{X_{-j}} \,\big|\, D=0\right).
\]
Benchmarking the bias factors $\rho_0$, $C_{0\Delta Y}$ and $C_{0D}$ separately allows researchers to entertain richer confounding scenarios. For example, one may posit that an unobserved confounder is comparable to $X_j$ in terms of explaining treatment selection and outcome evolution, while remaining agnostic about alignment. Further details can be found in Appendix~\ref{app:bench-covariates}.
\subsection{Comparing post-treatment bias against pre-treatment bias}
\label{sec:bench-pre-trend}
We now connect the OVB framework to pre-trend extrapolation approaches such as \citet{rambachan2023more}. Existing methods typically bound the post-treatment bias by extrapolating deviations from parallel trends observed before treatment. Our decomposition shows that these extrapolation assumptions can instead be interpreted as restrictions on the underlying confounding mechanism itself, and also suggests natural alternative ways of both interpreting and augmenting them. We consider the simplest case of one additional pre-treatment period, and leave extensions to multiple time periods to future work.
\subsubsection*{Quantifying pre-treatment bias}
Consider three time periods $t \in \{0, 1, 2\}$, with treatment occurring after
$t = 1$. To estimate the (placebo) effect in the pre-treatment period, we take
$t=0$ as the reference and estimate the effect at $t=1$. Let
$\Delta Y^{\text{pre}} := Y_1 - Y_0$ denote the observed outcome evolution prior
to treatment, and let $(X^{\text{pre}}, U^{\text{pre}})$ denote the covariates
that would make parallel trends hold in the pre-treatment period. Throughout, we
write $\theta^{\text{pre}}$, $\theta_s^{\text{pre}}\neq 0$, $\rho_0^{\text{pre}}$,
$C^{\text{pre}}_{0\Delta Y}$, $C^{\text{pre}}_{0D}$, and
$S_0^{\text{pre}} = \sigma_{0s}^{\text{pre}}\nu_{0s}^{\text{pre}}$ for the
estimands, bias factors, and scale factor of Section~\ref{sec:ovb}, computed with
$(\Delta Y^{\text{pre}}, X^{\text{pre}}, U^{\text{pre}})$ in place of
$(\Delta Y, X, U)$, with $D$ unchanged throughout.
Since no treatment has yet occurred, $\theta^{\text{pre}} = 0$, and the
\emph{pre-treatment bias} from omitting confounders reduces to
$\theta_s^{\text{pre}}$, the (average) deviation from parallel trends before
treatment. Applying Theorem~\ref{thm:main_npm} to the pre-treatment period gives
\[
|\rho^{\text{pre}}_0|\, C^{\text{pre}}_{0\Delta Y}\, C^{\text{pre}}_{0D}\,
S_0^{\text{pre}} = |\theta_s^{\text{pre}}|,
\]
so $\theta_s^{\text{pre}}$ does not identify the pre-treatment bias factors
individually, but constrains their product. We allow $(X^{\text{pre}},
U^{\text{pre}})$ to differ from $(X,U)$, but in many cases, such as in our
empirical application, they coincide. When that happens, the treatment selection component is then immediately transportable across periods, since
$C^{2,\text{pre}}_{0D}=C^{2}_{0D}$ and $\nu_{0s}^{2,\text{pre}} = \nu_{0s}^{2}$.
Therefore, two natural strategies emerge for extrapolating the pre-treatment bias to the post-treatment period in this setting. One is to directly transport the bias magnitude, as is currently done in pre-trend extrapolation approaches; another is to transport only (some of) the bias factors, while re-estimating the scaling factor $S_0$ from post-treatment data.
\subsubsection*{Extrapolating the bias magnitude}
As in \cite{rambachan2023more}, researchers may be willing to assume that the post-treatment bias is no larger than $k$ times the bias observed in the pre-treatment period. This corresponds to the assumption that $|\theta - \theta_s| \leq k|\theta_s^{\text{pre}}|$, and yields the following bounds
\begin{align*}
\theta_{\pm} = \theta_s \pm k |\theta_s^{\text{pre}}|,
\end{align*}
where both $\theta_s$ and $\theta_s^{\text{pre}}$ are estimable from the data. Mapping this to unobserved confounders, this approach is equivalent to imposing a constraint on the full product of bias factors and the scale factor, i.e.,
\[
|\rho_0|\, C_{0\Delta Y}\, C_{0D}\, S_0 \leq k |\rho^{\text{pre}}_0| C^{\text{pre}}_{0\Delta Y} C_{0D}^{\text{pre}}\, S_0^{\text{pre}} = k |\theta_s^{\text{pre}}|.
\]
This connection also clarifies how different values of $k$ could arise; for example, consider $|\theta-\theta_s|=k|\theta_s^{\text{pre}}|$. Writing
$k_\rho := |\rho_0|/|\rho_0^{\text{pre}}|$,
$k_{\Delta Y} := C_{0\Delta Y}/C_{0\Delta Y}^{\text{pre}}$,
$k_D := C_{0D}/C_{0D}^{\text{pre}}$, and
$k_S := S_0/S_0^{\text{pre}}$, we then have $k = k_\rho k_{\Delta Y} k_D k_S$, so the multiplier $k$ collects the relative changes in each component of the bias across periods (and, in fact, $k_S$ is estimable).
\subsubsection*{Extrapolating the bias factors}
The previous considerations suggest another natural approach for extrapolating the bias. Rather than transporting the full magnitude of the pre-treatment bias, we can transport only the bias factors associated with the unobserved confounding mechanism, while allowing the observed scale component to change across periods. This yields the following bounds
\begin{align*}
\theta_{\pm} &= \theta_s \pm k \left(\frac{S_0}{S_0^{\text{pre}}}\right) |\theta_s^{\text{pre}}| ,
\end{align*}
where $\theta_s$, $\theta_s^{\text{pre}}$, $S_0^{\text{pre}}$, and $S_0$ are estimable from the data, and $k = k_{\rho} k_{\Delta Y} k_{D}$.
Regardless of the choice of pre-trend extrapolation, the OVB decomposition of Section~\ref{sec:ovb}, together with benchmarking against observed covariates in Section~\ref{sec:bench-covariates}, allows researchers to clearly translate these assumptions into statements about the underlying strength of confounding. For example, one can ask what kind of unobserved confounder would be required to generate the extrapolated post-treatment bias, in terms of its strength in treatment selection, outcome trends, and the correlation between the two, and whether such strength of confounding is comparable in magnitude to that of observed covariates.
\section{Estimation and inference}
\label{sec:est-and-inf}
The previous sections characterize the OVB formula and the associated sensitivity parameters at the population level. We now discuss estimation and inference in finite samples. In what follows, we assume we have an i.i.d. sample from the observed data distribution. Appendix~\ref{app:simulations} reports extensive simulation exercises to verify the properties of our estimators and confidence intervals.
\subsection{Inference under direct restrictions on confounding}
Given plausibility judgments on the magnitude of sensitivity parameters $|\rho_0|$, $C_{0\Delta Y}$, $C_{0D}$, we have the following bounds on $\theta$:
\begin{align}
\theta_{\pm} &= \theta_s \pm |\rho_0|\,C_{0\Delta Y}\,C_{0D}\,S_0, \quad\text{with}\quad S_0^2 = \sigma_{0s}^2\,\nu_{0s}^2, \nonumber
\end{align}
where $\theta_s$ and $S_0^2$ are estimable from the observed data.
We estimate these components using debiased machine learning (DML), which combines Neyman orthogonal scores with cross-fitting to enable the use of machine learning methods for nuisance estimation \citep{chernozhukov2018double,chernozhukov2022long}. Specifically, let $Z := (\Delta Y,D,X)$, $p:=P(D=1)$ and $\pi:=P(D=1\mid X)$. Estimation proceeds via Algorithm~\ref{alg:DML}, with orthogonal scores
\begin{align}
\psi_{\theta_s}(Z; g_{0s}, \pi, p) &= \left(\frac{D}{p} - \frac{(1-D)}{(1-p)}\frac{O_{X}}{O}\right)(\Delta Y - g_{0s}) - \frac{D\theta_s}{p},
\label{eq:psi-theta} \\
\psi_{\sigma_{0s}^2}(Z;g_{0s},p) &= \frac{1-D}{1-p}\left((\Delta Y - g_{0s})^2 - \sigma_{0s}^2\right), \label{eq:psi-sigma} \\
\psi_{\nu_{0s}^2}(Z;\pi,p)
&= 2\frac{D}{p}\left(\frac{O_{X}}{O} - \nu_{0s}^2\right) - \frac{1-D}{1-p}\left(\left(\frac{O_{X}}{O}\right)^2 - \nu_{0s}^2\right).
\label{eq:psi-nu}
\end{align}
We denote the resulting estimates as
\begin{align*}
\widehat{\theta}_s := \text{DML}(\psi_{\theta_s}), \qquad
\widehat{\sigma}_{0s}^2 := \text{DML}(\psi_{\sigma_{0s}^2})\qquad
\widehat{\nu}_{0s}^2 := \text{DML}(\psi_{\nu_{0s}^2}).
\end{align*}
\begin{algorithm}[t]
\small
\caption{DML for estimable components: DML($\psi_{\beta}$)}
\begin{algorithmic}
\State \textbf{Input:} The Neyman orthogonal score $\psi_{\beta}(Z;\eta)$, sample $\{Z_i\}_{i=1}^n$, and number of folds $L$.
\State \textbf{Sample Splitting:} Randomly partition the sample into $L$ folds of approximately equal size. Denote each fold by $I_l$ and its complement by $I_l^c$, for $l = 1,\dots,L$.
\For{$l = 1,\dots,L$}
\State Estimate the nuisance parameters using observations in $I_l^c$. Denote the estimates by $\widehat{\eta}_l$.
\EndFor
\State \textbf{Estimate:} Construct the estimator $\widehat{\beta}$ as the solution of $0 = \frac{1}{n}\sum_{l=1}^L\sum_{i \in I_l}\psi_{\widehat{\beta}}(Z_i;\widehat{\eta}_l)$.
\State \textbf{Return:} The estimate $\widehat{\beta}$ and the estimated scores $\psi_{\widehat{\beta}}(Z_i;\widehat{\eta}_l)$ for each $i \in I_l$ and each $l$.
\end{algorithmic}
\label{alg:DML}
\end{algorithm}
Let $\beta \in \{\theta_s, \sigma^2_{0s},\nu^2_{0s} \}$ denote a generic target parameter and $\eta = (g_{0s}, \pi, p)$ the vector of nuisance parameters. The scores (\ref{eq:psi-theta})-(\ref{eq:psi-nu}) admit the representation $\psi_{\beta}(Z;\eta) = \psi_\beta^a(Z;\eta)\,\beta + \psi_\beta^b(Z;\eta)$. An application of the results in \citet{chernozhukov2018double} for linear score functions yields the following lemma, which establishes the asymptotic linearity and normality of the DML estimators, and characterizes their influence functions.
\begin{Lemma}[Asymptotic linearity of the DML estimators]
\label{lemma:Properties of the DML Bound Estimators}
Suppose that Assumptions 3.1 and 3.2 from \citet{chernozhukov2018double} hold for each of the scores $\psi_{\beta}$ above, and for the estimators of the nuisance parameters $\widehat{\eta}_l$. Then, the DML estimators $\widehat{\beta}$ are asymptotically linear and Gaussian with the following properties
\begin{align}
\sqrt{n}(\widehat{\beta} - \beta) &= \frac{1}{\sqrt{n}}\sum_{i=1}^n\underbrace{\frac{\psi_{\beta}(Z_i;\eta)}{-E[\psi_\beta^a(Z;\eta)]}}_{=: \varphi^0_{\beta}(Z_i)} + o_p(1) \overset{d}{\rightarrow} N(0,\sigma_{\varphi_{\beta}}^2), \text{ with }
\sigma_{\varphi_{\beta}}^2 := E[(\varphi^0_{\beta}(Z))^2], \nonumber
\end{align}
where $\varphi^0_{\beta}(\cdot)$ is the influence function. Following Theorem 3.2 in \citet{chernozhukov2018double}, we denote the estimate of the influence function as $\widehat{\varphi}_{\widehat{\beta}}(\cdot) := -\psi_{\widehat{\beta}}(\cdot;\widehat{\eta}_l)\big/E_n[\psi^a(Z;\widehat{\eta}_l)]$ and use $\widehat{\sigma}_{\varphi_{\beta}}^2 := E_n\left[(\widehat{\varphi}_{\widehat{\beta}}(Z))^2\right]$ as the estimator for $\sigma^2_{\varphi_\beta}$.
\end{Lemma}
\begin{remark}
Note that the DML estimator for $\nu^2_{0s}$ can be used to estimate the chi-squared divergence of observed covariates, since $\nu^2_{0s} -1 = \chi^2(P_{X|D=1}\|P_{X|D=0})$---see Corollary~\ref{thm:alt-selection}.
\end{remark}
The plug-in estimator for the bounds thus takes the following form,
\begin{equation*}
\widehat{\theta}_{\pm} := \widehat{\theta}_s \pm |\rho_0|\,C_{0\Delta Y}\,C_{0D}\,\sqrt{\widehat{\sigma}_{0s}^2\,\widehat{\nu}_{0s}^2}.
\end{equation*}
Since $\widehat{\theta}_{\pm}$ is a smooth function of asymptotically linear estimators, we can obtain confidence intervals for the bounds by applying the delta method.
\begin{Theorem}[Confidence intervals for the bounds] \label{thm: inference}
Under the Assumptions of Lemma \ref{lemma:Properties of the DML Bound Estimators}, the estimator $\widehat{\theta}_{\pm}$ is asymptotically linear and Gaussian with the following properties
\begin{align}
\sqrt{n}(\widehat{\theta}_{\pm} - \theta_{\pm}) &= \frac{1}{\sqrt{n}}\sum_{i=1}^n\varphi^0_{\theta_{\pm}}(Z_i) + o_p(1) \overset{d}{\rightarrow} N(0, \sigma^2_{\varphi_{\theta_{\pm}}}), \text{ with } \sigma^2_{\varphi_{\theta_{\pm}}} = E[(\varphi^0_{\theta_{\pm}}(Z))^2], \nonumber \\
\text{where } \varphi^0_{\theta_{\pm}}(Z) &= \varphi_{\theta_s}^0(Z) \pm |\rho_0|C_{0\Delta Y}C_{0D} \times \underbrace{\frac{1}{2S_0}\left(\sigma_{0s}^2\varphi^0_{\nu_{0s}^2}(Z) + \nu_{0s}^2\varphi^0_{\sigma_{0s}^2}(Z)\right)}_{=:\varphi_{S_0}^0(Z)}. \nonumber
\end{align}
Let $\Phi$ denote the CDF of the standard normal and $\alpha$ denote the significance level. The confidence interval
\begin{align}
[l, u] &= \left[\widehat{\theta}_{-} - \Phi^{-1}(1 - \alpha)\sqrt{\frac{E[(\varphi^0_{\theta_{-}}(Z))^2]}{n}}, ~\widehat{\theta}_{+} + \Phi^{-1}(1 - \alpha)\sqrt{\frac{E[(\varphi^0_{\theta_{+}}(Z))^2]}{n}}\right], \nonumber
\end{align}
satisfies $P(\theta_{-} \in [l,\infty)) \rightarrow 1-\alpha$ and $P(\theta_{+} \in (-\infty, u]) \rightarrow 1-\alpha$.
These results continue to hold if we replace $E[(\varphi^0_{\theta_{\pm}}(Z))^2]$ with $E_n[(\widehat{\varphi}_{\widehat{\theta}_{\pm}}(Z))^2]$.
\end{Theorem}
\subsection{Inference under benchmarking restrictions on confounding}
\label{sec:est-inf-bench}
In the benchmarking analysis introduced in Section~\ref{sec:benchmarking}, additional components characterizing the strength of observed confounding are estimable from the data. We now discuss how to account for the uncertainty in estimating these components. We begin with benchmarking against observed covariates.
Given plausibility judgments on the relative strength parameters $k_{0\Delta Y,j},k_{0D,j}$, and setting $|\rho_0|=|\rho_{0,j}|$, the plug-in estimator of the bias bound now takes the following form:
\begin{align*}
\widehat{\theta}_{\pm} = \widehat{\theta}_s \pm |\widehat{\rho}_{0,j}|\times \sqrt{k_{0\Delta Y,j}\widehat{G}_{0\Delta Y,j} \times \frac{k_{0D,j}\widehat{G}_{0D,j}}{1-k_{0D,j}\widehat{G}_{0D,j}}}\times \sqrt{\widehat{\sigma}_{0s}^2\,\widehat{\nu}_{0s}^2}
\end{align*}
where,
\begin{align*}
\widehat{G}_{0\Delta Y,j} := \frac{\widehat{\sigma}_{0s,-j}^2 - \widehat{\sigma}_{0s}^2}{\widehat{\sigma}_{0s}^2}, \quad
\widehat{G}_{0D,j} := \frac{\widehat{\nu}_{0s}^2 - \widehat{\nu}_{0s,-j}^2}{\widehat{\nu}_{0s,-j}^2}, \ \text{and} \quad
\widehat{\rho}_{0,j} = \frac{-(\widehat{\theta}_s - \widehat{\theta}_{s,-j})}{\sqrt{(\widehat{\sigma}_{0s,-j}^2 - \widehat{\sigma}_{0s}^2) \times (\widehat{\nu}_{0s}^2 - \widehat{\nu}_{0s,-j}^2)}}.
\end{align*}
Here $\theta_{s,-j}$ denotes the ATT estimand conditional on $X_{-j}$ alone, and $\sigma_{0s,-j}^2$ and $\nu_{0s,-j}^2$ are the scaling factors in our OVB decomposition of $\theta_s - \theta_{s,-j}$. The orthogonal scores for these parameters have the same form as (\ref{eq:psi-theta})-(\ref{eq:psi-nu}), with the only difference that $X_{-j}$ is used for estimation in place of $X$.
We now move to benchmarking against pre-trend bias. For fixed $k$, the first approach yields the following plug-in estimator of the bias bound
\begin{align*}
\widehat{\theta}_{\pm} &= \widehat{\theta}_s \pm k |\widehat{\theta}_s^{\text{pre}}|,
\end{align*}
whereas the second approach gives
\begin{align*}
\widehat{\theta}_{\pm} &= \widehat{\theta}_s \pm k \left(\frac{\widehat{S}_0}{\widehat{S}_0^{\text{pre}}} \right) |\widehat{\theta}_s^{\text{pre}}|,
\qquad \text{with}
\quad \widehat{S}_0^2 = \widehat{\sigma}_{0s}^2\,\widehat{\nu}_{0s}^2,
\qquad \text{and}
\quad \widehat{S}_0^{2,\text{pre}} = \widehat{\sigma}_{0s}^{2,\text{pre}}\,\widehat{\nu}_{0s}^{2,\text{pre}}.
\end{align*}
Under the regularity conditions stated in Appendix~\ref{app:bench}, $\widehat{\theta}_{\pm}$ is a smooth function of asymptotically linear estimators whose influence functions were characterized above. Thus, inference for the bounds follows from the delta method. See the Appendix for corresponding influence functions and asymptotic results.
\subsection{Sensitivity statistics for routine reporting}
\label{sec:statistics}
We now introduce two sensitivity statistics that measure the minimum strength of confounding required to invalidate the conclusions of a DiD study. These statistics require no assumptions about the strength of unobserved confounding; rather, they communicate what one must be prepared to believe in order to rule out confounding that would be problematic. They serve as quick summaries of the overall robustness of the ATT estimates against systematic biases \citep{cinelli2020making, cinelli2025omitted,chernozhukov2022long}.
For $\bar R^2_{\Delta Y},\bar R^2_D\in[0,1]$, let
\(
\text{CI}^{\max}_{1-\alpha,\bar R^2_{\Delta Y},\bar R^2_D}(\theta)
\)
denote the widest confidence interval constructed from Theorem~\ref{thm: inference} under the restrictions
\(
|\rho_0|\leq1,~
\eta^2_{\Delta Y\sim U\mid X,D=0}\leq\bar R^2_{\Delta Y},~
1-R^2_{O_{XU}\sim O_X\mid D=0}\leq\bar R^2_D.
\)
The first sensitivity statistic we propose asks: leaving alignment and trend entirely unrestricted, what is the minimum strength of confounding with treatment selection alone that would lead one to not reject $H_0:\theta=\theta^*$? This yields the \emph{extreme robustness value} (XRV),
\[
\text{XRV}_{\theta^*,\alpha}(\theta)
:=
\inf\left\{
\text{XRV}:
\theta^*\in
\text{CI}^{\max}_{1-\alpha,1,\text{XRV}}(\theta)
\right\}.
\]
Any confounding scenario with selection strength below $\text{XRV}_{\theta^*,\alpha}(\theta)$ is logically incapable of overturning the original conclusions, regardless of how much variation such confounders explain of the untreated trend.
Leaving the association of the confounder with the trend completely unrestricted may be too conservative. Thus, the second sensitivity statistic considers the minimum strength of confounding in explaining both the untreated outcome trend and treatment selection jointly. The \emph{robustness value} is
\[
\text{RV}_{\theta^*,\alpha}(\theta)
:=
\inf\left\{
\text{RV}:
\theta^*\in
\text{CI}^{\max}_{1-\alpha,\text{RV},\text{RV}}(\theta)
\right\}.
\]
Any unobserved confounding with trend and selection strength below $\text{RV}_{\theta^*,\alpha}(\theta)$ is incapable of overturning the results of the study.
Appendix~\ref{app:rv} provides the general definition and inferential properties of these quantities. In particular, it shows that $\text{RV}_{\theta^*,\alpha}(\theta)$ and $\text{XRV}_{\theta^*,\alpha}(\theta)$ are asymptotically valid lower confidence bounds for their population counterparts.
\section{Applying the OVB framework to the sensitivity of DiD}
\label{sec:app}
We now return to the minimum wage example of Section~\ref{sec:background}, and show how to deploy the results of Sections~\ref{sec:ovb} to \ref{sec:est-and-inf} to answer: (i) How strong would unobserved confounders have to be to overturn the estimated negative effect in 2007? (ii) How does this required strength compare to that of the observed covariates? and, (iii) What does the 2006 pre-trend imply about the nature of unobserved confounding when read through the OVB decomposition?
\subsection{Minimal sensitivity reporting}
Table~\ref{tab: MW-results-robustness_main} shows our proposal for minimal sensitivity reporting of DiD estimates. The first columns report usual estimates of the treatment effect in 2007 under the assumption that parallel trends holds conditionally on the observed covariates alone. In addition to these estimates, we propose researchers report the extreme robustness value and the robustness value \citep{cinelli2020making,cinelli2025omitted,chernozhukov2022long}. As discussed in Section~\ref{sec:statistics}, these statistics quickly convey how robust the estimated effect is to the presence of omitted variables.
The extreme robustness value ($\text{XRV}_{\theta^*=0,\ \alpha=0.05}$) of $0.2\%$ means that confounders that explain at most 0.2\% of the variation in treatment odds cannot explain away the estimated effect, at the 5\% significance level, even if such confounders were to explain all leftover variation of the outcome evolution. Alternatively, the same number can be interpreted as confounders that induce at most a $\approx0.2\%$ increase in the average odds of treatment among the treated. When the XRV is large, this may provide sufficient evidence for a causal effect in and of itself, by virtue of ruling out confounders that explain a large fraction of treatment assignment. This turns out not to be the case in our application, as it seems difficult to rule out confounders that change the odds of treatment by 0.2\%.
We thus move to examining the minimum \emph{joint} strength that confounders need to have, both in terms of explaining treatment selection and outcome evolution, in order to explain away the results.
\begin{table}[h]
\centering
{\footnotesize
\begin{tabular}{ccccc}
\toprule
\multicolumn{3}{c}{\textbf{Results Under Conditional Parallel Trends}}
& \multicolumn{2}{c}{\textbf{Robustness Values}} \\
\cmidrule(lr){1-3} \cmidrule(lr){4-5}
\textbf{Short Estimate ($\widehat{\theta}_{s,2007}$)}
& \textbf{Std. Error}
& \textbf{Confidence Interval}
& $\text{RV}_{\theta^*=0,\ \alpha=0.05}$
& $\text{XRV}_{\theta^*=0,\ \alpha=0.05}$\\
\midrule
-0.0366
& 0.0105
& [-0.0572;\;-0.0160]
& 4.38\%
& 0.20\% \\
\bottomrule
\end{tabular}
}
\caption{{\footnotesize Minimal Sensitivity Reporting.
}}
\label{tab: MW-results-robustness_main}
\end{table}
The robustness value of $\text{RV}_{\theta^*=0,\ \alpha=0.05}=4.38\%$ indicates that unobserved confounders explaining less than $4.38\%$ of the residual variation in both treatment odds and the outcome evolution of untreated units are not sufficiently strong to render the negative effect insignificant, in the sense of shifting the upper confidence bound to zero, at the $5\%$ significance level. As before, the alternative characterizations of selection, discussed in Section~\ref{sec:treatment-selection}, offer two additional interpretations of this value---it corresponds to a $4.58\%$ relative increase in the observed average treatment odds among treated units. For a single independent binary confounder, this corresponds to a standardized mean difference of about $0.2$---which is not implausible in light of observed imbalances, as we have seen in Section~\ref{sec:background}.
\subsection{Sensitivity contour plots and benchmarking}
Overall, neither the XRV nor the RV allowed us to quickly rule out confounding of magnitudes that would be problematic. We thus turn to the benchmarking procedure of Section~\ref{sec:bench-covariates}, which compares the gains in explanatory power from unobserved confounders with those from key observed covariates. Table~\ref{tab:benchmark_conf_bound} in the Appendix reports one-sided 95\% confidence bounds implied by an unobserved confounder comparable in strength and alignment to each observed covariate. These bounds account for the uncertainty in estimating the benchmark components, using the results of Section~\ref{sec:est-inf-bench}. We see that latent confounders comparable to observed covariates are not sufficiently strong to explain away the estimated effect. Nor would such confounders be able to reproduce the 2006 pre-trend deviation.
We now turn to examining the full range of estimates we could have obtained under different confounding scenarios, using sensitivity contour plots \citep{Imbens03,cinelli2020making,cinelli2025omitted,chernozhukov2022long}. Since the estimated effect is negative, we focus on the upper confidence bound, which corresponds to the direction of bias relevant for overturning the original conclusion. Figure~\ref{fig:mw_contour_all_main} shows the results, for two different alignment values, a moderate scenario $|\rho_0| = .3$ and an extreme scenario $|\rho_0|=1$. The horizontal axis measures confounding in treatment selection, whereas the vertical axis measures confounding in the untreated outcome trend, both on the same $R^2$ scale. Each contour line corresponds to confounding scenarios that yield the same upper confidence bound for $\theta$, under those hypothetical strengths, and a fixed chosen alignment value. The black triangle in the lower-left corner corresponds to the upper bound of the confidence interval reported in Table~\ref{tab: MW-results-robustness_main}. The red dashed line represents the threshold scenarios that render the effect insignificant. Any confounding scenario below that line is not capable of overturning the original conclusion. The virtue of the contour plot is that the investigator can examine robustness to postulated confounding of \emph{any} hypothetical strength, by simply reading off the value at the corresponding coordinates.
To better aid plausibility judgments, the red diamonds show the point estimates of the scenarios implied by unobserved confounders comparable in strength to observed covariates in explaining treatment selection and outcome evolution, with the alignment fixed at the value of each panel. In addition, the blue dashed line represents the scenario implied by extrapolating the 2006 pre-trend deviation, and it shows what kinds of omitted variables reproduce that amount of bias. (Table~\ref{tab:benchmark_pretrend_bounds} in the Appendix shows the same scenarios accounting for sampling uncertainty.) In accordance with the previous results, in a scenario of moderate alignment none of the benchmark points cross the critical threshold, let alone explain the pre-trend. If, however, we postulate a more conservative scenario, then we cannot rule out that adversarial confounders comparable to \textit{region}, median income (\textit{lmedinc}) or race (\textit{white}) would explain away the estimated effect. In that case, the 2006 pre-trend deviation would be rationalized by regional confounding, but would still not be reproducible by confounders comparable to race, poverty, or income.
\begin{figure}[t]
\centering
\begin{subfigure}[t]{0.49\textwidth}
\centering
\makebox[\linewidth][c]{
\includegraphics[width=1.4\linewidth]{plots/MW_contour_rho2_0.09.pdf}}
\caption{\scriptsize{Upper limit conf. bound, $|\rho_{0}|=0.3$.}}
\end{subfigure}\hfill
\begin{subfigure}[t]{0.49\textwidth}
\centering
\makebox[\linewidth][c]{
\includegraphics[width=1.4\linewidth]{plots/MW_contour_rho2_1.pdf}}
\caption{\scriptsize{Upper limit conf. bound, $|\rho_{0}|=1$.}}
\end{subfigure}
\caption{{\footnotesize Sensitivity contour plots for the minimum wage example at a significance level of $\alpha = 0.05$.
}}
\label{fig:mw_contour_all_main}
\end{figure}
Taken together, these results show that the minimum wage estimate is not completely immune to omitted confounding, but also that overturning the original analysis requires confounding of nontrivial magnitude. More importantly, the analysis clarifies exactly what one needs to believe in order to sustain the estimated effect. For example, unobserved confounders would need to be stronger---or at least more adversarial---than observed covariates to explain away the result. Reproducing the 2006 pre-trend deviation is more demanding still. Thus, a critic arguing that the estimate is not credible must articulate what plausible omitted variables remain that are not only comparable in magnitude to regional confounding, but also sufficiently adversarial. Conversely, someone defending the estimate must argue against unobserved confounders of that magnitude. Note also that sensitivity exercises such as the one above may help adjudicate between competing explanations for the 2006 deviation. If confounders of the required magnitude can be ruled out, anticipation becomes more credible---see Appendix~\ref{app:anticipation} for additional results under this scenario. While researchers may not have much confidence in answering these questions, the sensitivity analysis shifts the discussion from a generic concern about whether parallel trends might fail, to a more disciplined
discussion about the mechanisms required to overturn the conclusion, and the competing hypotheses that could explain pre-trend deviations.
\section{Conclusion}
\label{sec:conclusion}
In this paper we provide an omitted variable bias framework for sensitivity analysis of difference-in-differences designs when unobserved confounding induces violations of parallel trends. We show that the bias in the average treatment effect on the treated admits an exact decomposition into an estimable scale factor and three bias factors, measuring confounding in treatment selection, confounding in untreated outcome evolution, and the alignment between these two channels. We also provide alternative characterizations of the treatment selection component in terms of treatment odds and covariate imbalance. Building on these results, we derive sensitivity statistics for routine reporting, develop benchmarking procedures based on observed covariates and pre-treatment biases, and provide inference methods for the resulting bounds using debiased machine learning.
Possible extensions of our framework include accommodating multiple time periods with staggered treatment adoption. In these cases, the OVB analysis we showed here can be performed period-by-period and aggregated across cohorts and time. Likewise, it is possible to perform richer benchmarking exercises in these settings, including calibration against the full sequence of pre-trends and pre-treatment covariates. Finally, extending the analysis to continuous and multi-valued treatments is also an interesting direction for future work.
\section*{Acknowledgments}
This work was supported in part by the National Science Foundation Grant No. MMS-2417955.
\par
\spacingset{1.4}
\putbib
\end{bibunit}
\newpage