Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
85,112 characters · 10 sections · 66 citation commands
A joint test of unconfoundedness and common trends
\thispagestyle{empty}
Keywords: Common trends, selection on observables, unconfoundedness, overidentification
\setcounter{page}{1}
For panel data where outcomes are observed both before and after treatment, there exist two popular strategies that exploit pre-treatment outcomes to identify the average treatment effect on the treated (ATET) on post-treatment outcomes. The first approach uses pre-treatment outcomes as control variables and relies on the unconfoundedness assumption, which invokes that, conditional on pre-treatment outcomes and possibly additional covariates, the treatment is independent of mean potential outcome under non-treatment. The second approach utilizes the difference between pre- and post-treatment outcomes as dependent variable when assessing the ATET, a so-called Difference-in-Differences (DiD) method that relies on the common trends assumption. This assumption asserts that treatment is independent of the time trend in mean potential outcome under non-treatment, potentially conditional on covariates. Unconfoundedness and common trends are non-nested assumptions, as discussed in ChabeFerret2017 and xu2023causal.
For example, the unconfoundedness assumption allows pre-treatment outcomes to influence treatment but mandates that conditional on the pre-treatment outcome and possibly additional covariates, there are no unobserved confounders affecting treatment and post-treatment outcomes, not even time-constant ones. In contrast, the common trends assumption allows for time-constant unobserved confounders affecting treatment and outcome, as long as their effects are additively separable with regards to period fixed effects, thereby ruling out interaction effects between time-constant confounders and period fixed effects. However, the common trends assumption generally does not permit pre-treatment outcomes to influence treatment if the outcome is affected by time-varying (random) errors. The preference for one assumption over the other is typically not clearly motivated in applied work, as also acknowledged by NBERw31942 and in many empirical contexts, it appears far from clear which of the two assumptions is more plausible, if any.
Given the frequent ambiguity surrounding the plausibility of common trends and unconfoundedness assumptions, this paper proposes a method to jointly test the common trends and unconfoundedness assumptions for ATET identification. The test is based on the difference in estimated average non-treated counterfactuals of treated units under common trends and unconfoundedness assumptions. Estimation relies on doubly robust (DR) statistics, as discussed in Robins+94 and RoRo95, which is combined with machine learning to control for (possibly high-dimensional) covariates in a data-driven manner, following the double machine learning framework (DML) of Chetal2018. As shown in the appendix, our DR statistic satisfies Neyman1959-orthogonality, implying that the test statistic is asymptotically normal and $\sqrt{n}$-consistent under specific regularity conditions for machine learners applied to control for important covariates acting as confounders. We also discuss causal models in which either common trends or unconfoundedness hold, implying the failure of the null hypothesis tested by our procedure, or both assumptions hold jointly, such that the null hypothesis is satisfied. Furthermore, we provide a simulation study to investigate the finite sample properties.
Moreover, we apply the test to four different publicly available datasets, where the treatment is not randomly assigned. First, we consider the non-experimental LaLonde86 data, designed to estimate the effect of a job training program on real earnings, when comparing the original treatment group to a control group sourced from the Panel Study of Income Dynamics. Second, we investigate data from the Job Corps program, which supports disadvantaged youths through various vocational training opportunities. Youths enrolled in the program have the option to participate in educational programs. We apply the test to the effect of educational program participation on weeks employed and general health. Third, we analyze data from a natural experiment in the United States. CardKrueger1994 examined the effect of the minimum wage increase on employment, utilizing survey data from fast food restaurants in New Jersey, where the the policy was implemented, and Pennsylvania, where the minimum wage remained the same. Lastly, we apply the test to data from the United States' National Health and Nutrition Examination Survey to examine robustness of ATET regarding the effect of smoking cessation on weight using a textbook example from Hernan2020. Selection into smoking cessation is likely confounded. Yet, the panel provides a larger number of covariates, to leverage selection on observables. In our empirical findings, the test indicates that the null hypothesis of joint unconfoundedness and common trends does not hold in the Job Corps data. However, we cannot reject the null hypothesis in the remaining examples.
Our test follows the Durbin–Wu–Hausman framework, as outlined in Hausman1978, designed to test two distinct model specifications. Such tests have been utilized for assessing nested model assumptions, where one causal model is more efficient while the other is more robust in the sense that it relies on weaker assumptions. For example, it has been employed to compare the random effects model, assuming exogeneity of treatment with respect to time-constant unobserved confounders, against the fixed effects model, which permits treatment correlation with time-constant confounders, as discussed in Wooldridge02book. Additionally, the test has been used to compare OLS versus linear instrumental variable model assumptions, examining whether coefficients in one model significantly differ from those in the other. Our approach stands out by not imposing functional form assumptions such as linearity, thus accommodating heterogeneous treatment effects across observed or unobserved background characteristics. It also allows for controlling covariates and functional forms in a data-driven manner.
Moreover, our paper contributes to the literature on the bracketing relationship between estimation methods controlling for pre-treatment outcomes and DiD estimation. angrist_mostly_2008 demonstrate that if the true ATET is positive, falsely assuming common trends while unconfoundedness holds will overestimate the ATET, whereas falsely assuming unconfoundedness while common trends hold will underestimate the ATET. This relationship reverses if the true ATET is negative. Ding2019 extend this result to nonparametric and semiparametric settings, such as ours. However, our test does not identify which of the two assumptions is violated; it only indicates if one of them is not met.
Our method is most closely related to the Durbin-Wu-Hausman-type test suggested by DoHsLi2014. The latter utilizes inverse probability weighting (IPW) to assess the unconfoundedness assumption in the presence of a valid instrumental variable (IV) that influences the treatment but is unrelated to the outcome except through the treatment, conditional on covariates. In scenarios featuring one-sided noncompliance, where the treatment is only received when the instrument is present, their test compares the IV-based estimate of the local average treatment effect on the treated (LATT) with the ATET estimate under unconfoundedness, as both causal parameters coincide under such circumstances. While our approach shares similarities with this overidentification strategy, it also exhibits some distinctions. Firstly, we jointly test the unconfoundedness and common trend (rather than IV) assumptions and refrain from assuming that one is a priori more plausible or known to be satisfied. Secondly, instead of employing IPW, we utilize DML techniques to control for possibly high-dimensional covariates within a doubly robust framework tailored for panel data.
The remainder of this study is organized as follows. Section (ref) discusses the common trend and unconfoundedness assumptions required for ATET identification based on DiD or controlling for pre-treatment outcomes, respectively, and introduces our doubly robust test statistic. Section (ref) discusses the empirical implementation of our test based on DML. Section (ref) provides a simulation study. Section (ref) presents empirical applications, and Section (ref) concludes.
We aim at identifying the average treatment effect on the treated (ATET), i.e., the average causal impact of a binary treatment, denoted by $D$, on an outcome variable in a post-treatment period among those receiving the treatment. We will make use of capital letters for denominating random variables ($D$) and lowercase letters to express specific values of these variables ($d=0,1$ for not receiving or receiving the treatment, respectively). To distinguish observed outcomes in terms of pre- and post-treatment periods, we make use of time index $T$. The latter is equal to zero in the pre-treatment period, when no subject receives the treatment (yet), and one in the post-treatment period, after introducing the treatment in the treated group with $D=1$, but not in the nontreated group with $D=0$. We denote by $Y_t$ the observed outcome in period $T=t$, with $t=0,1$ for outcomes observed in the pre-treatment and post-treatment periods, respectively. We assume to have panel data available in which we observe both pre- and post-treatment outcomes of all subjects in those data.
We also assume to observe a set of covariates, denoted as $X$, which are allowed to jointly affect the treatment and outcome, but must not be affected by the treatment. In many treatment evaluations, $X$ includes variables measured prior to treatment assignment to rule out effect of $D$ on $X$. However, in numerous DiD applications, the covariates are measured in the same period as the (pre- or post-treatment) outcome, implying that for $T=1$, $X$ is measured after the introduction of the treatment. In this case, $X$ must not be influenced by $D$, otherwise treatment evaluation is generally biased if $X$ also affects $Y$ and/or is associated with unobservables affecting $Y$. See also the discussion in caetano2022difference on alternative evaluation strategies for time-varying covariates depending on whether the latter are or are not influenced by the treatment. In this paper, we will consider conditioning on pre-treatment covariates only. This implies that for panel data including subject-specfic pre- and post-treatment outcomes and covariates, $X$ only includes pre-treatment covariates measured in period $t=0$.
For discussing the identification of the ATET, we make use of the potential outcomes framework, see for instance Neyman23 and Rubin74. Notably, $Y_t(d)$ denotes the potential outcome that would be hypothetically realized under treatment assignment $D=d$ in period $t=T$, with $d,t$ $\in$ $\{0,1\}$. This permits defining the ATET in post-treatment period $t=1$, which we denote by $\Delta_{D=1}$:
Throughout the discussion, we assume that the `stable unit treatment value assumption' (SUTVA) holds, see Cox58 and Rubin80, imposing that any subject's potential outcomes are not affected by the treatment status of other subjects. Furthermore, we rule out any anticipation effects of the treatment on the pre-treatment outcomes, implying that subjects do not anticipate their treatment in a way that affects their pre-treatment outcomes.
We henceforth discuss two types of assumptions that both yield the identification of $\Delta_{D=1}$. We first consider the unconfoundedness assumption, a popular condition in the treatment evaluation literature as for instance discussed in Im04 and ImWo08. It implies treatment selection based on observables in the sense that the treatment is conditionally mean independent of the potential outcome under non-treatment, $Y_1(0)$, given the covariates $X$ and pre-treatment outcome $Y_0$:
The first line of Assumption (ref) formalizes the unconfoundedness assumption, stating that given $X$ and $Y_0$, no unobservables jointly affect the treatment and the (post-treatment) mean potential outcome under non-treatment. The second line imposes a common support condition, requiring that for any value of $X$ and $Y_0$ among the treated, non-treated subjects with comparable values in $X$ and $Y_0$ must exist.
Under Assumption (ref), $\Delta_{D=1}$ is identified by
because
where the first equality follows from the observational rule ($Y_1(0)=Y_1$ given $D=0$) and the second equality follows from unconfoundedness. Subsequently, we denote the ATET identified by Assumption (ref) as $\Delta_{Unconf}$.
The second type of assumptions imposes conditional common trends, as discussed in Abadie2005 and Lechner2010. It requires that that the treatment is conditionally mean independent of trends in the potential outcome under non-treatment, $Y_1(0)-Y_0(0)$, given the covariates $X$:
The first line in Assumption (ref) formalizes the conditional common trend assumption, stating that given $X$, no unobservables jointly affect the treatment and the trend of mean potential outcomes under non-treatment. This is a selection-on-observables assumption on $D$, however, w.r.t.\ the changes in mean potential outcomes under non-treatment over time, rather than the post-treatment levels of potential outcomes under non-treatment as under the previously discussed unconfoundedness assumption. Subsequently we use the terms conditional common trends and common trends interchangeably, when referring to (ref). The two types of selection-on-observables assumptions are not nested in the sense that neither implies the other. Furthermore, they cannot be fruitfully combined to develop a more general treatment effect model, as shown in ChabeFerret2017. The second line in Assumption (ref) imposes common support, requiring that for any value of $X$ among the treated, non-treated subjects with comparable values in $X$ must exist.
Under Assumption (ref), $\Delta_{D=1}$ is identified by the following difference-in-differences (DiD) approach in panel data:
This is because by the observational rule and the absence of anticipation effects,
while
where the first equality follows from the observational rule and the second equality follows from common trends. We label the ATET identified by Assumption (ref) as $\Delta_{DiD}$.
Figures (ref) to (ref) provide a causal models that help understanding the implications of the unconfoundedness and common trend assumptions based on directed acyclic graphs (DAG), see for instance Pearl00. Solid nodes represent observed variables, dashed nodes unobserved variables, and arrows average causal effects between variables. The subfigures on the left represent causal models in which the post-treatment outcome $Y_1$ is considered as outcome variable, and subfigures on the right correspond to causal models in which the difference $Y_1-Y_0$ is considered as outcome, like in DiD approaches. Considering the left graph of Figure (ref), we see that treatment $D$ affects the $Y_1$, which is the causal effect of interest, and is itself affected by the pre-treatment outcome $Y_0$. In addition, we have two time varying unobservables $V_0$ and $V_1$ that affect the outcomes $Y_0$ and $Y_1$, respectively, and the pre-treatment unobservable $V_0$ may also affect the post-treatment unobservable $V_1$. Furthermore, we assume an independent unobservable $Q$ that affects the treatment $D$. Finally, the unobserved time constant variable $U$ jointly affects the pre-treatment outcome $Y_0$, the treatment $D$, and the post-treatment outcome $Y_1$. In the right graph, the outcome change over time, $Y_1-Y_0$, is affected by the change in time varying unobservables, $V_1-V_0$, rather than $V_1$. $V_1-V_0$ is affected by $V_0$, as the difference between two variables in general may depend on both variables, which motivates the causal arrow from $V_0$ to $V_1-V_0$. Only under the specific serial dependence $E[V_1|V_0]=V_0$ would $V_0$ be uncorrelated with $V_1-V_0$, because $E[V_1-V_0|V_0]=E[V_1|V_0]-V_0=V_0-V_0=0$ would hold in this special case of a random walk. We note that we could additionally include observed covariates $X$ in the causal graphs that might jointly affect the treatment and the outcomes, which is omitted for the sake of simplicity. When one controls for such covariates $X$, the graphs can be interpreted as causal associations conditional on $X$.
In the causal framework given in Figure (ref), neither unconfoundedness, nor common trends hold. Unconfoundedness fails because conditional on $Y_0$ and possibly $X$, $U$ affects both $Y_1$ and $D$ in the left graph. Common trends fail for two reasons. First, the time constant unobservable $U$ affects both $D$ and the outcome change $Y_1-Y_0$ (implying heterogenous time trends in $U$). Second, the time varying unobservable $V_0$ affects both $D$ (via $Y_0$) and $Y_1-Y_0$ (via $V_1-V_0$), unless $E[V_1|V_0]=V_0$ holds, as also discussed in ghanem2022selection. Both assumptions impose different, non-nested conditions, see also the discussion in ChabeFerret2017. Unconfoundedness invokes that conditional on $Y_0$ (and $X$), no unobserved time constant characteristics like $U$ or time varying characteristics like $V_0$ jointly affect $D$ and $Y_1$. This is satisfied in the causal model presented in Figure (ref), where $U$ only affects the outcomes $Y_0$ and $Y_1$, but not directly $D$, such that $D$ is not affected by $U$ conditional on $Y_0$. Likewise, $V_0$, which affects $Y_1$ via $V_1$, does not affect $D$ conditional on $Y_0$. However, the common trend assumption fails because of $U$ jointly affecting $D$ (via $Y_0$) and $Y_1-Y_0$, as well as $V_0$ jointly affecting $D$ (via $Y_0$) and $Y_1-Y_0$ (via $V_1-V_0$). In the causal model in Figure (ref), it is the common trend assumption that holds. Here, taking the outcome difference $Y_1-Y_0$ removes the impact of $U$ if the average effect of $U$ on the outcome does not interact with (or is homogeneous in) the time trend, which holds if $U$ is additively separable in the outcome model. Furthermore, the pre-treatment outcome $Y_0$ does not affect $D$ such that $V_0$ does not jointly affect $Y_1-Y_0$ and $D$. In contrast, unconfoundedness fails as $U$ directly affects $D$ and $Y_1$. Finally, Figure (ref) provides a causal model in which both the unconfoundedness and common trends assumptions hold, such that identification based on Assumption (ref) and (ref) both yield the true ATET. Unconfoundedness holds because $U$, which influences $Y_1$, and $V_0$, which also influences $Y_1$ via $V_1$ both affect $D$ only via $Y_0$. Common trends hold because $U$ does not affect the outcome change $Y_1-Y_0$, while $V_0$ is not associated with $Y_1-Y_0$ via $V_1-V_0$, implying a specific serial correlation of $V_1$ and $V_0$, $E[V_1|V_0]=V_0$.
Comparing expressions (ref) and (ref) demonstrates that the identification results under common trends and unconfoundedness differ in terms of how the mean potential outcome under non-treatment, $E[Y(0)|D=1]$, is computed. We define the parameter $\theta$ as the difference between expressions (ref) and (ref), which is the target parameter of our testing approach:
Expression (ref) makes use of conditional mean outcomes for defining $\theta$. Alternatively, a numerically equivalent expression for $\theta$ is the following, so-called doubly robust (DR) statistic, which is based on both, conditional mean outcomes and conditional treatment probabilities:
with $\mu(x,y_0)=E[Y|X=x,Y_0=y_0,D=0]$, $m(x)=E[Y_1-Y_0|X=x,D=0]$ denoting the conditional mean outcomes or differences in outcomes under non-treatment, respectively, and $p(x,y_0)=\Pr(D=1|X=x, Y_0=y_0)$, $\pi(x)=\Pr(D=1|X=x)$ denoting the conditional treatment probabilities, also known as propensity scores. \footnote{By the law of iterated expectations and basic probability theory,
In an analogous manner,
For this reason, expressions (ref) and (ref) are equivalent.} Expression (ref) coincides with the difference in DR identification of the ATET based on DiD in panel data, see e.g.\ equation (2.6) of SantAnnaZhao2018, and DR identification of the ATET based on unconfoundedness when controlling for covariates $X$ and pre-treatment outcome $Y_0$, see e.g. equation (4.45) in huber2023causal. This test-statistic is our main result.
DR approaches bear the attractive property that estimation is consistent (under standard regularity conditions) if either the models for the conditional mean outcomes or the propensity scores, henceforth denoted as nuisance parameters are correctly specified, see e.g.\ RobinsMarkNewey1992, Robins+94, RoRoZa95, and BaRo05. This implies that estimation of $\theta$ based on expression (ref) is consistent if either $\mu(X,Y_0)$ or $p(X,Y_0)$ when imposing unconfoundedness, and either $m_T(X,0)$ or $\pi(X)$ when imposing common trends is correctly specified. Furthermore, DR estimators are quite robust to moderate misspecifications in all of the nuisance parameters, a property known as Neyman1959-orthogonality. The latter implies that the estimation of $\theta$ based on expression (ref) is first order insensitive to approximation errors in the estimation of the propensity scores and conditional mean outcomes. This appears attractive in big data contexts with high-dimensional covariates $X$, as it permits estimating the nuisance parameters by machine learning, which generally introduces regularization bias, but nevertheless permits $\sqrt{n}$-consistent estimation of $\theta$ under specific regularity conditions outlined in the next section.
This section proposes an estimator of $\theta$ as given in the DR expression (ref), based on double machine learning (DML) as suggested in Chetal2018. Let $\mathcal{W} = \{W_i|1\leq i \leq n\}$ with $W_i = (Y_{1i}, Y_{0i}, X_{i}, D_{i})$ for all $ i $ denote the set of observations in an i.i.d.\ sample of size $n$. We define $\eta=\{p(X, Y_0), \pi(X), \mu (X,Y_0,D), m(X,D))\}$ as the vector of nuisance parameter functions, i.e.\ the propensity scores and the conditional means of the outcomes or outcome differences, respectively. Their respective estimates are denoted as $\hat{\eta} =\{\hat p(X, Y_0), \hat\pi(X), \hat\mu (X,Y_0,D), \hat m(X,D))\}$ and the respective true functions by $\eta_0=\{p_0(X, Y_0), \pi_0(X), \mu_0 (X,Y_0,D), m_0(X,D))\}$.
We estimate $\theta$ by the following algorithm that combines the estimation of Neyman-orthogonal scores with sample splitting or cross-fitting and is $\sqrt{n}$-consistent under conditions outlined further below. \newline Algorithm 1: Estimation of $\theta$ based on expression (ref)
To see that the normalized estimator in step 4 corresponds in expectation to expression (ref) under consistent estimation of the nuisance parameters, we note that by the law of iterated expectations, $E\left[\sum_{i=1}^n D_i\right]=E\left[\sum_{i=1}^n \frac{(1-D_i)\cdot \hat p(X_i,Y_{0i})}{1-\hat p(X_i,Y_{0i})}\right]=E\left[\sum_{i=1}^n \frac{(1-D_i)\cdot \hat \pi(X_i)}{1-\hat \pi(X_i)}\right]=n \cdot \Pr(D=1)$.
In order to obtain root-n consistency for the estimation of $\theta$, we impose the following assumption consisting of regularity conditions w.r.t.\ the prediction quality of machine learning for estimating the nuisance parameters. Following Chetal2018, we introduce some further notation: let $(\delta_n)_{n=1}^{\infty}$ and $(\Delta_n)_{n=1}^{\infty}$ denote sequences of positive constants with $\lim_{n\rightarrow \infty} \delta_n = 0 $ and $\lim_{n\rightarrow \infty} \Delta_n = 0.$ Furthermore, let $c, \epsilon, C$ and $q$ be positive constants such that $q>2,$ and let $K \geq 2$ be a fixed integer. Also, for any random vector $R = (R_1,...,R_l)$, let $\left\| R \right\|_{q} = \max_{1\leq j \leq l}\left\| R_l \right\|_{q},$ where $\left \| R_l \right\|_{q} = \left( E\left[ \left| R_l \right|^q \right] \right)^{\frac{1}{q}}$. In order to ease notation, we assume that $n/K$ is an integer. For the sake of brevity we omit the dependence of probability $\Pr_P,$ expectation $E_P(\cdot),$ and norm $\left\| \cdot \right\|_{P,q}$ on the probability measure $P.$
Condition (a) in Assumption (ref) states that the distribution of outcomes or differences in outcomes does not have unbounded moments. Condition (b) refines the common support condition such that the treatment propensity scores are bounded away from $0$ and $1$. Condition (c) requires that the covariates $X$ do not perfectly predict the conditional means of the outcomes or differences in outcomes, respectively. Condition (d) is the only non-primitive condition and puts restrictions on the quality of the nuisance parameter estimators. We note that for the satisfaction of the last two restrictions of condition (d), $\left\| \hat \mu(X,Y_0,D)-\mu_0(X,Y_0,D)\right\|_{2} \times \left\| \hat p(X, Y_0)-p_0(X, Y_0)\right\|_{2} \leq \delta^{}_n n^{-1/2}$ and $\left\| \hat m(X,D)-m_0(X,D)\right\|_{2} \times \left\| \hat \pi(X)-\pi_{0}(X)\right\|_{2} \leq \delta^{}_n n^{-1/2}$, it suffices that the nuisance estimators converge with rate $ o(n^{-1/4})$, which is attainable by many commonly used machine learners under specific conditions (like approximate sparsity), such as lasso, random forests, boosting and neural nets, see for instance Bellonietal2014, huber2023testing, WagerAthey2018, and FarrellLiangMisra2018.
Theorem (ref) formally states the asymptotic normality and root-n consistency of the estimator $\hat \theta$. The proof is based on demonstrating that our estimation approach satisfies the conditions of DML estimation given in Chetal2018, namely their Assumption 3.1 on the linearity and Neyman-orthogonality of the DR expression underlying the identification of $\theta$ and Assumption 3.2 on regularity conditions and the quality of nuisance parameter estimators, related to Assumption (ref).
As DR expression (ref) corresponds to the difference in DR expressions for the ATET under Assumptions (ref) (common trends) and (ref) (unconfoundedness), respectively, Theorem (ref) follows from combining Theorem 3.1 in Chang2020 and Theorem 5.1 in Chetal2018 on ATET estimation by cross-fitting under common trends and unconfoundedness, respectively. Appendix (ref) provides a proof of the linearity and Neyman orthogonality of our DR expression, i.e., Assumption 3.1 in Chetal2018, and refers to Chang2020 and Chetal2018 for proofs of Assumption 3.2 in Chetal2018.
This section presents a simulation study designed to examine the finite sample properties of our proposed testing approach as outlined in Theorem (ref). The underlying data generating process (DGP) models pre-treatment and post-treatment outcomes $Y_0$ and $Y_1$, along with the treatment assignment $D$, as follows:
In this setup, the pre-treatment outcome $Y_0$ is a linear function of the fixed effect $U$ and the period-specific unobservable $V_0$. The binary treatment $D$ is a function of covariates $X$ for the coefficient vector $\beta_X \neq 0$, the fixed effect $U$ if $\gamma\neq 0$, the pre-treatment outcome $Y_0$ if $\delta\neq 0$, and further unobservables $Q$. The post-treatment outcome $Y_1$ is a linear function of treatment $D$, whose effect is 1, of the fixed effect $U$, covariates $X$ if $\beta_X \neq 0$, and the period-specific unobservable $V_1$. We note that for $\gamma \neq 0$, the unconfoundedness assumption is violated, due to the confounding of $D$ and $Y_1$ by $U$. For $\delta \neq 0$, the common trend assumption is violated. To see this, note that if $Y_0$, which is a function of the period-specific error $V_0$, affects $D$, then $D$ also becomes a function of $V_0$. Furthermore, the difference in outcomes over time, $Y_1 - Y_0$, is a function of the difference in the period-specific errors, $V_1 - V_0$, and thus, of $V_0$. For this reason, $V_0$ confounds $D$ and $Y_1 - Y_0$ if $Y_0$ affects $D$, such that common trends fail in our framework with independently drawn errors $V_1$ and $V_0$. \footnote{Only in the special case of $E(V_1 - V_0) = 0$ would common trends hold, but this would imply a strong serial correlation between $V_1$ and $V_0$, characterized by a random walk.}
The covariate vector $X$ follows a multivariate normal distribution with a zero mean and a covariance matrix $\sigma^2_X$, where the covariance between the $i$th and $j$th covariate is set to $0.5^{|i-j|}$. The variables $V_0$, $V_1$, $U$, and $Q$ each follow standard normal distributions and are independent of each other, as well as of $X$. The vector $\beta_X$ gauges the effects of the covariates on the treatment assignment $D$, on the post-treatment outcome $Y_1$, and the outcome trend $Y_1-Y_0$ alike, which is true, because $X$ does not influence $Y_0$. Thus, $\beta_X$ precisely captures confounding due to observables, yet it does not capture potential unobserved effects introduced by $U$ and $Q$. The $i$th element in the coefficient vector $\beta_X$ is set to $0.6/i$ for $i=1, \dots, p$. This implies a linear decay of covariate importance in terms of confounding.
We define three distinct scenarios to evaluate the performance of our approach, which mirror the causal models depicted in Figures (ref)-(ref):
To assess convergence rates, we vary the sample sizes, using either $n = 1000$ or $n = 4000$ and keep the remaining parameters constant. We maintain 1000 iterations and $p = 100$ covariates. The test statistic $\hat{\theta}$ is computed using Algorithm (ref). Under the null hypothesis both selection on observables assumptions hold. Thus, a significant deviation of $\hat{\theta}$ from zero indicates a violation of at least one assumption. We reject the null if the p-value is smaller than $0.05$.
We employ various models for training the nuisance parameters. Our parametric approach utilizes ordinary least squares (OLS) for estimating outcomes $\hat{\mu}$ and $\hat{m}$, and logistic regression for computing the propensity scores $\hat{\pi}$ and $\hat{p}$. Additionally, we explore four ML models: Lasso, Random Forest (RF), Support Vector Machine (SVM), and the Super Learner, which determines the optimal weighted average across an ensemble of Lasso, RF, and SVM. This approach facilitates the selection of the best model for the nuisance parameters without the need for manual performance comparisons and selection of the best model VanderLaanHubbard2007. To prevent extreme weights, we trim on the propensity score, and exclude observations where $\hat{\pi}(X_i)$ and $\hat{p}(X_i, Y_{0i})$ are greater than or equal to 0.99. The nuisance parameters are estimated through cross-fitting to ensure honest estimates, following the methodology proposed by AtheyImbens2016.
We first analyze the finite sample distribution of $\hat{\theta}$. Figure (ref) presents a histogram with the distribution of $\hat{\theta}$ for the three distinct simulation scenarios determined by the parameters $\delta$ and $\gamma$. These scenarios are visualized with $n = 1000$ observations in blue and $n = 4000$ observations in gray. Nuisance parameters for this analysis are estimated by the ensemble learner. The Figure effectively demonstrates the asymptotic properties of $\hat{\theta}$, showing that the estimator converges to a normal distribution, with decreasing varaince as N increases. Specifically, Panel (a) indicates that $\hat{\theta}$ aligns with an expectation of zero when both unconfoundedness and conditional common trends are met. Conversely, Panels (b) and (c) reveal that the expectation of $\hat{\theta}$ deviates from zero, when either assumption is compromised.
Next, we conduct an in-depth analysis of the test's performance. Table (ref) provides a comprehensive summary of the simulation results across the three specifications. It details the most important statistics of the test, under `Test results', and the two doubly robust approaches to estimate ATET, `Unconf' and `DiD'. The Table reports the test-statistic $\hat{\theta}$ along with the standard deviation, standard errors, and p-values, which equal the mean values across iterations within the same specification, and the corresponding rejection rates. Instead of reporting the point estimates for $\Delta_{Unconf}$ and $\Delta_{DiD}$, we present the bias, which can be interpreted as percentage deviation from $\Delta_{D=1} = 1$. Similar to the test results, we report the bias, RMSE, and standard errors as mean values across all iterations within the same parameter specification. Additionally, we include the empirical standard deviation as a robustness measure for the standard errors, which is expected to stabilize as the number of iterations increases.
In the first scenario, unconfoundedness and common trends are satisfied. The results are reported in the upper panel of Table (ref). The average of the DR-based difference $\hat{\theta}$, the test-statistic, is close to zero for either sample sizes. We find that the estimator seems to converge at $\sqrt{n}$-rate, as the standard deviation roughly reduces by half when quadrupling the sample size from 1000 to 4000. Moreover, the standard deviation is similar to the standard error, at least for the larger sample size. The average p-values are all larger than $0.3$ and the test rejects the null in 5-15$\%$ of the cases if $n = 1000$ and and 6-12$\%$ if $n = 4000$, respectively. This implies a decent empirical size for $n = 4000$. The bias of $\Delta_{Unconf}$ and $\Delta_{\text{DiD}}$ vanishes with increasing sample size and the empirical standard deviation resembles the mean standard errors.
In the second scenario, we set $\gamma = 0.25$ such that treatment assignment $D$ and post-treatment outcome $Y_1$ are confounded by the fixed effect $U$. The results are reported in the middle panel of Table (ref). Here, $\hat{\theta}$ is moderately different from zero for both sample sizes. The average estimate for $\hat{\theta}$ now amounts to roughly $0.2$ to $0.3$, across both sample sizes. In the small sample case, the rejection rate varies from 33-65$\%$. However, it increases to 92-100$\%$ in the large n case. Thus, in this specification the precision gain with larger sample sizes is more pronounced, due to the smaller standard errors. For this reason, also the average p-values are close to zero, if $n=4000$. The bias in $\Delta_{Unconf}$ remains persistent, while the bias in $\Delta_{DiD}$ is smaller with increasing sample size.
In the third scenario, where $\delta = 0.25$, the treatment assignment $D$ becomes a function of the period specific error term $V_0$, mediated by the pre-treatment outcome $Y_0$. The test robustly detects violations of the common trends assumption, even in smaller samples. For the small sample case $\hat{\theta}$ varies between $0.38$ to $0.60$ and does not change much if n increases. Thus, the precision gain from the smaller standard errors is less pronounced. The test correctly rejects the null hypothesis in 55-100$\%$ of the cases using the smaller sample, and almost 100$\%$ of the cases in the larger sample. Here, the bias in $\Delta_{Unconf}$ diminishes, while the bias in $\Delta_{DiD}$ remains present with increasing sample size.
It is crucial to recognize that the accuracy of nuisance parameters significantly influences the statistical power of the test. Specifically, either $\hat{p}(X, Y_0)$ or $\hat{\mu}(X, Y_0)$ and $\hat{\pi}(X)$ or $\hat{m}(X)$ must be accurately specified to yield a precise estimate of $\hat{\theta}$. While the test statistic is doubly robust with respect to the outcome and propensity score estimates, achieving true robustness requires that at least one set of nuisance parameters be approximately correctly specified for both $Unconf$ and $DiD$ scenarios. Applying the test to simulated data corroborates the previously derived asymptotic sample properties and underscores the critical importance of accurately estimated nuisance parameters for achieving a robust performance of the test. This foundation is crucial when extending the application of the test to real-world data.
This section presents an empirical examination of the test using publicly accessible datasets previously employed to estimate the average treatment effect on the treated (ATET). The test rejects the null hypothesis in two of the five cases. Emphasis is placed on scenarios with non-random assignment mechanisms into treatment and control group to highlight the practical applicability of the test.
In our first example, we apply the test to one of the most prominently used data from the field of labor economics, the LaLonde86 data. Within the framework of the National Supported Work Demonstration (NSW), LaLonde conducted a field experiment to evaluate the impact of a job training program on the real earnings of participants during the mid-1970s in the United States.
In the experiment, applicants that were randomly assigned to the treatment group were guaranteed employment during the program. In contrast, applicants that were randomly assigned to the control group did not receive any form of support from the NSW. The job training initiative started in January 1978, with male participants completing the program after nine and female participants after twelve months. Initial interviews with participants were conducted in 1974, and their real earnings were documented for the years 1974, 1975, and 1978. The treatment is participation in the job program and the outcome of interest are participant's real earnings in the year 1978.
To assess the accuracy with which results from experimental data could be replicated using observational data LaLonde replaced the original control group with observations from the Panel Study of Income Dynamics (PSID). The comparison between estimates based on the experimental LaLonde data and the nonexperimental control group from the PSID revealed significant discrepancies, leading to the formulation of what is now widely acknowledged as the LaLonde critique.
Using the experimental LaLonde data as a benchmark and PSID-data to assess the performance of causal estimates, have been used in various settings, including propensity score matching (DehejiaWahba99 and DeWa02), DiD-matching (HeIcTo98 and SmithTodd00), and finally, the doubly robust score function for $\Delta_{DiD}$ which enters the test-statistic described in SantAnnaZhao2018. Most recently, imbens2024lalonde use the LaLonde data to reassess the LaLonde critique focusing on the performance of doubly robust methods. They stress the importance of validation exercises, such as the evaluation of common support, to obtain credible estimates.
We perform the test using the experimental treatment group and the non-experimental PSID control group, which only covers male participants. Under the null hypothesis unconfoundedness and conditional common trends jointly hold. Our analysis considers real earnings in 1975 as the pre-treatment outcome. Included control variables are employment status in 1974 and 1975, alongside demographic factors such as as age, education, marital status, and race. We include real earnings in 1974 in the set of controls for estimating $\hat{p}(X, Y_0)$ and $\hat{\mu}(X, Y_0)$, where the previous earnings are likely an important predictor to account for dynamic selection bias. However, we exclude real earnings in 1974 for the estimation of $\hat{\pi}(X)$ and $\hat{m}(X)$ to avoid M-bias ChabeFerret2017\footnote{The DiD estimator assumes that selection bias remains constant over time. By including pre-treatment outcomes, we attempt to account for time-varying selection bias within the trend itself. However, this approach can lead to spurious correlation between the pre-treatment outcome and the treatment effect, as they are influenced by the same confounding factors that affect the treatment effect and result in bias. This correlation would not exist had we not opened the backdoor pathway in the first place.}. To estimate the nuisance parameters we use four different learners: Lasso, RF, and the parametric specification in combination with three-fold cross-validation. The sample consists of \(n = 2650\) observations out of which 185 are treated\footnote{We use the lalonde.psid dataset accessible via the causalsens R package \url{https://CRAN.R-project.org/package=causalsens}.}. We employ different learners for the nuisance parameters, which are reported in the table and the plots\footnote{Across all empirical examples we use Lasso, Random Forest, Support Vector Machine, an ensemble of the latter, and a parametric specification. If SVM does not converge we exclude it from the ensemble learner and as a single learner to obtain the nuisance parameter estimates.}. We find satisfactory overlap in the distribution of the propensity scores. The appendix (ref) provides plots for all propensity score estimates of our empirical examples.
Table (ref) reports the test results, the point estimates for $\hat{\theta}$, the standard errors, and the p-values; and the point estimates and standard errors for $\Delta_{Unconf}$ and $\Delta_{DiD}$ using the LaLonde-PSID sample. Across the chosen specifications, the test statistic ranges from 65 to 600 and the standard errors from 330 to 1400. As a consequence, we cannot reject the null hypothesis that unconfoundedness and conditional common trends jointly hold at the 5$\%$ significance level. However, the ensemble learner yields a comparably small p-value of $0.072$. The point estimates suggest that participation in the job training program leads to an approximate increase by $1,500\$ $ to $3,000\$ $ in annual real wages, one year after participation. Although, there is variation across the different learners, the estimates for $\Delta_{Unconf}$ and $\Delta_{DiD}$ lie in a similar range, and exhibit largely overlapping confidence intervals. A visual representation can be found in Figure (ref), where the point estimator and corresponding 95-$\%$ confidence bands for $\Delta_{Unconf}$ are color-coded in black, and the point estimator and corresponding 95-$\%$ confidence bands for $\Delta_{DiD}$ are color-coded in red. The p-values denote the significance-level that $\hat{\theta}$ is different from zero. To sum up, the distinct learners yield estimates for $\hat{\Delta}_{D = 1}$ in different ranges. Yet, the joint test does not reject in this example.
In our second example we use data from the United States largest education and training program for disadvantaged youths, the Job Corps program (JC). This program provides vocational training of different forms, such as academic education, health education, training in independent living, and assistance to find employment after the training was completed. JC was initiated in the 1960s and receives the largest share of funds from the US Department of Labor that targets training of young adults. To evaluate the impact of the program, the National Job Corps Study conducted the first nationally representative experiment in the beginning of the 1990s.
During the enrollment phase between November 1994 to December 1995, all youths that applied to JC were randomly assigned to treatment or control group. Youths in the treatment group were able to enroll in the training program. More precisely, the final group of participants consist of compliers within the randomly chosen treatment group. ScBuGl01 provide a detailed description of the experiment. Applications of the data can be found in FlFl09 who investigate the effect of the program on earnings, a discussion about the programs success FlFl10, a decomposition of the ATE on general health using mediation analysis via inverse probability weighting Huber2012, a discussion about the long-run effect of the program ScBuMc2008.
We are interested in the effect of the program on the employment rates, measured as proportion of weeks employed, after one year. The sample consists of \(n = 9240\) observations out of which 6574 are enrolled in an education program in the first year of JC\footnote{We use the JC dataset accessible via the causalweight R package \url{https://CRAN.R-project.org/package=causalweight}.}. The pre-treatment outcome is proportion of weeks employed before the program started, and we control for gender, age, ethnicity, education, GED degree, highschool degree, mother's and father's education, English as mother tongue, marital or cohabiting status, if participant has at least one child, and ever worked at the time of assignment, average weekly gross earnings, household size, welfare receipt during childhood, general health, extent of smoking and alcohol consumption.
The results are summarized in Table (ref). The test statistic $\hat{\theta}$ ranges between $3.2$ to $4.1$ with corresponding standard errors below $0.78$. Consequently, the test rejects across all implemented specifications at a lower than $0.001$ level. The variation in the point estimates for $\hat{\Delta}_{Unconf}$ and $\hat{\Delta}_{DiD}$ is demonstrated explicitly in Figure (ref). Thus, the unconfoundedness and conditional common trends assumption do not jointly hold in this application.
As the JC program had multiple objectives, we employ a second variable, health outcomes, as the dependent variable employing the same vector of control variables. Similar to proportion of weeks employed, the test rejects the null hypothesis, as $\hat{\theta}$ is significantly different from zero. The detailed results are summarized in Table (ref) and Figure (ref).
This section investigates a natural experiment that examines the impact of minimum wage increases on employment. In April 1992, the state of New Jersey raised its minimum wage from $\$4.25$ to $\$5.05$. CardKrueger1994 used this policy change to investigate the hypothesis that an increase in the minimum wage leads to decreased employment. Their seminal study focused on the fast food industry, which is the largest employer of low-wage workers. Utilizing fast food restaurants from the neighboring state of Pennsylvania, the authors established control observations that were not subject to the policy shift. In total, they conducted two survey waves involving 410 establishments: the initial one in early spring 1992, just before the wage increase took effect, and the second in late autumn of the same year.
We analyze a subset of the data comprising $n = 334$ observations with complete records for the employment of full-time and part-time workers, the number of managers, the percentage of employees affected by the wage increase, and the initial wage upon starting the new job\footnote{The data and replication code from the original paper are available under \url{https://davidcard.berkeley.edu/data_sets.html}.}. Following the methodology of CardKrueger1994, we assign full-time employment as dependent variable. The treated units are precisely identified as those fast-food restaurants in New Jersey that were mandated to increase wages due to the policy change. We incorporate a total of 27 control variables, accounting for potential confounding factors such as the composition of workers employed, product pricing, operating hours, co-ownership, and the type of fast-food chain.
The test results are displayed in Table (ref) and indicate that we cannot reject the null hypothesis as $\hat{\theta}$ is close to zero across all specifications. Consistent with the findings of the original study, the two doubly robust point estimates for $\Delta_{D=1}$ suggest that the increase in the minimum wage does not significantly affect full-time employment. The point estimates along with their confidence intervals, and the p-values of the test are displayed in Figure (ref).
Next, we assess whether parallel trends and unconfoundedness hold when estimating the effect of smoking cessation on weight, following the analysis from a textbook example by Hernan2020. The authors used data from the National Health and Nutrition Examination Survey I Epidemiologic Follow-up Study (NHEFS), initiated by the National Center for Health Statistics, National Institute on Aging, and United States Public Health Service. This survey collected various health measures over time, such as pulse rate, weight, and blood pressure, and documented hospital and nursing home stays, as well as alcohol and cigarette consumption, resulting in a rich panel dataset.
We use a subset of the NHEFS data to compare the ATET of smoking cessation on weight, measured in kilograms. Our final sample includes \(n = 1566\) participants who were smokers in the first wave of the survey in the early 1970s, with complete individual-level characteristics. The treatment is defined as smoking cessation between the initial survey and the follow-up visit ten years later. The sample includes 403 quitters. We control for 47 variables reflecting socioeconomic characteristics, health measures, and cigarette expenses. The data is publicly available\footnote{We use a subsample from the causaldata R package at \url{https://CRAN.R-project.org/package=causaldata}. The full sample can be accessed at \url{https://wwwn.cdc.gov/nchs/nhanes/nhefs/}.}.
The results are summarized Table (ref) and indicate that we cannot reject the null hypothesis at any sensible level, as $\hat{\theta}$ is close to zero. A visual representation can be found in Figure (ref). This observational study provides an example, where selection bias is likely present, as quitters probably self-select into the treatment group. However, the high-dimensional controls seem to account for confounding effects.
The applications in this section demonstrate that the test can serve as a robustness check to causal estimates using panel data, when using an ex-post matched control group, the treatment group consists of compliers, the data derives from a natural experiment, or has no degree of randomness to it.
We introduce a Durbin–Wu–Hausman type test leveraging selection on observables to assess the joint validity of unconfoundedness and conditional common trends in panel data. Our test statistic is derived from the difference between counterfactual estimation when applying doubly robust estimation to control for pre-treatment outcomes and covariates, and a doubly robust DiD approach controlling for covariates. The former assumes unconfoundedness, while the latter assumes conditional common trends. A significant deviation of the test statistic from zero indicates a violation of at least one assumption.
We establish Neyman-orthogonality of our test statistic, which implies asymptotic normality and $\sqrt{n}$-consistency of estimation under specific regularity conditions. The latter permit controlling for possibly high dimensional covariates in a data-adaptive manner based on a double machine learning (DML) approach, as we suggest in the paper. We also investigate the finite sample performance of the DML-based testing procedure in a simulation study and find it to perform satisfactorily across the considered designs. Additionally, we apply our test to five empirical examples in which the treatment is not randomly assigned and find the test to reject the null hypothesis in two examples.
{ \setcounter{equation}{0}