Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
57,979 characters · 13 sections · 48 citation commands
On a Debiased and Semiparametric Efficient Changes-in-Changes Estimator
Do mass shootings influence voter behavior in the United States? Recent years have seen a tragic rise in mass shootings, prompting renewed interest in how such salient events shape public opinion and political outcomes peterson2024epidemiology. While some studies suggest that mass shootings lead to decreased Republican vote share in affected areas yousaf2021sticking, others argue these findings may reflect methodological artifacts, such as violations of key identification assumptions hassell2025navigating. At stake is not only a question of electoral dynamics but also our understanding of how collective trauma and public safety crises influence democratic accountability.
Although the electoral effects of mass shootings offer a timely and consequential case study, they reflect a more general challenge in causal inference: how to credibly estimate treatment effects from panel data when treatment is non-randomly assigned, and unmeasured confounding may be present. These problems arise widely in economics, political science, public health, and policy evaluation streeter2017adjusting, tchetgen2020introduction, chiu2023causal,arkhangelsky2024causal, where researchers seek to understand the effects of interventions, shocks, or exposures in settings where randomized experiments are infeasible. Conventional methods like difference-in-differences (DiD) require strong assumptions—such as parallel trends—that may not hold in practice, particularly when treatment assignment or outcome evolution is driven by latent, possibly complex and high-dimensional factors.
To address these challenges, this paper develops a general framework for robust causal inference in panel data settings with unmeasured confounding and flexible covariate structures. We build on and extend the influential Changes-in-Changes (CiC) framework of athey2006identification. However, existing CiC methods typically assume a single scalar unmeasured confounder and a monotonic relationship with the outcome. These assumptions may be overly restrictive in settings characterized by complex, nonlinear confounding structures.
Our first contribution is to establish new nonparametric identification results under a relaxed set of assumptions that accommodate high-dimensional, non-monotonic unmeasured confounders. These results enable the identification of the average treatment effect on the treated (ATT), as well as other distributional causal estimands such as quantile treatment effects, without relying on linear trend or monotonicity assumptions.
Our second contribution, building on this foundation, is to develop a semiparametric estimation strategy that achieves both efficiency and robustness in high-dimensional settings. We derive the efficient influence functions (EIFs) for the ATT and for general causal estimands defined through moment conditions, and use them to construct estimators that are Neyman orthogonal to infinite-dimensional nuisance components. This orthogonality ensures that errors in estimating nuisance functions, even when using flexible machine learning methods, do not have a first-order impact on the estimation of the target parameter. To further reduce overfitting bias, we incorporate sample splitting and cross-fitting.
We illustrate the practical utility of our approach through illustrative simulation studies and an empirical application investigating the electoral consequences of mass shootings. The proposed estimator consistently demonstrates strong performance in terms of coverage, bias, and efficiency. While this application serves as a motivating example, the proposed framework is broadly applicable to a wide range of empirical problems in observational panel data settings.
Our work contributes to a broad and rapidly evolving literature on causal estimation in panel data, particularly within the difference-in-differences (DiD) framework. The classical DiD approach relies on the “parallel trends” assumption, which states that in the absence of treatment, treated and control units would have followed parallel outcome trajectories angrist2009mostly, baker2025difference. Extensions incorporating conditioning on observed covariates relax this assumption somewhat, allowing for heterogeneous trends heckman1997matching, abadie2005semiparametric, sant2020doubly, callaway2021difference. Still, the plausibility of parallel trends is often questionable, especially when outcomes are bounded or display nonlinear dynamics roth2023parallel.
To address these issues, researchers have proposed various alternative strategies:
A complementary line of research is the Nonlinear DiD approach, which impose structure on transformed outcomes or recast identification assumptions on alternative scales bonhomme2011recovering, puhani2012treatment, callaway2019quantile, park2022universal, wooldridge2023simple, roth2023parallel, fernandez2024distribution. The present work contributes to this growing literature by extending the Changes-in-Changes (CiC) framework.
The CiC method introduced by athey2006identification was a landmark in allowing nonparametric identification of ATT and distributional causal effects without relying on parallel trends. Its identification strategy is built on a “production function” model and is invariant to monotonic transformations of the outcome. However, standard CiC implementations typically assume a scalar unmeasured confounder and a monotonic relationship between the confounder and the outcome. Our work builds on this foundation by relaxing these assumptions. We show that identification is still possible under a more general confounding structure, thereby broadening the scope of CiC-style methods in practice.
Additionally, covariates are often crucial for reducing bias from observed confounding. athey2006identification outlined a restrictive semiparametric extension that incorporated covariates linearly, without allowing treatment effects to vary across covariate levels. Later approaches melly2015changes, thome2025estimating introduced more flexible semiparametric adjustments but continued to rely on linear functional form assumptions, leaving room for model misspecification. By contrast, our framework is fully nonparametric with respect to the observed data, enabling flexible covariate adjustment and allowing treatment effects to vary flexibly, while attaining the semiparametric efficiency bound.
Let $(\Omega, \mathcal{F}, \mathbb{P})$ denote the underlying complete probability space. We observe $n$ independent and identically distributed observations $W_i = (Y_{0,i}, Y_{1,i}, A_i, L_i)$, for $i = 1, \dots, n$, where $Y_{t,i} \in \mathbb{R}$ denotes the outcome at time $t \in \{0,1\}$, $A_i \in \{0,1\}$ is the binary treatment indicator (1 for treated, 0 for control), and $L_i \in \mathbb{R}^p$ is a vector of pre-treatment covariates.
To accommodate unmeasured confounding, we allow for an unobserved random element $U_i$, which may be of arbitrary type (e.g., vector- or function-valued). Under the potential outcomes framework, let $Y_t^{a = \tilde{a}}$ denote the outcome at time $t$ had the unit been assigned treatment $A = \tilde{a}$, possibly contrary to fact.
The parameter of interest is the average treatment effect on the treated (ATT) at time \( t = 1 \), defined as \[ \theta = \mathbb{E}\left[Y_1^{a = 1} - Y_1^{a = 0} \,\middle|\, A = 1\right], \] where \( \theta \in \Theta \subset \mathbb{R} \). The ATT is particularly relevant in policy evaluation, as it captures the causal effect for the subpopulation that actually received the intervention, without requiring additional assumptions needed for identifying the average treatment effect (ATE).
Let \(P\) be the true distribution of the observed data. We denote the collection of relevant nuisance parameters by \(\eta \in \mathcal{H} \subseteq \tilde{\mathcal{H}}\), where \(\tilde{\mathcal{H}}\) is a normed space equipped with the \(L^2(P)\) norm, and $\mathcal{H}$ is a convex subset.
For a real-valued random variable \( X \in \mathbb{R} \), we denote its cumulative distribution function (CDF) and quantile function by \[ F_X(x) \equiv \Pr(X \le x), \quad Q_X(u) \equiv \inf\{x \in \mathbb{R} : F_X(x) \ge u\}, \quad u \in (0,1). \] When \( X \) admits a density with respect to a dominating measure (e.g., the Lebesgue measure for continuous outcomes), we denote it by \( f_X(x) \). Conditional versions given a random element $V$ are $F_{X \mid V}$, $f_{X \mid V}$, and $Q_{X \mid V}$. The support of $X$ under $P$ is denoted by $\mathrm{Supp}(X)$.
We begin by outlining the proposed identification strategy for the ATT $\theta$. We then extend the results to distributional causal estimands including the quantile treatment effect on the treated and the counterfactual distribution. We then illustrate two common data-generating scenarios that satisfy the identification assumptions. All technical proofs are deferred to the Supplementary Material (ref).
Assumption (ref) defines our setting and are standard in the causal inference hernan2020causal and difference-in-differences literature baker2025difference.
Assumption (ref) links observed outcomes to their corresponding potential outcomes. Assumption (ref) implies that, conditional on both observed covariates \( L \) and latent variables \( U \), treatment assignment is independent of the untreated potential outcome. This contrasts with the general case where \( A \) is not conditionally independent of \( Y_t^{a=0} \) given \( L \) alone. Assumption (ref) ensures sufficient overlap in treatment assignment across values of $L$ and $U$, enabling regular identification and estimation. In some settings, Assumption (ref) may be redundant if one is willing to posit an underlying graphical causal model, such as a single world intervention graph (SWIG) richardson2013single or a nonparametric structural equation model with independent errors pearl2009causality. These frameworks often imply Assumption (ref) through the temporal ordering of variables, particularly when treatment is assigned strictly between the pre- and post-treatment periods piccininni2025refining.
Assumption (ref) pertains to the distributional form of the outcomes and reflects the continuous nature of the outcome considered in this paper.
Assumption (ref) plays a central role in our identification strategy. It allows for causal inference despite the presence of unmeasured confounding, and serves as a relaxation of the traditional parallel trends assumption used in DiD. Let “$\overset{d}{=}$” denote equality in distribution.
Assumption (ref) states that, conditional on covariates and latent variables, the distribution of the untreated potential outcome at time $t = 1$ can be characterized by a monotonic transformation of the untreated baseline outcome and observed covariates. This distributional assumption is notably weaker than a rank-preservation condition, such as requiring a strictly monotonic relationship between potential outcomes $Y_0^{a=0}$ and $Y_1^{a=0}$ for each unit.
The distributional bridge condition is a stronger analog of the outcome confounding bridge function used in proximal causal inference miao2024confounding, cui2024semiparametric. In that framework, \(Y_0^{a=0}\) is treated as a negative control outcome or an outcome confounding proxy, and one assumes the existence of a square-integrable function \(g(y,l)\) such that \[ \mathbb{E}[g(Y_0^{a=0}, L) \mid L, U] = \mathbb{E}[Y_1^{a=0} \mid L, U]. \] A distributional bridge function satisfies this with \(g = \gamma\), but further imposes that for any measurable function \(h(y,l)\), \[ \mathbb{E}[h(\gamma(Y_0^{a=0}, L), L) \mid L, U] = \mathbb{E}[h(Y_1^{a=0}, L) \mid L, U], \] a condition that recovers the full conditional distribution of \(Y_1^{a=0}\).
This added stringency is not merely stronger---it is essential for identifying distributional causal estimands such as quantile treatment effects at arbitrary levels, which require recovering aspects of the outcome distribution beyond its mean. Additionally, it enables identification without requiring a separate treatment-confounding proxy, which is typically needed in the proximal causal inference framework.
The next result offers an equivalent and more concrete statement of Assumption (ref).
This characterization shows that Assumption (ref) implies a form of invariance of conditional optimal transport map with respect to unmeasured confounders $U$. We further elaborate on this point in Example (ref), which introduces a general class of semiparametric transformation models that comply with Assumption (ref).
Note that Assumptions (ref)-(ref) are invariant to monotonic transformations of the outcomes.
We now present our main identification result.
Following the main identification theorem, a couple of remarks are warranted.
While ATT serves as our primary estimand, the same assumptions allow identification of other causal estimands that may be of independent interest.
To illustrate the scope of our framework, we next describe two classes of data-generating processes (DGPs) that satisfy Assumptions (ref), (ref), and (ref). These examples concretely demonstrate the conditions under which ATT is identified as per Theorem (ref). Formal verification is deferred to Proposition (ref) in the Supplementary Material (ref).
The identification results established in the previous section provide a foundation for causal inference under our proposed framework. However, identification alone is not sufficient for developing practical estimators with desirable statistical properties, especially when there are (possibly high-dimensional) continuous covariates. To this end, we turn to modern semiparametric efficiency theory. We develop an efficient influence function-based estimator for the causal estimands and establish its theoretical properties.
A central tool in this theory is the efficient influence function (EIF), which serves multiple purposes. First, it provides straightforward characterization of the semiparametric efficiency bound—the lowest possible asymptotic variance—for regular and asymptotically linear (RAL) estimators of the target parameter under the specified model. Second, the EIF allows for the construction of a RAL estimator that attains the efficiency bound via solving estimating equations. Third, and crucially in modern applications, efficient-influence-function-based estimators are Neyman orthogonal neyman1959optimal to nuisance parameters, making them robust to estimation errors in nuisance components. This robustness is critical when using flexible machine learning methods to estimate high-dimensional or nonparametric nuisance functions, as it mitigates the impact of slow convergence rates. These advantages have been widely discussed in foundational and recent literature (e.g., newey1990semiparametric; bickel1993efficient; robins2008higher; chernozhukov2018double).
In what follows, we derive the EIF for ATT based on the identification results in Theorem (ref). We then generalize the approach to a class of causal parameters defined via moment restrictions, including counterfactual outcome distributions and quantile treatment effects on the treated (CDT and QTT) as examples.
We use the notation \(\mathbb{IF}(W; \theta, \eta)\) to denote the influence function associated with the target parameter \(\theta\), evaluated at a data point \(W\), where \(\eta\) represents a collection of nuisance functions. Define \(\gamma(y, l) \equiv Q_{Y_1 \mid A = 0, L=l} \circ F_{Y_0 \mid A=0, L=l}(y)\) and \(\nu(x,l) \equiv \frac{\Pr(A = 1 \mid \gamma(Y_0, l) = x, L=l)}{\Pr(A = 0 \mid \gamma(Y_0, l) = x, L=l)}\), and let \(\pi \equiv \Pr(A = 1)\). Then we have the following result.
We now extend the previous result to a broad class of estimands on the treated, denoted by \(\vartheta\), which are identified through a moment condition of the form:
where the function \(g(W; \vartheta, \gamma)\) takes the form \[ g(W; \vartheta, \gamma) = \tilde{g}(\gamma(Y_0, L), \vartheta), \] with \(\tilde{g}(\cdot, \cdot)\) known, right-continuous, and of bounded variation in its first argument. Several important causal estimands fall into this framework. For example:
The following theorem characterizes the efficient influence function (EIF) for this general class of estimands.
We now specialize this result to the estimands in Corollary (ref). Define the sign function \(\operatorname{sign}(x) = \mathbb{1}\{x > 0\} - \mathbb{1}\{x < 0\}\), and introduce the compound sign function \[ \chi(x, W; \gamma) \equiv \operatorname{sign}(Y_1 - \gamma(Y_0, L)) \cdot \mathbb{1}\left\{ \min(Y_1, \gamma(Y_0, L)) \le x \le \max(Y_1, \gamma(Y_0, L)) \right\}, \] which takes values in \(\{-1, 0, 1\}\) and simplifies the representation of the EIF.
We construct an estimator for the average treatment effect on the treated (ATT), denoted by \(\theta\). Since the efficient influence function (EIF) \(\mathbb{IF}(W; \theta, \eta)\) satisfies the moment condition \(\mathbb{E}[\mathbb{IF}(W; \theta, \eta)] = 0\), it can be used as a valid estimating function. Accordingly, we define \(\psi(W; \tilde{\theta}, \tilde{\eta})\) to share the same functional form as \(\mathbb{IF}(W; \theta, \eta)\), where \(\tilde{\theta} \in \Theta\) and \(\tilde{\eta} \in \mathcal{H}\) represent candidate values for the target parameter and nuisance components, respectively. Estimators for more general causal estimands defined by Equation (ref) can be constructed analogously using their corresponding EIFs from Theorem (ref). For brevity, we focus here on ATT.
Our estimation strategy employs sample splitting and cross-fitting schick1986asymptotically, chernozhukov2018double to mitigate overfitting bias. Moreover, the use of the Neyman-orthogonal EIF from Theorem (ref) helps to control regularization bias arising from the estimation of complex or high-dimensional nuisance functions. The resulting estimator is semiparametric efficient and supports valid inference even when nuisance quantities are learned via flexible machine learning methods with convergence rates slower than \(\sqrt{n}\), under mild regularity conditions.
To formalize the construction, let \(\{W_i\}_{i=1}^n\) denote an i.i.d. sample. Define the index set \([n] \equiv \{1, \ldots, n\}\), and fix an integer \(K\) denoting the number of folds used for cross-fitting. Assume for simplicity that \(n\) is divisible by \(K\), and let \((\mathcal{I}_k)_{k \in [K]}\) be a random, equal-sized partition of \([n]\). For each \(k\), let \(\mathcal{I}_{(-k)} \equiv \bigcup_{k' \neq k} \mathcal{I}_{k'}\) denote the training set excluding fold \(k\).
Nuisance functions are estimated using only data from \(\mathcal{I}_{(-k)}\), yielding estimates \[\hat{\eta}_k = \hat{\eta}((W_i)_{i \in \mathcal{I}_{(-k)}}; \zeta),\] where \(\hat{\eta}\) may be obtained via machine learning or nonparametric methods, and \(\zeta\) is a tuning parameter chosen either a priori or via cross-validation. Let \(z_{\alpha}\) denote the \((1 - \alpha)\)-quantile of the standard normal distribution \(\mathcal{N}(0,1)\). The full estimation and inference procedure is summarized in Algorithm (ref).
We now analyze the asymptotic properties of the proposed estimator for the average treatment effect on the treated (ATT), denoted by \(\theta\). These results justify using the efficient influence function (EIF) for point estimation and Wald-type inference. While we focus on ATT for brevity of exposition, analogous asymptotic guarantees extend to general estimands defined by Equation (ref), including quantile treatment effects, as a consequence of the Neyman orthogonality of the EIFs established in Theorem (ref); see chernozhukov2018double. Alternatively, confidence intervals can be obtained by inverting test statistics based on the EIF, as recently illustrated by lee2025inference in the instrumental variable setting.
Recall that \(\eta = (\gamma, \nu, \pi)\) denotes the true nuisance functions, and let \(\tilde{\eta} = (\tilde{\gamma}, \tilde{\nu}, \tilde{\pi})\) represent a generic element in \(\mathcal{H} \subseteq \tilde{\mathcal{H}}\). We endow \(\tilde{\mathcal{H}}\) with the \(L^2(P)\) norm defined by: \[ \|\tilde{\eta}\| \equiv \|\tilde{\eta}\|_{L^2(P)} = \left( \mathbb{E}\left[ \tilde{\gamma}(Y_0, L)^2 + \left\{ \tilde{\nu}(\gamma(Y_0, L), L) \right\}^2 + \tilde{\pi}^2 \right] \right)^{1/2}. \] We also define the norms for each component relative to their respective function spaces: \[ \|\tilde{\gamma}\| = \left( \mathbb{E}\left[ \tilde{\gamma}(Y_0, L)^2 \right] \right)^{1/2}, \quad \|\tilde{\nu}\| = \left( \mathbb{E}\left[ \left\{ \tilde{\nu}(\gamma(Y_0, L), L) \right\}^2 \right] \right)^{1/2}, \quad \|\tilde{\pi}\| = |\tilde{\pi}|. \] We begin by formalizing conditions on the nuisance function class and the estimation rate of the nuisance estimators.
To clarify the robustness of the estimator, we first characterize the bias induced by deviation from the true nuisance functions via the first and second Gateaux derivatives of the moment function \(\mathbb{E}[\psi(W; \theta, \eta)]\). Let $A \lesssim B$ denote that $A \le c\cdot B$ for some constant $c > 0$.
We now present the main asymptotic result. Let \(\mathcal{P}\) denote the collection of data-generating processes that satisfy Assumptions (ref)--(ref).
Hypothesis testing for $\theta$ can proceed via the constructed confidence interval. Bootstrap procedures could offer further refinements; we leave their investigation to future work.
\paragraph{Data-Generating Process} We simulate data from the semiparametric transformation model introduced in Example (ref) to assess the performance of various estimators. Complete simulation specifications, including all functional forms and parameter settings, are provided in Supplementary Material (ref).
Covariates \( L \in \mathbb{R}^p \) with \( p = 6 \) are independently drawn from a standard multivariate normal distribution. Treatment assignment is governed by a nonlinear function of both observed and unobserved variables:
where \( U_1, U_2 \sim \mathcal{N}(0,1) \) are independent unmeasured confounders, and \( \operatorname{expit}(x) \equiv 1/(1 + e^{-x}) \) denotes the logistic function.
Potential outcomes under no treatment evolve over two time periods \( t = 0,1 \) according to:
where \( \varepsilon_t \sim \mathcal{N}(0,1) \) are independent error terms, and \( U \equiv (U_1, U_2) \). The functions \( k_t(\cdot) \) and \( m(\cdot, \cdot) \) capture nonlinear relationships with covariates and unobservables.
To simplify exposition, we impose the sharp null hypothesis for treated potential outcomes: \( Y_t^{a=1} = Y_t^{a=0} \). Under this assumption, the true average treatment effect on the treated (ATT) is zero in the simulated data-generating process.
\paragraph{Setup} We simulate datasets of sizes \( n = 500,\, 1000,\, 2000 \), repeating each scenario 1000 times. The following three estimators are compared:
\paragraph{Results} Table (ref) and Figure (ref) summarize the performance of the estimators across key metrics: coverage probability at the 95% nominal level, bias, average estimated asymptotic standard error (SE), and empirical standard deviation (SD) of the estimates across 1000 replications.
The Debiased CiC estimator consistently achieves coverage close to the nominal level (93.5% - 94.0%), exhibits negligible bias, and produces standard error estimates that closely align with empirical standard deviations—even at smaller sample sizes.
In contrast, both CiC and DiD exhibit deteriorating coverage as \( n \) increases. The DiD estimator suffers from persistent bias stemming from violations of the parallel trends assumption, resulting in coverage probabilities that approach zero at larger sample sizes. While the CiC estimator shows diminishing bias as \( n \) increases, the convergence is too slow to ensure valid inference, leading to substantial undercoverage across all sample sizes considered ($80.8\% - 86.5\%$).
Overall, these results highlight the advantages of the proposed Debiased CiC estimator in delivering reliable inference and low bias under complex, nonlinear, and confounded data-generating processes.
\paragraph{Background} Mass shootings are among the most visible and devastating forms of gun violence in the United States, inflicting profound harm on communities and dominating national political discourse peterson2024epidemiology. Despite widespread public support for gun control, policy responses remain limited and inconsistent. This disconnect between tragedy and legislative action presents a central puzzle in American politics hassell2025navigating. These events, amplified by media coverage, often galvanize political attention and mobilize voters, making them more than isolated tragedies. They can shift public attitudes, influence political behavior, and alter electoral incentives. Studying their electoral impact thus provides a lens into broader structural forces shaping gun policy.
Prior studies use panel data and difference-in-differences (DiD) designs to estimate these effects, but results are mixed hassell2020mobilize, yousaf2021sticking, garcia2022violence, hassell2025navigating. Many rely on the strong parallel trends assumption and face challenges from confounding. Counties vary across both observable traits (e.g., demographics, economics, geography) and characteristics that are difficult to measure or quantify (e.g., media narratives, community activism, local gun culture). While the former can be controlled for, the latter are often high-dimensional and may influence outcomes in complex, nonlinear ways.
We address these challenges with the proposed approach that offers two key advantages: (1) it relaxes the parallel trends assumption, enabling more credible causal inference; and (2) it leverages machine learning to incorporate a rich set of covariates, thereby mitigating confounding bias. The empirical analysis below demonstrates the practical value of this methodology in addressing urgent, real-world policy questions.
\paragraph{Data Description & Setup} We re-analyze the dataset from yousaf2021sticking, yousaf2022replication, extending the original specification by incorporating a richer set of relevant covariates. The sample consists of approximately 3,000 U.S. counties observed in the 2004 (pre-treatment) and 2008 (post-treatment) presidential election years, a cycle selected for its consistency and completeness across data sources. The treatment indicator equals one for counties that experienced a mass shooting—successful or failed—between 2005 and 2008, where a successful mass shooting follows the FBI definition of an event resulting in four or more deaths at a single location. In total, 65 counties were classified as treated. The outcome is the county-level Republican vote share in the presidential elections. Covariates, measured prior to treatment, include economic indicators (household income, Gini coefficient), demographic characteristics (population size, proportion never married, racial Herfindahl-Hirschman index), education (proportion of college dropouts), health measures (proportion of residents reporting mental health issues), and geographic location (state).
\paragraph{Estimation & Results}
We estimate the effect of mass shootings using two methods: (a) the proposed Changes-in-Changes (Debiased CiC) estimator and (b) the standard Difference-in-Differences (DiD) estimator, which relies on the (conditional) parallel trends assumption and is implemented in the R package did callaway2021did. For each method, we consider two specifications: with and without covariates. Notably, Debiased CiC without covariates is asymptotically equivalent to the CiC estimator described in Section 5 of athey2006identification.
Table (ref) reports point estimates, standard errors (SE), and 95% confidence intervals for the ATT using the four estimators described above. Our covariate-adjusted Debiased CiC estimate — smaller in magnitude and statistically indistinguishable from zero — suggests a muted electoral response. A comparison across estimators highlights three key insights:
First, Debiased CiC consistently yields smaller estimated effects than DiD, both with and without covariates. This attenuation supports the concern that DiD may overstate treatment effects when the parallel trends assumption is violated. In contrast, Debiased CiC relaxes this assumption, allowing for more flexible, nonlinear trends, and potentially producing more credible estimates.
Second, covariate adjustment substantially influences the estimated effects. Without covariates, the ATT is estimated at $-2.60\%$ under DiD and $-1.63\%$ under Debiased CiC. Including covariates reduces the magnitudes of both estimates by over one percentage point, underscoring the importance of addressing confounding.
Third, Debiased CiC with covariates achieves a lower standard error (SE = 0.46%) compared to that without covariates (SE = 0.53%) and is comparable to DiD with covariates (SE = 0.47%). Notably, DiD often adopts a parametric linear model by default, such as in the did package, which may not efficiently handle high-cardinality categorical variables like state (with 50 levels). This can inflate variance and reduce efficiency. By contrast, our Debiased CiC framework leverages machine learning to flexibly and adaptively incorporate rich covariate information, improving precision without imposing restrictive functional form assumptions.
This paper introduces a novel extension of the Changes-in-Changes (CiC) framework that accommodates high-dimensional, non-monotonic unmeasured confounding and allows for continuous covariates. By deriving the efficient influence function and constructing Neyman-orthogonal estimators, we provide a semiparametrically efficient estimation strategy that supports valid inference with flexible, machine learning-based nuisance estimation. This work fills a key methodological gap in the Difference-in-Differences literature, where CiC has served as a foundational tool.
While we focused on the two-period setting with continuous outcomes, the proposed method can be adapted to multiple time periods, discrete outcomes, and staggered adoption designs. Future work may also extend our approach to enhance existing CiC-based methods—such as those for mediation huber2022direct or multivariate outcomes torous2024optimal—which often assume no covariates.
When multiple pre-treatment periods are available, the key distributional bridge assumption can, in principle, be assessed empirically, providing a basis for a falsification or placebo test. We leave the development of formal testing procedures to future work. While our simulations highlight the robustness and efficiency of the proposed estimator, careful implementation and sensitivity analyses remain essential in applied settings.
The authors would like to acknowledge the National Institutes of Health (NIH) for their generous funding and support.
Supplementary Material provides additional technical results, complete proofs, and further simulation details. The data used in the application section are publicly available from yousaf2022replication. Code for replication is available from the first author upon request.