Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
118,808 characters · 22 sections · 98 citation commands
Doubly Robust Instrumented Difference-in-Differences
Keywords: Difference-in-differences, instrumental variables, causal inference, semiparametric inference, instrumented difference-in-differences, double machine learning
Difference-in-differences (DiD) is a central tool in applied econometrics for estimating causal effects in non-experimental settings. Recent work, often referred to as the “DiD renaissance”, has exposed important limitations of classical DiD methods, particularly under treatment effect heterogeneity across cohorts and over time. In such settings, coefficients from standard regression specifications can lack a clear causal interpretation.
Two broad responses have emerged. One is to refine the underlying regression model. The other is to begin with a well-defined, interpretable statistical target, an estimand, rather than starting from a regression model. Under suitable identifying assumptions, this estimand can be linked to observable quantities and estimated in a way that admits a transparent causal interpretation. A central object in this approach is the efficient influence function (EIF), which is used to construct semiparametrically efficient estimators and valid inference. This perspective is adopted, for example, in the work of sazhao and csa.
However, even modern DiD approaches can fail when the treatment is endogenous. A classical remedy is the use of instrumental variables (IV). The integration of IV with DiD remained underdeveloped until miyaji, who introduced the instrumented difference-in-differences (IDiD) framework with staggered exposure to the instrument.
Adopting an estimand-based perspective, miyaji defines the causal parameter of interest, the cohort-specific time-varying local average treatment effect on the treated ($LATT(e, t)$), and links it to an estimable quantity under identifying assumptions. However, the framework does not incorporate covariates, which are often essential in empirical applications to support the plausibility of the identifying assumptions.
In related work, csax identify the $LATT$ parameter of miyaji with the inclusion of covariates as a special case of their sequential conditional moment restriction framework for DiD. They derive the efficient influence function of the parameter in the panel data setting and construct a corresponding doubly robust estimator. However, their analysis focuses on the case with a single exposure date to the instrument. Moreover, they leave the case of repeated cross-sections for future work.
This paper studies the general case with covariates and staggered exposure to the instrument in both panel data and repeated cross-sections. We derive the efficient influence function for the $LATT$ parameter in both data settings. The repeated cross-sections case is empirically relevant in settings where balanced panels are unavailable, for example in “trimmed panel” data or repeated survey samples. A novel contribution is that the EIF is derived explicitly using the approach developed by kennedy. Moreover, as done in miyaji, the framework of csa is extended to the IDiD case allowing the use of not-yet-exposed units as controls, analogous to the not-yet-treated control group in the staggered DiD setting. Both control variables are handled generally by invoking either of the corresponding identifying assumptions.
Using the derived EIFs, we construct doubly robust estimands and corresponding estimators for the $LATT$ parameter in both data settings. The construction follows the estimating equation approach: the estimand solves the population moment condition implied by the EIF, while the estimator solves its empirical counterpart. This construction also underlies the doubly robust estimand and estimators in sazhao (and, by extension, csa), although implicitly. The resulting doubly robust estimators have a structure closely related to those proposed by sazhao and csa. In particular, the estimator takes the form of a ratio of two doubly robust estimators for ATT-type parameters. The numerator corresponds to the outcome of interest and the denominator to the treatment variable, with the instrument exposure variable playing the role of the treatment indicator. This structure arises naturally because the identified $LATT$ parameter itself can be written as a ratio of two ATT-type parameters. The estimators reduce to those of miyaji in the case of no covariates.
The asymptotic behavior of the estimator is established via a decomposition into an influence function term, an empirical process term, and a remainder term (cf. kennedy). The remainder term is handled directly in the proofs. The empirical process term is controlled either via Donsker class assumptions or by employing cross-fitting, which permits the use of flexible machine learning methods. We derive DML estimators, which we show are equivalent to the cross-fitted estimator in kennedy, for both data settings.
A simulation study illustrates the finite-sample properties of the proposed estimators. A freely available implementation is provided in the Python package idid.
\paragraph{Contributions} This paper makes three contributions. First, we derive the efficient influence function for the $LATT$ parameter in the IDiD framework with covariates and staggered instrument exposure. The derivation covers both panel data and repeated cross-sections. Second, using the EIF, we construct doubly robust estimands and estimators for the $LATT$ parameter. Third, the paper illustrates how the influence function derivation strategy of kennedy can be applied to econometric target parameters. This approach provides a practical alternative to classical tangent space calculations newey1990,tsiatis and ensures that the resulting estimators target precisely the interpretable estimand specified by the researcher. We also construct Neyman-orthogonal scores for DML estimation using the EIFs and relate IDiD with staggered instrument exposure to DiD with staggered treatment via a Bloom-type result.
\paragraph{Related literature} This paper relates to several strands of the literature.
First, it contributes to the literature on IDiD. miyaji introduce the IDiD framework and define the $LATT$ parameter. csax extend this framework to allow for covariates and derive a doubly robust estimator in the panel setting with a single exposure date. The present paper studies the general case with covariates and staggered exposure to the instrument in both panel data and repeated cross-sections, and derives corresponding doubly robust estimands and estimators.
Second, the paper relates to the literature on semiparametric DiD estimation. sazhao and csa develop doubly robust estimators for ATT-type parameters in two-period and staggered adoption settings. The estimators proposed here have a similar structure but tailored to the IDiD framework. Moreover, in deriving the influence function, doubly robust estimand, and estimators for the LATT parameter, we simultaneously derive the corresponding results for the ATT parameter, thereby recovering the results of sazhao and csa from first principles.
Third, the paper connects to recent work at the intersection of modern semiparametric methods and DiD/IDiD, as well as to recent work on doubly robust IV estimators. For example, chang2020 develop DML estimators for DiD in the case of panel data and repeated cross-sections, deng develop a TMLE estimator for the two-period DiD parameter, while metalearner propose a meta-learner algorithm for the $LATT$ in two-period IDiD with panel data. sloczynskiDoublyRobustEstimation2022 develop doubly robust estimators for LATE and LATT in cross-sectional IV settings, whose structure resembles the panel data estimator derived here.
More broadly, the paper relates to recent work emphasizing clearly defined target parameters in causal inference and IV analysis, including mogstad2024instrumental, who distinguish between forward and reverse engineering approaches to IV parameters. In this paper, we take the forward engineering approach. It also connects to the targeted learning literature vdlrose,van2015statistics, which emphasizes aligning the estimand with the underlying scientific question and constructing estimators under minimal modeling assumptions, rather than relying on potentially misspecified parametric models. Finally, it relates to the double machine learning literature chernozhukovDoubleDebiasedMachine2018a, as we derive two DML estimators.
\paragraph{Organization of the paper} (ref) introduces IDiD, the causal estimands, and the main identification results. (ref) studies aggregated effects. (ref) develops the doubly robust and DML estimators and establishes their asymptotic properties. (ref) reports the simulation results. Proofs and derivations are collected in the appendix. (ref) provides a detailed account of how the influence function calculations fit into the development of the paper.
\paragraph{Notation} For a measurable function $f$, let $\Vert f \Vert_{q,P} = (\int |f|^q dP)^{1/q}$ denote its $L^q(P)$ norm. Write $P_n f = n^{-1}\sum_{i=1}^n f(X_i)$ for the empirical average and $Pf = \int f\, dP$ for the expectation under $P$. We also write $E[f(X)]$ for the expectation (under the relevant distribution) and use the notation interchangeably where convenient. If $f$ depends on a parameter $\tau$ and some nuisance $\eta$, we write $P_n f(\cdot;\tau, \eta) = n^{-1}\sum_{i=1}^n f(X_i;\tau, \eta)$ and $P f(\cdot;\tau, \eta) = \int f(\cdot;\tau, \eta)\, dP$. The empirical process is written $\sqrt{n}(P_n - P)[f]$. Calligraphic letters denote supports of random variables, e.g.\ $\mathcal{E}$, $\mathcal{D}$, and $\mathcal{Z}$ for $E$, $D$, and $Z$. For an event $A$, $\mathbf{1}\{A\}$ denotes its indicator. We write $O$ for a generic tuple of random variables that is context-dependent. See also (ref) for details on the notation for the influence function operator, $\mathbb{IF}$.
We first introduce the notation used throughout the article, building on miyaji. We consider the general case of $\mathcal{T}$ periods. Let $D_{t} \in \{0, 1\}$ denote treatment status and $Z_{t} \in \{0, 1\}$ the instrument status. Moreover, let $D = (D_{1}, D_{2}, \ldots, D_{\mathcal{T}})$ and $Z = (Z_{1}, Z_{2}, \ldots, Z_{\mathcal{T}})$ denote the treatment and instrument paths.
We make the following assumption on the instrument:
(ref) enforces that no units are exposed to the instrument in the first period and all units that are exposed in some period stay exposed\footnote{ This is analogous to Assumption 1 on the treatment variable assumed by csa. }. (ref) implies that the time period where the instrument switches on characterizes the instrument path $Z$ completely. Because of this we define the cohort exposure variable $E := \min \{t \mid Z_{t} = 1\}$ and $E_{e} := \mathbf{1}\{E = e\}$.
To identify the target parameter, valid control groups are needed. In the DiD literature, commonly used control groups are never-treated and not-yet-treated units. Here, we construct similar control groups but based on the instrument instead of the treatment variable. Hence, we define
for the never-exposed and not-yet-exposed control groups, respectively\footnote{ The control variable $C^{nye}_{e,s}$ is analogous to $(1-D_s)(1-G_g)$ in csa, but here corresponds to units not in the cohort exposed at $E = e$ and not yet exposed at time $s$, i.e., $Z_s = 0$. In the appendix, we provide a table comparing the different objects in the two-period DiD, staggered adoption DiD and our case of staggered IDiD; see (ref). }. Let $\bar{e} := \max_{i}E_{i}$; in the case of never-exposed units, $\bar{e} = \infty$, and in the case of only not-yet-exposed units, $\bar{e} < \infty$. Denote the support of the exposure variable excluding $\bar{e}$ as $\mathcal{E} := \mathrm{supp}(E) \setminus \bar{e} \subseteq \{2, 3, \ldots, \mathcal{T}\}$\footnote{ Analogous to csa, when there is a never-exposed cohort, $E = \infty$, $\mathcal{E}$ only excludes this. In the case of not-yet-exposed control groups only, we exclude the last exposed cohort because there are no available control groups for this cohort. }. Denote the generalized propensity scores corresponding to the control variable in use as
The conditioning on $E_{e} + C = 1$ restricts attention to the relevant $2 \times 2$ comparison: units are either in the exposed cohort, $E_{e} = 1$, or in the control group, $C = 1$. Within each such slice, the objects behave as in the corresponding two-period setup.
\paragraph{Potential treatment and outcomes} Let $D_{t}(\infty)$ denote a unit's unexposed potential treatment at time $t$ if they remain untreated through time period $\mathcal{T}$, i.e., if they were not to be exposed to the instrument across all available time periods. For $e = 2, 3, \ldots, \mathcal{T}$, let $D_{t}(e)$ denote the potential treatment that a given unit would experience at time $t$ if they were to first become exposed to the instrument in time period $e$. The observed and potential treatment for a given unit are related through
i.e., we only observe one potential treatment path for each unit.
Let $Y_{t}(d, z)$ be a given unit's potential outcome in period $t$ had they been given treatment path $D = d$ and instrument path $Z = z$. The analogy between the treatment and instrument for the DiD and IDiD frameworks differs in the sense that the instrument does not affect the outcome directly but only through the treatment. Specifically, the instrument creates exogenous variation in the treatment that allows us to identify the effect of the treatment on the outcome in the presence of hidden confounders. This relation between the instrument and the potential outcomes is enforced in the following assumption:
(ref) means that the potential outcomes at time $t$ only depend on the treatment variable at time $t$, and that they do not depend directly on the instrument. The latter is analogous to the exclusion restriction in the simple cross-sectional IV design.
To arrive at an expression for the observed outcome, let $Y_{t}(0)$ denote a given unit's potential outcome at time $t$ if they are untreated at time period $t$, and $Y_{t}(1)$ if they are treated at time period $t$. With (ref) we can write the observed outcome in terms of the potential outcomes and treatments as:
Using (ref), we define the exposed/unexposed outcomes as:
The exposed/unexposed outcomes\footnote{ As noted by miyaji, the concept of exposed and unexposed outcomes is not new. In the standard cross-sectional binary IV setup with potential outcomes $Y(0), Y(1)$, potential treatments $D(0), D(1)$, and instrument $Z$, the observed treatment can be written as $D = D(0) + [D(1) - D(0)]Z$. Substituting this into the observed outcome equation $Y = Y(0) + [Y(1) - Y(0)]D$ yields the exposed and unexposed outcomes $Y(D(Z)) := Y(0) + [Y(1) - Y(0)]D(Z)$ for $Z \in \{0,1\}$. This is analogous to the staggered exposure setting in (ref), which leads to the exposed and unexposed outcomes in (ref). } play a key role in the identification results, (ref).
\paragraph{Sampling assumption} Our results apply to both panel data and repeated cross-sections, which are covered by the following assumption. Let $T \in \{1,\ldots,\mathcal{T}\}$ denote the period a unit is observed in the repeated cross-sections case.
(ref) implies that the observed data consist of panel data, whereas (ref) implies that the observed data are i.i.d. draws from the mixture distribution
where $\lambda_{t} := P(T_{t} = 1)$ and $T_{t} := \mathbf{1}\{T = t\}$. This mixture distribution setup is analogous to abadieSemiparametricDifferenceinDifferencesEstimators2005a,sazhao,csa but here for the IDiD design with staggered exposure.
The target parameter in this paper is the cohort-specific time-varying local average treatment effect on the treated of miyaji
(ref) is the treatment effect, $Y_{t}(1) - Y_{t}(0)$, averaged over the subpopulation of compliers, $D_{t}(e) > D_{t}(\infty)$, and the units exposed to the instrument in period $e$, $E_{e} = 1$. The parameter varies across cohorts $E$ and time $t$; hence it allows us to answer questions related to the heterogeneity across cohorts and time. In (ref) we show how to aggregate the parameters into aggregated effects; similar to how csa aggregates their $ATT(g, t)$ parameters.
(ref) is analogous to the monotonicity assumption in cross-sectional IV and requires that the instrument, here the exposure cohort variable $E$, affects the treatment in only one direction.
(ref) is analogous to the standard no-anticipation assumption in DiD, with the exposure variable replacing the treatment, and requires that exposure does not affect the treatment prior to the exposure date.
The following two assumptions use the never-exposed control group $C^{nev}$.
The following two assumptions are analogous to the previous two, but use the not-yet-exposed control group $C^{nye}_{e, s}$.
In the identification proofs we invoke either pair of the four assumptions above depending on the control variable used. If using the never-exposed control group $C^{nev}$, we assume (ref), and if using the not-yet-exposed control group $ C^{nye}_{e,t}$, we assume (ref). Writing $C$ for a generic control variable allows us to encompass both control variables in the identification arguments.
(ref) is a standard overlap condition. It requires that, for each exposure cohort, the probability of exposure is bounded away from zero, and from one conditional on covariates. This ensures that all relevant subpopulations have a non-negligible probability of being both exposed and unexposed, which is necessary for identification and stable estimation.
To derive the main identification result, we introduce the central objects and notation below. Let $V_t$ be a generic random variable and define \[ \Delta_{t-e+1} V_t := V_t - V_{e-1}, \] i.e.\ the change of $V_{t}$ between period $t$ and the pre-exposure period $e-1$.
Panel data. For panel data, define the outcome regression functions for treated, never-exposed, and not-yet-exposed units as
Analogously, define $g_{e, t}^{trt, p}(X)$, $g_{e, t}^{nev, p}(X)$, and $g_{e, t}^{nye, p}(X)$ with outcome $\Delta_{t-e+1} D_t$.
Repeated cross-sections. For repeated cross-sections, the outcome regression functions are defined as
with corresponding definitions $g_{e, t}^{trt, rc}(X)$, $g_{e, t}^{nev, rc}(X)$, and $g_{e, s, t}^{nye, rc}(X)$ obtained by replacing $Y$ with $D$.
\paragraph{Encompassing both never exposed and not-yet-exposed} A generic control variable is written as $C$ and the corresponding propensity as $p(X)$. Below we write $m_{e, t}^{c, p}(X), g_{e, t}^{c, p}(X)$ for a generic control variable outcome regression function in the case of panel data, and similarly $m_{e, s, t}^{c, rc}(X), g_{e, s, t}^{c, rc}(X)$ in the case of repeated cross-sections.
The proofs for the not-yet-exposed and never-exposed cases are identical, differing only in the control indicator and propensity score. To avoid repetition, we present a unified argument covering both cases. The results follow under (ref) for the never-exposed case and under (ref) for the not-yet-exposed case. This is done for both the panel-data and repeated-cross-sections settings.
Proof in (ref).
\paragraph{Influence functions and construction of the DR Estimands} Below we derive DR estimands and corresponding DR estimators for the $LATT(e, t)$ in both data settings. The steps are as follows:
As the above procedure shows, the EIF of the target parameter is central to deriving the DR estimands. A key feature is that the identified LATT parameters are ratios of ATT components and that the staggered exposure setting reduces to two-period comparisons. This allows us to start from the canonical two-period DiD case targeting the ATT parameter and derive the DR estimands and EIFs in this setting. As part of this derivation, we recover the DR DiD estimands of sazhao from first principles. We then exploit the ratio structure of the LATT parameters together with the EIF machinery of kennedy developed for our setting in (ref) to combine the components. Constructing the resulting estimators requires additional work and is taken up in (ref).
\paragraph{Weights definition} As noted in Step (ref) of the procedure, the doubly robust estimands and their corresponding EIFs rely on Hájek-type normalized weights. The construction of these weights in both sampling settings follows the same principle as in sazhao,csa: the control weights are normalized to sum to one in sample (in contrast to the unnormalized control weights appearing in the EIF derived in Step (ref)).
The weights in the panel data setting are:
The weights in the repeated cross-sections setting are:
In the repeated cross-sections setting, we further define:
for $w^{c,rc}_{e, t, s}$ a generic control-weight.
\paragraph{Estimand and EIF}
Proof in (ref).
\paragraph{DR estimands and estimator decomposition} In the following, we denote a generic estimand by $\tau$ and its plug-in estimator by $\hat{\tau}$. A generic decomposition of this plug-in estimator into a CLT, empirical process and remainder term, similar to kennedy\footnote{ We depart from the notation in kennedy by indexing the EIF explicitly by the target parameter and nuisance functions, $\varphi(\cdot; \tau, \eta)$, rather than by the distribution $P$ alone. }, can be written as:
where $\varphi(\cdot; \tau, \eta)$ is the influence function and $\eta$ is a tuple of nuisance parameters. We derive our doubly robust estimands by using the EIF, $\varphi(\cdot; \tau, \eta)$, corresponding to $\tau$ and solving
for the new target parameter $\tau^{dr}$, where the superscript "dr" means doubly robust. This is also how the estimands of sazhao are derived (although not explicitly shown in their paper). In (ref) we show this from first principles, hence demystifying where the estimands come from, building up to our DR estimands for the $LATT$ parameter in both data settings. The resulting estimators are of the "estimating equation form", i.e., for $\hat{\eta}$ a generic estimator of the nuisance, the estimator, $\hat{\tau}$, solves:
Note that the DR estimands, e.g. the $\tau^{dr}$ estimand found through (ref), have their own influence functions, which differ from those of the original parameter, e.g. $\tau$. Denote this influence function for the DR estimand as $\varphi^{dr}(\cdot; \tau^{dr}, \eta^{dr})$. We use this influence function to conduct inference for the DR estimators. This is also what sazhao does. However, sazhao also study the estimation effects arising from the nuisance function estimators of $\eta^{dr}$. The estimation effects arise from their linearization of their estimators and entails Taylor expanding the components building up to their influence function. In the IDiD setting, the estimand and estimators are ratios and hence the Taylor expansion becomes significantly more tedious to derive. Hence, in this paper, we do not pursue this approach. Instead, we take the more general approach of either assuming Donsker conditions or using cross-fitting to tame the empirical process term in (ref) cf. kennedy. The remainder term is handled explicitly in the proofs; see (ref). A drawback of taking the more general approach is that our estimator does not inherit the DR-for-inference property as in sazhao; we leave this for future work (see also dukesvansteelandt).
\paragraph{DR Estimand and EIF}
Proof in (ref).
Proof in (ref).
Proof in (ref).
\paragraph{Estimand and EIF} In the repeated cross-sections case, we denote generic mean functions for both control groups in (ref) by $m^{c,rc}_{e,s,T}(X)$ and $g^{c,rc}_{e,s,T}(X)$ (again, in the never-exposed case, the index $s$ is not used). This notation allows us to treat both cases jointly. Using this notation, we define the control mean functions unified across periods $e-1$ and $t$ as:
and
where implicitly, because we only consider $2 \times 2$ comparisons, $\mathbf{1}\{T = e - 1\} + \mathbf{1}\{T = t\} = 1$.
Proof in (ref).
\paragraph{DR Estimand and EIF}
Proof in (ref).
Proof in (ref).
Proof in (ref).
In the special case of absorbing treatment (defined below) and one-sided compliance (i.e., no units unexposed to the instrument are treated), the $LATT(e,t)$ parameter can be related to the instrument-exposure-cohort-specific $ATT(g,t)$ parameters of csa. The following result shows this:
Proof in (ref).
Researchers often compare the target parameter across subgroups. This can be handled explicitly by expressing subgroup parameters as functionals of the full population distribution. Differences are then represented by embedding the corresponding influence functions in a common full-sample framework, weighted by the group indicator and its probability. To see this, let $B \in \{m, f\}$ denote a group indicator. We consider the difference in subgroup-specific LATTs, $LATT^{\Delta}(e,t) := LATT^{m}(e,t) - LATT^{f}(e,t)$, for example based on the DR estimand in the panel data setting:
where each component admits an influence function under the corresponding subgroup distribution, $\varphi^{dr, p, m}(O; \tau^{dr, p, m}_{e, t}, \eta^{m})$ and $\varphi^{dr, p, f}(O; \tau^{dr, p, f}_{e, t}, \eta^{f})$, respectively. These subgroup influence functions can be embedded into the full population as:
where $\eta^{dr, p, \Delta}_{e, t} = (\eta^{dr, p, m}_{e, t}, \eta^{dr, p, f}_{e, t})$ collects the subgroup-specific nuisance functions. Thus, inference on (ref) can be based on (ref). The group-level difference estimand in the case of repeated cross-sections is constructed in the exact same way and is omitted for brevity.
As shown in miyaji, the $LATT(e,t)$ parameters can be aggregated into estimands of interest, in a manner entirely analogous to the aggregation of $ATT(g,t)$ parameters in csa. In this section, we apply the EIF machinery developed in (ref) to derive the influence function for a general weighted estimand $\theta$, defined below.
Consider an estimand that aggregates the $LATT(e,t)$ parameters:
where $w(e,t)$ denotes a (possibly data-dependent) weighting scheme. To ease notation, define $\tau_{e, t} := LATT(e, t)$. Let $\mathcal{I} := \{(e, t) \in \mathcal{E} \times \{2, 3, \ldots, \mathcal{T}\} \mid A_{e, t} \}$ be a set of indices that picks out the correct cohorts and time indices related to the summary parameter $\theta$. By the product rule (ref), the EIF of the weighted estimand (ref) equals:
\paragraph{Weighted estimands and weights} Define the cohort-specific average exposed effect on the treated in the first stage, i.e., the share of the compliers in cohort $e$ in period $t$:
(ref) shows the weighted estimands and their corresponding weights. For comparison, we also include the same aggregation procedure as in csa, but without the cohort-specific time-varying complier weights. This allows researchers to choose whether to weight by complier shares or not, although the complier-weighted version is the most natural when the target parameter is the LATT.
The weights are themselves estimands, so the tools developed in (ref) can be applied to derive influence functions for each weight in (ref). With these in hand, we apply (ref) to obtain the influence function of the weighted estimand.
As is evident from (ref), most of the weights consist of sums of cohort probabilities conditional on different events, and the time-varying complier share $AET(e, t)$. For instance, for the weighted estimand $\theta^{o,IV}_{W}$, the main component of the weight has influence function (using the quotient rule (ref))
with $\mathbb{IF}\left(P(E \leq \mathcal{T})\right) = \sum_{e \leq \mathcal{T}} \mathbb{IF}\left(P(E = e)\right)$ and $\mathbb{IF}\left(P(E = e)\right) = \mathbf{1}\{E = e\} - P(E = e)$ cf. (ref). The influence function of the weight can now be derived using (ref) and the quotient-rule (ref) summing the correct terms.
\paragraph{Event study at horizon $l$ parameter, $\theta_{es}^{IV}(l)$} Here, we focus on the event study at horizon $l$ target parameter, which aggregates the $LATT(e, t)$ parameter at time $t = e + l$, i.e., $l$ periods after exposure to the instrument. The target parameter is:
where the weight $w^{IV}_{es(l)}(e, t)$ is defined in the upper part of (ref). \paragraph{Influence function of $\theta_{es}^{IV}(l)$} To derive the influence function of the $\theta_{es}^{IV}(l)$ estimand, we apply the product rule (ref) and derive the influence function of each component separately. First, we find the influence function of the weight $w^{IV}_{es(l)}(e, t)$. We apply the product rule (ref) to get the EIF of a generic term in the numerator and denominator, and then use the quotient rule to get the EIF of the weight. Towards this end, define
For the conditional probability we can use the block (ref), and the EIF for $AET(e, t)$ is the IF of the denominator for our (estimable) target parameter at hand. Hence, applying the quotient rule gives:
Let $\mathcal{I}_{l} := \{(e, t) \in \mathcal{E} \times \{2, 3, \ldots, \mathcal{T}\} \mid e + l \leq \mathcal{T} \wedge t = e + l\}$. Thus, using (ref), and writing $\varphi_{e, t}(O; \tau_{e, t}, \eta)$ for the influence function of the generic $\tau_{e, t}$, we get using (ref):
(ref) is the influence function we use to conduct inference with on the horizon $l$ parameters, $\theta^{IV}_{es(l)}$.
Note that we chose $\theta^{IV}_{es(l)}$ over $\theta^{IV}_{bal,es}(l,l')$ for simplicity. The point on compositional changes making the interpretation of $\theta^{IV}_{es(l)}$ harder, as pointed out in csa, still applies, and it is up to the econometrician to make the trade-off of interpretability vs efficiency (more observations).
\paragraph{Multiplier Bootstrap and Simultaneous Confidence Bands} Researchers often interpret multiple $\theta^{IV}_{es(l)}$ jointly. In this case, simultaneous confidence bands are appropriate. csa show how to construct such bands via the multiplier bootstrap and establish validity for the vector of $ATT(g,t)$ in their Theorem 3. An analogous result applies to the vector of $LATT(e,t)$, allowing us to use their Algorithm 1 for inference on our parameters of interest. In particular, their Corollary 1 implies that the resulting bands achieve correct asymptotic coverage, so simultaneous inference on $\{\theta^{IV}_{es}(l) : l \in \{0,1,\ldots,h\}\}$ is valid. The multiplier bootstrap is applied to the influence function (ref) for each $l$. We verify the procedure in (ref).
We consider plug-in estimators of $\tau^{dr,p}_{e,t}$ and $\tau^{dr,rc}_{e,t}$ in (ref). We take two approaches. First, we estimate the nuisance functions using simple parametric models (e.g.\ linear or logistic regression) and invoke Donsker conditions to control the empirical process term. Second, we allow for flexible machine learning methods, avoid empirical process assumptions, and instead use DML. We show below that the resulting DML estimators coincide with the cross-fitted estimator of kennedy.
\paragraph{Asymptotically linear representation} The plug-in estimators of $\tau^{dr,p}_{e,t}$ and $\tau^{dr,rc}_{e,t}$ admit the decomposition in (ref). kennedy shows that the general plug-in estimator has first-order bias equal to $- \int \varphi(\cdot; \hat{\tau}, \hat{\eta}) \, dP$. This bias can be removed by imposing the moment condition $P_n[\varphi(\cdot; \hat{\tau}, \hat{\eta})] = 0$ and solving for the estimator, yielding an estimating equation estimator\footnote{ We show in (ref) that the cross-fitted estimating equation estimator coincides with the DML estimator under the DML2 algorithm; see (ref). }\footnote{ Alternative approaches include one-step estimators and targeted maximum likelihood estimation (TMLE); see kennedy. }.
As discussed in (ref), the estimators based on the DR estimands in this paper are of this type and therefore eliminate the first-order bias by construction\footnote{ Although the plug-in estimator of (ref) solves the estimating equation based on (ref), $P_n \varphi^{p}(O; \hat{\tau}^{dr,p}_{e,t}, \hat{\eta}^{p}_{e,t}) = 0$, rather than the IF $\varphi^{dr,p}(\cdot; \tau^{dr,p}_{e,t}, \eta^{dr,p}_{e,t})$ in (ref), the plug-in bias is still zero since $P_n \varphi^{dr,p}(O; \hat{\tau}^{dr,p}_{e,t}, \hat{\eta}^{dr,p}_{e,t}) = 0$ as well. The same holds for repeated cross-sections. This follows by direct calculation; see (ref). }. This justifies the decomposition in (ref) for our plug-in estimators. To obtain the asymptotic linear representation
it remains to control the empirical process and remainder terms. Specifically, to go from (ref) to (ref) we require
The first condition is ensured by cross-fitting or Donsker assumptions, while the second requires case-specific arguments kennedy, formalized in (ref).
The following proposition provides conditions under which the remainder term is asymptotically negligible, i.e., satisfies the right-hand side of (ref), in the panel data setting.
Proof in (ref).
The following proposition provides conditions under which the remainder term is asymptotically negligible, i.e., satisfies the right-hand side of (ref), in the repeated cross-sections setting.
Proof in (ref).
In the following, we derive asymptotic linear representations for plug-in estimators of $\tau^{dr,p}_{e,t}$ and $\tau^{dr,rc}_{e,t}$ under Donsker class conditions, building on the remainder results in (ref). Such conditions are often restrictive in high-dimensional settings. Accordingly, (ref) introduces a DML estimator that avoids these assumptions by controlling the empirical process term via cross-fitting.
\paragraph{Panel Data}
\paragraph{Repeated Cross-Sections}
In this section, we present DML estimators for both data settings. In each setting, the Neyman-orthogonal score is based on the corresponding EIF in (ref). The resulting estimators coincide with the cross-fitted estimating equation estimator of kennedy. This follows because DML applies cross-fitting to Neyman-orthogonal scores; when these scores are derived from the EIFs in (ref), the procedure is equivalent to cross-fitting the corresponding estimating equation estimator.
\paragraph{Regularization and overfitting bias} The Donsker assumptions made in the previous section are often unrealistic in settings where the nuisances are high-dimensional and complex. Here, flexible machine learning estimators are more appropriate. Estimating nuisance functions with machine learning induces two forms of bias, often referred to as regularization and overfitting bias in the DML literature chernozhukovDoubleDebiasedMachine2018a. The remedy is to leverage a Neyman-orthogonal score for the DML estimator and apply cross-fitting chernozhukovDoubleDebiasedMachine2018a. This is what we will do below.
\paragraph{DML Algorithm} We construct our DML estimators using the DML2 procedure of chernozhukovDoubleDebiasedMachine2018a,bach2022doubleml. The key components are:
Hence, we first construct the Neyman-orthogonal scores. For this we can use the influence functions (ref).
\paragraph{Panel Data} Using the EIF for the panel data setting in (ref), we obtain:
where we use DML notation for the score components, i.e. $\psi$, $\psi_{a}$ and $\psi_{b}$. Hence, we see that the influence function is proportional to a score (i.e. function that identifies the target parameter) that is linear in the estimand $\tau^{p}_{e, t}$.
Let $\{O_i\}_{i=1}^n$ be partitioned into $K$ folds of equal size $n_k = n/K$, and denote the empirical measure on fold $k$ by $P_n^k$\footnote{ For equal-sized folds: $K^{-1}P_{n}^{k}[f] = K^{-1}n_{k}^{-1}\sum_{i=1}^{n} f(O_{i}) = n^{-1}\sum_{i=1}^{n}f(O_{i})$. }. Let $\hat{\eta}^{p,-k}_{e,t}$ be the nuisance estimator trained on all observations outside fold $k$ (the superscript $-k$ indicates this). Hats denote estimated quantities.
The DML estimator solves:
Hence we have the result:
\paragraph{Repeated Cross-Sections} Using the EIF for the repeated cross-sections setting in (ref), we obtain:
Again, the influence function is proportional to a score that is linear in the estimand $\tau^{rc}_{e, t}$.
Similarly to the panel data case, the DML estimator solves:
where the score components are defined in (ref). Hence we have the result:
\paragraph{Estimators of aggregated effects} For completeness, we state below a proposition for estimators of the weighted estimands in (ref).
(ref) is analogous to Corollary 1 in miyaji and Corollary 2 of csa, and the proof is omitted.
In this section, we present three simulation experiments that assess the finite-sample performance and double robustness of the estimators in (ref).
The first experiment verifies double robustness in a simple two-period DGP using the never-exposed control group. We compare the main estimators, namely the DR estimators in (ref) and the DML estimators in (ref), to three non-doubly robust alternatives: inverse probability weighting (IPW), standardized IPW (IPWS), and outcome regression (REG), each constructed by exploiting the ratio-of-ATT-parameters structure in (ref). The standardization in IPWS corresponds to the normalization of control weights as done in (ref). This simulation setup follows the one in sazhao.
The second experiment considers staggered exposure to the instrument for both control groups in (ref), in both data settings, and in the presence of an unobserved confounder affecting both treatment and outcome. We focus on the $LATT(e,t)$ estimates varying across cohorts and time, and aggregate these into estimates of the weighted estimand in (ref). Finally, we introduce a group indicator under which $LATT$ differs across the two groups. We estimate group-specific effects via (ref) and aggregate these to horizon-specific effects as in (ref), capturing average dynamic differences across groups for each horizon $l$.
The third experiment follows the same setup as in the second experiment (without the group indicator) but now imposes absorbing treatment (ref). This setting allows us to verify the Bloom-type result in (ref) and compare our estimators to those of csa. As before, a hidden confounder jointly determines treatment and outcome, so the DiD estimator is biased, whereas the IDiD estimator is not.
We compare the DR and DML estimators to the IPW, IPWS and REG estimators through a simulation experiment similar to sazhao. To assess the double robustness properties, we simulate four DGPs with varying degrees of misspecification. The first DGP has both the outcome regression and propensity score correctly specified. The second DGP has the outcome regression correctly specified and the propensity score misspecified. The third DGP has the outcome regression misspecified and the propensity score correctly specified. The fourth and last DGP has both the outcome regression and propensity score misspecified.
We focus on a two-period setup with covariates using a never-exposed control group with panel data. Let
and
and define
Here, $h_{\ell}$ corresponds to correct specification and $h_{n}$ to misspecification. Exposure occurs in period $2$ with probability
Treatment propensities are
and treatment states are generated as
Untreated potential outcomes follow
while treated outcomes are
where $\varepsilon_{1}, \varepsilon_{2} \sim N(0,0.2^{2})$. The observed treatment is given by (ref) and the observed outcome by (ref). Note that the treatment and outcome evolutions depend on the same specification $s(X)$. Also, the true local average treatment effect on the treated for those exposed in period $2$ equals $LATT(2, 2) = \tau$.
In the experiment, we set $\kappa, \tau = 1$, $n = 5000$, $\mathcal{T} = 2$, and repeat the experiment $B=4999$ times. For the repeated cross-sections case, we set $\lambda = 0.5$. For the DML estimators, we use $K=5$ folds for cross-fitting. Outcome regressions are estimated by linear regression, and propensity scores, when required, by logistic regression. The results are shown in (ref) for the panel and repeated cross-sections settings, respectively.
The panel results confirm the expected double robustness and efficiency properties of the proposed estimators. When both nuisance components are correctly specified (DGP1), all estimators are essentially unbiased with coverage close to the nominal level. The IPW estimators exhibit substantially higher variance than the other estimators. When only the outcome regression is correctly specified (DGP2), the DR, DML and regression estimators remain approximately unbiased, whereas the IPW estimators become severely biased and coverage deteriorates. Conversely, when only the propensity score is correctly specified (DGP3), the DR and DML estimators again remain unbiased, while the regression estimator is biased, illustrating the double robustness property. When both nuisance components are misspecified (DGP4), all estimators are biased, as expected.
The repeated cross-section results display the same qualitative double robustness pattern but with larger dispersion. In DGP1-DGP3, the DR and DML estimators remain approximately unbiased with reasonable coverage, again verifying the double robustness property, although their variance is higher than in the panel case. In contrast, the IPW estimators, even in DGP1 under correct specification, exhibit extreme variability, reflected in very large RMSE and asymptotic variance, driven by instability in the weighting scheme. In particular, the non-standardized IPW estimator explodes in variance, as was also the case for the two-period DiD experiment in sazhao, although here further amplified by the ratio structure of the estimators. When the propensity score is misspecified (DGP2), the DR and DML estimators remain relatively robust, while the regression estimator is biased. Finally, when both nuisance components are misspecified (DGP4), all estimators are biased.
Overall, the results highlight the double robustness of the proposed estimators. While this property holds in both data settings, the repeated cross-section estimators are less stable in finite samples due to noisier estimation of the underlying components; larger sample sizes would mitigate this. The DR and DML estimators perform similarly in this experiment, although the DML estimator exhibits slightly higher variance, likely due to cross-fitting. In more complex settings, DML, combined with more flexible nuisance models, is expected to outperform the simpler DR estimator.
In this simulation experiment we simulate a panel with staggered exposure and introduce the time-varying hidden confounder $H_{t}$ determining both the treatment $D_{t}$ and outcome $Y_{t}$. The setup for the simulation is as follows. We simulate a time-invariant covariate and time-varying hidden confounder as $X, H_{t} \sim N(0, 1),$ and exposure cohorts from
with $\beta_{e} \in \{0.1, \ldots, 0.3\}$, $e \in \mathrm{supp}(E)$, equally spaced between $0.1$ and $0.3$\footnote{ For $P(E = e \mid X)$ simulated this way, the generalized propensity (ref) in the case of the never-exposed control group equals $p := P(E = e \mid X, E_{e} + C^{nev} = 1) = \mathrm{expit}([\beta_{e} - \beta_{\infty}]X), $ where $\beta_{\infty}$ is the parameter corresponding to the never-exposed control group; hence, $p / (1 - p) = \exp([\beta_{e} - \beta_{\infty}]X) \sim \mathrm{LogNormal}(0, [\beta_{e} - \beta_{\infty}]^{2}). $ Thus, in order to stabilize the propensity in the simulations, we constrain the parameters $\beta_{e}$ to $\{0.1, \ldots, 0.3\}$ so the maximum difference is $0.2$. }. The potential treatments and outcomes are simulated as:
where $\eta, \varepsilon_{t}, \nu_{t} \sim N(0, 1)$. Note that $H_{t}$ determines both the treatment and outcome cf. (ref). Again, the observed treatment is given by (ref) and the observed outcome by (ref).
For simplicity, we focus on the DR estimators (ref) and evaluate their finite-sample performance under panel and repeated cross-section sampling, respectively, using never-exposed and not-yet-exposed control groups. We set $n = 10{,}000$, $\mathcal{T} = 5$, $\tau_t = 1$ for all $t$, $\delta = 1$, $n_{E} = 5$ and perform $B = 1499$ simulation draws. In the not-yet-exposed experiment the $E = \infty$ is removed from (ref). The average evolution of the treatment $D_t$ and outcome $Y_t$, conditional on the exposure cohort $E$, for a simulated dataset is shown in (ref).
\noindentCohort-time effects. (ref) reports the distribution of $\hat{\tau}^{dr, p}_{e,t}$ and $\hat{\tau}^{dr, rc}_{e,t}$ across all $(e, t)$ pairs. The estimates are centered around the true value in all designs, showing the robustness of the IDiD procedure to hidden confounding. The panel case is most tightly concentrated, while the repeated cross-sections case exhibits higher dispersion; both show occasional large realizations due to the ratio structure. In both designs, the not-yet-exposed estimates, compared to the never-exposed, become increasingly dispersed at longer horizons, reflecting the corresponding reduction in the size of the control group.
\noindentEvent study aggregation. (ref) shows the aggregated effects $\hat{\theta}^{IV}_{es(l)}$. The same pattern emerges: estimates remain centered, but variability increases in repeated cross-sections. Again, for the not-yet-exposed controls, dispersion increases with the horizon due to the shrinking control group.
We conduct a multiplier bootstrap and construct simultaneous confidence bands for $\{\theta^{IV}_{es}(l) \mid l = 0,1,2\}$ in (ref). Let $\hat{C}^{IV}_{es}(l)$ denote the interval for the horizon-$l$ effect. Then, as in csa, the intervals are simultaneous in the sense that\footnote{ The general simultaneous confidence band for $LATT(e, t)$, $t \geq e$, is simultaneous in the sense that $ P (LATT(e, t) \in \hat{C}(e, t) \, \forall (e,t) \in \mathcal{E} \times \{2, 3, \ldots, \mathcal{T}\} : t \geq e) \to 1 - \alpha $; see csa for details. }
We verify this in (ref) with $\alpha = 0.05$. The block “Pooled” reports coverage rates (ref) computed from simultaneous confidence bands (“Simultaneous”) obtained via the multiplier bootstrap and from pointwise confidence intervals (“Pointwise”).
As expected, the simultaneous coverage is close to $0.95$, whereas the pointwise coverage is too low. The repeated cross-section estimates are noisier, which propagates to the simultaneous bands, yielding slightly conservative (too high) coverage.
Group-specific aggregation. Finally, we simulate a new dataset with a binary group indicator $F \sim \mathrm{Bern}(0.5)$ splitting the units into two groups. The groups have different $LATT$s with $F = 1$ being affected less. We set $n=20,000$ so there are two groups of approximately $10,000$. Our aim is to estimate the group-wise estimand $\tau^{dr,p,\Delta}_{e,t}$ as defined in (ref) and aggregate the effects into dynamic horizon-$l$ effects as in (ref). The treated state in (ref) now becomes
where those with $F = 1$ have $LATT(e, t) = 0.5 \tau_{t}$.
(ref) reports the aggregated group-specific horizon effects and their difference, i.e., the aggregated estimates of $\tau^{dr,p,\Delta}_{e,t}$. The estimator tracks the targets well in the panel setting, whereas repeated cross-sections exhibit higher variability, particularly for the difference estimand. Again, the not-yet-exposed case has higher variability because of the shrinking control group.
We again verify the validity of the multiplier bootstrap to construct simultaneous confidence bands for the aggregated effects, here for the aggregated horizon-$l$ effects of $\tau^{dr,p,\Delta}_{e,t}$. The results are reported in block “Group” of (ref). For panel data, the simultaneous coverage is close to $0.95$ for both control groups. In repeated cross-sections, coverage is again slightly conservative (above $0.95$), reflecting the higher variability of this design, compounded by the additional noise from differencing. Pointwise coverage is too low across all simulations, as expected.
We simulate a setting with staggered exposure to the instrument and one-sided compliance in the presence of an unobserved confounder affecting both treatment and outcomes to verify (ref). The setup follows Simulation Experiment 2, except that treatment is an absorbing state, i.e. (ref), and the treatment effects are now heterogeneous across treatment cohorts and time. The potential outcomes of (ref) in the treated state are now given by
where the dependence on $g$ reflects the absorbing treatment; the $ATT(g,t,e)$ parameters increases linearly over post-treatment periods as $1/2,\,1,\,3/2,\ldots$ and are invariant across exposure cohorts $e \in \mathcal{E}$. The potential outcomes in the untreated states are as in (ref).
In the simulation experiment, we set $n = 10,000$, $\mathcal{T}=5$, $\mathrm{supp}(E)=\{\infty,2,3,4,5\}$, use as control group the never-exposed units, $C^{nev}$, and repeat the experiment $B=1499$ times. The implied $ATT(g, t)$ parameters are $ \{\tau_{2,2}, \tau_{2,3}, \tau_{2,4}, \tau_{2,5}, \tau_{3,3}, \tau_{3,4}, \tau_{3,5}, \tau_{4,4}, \tau_{4,5}, \tau_{5,5}\}, $ with true values given by (ref).
For each draw, we estimate $LATT(e,t)$ using our DR estimators in each data setting. We also construct the treatment cohort variable $G = \min\{t \mid D_t = 1\}$, enabling two quantities: (i) the true $ATT(g,t,e)$, computed from the simulated (unobserved) potential outcomes, and (ii) $ATT(g,t)$ estimated via the procedure of csa.
By (ref), for each $(e,t) \in \mathcal{E} \times \{2, 3, \ldots, \mathcal{T}\}$ with $t \ge e$, $\hat{LATT}(e,t)$ should approximately equal $\sum_{g \le t} \hat{ATT}(g,t,e)\,\hat{P}(G = g \mid D_t = 1, E_e = 1),$ cf. (ref). We compute this quantity using Oracle estimates for $\hat{ATT}(g,t,e)$ (the CSA estimators won't work because of the hidden confounder) and empirical conditional probabilities for the weights.
Results are reported in (ref). For horizon $l=t-e=0$, the $LATT(e, t)$ estimates coincide exactly with the exposure-cohort-specific $ATT(g, t)$, so the aggregated $l=0$ effect equals the cohort-weighted $ATT(g, t)$, which equals $1/2$ for each pair $(e, t) \in \{(2, 2), (3, 3), \ldots, (5, 5)\}$. For $l>0$, the estimates of $LATT(e,t)$ are a convex combination of the underlying $ATT(g, t, e)$ parameters. For instance,
use the first two entries in row $(e,t) = (2,4)$ of (ref) (the oracle and weight estimates are not reported in the table).
The weights $\hat{P}(G_g=1 \mid D_4=1, E_2=1)$, $g = 2, 3, 4$, are decreasing, reflecting that treatment is absorbing. Among units first exposed at $t=2$, a large share is treated at $t=2$, a smaller share at $t=3$, and only a small fraction remains untreated until $t=4$. The heterogeneity in $\hat{ATT}(g,4,2)$ captures the different treatment horizons.
Comparing the Oracle and $LATT(e,t)$ columns across both data settings, we see that the columns are approximately equal, verifying the result in (ref). In contrast, the DiD estimates of $ATT(g,t)$ in the CSA column are biased because of the unobserved confounder $H_t$ affecting both the treatment and outcome, cf. (ref). The instrumental variable component of the IDiD estimator addresses this source of bias and remains unbiased, illustrating the robustness of the IDiD design to hidden confounding, and how the IDiD estimators can be leveraged in settings where DiD estimators fail.
This paper develops doubly robust estimands for the $LATT(e, t)$ parameter in the IDiD setting with covariates, covering both panel data and repeated cross-sections, and allowing for never-exposed and not-yet-exposed control groups. We also construct corresponding DR and DML estimators.
Our approach is estimand-based: using the influence function machinery of kennedy, adapted to our setting in (ref), we derive the DR estimands from first principles. The simulation results confirm the validity of the corresponding estimators across both data settings and control groups, and verify the group-difference estimands and the Bloom result under absorbing treatment, linking IDiD with staggered instrument exposure to DiD with staggered treatment. At the same time, the analysis highlights an important distinction between the two data settings: while both admit doubly robust identification and estimation, the repeated-cross-sections case does not inherit the same remainder-term behavior as the panel case because of the additional nuisance components.
More broadly, the paper underscores the value of an estimand-based approach: it replaces reverse-engineering of regression parameters with the specification of interpretable estimands, and enables modular reuse of components across settings, in particular when deriving the corresponding EIFs. It also provides another example of constructing Neyman-orthogonal scores from EIFs (cf. chernozhukovDoubleDebiasedMachine2018a; see also chen2026equivalence) and of two cases in which the cross-fitted estimator coincides with the DML estimator. Practically, the paper delivers both the DR and DML estimators, together with a software implementation, making the proposed methods directly usable in applied work. We hope that the approach and components developed here, together with the work on which they build, serve as a useful guide and toolbox for future work.