EconBase
← Back to paper

Difference-in-Differences with Time-varying Continuous Treatments using Double/Debiased Machine Learning

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

93,979 characters · 9 sections · 72 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Difference-in-Differences with Time-varying Continuous Treatments Using Double/Debiased Machine Learning

\def\spacingset#1{ {#1}} \spacingset{1}

\if11 \fi

abstractWe propose a difference-in-differences (DiD) framework designed for time-varying continuous treatments across multiple periods. Specifically, we estimate the average treatment effect on the treated (ATET) by comparing distinct non-zero treatment intensities. Identification rests on a conditional parallel trends assumption that accounts for observed covariates and past treatment histories. Our approach allows for lagged treatment effects and, in repeated cross-sectional settings, accommodates compositional changes in covariates. We develop kernel-based ATET estimators for both repeated cross-sections and panel data, leveraging the double/debiased machine learning framework to handle potentially high-dimensional covariates and histories. We establish the asymptotic properties of our estimators under mild regularity conditions and demonstrate via simulations that their undersmoothed versions perform well in finite samples. As an empirical illustration, we apply our estimator to assess the effect of the second-dose COVID-19 vaccination rate in Brazil and find that higher vaccination rates reduce COVID-19-related mortality after a lag of several weeks.

{\it Keywords:} Causal inference, Continuous treatment, Difference-in-differences, Machine learning, Parallel trends.

\thispagestyle{empty} \setcounter{page}{1}

\spacingset{1.45}

Introduction

Difference-in-differences (DiD) is a cornerstone method for causal inference, i.e., the evaluation of the impact of a treatment (such as vaccination) on an outcome of interest (such as mortality), in observational studies, provided that outcomes may be observed both before and after the treatment introduction. In such studies, the observed outcomes of treated individuals before and after a treatment typically do not allow researchers to directly infer the treatment effect. This is due to confounding time trends in the treated individuals' counterfactual outcomes that would have occurred without the treatment. The DiD approach tackles this issue based on a control group not receiving the treatment and the so-called parallel trends assumption, imposing that the treated group's unobserved counterfactual outcomes under non-treatment follow the same observed mean outcome trend of the control group. While the canonical DiD setup considers a binary treatment definition (treatment versus no treatment), in many empirical applications, treatments are continuously distributed - e.g., the vaccination rate in a region. Furthermore, the treatment intensity might vary over time and a strict control group with a zero treatment might not be available across all or even any treatment period, as also argued in dechaisemartin2023differenceindifferences, who point to taxes, tariffs, or prices as possible treatment variables with strictly non-zero doses.

To address these limitations, we propose an DiD approach for a time-varying continuous treatment across multiple time periods, enabling the assessment of the average treatment effect on the treated (ATET) when comparing two non-zero treatment doses, such as a higher versus a lower vaccination rate in a region. Identification relies on a conditional parallel trend assumption imposed on the mean potential outcome under the lower dose, given observed covariates and past treatment histories. This assumption requires that, conditional on covariates and previous treatments, the counterfactual outcomes of the treatment group (receiving the higher treatment dose) under the lower dose follow the same mean outcome trend as in the control group (receiving the lower treatment dose). We note that this parallel trend assumption on counterfactual outcomes under a positive treatment dose is more stringent than that on non-treated outcomes, see e.g., the discussion in Fricke2017, which is the trade-off for DiD-based identification of causal effects based on comparisons of non-zero treatment doses.

We suggest novel semiparametric kernel-based DiD estimators for both repeated cross-sections, in which subjects differ across time periods, and panel data, in which the same subjects are repeatedly observed across periods. The estimators adopt the double/debiased machine learning (DML) framework introduced in Chetal2018 to control for covariates and past treatment histories in a data-adaptive manner, which appears attractive in applications where the number of potential control variables is large relative to the sample size. Thus, we estimate the models for the conditional mean outcomes and conditional treatment densities - also known as generalized treatment propensity scores, based on machine learning and use these nuisance parameters as plug-ins in doubly robust (DR) score functions (see for instance Robins+94 and RoRo95, which are tailored to DiD-based ATET estimation).

These DR scores include kernel functions on the continuous treatment as also considered in Kennedyetal2017 and colangelo2023doubledebiasedmachinelearning, when imposing a selection-on-observables-assumption (rather than a parallel trend assumption) for identification. Under an appropriately shrinking bandwidth as the sample size increases, we demonstrate that our estimators satisfy Neyman1959-orthogonality, implying that they are substantially robust to approximation errors in the nuisance parameters under specific regularity conditions - e.g., approximate sparsity when using lasso regression Tibshirani96 for nuisance parameter estimation as considered in Bellonietal2014). Following Chetal2018, we also apply cross-fitting such that the nuisance parameters and scores are not simultaneously estimated in the same parts of the data, to avoid overfitting. To account for the biases introduced by the kernel functions, we use undersmoothed bandwidths, implying that the biases vanish asymptotically. We show that under particular regularity conditions, the DML kernel estimators satisfy asymptotic normality, which also provides us with an asymptotic formula for variance estimation.

We also conduct a simulation study, which points to a compelling finite sample behavior of undersmoothed versions of our estimators under a sample size of several thousand observations. As an empirical application, we use Brazilian municipality-level panel data spanning March 2020 to May 2022 to estimate the effect of the continuously distributed second-dose COVID-19 vaccination rate on mortality. Comparing vaccination rates of 60% versus 40%, we find that higher vaccination coverage does not have a statistically significant effect on COVID-19-related mortality in the first two weeks. However, it leads to a statistically significant reduction in mortality after approximately four weeks, consistent with the expected time lag between infection and death.

Our study adds to a growing literature on the semi- and nonparametric DiD-based identification of causal effects of continuously distributed treatments. This framework avoids potential misspecification errors of classical linear two-way fixed effects models - see, for instance, the discussion in de2020two. callaway2024difference consider a treatment that has a mass at zero (constituting control group), but might also take continuously distributed non-zero doses. This permits identifying the ATET of a specific non-zero treatment dose versus no treatment, under the parallel trend assumption that the treatment group (receiving a non-zero dose) and the control group share the same trend in mean outcomes under non-treatment. However, callaway2024difference do not consider semi- or nonparametric estimation conditional on covariates. zhang2025 adopts a conditional parallel trends assumption in the continuous treatment context, implying that parallel trends are assumed to hold conditional on observed covariates, as previously introduced by abadie2005 for binary treatments. zhang2025 studies the canonical two-period model (with one pre- and one post-treatment period) in both the panel and repeated cross-section data, providing identification, estimation, and inference results using the double/debiased machine learning (DML) framework. In a related work, Hettingeretal2025 combine two parallel trends assumptions to identify the average dose effect on the entire treated group in the panel setting. They adopt a similar framework as Kennedyetal2017 to establish a multiply-robust estimator.

Our paper extends these results in several ways. First, we permit comparisons between two non-zero treatment doses under our stronger parallel trends assumption. Second, we allow for more than two time periods with possibly non-zero treatment doses that may vary over time, such as in a staggered adoption scenario - as e.g.,\ considered for binary treatments in GoodmanBacon2018, SantAnnaZhao2018, sun2021estimating, and borusyak2024revisiting), which motivates conditioning on past treatment histories when assessing the ATET. Third, our approach for repeated cross-sections is capable of appropriately tackling covariate distributions that change within treatment groups over time, while many established approaches require them to be time-invariant, as pointed out in Hong2013, which concerns e.g.,\ the DID methods for binary or discrete treatments in abadie2005 using inverse probability weighting, SantAnnaZhao2018 using DR, or Chang2020 using DML. A notable exception is the DML approach in Zimmert2018, which allows for time-varying covariate distributions within treatment groups, nonetheless applying to binary rather than continuous treatments.

We also distinguish our results from dechaisemartin2023differenceindifferences, which considers treatments with strictly non-zero doses, and discuss effect identification and estimation under parallel trend assumptions related to those considered in the present paper. The authors focus on evaluating an average ATET across the treatment margins of all subjects of changing their treatment dose, which are considered as treatment dose - e.g., what is the average effect of marginally increasing the vaccination rate on mortality among all regions that increase the vaccination rate by possibly different margins. In contrast, our study suggests estimators for assessing the ATET based on comparing two specific treatment doses rather than averaging across many doses - e.g., the effect of a vaccination rate of 40% versus 30%. As a further distinction, dechaisemartin2023differenceindifferences do not consider controlling for covariates when imposing parallel trends, while our estimators permit controlling for covariates in a data-driven manner through machine learning.

The remainder of the paper is organized as follows. Section (ref) motivates and establishes the identification of the instantaneous ATET for continuous treatments in repeated cross-sections. Section (ref) generalizes this framework to time-varying treatments, accommodating lagged effects and dynamic responses. Section (ref) addresses identification in panel data settings. Section (ref) introduces our proposed estimators based on double machine learning (DML) with cross-fitting, while Section (ref) derives their asymptotic properties. Section (ref) assesses finite-sample performance through a simulation study. Section (ref) presents an empirical application evaluating the effectiveness of the second-dose COVID-19 vaccination rate in Brazil on mortality. Section (ref) concludes.

Instantaneous effects in repeated cross-sections

This section discusses the identification of the instantaneous ATET, evaluated shortly after treatment introduction, in the repeated cross-sectional setting. For intuition, we first consider a discrete treatment, denoted by $D$. We first assume that there is a single point in time at which the treatment switches from zero to a non-zero treatment dose for a subset of the population (i.e., the treated subjects), while it remains at zero for another subset which constitutes the control group. Accordingly, let $T$ denote the time period, where $T=0$ is the baseline period prior to treatment assignment, while $T>0$ is a post-treatment period after the introduction of a positive treatment dose among the treated. The treatment dose may differ across subjects and for now it is assumed to remain constant across post-treatment periods. Let $Y_t$ denote the observed outcome at period $T=t$, while $Y_t(d)$ is the potential outcome at period $t$ when hypothetically setting the treatment to $D=d$, see e.g.,\ Neyman23 and Rubin74 for further discussion on the potential outcome framework. Throughout the present work, we impose the stable unit treatment value assumption (SUTVA), meaning that the potential outcomes of a subject are not affected by the treatment of others - see, for instance, the discussion in Rubin80 and Cox58. This permits defining the average treatment effect on the treated (ATET) of treatment dose $d$ in period $t$ as

align[align omitted — 67 chars of source]

Furthermore, let $X$ denote the observed covariates. We impose the following identifying assumptions, which are closely related to those in Lechner2010.

assumption{\bf (Conditional parallel trends):}\\ $E[Y_t(0)-Y_0(0)|D=d,T=t,X]=E[Y_t(0)-Y_0(0)|D=0,T=t,X]$ for $d$ in the support of $D$ and $t>0$ in the support of $T$.
assumption{\bf (No anticipation):}\\ $E[Y_0(d)-Y_0(0)|D=d,T=0,X]=0$ for $d > 0$ in the support of $D$.
assumption{\bf (Common support):}\\ $\Pr(D = d,\, T = t \mid X,\,(D, T) \in \{(d', t'), (d, t)\}) < 1 \quad \text{for } d > 0 \text{ in the support of } D,\; t > 0 \text{ in the support of } T,\; \text{and } (d', t') \in \{(d, 0), (0, t), (0, 0)\}.$
assumption{\bf (Exogenous covariates):}\\ $X(d)=X$ for all $d$ in the support of $D$.

Assumption (ref) imposes parallel trends in the mean potential outcome under non-treatment across treated and control groups, conditional on covariates $X$. Assumption (ref) imposes that conditional on $X$, the ATET must be zero in the pre-treatment period $T=0$, which on average rules out non-zero anticipation effects among the treated in period $T=0$. Assumption (ref) requires that for any subject receiving treatment dose $d$ in a post-treatment period $t > 0$, there exist in terms of $X$ comparable treated subjects in the baseline period, non-treated subjects in the post-treatment period, and non-treated subjects in the baseline period. Lastly, Assumption (ref) imposes that the covariates are not causally affected by the treatment. We note that this may be a relevant restriction if the researchers include post-treatment variables - as it is frequently the case in repeated cross-sections (see, for instance, the discussion in caetano2022difference.

Under these assumptions, the conditional ATET of treatment dose $d>0$ in period $t>0$ given $X$ is identified:

eqnarray[eqnarray omitted — 470 chars of source]

The first equality in (ref) follows from subtracting and adding $E[Y_0(0)|D=d,X]$, and the second from Assumption (ref), implying that $E[Y_0(0)|D=d,T-0,X]=E[Y_0(d)|D=d,T=0,X]$. The third equality follows from the conditional parallel trends assumption (Assumption (ref)). The fourth equality follows from the fact that $Y_T=Y_t(d)$ conditional on $D=d$ and $T=t$, which is known as the observational rule. By Assumption (ref), the effect of the treatment on the outcome $Y_t$ is not mediated by $X$, such that the conditional ATET in expression (ref) captures the total causal effect among the treated given $X$, rather than only a partial effect not mediated by $X$. Averaging the conditional ATET over the distribution of $X$ among the treated in a post-treatment period $T=t$ yields the ATET in that period:

eqnarray[eqnarray omitted — 117 chars of source]

where we use the short hand notation $\mu_d(t,x)=E[Y_T|D=d,T=t,X=x]$ for the conditional mean outcome given the treatment, the time period, and the covariates.

If the treatment $d$ consisted of discrete treatment doses, the ATET can alternatively be assessed based on the following doubly robust (DR) expression, which is in analogy to Zimmert2018 for binary treatments:

align[align omitted — 496 chars of source]

where $\Pi=\Pr(D=d,T=t)$ denotes the unconditional probability of receiving treatment dose $d$ at period $t>0$, $\rho_{d,t}(X)=\Pr(D=d,T=t|X)$ denotes the conditional probability of a specific treatment group-period-combination $D=d,T=t$ given $X$, and $I\{\cdot\}$ denotes the indicator function.

To see that the DR expression (ref) indeed corresponds to the ATET in equation (ref), note that the first term on the right of equation (ref), by the law of iterated expectation, corresponds to the ATET:

align*[align* omitted — 249 chars of source]

Moreover, again by the law of iterated expectation, we have for any $d'$ and $t'$ in the support of $D$ and $T$,

align*[align* omitted — 457 chars of source]

Therefore, the remaining terms in equation (ref) are all equal to zero. This demonstrates that (ref) and (ref) are equivalent and both yield the ATET under our identifying assumptions.

To adapt the previous results based on discrete treatments into a continuous treatment, the indicator functions for non-zero treatment doses $d$ need to be replaced by smooth weighting functions with a bandwidth approaching zero to obtain identification. Thus, we denote by $\omega(D;d,h)$ a weighting function that depends on the distance between $D$ and the non-zero reference value $d$ as well as a non-negative bandwidth parameter $h$: $\omega(D;d,h)=\frac{1}{h}\mathcal{K}\left(\frac{D-d}{h}\right)$, with $\mathcal{K}$ denoting a well-behaved kernel function. The closer $h$ is to zero, the less weight is given to realizations of $D$ that are further away from $d$. Moreover, we note that the previous (conditional) probabilities $\Pi_{d,t}$ and $\rho_{d,t}(X)$ for $d>0$ and $t\geq0$ should be replaced by (conditional) density functions if the treatment is continuous. Note that the conditional treatment density is also referred to as generalized propensity score in the literature, e.g., ImaivanDyk2004 and HiranoImbens2005. In fact, as observed in fan1996estimation the conditional densities can also be expressed as the limit of the conditional expectations of kernel functions: $\rho_{d,t}(X)=\lim_{h \rightarrow 0} E[\omega(D;d,h) \cdot I\{T=t\}|X]$. Similarly, $\Pi_{d,t}=\lim_{h \rightarrow 0} E[\omega(D;d,h) \cdot I\{T=t\}]$. Applying these modifications to DR expression (ref) yields

align[align omitted — 550 chars of source]

where $\epsilon_{d,t,x}=(Y_T-\mu_d(t,x))$ denotes a regression residual. It is worth mentioning that in a related study, zhang2025 also provides an DR expression for $\Delta_{d,t}$ with continuous treatment doses $d$ that in the repeated cross-sections setting. However, in contrast to our results, his approach is not applicable to settings in which the distribution of covariates $X$ changes over time.

Section (ref) extends the analysis to the identification of ATETs in later periods following treatment, allowing for time lags and potentially time-varying treatment doses.

Time-varying treatment doses in repeated cross-sections

In this section, we extend our DiD framework to accommodate time-varying treatment doses and multiple treatment (and, later, outcome) periods. For notation, we denote the treatment dose in period $t$ by $D_t$, and we use $\mathbf{D}_{t}=\{D_0,...,D_t\}$ to denote the sequence of treatment doses up to and including period $t$. We replace Assumption (ref), which imposes parallel trends on the mean potential outcome under non-treatment, with Assumption (ref), which imposes parallel trends in mean potential outcomes between non-zero treatment doses.

assumption{\bf (Conditional parallel trends with non-zero treatment doses):}\\ $E[Y_t(d_t')-Y_{t-1}(d_{t-1}'')|D_t=d_t,T=t,\mathbf{D}_{t-1}=d_{t-1}'',X] \\ = E[Y_t(d_t')-Y_{t-1}(d_{t-1}'')|D_t=d_t',T=t,\mathbf{D}_{t-1}=d_{t-1}'',X]$ for $d_t>d_t'\geq 0$ in the support of $D_t$, $\mathbf{d}_{t-1}''$ in the support of $\mathbf{D}_{t-1}$, and $t>0$ in the support of $T$.

Assumption (ref) states that for two groups that receive a higher treatment dose $d_t$ (treatment group) and a lower treatment dose $d_t'$ (control group), respectively, in period $t$, and both receive treatment dose $d_{t-1}''$ in period $t-1$, the trend in mean potential outcomes when moving from dose $d_{t-1}''$ in period $t-1$ to dose $d_t'$ in period $t$ is the same across groups, conditional on $X$. For the special case where $d_t' = d_{t-1}''$, see for example dechaisemartin2023differenceindifferences and dechaisemartin2025treatmenteffectestimationcomplexdesigns, this implies that if no group changed the treatment dose between periods $t-1$ and $t$, then the treatment and control groups would experience the same trend in their mean outcomes. Moreover, we note that Assumption (ref) is stronger than Assumption (ref) because it imposes parallel trends with respect to non-zero treatment doses.

This generally imposes restrictions on treatment effect heterogeneity; see the discussions in Fricke2017 and deChaisemartin2022did. For instance, if Assumption (ref) is plausibly satisfied, then Assumption (ref) will additionally hold only if the change in the average treatment effects of receiving dose $d_t'$ (relative to no treatment, $d_t=0$) in period $t$ and receiving dose $d_{t-1}''$ (relative to $d_{t-1}=0$) in period $t-1$ is identical across the treatment group receiving $d_t$ and the control group receiving $d_t'$ in period $t$, conditional on $\mathbf{D}_{t-1}=d_{t-1}''$ and $X$. Two special cases satisfying this condition are that (i) average treatment effects are homogeneous across groups in each period, or (ii) average treatment effects are time-invariant within groups (though levels may differ across groups).\footnote{To show this formally, consider the conditional expectation in the first line of Assumption (ref), $E[Y_t(d_t')-Y_{t-1}(d_{t-1}'') \mid D_t=d_t,T=t,\mathbf{D}_{t-1}=d_{t-1}'',X]$, and add and subtract $Y_t(0)$ as well as $Y_{t-1}(0)$, which yields \[ \underbrace{E[Y_t(0) - Y_{t-1}(0) \mid D_t=d_t, \mathbf{D}_{t-1}=d_{t-1}'', X]}_{=0\textrm{ under Assumption \ref{ass1}}} + \tau_{d_t,d_{t-1}'', X}^t(d_t') - \tau_{d_t,d_{t-1}'', X}^{t-1}(d_t''), \] where $\tau_{d_t,d_{t-1}'', X}^t(d_t') = E[Y_t(d_t') - Y_t(0) \mid D_t=d_t, \mathbf{D}_{t-1}=d_{t-1}'', X]$ denotes the conditional average treatment effect of moving from $0$ to $d_t'$ at time $t$. Conditional on Assumption (ref), Assumption (ref) holds if and only if the change in the average treatment effects over time is identical across treatment groups $d_t$ and $d_t'$: \[ (\tau_{d_t,d_{t-1}'', X}^t(d_t') - \tau_{d_t,d_{t-1}'', X}^{t-1}(d_t'')) - (\tau_{d_t',d_{t-1}'', X}^t(d_t') - \tau_{d_t',d_{t-1}'', X}^{t-1}(d_t'')) = 0. \] This is satisfied in two special cases: (i) effect homogeneity across groups within each period, i.e.\ $\tau_{d_t,d_{t-1}'', X}^t(d_t')=\tau_{d_t',d_{t-1}'', X}^t(d_t')$ and $\tau_{d_t,d_{t-1}'', X}^{t-1}(d_t'')=\tau_{d_t',d_{t-1}'', X}^{t-1}(d_t'')$; or (ii) time-invariant effects within groups, i.e.\ $\tau_{d_t,d_{t-1}'', X}^t(d_t')=\tau_{d_t,d_{t-1}'', X}^{t-1}(d_t'')$ and $\tau_{d_t',d_{t-1}'', X}^t(d_t')= \tau_{d_t',d_{t-1}'', X}^{t-1}(d_t'')$.}

Furthermore, we replace the previous Assumption (ref) (ruling out anticipation) by the following condition:

assumption{\bf (No anticipation):}\\ $E[Y_{t-1}(d_t)-Y_{t-1}(d_t')|D=d_t,T=t-1,\mathbf{D}_{t-1},X]=0$ for all $d_t\neq d_t'$ in the support of $D_t$.

Lastly, the previous common support condition given in Assumption (ref) is replaced by the following assumption, where $f(A|B)$ denotes the conditional density of $A$ given $B$:

assumption{\bf (Common support):}\\ $f(D_t = d_t,\, T = t \mid \mathbf{D}_{t-1}, X,\, (D_t, T) \in \{(d^*_t, t'), (d_t, t)\}) < 1 \quad \text{for } d_t > d'_t \text{ in the support of } D_t,\; t > 0 \text{ in the support of } T,\; \text{and } (d^*_t, t') \in \{(d_t, t-1), (d'_t, t), (d'_t, t-1)\}.$

We note that our modified set of Assumptions (ref) to (ref) does not rule out that treatment doses in earlier periods affect outcomes in later periods. Thus, the treatment sequence $\mathbf{D}_{t-1}$ may influence the outcome in period $t$. In this case, we may nevertheless identify the ATET of treatment dose $d_t$ versus $d_t'$ in period $t$ net of any direct effects of the previous history of treatment doses on the outcome in period $t$, formally defined as

align[align omitted — 87 chars of source]

It is worth noticing that under Assumptions (ref) to (ref), the conditional ATET of treatment dose $d_t$ versus $d'_t$ in period $t$ given covariates $X$ and the previous treatment history $\mathbf{D}_{t-1}$ is identified analogously as expression (ref):

eqnarray[eqnarray omitted — 307 chars of source]

Integrating the conditional ATET over the distribution of $X$ and $\mathbf{D}_{t-1}$ yields the ATET at a treatment dose $D_t = d_t$ in that period:

eqnarray[eqnarray omitted — 230 chars of source]

where $\mu_{d_t}(t',\mathbf{D}_{t-1},x)=E[Y_T|D_t=d_t,T=t',\mathbf{D}_{t-1}=\mathbf{D}_{t-1},X=x]$ is the conditional mean outcome in period $t'$ given the treatment received in period $t$, the previous treatment history up to period $t-1$, and the covariates. We note that for the case that $d_t'=d_{t-1}''$, the ATET corresponds to a causal effect of switching from the previous treatment dose $d_t'$ to a higher treatment dose $d_t$ among switchers. As a comparison, dechaisemartin2023differenceindifferences consider average effects among switchers across different treatment margins, rather than exclusively switching to dose $d_t$.

As an alternative to equation (ref), the ATET may be formalized by the following, equivalent DR expression:

align[align omitted — 890 chars of source]

where $\epsilon_{{d_t},t,\mathbf{D}_{t-1},x}=(Y_T-\mu_{d_t}(t,\mathbf{D}_{t-1},x))$ denotes a regression residual.

As a further modification in our DiD framework, we might also be interested in the ATET of a lagged treatment dose on outcomes in later periods, rather than only evaluating the the instantaneous impact of a treatment dose on an outcome in the same period. Hence, let us denote by $s\geq 1$ the time lag between the treatment and the outcome period, which permits defining the ATET of a treatment dose in period $t-s$ on an outcome in period $t$:

align[align omitted — 107 chars of source]

with time lag $s\geq 1$. We note that the ATET $\Delta_{d_{t-s},d'_{t-s}, t}$ comprises both the direct effect effect of the treatment dose in period $t-s$ on the outcome in $t$, as well as the indirect (or mediated) effect of the treatment dose in period $t-s$ on the treatments in later periods, that in turn also affect the outcome in period $t-s$. $\Delta_{d_{t-s},d'_{t-s}, t}$ is identified when replacing Assumption (ref) by the following parallel trend assumption:

assumption{\bf (Conditional parallel trends with time lags):}\\ $E[Y_t(d_{t-s}')-Y_{t-s-1}(d_{t-s-1}'')|D_{t-s}=d_{t-s},T=t, \mathbf{D}_{t-s-1}=\mathbf{d}_{t-s-1}'',X] \\ = E[Y_t(d_{t-s}')-Y_{t-s-1}(d_{t-s-1}'')|D_{t-s}=d_{t-s}', T=t, \mathbf{D}_{t-s-1}=\mathbf{d}_{t-s-1}'',X]$ for $d_{t-s}>d_{t-s}'\geq 0$ in the support of $D_{t-s}$, $\mathbf{d}_{t-s-1}''\geq \mathbf{0}$ in the support of $\mathbf{D}_{t-s-1}$, $t>1$ in the support of $T$, and $1 \leq s < t$.

Assumption (ref) is stronger than Assumption (ref), because it requires that changes in the treatment dose after period $t-s$ up to period $t$ do not evolve endogenously. This means that unobserved confounders jointly affecting the treatment doses in period $t-s$ and subsequent $s$ periods are ruled out. Otherwise, the parallel trend assumption would generally fail, as the outcome (trend) is typically also a function of the treatments after period $t-s$. Therefore, the presence of unobserved confounders jointly affecting the treatment doses in period $t-s$ and subsequent periods would ultimately imply confounding of treatment $D_{t-s}$ and outcome $Y_t$. In contrast, Assumption (ref) does not put restrictions on confounders of the treatment dose over time, with the caveat that the ATET may only be assessed on outcomes in the same period as the treatment dose to be evaluated, but not on outcomes in later periods.

To better see the implications of Assumption (ref) in terms of ruling out confounders that jointly affect treatments in different periods, consider the following example. Suppose the outcome in time period $T$ is a (possibly unknown) function, denoted by $\mathcal{F}_Y$, of the treatments in periods 1 and 2, $D_1$ and $D_2$, covariates $X$, and time $T$ (pre-treatment histories are omitted for simplicity), together with an additively separable, time-constant function $\mathcal{F}_U$ of time-invariant unobservables: \[ Y_T = \mathcal{F}_Y(D_1, D_2, X, T) + \mathcal{F}_U(U). \] Further, assume that the first treatment is a function of $X$ and $U$, \[ D_1 = \mathcal{F}_{D_1}(X, U), \] and that the second treatment depends on the first treatment and $X$, \[ D_2 = \mathcal{F}_{D_2}(D_1, X). \] We may add random, time-varying, and additively separable error terms to the models for $Y_T$, $D_1$, and $D_2$, but these are omitted here for simplicity.

Assume we are interested in the ATET of the treatment in period $T=1$ on the outcome in period $T=2$. Differencing average potential outcomes under treatment value $d_1'$ between period $T=2$ and the pre-treatment period $T=0$, conditional on $X=x$, eliminates the time-constant component $\mathcal{F}_U(U)$ and yields

align[align omitted — 140 chars of source]

where $D_2(d_1')$ denotes the potential second-period treatment under a first-period dose $d_1'$.

Under the maintained model for $D_2$ (i.e., $D_2 = \mathcal{F}_{D_2}(D_1, X)$), the potential second-period treatment is $D_2(d_1') = \mathcal{F}_{D_2}(d_1', X)$ and hence does not depend on $U$. Consequently, $U$, while affecting $D_1$, does not affect the quantity on the right-hand side of (ref); therefore, Assumption (ref) is satisfied in this setup.

By contrast, if the second treatment also depends on $U$, say \[ D_2 = \mathcal{F}_{D_2}(D_1, X, U), \] then the potential second-period treatment becomes $D_2(d_1') = \mathcal{F}_{D_2}(d_1', X, U)$, and $U$ enters the right-hand side of the differenced expression. In that case

align[align omitted — 179 chars of source]

so $U$ affects the mean potential outcome difference via the second treatment and also through the first treatment (since $D_1$ depends on $U$). Hence $U$ acts as a confounder for the effect of the first-period treatment on the differenced outcome, and Assumption (ref) is violated. In summary, Assumption (ref) rules out precisely such scenarios in which unobservables affect both the early-period treatment $D_{t-s}$ and subsequent treatments that also affect the outcome.

To identify the ATET under Assumptions (ref), (ref), and (ref) (when considering $t-s$ rather than $t$ as treatment period), as well as Assumption (ref) using a DR expression, we need to appropriately modify equation (ref) to include time lags when defining the pre- and post-treatment periods of interest:

align[align omitted — 1,058 chars of source]

Appendix (ref) demonstrates that DR expressions (ref), (ref), and (ref) satisfy Neyman orthogonality as the kernel bandwidth goes to zero. Therefore, these expressions can be employed to construct DML estimators under certain regularity conditions outlined in Chetal2018.

We note that the identification of $\Delta_{d_{t-s},d'_{t-s}, t}$ under parallel trends as imposed in Assumption (ref) is distinct from the scenario considered in de2024difference, where the counterfactual is defined in terms of a fixed treatment dose $d'$ over all post-treatment periods $t-s$ to $t$, that is, conditional on being an non-switcher in terms of treatment doses. This scenario allows considering the following ATET:

align[align omitted — 213 chars of source]

where $Y_t(\mathbf{d'}_{t-s:t})$ denotes the lagged post-treatment counterfactual outcome under the sequence of fixed treatment doses $\mathbf{d'}_{t-s:t} = (d'_{t-s}, \ldots, d'_t)$. Identification of the ATET in (ref) requires yet another parallel trend assumption, different from Assumption (ref) and more closely related to the one considered in de2024difference:

assumption{\bf (Conditional parallel trends under a fixed treatment sequence):}\\ $E\left[\,Y_t(d'_{t-s}, \ldots, d'_t) - Y_{t-s-1}(d''_{t-s-1}) \mid D_{t-s}=d_{t-s}, T=t, \mathbf{D}_{t-s-1}=\mathbf{d}''_{t-s-1}, X \right] \\ = E\left[\,Y_t(d'_{t-s}, \ldots, d'_t) - Y_{t-s-1}(d''_{t-s-1}) \mid D_{t-s}=d'_{t-s}, T=t, \mathbf{D}_{t-s-1}=\mathbf{d}''_{t-s-1}, X \right]$ for $d_{t-s}>(d_{t-s}',...,d'_t)\geq 0$ in the support of $(D_{t-s},...,D_t)$, $\mathbf{d}_{t-s-1}''\geq \mathbf{0}$ in the support of $\mathbf{D}_{t-s-1}$, $t>1$ in the support of $T$, and $1 \leq s < t$.

Assumption (ref) requires that, conditional on covariates $X$ and the pre-period treatment history $\mathbf{D}_{t-s-1}$, the mean trend in potential outcomes under the sequence of fixed treatment doses $\mathbf{d}'_{t-s:t}$ is the same for units whose realized treatment at time $t-s$ equals $d_{t-s}$ and for those with $d'_{t-s}$. In contrast to Assumption (ref), Assumption (ref) allows for unobserved factors that jointly affect treatments in different time periods (provided that these factors do not also affect the differences in mean potential outcomes between the pre- and post-treatment periods). In the example given in equation (ref), which entails the violation of Assumption (ref), Assumption (ref) is satisfied because, conditional on both the first- and second-period treatments, time-invariant unobserved confounders are differenced out from the mean potential outcome equations:

align[align omitted — 141 chars of source]

In addition to a modification of the parallel trend assumption, the nonparametric identification of the ATET in (ref) requires strengthening the previous common support condition to hold with respect to the conditional density for the sequence:

assumption{\bf (Common support under a fixed treatment sequence):}\\ $f(D_{t-s} = d_{t-s},\, T = t \mid \mathbf{D}_{t-s-1}, X,\, (D_{t-s}, T) \in \{(d^*_t, t'), (d_{t-s}, t)\}) < 1 \\ \text{for } d_{t-s} > d'_{t-s} \text{ in the support of } D_{t-s},\; t > 0 \text{ in the support of } T,\; \text{and } (d^*_t, t') \in \{(d_{t-s}, t-s-1), (\mathbf{d}'_{t-s:t}, t), (\mathbf{d}'_{t-s:t}, t-s-1)\}.$

This common support condition becomes more stringent as the number of post-treatment periods increases.

Under Assumptions (ref), (ref), (ref), and (ref), the ATET $\Delta_{d_{t-s},\mathbf{d'}_{t-s:t}, t}$ is identified by appropriately modifying DR expression (ref). To account for a fixed treatment sequence $\mathbf{d}'_{t-s:t}$, we define the product kernel

align[align omitted — 187 chars of source]

where $\mathbf{h} = (h_0, \ldots, h_s)$ is a vector of bandwidths and $\mathcal{K}$ is a univariate kernel. Furthermore, we denote by $\mu_{\mathbf{d}'_{t-s:t}}(t, \mathbf{D}_{t-s-1}, X) = E[Y_t \mid \mathbf{D}'_{t-s:t}=\mathbf{d}'_{t-s:t}, \mathbf{D}_{t-s-1}, X]$ the conditional mean outcome under the treatment sequence, and by $\epsilon_{\mathbf{d}'_{t-s:t},t,\mathbf{D}_{t-s-1},X} = Y_t - \mu_{\mathbf{d}'_{t-s:t}}(t, \mathbf{D}_{t-s-1}, X)$ the corresponding residual. Finally, the conditional density of the treatment sequence is given by $ \rho_{\mathbf{d}'_{t-s:t},t}(\mathbf{D}_{t-s-1},X)= f(\mathbf{D}'_{t-s:t}=\mathbf{d}'_{t-s:t} \mid \mathbf{D}_{t-s-1}, X, T=t)$. Then, the following DR expression identifies the ATET defined in (ref):

align[align omitted — 1,202 chars of source]

We note that in empircal applications, the use of product kernels in $(s+1)$ dimensions (related to the number of post-treatment periods) introduces the usual curse of dimensionality, since the effective sample size scales with $\prod_{r=0}^s h_r$.

Identification in panel data

This section discusses identification in panel data under the identifying assumptions proposed in the previous section. We first consider the identification of $\Delta_{d_t, d_t', t}$, i.e., the ATET of treatment dose $d_t$ versus $d_t'$ in period $t$ (net of direct effects of previous treatment doses) as defined in equation (ref). As panel data permits taking outcome differences within subjects over time, the identification result in equation (ref) simplifies as follows:

eqnarray[eqnarray omitted — 226 chars of source]

It is worth noticing that dropping $T=t$ from the conditioning follows from the fact that, in panel data, the same subjects are observed in periods $t$ and $t-1$, such that distributions of $\mathbf{D}_{t-1}$ and $X$ are constant across periods by design. To ease notation, we denote the conditional mean differences in outcomes by $m_{d_t}(t,\mathbf{D}_{t-1},X)=E[Y_{t}-Y_{t-1}|D_t=d_t,\mathbf{D}_{t-1},X]$. Moreover, we use $P_{d_t}=f(D_t=d_t)$ and $p_{d_t}(\mathbf{D}_{t-1},X)=f(D_t=d_t|\mathbf{D}_{t-1},X)$ to denote the (conditional) density functions of treatment $D_t=d_t$. Specifically, it can be shown that $P_{d_t}=\lim_{h \rightarrow 0} E[\omega(D;d_t,h)]$ and $p_{d_t}(\mathbf{D}_{t-1},X)$$=\lim_{h \rightarrow 0} E[\omega(D;d_t,h)|\mathbf{D}_{t-1},X]$ hold. We propose an DR expression that is a straightforward modification of the one provided in equation (3.8) of zhang2024. The key differences are that the treatment dose in the control group is non-zero, thus involving kernel functions with respect to both $d_t'$ and $d_t$, and that the conditioning set additionally includes the treatment history $\mathbf{D}_{t-1}$ due to the multiple-period setup. Our proposed DR expression for equation (ref) for ATET identification is given by:

align[align omitted — 380 chars of source]

To see the equivalence of equations (ref) and (ref), we note that by the law of iterated expectation,

align*[align* omitted — 283 chars of source]

which corresponds to the ATET of interest. In addition, again by the law of iterated expectations, we have that

align*[align* omitted — 496 chars of source]

which demonstrates that both equations (ref) and (ref) identify the ATET under Assumptions (ref) through (ref) (without conditioning on $T=t$).

Next, we consider evaluating $\Delta_{d_{t-s},d'_{t-s}, t}$, the ATET of receiving the treatment dose $d$ versus $d'$ in period $t-s$ on a later outcome in period $t$, as defined in equation (ref). Under Assumptions (ref), (ref), (ref), and (ref), a DR expression in the panel setting with time lags, analogous to (ref) for repeated cross sections, is given by the following expression:

align[align omitted — 436 chars of source]

where $m_{d_{t-s}'}(t,\mathbf{D}_{t-s-1},X)=E[Y_{t}-Y_{t-s-1}|D_{t-s}=d_{t-s},\mathbf{D}_{t-s-1},X]$. Appendix (ref) demonstrates that DR expressions (ref) and (ref) satisfy Neyman orthogonality.

Finally, when assessing the ATET $\Delta_{d_{t-s},\mathbf{d'}_{t-s:t}, t}$ under the fixed treatment sequence defined in equation (ref), equation (ref) needs to be modified by making use of the product kernel defined in equation (ref), the conditional mean outcome differences under fixed treatment sequences, $m_{\mathbf{d}_{t-s:t}'}(t,\mathbf{D}_{t-1},X)=E[Y_{t}-Y_{t-1} \mid \mathbf{D}_{t-s:t}'=\mathbf{d}_{t-s:t}',\mathbf{D}_{t-1},X]$, and the corresponding conditional densities of fixed treatment sequences, $p_{\mathbf{d}_{t-s:t}'}(\mathbf{D}_{t-1},X)=f(\mathbf{D}_{t-s:t}'=\mathbf{d}_{t-s:t}' \mid \mathbf{D}_{t-1},X)$. Under Assumptions (ref), (ref), (ref), and (ref), the following expression identifies the ATET in panel data:

align[align omitted — 512 chars of source]

Estimation

In this section, we propose DiD estimators based on the DML framework as discussed in Chetal2018, using the sample analogs of the DR identification results in sections (ref) and (ref). We first consider estimating the ATET $\Delta_{d_t,d_t', t}$ as defined in equation (ref) with the sample version of equation (ref). Thus, let $\mathcal{W} = \{W_i|1\leq i \leq n\}$ with $W_i = (Y_{i,T}, D_{i,t}, \mathbf{D}_{i,t-1} X_i, T_i)$ for $i=1,\ldots, n$ denote the set of observations in an i.i.d.\ sample of size $n$. Moreover, let

align[align omitted — 295 chars of source]

denote the set of nuisance parameters, i.e., the conditional mean outcomes and treatment densities. Their respective estimates are denoted by

align[align omitted — 326 chars of source]

The following algorithm formally sumarizes the estimation procedure for $\Delta_{d_t,d_t',t}$: \newline DML algorithm:

enumerate• Split $\mathcal{W}$ in $ K $ subsamples. For each subsample $k$, let $n_k$ denote its size, $\mathcal{W}_k$ the set of observations in the sample and $\mathcal{W}_k^{C}$ the complement set of all observations not in $\mathcal{W}_k$. • For each $k$, use $\mathcal{W}_k^{C}$ to estimate the model parameters of the nuisance parameters $\eta$ by machine learning in order to predict these nuisance parameters in $\mathcal{W}_k$, where the predictions are denoted by $\hat{\eta}^k$. • Stack the fold-specific estimates of the nuisance parameters $\hat{\eta}^1,...,\hat{\eta}^K$ to generate a matrix of nuisance estimates $\hat{\eta}$ for the entire sample: $\hat{\eta} = \begin{pmatrix} \hat{\eta}^1 \\ \vdots \\ \hat{\eta}^K \end{pmatrix}$ • Plug the nuisance parameters into the normalized sample analog of equation (ref) to obtain an estimate of the ATET using bandwidth $h$, denoted by $\hat\Delta_{d_t,d_t', t}^h$: \begin{flalign} \hat\Delta_{d_t,d_t',t}^h& = \sum_{i=1}^{n} \left[\biggl(\sum_{j=1}^n \omega(D_{j,t};d_t,h) \cdot I\{T_j=t\}\biggr)^{-1} \cdot (\omega(D_{i,t};d_t, h) \cdot I\{T_i=t\} \right. \notag \\ &\cdot \left[Y_{i,T}-\hat\mu_{d_{t}}(t-1,\mathbf{D}_{i,t-1},X_i)-\hat\mu_{d'_t}(t,\mathbf{D}_{i,t-1},X_i)+\hat\mu_{d'_t}(t-1,\mathbf{D}_{i,t-1},X_i)\right])\notag \\ &- \frac{ \omega(D_{i,t};d_t, h) \cdot I\{T_i=t-1\}\cdot \hat\rho_{d_t,t}(\mathbf{D}_{i,t-1},X_i)\cdot \hat\epsilon_{i,d_t,t-1,\mathbf{D}_{i,t-1},X_i}/ \hat\rho_{d_t,t-1}(\mathbf{D}_{i,t-1},X)}{ (\sum_{j=1}^n \omega(D_{j,t};d_t, h) \cdot I\{T_j=t-1\})\cdot \hat\rho_{d_t,t}(\mathbf{D}_{i,t-1},X_i)/ \hat\rho_{d_t,t-1}(\mathbf{D}_{i,t-1},X)}\notag \\ &- \left(\frac{\omega(D_{i,t};d'_t, h)\cdot I\{T_i=t\}\cdot \hat\rho_{d_t,t}(\mathbf{D}_{i,t-1},X)\cdot \hat\epsilon_{i,d'_t,t,\mathbf{D}_{i,t-1},X_i}/\hat\rho_{d'_t,t}(\mathbf{D}_{i,t-1},X_i)}{(\sum_{j=1}^n \omega(D_{j,t};d'_t, h)\cdot I\{T_j=t\})\cdot \hat\rho_{d_t,t}(\mathbf{D}_{i,t-1},X_i)/\hat\rho_{d'_t,t}(\mathbf{D}_{i,t-1},X_i)} \right. \notag \\ &- \left.\left.\frac{\omega(D_{i,t};d'_t, h)\cdot I\{T_i=t-1\}\cdot \hat\rho_{d_t,t}(\mathbf{D}_{i,t-1},X)\cdot \hat\epsilon_{i,d'_t,t-1,\mathbf{D}_{i,t-1},X_i}/\hat\rho_{d_t',t-1}(\mathbf{D}_{i,t-1},X_i)}{(\sum_{j=1}^n \omega(D_{j,t};d'_t, h)\cdot I\{T_j=t-1\})\cdot \hat\rho_{d_t,t}(\mathbf{D}_{i,t-1},X_i)/\hat\rho_{d_t',t-1}(\mathbf{D}_{i,t-1},X_i)}\right) \right],\notag \end{flalign} where $\hat\epsilon_{i,{d_t},t,\mathbf{D}_{i,t-1},X_i}=(Y_{i,T}-\hat\mu_{d_t}(t,\mathbf{D}_{i,t-1},X_i))$ denotes the estimated regression residual. \newline

For estimation in panel data, we assume an i.i.d. sample of $n$ subjects with $W_i = (Y_{i,t}, Y_{i,t-1}, D_{i,t}, \mathbf{D}_{i,t-1} X_i)$ for $i=1,\ldots, n$. Then, the nuisance parameter vector corresponds to

align*[align* omitted — 125 chars of source]

and its estimate to

align*[align* omitted — 142 chars of source]

ATET estimation proceeds analogously as outlined in the algorithm above, although using the sample analog of equation (ref) in step 4 of ATET estimation:

align[align omitted — 551 chars of source]

Furthermore, the algorithm can be easily modified to estimate $\Delta_{d_{t-s},d'_{t-s}, t}=E[Y_t(d_{t-s})-Y_t(d'_{t-s})|D_{t-s}=d_{t-s}]$, the ATET of a lagged treatment dose on a later outcome as defined in (ref), by using of the sample analogs of DR expressions (ref) or (ref) in the case of repeated cross-sections or panel data, respectively.

For the variances of our estimators, we adapt the asymptotic variance approximation suggested in zhang2025 to our settings with non-zero treatment doses in both the treatment and control groups. Considering the case of repeated cross-sections, we define

align[align omitted — 851 chars of source]

with $\Delta_{d_t,d_t', t}^h= E\left[ \psi^h(W,\Delta_{d_t,d_t',t}^h, \Pi_{d_t,t},\eta) \right]$ being the smoothed ATET with bandwidth $h$. Then $\phi^h(W,\Delta_{d_t,d_t',t}^h, \Pi_{d_t,t},\eta):=\psi^h(W,\Delta_{d_t,d_t',t}^h, \Pi_{d_t,t},\eta)-\Delta_{d_t,d_t', t}^h$ is the smoothed Neyman-orthogonal score function for $\Delta_{d_t,d_t', t}^h$ with bandwidth $h$. The asymptotic variance, denoted by $\sigma_h^2$, of our DML estimator of $\Delta_{d_t,d_t', t}$ corresponds to

align[align omitted — 252 chars of source]

It is worth noting that the second term in the asymptotic variance in equation (ref) \[ \frac{\Delta_{d_t, d_t', t}^h}{\Pi_{d_t,t}}\left(\omega(D;d,h)\cdot I\{T=t\} - E[\omega(D;d,h)\cdot I\{T=t\}]\right) \] arises from a Taylor expansion of the score w.r.t.\ the (low dimensional) density $P(D_t=d_t, T=t)$.

A cross-fitted variance estimator can be constructed based on the sample analog of equation (ref):

align[align omitted — 321 chars of source]

where $\hat\Delta_{d_t,d_t', t}^h$ is the cross-fitted ATET estimate (see step 4 of the algorithm). The term $\phi^h(W_i,\hat{\Delta}_{d_t,d_t',t}^h, \hat{\Pi}_{k,d_t,t},\hat{\eta}_k)$ denotes the estimate of score function $\phi^h(W,\Delta_{d_t,d_t',t}^h, \Pi_{d_t,t},\eta)$ for subject $i$, which is computed using the subsample estimates of the nuisance parameters $\hat{\eta}_k$ (see step 3 of the algorithm). Similarly $\hat{\Pi}_{k,d_t,t} =|\mathcal{W}_k^{C}|^{-1}\sum_{i} \omega(D_i;d,h)\cdot I\{T_i=t\}$ is the subsample estimate of the density $\Pi_{d_t,t}$. In zhang2025, such types of variance estimators are shown to be consistent, provided that the kernel bandwidth $h$ satisfies $h^{-2}\varepsilon_n^2 + h^{-3}n^{-1} = o(1)$ for $\varepsilon_n = o(n^{-1/4})$.

In the case of panel data, the results in zhang2025 imply that the asymptotic variance of the ATET estimator of $\Delta_{d_t,d_t', t}$ takes the following form:

align[align omitted — 222 chars of source]

where $\Delta_{d_t,d_t',t}^h=E[\psi^h(W,\Delta_{d_t,d_t',t}^h, P_{d_t},\eta)]$ is the smoothed ATET under bandwidth $h$ and $\phi^h(W,\Delta_{d_t,d_t',t}^h, P_{d_t},\eta)=\psi^h(W,\Delta_{d_t,d_t',t}^h, P_{d_t},\eta)-\Delta_{d_t,d_t',t}^h$ is the score function, with

align[align omitted — 345 chars of source]

A cross-fitted estimator of the variance in equation (ref) can be constructed analogously as outlined for the case of repeated cross-sections.

Lastly, the corresponding estimators and variance expressions for the ATET with lagged treatment $\Delta_{d_{t-s},d'_{t-s}, t}=E[Y_t(d_{t-s})-Y_t(d'_{t-s})|D_{t-s}=d_{t-s}]$ can be easily obtained by substituting the expressions shown earlier with the lagged versions.

As an alternative approach to the asymptotic variance approximation, one can employ a multiplier bootstrap procedure for constructing confidence intervals. Let $\{\xi_i\}_{i=1}^n$ be an i.i.d.\ sequence of sub-exponential random variables independent from the data $\{W_i\}_{i=1}^n$, such that $E[\xi] = Var(\xi) = 1$. In each bootstrap replication $b = 1, \cdots, B$, one draws such a sequence $\{\xi_i\}_{i=1}^n$ and then constructs a bootstrap-specific ATET estimators, denoted by $\hat{\Delta}_{d_t, d_t', t}^{b,h}$, using the following expression:

align[align omitted — 210 chars of source]

Let $\hat{c}_{\alpha}$ denote the $\alpha$-th quantile of the difference $\{\hat{\Delta}_{d_t, d_t', t}^{b,h} - \hat{\Delta}_{d_t, d_t', t}^h\}_{b=1}^B$, a $1-\alpha$ confidence interval may be constructed as $[\hat{\Delta}_{d_t, d_t', t}^h - \hat{c}_{1-\alpha/2}, \hat{\Delta}_{d_t, d_t', t}^h - \hat{c}_{\alpha/2}]$.

Moreover, in the continuous treatment setting, it is also useful to consider uniform inference, and the bootstrap estimator defined in ((ref)) can be used to construct the uniform confidence bands. For simplicity, we focus on the two-period repeated cross-sectional case comparing the treated group with intensity $d$ and the control group with intensity $0$. In particular, this extends the repeated cross-sectional results in zhang2025 to allow time-varying covariates using efficient scores.

Recall that our causal parameter in this setting is defined as

align[align omitted — 74 chars of source]

and we aim to establish valid uniform confidence bands for $\Delta_{d_t, 0, t}$ over the support of $D_t$. Let $\hat{\Delta}_{d_t, 0, t}^h$, $\hat{\Delta}_{d_t, 0, t}^{b,h}$, and $\hat{\sigma}_h^2(d_t,0,t)$ denote our DML estimator, the multiplier bootstrap estimator, and the cross-fitted variance estimator respectively. We consider the following procedure that follows CCK2014a, CCK2014b, FHLZ22, and zhang2025 to establish valid uniform confidence bands.

itemize• Estimate $\hat{\Delta}_{d_t, 0, t}^h$ and $\hat{\sigma}_h^2(d_t,0,t)$ on a finite grid of values $d_t\in \bar{\mathcal{D}}_t \subset supp(D_t)$. • For each $b=1, \cdots, B$, draw an i.i.d. sequence of multipliers $\{\xi\}_{i=1}^N$ from a $N(1,1)$ distribution, and construct $\hat{\Delta}_{d_t, 0, t}^{b,h}$ for all $d_t\in \bar{\mathcal{D}}_t$. • Compute $\hat{c}(1-\alpha)$, which we denote as the $1-\alpha$-quantile of \[ \Bigg\{\max_{d_t\in \bar{\mathcal{D}}_t} \frac{\sqrt{N}|\hat{\Delta}_{d_t, 0, t}^h - \hat{\Delta}_{d_t, 0, t}^{b,h}|}{\hat{\sigma}_h^2(d_t,0,t)} \Bigg \}_{b=1}^B. \] • For all $d_t\in supp(D_t)$, construct the $1-\alpha$ uniform confidence band as \[ [\hat{\Delta}_{d_t, 0, t}^h - \hat{c}(1-\alpha)\hat{\sigma}_h^2(d_t,0,t)/\sqrt{N}, \quad \hat{\Delta}_{d_t, 0, t}^h + \hat{c}(1-\alpha)\hat{\sigma}_h^2(d_t,0,t)/\sqrt{N}]. \]

We omit the formal theoretical discussion here and point the readers to zhang2025, which establishes valid uniform asymptotic linear expansion of the bootstrap estimators analogous to ours, and the validity of the uniform confidence bands constructed above follows from the results in CCK2014a,CCK2014b and FHLZ22.

Asymptotic Theory

This section establishes the asymptotic properties of our estimators. The results largely follow zhang2025 and we sketch the main idea of the proofs in Appendix (ref). First, we impose a set of regularity conditions on the kernel function and the boundedness and smoothness of the model parameters. These conditions are used to establish three main results. First, the bias of ATET from using a kernel approximation becomes negligible asymptotically if the kernel bandwidth is undersmoothed. Second, the ATET estimators are asymptotically normal. Third, the variance estimators based on the asymptotic variance formulas proposed in Section (ref) are consistent. We focus on the case without time lags, and the results for the case with time lags follow immediately by appropriately defining treatment and outcome periods according to the time lag $t-s$. Recall that in the previous sections, we define the weighting function as $\omega(D;d,h)=\frac{1}{h}\mathcal{K}\left(\frac{D-d}{h}\right)$ with $h$ being the bandwidth. Using this weighting function, we consider the bandwidth-dependent approximations of the generalized propensitiy scores, denoted by $\rho_{d,t}^h(\mathbf{D}_{t-1},X)= E[\omega(D;d,h) \cdot I\{T=t\}|\mathbf{D}_{t-1},X]$ and $p_{d}^h(\mathbf{D}_{t-1},X)= E[\omega(D;d,h)|\mathbf{D}_{t-1},X]$, for the cross-sectional and panel settings respectively. In addition, we use $\hat\rho_{d,t}^h(\mathbf{D}_{t-1},X)$ and $\hat p_{d}^h(\mathbf{D}_{t-1},X)$ to denote the estimates of $\rho_{d,t}^h(\mathbf{D}_{t-1},X)$ and $p_{d}^h(\mathbf{D}_{t-1},X)$.

assumption{\bf (Kernel):}\\ The kernel function $\mathcal{K}(\cdot)$ satisfies: (a) $\mathcal{K}(\cdot)$ is bounded and differentiable; (b) $\int \mathcal{K}(u) du = 1$, $\int u\mathcal{K}(u)du = 0$, $0<\int u^2 \mathcal{K}(u) du <\infty$.
assumption{\bf (Bounds and smoothness, repeated cross-sections):} \\ (a) There exists $0<c<1$ and $0<C<\infty$ such that $\Pi_{d_t,t}>c$, $c<\rho_{d_t, t-1}^h(\mathbf{D}_{t-1},X)<C$, $\rho_{d_t', t}^h(\mathbf{D}_{t-1},X)>c$, $\rho_{d_t', t-1}^h(\mathbf{D}_{t-1},X)>c$, $|Y_T| < C$, $|\mu_{d_t}(t-1,\mathbf{D}_{t-1},X)|<C$, $|\mu_{d_t'}(t,\mathbf{D}_{t-1},X)|<C$, and $|\mu_{d_t'}(t-1,\mathbf{D}_{t-1},X)|<C$ almost surely; (b) $\Pi_{d_t,t} \in C^2(\text{supp}(D_t))$ and $|\Pi_{d_t,t}^{(2)}| < \infty$; (c) $\rho_{d_t, t-1}(\mathbf{D}_{t-1},d) \in C^2(\text{supp}(D_t))$ for all $(\mathbf{D}_{t-1}, x)$, and $\sup_{d_t, d_{t-1}, x} |\rho_{d_t, t-1}^{(2)}(\mathbf{D}_{t-1},x)| < \infty$; (d) joint density $f_{Y_T, D_t}(y,d_t) \in C^2(\text{supp}(Y_T))$ and $\sup_{y, d_t \in \text{supp}(Y_T, D_t)} |f_{Y_T, D_t}^{(2)}(y,d_t)| < \infty$.
assumption{\bf (Rates, repeated cross-sections):}\\ (a) The kernel bandwidth $h = h_N\to 0$ satisfies $Nh\to\infty$ and $\sqrt{Nh^5} = o(1)$; (b) there exists a sequence $\varepsilon_n\to 0$, such that $h^{-1}\varepsilon_n^2 = o(1)$; (c) with probability tending to $1$, $\|\hat{\rho}_{d_t, t}^h(\mathbf{D}_{t-1},X) - \rho_{d_t, t}^h(\mathbf{D}_{t-1},X)\|_{P,2}\leq h^{-1/2}\varepsilon_n$, $\|\hat{\rho}_{d_t, t-1}^h(\mathbf{D}_{t-1},X) - \rho_{d_t, t-1}^h(\mathbf{D}_{t-1},X)\|_{P,2}\leq \varepsilon_n$, $\|\hat{\rho}_{d_t', t-1}^h(\mathbf{D}_{t-1},X) - \rho_{d_t', t-1}^h(\mathbf{D}_{t-1},X)\|_{P,2}\leq \varepsilon_n$, $\|\hat{\mu}_{d_t}(t-1,\mathbf{D}_{t-1},X) - \mu_{d_t}(t-1,\mathbf{D}_{t-1},X)\|_{P,2}\leq \varepsilon_n$, $\|\hat{\mu}_{d_t'}(t-1,\mathbf{D}_{t-1},X) - \mu_{d_t'}(t-1,\mathbf{D}_{t-1},X)\|_{P,2}\leq \varepsilon_n$, $\|\hat{\mu}_{d_t'}(t,\mathbf{D}_{t-1},X) - \mu_{d_t'}(t,\mathbf{D}_{t-1},X)\|_{P,2}\leq \varepsilon_n$; (d) with probability tending to $1$, $ c< \hat{\rho}_{d_t, t-1}^h(\mathbf{D}_{t-1},X) <C$, $ \hat{\rho}_{d_t', t}^h(\mathbf{D}_{t-1},X) >c$, and $\hat{\rho}_{d_t', t-1}^h(\mathbf{D}_{t-1},X) >c$ almost surely; (e) with probability tending to 1, $\|\hat{\mu}_{d_t}(t-1,\mathbf{D}_{t-1},X)\|_{P,\infty}<C$, $\|\hat{\mu}_{d_t'}(t-1,\mathbf{D}_{t-1},X)\|_{P,\infty}<C$, and $\|\hat{\mu}_{d_t'}(t,\mathbf{D}_{t-1},X)\|_{P,\infty}<C$.
assumption{\bf (Bounds and smoothness, panel):} \\ (a) There exist constants $0<c<1$ and $0<C<\infty$, $P_{d_t}>c$, $|Y_{t-1}| < C$, $|Y_{t}| < C$, $p_{d_t'}^h(\mathbf{D}_{t-1},X)> c$, $|p_{d_t}^h(\mathbf{D}_{t-1},X)|<C$, and $|m_{d_t'}(t,\mathbf{D}_{t-1}, X)|<C$ almost surely; (b) $P_{d_t} \in C^2(\text{supp}(D_t))$ and $|P_{d_t}^{(2)}| < \infty$; (c) $p_{d_t}(\mathbf{D}_{t-1},x) \in C^2(\text{supp}(D_t))$ for all $(\mathbf{D}_{t-1},x)$ and $\sup_{\text{supp}(\mathbf{D}_{t-1}, X)} |p_{d_t}^{(2)}(\mathbf{D}_{t-1},X)| < \infty$; (d) for $\Delta Y = Y_t - Y_{t-1}$, the joint density $f_{\Delta Y, D_t}(s, d_t) \in C^2(\text{supp}(\Delta Y))$ and $\sup_{s,d_t \in \text{supp}(\Delta Y, D_t)} |f_{\Delta Y, D_t}^{(2)}(s, d_t)| < \infty$.
assumption{\bf (Rates, panel):} \\ (a) The kernel bandwidth $h = h_N\to 0$ satisfies $Nh\to\infty$ and $\sqrt{Nh^5} = o(1)$; (b) there exists a sequence $\varepsilon_n\to 0$ such that $h^{-1}\varepsilon_n^2 = o(1)$; (c) with probability tending to $1$, $\|\hat{p}_{d_t}^h(\mathbf{D}_{t-1},X) - p_{d_t}^h(\mathbf{D}_{t-1},X)\|_{P,2}\leq h^{-1/2}\varepsilon_n$, $\|\hat{p}_{d_t'}^h(\mathbf{D}_{t-1},X) - p_{d_t'}^h(\mathbf{D}_{t-1},X)\|_{P,2}\leq \varepsilon_n$, $\|\hat{m}_{d_t'}(t,\mathbf{D}_{t-1}, X) - m_{d_t'}(t,\mathbf{D}_{t-1}, X)\|_{P,2}\leq \varepsilon_n$; (d) with probability tending to $1$, $\sup_{\text{supp}(\mathbf{D}_{t-1}, X)} \hat{p}_{d_t'}^h(\mathbf{D}_{t-1},X) > c$, $\|\hat{p}_{d_t}^h(\mathbf{D}_{t-1},X)\|_{P,\infty}<C$, $\|\hat{m}_{d_t'}(t,\mathbf{D}_{t-1}, X)\|_{P,\infty}<C$.

Assumption (ref) on the kernel function is standard and many commonly used smoothing functions satisfy these conditions, e.g., the Gaussian or Epanechnikov kernels. Moreover, we impose two sets of regularity conditions in both repeated cross-sections and panel settings. First, we require that bandwidth-dependent nuisance parameters in the population are bounded and satisfy certain smoothness conditions, see Assumptions (ref) and (ref). Second, we impose high-level assumptions on the quality of the estimators of the nuisance parameters, permitting those estimators to converge at a slower rate compared to a setting without Neyman orthogonality, see Assumptions (ref) and (ref). A first result obtained from these conditions is that the biases in repeated cross-sections and panel settings are of order $O(h^2)$, such that they vanish asymptotically when using an undersmoothed kernel bandwidth for ATET estimation.

lemmaSuppose Assumptions (ref), (ref), (ref) hold for the repeated cross-sections case, and Assumptions (ref), (ref), (ref) hold for the panel case. Then, the bias satisfies $B(h) = \Delta_{d_t,d_t',t} - \Delta_{d_t,d_t',t}^h = O(h^2)$ for $d_t, d_t'\in \text{supp}(D_t)$.

The next two theorems are the main results of this section. The first theorem states that our proposed ATET estimators are asymptotically normal and converge at $\sqrt{nh}$ rate to the true ATET, while the second theorem states that our cross-fitted variance estimators proposed in the previous section are consistent.

theoremSuppose Assumptions (ref), (ref), (ref), (ref) hold. Moreover, suppose Assumptions (ref), (ref), (ref) hold for the repeated cross-sections case, and Assumptions (ref), (ref), (ref) hold for the panel case. Then for $\varepsilon_n = o(n^{-1/4})$, the ATET estimators defined in Section (ref) satisfy \begin{align} \frac{\hat{\Delta}_{d_t,d_t',t} - \Delta_{d_t,d_t',t}}{\sigma_h/\sqrt{n}} \quad \rightarrow^d \quad N(0,1) \end{align} where $\sigma_h$ is defined as in ((ref)) and ((ref)) for the repeated cross-sections case and panel case respectively.
theoremAssume the conditions in Theorem (ref) hold and assume that $h^{-2}\varepsilon_n^2 + h^{-3}n^{-1} = o(1)$. Then the cross-fitted variance estimators defined in Section (ref) are consistent for the asymptotic variance $\sigma_h$ defined in ((ref)) and ((ref)) for the repeated cross-sections case and panel case respectively.

A sketch of the proofs is provided in Appendix (ref).

Simulation study

This section provides a simulation study to investigate the finite sample behavior of our DiD estimators for repeated cross-sections and panel data. Starting with the case of repeated cross-sections, we consider the following data generating process (DGP):

eqnarray*[eqnarray* omitted — 250 chars of source]

The vector $X$ consists of time varying covariates $X_j$ with $j \in \{1,...,p\}$, where $p$ is the dimension of $X$. Each covariate $X_j$ depends on the time index $T$, which is either 0 for the pre-treatment and 1 for the post-treatment period (each with a probability of 50%), as well as the random error $Q_j$, which is independently uniformly distributed with its support ranging from 0 to 2.

The continuous treatment $D$ is a function of covariates $X$ if the coefficient vector $\beta $ is different from zero and two uniformly distributed error terms, namely the fixed effect $U$ and a random component $V$. Both the covariates $X$ and the fixed effect $U$ also affect the outcome $Y_T$ and are, therefore, confounders of the treatment-outcome relation that is tackled by our covariate-adjusted DiD approach. In addition, the outcome is affected by the period indicator $T$, which reflects a time trend, and a random uniformly distributed error term $W$. Lastly, the treatment has a nonlinear effect on the post-treatment outcome, which is modeled by the interaction of $D^2$ and $T$.

In our simulations, we set the number of covariates to $p=100$ and the $j$th element in the coefficient vector $\beta$ to $0.4/i^2$ for $j \in \{1,...,p\}$, which implies a quadratic decay of covariate importance in terms of confounding. We consider two sample sizes of $n=2000, 4000$. DiD estimation is based on the sample analog of equation (ref). In our simulation design, there are only two periods $t$ such that $T=t$ and $T=t-1$ in equation (ref) correspond to $T=1$ and $T=0$, respectively. Furthermore, there is no history of previous treatments in $T=0$ such that $\mathbf{D}_{t-1}=\emptyset$ in our simulations. However, from an econometric perspective, there is no difference between controlling for the treatment history $\mathbf{D}_{t-1}$ or covariates $X$, which are both time-varying sets of confounders. For this reason, not considering $\mathbf{D}_{t-1}$ in our simulations comes without loss of generality as we could easily relabel some elements in $X$ to represent $\mathbf{D}_{t-1}$.

For estimation, we apply the didcontDML function from the causalweight package by BodoryHuber2018 using the statistical software R (R2020). We implement this DML-based DiD estimator for repeated cross-sections with two-fold cross-fitting and estimate the nuisance parameters via cross-validated lasso regression. We employ a second-order Epanechnikov kernel as the weighting function $\omega$ for smoothing over the continuous treatment, with a bandwidth given by $2.34 \cdot n^{-0.25}$. Choosing 2.34 as the bandwidth constant corresponds to a Silverman86-type rule of thumb for second-order Epanechnikov kernels, as also applied, for instance, in huber2020direct, while the rate $n^{-0.25}$ ensures undersmoothing to avoid asymptotic bias induced by kernel smoothing. In addition to estimation based on this rule, we consider a specification with stronger undersmoothing by multiplying the resulting bandwidth by 0.5 (which amounts to halving the constant in the bandwidth formula), yielding $2.34 \cdot n^{-0.25} \cdot 0.5$.

Furthermore, when estimating the generalized propensity scores $\rho_{d,t}(X)$ via lasso, we consider both linear and log-linear specifications that assume normally distributed unobservables $V$ (an assumption that is violated here, as $V$ is uniformly distributed). The ATET estimator additionally imposes a trimming rule to safeguard against disproportionately influential observations in the inverse propensity weighting (IPW). Specifically, observations receiving a weight larger than 10% in the IPW-based computation of any conditional mean outcome given the treatment state and time period are dropped. Such a trimming rule to ensure common support are also considered in HuLeWu10. Finally, standard errors for the ATET estimates are obtained using the estimated asymptotic variance approximation given in equation (ref) of Section (ref).

The upper panel of Table (ref) provides the simulation results for repeated cross-sections, when estimating the ATET $\Delta_{3,2,1}=E[Y_1(3)-Y_0(2)|D_1=3]=3^2-2^2=5$, based on linear lasso regression with weaker undersmoothing (`lasso'), loglinear lasso regression with weaker undersmoothing (`lnorm'), linear lasso regression with stronger undersmoothing (`under'), and loglinear lasso regression with stronger undersmoothing (`ln under'). The columns provide the estimation method (`method'), the bias of the estimator (`bias'), its standard deviation (`std'), its root mean squared error (`rmse'), and the average standard error (`avse'), respectively, for either sample size.

For $n=2000$, we observe that the ATET estimators with weaker undersmoothing are non-negligibly biased, whereas those with stronger undersmoothing are less biased while having a slightly higher standard deviation. Overall, the strongly undersmoothed versions have a superior performance in terms of a lower root mean squared error (RMSE), which measures the overall estimation error as the square root of the sum of the squared bias and variance. Moreover, the ATET estimates are very similar when using normal or lognormal models for the estimation of the generalized propensity scores. Concerning statistical inference, the average standard error (`avse') is generally close to the actual standard deviation (`std'), suggesting that asymptotic variance estimation as outlined in Section (ref) performs well in our simulations. When increasing the sample size to $n=8000$, the biases and standard deviations of all estimators decrease, but again, it is the strongly undersmoothed versions of the estimators that perform substantially better in terms of the RMSE than those with weaker undersmoothing.

table[table omitted — 2,068 chars of source]

Subsequently, we investigate the DiD estimator for panel data, considering a slightly modified DGP in which $X_j$ only affects the outcome in the second period, but not in the first period:

eqnarray*[eqnarray* omitted — 218 chars of source]

We apply the didcontDMLpanel function of the causalweight package by BodoryHuber2018 and consider the same choices in terms of kernel functions, bandwidth selection, undersmoothing, generalized propensity score modeling, and trimming as for the repeated cross-sections case.

The lower panel of Table (ref) provides the simulation results for the panel case when estimating the ATET $\Delta_{3,2,1}=5$. The absolute magnitudes of the biases and the RMSEs of the ATET estimators with weaker undersmoothing are again considerably larger that those of the strongly undersmoothed versions for either sample size. Moreover, the average standard errors based on asymptotic variance approximation for the panel case as outlined in Section (ref) closely match the actual standard deviations of the respective estimators. For this reason, the simulation results for the panel and repeated cross-section cases are qualitatively similar.

Empirical application

In this section, we present an empirical application of our method in the context of the COVID-19 pandemic in Brazil—one of the countries most severely affected by the crisis. We estimate the causal effect of second-dose COVID-19 vaccination rates on mortality across Brazilian municipalities. Our data are compiled from multiple sources providing municipality-level health and socio-economic information, including the laboratory DB Molecular – Diagnósticos do Brasil, the Brazilian Institute of Geography and Statistics (IBGE), DataSUS (Ministry of Health, Brazil), and the monitoring of COVID-19 cases and deaths by CotaCovid19br2020. After cleaning and harmonizing the data, we obtain a balanced panel of 213 municipalities observed daily from August 4, 2020 to March 23, 2022, covering 597 days and yielding 127{,}161 observations.

We define the continuous and time-varying treatment variable as the municipality-specific vaccination rate, measured as the ratio of cumulative dose-2 vaccinations administered in a municipality up to a given day to the municipality’s population size according to the 2020 census. Figure (ref) displays vaccination rates over time (across 597 days) during the sample period, highlighting considerable heterogeneity in the pace of vaccination across municipalities. Given the limited number of municipalities in our sample, we adopt a stacked design that permits pooling across multiple periods to increase statistical power. Specifically, we designate the first day on which any municipality reaches a $70\%$ (or 0.7) dose-2 vaccination rate as the reference day $t_0$, as indicated by the dashed vertical line in Figure (ref). We then stack the data at seven-day intervals over five periods, $(t_0, t_7, t_{14}, t_{21}, t_{28})$, which increases the effective evaluation sample five-fold relative to a single day, yielding $5 \cdot 213 = 1065$ panel units. Furthermore, this design avoids common support issues at the treatment intensities under comparison, which are chosen to be 60% (or 0.6) and 40% (or 0.4), as indicated by the horizontal dashed lines in Figure (ref).

figure[figure omitted — 171 chars of source]

Our outcome variable is defined as the number of new COVID-19-related deaths per $100{,}000$ population in a given municipality at specific post-treatment horizons—namely one, two, four, seven, ten, fifteen, thirty, and sixty days after treatment. This allows us to study the dynamic evolution of the treatment effect. We control for the following covariates in a data-adaptive manner, conditional on which the parallel trends assumption is required to hold: (1) 45 demographic and socioeconomic characteristics measured at the municipal level, such as fertility, education, income, and life expectancy; (2) a set of dummy variables indicating the stacked reference days; (3) treatment histories capturing dose-2 vaccination rates one, two, seven, and fourteen days prior to the treatment day; and (4) dose-1 vaccination rates on the treatment day as well as their corresponding histories (dose-1 vaccination rates one, two, seven, and fourteen days prior to the treatment day), to avoid confounding from dose-1 vaccination.

Table (ref) reports descriptive statistics (mean, standard deviation, minimum, and maximum) for several key variables - including the outcome, the treatment, and a selected set of covariates reflecting the dose-1 vaccination rate, socioeconomic conditions, and infrastructure - computed across the 213 municipalities and all 597 days of the balanced panel. While the mortality outcome and vaccination rates can vary at the daily level, municipalities' socioeconomic and infrastructure characteristics are measured only once in 2022 and are therefore time-invariant (although they may exert time-varying influences on outcomes, which is why they are considered as potential control variables). The table shows that municipalities represent diverse socioeconomic contexts, with substantial variation in per capita income and infrastructure (such as sewerage network coverage and access to piped water), which are relevant factors for healthcare capacity and pandemic management. See Appendix (ref) for a more comprehensive set of descriptive statistics covering a broader range of variables.

table[table omitted — 924 chars of source]

As in the simulation study in the previous section, we use the didcontDML function from the causalweight package by BodoryHuber2018 to estimate the causal effect of the vaccination rate, employing two-fold cross-fitting to mitigate overfitting. To flexibly adjust for high-dimensional control variables, the nuisance parameters are estimated nonparametrically using random forests as the machine learner (see, e.g., Ho1995, Breiman2001), assuming normally distributed unobservables in the generalized propensity score model. The kernel bandwidth used for smoothing over treatment intensity is set to $2.34 \cdot n^{-0.25} \cdot 0.7$, that is, the rule-of-thumb constant for bandwidth selection multiplied by $0.7$ to implement undersmoothing based on $n^{-0.25}$. This choice lies between the weaker and stronger undersmoothing approaches considered in the simulation study. Finally, we apply a trimming rule that drops observations receiving more than 10% weight in the IPW-based estimation of the mean potential outcome under non-treatment.

Table (ref) reports the estimated ATETs (measured in COVID-19 deaths per $100{,}000$ population) based on a comparison of a higher dose-2 vaccination rate of 60% ($d_{\text{treat}}=0.6$) versus a lower rate of 40% ($d_{\text{control}}=0.4$) across post-treatment horizons. The corresponding standard errors (SE) and p-values, clustered at the municipality level, are also reported. The results indicate that higher dose-2 coverage has no statistically significant short-run effects from one to fifteen days after treatment. However, statistically significant negative effects emerge at longer horizons, after thirty days (at the 5% level) and sixty days (at the 10% level). Specifically, the DML estimates suggest a reduction of approximately 1.5 COVID-19 deaths per $100{,}000$ inhabitants in municipalities with a 60% vaccination rate relative to a rate of 40%, after one to two months. These findings are consistent with the expected time lag between infection and death documented in the COVID-19 literature; see, e.g., huber2020timing.

table[table omitted — 569 chars of source]

Conclusion

In this paper, we have developed a generalized difference-in-differences framework to estimate the causal effects of time-varying continuous treatments. Our approach focused on identifying the Average Treatment Effect on the Treated (ATET) by comparing distinct treatment intensities, relying on a conditional parallel trends assumption that robustly accounts for complex treatment histories and covariate evolution. By integrating double/debiased machine learning with kernel-based weighting, we provided a rigorous estimation strategy that remains consistent even with high-dimensional nuisance parameters. Finally, we illustrated our method by applying it to Brazilian municipality-level data to estimate the effect of second-dose COVID-19 vaccination rates on mortality. The results indicate that higher vaccination coverage does not significantly affect mortality in the very short term but leads to a statistically significant reduction in COVID-19 deaths after several weeks, consistent with the expected time lag between infection and fatal outcomes.

{ \setcounter{equation}{0}