Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
84,224 characters · 0 sections · 50 citation commands
Difference-in-Differences for Continuous Treatments and Instruments with Stayers
\@startsection{section}{2}{0mm}{-1\baselineskip}{1\baselineskip}{\normalfont}{Introduction}
To estimate a treatment's effect, researchers often run a two-way fixed effects (TWFE) regression: $$Y_{i,t}=\alpha_i+\gamma_t+\beta_{TWFE}D_{i,t}+u_{i,t},$$ where $D_{i,t}$ is the treatment of unit $i$ at time $t$, and $\alpha_i$ and $\gamma_t$ are unit and period fixed effects. de2020difference find that 26 of the 100 most cited papers published by the American Economic Review from 2015 to 2019 estimate TWFE regressions. chiu2023and find that 52 papers published by the top three political science journals from 2017 to 2022 estimate a TWFE regression. Researchers have long believed that those regressions only rely on a parallel-trends assumption, requiring that without treatment, all units would have experienced the same outcome evolutions. However, recent papers have shown that under that assumption alone, TWFE regressions may estimate a non-convex sum of treatment effects across units and periods deChaisemartin2018twowayfev4arxiv,goodman2021difference,borusyak2020revisiting. Then, $\beta_{TWFE}$ could be, say, negative, even if the effect is positive for every $(i,t)$.
In reaction to this finding, several “heterogeneity-robust difference-in-difference” estimators, consistent for convex combinations of $(i,t)$-specific effects under a parallel-trends assumption, have been proposed. However, none can be applied to treatments continuously distributed at every period. Yet, such treatments are commonly found. In economics, taxes li2014gasoline or tariffs fajgelbaum2020return are often continuously distributed at all periods. Similarly, peterson2021paper, aklin2021side, and mcnamara2022estimating are examples of papers in political sciences and epidemiology studying treatments continuous in all periods. With continuous treatments, due to the lack of an untreated group, TWFE regressions often estimate a highly non-convex combination of effects. For instance, with two time periods, exactly 50% of $(i,t)$-specific effects are weighted negatively de2021more. Then, proposing robust estimators for such treatments is particularly important.
We assume that we have a two-periods panel data set. From period one to two, the treatment of some units, referred to as the switchers, changes, while the treatment of some other units, the stayers, does not change. We propose a novel parallel-trends assumption, which requires that switchers and stayers with the same period-one treatment would have experienced the same average outcome evolutions from period one to two, if switchers' treatment had not changed. Our assumption has two desirable features. First, because it conditions on units' period-one treatment, it does not restrict effects' heterogeneity. Instead, as explained later, assuming parallel trends for switchers and stayers with different period-one treatments would essentially amount to assuming that the treatment effect is constant over time. Second, when another prior period of data, period zero, is available, the plausibility of our assumption can be assessed via a “pre-trends” or “placebo” test, comparing the period-zero-to-one outcome evolutions of period-one-to-two switchers and stayers.\footnote{As explained later, the pre-trends estimator has to be computed in the subsample of units whose treatment does not change from period zero to one. Then, one may also restrict the main estimators to that subsample.} Thus, one can assess if switchers and stayers are on parallel trends before switchers' treatment changes.
Then, we consider two target parameters. The first is the average slope of switchers' period-two potential outcome function, from their period-one to their period-two treatment, referred to as the Average of Slopes (AS). The second is a weighted average of switchers' slopes, where switchers receive a weight proportional to the absolute value of their treatment change, referred to as the Weighted Average of Slopes (WAS). The AS and WAS serve different purposes. Under shape restrictions on the potential outcome function, the AS can be used to identify or bound the effect of other treatment changes than those that took place. Instead, the WAS can be used to conduct a cost-benefit analysis of the treatment changes that took place. Under our parallel-trends assumption, whose plausibility can be assessed using a pre-trends test, the AS and the WAS are identified by DID estimands that compare the outcome evolution of switchers and stayers with the same period-one treatment. This is in contrast with other targets considered in the literature like the dose-response function or the average marginal effect, which can only be identified under assumptions whose plausibility cannot be assessed via a pre-trends test.
Turning to estimation, the challenge with a continuous treatment is that the sample does not contain switchers and stayers with the same period-one treatment. In other words, the estimands identifying the AS and WAS depend on nonparametric regression functions. Therefore, we derive doubly-robust moment conditions identifying the AS and WAS, and we propose doubly-robust estimators. We show that the AS estimator converges at the parametric rate, provided switchers cannot experience arbitrarily small treatment changes. We then show that the WAS can always be estimated at the parametric rate. Under some conditions, the asymptotic variance of the WAS estimator is strictly lower than that of the AS estimator.
We extend those results in several directions. First, we consider the instrumental-variable (IV) case, where one makes a parallel-trends assumption with respect to an instrument rather than the treatment. For instance, one may be interested in estimating the price-elasticity of a good. If prices respond to demand shocks, the counterfactual consumption trends of units experiencing and not experiencing a price change may not be the same, so a parallel-trends assumption with respect to prices may not hold. But instead, one can make a parallel-trends assumption with respect to taxes, used as an instrument for prices. We start by noting that parallel-trends with respect to an instrument restricts treatment-effect heterogeneity. Then, we show that under that assumption, the reduced-form WAS effect of the instrument on the outcome divided by the first-stage WAS effect is equal to a weighted average of switchers' outcome-slope with respect to the treatment, where switchers with a larger first-stage effect receive more weight.
We also consider other extensions. First, we consider applications with more than two periods. Then, our estimators start by comparing, for every pair of consecutive time periods $(t-1,t)$, the outcome evolutions of units whose treatment switches / does not switch from $t-1$ to $t$, controlling for their period-$t-1$ treatment. Then, they aggregate those comparisons across $t$. Second, we propose a pre-trends estimator to assess the plausibility of our parallel-trends assumption. Essentially, instead of comparing the $t-1$-to-$t$ outcome evolutions of $t-1$-to-$t$ switchers and stayers, it compares their $t-2$-to-$t-1$ outcome evolutions. Third, with more than two periods, our estimators rule out dynamic effects: they assume that units' outcome at period $t$ only depends on their period-$t$ treatment, not on prior treatments. To alleviate this concern, we propose a modified version of our estimators, robust to dynamic effects up to a pre-specified treatment lag. Finally, we propose estimators allowing for control variables. The did_multiplegt_stat Stata package computes our estimators.
As an illustration, we use the yearly, 1966 to 2008 US state-level panel dataset of li2014gasoline to estimate the effect of gasoline taxes on gasoline consumption and prices. Using the WAS estimators, we find a significantly negative effect of taxes on gasoline consumption, and a significantly positive effect on prices. The AS estimators are close to, and not significantly different from, the WAS estimators, but they are insignificant because they are markedly less precise: their standard errors are almost three times larger. We compute an IV estimator of the price elasticity of gasoline consumption, and find a significantly negative elasticity of -0.66, which is more negative than, but not significantly different from, that reported in recent literature coglianese2017anticipation. Our pre-trend estimators are small and insignificant, thus lending credibility to our parallel-trends assumption.
\@startsection{subsection}{2}{0mm}{-1\baselineskip}{0.7\baselineskip}{\normalfont}*{Related Literature.}
Our paper builds upon several papers in the panel data literature. chamberlain1982multivariate seems to be the first paper to have proposed an estimator of the AS parameter. Under the assumption of no counterfactual time trend, the estimator therein is a before-after estimator. Our paper is also closely related to the work of graham2012identification, who propose DID estimators of the AS (see their Equation (21)) when the treatment is continuously distributed at every time period. Their estimators rely on a linear effect assumption and assume that units experience the same evolution of their treatment effect over time. By contrast, our estimator of the AS does not place any restriction on treatment effects. But our main contribution to this literature is to introduce the WAS, and to contrast the different purposes served by the two parameters.
With respect to the aforementioned heterogeneity-robust DID literature, we make two contributions. First, in the non-IV case we propose estimators that can be used even if units receive heterogeneous treatment doses at baseline. Thus we complement previous literature, that has mostly focused on the case where all units receive the same dose at baseline abraham2018,callaway2020difference,liu2024practical,borusyak2020revisiting,de2024two,callaway2021difference. One exception predating this paper is deChaisemartin2018twowayfev4arxiv, who, in their web appendix, consider designs where units receive heterogeneous doses of a non-binary discrete treatment at baseline. With a discrete treatment, they propose DID estimators comparing switchers and stayers with the exact same baseline treatment. This solution is no longer applicable with a continuously distributed treatment, hence the estimation challenge we tackle in this paper. Another exception, contemporaneous of this paper, is de2020difference, who propose estimators robust to dynamic effects up to any lag. While they focus mostly on discrete treatments, in Section 1.10 of their web appendix they build upon the current paper to sketch an extension of their estimators to continuous treatments. Contrary to here, the estimators therein are parametric and not doubly-robust, and the estimators' asymptotic distribution is not derived. Moreover, with unrestricted dynamic effects, instead of comparing $t-1$-to-$t$ switchers and stayers, DID estimators need to compare units switching treatment for the first time at $t$ to units whose treatment has not switched yet at $t$. This reduces the sample of units that can be used in the estimation, which may come with high costs in terms of external validity and statistical precision. We illustrate this in our empirical application. There, the estimators of de2020difference apply to a sample more than eight times smaller than the estimators in this paper. While point estimates are similar, the standard errors of the estimators of de2020difference are more than nine times larger than the standard errors of the estimators proposed in this paper. Thus, this paper's estimators may be appealing when it is a priori plausible to rule out dynamic effects, at least up to a pre-specified treatment lag, and/or when the tests of the no-dynamic effects assumption recently proposed by liu2024practical and de2025treatmenteffect are not rejected.\footnote{On the other hand, our results are not directly related to those in TVB. That paper considers setups where the switchers and stayers start from different treatment levels, but focuses on designs with discrete rather than continuous treatments. Moreover, they analyze the performance of standard TWFE estimators, while we propose nonparametric heterogeneity-robust DID estimators.}
Second, in the IV case we extend de2010note and hudson2017interpreting, who consider simple designs with two periods and a binary instrument and treatment. Our paper is also the first to note that parallel-trends with respect to an instrument restricts treatment-effect heterogeneity.
\paragraph{Organization of the paper.} Section (ref) presents the set-up, our main assumptions, and a building-block identification result. Section (ref) considers identification and estimation of the AS. Section (ref) turns to the WAS. Section (ref) presents some extensions. Finally, Section (ref) presents our application. The web appendix also collects all proofs.
\@startsection{section}{2}{0mm}{-1\baselineskip}{1\baselineskip}{\normalfont}{Set-up, assumptions, and building-block identification result}
\@startsection{subsection}{2}{0mm}{-1\baselineskip}{0.7\baselineskip}{\normalfont}{Set-up}
A representative unit is drawn from an infinite super population, and observed at two time periods. This unit could be an individual or a firm, but it could also be a geographical unit, like a county or a region.\footnote{In that case, one may want to weight the estimation by counties' or regions' populations. Extending the estimators we propose to allow for such weighting is a mechanical extension.} All expectations below are taken with respect to the distribution of variables in the super population. We are interested in the effect of a continuous and scalar treatment variable on that unit's outcome. Let $D_1$ (resp. $D_2$) denote the unit's treatment at period 1 (resp. 2), and let $\mathcal{D}_1$ (resp. $\mathcal{D}_2$) be its support. Let $S=1\{D_2\ne D_1\}$ be an indicator equal to 1 if the unit's treatment changes from period one to two, i.e. if they are a switcher.
For all $(d_1,d_2)\in \mathcal{D}_1\times \mathcal{D}_2$, let $Y_1(d_1,d_2)$ and $Y_2(d_1,d_2)$ respectively denote the unit's potential outcomes at periods 1 and 2 if $(D_1,D_2)=(d_1,d_2)$, and let $Y_1$ and $Y_2$ denote their observed outcomes.
In what follows, all equalities and inequalities involving random variables are required to hold almost surely. Finally, for any random variable observed at the two time periods $(X_1,X_2)$, let $\Delta X=X_2-X_1$ denote the change of $X$ from period $1$ to $2$.
\@startsection{subsection}{2}{0mm}{-1\baselineskip}{0.7\baselineskip}{\normalfont}{Assumptions}
We make the following assumptions.
Assumption (ref) requires that the period-one outcome does not depend on units' period-two treatment, thus ruling out anticipation effects, a commonly-made assumption in the DID literature. Assumption (ref) requires that the period-two outcome does not depend on units' period-one treatment, thus ruling out dynamic effects. This is not of essence in the two-periods case we focus on for now. This is of essence when we extend our estimators to the case where $T>2$, as we will explain in Section (ref). Therein, we propose a modified version of our estimators, which partly alleviates this concern by allowing for dynamic effects up to a pre-specified treatment lag.
Assumption (ref) is a parallel-trends assumption, requiring that $\Delta Y(d_1)$ be mean independent of $D_2$, conditional on $D_1=d_1$.
Assumption (ref) ensures that all the expectations below are well defined. It requires that the set of values that the period-one and period-two treatments can take be bounded. It also requires that the potential outcome functions be Lipschitz (with a unit-specific Lipschitz constant). This holds if $d\mapsto Y_2(d)$ is differentiable with respect to $d$ and has a bounded derivative.
Finally, for estimation and inference we assume we observe an iid sample with the same distribution as $(Y_{1},Y_{2},D_{1},D_{2})$:
Importantly, Assumption (ref) allows for the possibility that $Y_{1}$ and $Y_{2}$ (resp. $D_{1}$ and $D_{2}$) are serially correlated, as is commonly assumed in DID studies bertrand2004.
\@startsection{subsection}{2}{0mm}{-1\baselineskip}{0.7\baselineskip}{\normalfont}{Building-block identification result}
Assumption (ref) implies the following lemma, our building-block identification result.
Proof: we have
where the second equality follows from Assumption (ref). This proves the result $_\Box$
Intuitively, under Assumption (ref) the counterfactual outcome evolution switchers would have experienced if their treatment had not changed is identified by the outcome evolution of stayers with the same period-one treatment. If a unit's treatment changes from two to five, we can recover its counterfactual outcome evolution if its treatment had not changed, by using the average outcome evolution of all stayers with a baseline treatment of two. Then, a DID estimand comparing switchers' and stayers' outcome evolutions identifies $E\left(Y_2(d_2) - Y_2(d_1)\middle|D_1=d_1,D_2=d_2\right)$, and we can scale that effect by $d_2-d_1$ to identify a slope rather than an unnormalized effect.
In a canonical DID design where $\mathcal{D}_1=\{0\}$ and $\mathcal{D}_2\in \{0,1\}$, Lemma (ref) only applies to $(d_1,d_2)=(0,1)$, $\text{TE}(0,1|0,1)$ reduces to the average treatment on the treated (ATT), and the estimand reduces to the canonical DID estimand comparing the outcome evolutions of treated and untreated units. Also, if all units are untreated at period one ($\mathcal{D}_1=\{0\}$), we have $$\text{TE}(0,d_2|0,d_2)=E\left(\frac{Y_2(d_2) - Y_2(0)}{d_2}\middle|D_2=d_2\right),$$ an effect related to the $\Delta_{d0|D=d}$ effect in fricke2017identification, or to the $ATT(d|d)$ effect in callaway2021difference. Thus, the effects we consider are extensions of those effects.
In the next sections, our target parameters will be two averages of the slopes $\text{TE}(d_1,d_2|d_1,d_2)$, across both $d_1$ and $d_2$, which can be estimated nonparametrically at the $\sqrt{n}-$parametric rate. We do not propose estimators of $(d_1,d_2)\mapsto \text{TE}(d_1,d_2|d_1,d_2)$, for two reasons. First, variability in $\text{TE}(d_1,d_2|d_1,d_2)$ across values of $(d_1,d_2)$ conflates a dose-response relationship that may be of interest, and a selection bias, typically not of interest, due to the fact that units with different period one and two treatments may have different treatment effects callaway2021difference. Moreover, Lemma (ref) shows that estimating $\text{TE}(d_1,d_2|d_1,d_2)$ requires estimating two bivariate nonparametric regressions. Unless one is willing to make parametric functional-form assumptions, the resulting estimator converges at most at the $n^{1/3}$ rate stone1982optimal, and may therefore be very imprecise, given that DID applications often have a small number of units (e.g. the 50 US states in our empirical application). Similarly, while nonparametric estimators of $d_1\mapsto E\left(\frac{Y_2(D_2)- Y_2(d_1)}{D_2-d_1}\middle|D_1=d_1,D_2\ne d_1\right)$ or $\delta \mapsto E\left(\frac{Y_2(D_2)- Y_2(D_1)}{D_2-D_1}\middle|D_2-D_1=\delta >0\right)$ may converge at a faster rate than $n^{1/3}$, they cannot converge at the $\sqrt{n}$ rate.\footnote{The did_multiplegt_stat Stata package can still compute discretized versions of $d_1\mapsto E\left[(Y_2(D_2)- Y_2(d_1))/(D_2-d_1)\middle|D_1=d_1,D_2\ne d_1\right]$ (resp. $\delta \mapsto E[(Y_2(D_2)- Y_2(D_1))/(D_2-D_1)|D_2-D_1=\delta >0]$), when the option by_baseline (resp. by_fd) is specified, see the help file for further details.} In a recent paper, posterior to ours, haddad2024difference consider a parallel-trends assumption similar to Assumption (ref), and propose an estimator of a functional parameter similar to $(d_1,d_2)\mapsto \text{TE}(d_1,d_2|d_1,d_2)$.
Instead of averages of the slopes $\text{TE}(d_1,d_2|d_1,d_2)$, one may be interested in other parameters, like the average dose-slope function $(d,d')\mapsto \text{TE}(d,d'):=E\left((Y_2(d) - Y_2(d'))/(d-d')\right),$ or the average marginal effect $E\left(Y'_2(D_2)\right)$. However, while averages of $\text{TE}(d_1,d_2|d_1,d_2)$ can be identified under an assumption whose plausibility can be assessed via a pre-trends test, $\text{TE}(d,d')$ and $E\left(Y'_2(D_2)\right)$ cannot be identified under assumptions whose plausibility can be assessed via a pre-trends test, as explained in Section (ref) below.
\@startsection{subsection}{2}{0mm}{-1\baselineskip}{0.7\baselineskip}{\normalfont}{Alternative parallel-trends assumption}
The DID estimands in Lemma (ref) compare switchers and stayers with the same period-one treatment. Instead, one could propose estimands comparing switchers and stayers, without conditioning on their period-one treatment. To recover the counterfactual outcome trend of a switcher going from two to five units of treatment, one could use a stayer with treatment equal to three at both dates. On top of Assumption (ref), such estimands rest on two supplementary conditions:
(i) requires that all units experience the same evolution of their potential outcome with treatment $d$, while Assumption (ref) only imposes that requirement for units with the same baseline treatment. Assumption (ref) may be more plausible: units with the same period-one treatment may be more similar and more likely to be on parallel trends than units with different period-one treatments. (ii) requires that the trend affecting all potential outcomes be the same: to rationalize a DID estimand comparing a switcher going from two to five units of treatment to a stayer with treatment equal to three, $E(\Delta Y(2))$ and $E(\Delta Y(3))$ should be equal. Rearranging, (ii) is equivalent to $E(Y_2(d)-Y_2(d'))=E(Y_1(d)-Y_1(d'))$, so the treatment effect should be constant over time, a strong restriction on treatment effect heterogeneity. In contrast, our Assumption (ref) does not restrict treatment effect heterogeneity.
\@startsection{section}{2}{0mm}{-1\baselineskip}{1\baselineskip}{\normalfont}{Estimating the average of switchers' slopes}
\@startsection{subsection}{2}{0mm}{-1\baselineskip}{0.7\baselineskip}{\normalfont}{Target parameter}
In this section, our target parameter is
the average of the slopes of switchers' potential outcome functions, between their period-one and their period-two treatments. Hereafter, $\delta_1$ is referred to as the Average of Slopes (AS).
The AS is a local effect: it only applies to switchers, and it measures the effect of changing their treatment from its period-one to its period-two value, not of other changes of their treatment. Still, the AS can be used to point or partially identify the effect of other treatment changes under shape restrictions. First, assume that the potential outcomes are linear: for $t\in \{1,2\}$, $$Y_t(d)=Y_t(0)+B_td,$$ where $B_t$ is a slope that may vary across units and may change over time. Then, $\delta_1=E\left(B_2\middle| S=1\right)$: the AS is equal to the average, across switchers, of the slopes of their potential outcome functions at period 2. Therefore, for all $d\ne d'$, $$E(Y_2(d)-Y_2(d')|S=1)=(d-d')\delta_1:$$ under linearity, knowing the AS is sufficient to recover the ATE of any uniform treatment change among switchers. Of course, this only holds under linearity, which may not be a plausible assumption. Assume instead that $d\mapsto Y_2(d)$ is convex. Then, for any $\epsilon>0$, $$E\left(Y_2(D_2+\epsilon) - Y_2(D_2)\middle| S=1\right)\geq \epsilon \delta_1.$$ Accordingly, under convexity one can use the AS to obtain lower bounds of the effect of changing the treatment from $D_2$ to larger values than $D_2$. For instance, in fajgelbaum2020return, one can use this strategy to derive a lower bound of the effect of increasing tariffs' to even higher levels than those imposed in 2018 by the Trump administration. Under convexity, one can also use the AS to derive an upper bound of the effect of changing the treatment from $D_1$ to a lower value. And under concavity, one can derive an upper (resp. lower) bound of the effect of changing the treatment from $D_2$ (resp. $D_1$) to a larger (resp. lower) value.\footnote{See d2021nonparametric for bounds of the same kind obtained under concavity or convexity.} The AS is identified even if those linearity or convexity/concavity conditions fail, but those conditions are necessary to use the AS to identify or bound the effects of alternative policies.
\@startsection{subsection}{2}{0mm}{-1\baselineskip}{0.7\baselineskip}{\normalfont}{Identification}
The AS is identified under the two following assumptions on the design.
The estimand identifying the AS compares the outcome evolutions of switchers and stayers with the same period-one treatment. Then, this estimand requires that there be no value of $D_1$ such that only switchers have that value, as assumed in Assumption (ref). Assumption (ref) requires that there are no quasi-stayers: the treatment of all switchers changes by at least $\kappa>0$.
Define the functions: \[\mu_0(d):=E[\Delta Y|S=0,D_1=d]\quad \text{ and } \alpha_0(s,d):=E\left(\left.\frac{S}{\Delta D}\right\vert D_1=d\right)\frac{1-s}{E(1-S|D_1=d)}.\]
It readily follows from Lemma (ref) and the law of iterated expectations that $\delta_1$ is identified by $$E\left[\left(\frac{\Delta Y-E[\Delta Y|S=0,D_1]}{\Delta D}\right)\middle|S=1\right]=E\left[\left(\frac{\Delta Y-\mu_0(D_1)}{\Delta D}\right)\middle|S=1\right],$$ where the numerator of the ratio inside the expectation is a DID comparing the outcome evolutions of switchers and stayers with the same $D_1$. Theorem (ref) implies that a related estimand, $$E\left(\frac{1}{E[S]}\left(\frac{S}{\Delta D}-\alpha_0(S,D_1)\right)\left[\Delta Y-\mu_0(D_1)\right]\right),$$ also identifies $\delta_1$, and is doubly-robust in the functions $\mu_0(D_1)$ and $\alpha_0(S,D_1)$.\footnote{We are grateful to Andres Santos for pointing out to us that the AS is identified by a doubly-robust estimand.}
If there are quasi-stayers, the AS is still identified, but estimation may be challenging. For any $\eta>0$, let $S_\eta=1\{|\Delta D|>\eta\}$ be an indicator for switchers whose treatment changes by at least $\eta$ from period one to two.
If there are quasi-stayers whose treatment change is arbitrarily close to $0$ (i.e. $f_{|\Delta D||S=1}(0)>0$), the denominator of $(\Delta Y - E(\Delta Y | D_1, S=0))/\Delta D$ is close to $0$ for them. On the other hand,
so the ratio's numerator may not be close to $0$. Then, under weak conditions, $$E\left(\left|\frac{\Delta Y - E(\Delta Y | D_1, S=0)}{\Delta D}\right|\; \middle| S=1\right)=+\infty.$$ Therefore, we need to trim quasi-stayers from the estimand in Theorem (ref), and let the trimming go to 0, as in graham2012identification who consider a related estimand with some quasi-stayers. Accordingly, with quasi-stayers the AS is irregularly identified by a limiting estimand. Then, we conjecture that the AS cannot be estimated at the $\sqrt{n}-$rate, as in graham2012identification, and as is often the case with target parameters identified by limiting estimands. Therefore, we do not consider estimation of the AS with quasi stayers.
\@startsection{subsection}{2}{0mm}{-1\baselineskip}{0.7\baselineskip}{\normalfont}{Estimation and inference}
Let \[g_0(d):=E[S/\Delta D|D_1=d]\quad \text{and } p_0(d):=E[1-S|D_1=d].\] It follows from Theorem (ref) that estimating $\delta_1$ requires estimating the nuisance functions $g_0(d)$, $p_0(d)$, and $\mu_0(d)$ in a first step, before averaging the variable $$\left(\frac{S_i}{\Delta D_i}-\frac{\hat g(D_{i,1})}{\hat p(D_{i,1})}(1-S_i)\right)\left(\Delta Y_i-\hat\mu(D_{i,1})\right).$$ We propose to use cross-fitting to perform this two-step estimation. For any set $A$, let $\# A$ denote its cardinality. Let the sample be split randomly into two sub-samples, $\mathcal{I}_1$ and $\mathcal{I}_2$, and let $I_k=\# \mathcal{I}_k$ for $k=1,2$. We let $\hat g^{(2)}$, $\hat p^{(2)}$ and $\hat \mu^{(2)}$ be nonparametric estimators of $g_0$, $p_0$ and $\mu_0$ based on $\mathcal{I}_2$. Then the estimator $\hat\delta_{1,\mathsf{DR}}^{(1)}$ is computed in $\mathcal{I}_1$ as follows: \[\hat\delta_{1,\mathsf{DR}}^{(1)}=\frac{1}{n_S^{(1)}}\sum_{i\in \mathcal{I}_1}\left(\frac{S_i}{\Delta D_i}-\frac{\hat g^{(2)}(D_{i,1})}{\hat p^{(2)}(D_{i,1})}(1-S_i)\right)\left(\Delta Y_i-\hat\mu^{(2)}(D_{i,1})\right),\] where $n_S^{(1)}:=\sum_{i\in \mathcal{I}_1} S_i$ and we use the convention that $\hat\delta_{1,\mathsf{DR}}^{(1)}=0$ if $n_S^{(1)}=0$. The estimator $\hat\delta_{1,\mathsf{DR}}^{(2)}$ is computed similarly, permuting the roles of $\mathcal{I}_1$ and $\mathcal{I}_2$. Finally the cross-fitting doubly-robust estimator is computed as: \[\hat{\delta}_{1,\mathsf{DR}}=\frac{n_S^{(1)}}{n_S^{(1)}+n_S^{(2)}}\hat\delta_{1,\mathsf{DR}}^{(1)}+ \frac{n_S^{(2)}}{n_S^{(1)}+n_S^{(2)}}\hat\delta_{1,\mathsf{DR}}^{(2)}.\] We impose the following condition:
We also require that the estimators of the nonparametric nuisance functions satisfy some convergence-rate requirements. For any function $h(\cdot)$, $\hat h^{(1)}(.), \hat h^{(2)}(.)$ estimators of $h(.)$ based on subsample $\mathcal{I}_1$ and $\mathcal{I}_2$ respectively and $k\in\{1,2\}$, define: \[\left\Vert\hat{h}-h\right\Vert_{2,\mathcal{I}_k}:=\left(\frac{1}{I_k}\sum_{i\in\mathcal{I}_k}(\hat h^{(3-k)}(D_{1,i})-h(D_{1,i}))^2\right)^{1/2},\quad \left\Vert\hat h-h\right\Vert_{k,\infty}:=\sup_{d\in \mathcal{D}_1}\left\vert\hat h^{(3-k)}(d)-h(d)\right\vert.\]
As all nuisance functions are conditional expectations with respect to a single variable, these convergence-rate requirements are satisfied by classical nonparametric estimators. For instance, one can use series estimators with a polynomial order chosen by cross validation within subsample $\mathcal{I}_k$ to estimate $\hat\mu^{(k)}(d)$, $\hat{g}^{(k)}(d)$, and $\hat p^{(k)}(d)$. Results in Section 4 of andrews1991asymptotic imply that those estimators satisfy the rate conditions in Assumption (ref), as long as the population functions are twice differentiable. Note also that the mean-square convergence requirements can be established under regularity conditions and appropriate choices of tuning parameters for machine learning estimators such as random forests Biau_2012_JMLR,Wager-Athey_2018_JASA,Syrgkanis-Zampetakis_2020 and neural networks Chen_White_1999_IEEE,Farrell-etal_2021_ECMA.\footnote{We rely on uniform convergence of the propensity score to ensure that the inverse propensity score is consistent. While uniform convergence may be too strong for machine learning estimators, this condition may be relaxed using automatic debiased machine learning, as proposed in Cherno-etal_2024_wp.}
\@startsection{section}{2}{0mm}{-1\baselineskip}{1\baselineskip}{\normalfont}{Estimating a weighted average of switchers' slopes}
\@startsection{subsection}{2}{0mm}{-1\baselineskip}{0.7\baselineskip}{\normalfont}{Target parameter}
In this section, our target parameter is
a weighted average of the slopes of switchers' potential outcome functions from their period-one to their period-two treatments, where slopes receive a weight proportional to switchers' absolute treatment change from period one to two. We refer to $\delta_{2}$ as the Weighted Average of Slopes (WAS). $\delta_2=\delta_1$ if and only if switchers' slopes are uncorrelated with $|D_2-D_1|$.
The AS and WAS serve different purposes. Unlike the AS, the WAS may not be used to identify or bound the effect of other treatment changes than the actual changes that took place, even under shape restrictions on the potential outcome function. Instead, the WAS may be used to conduct a cost-benefit analysis of the treatment changes that took place from period one to two. To simplify the discussion, let us assume in the remainder of this paragraph that $D_2\ge D_1$. Assume also that the outcome is a measure of output, such as agricultural yields or wages, expressed in monetary units. Finally, assume that the treatment is costly, with a cost linear in dose, uniform across units, and known to the analyst: the cost of giving $d$ units of treatment to a unit at period $t$ is $c_t\times d$ for some known $(c_t)_{t\in \{1,2\}}$. Then, $D_2$ is beneficial relative to $D_1$ if and only if $E(Y_2(D_2)-c_2D_2)>E(Y_2(D_1)-c_2D_1)$, namely if and only if $\delta_2>c_2:$ comparing $\delta_2$ to $c_2$ is sufficient to evaluate if changing the treatment from $D_1$ to $D_2$ was beneficial.
\@startsection{subsection}{2}{0mm}{-1\baselineskip}{0.7\baselineskip}{\normalfont}{Identification}
Let $S_+=1\{D_2-D_1>0\}$ and $S_-=1\{D_2-D_1<0\}$ respectively be indicators for “switchers up” and “switchers down” and define the function: \[\gamma_0(s,d):=\frac{E(S_+|D_1=d)-E(S_-|D_1=d)}{E(1-S|D_1=d)}(1-s).\]
Here are the steps leading to Theorem (ref). One has
The fourth equality follows from adding and subtracting $Y_1(D_1)$. The fifth equality follows from the law of iterated expectations. The sixth follows from Assumption (ref) and the fact that $\Delta Y=\Delta Y(D_1)$ for stayers. The seventh follows from the law of iterated expectations, and the last follows from the definition of $\mu_0$. The numerator of the estimand identifying $\delta_2$ in the previous display is a conditional DID comparing the outcome evolutions of switchers and stayers with the same $D_1$. Following similar steps, one can show that $\delta_2$ is also identified by $$\frac{E\left[\left(S_+-S_-\right)\Delta Y\right]-E\left[\frac{E(S_+|D_1=d)-E(S_-|D_1=d)}{E(1-S|D_1=d)}(1-S)\Delta Y\right]}{E[\left\vert\Delta D\right\vert]}=\frac{E\left[\left(S_+-S_--\gamma_0(S,D_1)\right)\Delta Y\right]}{E[\left\vert\Delta D\right\vert]},$$ an estimand whose numerator is an unconditional DID with propensity score reweighting. Theorem (ref) implies that an estimand combining those two estimands, $$\frac{E\left[\left(S_+-S_--\gamma_0(S,D_1)\right)(\Delta Y-\mu_0(D_1))\right]}{E[\left\vert\Delta D\right\vert]},$$ also identifies $\delta_2$, and is doubly-robust in the functions $\mu_0(D_1)$ and $\gamma_0(S,D_1)$.
\@startsection{subsection}{2}{0mm}{-1\baselineskip}{0.7\baselineskip}{\normalfont}{Estimation and inference}
Let $h_0(d):=E[S_+-S_-|D_1=d]$. As before, split the sample into subsamples $\mathcal{I}_1$ and $\mathcal{I}_2$, and let \[\hat\delta_{2,\mathsf{DR}}^{(1)}:=\frac{1}{\Delta_S^{(1)}}\sum_{i\in \mathcal{I}_1}\left(S_{i+}-S_{i-}-\frac{\hat h^{(2)}(D_{i,1})}{\hat p^{(2)}(D_{i,1})}(1-S_i)\right) (\Delta Y_i -\hat\mu^{(2)}(D_{i,1})),\] where $\Delta_S^{(1)}:=\sum_{i\in \mathcal{I}_1} \left\vert\Delta D_i\right\vert$ and we use the convention that $\hat\delta_{2,\mathsf{DR}}^{(1)}=0$ if $\Delta_S^{(1)}=0$. The estimator $\hat\delta_{2,\mathsf{DR}}^{(2)}$ is constructed analogously, permuting the roles of $\mathcal{I}_1$ and $\mathcal{I}_2$. Finally, let \[\hat\delta_{2,\mathsf{DR}}:=\frac{\Delta_S^{(1)}}{\Delta_S^{(1)}+\Delta_S^{(2)}}\hat\delta_{2,\mathsf{DR}}^{(1)}+\frac{\Delta_S^{(2)}}{\Delta_S^{(1)}+\Delta_S^{(2)}}\hat\delta_{2,\mathsf{DR}}^{(2)}.\] We impose the following rate conditions.
Finally, we now show that under some assumptions, the asymptotic variance of the WAS estimator is lower than that of the AS estimator.
The constant effect and homoscedasticity assumptions underlying Proposition (ref) are admittedly strong, but ranking estimators' variances typically requires such conditions. The question then is whether the ranking in Proposition (ref) still holds in real-life applications, where those assumptions likely fail. In our application, the variance of $\widehat{\delta}_{1,\mathsf{DR}}$ is indeed much larger than that of $\widehat{\delta}_{2,\mathsf{DR}}$.
\@startsection{section}{2}{0mm}{-1\baselineskip}{1\baselineskip}{\normalfont}{Extensions}
\@startsection{subsection}{2}{0mm}{-1\baselineskip}{0.7\baselineskip}{\normalfont}{Instrumental-variable estimators}
In this section, we consider instances where Assumption (ref) is implausible, but one has an instrument that satisfies a similar parallel-trends condition.
\@startsection{subsubsection}{3}{0mm}{-0.8\baselineskip}{0.4\baselineskip}{\normalfont}{Notation and assumptions}
Let $Z_t$ denote the instrumental variable at period $t\in\{1,2\}$. Throughout this section, we rule out anticipatory and dynamic effects of the instrument on the treatment. Then, let $D_t(z)$ denote the unit's potential treatment at $t$ if $Z_t=z$. Let also $SC=1\{D_2(Z_2)\ne D_2(Z_1)\}$ be an indicator equal to 1 for switchers-compliers, namely units whose instrument changes from period one to two and whose treatment is affected by that change. Finally, let $\mathcal{Z}_t$ denote the support of $Z_t$.
We replace Assumption (ref) by the following assumption.\footnote{Our notation, where potential outcomes do not depend on $z$, implicitly imposes the usual exclusion restriction.}
Point 1 of Assumption (ref) requires that $Y_2(D_2(z))-Y_1(D_1(z))$, units' outcome evolutions in the counterfactual where their instrument does not change from period one to two, be mean independent of $Z_2$, conditional on $Z_1$ and $D_1$. This condition imposes some restrictions on treatment effect heterogeneity, and the goal of conditioning on $D_1$ is to alleviate those restrictions. To see this, assume that
(ref) requires that $Y_2(D_1(z))-Y_1(D_1(z))$, units' outcome evolutions in the counterfactual where their instrument and their treatment does not change from period one to two, be mean independent of $Z_2$, conditional on $Z_1$ and $D_1$. Thanks to the conditioning on $D_1$, (ref) is a standard parallel-trends assumption that does not impose any restriction on treatment effect heterogeneity, like Assumption (ref). If $D_1$ was not conditioned upon, (ref) would require parallel trends among units with different baseline treatments, which implicitly assumes homogeneous treatment effects over time, as discussed in Section (ref). Under (ref), Point 1 of Assumption (ref) holds if and only if
(ref) is a restriction on treatment effect heterogeneity across units. It requires that switching the treatment from $D_1(Z_1)$ to $D_2(Z_1)$, the natural change happening over time without a change in the instrument, has an effect that is mean independent of $Z_2$ conditional on $Z_1$ and $D_1$. Thus, it is only if (ref) fails that Point 1 of Assumption (ref) does not restrict treatment-effect heterogeneity. However, it may not be plausible to assume that Assumption (ref) holds but (ref) fails.
de2010note and hudson2017interpreting also consider IV-DID estimands, in classical designs with two periods and a binary instrument that turns on for some units at period two. Both papers introduce a “reduced-form” parallel-trends assumption similar to Point 1 of Assumption (ref), but without noting that it imposes restrictions on effects' heterogeneity, even in the simple designs considered by those papers. Turning to Point 2 of Assumption (ref), this condition requires that units' treatment evolutions under $Z_1$ be mean independent of $Z_2$, conditional on $Z_1$ and $D_1$. Because $D_1$ is conditioned upon, this parallel trends condition is equivalent to a sequential exogeneity assumption robins1986new,bojinov2020panel.
We also make the following assumptions.
Point i) of Assumption (ref) is a monotonicity assumption similar to that in Imbens94. It requires that increasing the period-two instrument weakly increases the period-two treatment. This condition is plausible when the instrument is taxes and the treatment is prices, as in our application. Point ii) requires that the instrument has a strictly positive first stage.
Assumptions (ref) and (ref) are adaptations of Assumptions (ref) and (ref) to the IV setting.
\@startsection{subsubsection}{3}{0mm}{-0.8\baselineskip}{0.4\baselineskip}{\normalfont}{Target parameter}
In this section, our target parameter is
$\delta_{\mathsf{IV}}$ is a weighted average of the slopes of compliers-switchers' period-two potential outcome functions, from their period-two treatment under their period-one instrument, to their period-two treatment under their period-two instrument. Slopes receive a weight proportional to the absolute value of compliers-switchers' treatment response to the instrument change. Under Assumption (ref), $\delta_{\mathsf{IV}}$ is equal to the reduced-form WAS effect of the instrument on the outcome $$E\left(\frac{|Z_2-Z_1|}{E(|Z_2-Z_1||Z_2\ne Z_1)}\times \frac{Y_2(D_2(Z_2))-Y_2(D_2(Z_1))}{Z_2-Z_1}\middle|Z_2\ne Z_1\right),$$ divided by the first-stage WAS effect of the instrument on the treatment $$E\left(\frac{|Z_2-Z_1|}{E(|Z_2-Z_1||Z_2\ne Z_1)}\times \frac{D_2(Z_2)-D_2(Z_1)}{Z_2-Z_1}\middle|Z_2\ne Z_1\right).$$ With a binary instrument, such that $Z_1=0$ and $Z_2 \in \{0,1\}$, our IV-WAS effect coincides with that identified in Corollary 2 of angrist2000, in a cross-sectional IV model.
We could also consider a reduced-form AS divided by a first-stage AS. The resulting target is a weighted average of the slopes $(Y_2(D_2(Z_2))-Y_2(D_2(Z_1)))/(D_2(Z_2)-D_2(Z_1))$, with weights proportional to $(D_2(Z_2)-D_2(Z_1))/ (Z_2-Z_1)$. It seems more natural to us to weight compliers-switchers' slopes by the absolute value of their first-stage than by the slope of their first-stage.\footnote{If the first-stage effect is homogenous and linear, the weights in the IV-AS effect reduce to one, and one recovers a standard AS effect. However, linearity and homogeneity of the first-stage effect are strong assumptions.}
\@startsection{subsubsection}{3}{0mm}{-0.8\baselineskip}{0.4\baselineskip}{\normalfont}{Identification}
Let $S^I=1\{Z_2-Z_1\ne 0\}$, $S^I_{+}=1\{Z_2-Z_1>0\}$, $S^I_{-}=1\{Z_2-Z_1<0\}$ and
The following condition is similar to Assumption (ref).
The estimand identifying $\delta_{\mathsf{IV}}$ is equal to the estimand identifying the reduced-form WAS effect of the instrument on the outcome controlling for $D_1$, divided by the estimand identifying the first-stage WAS effect of the instrument on the treatment controlling for $D_1$. One can show that the estimand identifying $\delta_{\mathsf{IV}}$ is doubly robust in $\gamma^I_0(S^I,Z_1,D_1)$ and $(\mu^{Y}_0(Z_1,D_1),\mu^{D}_0(Z_1,D_1))$.
\@startsection{subsubsection}{3}{0mm}{-0.8\baselineskip}{0.4\baselineskip}{\normalfont}{Estimation and inference}
Estimating $\delta_{\mathsf{IV}}$ requires estimating the nuisance functions $\mu^{Y}_0(z,d)$, $\mu^{D}_0(z,d)$, $E(S^I_+|Z_1=z,D_1=d)$, $E(S^I_-|Z_1=z,D_1=d)$, and $E(1-S^I|Z_1=z,D_1=d)$, before averaging the variables $(S^I_{i+}-S^I_{i-}-\widehat{\gamma}^I(S^I_i,Z_{i,1},D_{i,1}))(\Delta Y_i- \widehat{\mu}^{Y}_0(Z_{i,1},D_{i,1}))$ and $(S^I_{i+}-S^I_{i-}-\widehat{\gamma}^I(S^I_i,Z_{i,1},D_{i,1}))(\Delta D_i-\widehat{\mu}^{D}_0 (Z_{i,1},D_{i,1}))$, and taking the ratio between these two averages. We propose to use cross-fitting to perform this two-step estimation, and we let $\widehat{\delta}_{\mathsf{IV},\mathsf{DR}}$ denote the resulting estimator. As the definition of $\widehat{\delta}_{\mathsf{IV},\mathsf{DR}}$ is similar to that of $\widehat{\delta}_{2,\mathsf{DR}}$, the full definition is omitted to preserve space. Since all the nuisance functions in Theorem (ref) are conditional expectations with respect to two variables, one can still use series estimators with a polynomial order chosen by cross validation within each subsample to estimate them. Results in Section 4 of andrews1991asymptotic imply that those estimators converge fast enough to have that $\widehat{\delta}_{\mathsf{IV},\mathsf{DR}}$ is $\sqrt{n}-$consistent and asymptotically normal, as long as the population functions are twice differentiable. Finally, for $U\in\{D,Y\}$, let
Then, let $\psi_{\mathsf{IV}}=(\psi_{Y}-\delta_{\mathsf{IV}}\psi_{D})/\delta_{D}$. Under technical conditions similar to those in Section (ref), one can show that $\psi_{\mathsf{IV}}$ is the influence function of $\widehat{\delta}_{\mathsf{IV},\mathsf{DR}}$.
\@startsection{subsection}{2}{0mm}{-1\baselineskip}{0.7\baselineskip}{\normalfont}{Alternative target parameters}
One may be interested in other parameters than the AS and WAS we consider in the paper, like the average dose-slope function $$(d,d')\mapsto \text{TE}(d,d'):=E\left(\frac{Y_2(d) - Y_2(d')}{d-d'}\right),$$ or the average marginal effect $E\left(Y'_2(D_2)\right)$. What is the appeal of considering the AS and the WAS, namely averages of $\text{TE}(d_1,d_2|d_1,d_2)$, rather than those other parameters? Conditional on $D_1=d_1,D_2=d_2$, $Y_2(d_2)$ is observed, so estimating $\text{TE}(d_1,d_2|d_1,d_2)$ only requires estimating $Y_2(d_1)$, switchers' counterfactual outcomes at period $2$ if their treatment had not changed. By definition, $Y_1(d_1)$ is observed. If the data contains a third period $0$ and the treatment of some units does not change from period $0$ to $1$, then $Y_0(d_1)$ is also observed for some switchers and stayers. Then, as explained in Section (ref) below, one can run a pre-trends test of Assumption (ref), by comparing the outcome evolutions of switchers and stayers from period $0$ to $1$. This shows that averages of $\text{TE}(d_1,d_2|d_1,d_2)$ can be identified under a parallel-trends assumption whose plausibility can be assessed via a pre-trends test. On the other hand, estimating $\text{TE}(d,d')$ requires estimating, for most units, two unobserved counterfactual outcomes. This cannot be achieved under an assumption whose plausibility can be assessed via a pre-trends test, as we only observe one potential outcome at each date. Similarly, estimating $E[(Y_2(D_2) - Y_2(d'))/(D_2-d')]$ requires estimating $Y_2(d')$, which cannot be achieved under an assumption whose plausibility can be assessed via a pre-trends test, because $Y_1(d')$ is not observed for all units. As $Y'_2(D_2)=\lim_{d'\rightarrow D_2}(Y_2(D_2) - Y_2(d'))/(D_2-d')$, this concern also applies to $E\left(Y'_2(D_2)\right)$.
\@startsection{subsection}{2}{0mm}{-1\baselineskip}{0.7\baselineskip}{\normalfont}{Other extensions}
We return to the case where the treatment satisfies a parallel-trends condition.
\@startsection{subsubsection}{3}{0mm}{-0.8\baselineskip}{0.4\baselineskip}{\normalfont}{Estimators with more than two time periods}
This section sketches the extension of our estimators to datasets with three periods or more. For further details, see Section (ref) of the web appendix. With $T>2$ time periods, we propose a simple extension of our estimators. Let $\widehat{\delta}_{1,t,\mathsf{DR}}$ and $\widehat{\delta}_{2,t,\mathsf{DR}}$ be defined exactly as $\hat{\delta}_{1,\mathsf{DR}}$ and $\hat{\delta}_{2,\mathsf{DR}}$ but on a pair of consecutive time periods $(t-1,t)$. Then, we consider a weighted average across $t$ of $\widehat{\delta}_{1,t,\mathsf{DR}}$, with weights proportional to the proportion of switchers from $t-1$ to $t$, thus yielding an AS estimator with more than two periods, denoted $\widehat{\delta}^{T\geq 3}_{1}$. We also consider a weighted average across $t$ of $\widehat{\delta}_{2,t,\mathsf{DR}}$, with weights proportional to the average absolute treatment change from $t-1$ to $t$, thus yielding a WAS estimator with more than two periods, denoted $\widehat{\delta}^{T\geq 3}_{2}$.
Those estimators are consistent under the following parallel-trends assumption: $t-1$-to-$t$ switchers and stayers with the same $t-1$ treatment would have experienced the same average $t-1$-to-$t$ outcome evolutions if switchers had not switched. Because it is conditional on the lagged treatment, this assumption cannot be “chained” across pairs of periods: it only requires that some groups be on parallel trends over consecutive periods, not over the entire duration of the panel. We propose “pre-trends” or “placebo” estimators, to assess the plausibility of that assumption. The placebo AS and WAS estimators are analogous to $\widehat{\delta}^{T\geq 3}_{1}$ and $\widehat{\delta}^{T\geq 3}_{2}$, but they compare the $t-2$-to-$t-1$ outcome evolutions of $t-1$-to-$t$ switchers and stayers, restricting the sample to $t-2$-to-$t-1$ stayers to avoid that those comparisons be contaminated by the treatment's effect. If one finds that from $t-2$-to-$t-1$, $t-1$-to-$t$ switchers and stayers are on parallel trends, this lends credibility to the parallel-trends assumption underlying $\widehat{\delta}^{T\geq 3}_{1}$ and $\widehat{\delta}^{T\geq 3}_{2}$.
With $T>2$ time periods, the assumption that units' potential outcomes does not depend on their past treatments is of essence. If there are dynamic treatment effects, $\widehat{\delta}^{T\geq 3}_{1}$ and $\widehat{\delta}^{T\geq 3}_{2}$ may be biased. Those estimators compare the outcome evolutions of $t-1$-to-$t$ switchers and stayers. But there may be, say, $t-1$-to-$t$ stayers whose treatment changed from $t-2$ to $t-1$. If lagged treatments affect units' current outcome, that change could still affect the $t-1$-to-$t$ outcome evolution of those stayers, thus leading them to violate the parallel-trends assumption underlying $\widehat{\delta}^{T\geq 3}_{1}$ and $\widehat{\delta}^{T\geq 3}_{2}$. To mitigate this concern, we propose the following robustness check. One can recompute $\widehat{\delta}^{T\geq 3}_{1}$ and $\widehat{\delta}^{T\geq 3}_{2}$, restricting, for each pair of consecutive periods $(t-1,t)$, the estimation sample to $t-2$-to-$t-1$ stayers, as in the placebo analysis. We show that the resulting estimator is robust to dynamic effects up to one treatment lag. Similarly, if one wants to allow for effects of the first and second treatment lags on the outcome, one just needs to restrict the estimation sample to $t-3$-to-$t-1$ stayers. However, the more robustness to dynamic effects one would like to have, the smaller the estimation sample becomes.
\@startsection{subsubsection}{3}{0mm}{-0.8\baselineskip}{0.4\baselineskip}{\normalfont}{Identification and estimation with covariates}
This section sketches how our estimators can be generalized to incorporate covariates, see Section (ref) of the web appendix for more details. Consider for simplicity the case with two periods, $T=2$, where $X_t$ is a $c$-dimensional vector of covariates. We show in Section (ref) of the web appendix that the AS and WAS can be identified and estimated under an alternative parallel trends assumption that holds conditional on the first-period covariates $X_1$ (Assumption (ref)). Because our estimation and inference methods are fully nonparametric, they can be readily applied to the case with covariates after redefining the nuisance functions to depend on the covariates, for instance, $\mu_0(x,d)=E(\Delta Y|X_1=x,D_1=d,S=0)$ and so on.
\@startsection{section}{2}{0mm}{-1\baselineskip}{1\baselineskip}{\normalfont}{Application}
\@startsection{subsection}{2}{0mm}{-1\baselineskip}{0.7\baselineskip}{\normalfont}{Data and design}
We use the yearly 1966-to-2008 panel dataset of li2014gasoline, covering 48 US states (Alaska and Hawaii are excluded). This is a long panel, but recall that our estimators assume parallel trends only across pairs of consecutive years. For each state$\times$year cell $(i,t)$, the data contains $Z_{i,t}$, the total (state plus federal) gasoline tax in cents per gallon, $D_{i,t}$, the log tax-inclusive price of gasoline, and $Y_{i,t}$, the log gasoline consumption per adult. We seek to estimate the effect of taxes on gasoline consumption and prices, and the price-elasticity of gasoline consumption.\footnote{Instead, li2014gasoline jointly estimate the effect of gasoline taxes and tax-exclusive prices on consumption, using a TWFE regression with two treatments. Between each pair of periods, the tax-exclusive price changes in all states, so this treatment does not have stayers and its effect cannot be estimated using our estimators.}
Let $\mathcal{S}$ be the set of switching $(i,t)$ cells such that $Z_{i,t}\ne Z_{i,t-1}$ but $Z_{i',t}=Z_{i',t-1}$ for some $i'$. The second condition drops seven pairs of consecutive periods between which the federal gasoline tax changed, thus implying that taxes changed in all states. $\mathcal{S}$ includes 384 cells, namely 19% of the 2,016 state$\times$year cells for which $Z_{i,t}-Z_{i,t-1}$ can be computed. Table (ref) in the web appendix shows that cells in $\mathcal{S}$ are similar to other cells on a number of observable characteristics. Thus, there is no strong indication that the effects below apply to a selected subgroup.
As an example, the top panel of Figure (ref) below shows the distribution of $Z_{i,1987}$ for 1987-to-1988 stayers, while the bottom panel shows the distribution for 1987-to-1988 switchers. The figure shows that there are many values of $Z_{i,1987}$ such that only one or two states have that value, so $Z_{i,1987}$ is close to being continuously distributed. Moreover, all switchers $i$ are such that $$\min_{i':Z_{i',1988}=Z_{i',1987}}Z_{i',1987}\leq Z_{i,1987}\leq \max_{i':Z_{i',1988}=Z_{i',1988}}Z_{i',1987}.$$ Thus, Assumption (ref) seems to hold for this pair of years. $(1987,1988)$ is not atypical: there are many other years where $Z_{i,t}$ is close to being continuously distributed, and almost 95% of cells in $\mathcal{S}$ are such that $\min_{i':Z_{i',t}=Z_{i',t-1}}Z_{i',t-1}\leq Z_{i,t-1}\leq \max_{i':Z_{i',t}=Z_{i',t-1}}Z_{i',t-1}$. Dropping the few cells that do not satisfy this condition barely changes the results presented below.
Turning to the distribution of the tax changes $Z_{i,t}-Z_{i,t-1}$, 90% of the cells in $\mathcal{S}$ experience an increase in their taxes, and 10% experience a decrease. The average value of $|Z_{i,t}-Z_{i,t-1}|$ is equal to 1.61 cents, while prior to the tax change, switchers' average gasoline price is equal to 112 cents: our estimators leverage small changes in taxes relative to gasoline prices. Finally, $\min_{(i,t)\in \mathcal{S}}|Z_{i,t}-Z_{i,t-1}| = 0.05$ : some switchers experience a very small change in their taxes.
\@startsection{subsection}{2}{0mm}{-1\baselineskip}{0.7\baselineskip}{\normalfont}{Reduced-form and first-stage estimates}
Table (ref) below shows the AS and WAS estimates of the reduced-form (Panel A) and first-stage (Panel B) effects of taxes on gasoline quantities and prices. We control for lagged prices $D_{t-1}$, to ensure that the resulting IV is robust to heterogeneous effects over time, as explained in Section (ref). Estimates are computed with cross-fitting, using 10 splits. We use a polynomial of order 1 in $(Z_{t-1},D_{t-1})$ to estimate $E(\Delta Y_t|Z_{t-1},D_{t-1}, S_t=0)$, $E(\Delta D_t|Z_{t-1},D_{t-1},S_t=0)$, $E[S/\Delta D|Z_{t-1},D_{t-1}]$, $P(S_{+,t}=1|Z_{t-1},D_{t-1})$, $P(S_{-,t}=1|Z_{t-1},D_{t-1})$, and $P(S_{t}=0|Z_{t-1},D_{t-1})$. 10-folds cross-validation selects a polynomial of order one for all conditional expectations, except for $E(\Delta D_t|Z_{t-1},D_{t-1},S_t=0)$ where it selects a polynomial of order two. Thus, polynomials of order one are in line with those selected by cross validation.\footnote{In a previous version of this paper de2022difference, we had shown results very similar to those below with polynomials of order two and without cross fitting. With polynomials of order two and cross fitting, estimates become noisy: due to the small sample size, $\widehat{P}(S_{t}=0|Z_{t-1},D_{t-1})$ is close to zero for some $(i,t)$.} Standard errors clustered at the state level, computed following (ref) and (ref) in the web appendix, are shown between parentheses. Estimations use 1632 (48$\times$ 35) first-difference observations: 7 periods are excluded as they do not have stayers. The last line of each panel shows the p-value of a test that AS=WAS.
In Panel A, the AS estimate indicates that increasing gasoline tax by 1 cent decreases quantities consumed by 0.43 percent on average for the switchers. That effect is insignificant. The WAS estimate indicates that increasing gasoline tax by 1 cent decreases quantities consumed by 0.36 percent, an effect that is significant. In Panel B, the AS estimate of the first-stage effect is positive and insignificant. The WAS estimate is significant, and indicates that if gasoline tax increases by 1 cent on average, prices increase by 0.58 percent. In both panels, the AS estimates are close to, and not significantly different from, the WAS estimates, but they are insignificant because they are markedly less precise: their standard errors are almost three times larger, as predicted by Proposition (ref). Therefore, in what follows we focus on the WAS estimates.
\@startsection{subsection}{2}{0mm}{-1\baselineskip}{0.7\baselineskip}{\normalfont}{Placebo analysis}
Table (ref) below shows “pre-trends” or “placebo” WAS estimates of the reduced-form and first-stage effects. The placebos are analogous to the actual estimates, but they replace $\Delta Y_t$ by $\Delta Y_{t-1}$, and they restrict the sample to states whose taxes did not change between $t-2$ and $t-1$. This reduces the sample size by around a third. The placebos are small and insignificant, both for quantities and prices. This shows that before switchers change their gasoline taxes, switchers' and stayers' consumption and prices do not follow detectably different evolutions. In Table (ref) of the web appendix we reestimate the WAS in the placebo subsample. In that subsample, the WAS remains valid if the first lag of taxes affect current gasoline prices and quantities, as discussed in Section (ref). The estimates therein are close to those in Table (ref).
\@startsection{subsection}{2}{0mm}{-1\baselineskip}{0.7\baselineskip}{\normalfont}{Comparison with estimates robust to unrestricted dynamic effects of taxes}
While results seem robust to allowing for dynamic effects up to one lag, one may want to allow for unrestricted dynamic effects. For that purpose, we re-estimate the effect of current taxes on consumption and prices, using the estimators of de2020difference. Those estimators are applicable to continuous treatments, and they are robust to dynamic effects up to any lag. Results are shown in Table (ref) of the web appendix. While estimates are of the same sign as those in Table (ref), only 215 $(g,t)$ cells are used in the estimation. With unrestricted dynamic effects, instead of comparing states whose taxes change/do not change from $t-1$ to $t$, DID estimators need to compare states whose taxes change for the first time at $t$ to states whose taxes have not changed yet at $t$. This reduces the estimation sample to 46 first-time-switcher cells, to which the estimated effects apply, and 169 never-switcher cells used as the control group. Accordingly, estimated effects are insignificant: the standard error of the estimated effect on consumption is 13 times larger than that of the WAS estimate in Table (ref),\footnote{The target parameters in de2020difference are generalizations of the WAS to models allowing for unrestricted dynamic effects.} while the standard error of the estimated effect on prices is more than nine times larger. Thus, in this application allowing for unrestricted dynamic effects comes with large costs in terms of external validity and statistical precision. As the data is at the yearly level, it may be reasonable to rule out effects of taxes in previous years on current consumption and prices, at least for taxes that prevailed two years or more ago.
\@startsection{subsection}{2}{0mm}{-1\baselineskip}{0.7\baselineskip}{\normalfont}{IV estimate}
Table (ref) shows an IV-WAS estimate of the price-elasticity of gasoline consumption. As the instrument's first stage is not very strong and the sample only has 48 observations, asymptotic approximations may not be reliable for inference. In line with that conjecture, the bootstrap distribution of the estimator in Table (ref) is non-normal, with some outliers. Therefore, for inference we use a percentile bootstrap confidence interval, clustered at the state level. Reassuringly, that confidence interval has close to nominal coverage in simulations tailored to our application.\footnote{Here is the DGP used in our simulations. We estimate TWFE regressions of $Y_{i,t}$ on state and year fixed effects and $Z_{i,t}$, and of $D_{i,t}$ on state and year fixed effects and $Z_{i,t}$. We let $\hat{\gamma}^Y_i+\hat{\lambda}^Y_t+\hat{\beta}^YZ_{i,t}+\epsilon^Y_{i,t}$ and $\hat{\gamma}^D_i+\hat{\lambda}^D_t+\hat{\beta}^DZ_{i,t}+\epsilon^D_{i,t}$ denote the resulting regression decompositions. In each simulation, the simulated instrument is just the actual instrument, while the simulated outcomes and treatments are respectively equal to $Y^s_{i,t}=\hat{\gamma}^Y_i+\hat{\lambda}^Y_t+\hat{\beta}^YZ_{i,t}+\epsilon^{Y,s}_{i,t},$ and $D^s_{i,t}=\hat{\gamma}^D_i+\hat{\lambda}^D_t+\hat{\beta}^DZ_{i,t}+\epsilon^{D,s}_{i,t},$ where the vector of simulated residuals $(\epsilon^{Y,s}_{i,1},..., \epsilon^{Y,s}_{i,T},\epsilon^{D,s}_{i,1},..., \epsilon^{D,s}_{i,T})$ is drawn at random and with replacement from the estimated vectors of residuals $((\epsilon^{Y}_{i',1},..., \epsilon^{Y}_{i',T},\epsilon^{D}_{i',1},..., \epsilon^{D}_{i',T}))_{i'\in \{1,...,n\}}$. Thus, the first-stage and reduced-form effects, the correlation between the reduced-form and first-stage residuals, and the residuals' serial correlation are the same as in the sample.} The IV-WAS estimate is negative and significant. It is more negative than, but not significantly different from, estimates reported in recent literature coglianese2017anticipation.
\@startsection{section}{2}{0mm}{-1\baselineskip}{1\baselineskip}{\normalfont}{Conclusion}
We propose new, doubly-robust difference-in-difference (DID) estimators for treatments continuously distributed at every time period. We assume that between pairs of consecutive periods, the treatment of some units, the switchers, changes, while the treatment of other units, the stayers, does not change. We propose a parallel-trends assumption on the outcome evolution of switchers and stayers with the same baseline treatment. Under that assumption, two target parameters can be estimated. Our first target is the average slope of switchers' period-two potential outcome function, from their period-one to their period-two treatment, referred to as the AS. Our second target is a weighted average of switchers' slopes, where switchers receive a weight proportional to the absolute value of their treatment change, referred to as the WAS. The AS and WAS serve different purposes. When it comes to estimation, the WAS has two advantages. First, it can be estimated at the parametric rate even if units can experience an arbitrarily small treatment change. Second, under some conditions, the asymptotic variance of the WAS estimator is strictly lower than that of the AS. In our application, we use US-state-level panel data to estimate the effect of gasoline taxes on gasoline consumption. The standard error of the WAS is almost three times smaller than that of the AS, and the two estimates are close.
We also consider the instrumental-variable case, as there are instances where units experiencing/not experiencing a treatment change are unlikely to be on parallel trends, but one has at hand an instrument such that units experiencing/not experiencing an instrument change are more likely to be on parallel trends. Then, we propose widely applicable IV-DID estimators.
Throughout, we assume that there are some stayers, namely units whose treatment does not change between consecutive periods. de2023nostayers discuss the extension of the results in this paper to applications without stayers.