The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.
84,165 characters
Difference-in-Differences for Continuous Treatments and Instruments with Stayers
\title{Difference-in-Differences for Continuous Treatments and Instruments with Stayers\thanks{We thank Matias Cattaneo, Max Farrell, Joachim Freyberger, Thomas Le Barbanchon, and Andres Santos for helpful comments. Chaisemartin was funded by the European Union (ERC, REALLYCREDIBLE, GA Number 101043899). Views and opinions expressed are those of the authors and do not reflect those of the European Union or the European Research Council Executive Agency. Neither the European Union nor the granting authority can be held responsible for them.}}
\author{Cl\'{e}ment de Chaisemartin\thanks{Sciences Po Paris, [email removed]
} \and Xavier D'Haultf\oe{}uille
\thanks{CREST-ENSAE, [email removed].}
\and F\'{e}lix Pasquier\thanks{CREST-ENSAE, [email removed].} \and
Doulo Sow\thanks{Sciences Po, [email removed].} \and Gonzalo
Vazquez-Bare\thanks{University of California, Santa Barbara,
[email removed].} }
\maketitle ~\vspace{-1.5cm}
\begin{abstract}
When one studies the effects of taxes, tariffs, or prices using panel data, the treatment is often continuously distributed in every period. We propose difference-in-differences (DID) estimators for such cases. We assume that between consecutive periods, the treatment of some units, the switchers, changes, while the treatment of other units, the stayers, remains constant. We show that under a parallel-trends assumption, the slopes of switchers' potential outcomes are nonparametrically identified by difference-in-differences estimands comparing the outcome evolutions of switchers and stayers with the same baseline treatment. Controlling for the baseline treatment ensures that our estimands remain valid if the treatment's effect changes over time. We consider two weighted averages of switchers' slopes, and discuss their respective advantages. For each weighted average, we propose a doubly-robust, nonparametric, and $\sqrt{n}$-consistent estimator. We generalize our results to the instrumental-variable case. We apply our method to estimate the price-elasticity of gasoline consumption. \\
\textbf{Keywords:} differences-in-differences, continuous treatment, instrumental variable.
\end{abstract}
\newpage
\@startsection{section}{2}{0mm}{-1\baselineskip}{1\baselineskip}{\normalfont\large\bfseries}{Introduction}
To estimate a treatment's effect, researchers often run a two-way fixed
effects (TWFE) regression:
$$Y_{i,t}=\alpha_i+\gamma_t+\beta_{TWFE}D_{i,t}+u_{i,t},$$
where $D_{i,t}$ is the treatment of unit $i$ at time $t$, and $\alpha_i$ and
$\gamma_t$ are unit and period fixed effects. \cite{de2020difference} find
that 26 of the 100 most cited papers published by the American Economic
Review from 2015 to 2019 estimate TWFE regressions. \cite{chiu2023and} find
that 52 papers published by the top three political science journals from
2017 to 2022 estimate a TWFE regression. Researchers have long believed that
those regressions only rely on a parallel-trends assumption, requiring that
without treatment, all units would have experienced the same outcome
evolutions. However, recent papers have shown that under that assumption
alone, TWFE regressions may estimate a non-convex sum of treatment effects
across units and periods \cite[see,
e.g.,][]{deChaisemartin2018twowayfev4arxiv,goodman2021difference,borusyak2020revisiting}. Then,
$\beta_{TWFE}$ could be, say, negative, even if the effect is
positive for every $(i,t)$.
\medskip
In reaction to this finding, several ``heterogeneity-robust
difference-in-difference'' estimators, consistent
for convex combinations of $(i,t)$-specific effects under a parallel-trends
assumption, have been proposed.
However, none can be applied to treatments continuously distributed at every
period. Yet, such treatments are commonly
found. In economics, taxes \citep[see][]{li2014gasoline} or tariffs
\citep[see][]{fajgelbaum2020return} are often continuously distributed at all
periods. Similarly, \cite{peterson2021paper}, \cite{aklin2021side}, and \cite{mcnamara2022estimating} are examples of papers in political sciences and epidemiology studying treatments
continuous in all periods.
With continuous treatments, due to the lack of an untreated group, TWFE regressions often estimate a highly non-convex combination of effects. For instance, with two time periods,
exactly 50\% of $(i,t)$-specific effects are weighted negatively \citep{de2021more}. Then, proposing robust estimators for such treatments is particularly important.
\medskip
We assume that we have a two-periods panel data set. From period one to two,
the treatment of some units, referred to as the switchers, changes, while the
treatment of some other units, the stayers, does not change.
We propose a novel parallel-trends assumption, which requires that switchers
and stayers with the same period-one treatment would have experienced the
same average outcome evolutions from period one to two, if switchers'
treatment had not changed. Our assumption has two desirable
features. First, because it conditions on units' period-one treatment, it
does not restrict effects' heterogeneity. Instead, as explained later,
assuming parallel trends for switchers and stayers with different period-one
treatments would essentially amount to assuming that the treatment effect is
constant over time. Second, when another prior period of data, period zero,
is available, the plausibility of our assumption can be assessed via a ``pre-trends'' or ``placebo'' test,
comparing the period-zero-to-one outcome evolutions of period-one-to-two
switchers and stayers.\footnote{As explained later, the pre-trends estimator has
to be computed in the subsample of units whose treatment does not change from
period zero to one. Then, one may also restrict the main estimators to that
subsample.} Thus, one can assess if switchers and stayers are on parallel
trends \textit{before} switchers' treatment changes.
\medskip
Then, we consider two target parameters. The first is the average
slope of switchers' period-two potential outcome function, from their
period-one to their period-two treatment, referred to as the Average of
Slopes (AS). The second is a weighted average of switchers' slopes,
where switchers receive a weight proportional to the absolute value of their
treatment change, referred to as the Weighted Average of Slopes (WAS). The AS
and WAS serve different purposes. Under shape restrictions on the potential
outcome function, the AS can be used to identify or bound the effect of other
treatment changes than those that took place. Instead, the WAS can be used to
conduct a cost-benefit analysis of the treatment changes that took place.
Under our parallel-trends assumption, whose plausibility can be assessed using a pre-trends test, the AS and the WAS are
identified by DID estimands that compare the outcome evolution of switchers
and stayers with the same period-one treatment. This is in contrast with other
targets considered in the literature like the dose-response function or the average marginal effect, which can
only be identified under assumptions whose plausibility cannot be assessed via a pre-trends test.
\medskip
Turning to estimation, the challenge with a continuous treatment is that the
sample does not contain switchers and stayers with the same period-one
treatment. In other words, the estimands identifying the AS and WAS depend on
nonparametric regression functions. Therefore, we derive doubly-robust moment
conditions identifying the AS and WAS, and we propose doubly-robust
estimators. We show that the AS estimator converges at the parametric rate, provided
switchers cannot experience arbitrarily small treatment changes. We then
show that the WAS can always be estimated at the parametric rate. Under some
conditions, the asymptotic variance of the WAS estimator is strictly lower
than that of the AS estimator.
\medskip
We extend those results in several directions. First, we consider the instrumental-variable (IV) case, where one makes a parallel-trends assumption with respect to an instrument rather than the treatment. For instance, one may be interested in estimating the price-elasticity of a good. If prices respond to demand shocks, the counterfactual consumption trends of units experiencing and not experiencing a price change may not be the same, so a parallel-trends assumption with respect to prices may not hold. But instead, one can make a parallel-trends assumption with respect to taxes, used as an instrument for prices. We start by noting that parallel-trends with respect to an instrument restricts treatment-effect heterogeneity. Then, we show that under that assumption, the reduced-form WAS effect of the instrument on the outcome divided by the first-stage WAS effect is equal to a weighted average of switchers' outcome-slope with respect to the treatment, where switchers with a larger first-stage effect receive more weight.
\medskip
We also consider other extensions. First, we consider
applications with more than two periods. Then, our estimators start by comparing, for every pair of consecutive time periods $(t-1,t)$, the outcome evolutions of units whose treatment switches / does not switch from $t-1$ to $t$, controlling for their period-$t-1$ treatment. Then, they aggregate those comparisons across $t$. Second, we propose a pre-trends estimator to assess the plausibility of our parallel-trends assumption. Essentially, instead of comparing the $t-1$-to-$t$ outcome evolutions of $t-1$-to-$t$ switchers and stayers, it compares their $t-2$-to-$t-1$ outcome evolutions. Third, with more than two
periods, our estimators rule out dynamic effects: they assume that units'
outcome at period $t$ only depends on their period-$t$ treatment, not on
prior treatments. To alleviate this concern, we propose a modified
version of our estimators, robust to dynamic effects up to a pre-specified
treatment lag. Finally, we propose estimators allowing for control
variables. The \texttt{did\_multiplegt\_stat} Stata package computes
our estimators.
\medskip
As an illustration, we use the yearly, 1966 to 2008 US state-level panel
dataset of \cite{li2014gasoline} to estimate the effect of gasoline taxes on
gasoline consumption and prices. Using the WAS estimators, we find a
significantly negative effect of taxes on gasoline consumption, and a
significantly positive effect on prices. The AS estimators are close to, and
not significantly different from, the WAS estimators, but they are insignificant because they are
markedly less precise: their standard errors are almost three times larger.
We compute an IV estimator of the price
elasticity of gasoline consumption, and find a significantly negative
elasticity of -0.66, which is more negative than, but not significantly different
from, that reported in recent literature \citep[see,
e.g.,][]{coglianese2017anticipation}. Our pre-trend estimators are small and
insignificant, thus lending credibility to our parallel-trends
assumption.
\@startsection{subsection}{2}{0mm}{-1\baselineskip}{0.7\baselineskip}{\normalfont\normalsize\bfseries}*{Related Literature.}
Our paper builds upon several papers in the panel data literature.
\cite{chamberlain1982multivariate} seems to be the first paper to have
proposed an estimator of the AS parameter. Under the assumption of no
counterfactual time trend, the estimator therein is a before-after estimator.
Our paper is also closely related to the work of
\cite{graham2012identification}, who propose DID estimators of the AS (see
their Equation (21)) when the treatment is continuously distributed at every
time period. Their estimators rely on a linear effect assumption and assume
that units experience the same evolution of their treatment effect over time.
By contrast, our estimator of the AS does not place any restriction on
treatment effects. But our main contribution to this literature is to
introduce the WAS, and to contrast the different purposes served by the two
parameters.
\medskip
With respect to the aforementioned heterogeneity-robust DID literature, we make two contributions. First, in the non-IV case we propose estimators that can be used even if units receive heterogeneous treatment doses at baseline. Thus we complement previous literature, that has mostly focused on the case where all units receive the same dose at baseline \citep[see][]{abraham2018,callaway2020difference,liu2024practical,borusyak2020revisiting,de2024two,callaway2021difference}. One exception predating this paper is \cite{deChaisemartin2018twowayfev4arxiv}, who, in their web appendix, consider designs where units receive heterogeneous doses of a non-binary discrete treatment at baseline. With a discrete treatment, they propose DID estimators comparing switchers and stayers with the exact same baseline treatment. This solution is no longer applicable with a continuously distributed treatment, hence the estimation challenge we tackle in this paper. Another exception, contemporaneous of this paper, is \cite{de2020difference}, who propose estimators robust to dynamic effects up to any lag. While they focus mostly on discrete treatments, in Section 1.10 of their web appendix they build upon the current paper to sketch an extension of their estimators to continuous treatments. Contrary to here, the estimators therein are parametric and not doubly-robust, and the estimators' asymptotic distribution is not derived. Moreover, with unrestricted dynamic effects, instead of comparing $t-1$-to-$t$ switchers and stayers, DID estimators need to compare units switching treatment for the first time at $t$ to units whose treatment has not switched yet at $t$. This reduces the sample of units that can be used in the estimation, which may come with high costs in terms of external validity and statistical precision. We illustrate this in our empirical application. There, the estimators of \cite{de2020difference} apply to a sample more than eight times smaller than the estimators in this paper. While point estimates are similar, the standard errors of the estimators of \cite{de2020difference} are more than nine times larger than the standard errors of the estimators proposed in this paper. Thus, this paper's estimators may be appealing when it is a priori plausible to rule out dynamic effects, at least up to a pre-specified treatment lag, and/or when the tests of the no-dynamic effects assumption recently proposed by \cite{liu2024practical} and \cite{de2025treatmenteffect} are not rejected.\footnote{On the other hand, our results are not directly related to those in \citet{TVB}. That paper considers setups where the switchers and stayers start from different treatment levels, but focuses on designs with discrete rather than continuous treatments. Moreover, they analyze the performance of standard TWFE estimators, while we propose nonparametric heterogeneity-robust DID estimators.}
\medskip
Second, in the IV case we extend \cite{de2010note} and \cite{hudson2017interpreting}, who consider simple designs with two periods and a binary instrument and treatment. Our paper is also the first to note that parallel-trends with respect to an instrument restricts treatment-effect heterogeneity.
\paragraph{Organization of the paper.} Section \ref{sec:setup} presents
the set-up, our main assumptions, and a building-block identification result.
Section \ref{sec:AS} considers identification and estimation of the AS. Section
\ref{sec:WAS} turns to the WAS. Section \ref{sec:extensions} presents some
extensions. Finally, Section \ref{sec:appli}
presents our application. The web appendix also collects all proofs.
\@startsection{section}{2}{0mm}{-1\baselineskip}{1\baselineskip}{\normalfont\large\bfseries}{Set-up, assumptions, and building-block identification result}
\label{sec:setup}
\@startsection{subsection}{2}{0mm}{-1\baselineskip}{0.7\baselineskip}{\normalfont\normalsize\bfseries}{Set-up}
A representative unit is drawn from an infinite super population, and
observed at two time periods. This unit could be an individual or a firm, but
it could also be a geographical unit, like a county or a region.\footnote{In
that case, one may want to weight the estimation by counties' or regions'
populations. Extending the estimators we propose to allow for such weighting
is a mechanical extension.} All expectations below are taken with respect to
the distribution of variables in the super population. We are interested in
the effect of a continuous and scalar treatment variable on that unit's
outcome. Let $D_1$ (resp. $D_2$) denote the unit's treatment at period 1
(resp. 2), and let $\mathcal{D}_1$ (resp. $\mathcal{D}_2$) be its support. Let $S=1\{D_2\ne
D_1\}$ be an indicator equal to 1 if the unit's treatment changes from period
one to two, i.e. if they are a switcher.
\medskip
For all $(d_1,d_2)\in \mathcal{D}_1\times \mathcal{D}_2$, let $Y_1(d_1,d_2)$ and $Y_2(d_1,d_2)$
respectively denote the unit's potential outcomes at periods 1 and 2 if $(D_1,D_2)=(d_1,d_2)$, and let $Y_1$ and $Y_2$ denote their observed outcomes.
\medskip
In what follows, all equalities and inequalities involving random variables
are required to hold almost surely. Finally, for any random variable observed
at the two time periods $(X_1,X_2)$, let $\Delta X=X_2-X_1$ denote the change
of $X$ from period $1$ to $2$.
\@startsection{subsection}{2}{0mm}{-1\baselineskip}{0.7\baselineskip}{\normalfont\normalsize\bfseries}{Assumptions}
We make the following assumptions.
\begin{hyp}\label{hyp:noanticipation} (No anticipation)
For all $(d_1,d_2)\in \mathcal{D}_1\times \mathcal{D}_2$, $Y_1(d_1,d_2)=Y_1(d_1).$
\end{hyp}
\begin{hyp}\label{hyp:nodynamic} (No dynamic effects)
For all $(d_1,d_2)\in \mathcal{D}_1\times \mathcal{D}_2$, $Y_2(d_1,d_2)=Y_2(d_2).$
\end{hyp}
Assumption \ref{hyp:noanticipation} requires that the period-one outcome does not depend on units' period-two treatment, thus ruling out anticipation effects, a commonly-made
assumption in the DID literature. Assumption \ref{hyp:nodynamic} requires that the period-two outcome does not depend on units' period-one treatment, thus ruling out dynamic effects. This is
not of essence in the two-periods case we focus on for now. This is of essence
when we extend our estimators to the case where $T>2$, as we will explain in
Section \ref{sec:extensions}. Therein, we propose a modified version of our
estimators, which partly alleviates this concern by allowing for dynamic effects
up to a pre-specified treatment lag.
\begin{hyp}\label{hyp:parallel_trends} (Parallel trends)
$\forall$ $d_1\in \mathcal{D}_1$, $E(\Delta Y(d_1)|D_1=d_1,D_2)=E(\Delta
Y(d_1)|D_1=d_1)$.
\end{hyp}
Assumption \ref{hyp:parallel_trends} is a parallel-trends assumption, requiring that $\Delta Y(d_1)$ be mean independent of $D_2$, conditional on $D_1=d_1$.
\begin{hyp}\label{hyp:regularity} (Bounded treatment and regularity of potential outcomes)
\begin{enumerate}
\item $\mathcal{D}_1$ and $\mathcal{D}_2$ are bounded subsets of
$\mathbb{R}$.
\item $\sup_{d_1\in\mathcal{D}_1} E[(\Delta Y(d_1))^2|D_1=d_1]<\infty$.
\item $\forall$ $t\in \{1,2\}$ and $\forall$ $(d,d')\in \mathcal{D}_t^2$,
there is a random variable $\overline{Y}\ge 0$ such that
$|Y_t(d)-Y_t(d')|\leq \overline{Y}|d-d'|$, with $\sup_{(d_1,d_2)\in
\text{Supp}(D_1,D_2)}E[\overline{Y}^2|D_1=d_1, D_2=d_2]<\infty$.
\end{enumerate}
\end{hyp}
Assumption \ref{hyp:regularity} ensures that all the expectations below are well defined. It requires that the set of values that the period-one and period-two treatments can take be bounded. It also requires that the potential outcome functions be Lipschitz (with a unit-specific Lipschitz constant). This holds if $d\mapsto Y_2(d)$ is differentiable with respect to $d$ and has a bounded derivative.
\medskip
Finally, for estimation and inference we assume we observe an iid sample with
the same distribution as $(Y_{1},Y_{2},D_{1},D_{2})$:
\begin{hyp}\label{hyp:iid} (iid sample)
We observe $(Y_{i,1},Y_{i,2},D_{i,1},D_{i,2})_{1\leq i\leq n}$, which are
independent and identically distributed vectors with the same probability
distribution as $(Y_{1},Y_{2},D_{1},D_{2})$.
\end{hyp}
Importantly, Assumption \ref{hyp:iid} allows for the possibility that $Y_{1}$
and $Y_{2}$ (resp. $D_{1}$ and $D_{2}$) are serially correlated, as is
commonly assumed in DID studies \citep[see][]{bertrand2004}.
\@startsection{subsection}{2}{0mm}{-1\baselineskip}{0.7\baselineskip}{\normalfont\normalsize\bfseries}{Building-block identification result}
Assumption \ref{hyp:parallel_trends} implies the following lemma, our
building-block identification result.
\begin{lem}\label{lem:building_block}
If Assumptions \ref{hyp:noanticipation}-\ref{hyp:parallel_trends} hold, then $\forall$ $(d_1,d_2)\in
\mathcal{D}_1\times \mathcal{D}_2$ such that $d_1\ne d_2$ and
$P(S|D_1=d_1)<1$,
\begin{align*}
\text{TE}(d_1,d_2|d_1,d_2):=&E\left(\frac{Y_2(d_2) - Y_2(d_1)}{d_2-d_1}\middle|D_1=d_1,D_2=d_2\right)\\
=&E\left(\frac{\Delta Y - E(\Delta Y | D_1=d_1, S=0)}{d_2-d_1}\middle|D_1=d_1,D_2=d_2\right).
\end{align*}
\end{lem}
\textbf{Proof:} we have
\begin{align*}
&E\left(Y_2(d_2) - Y_2(d_1)\middle|D_1=d_1,D_2=d_2\right)\\
=&E\left(\Delta Y \middle|D_1=d_1,D_2=d_2\right)-E\left(\Delta Y(d_1)\middle|D_1=d_1,D_2=d_2\right)\\
=&E\left(\Delta Y \middle|D_1=d_1,D_2=d_2\right)-E\left(\Delta Y(d_1)\middle|D_1=d_1,D_2=d_1\right)\\
=&E\left(\Delta Y \middle|D_1=d_1,D_2=d_2\right)-E\left(\Delta Y\middle|D_1=d_1,S=0\right)\\
=&E\left(\Delta Y -E\left(\Delta Y\middle|D_1=d_1,S=0\right)\middle|D_1=d_1,D_2=d_2\right),
\end{align*}
where the second equality follows from Assumption \ref{hyp:parallel_trends}.
This proves the result $_\Box$
\medskip
Intuitively, under Assumption \ref{hyp:parallel_trends} the counterfactual
outcome evolution switchers would have experienced if their treatment had not
changed is identified by the outcome evolution of stayers with the same
period-one treatment. If a unit's treatment changes from two to five, we can
recover its counterfactual outcome evolution if its treatment had not
changed, by using the average outcome evolution of all stayers with a
baseline treatment of two. Then, a DID estimand comparing switchers' and
stayers' outcome evolutions identifies $E\left(Y_2(d_2) -
Y_2(d_1)\middle|D_1=d_1,D_2=d_2\right)$, and we can scale that effect by
$d_2-d_1$ to identify a slope rather than an unnormalized effect.
\medskip
In a canonical DID design where $\mathcal{D}_1=\{0\}$ and $\mathcal{D}_2\in \{0,1\}$, Lemma \ref{lem:building_block} only applies to $(d_1,d_2)=(0,1)$, $\text{TE}(0,1|0,1)$ reduces to the average treatment on the treated (ATT),
and the estimand reduces to the canonical DID estimand comparing the outcome evolutions of treated and untreated units. Also, if all units are untreated at period one ($\mathcal{D}_1=\{0\}$), we have
$$\text{TE}(0,d_2|0,d_2)=E\left(\frac{Y_2(d_2) - Y_2(0)}{d_2}\middle|D_2=d_2\right),$$ an effect related to the $\Delta_{d0|D=d}$ effect in \cite{fricke2017identification}, or to the $ATT(d|d)$ effect in \cite{callaway2021difference}. Thus, the effects we consider are extensions of those effects.
\medskip
In the next sections, our target parameters will be two averages of the slopes $\text{TE}(d_1,d_2|d_1,d_2)$, across both $d_1$ and $d_2$, which can be estimated nonparametrically at the $\sqrt{n}-$parametric rate. We do not propose estimators of $(d_1,d_2)\mapsto
\text{TE}(d_1,d_2|d_1,d_2)$, for two reasons. First, variability in $\text{TE}(d_1,d_2|d_1,d_2)$ across values of $(d_1,d_2)$ conflates a dose-response relationship that may be of interest, and a selection bias, typically not of interest, due to the fact that units with different period one and two treatments may have different treatment effects \citep{callaway2021difference}. Moreover, Lemma
\ref{lem:building_block} shows that estimating $\text{TE}(d_1,d_2|d_1,d_2)$ requires estimating two bivariate nonparametric regressions. Unless one is willing to make parametric functional-form assumptions, the resulting estimator converges at most at the $n^{1/3}$ rate \citep{stone1982optimal},
and may therefore be very imprecise, given that DID applications often have a small number of units (e.g. the 50 US states in our empirical application). Similarly, while nonparametric estimators of $d_1\mapsto E\left(\frac{Y_2(D_2)- Y_2(d_1)}{D_2-d_1}\middle|D_1=d_1,D_2\ne d_1\right)$ or $\delta \mapsto E\left(\frac{Y_2(D_2)- Y_2(D_1)}{D_2-D_1}\middle|D_2-D_1=\delta >0\right)$ may converge at a faster rate than $n^{1/3}$, they cannot converge at the $\sqrt{n}$ rate.\footnote{The \texttt{did\_multiplegt\_stat} Stata package can still compute discretized versions of $d_1\mapsto E\left[(Y_2(D_2)- Y_2(d_1))/(D_2-d_1)\middle|D_1=d_1,D_2\ne d_1\right]$ (resp. $\delta \mapsto E[(Y_2(D_2)- Y_2(D_1))/(D_2-D_1)|D_2-D_1=\delta >0]$), when the option \texttt{by\_baseline} (resp. \texttt{by\_fd}) is specified, see the help file for further details.} In a recent paper, posterior to ours, \cite{haddad2024difference} consider a parallel-trends assumption similar to Assumption \ref{hyp:parallel_trends}, and propose an estimator of a functional parameter similar to $(d_1,d_2)\mapsto \text{TE}(d_1,d_2|d_1,d_2)$.
\medskip
Instead of averages of the slopes $\text{TE}(d_1,d_2|d_1,d_2)$,
one may be interested in other parameters, like the average dose-slope
function
$(d,d')\mapsto \text{TE}(d,d'):=E\left((Y_2(d) - Y_2(d'))/(d-d')\right),$
or the average marginal effect $E\left(Y'_2(D_2)\right)$. However,
while averages of $\text{TE}(d_1,d_2|d_1,d_2)$ can be identified under an assumption whose plausibility can be assessed via a pre-trends test, $\text{TE}(d,d')$ and $E\left(Y'_2(D_2)\right)$
cannot be identified under assumptions whose plausibility can be assessed via a pre-trends test, as explained in Section \ref{appendixsec:alternatives} below.
\@startsection{subsection}{2}{0mm}{-1\baselineskip}{0.7\baselineskip}{\normalfont\normalsize\bfseries}{Alternative parallel-trends assumption}\label{sub:alt_PT}
The DID estimands in Lemma \ref{lem:building_block} compare switchers and
stayers with the same period-one treatment. Instead, one could propose
estimands comparing switchers and stayers, without conditioning on their
period-one treatment. To recover the counterfactual outcome trend of a
switcher going from two to five units of treatment, one could use a stayer
with treatment equal to three at both dates. On top of Assumption
\ref{hyp:parallel_trends}, such estimands rest on two supplementary
conditions:
\begin{enumerate}[(i)]
\item $E(\Delta Y(d)|D_1=d)=E(\Delta Y(d))$.
\item $\forall$ $(d,d')\in \mathcal{D}_1^2$, $E(\Delta Y(d))=E(\Delta
Y(d'))$.
\end{enumerate}
(i) requires that all units experience the same evolution of their potential
outcome with treatment $d$, while Assumption \ref{hyp:parallel_trends} only
imposes that requirement for units with the same baseline treatment.
Assumption \ref{hyp:parallel_trends} may be more plausible: units with the
same period-one treatment may be more similar and more likely to be on
parallel trends than units with different period-one treatments. (ii)
requires that the trend affecting all potential outcomes be the same: to
rationalize a DID estimand comparing a switcher going from two to five units
of treatment to a stayer with treatment equal to three, $E(\Delta Y(2))$ and
$E(\Delta Y(3))$ should be equal. Rearranging, (ii) is equivalent to $E(Y_2(d)-Y_2(d'))=E(Y_1(d)-Y_1(d'))$,
so the treatment effect should be constant over time, a strong restriction on
treatment effect heterogeneity. In contrast, our Assumption \ref{hyp:parallel_trends} does not restrict treatment effect heterogeneity.
\@startsection{section}{2}{0mm}{-1\baselineskip}{1\baselineskip}{\normalfont\large\bfseries}{Estimating the average of switchers' slopes}\label{sec:AS}
\@startsection{subsection}{2}{0mm}{-1\baselineskip}{0.7\baselineskip}{\normalfont\normalsize\bfseries}{Target parameter}\label{subsec:ASdef}
In this section, our target parameter is
\begin{align}\label{eq:ampos_definition}
\delta_1& :=E\left(\frac{Y_2(D_2) - Y_2(D_1)}{D_2-D_1} \middle| S=1\right),
\end{align}
the average of the slopes of switchers' potential outcome functions, between
their period-one and their period-two treatments. Hereafter, $\delta_1$ is
referred to as the Average of Slopes (AS).
\medskip
The AS is a local effect: it only applies to switchers, and it measures the
effect of changing their treatment from its period-one to its period-two
value, not of other changes of their treatment. Still, the AS can be used to
point or partially identify the effect of other treatment changes under shape
restrictions. First, assume that the potential outcomes are linear: for $t\in
\{1,2\}$,
$$Y_t(d)=Y_t(0)+B_td,$$
where $B_t$ is a slope that may vary across units and may change over time.
Then, $\delta_1=E\left(B_2\middle| S=1\right)$: the AS is equal to the
average, across switchers, of the slopes of their potential outcome functions
at period 2. Therefore, for all $d\ne d'$,
$$E(Y_2(d)-Y_2(d')|S=1)=(d-d')\delta_1:$$ under linearity, knowing the AS is
sufficient to recover the ATE of any uniform treatment change among
switchers. Of course, this only holds under linearity, which may not be a
plausible assumption. Assume instead that $d\mapsto Y_2(d)$ is convex. Then,
for any $\epsilon>0$,
$$E\left(Y_2(D_2+\epsilon) - Y_2(D_2)\middle| S=1\right)\geq \epsilon \delta_1.$$
Accordingly, under convexity one can use the AS to obtain lower bounds of the
effect of changing the treatment from $D_2$ to larger values than $D_2$. For
instance, in \cite{fajgelbaum2020return}, one can use this strategy to derive
a lower bound of the effect of increasing tariffs' to even higher levels than
those imposed in 2018 by the Trump administration. Under convexity, one can
also use the AS to derive an upper bound of the effect of changing the
treatment from $D_1$ to a lower value. And under concavity, one can derive an
upper (resp. lower) bound of the effect of changing the treatment from $D_2$
(resp. $D_1$) to a larger (resp. lower) value.\footnote{See
\cite{d2021nonparametric} for bounds of the same kind obtained under
concavity or convexity.} The AS is identified even if those linearity or
convexity/concavity conditions fail, but those
conditions are necessary to use the AS to identify or bound the effects of
alternative policies.
\@startsection{subsection}{2}{0mm}{-1\baselineskip}{0.7\baselineskip}{\normalfont\normalsize\bfseries}{Identification}\label{subsec:ASidentification}
The AS is identified under the two following assumptions on the design.
\begin{hyp}\label{hyp:support_condition0} (Overlap condition for AS identification)
$P(S=1)>0$ and $P(S=0|D_1)>\zeta$ for some $\zeta>0$.
\end{hyp}
\begin{hyp}\label{hyp:noquasi-stayers} (No quasi-stayers)
$P(|\Delta D|>\kappa|S=1)=1$ for some $\kappa>0$.
\end{hyp}
The estimand identifying the AS compares the outcome evolutions of switchers
and stayers with the same period-one treatment. Then, this estimand requires
that there be no value of $D_1$ such that only switchers have that value, as
assumed in Assumption \ref{hyp:support_condition0}. Assumption
\ref{hyp:noquasi-stayers} requires that there are no quasi-stayers: the
treatment of all switchers changes by at least $\kappa>0$.
\medskip
Define the functions:
\[\mu_0(d):=E[\Delta Y|S=0,D_1=d]\quad \text{ and } \alpha_0(s,d):=E\left(\left.\frac{S}{\Delta D}\right\vert D_1=d\right)\frac{1-s}{E(1-S|D_1=d)}.\]
\begin{thm}\label{thm:main0}
If Assumptions \ref{hyp:noanticipation}-\ref{hyp:regularity},
\ref{hyp:support_condition0}, and \ref{hyp:noquasi-stayers} hold, then for
any functions $(s,d)\mapsto\alpha(s,d)$ and $d\mapsto\mu(d)$, respectively
defined on $\{0,1\}\times \mathcal{D}_1$ and $\mathcal{D}_1$, such that $\alpha$ is bounded and $E[|\mu(D_1)|]<\infty$:
\begin{align*}
\delta_1=& E\left(\frac{1}{E[S]}\left(\frac{S}{\Delta D}-\alpha(S,D_1)\right)\left[\Delta Y-\mu_0(D_1)\right]\right)\\
=& E\left(\frac{1}{E[S]}\left(\frac{S}{\Delta D}-\alpha_0(S,D_1)\right)\left[\Delta Y-\mu(D_1)\right]\right).
\end{align*}
\end{thm}
It readily follows from Lemma \ref{lem:building_block} and the law of
iterated expectations that $\delta_1$ is identified by
$$E\left[\left(\frac{\Delta Y-E[\Delta Y|S=0,D_1]}{\Delta D}\right)\middle|S=1\right]=E\left[\left(\frac{\Delta Y-\mu_0(D_1)}{\Delta D}\right)\middle|S=1\right],$$
where the numerator of the ratio inside the expectation is a DID comparing
the outcome evolutions of switchers and stayers with the same $D_1$. Theorem
\ref{thm:main0} implies that a related estimand,
$$E\left(\frac{1}{E[S]}\left(\frac{S}{\Delta D}-\alpha_0(S,D_1)\right)\left[\Delta Y-\mu_0(D_1)\right]\right),$$
also identifies $\delta_1$, and is doubly-robust in the functions
$\mu_0(D_1)$ and $\alpha_0(S,D_1)$.\footnote{We are grateful to Andres Santos
for pointing out to us that the AS is identified by a doubly-robust
estimand.}
\medskip
If there are quasi-stayers, the AS is still identified, but estimation may be
challenging. For any $\eta>0$, let $S_\eta=1\{|\Delta D|>\eta\}$ be an
indicator for switchers whose treatment changes by at least $\eta$ from
period one to two.
\begin{thm}\label{thm:main0_quasistayers}
If Assumptions \ref{hyp:noanticipation}-\ref{hyp:regularity} and
\ref{hyp:support_condition0} hold,
\begin{align*}
\delta_1& = \lim_{\eta\downarrow 0} E\left(\frac{\Delta Y - E(\Delta Y | D_1, S=0)}{\Delta D} \middle| S_\eta=1\right).
\end{align*}
\end{thm}
If there are quasi-stayers whose treatment change is arbitrarily close to $0$
(i.e. $f_{|\Delta D||S=1}(0)>0$), the denominator of $(\Delta Y - E(\Delta Y
| D_1, S=0))/\Delta D$ is close to $0$ for them. On the other hand,
\begin{align*}
\Delta Y - E(\Delta Y | D_1, S=0)&=Y_2(D_2)-Y_2(D_1)+\Delta Y(D_1)- E(\Delta Y(D_1)| D_1, S=0)\\
&\approx\Delta Y(D_1)- E(\Delta Y(D_1)| D_1, S=0),
\end{align*}
so the ratio's numerator may not be close to $0$. Then, under weak
conditions,
$$E\left(\left|\frac{\Delta Y - E(\Delta Y | D_1, S=0)}{\Delta D}\right|\; \middle| S=1\right)=+\infty.$$
Therefore, we need to trim quasi-stayers from the estimand in Theorem
\ref{thm:main0}, and let the trimming go to 0, as in
\cite{graham2012identification} who consider a related estimand with some
quasi-stayers. Accordingly, with quasi-stayers the AS is irregularly
identified by a limiting estimand. Then, we conjecture that the AS cannot be
estimated at the $\sqrt{n}-$rate, as in \cite{graham2012identification}, and
as is often the case with target parameters identified by limiting estimands.
Therefore, we do not consider estimation of the AS with quasi stayers.
\@startsection{subsection}{2}{0mm}{-1\baselineskip}{0.7\baselineskip}{\normalfont\normalsize\bfseries}{Estimation and inference}\label{subsec:ASestimation}
Let
\[g_0(d):=E[S/\Delta D|D_1=d]\quad \text{and } p_0(d):=E[1-S|D_1=d].\]
It follows from Theorem \ref{thm:main0} that estimating $\delta_1$ requires
estimating the nuisance functions $g_0(d)$, $p_0(d)$, and $\mu_0(d)$ in a first
step, before averaging the variable
$$\left(\frac{S_i}{\Delta D_i}-\frac{\hat g(D_{i,1})}{\hat p(D_{i,1})}(1-S_i)\right)\left(\Delta Y_i-\hat\mu(D_{i,1})\right).$$
We propose to use cross-fitting to perform this two-step
estimation. For any set $A$, let $\# A$ denote its cardinality. Let the sample be split randomly into two sub-samples,
$\mathcal{I}_1$ and $\mathcal{I}_2$, and let $I_k=\# \mathcal{I}_k$ for
$k=1,2$. We let $\hat g^{(2)}$, $\hat p^{(2)}$ and $\hat \mu^{(2)}$ be
nonparametric estimators of $g_0$, $p_0$ and $\mu_0$ based on $\mathcal{I}_2$.
Then the estimator $\hat\delta_{1,\mathsf{DR}}^{(1)}$ is computed in
$\mathcal{I}_1$ as follows:
\[\hat\delta_{1,\mathsf{DR}}^{(1)}=\frac{1}{n_S^{(1)}}\sum_{i\in \mathcal{I}_1}\left(\frac{S_i}{\Delta D_i}-\frac{\hat g^{(2)}(D_{i,1})}{\hat p^{(2)}(D_{i,1})}(1-S_i)\right)\left(\Delta Y_i-\hat\mu^{(2)}(D_{i,1})\right),\]
where $n_S^{(1)}:=\sum_{i\in \mathcal{I}_1} S_i$ and we use the convention
that $\hat\delta_{1,\mathsf{DR}}^{(1)}=0$ if $n_S^{(1)}=0$. The estimator
$\hat\delta_{1,\mathsf{DR}}^{(2)}$ is computed similarly, permuting the roles
of $\mathcal{I}_1$ and $\mathcal{I}_2$. Finally the cross-fitting
doubly-robust estimator is computed as:
\[\hat{\delta}_{1,\mathsf{DR}}=\frac{n_S^{(1)}}{n_S^{(1)}+n_S^{(2)}}\hat\delta_{1,\mathsf{DR}}^{(1)}+
\frac{n_S^{(2)}}{n_S^{(1)}+n_S^{(2)}}\hat\delta_{1,\mathsf{DR}}^{(2)}.\]
We impose the following condition:
\begin{hyp}\label{hyp:subsamples}
$\mathcal{I}_1 \perp \!\!\! \perp (D_{i,2},Y_{i,1},Y_{i,2})_{i=1,...,n}|(D_{i,1})_{i=1,...,n}$ and
$I_1/n\to \pi\in (0,1)$.
\end{hyp}
We also require that the estimators of the nonparametric nuisance functions
satisfy some convergence-rate requirements. For any function $h(\cdot)$, $\hat h^{(1)}(.), \hat h^{(2)}(.)$ estimators of $h(.)$ based on subsample $\mathcal{I}_1$ and $\mathcal{I}_2$ respectively and $k\in\{1,2\}$, define:
\[\left\Vert\hat{h}-h\right\Vert_{2,\mathcal{I}_k}:=\left(\frac{1}{I_k}\sum_{i\in\mathcal{I}_k}(\hat h^{(3-k)}(D_{1,i})-h(D_{1,i}))^2\right)^{1/2},\quad \left\Vert\hat h-h\right\Vert_{k,\infty}:=\sup_{d\in \mathcal{D}_1}\left\vert\hat h^{(3-k)}(d)-h(d)\right\vert.\]
\begin{hyp}\label{hyp:rate_DR_AS}
For $k=1,2$,
\begin{enumerate}[(a)]
\item $\left\Vert\hat g-g_0\right\Vert_{2,\mathcal{I}_k}=o_P(1)$, $\left\Vert\hat
\mu-\mu_0\right\Vert_{2,\mathcal{I}_k}=o_P(1)$, $\left\Vert\hat p-p_0\right\Vert_{k,\infty}=o_P(1)$.
\item $\sqrt{n}\left\Vert\hat g-g_0\right\Vert_{2,\mathcal{I}_k}\times \left\Vert\hat
\mu-\mu_0\right\Vert_{2,\mathcal{I}_k}=o_P(1)$, $\sqrt{n}\left\Vert\hat
p-p_0\right\Vert_{2,\mathcal{I}_k}\times \left\Vert\hat
\mu-\mu_0\right\Vert_{2,\mathcal{I}_k}=o_P(1)$.
\end{enumerate}
\end{hyp}
As all nuisance functions are conditional expectations with respect to a
single variable, these convergence-rate requirements are satisfied by
classical nonparametric estimators. For instance, one can use series
estimators with a polynomial order chosen by cross validation within
subsample $\mathcal{I}_k$ to estimate $\hat\mu^{(k)}(d)$, $\hat{g}^{(k)}(d)$,
and $\hat p^{(k)}(d)$. Results in Section 4 of \cite{andrews1991asymptotic}
imply that those estimators satisfy the rate conditions in Assumption
\ref{hyp:rate_DR_AS}, as long as the population functions are twice
differentiable. Note also that the mean-square convergence requirements can
be established under regularity conditions and appropriate choices of tuning
parameters for machine learning estimators such as random forests
\citep{Biau_2012_JMLR,Wager-Athey_2018_JASA,Syrgkanis-Zampetakis_2020} and
neural networks
\citep{Chen_White_1999_IEEE,Farrell-etal_2021_ECMA}.\footnote{We rely on
uniform convergence of the propensity score to ensure that the inverse
propensity score is consistent. While uniform convergence may be too strong
for machine learning estimators, this condition may be relaxed using
automatic debiased machine learning, as proposed in
\cite{Cherno-etal_2024_wp}.}
\begin{thm}\label{thm:ASasnormality} If Assumptions \ref{hyp:noanticipation}-\ref{hyp:rate_DR_AS} hold, $\sqrt{n}(\hat\delta_{1,\mathsf{DR}}-\delta_1)\stackrel{d}{\longrightarrow} \mathcal{N}(0,V(\psi_1))$, where
\[\psi_1:=\frac{1}{E[S]}\left\{\left(\frac{S}{\Delta D}-E\left(\left.\frac{S}{\Delta D}\right\vert D_1\right)\frac{1-S}{E(1-S|D_1)}\right)\left[\Delta Y-E(\Delta Y|S=0,D_1)\right]-\delta_1S\right\}.\]
\end{thm}
\@startsection{section}{2}{0mm}{-1\baselineskip}{1\baselineskip}{\normalfont\large\bfseries}{Estimating a weighted average of switchers' slopes}\label{sec:WAS}
\@startsection{subsection}{2}{0mm}{-1\baselineskip}{0.7\baselineskip}{\normalfont\normalsize\bfseries}{Target parameter}\label{subsec:WASdef}
In this section, our target parameter is
\begin{align*}
\delta_2 :=&E\left(\frac{|D_2-D_1|}{E(|D_2-D_1||S=1)} \times \frac{Y_2(D_2)-Y_2(D_1)}{D_2-D_1}\middle|S=1\right),
\end{align*}
a weighted average of the slopes of switchers' potential
outcome functions from their period-one to their period-two treatments, where
slopes receive a weight proportional to switchers' absolute treatment change
from period one to two. We refer to $\delta_{2}$ as the Weighted
Average of Slopes (WAS). $\delta_2=\delta_1$ if and only if switchers' slopes are uncorrelated
with $|D_2-D_1|$.
\medskip
The AS and WAS serve different purposes. Unlike the AS, the WAS may not be used to identify or bound the effect of other treatment changes than the actual changes that took place, even under shape
restrictions on the potential outcome function. Instead, the WAS may be used to conduct a cost-benefit analysis of the treatment changes that took place from period
one to two. To simplify the discussion, let us assume in the remainder of
this paragraph that $D_2\ge D_1$. Assume also that the outcome is a measure
of output, such as agricultural yields or wages, expressed in monetary units.
Finally, assume that the treatment is costly, with a cost linear in dose,
uniform across units, and known to the analyst: the cost of giving $d$ units
of treatment to a unit at period $t$ is $c_t\times d$ for some known
$(c_t)_{t\in \{1,2\}}$. Then, $D_2$ is beneficial relative to $D_1$ if and
only if $E(Y_2(D_2)-c_2D_2)>E(Y_2(D_1)-c_2D_1)$, namely if and only if
$\delta_2>c_2:$ comparing $\delta_2$ to $c_2$ is sufficient to evaluate if
changing the treatment from $D_1$ to $D_2$ was beneficial.
\@startsection{subsection}{2}{0mm}{-1\baselineskip}{0.7\baselineskip}{\normalfont\normalsize\bfseries}{Identification}\label{subsec:WASidentification}
Let $S_+=1\{D_2-D_1>0\}$ and $S_-=1\{D_2-D_1<0\}$ respectively be indicators
for ``switchers up'' and ``switchers down''
and define the function:
\[\gamma_0(s,d):=\frac{E(S_+|D_1=d)-E(S_-|D_1=d)}{E(1-S|D_1=d)}(1-s).\]
\begin{thm}\label{thm:main}
If Assumptions \ref{hyp:noanticipation}-\ref{hyp:regularity}, and
\ref{hyp:support_condition0} hold, then for any functions $(s,d)\mapsto\gamma(s,d)$ and $d\mapsto\mu(d)$, respectively defined on $\{0,1\}\times \mathcal{D}_1$ and $\mathcal{D}_1$ and such that $\gamma$ is bounded and $E[|\mu(D_1)|]<\infty$:
\begin{align*}
\delta_2=& \frac{E\left[\left(S_+-S_--\gamma(S,D_1)\right)(\Delta Y-\mu_0(D_1))\right]}{E[\left\vert\Delta D\right\vert]}\\
=& \frac{E\left[\left(S_+-S_--\gamma_0(S,D_1)\right)(\Delta Y-\mu(D_1))\right]}{E[\left\vert\Delta D\right\vert]}.
\end{align*}
\end{thm}
Here are the steps leading to Theorem \ref{thm:main}. One has
\begin{align*}
\delta_2 =&E\left(\frac{|D_2-D_1|}{E(|D_2-D_1||S=1)} \times \frac{Y_2(D_2)-Y_2(D_1)}{D_2-D_1}\middle|S=1\right) \\
=& \frac{E\left(\text{sgn}(D_2-D_1) (Y_2(D_2)-Y_2(D_1))\middle|S=1\right)}{E(|D_2-D_1||S=1)}
\end{align*}
\begin{align*}
\qquad =& \frac{E\left((S_+-S_-) (Y_2(D_2)-Y_2(D_1))\right)}{E(|\Delta D|)}\\
=& \frac{E\left((S_+-S_-) (\Delta Y-\Delta Y(D_1))\right)}{E(|\Delta D|)}\\
=& \frac{E\left((S_+-S_-) (E(\Delta Y|D_1,D_2)-E(\Delta Y(D_1)|D_1,D_2))\right)}{E(|\Delta D|)}\\
=& \frac{E\left((S_+-S_-) (E(\Delta Y|D_1,D_2)-E(\Delta Y|D_1,S=0))\right)}{E(|\Delta D|)}\\
=& \frac{E\left[\left(S_+-S_-\right)(\Delta Y-E(\Delta Y|D_1,S=0))\right]}{E[\left\vert\Delta D\right\vert]}\\
=&\frac{E\left[\left(S_+-S_-\right)(\Delta Y-\mu_0(D_1))\right]}{E[\left\vert\Delta D\right\vert]}.
\end{align*}
The fourth equality follows from adding and subtracting $Y_1(D_1)$. The fifth
equality follows from the law of iterated expectations. The sixth follows
from Assumption \ref{hyp:parallel_trends} and the fact that $\Delta Y=\Delta
Y(D_1)$ for stayers. The seventh follows from the law of iterated
expectations, and the last follows from the definition of $\mu_0$. The
numerator of the estimand identifying $\delta_2$ in the previous display is a
conditional DID comparing the outcome evolutions of switchers and stayers
with the same $D_1$. Following similar steps, one can show that $\delta_2$ is
also identified by
$$\frac{E\left[\left(S_+-S_-\right)\Delta Y\right]-E\left[\frac{E(S_+|D_1=d)-E(S_-|D_1=d)}{E(1-S|D_1=d)}(1-S)\Delta Y\right]}{E[\left\vert\Delta D\right\vert]}=\frac{E\left[\left(S_+-S_--\gamma_0(S,D_1)\right)\Delta Y\right]}{E[\left\vert\Delta D\right\vert]},$$
an estimand whose numerator is an unconditional DID with propensity score
reweighting. Theorem \ref{thm:main} implies that an estimand combining those
two estimands,
$$\frac{E\left[\left(S_+-S_--\gamma_0(S,D_1)\right)(\Delta Y-\mu_0(D_1))\right]}{E[\left\vert\Delta D\right\vert]},$$
also identifies $\delta_2$, and is doubly-robust in the functions
$\mu_0(D_1)$ and $\gamma_0(S,D_1)$.
\@startsection{subsection}{2}{0mm}{-1\baselineskip}{0.7\baselineskip}{\normalfont\normalsize\bfseries}{Estimation and inference}\label{subsec:WASestimation}
Let $h_0(d):=E[S_+-S_-|D_1=d]$. As before, split the sample into
subsamples $\mathcal{I}_1$ and $\mathcal{I}_2$, and let
\[\hat\delta_{2,\mathsf{DR}}^{(1)}:=\frac{1}{\Delta_S^{(1)}}\sum_{i\in \mathcal{I}_1}\left(S_{i+}-S_{i-}-\frac{\hat h^{(2)}(D_{i,1})}{\hat p^{(2)}(D_{i,1})}(1-S_i)\right) (\Delta Y_i -\hat\mu^{(2)}(D_{i,1})),\]
where $\Delta_S^{(1)}:=\sum_{i\in \mathcal{I}_1} \left\vert\Delta D_i\right\vert$ and we use
the convention that $\hat\delta_{2,\mathsf{DR}}^{(1)}=0$ if
$\Delta_S^{(1)}=0$. The estimator $\hat\delta_{2,\mathsf{DR}}^{(2)}$ is
constructed analogously, permuting the roles of $\mathcal{I}_1$ and
$\mathcal{I}_2$. Finally, let
\[\hat\delta_{2,\mathsf{DR}}:=\frac{\Delta_S^{(1)}}{\Delta_S^{(1)}+\Delta_S^{(2)}}\hat\delta_{2,\mathsf{DR}}^{(1)}+\frac{\Delta_S^{(2)}}{\Delta_S^{(1)}+\Delta_S^{(2)}}\hat\delta_{2,\mathsf{DR}}^{(2)}.\]
We impose the following rate conditions.
\begin{hyp}\label{hyp:rate_DR_WAS}
For $k=1,2$,
\begin{enumerate}[(a)]
\item $\left\Vert\hat h-h_0\right\Vert_{2,\mathcal{I}_k}=o_P(1)$, $\left\Vert\hat
\mu-\mu_0\right\Vert_{2,\mathcal{I}_k}=o_P(1)$, $\left\Vert\hat p-p_0\right\Vert_{k,\infty}=o_P(1)$.
\item $\sqrt{n}\left\Vert\hat h-h_0\right\Vert_{2,\mathcal{I}_k}\times \left\Vert\hat
\mu-\mu_0\right\Vert_{2,\mathcal{I}_k}=o_P(1)$, $\sqrt{n}\left\Vert\hat
p-p_0\right\Vert_{2,\mathcal{I}_k}\times \left\Vert\hat
\mu-\mu_0\right\Vert_{2,\mathcal{I}_k}=o_P(1)$.
\end{enumerate}
\end{hyp}
\begin{thm}\label{thm:WASasnormality_nonmonotonictreatment}
If Assumptions \ref{hyp:noanticipation}-\ref{hyp:support_condition0},
\ref{hyp:subsamples}, and \ref{hyp:rate_DR_WAS} hold, $\sqrt{n}\left(\widehat{\delta}_{2,\mathsf{DR}} - \delta_{2}\right) \stackrel{d}{\longrightarrow} \mathcal{N}(0, V(\psi_{2}))$,
where
\begin{align*}
\psi_{2}&:=\frac{1}{E(\left\vert\Delta D\right\vert)}\left\{\left(S_+-S_--E(S_+-S_-|D_1)\frac{(1-S)}{E(1-S|D_1)}\right)\right.
(\Delta Y -E(\Delta Y | D_{1}, S=0))-\delta_{2}\left\vert\Delta D\right\vert\bigg\}.
\end{align*}
\end{thm}
\medskip
Finally, we now show that under some assumptions, the asymptotic variance of
the WAS estimator is lower than that of the AS estimator.
\begin{prop}\label{prop:variance_comparison}
If Assumptions \ref{hyp:noanticipation}-\ref{hyp:parallel_trends} hold,
$(Y_2(D_2)-Y_2(D_1))/(D_2-D_1)=\delta$, $V(\Delta Y(D_1)|D_1,D_2)=\sigma^2$
for some real number $\sigma^2>0$, $D_2\ge D_1$, and $\Delta D \perp \!\!\! \perp D_1$,
then
\begin{align*}
V(\psi_{1})=&\sigma^2\left[\frac{E(1/(\Delta D)^2|S=1)}{P(S=1)}+\frac{\left(E(1/\Delta D|S=1)\right)^2}{P(S=0)}\right]\\
\ge & \sigma^2\frac{1}{\left(E(\Delta D|S=1)\right)^2}\left[\frac{1}{P(S=1)}+\frac{1}{P(S=0)}\right]=V(\psi_{2}),
\end{align*}
with equality if and only if $V(\Delta D|S=1)=0$.
\end{prop}
The constant effect and homoscedasticity assumptions underlying Proposition
\ref{prop:variance_comparison} are admittedly strong, but ranking estimators' variances typically requires such conditions. The question then is whether the ranking in Proposition
\ref{prop:variance_comparison} still holds in real-life applications, where
those assumptions likely fail. In our application,
the variance of $\widehat{\delta}_{1,\mathsf{DR}}$ is indeed
much larger than that of $\widehat{\delta}_{2,\mathsf{DR}}$.
\@startsection{section}{2}{0mm}{-1\baselineskip}{1\baselineskip}{\normalfont\large\bfseries}{Extensions}
\label{sec:extensions}
\@startsection{subsection}{2}{0mm}{-1\baselineskip}{0.7\baselineskip}{\normalfont\normalsize\bfseries}{Instrumental-variable estimators}
\label{subsec:IV}
In this section, we consider instances where Assumption \ref{hyp:parallel_trends} is implausible, but one has an instrument that satisfies a similar parallel-trends condition.
\@startsection{subsubsection}{3}{0mm}{-0.8\baselineskip}{0.4\baselineskip}{\normalfont\normalsize\itshape}{Notation and assumptions}
Let $Z_t$ denote the instrumental variable at period $t\in\{1,2\}$. Throughout this section, we rule out anticipatory and dynamic effects of the instrument on the treatment. Then, let
$D_t(z)$ denote the unit's potential treatment at $t$ if $Z_t=z$. Let also $SC=1\{D_2(Z_2)\ne D_2(Z_1)\}$ be an indicator equal to 1 for switchers-compliers, namely units whose instrument changes from period one to two and whose treatment is affected by that change. Finally, let $\mathcal{Z}_t$ denote the support of $Z_t$.
\medskip
We replace Assumption \ref{hyp:parallel_trends} by the following assumption.\footnote{Our notation, where potential outcomes do not depend on $z$, implicitly imposes the usual exclusion restriction.}
\begin{hyp}\label{hyp:parallel_trends_IV} (Reduced-form and first-stage parallel trends)
For all $z\in \mathcal{Z}_1$,
\begin{enumerate}
\item $E(Y_2(D_2(z))-Y_1(D_1(z))|Z_1=z,Z_2,D_1)=E(Y_2(D_2(z))-Y_1(D_1(z))|Z_1=z,D_1)$.
\item $E(D_2(z)-D_1(z)|Z_1=z,Z_2,D_1)=E(D_2(z)-D_1(z)|Z_1=z,D_1)$.
\end{enumerate}
\end{hyp}
Point 1 of Assumption \ref{hyp:parallel_trends_IV} requires that $Y_2(D_2(z))-Y_1(D_1(z))$, units' outcome evolutions in the counterfactual where their instrument does not change from period one to two, be mean independent of $Z_2$, conditional on $Z_1$ and $D_1$. This condition imposes some restrictions on treatment effect heterogeneity, and the goal of conditioning on $D_1$ is to alleviate those restrictions. To see this, assume that
\begin{align}
&E(Y_2(D_1(z))-Y_1(D_1(z))|Z_1=z,Z_2,D_1)=E(Y_2(D_1(z))-Y_1(D_1(z))|Z_1=z,D_1).\label{eq:SC_PT_IV1}
\end{align}
\eqref{eq:SC_PT_IV1} requires that $Y_2(D_1(z))-Y_1(D_1(z))$, units' outcome evolutions in the counterfactual where their instrument and their treatment does not change from period one to two, be mean independent of $Z_2$, conditional on $Z_1$ and $D_1$. Thanks to the conditioning on $D_1$, \eqref{eq:SC_PT_IV1} is a standard parallel-trends assumption that does not impose any restriction on treatment effect heterogeneity, like Assumption \ref{hyp:parallel_trends}. If $D_1$ was not conditioned upon, \eqref{eq:SC_PT_IV1} would require parallel trends among units with different baseline treatments, which implicitly assumes homogeneous treatment effects over time, as discussed in Section \ref{sub:alt_PT}. Under \eqref{eq:SC_PT_IV1}, Point 1 of Assumption \ref{hyp:parallel_trends_IV} holds if and only if
\begin{align}
&E(Y_2(D_2(z))-Y_2(D_1(z))|Z_1=z,Z_2,D_1)=E(Y_2(D_2(z))-Y_2(D_1(z))|Z_1=z,D_1).\label{eq:SC_PT_IV2}
\end{align}
\eqref{eq:SC_PT_IV2} is a restriction on treatment effect heterogeneity across units. It requires that switching the treatment from $D_1(Z_1)$ to $D_2(Z_1)$, the natural change happening over time without a change in the instrument, has an effect that is mean independent of $Z_2$ conditional on $Z_1$ and $D_1$. Thus, it is only if \eqref{eq:SC_PT_IV1} fails that Point 1 of Assumption \ref{hyp:parallel_trends_IV} does not restrict treatment-effect heterogeneity. However, it may not be plausible to assume that Assumption \ref{hyp:parallel_trends_IV} holds but \eqref{eq:SC_PT_IV1} fails.
\medskip
\cite{de2010note} and \cite{hudson2017interpreting} also consider IV-DID estimands, in classical designs with two periods and a binary instrument that turns on for some units at period two. Both papers introduce a ``reduced-form'' parallel-trends assumption similar to Point 1 of Assumption \ref{hyp:parallel_trends_IV}, but without noting that it imposes restrictions on effects' heterogeneity, even in the simple designs considered by those papers.
Turning to Point 2 of Assumption \ref{hyp:parallel_trends_IV}, this condition requires that units' treatment evolutions under $Z_1$ be mean independent of $Z_2$, conditional on $Z_1$ and $D_1$. Because $D_1$ is conditioned upon, this parallel trends condition is equivalent to a sequential exogeneity assumption \citep[see][]{robins1986new,bojinov2020panel}.
\medskip
We also make the following
assumptions.
\begin{hyp}\label{hyp:monotonicity} (Monotonicity and strictly positive first-stage)
i) For all $(z,z')\in \mathcal{Z}_2^2$, $z\ge z' \Rightarrow D_2(z)\ge D_2(z')$, and ii) $E(|D_2(Z_2)-D_2(Z_1)|)>0$.
\end{hyp}
Point i) of Assumption \ref{hyp:monotonicity} is a monotonicity assumption similar to that in \cite{Imbens94}. It requires that increasing the period-two instrument weakly increases the period-two treatment. This condition is plausible when the instrument is taxes and the treatment is prices, as in our application. Point ii) requires that the instrument has a strictly positive first stage.
\begin{hyp}\label{hyp:regularity_IV} (Bounded instrument and regularity of potential outcomes and treatments)
\begin{enumerate}
\item $\mathcal{Z}_1$ and $\mathcal{Z}_2$ are bounded subsets of $\mathbb{R}$.
\item $\sup_{z_1\in\mathcal{Z}_1} E[(Y_2(D_2(z_1))-Y_1(D_1(z_1)))^2|Z_1=z_1]<\infty$, $\sup_{z_1\in\mathcal{Z}_1} E[(\Delta D(z_1))^2|Z_1=z_1]<\infty$.
\item For all $t\in \{1,2\}$ and for all $(z,z')\in \mathcal{Z}_t^2$, there is a random variable $\overline{Y}\ge 0$ such that $|Y_t(D_t(z))-Y_t(D_t(z'))|\leq \overline{Y}|z-z'|$, with $\sup_{(z_1,z_2)\in \text{Supp}(Z_1,Z_2)}E[\overline{Y}^2|Z_1=z_1, Z_2=z_2]<\infty$.
\item For all $t\in \{1,2\}$ and for all $(z,z')\in \mathcal{Z}_t^2$, there is a random variable $\overline{D}\ge 0$ such that $|D_t(z)-D_t(z')|\leq \overline{D}|z-z'|$, with $\sup_{(z_1,z_2)\in \text{Supp}(Z_1,Z_2)}E[\overline{D}^2|Z_1=z_1, Z_2=z_2]<\infty$.
\end{enumerate}
\end{hyp}
\begin{hyp}\label{hyp:iid_IV} (iid sample)
We observe $(Y_{i,1},Y_{i,2},D_{i,1},D_{i,2},Z_{i,1},Z_{i,2})_{1\leq i\leq n}$, that are independent and identically distributed with the same probability distribution as $(Y_{1},Y_{2},D_{1},D_{2},Z_1,Z_2)$.
\end{hyp}
Assumptions \ref{hyp:regularity_IV} and \ref{hyp:iid_IV} are adaptations of Assumptions \ref{hyp:regularity} and \ref{hyp:iid} to the IV setting.
\@startsection{subsubsection}{3}{0mm}{-0.8\baselineskip}{0.4\baselineskip}{\normalfont\normalsize\itshape}{Target parameter}
In this section, our target parameter is
\begin{align*}
\delta_{\mathsf{IV}} :=&E\left(\frac{|D_2(Z_2)-D_2(Z_1)|}{E(|D_2(Z_2)-D_2(Z_1)||SC=1)}\times \frac{Y_2(D_2(Z_2))-Y_2(D_2(Z_1))}{D_2(Z_2)-D_2(Z_1)}\middle|SC=1\right).
\end{align*}
$\delta_{\mathsf{IV}}$ is a weighted average of the slopes of compliers-switchers' period-two potential outcome functions, from their period-two treatment under their period-one instrument, to their period-two treatment under their period-two instrument. Slopes receive a weight proportional to the absolute value of compliers-switchers' treatment response to the instrument change. Under Assumption \ref{hyp:monotonicity}, $\delta_{\mathsf{IV}}$ is equal to the reduced-form WAS effect of the instrument on the outcome
$$E\left(\frac{|Z_2-Z_1|}{E(|Z_2-Z_1||Z_2\ne Z_1)}\times \frac{Y_2(D_2(Z_2))-Y_2(D_2(Z_1))}{Z_2-Z_1}\middle|Z_2\ne Z_1\right),$$
divided by the first-stage WAS effect of the instrument on the treatment
$$E\left(\frac{|Z_2-Z_1|}{E(|Z_2-Z_1||Z_2\ne Z_1)}\times \frac{D_2(Z_2)-D_2(Z_1)}{Z_2-Z_1}\middle|Z_2\ne Z_1\right).$$
With a binary instrument, such that $Z_1=0$ and $Z_2 \in \{0,1\}$, our IV-WAS effect coincides with that identified in Corollary 2 of \cite{angrist2000}, in a cross-sectional IV model.
\medskip
We could also consider a reduced-form AS divided by a first-stage AS. The resulting target is a weighted average of
the slopes $(Y_2(D_2(Z_2))-Y_2(D_2(Z_1)))/(D_2(Z_2)-D_2(Z_1))$, with weights proportional to $(D_2(Z_2)-D_2(Z_1))/
(Z_2-Z_1)$. It seems more natural to us to weight compliers-switchers' slopes by the absolute value of their first-stage than by the slope of their first-stage.\footnote{If the first-stage effect is homogenous and linear, the weights in the IV-AS effect reduce to one, and one recovers a standard AS effect. However, linearity and homogeneity of the first-stage effect are strong assumptions.}
\@startsection{subsubsection}{3}{0mm}{-0.8\baselineskip}{0.4\baselineskip}{\normalfont\normalsize\itshape}{Identification}
Let $S^I=1\{Z_2-Z_1\ne 0\}$, $S^I_{+}=1\{Z_2-Z_1>0\}$,
$S^I_{-}=1\{Z_2-Z_1<0\}$ and
\begin{align*}
&\mu^U_0(z,d):=E[\Delta U|S^I=0,Z_1=z,D_1=d],\quad U\in\{D,Y\},\\
&\gamma^I_0(s,z,d):=\frac{E(S^I_+|Z_1=z,D_1=d)-E(S^I_-|Z_1=z,D_1=d)}{E(1-S^I|Z_1=z,D_1=d)}(1-s).
\end{align*}
The following condition is similar to Assumption \ref{hyp:support_condition0}.
\begin{hyp}\label{hyp:support_condition_IV} (Overlap conditions for IV-WAS identification)
$P(S^I=1)>0$ and
$P(S^I=0|Z_1,D_1)>\zeta$ for some $\zeta>0$.
\end{hyp}
\begin{thm}\label{thm:main_IV}
If Assumptions \ref{hyp:noanticipation}, \ref{hyp:nodynamic}, \ref{hyp:parallel_trends_IV}-\ref{hyp:regularity_IV} and \ref{hyp:support_condition_IV} hold, then
\begin{align*}
\delta_{\mathsf{IV}}=& \frac{E\left[\left(S^I_+-S^I_--\gamma^I_0(S^I,Z_1,D_1)\right)(\Delta Y-\mu^{Y}_0(Z_1,D_1))\right]}{E\left[\left(S^I_+-S^I_--\gamma^I_0(S^I,Z_1,D_1)\right)(\Delta D-\mu^{D}_0(Z_1,D_1))\right]}.
\end{align*}
\end{thm}
The estimand identifying $\delta_{\mathsf{IV}}$ is equal to the estimand identifying the reduced-form WAS effect of the instrument on the outcome controlling for $D_1$, divided by the estimand identifying the first-stage WAS effect of the instrument on the treatment controlling for $D_1$.
One can show that the estimand identifying $\delta_{\mathsf{IV}}$ is doubly robust in $\gamma^I_0(S^I,Z_1,D_1)$ and $(\mu^{Y}_0(Z_1,D_1),\mu^{D}_0(Z_1,D_1))$.
\@startsection{subsubsection}{3}{0mm}{-0.8\baselineskip}{0.4\baselineskip}{\normalfont\normalsize\itshape}{Estimation and inference}
Estimating $\delta_{\mathsf{IV}}$ requires estimating the nuisance functions $\mu^{Y}_0(z,d)$, $\mu^{D}_0(z,d)$, $E(S^I_+|Z_1=z,D_1=d)$, $E(S^I_-|Z_1=z,D_1=d)$, and $E(1-S^I|Z_1=z,D_1=d)$, before averaging the variables $(S^I_{i+}-S^I_{i-}-\widehat{\gamma}^I(S^I_i,Z_{i,1},D_{i,1}))(\Delta Y_i-
\widehat{\mu}^{Y}_0(Z_{i,1},D_{i,1}))$ and $(S^I_{i+}-S^I_{i-}-\widehat{\gamma}^I(S^I_i,Z_{i,1},D_{i,1}))(\Delta D_i-\widehat{\mu}^{D}_0 (Z_{i,1},D_{i,1}))$, and taking the ratio between these two averages. We propose to use cross-fitting to perform this two-step estimation, and we let $\widehat{\delta}_{\mathsf{IV},\mathsf{DR}}$ denote the resulting estimator. As the definition of $\widehat{\delta}_{\mathsf{IV},\mathsf{DR}}$ is similar to that of $\widehat{\delta}_{2,\mathsf{DR}}$, the full definition is omitted to preserve space. Since all the nuisance functions in Theorem \ref{thm:main_IV} are conditional expectations with respect to two variables, one can still use series estimators with a polynomial order chosen by cross validation within each subsample to estimate them. Results in Section 4 of \cite{andrews1991asymptotic} imply that those estimators converge fast enough to have that $\widehat{\delta}_{\mathsf{IV},\mathsf{DR}}$ is $\sqrt{n}-$consistent and asymptotically normal, as long as the population functions are twice differentiable. Finally, for $U\in\{D,Y\}$, let
\begin{align*}
\delta_{U}&=E\left[\text{sgn}(\Delta Z)\left(\Delta U - E(\Delta U |S^I=0,Z_1,D_1)\right)\right],\\
\psi_{U}&=\frac{1}{E(\left\vert\Delta Z\right\vert)}\left\{\left(S^I_+-S^I_--E(S^I_+-S^I_-|Z_1,D_1)\frac{(1-S^I)}{E(1-S^I|Z_1,D_1)}\right)\right. \\
&\hspace{2.3cm}\times (\Delta U -E(\Delta U | Z_{1},D_1, S^I=0))-\delta_{U}\left\vert\Delta Z\right\vert\bigg\}.
\end{align*}
Then, let $\psi_{\mathsf{IV}}=(\psi_{Y}-\delta_{\mathsf{IV}}\psi_{D})/\delta_{D}$. Under technical conditions similar to those in Section \ref{sec:WAS}, one can show that $\psi_{\mathsf{IV}}$ is the influence function of $\widehat{\delta}_{\mathsf{IV},\mathsf{DR}}$.
\@startsection{subsection}{2}{0mm}{-1\baselineskip}{0.7\baselineskip}{\normalfont\normalsize\bfseries}{Alternative target parameters}\label{appendixsec:alternatives}
One may be interested in other parameters than the AS and WAS we consider in the paper, like the average dose-slope
function
$$(d,d')\mapsto \text{TE}(d,d'):=E\left(\frac{Y_2(d) - Y_2(d')}{d-d'}\right),$$
or the average marginal effect $E\left(Y'_2(D_2)\right)$. What is the appeal
of considering the AS and the WAS, namely averages of $\text{TE}(d_1,d_2|d_1,d_2)$, rather than those
other parameters? Conditional on $D_1=d_1,D_2=d_2$, $Y_2(d_2)$ is observed,
so estimating $\text{TE}(d_1,d_2|d_1,d_2)$ only requires estimating
$Y_2(d_1)$, switchers' counterfactual outcomes at period $2$ if their
treatment had not changed. By definition, $Y_1(d_1)$ is observed. If the data
contains a third period $0$ and the treatment of some units does not change
from period $0$ to $1$, then $Y_0(d_1)$ is also observed for some switchers
and stayers. Then, as explained in Section
\ref{sec:extensions} below, one can run a pre-trends test of Assumption
\ref{hyp:parallel_trends}, by comparing the outcome evolutions of switchers
and stayers from period $0$ to $1$. This shows that averages of
$\text{TE}(d_1,d_2|d_1,d_2)$ can be identified under a
parallel-trends assumption whose plausibility can be assessed via a pre-trends test. On the other hand, estimating $\text{TE}(d,d')$
requires estimating, for most units, \textit{two} unobserved counterfactual
outcomes. This cannot be achieved under an assumption whose plausibility can be assessed via a pre-trends test, as we
only observe \textit{one} potential outcome at each date. Similarly,
estimating $E[(Y_2(D_2) - Y_2(d'))/(D_2-d')]$ requires estimating $Y_2(d')$,
which cannot be achieved under an assumption whose plausibility can be assessed via a pre-trends test, because
$Y_1(d')$ is not observed for all units. As $Y'_2(D_2)=\lim_{d'\rightarrow
D_2}(Y_2(D_2) - Y_2(d'))/(D_2-d')$, this concern also applies to $E\left(Y'_2(D_2)\right)$.
\@startsection{subsection}{2}{0mm}{-1\baselineskip}{0.7\baselineskip}{\normalfont\normalsize\bfseries}{Other extensions}
We return to the case where the treatment satisfies a parallel-trends condition.
\@startsection{subsubsection}{3}{0mm}{-0.8\baselineskip}{0.4\baselineskip}{\normalfont\normalsize\itshape}{Estimators with more than two time periods}\label{subsec:Tgeq3}
This section sketches the extension of our estimators to datasets with three
periods or more. For further details, see Section \ref{appendixsec:Tgeq3} of
the web appendix. With $T>2$ time periods, we propose a simple extension of our estimators. Let $\widehat{\delta}_{1,t,\mathsf{DR}}$ and $\widehat{\delta}_{2,t,\mathsf{DR}}$ be defined exactly as $\hat{\delta}_{1,\mathsf{DR}}$ and $\hat{\delta}_{2,\mathsf{DR}}$ but on a pair of consecutive time periods $(t-1,t)$. Then, we consider a weighted average across $t$ of $\widehat{\delta}_{1,t,\mathsf{DR}}$, with weights proportional to the proportion of switchers from $t-1$ to $t$, thus yielding an AS estimator with more than two periods, denoted $\widehat{\delta}^{T\geq 3}_{1}$. We also consider a weighted average across $t$ of
$\widehat{\delta}_{2,t,\mathsf{DR}}$, with weights proportional to the
average absolute treatment change from $t-1$ to $t$, thus yielding a WAS
estimator with more than two periods, denoted $\widehat{\delta}^{T\geq
3}_{2}$.
\medskip
Those estimators are consistent under the following parallel-trends
assumption: $t-1$-to-$t$ switchers and stayers with the same $t-1$ treatment
would have experienced the same average $t-1$-to-$t$ outcome evolutions if
switchers had not switched. Because it is conditional on the lagged treatment, this assumption cannot be ``chained'' across pairs of periods: it only requires that some groups be on parallel trends over consecutive periods, not over the entire duration of the panel.
We propose ``pre-trends'' or ``placebo'' estimators, to assess the plausibility of that assumption.
The placebo AS and WAS estimators are analogous to $\widehat{\delta}^{T\geq
3}_{1}$ and $\widehat{\delta}^{T\geq 3}_{2}$, but they compare the
$t-2$-to-$t-1$ outcome evolutions of $t-1$-to-$t$ switchers and stayers,
restricting the sample to $t-2$-to-$t-1$ stayers to avoid that those
comparisons be contaminated by the treatment's effect. If one finds that from
$t-2$-to-$t-1$, $t-1$-to-$t$ switchers and stayers are on parallel trends,
this lends credibility to the parallel-trends assumption underlying
$\widehat{\delta}^{T\geq 3}_{1}$ and $\widehat{\delta}^{T\geq 3}_{2}$.
\medskip
With $T>2$ time periods, the assumption that units' potential outcomes does
not depend on their past treatments is of essence. If there are dynamic
treatment effects, $\widehat{\delta}^{T\geq 3}_{1}$ and
$\widehat{\delta}^{T\geq 3}_{2}$ may be biased. Those estimators compare the
outcome evolutions of $t-1$-to-$t$ switchers and stayers. But there may be,
say, $t-1$-to-$t$ stayers whose treatment changed from $t-2$ to $t-1$. If
lagged treatments affect units' current outcome, that change could still
affect the $t-1$-to-$t$ outcome evolution of those stayers, thus leading them
to violate the parallel-trends assumption underlying $\widehat{\delta}^{T\geq
3}_{1}$ and $\widehat{\delta}^{T\geq 3}_{2}$. To mitigate this concern, we
propose the following robustness check. One can recompute
$\widehat{\delta}^{T\geq 3}_{1}$ and $\widehat{\delta}^{T\geq 3}_{2}$,
restricting, for each pair of consecutive periods $(t-1,t)$, the estimation
sample to $t-2$-to-$t-1$ stayers, as in the placebo analysis. We show that
the resulting estimator is robust to dynamic effects up to one treatment lag.
Similarly, if one wants to allow for effects of the first and second
treatment lags on the outcome, one just needs to restrict the estimation
sample to $t-3$-to-$t-1$ stayers. However, the more robustness to dynamic
effects one would like to have, the smaller the estimation sample becomes.
\@startsection{subsubsection}{3}{0mm}{-0.8\baselineskip}{0.4\baselineskip}{\normalfont\normalsize\itshape}{Identification and estimation with covariates}\label{subsec:covariates}
This section sketches how our estimators can be generalized to incorporate covariates, see Section \ref{appendixsec:covariates} of the
web appendix for more details. Consider for
simplicity the case with two periods, $T=2$, where $X_t$ is a $c$-dimensional
vector of covariates. We show in Section \ref{appendixsec:covariates} of the
web appendix that the AS and WAS can be identified and estimated under an
alternative parallel trends assumption that holds conditional on the
first-period covariates $X_1$ (Assumption \ref{hyp:parallel_trendsX}).
Because our estimation and inference methods are fully nonparametric, they
can be readily applied to the case with covariates after redefining the
nuisance functions to depend on the covariates, for instance,
$\mu_0(x,d)=E(\Delta Y|X_1=x,D_1=d,S=0)$ and so on.
\@startsection{section}{2}{0mm}{-1\baselineskip}{1\baselineskip}{\normalfont\large\bfseries}{Application}
\label{sec:appli}
\@startsection{subsection}{2}{0mm}{-1\baselineskip}{0.7\baselineskip}{\normalfont\normalsize\bfseries}{Data and design}
We use the yearly 1966-to-2008 panel dataset of \cite{li2014gasoline},
covering 48 US states (Alaska and Hawaii are excluded). This is a long panel, but recall that our estimators
assume parallel trends only across pairs of consecutive years.
For each state$\times$year cell $(i,t)$, the data
contains $Z_{i,t}$, the total (state plus federal) gasoline tax in cents per
gallon, $D_{i,t}$, the log tax-inclusive price of gasoline, and $Y_{i,t}$,
the log gasoline consumption per adult. We seek to estimate the effect of
taxes on gasoline consumption and prices, and the
price-elasticity of gasoline consumption.\footnote{Instead, \cite{li2014gasoline} jointly estimate the effect of gasoline taxes
and tax-exclusive prices on consumption, using a TWFE regression with two
treatments. Between each pair of periods, the tax-exclusive price
changes in all states, so this treatment does not have stayers and its effect
cannot be estimated using our estimators.}
\medskip
Let $\mathcal{S}$ be the set of switching $(i,t)$ cells such that $Z_{i,t}\ne
Z_{i,t-1}$ but $Z_{i',t}=Z_{i',t-1}$ for some $i'$. The second condition
drops seven pairs of consecutive periods between
which the federal gasoline tax changed, thus implying that taxes changed in all states. $\mathcal{S}$ includes 384 cells, namely
19\% of the 2,016 state$\times$year cells for which $Z_{i,t}-Z_{i,t-1}$ can
be computed. Table \ref{table:balancing} in the web appendix shows that cells in $\mathcal{S}$ are
similar to other cells on a number of observable characteristics. Thus, there is no strong indication that the effects below
apply to a selected subgroup.
\medskip
As an example, the top panel of Figure \ref{fig:support} below shows the
distribution of $Z_{i,1987}$ for 1987-to-1988 stayers, while the bottom panel
shows the distribution for 1987-to-1988 switchers. The figure shows that
there are many values of $Z_{i,1987}$ such that only one or two states have
that value, so $Z_{i,1987}$ is close to being continuously distributed.
Moreover, all switchers $i$ are such that
$$\min_{i':Z_{i',1988}=Z_{i',1987}}Z_{i',1987}\leq Z_{i,1987}\leq \max_{i':Z_{i',1988}=Z_{i',1988}}Z_{i',1987}.$$
Thus, Assumption \ref{hyp:support_condition0} seems to hold for this pair of
years. $(1987,1988)$ is not atypical: there are many other years where
$Z_{i,t}$ is close to being continuously distributed, and almost 95\%
of cells in $\mathcal{S}$ are such that
$\min_{i':Z_{i',t}=Z_{i',t-1}}Z_{i',t-1}\leq Z_{i,t-1}\leq
\max_{i':Z_{i',t}=Z_{i',t-1}}Z_{i',t-1}$. Dropping the few cells that do not
satisfy this condition barely changes the results presented below.
\begin{figure}[H]
\centering
\includegraphics[scale=0.38]{Figures/Support_Lagged_Treatment_1988.pdf}
\caption{Gasoline tax in 1987 among 1987-to-1988 switchers and stayers}
\label{fig:support}
\normalsize
\end{figure}
Turning to the distribution of the tax changes $Z_{i,t}-Z_{i,t-1}$, 90\% of the cells in $\mathcal{S}$ experience an increase in
their taxes, and 10\% experience a decrease. The average value of
$|Z_{i,t}-Z_{i,t-1}|$ is equal to 1.61 cents, while prior to the tax change,
switchers' average gasoline price is equal to 112 cents: our estimators
leverage small changes in taxes relative to gasoline prices. Finally,
$\min_{(i,t)\in \mathcal{S}}|Z_{i,t}-Z_{i,t-1}| = 0.05$ : some switchers
experience a very small change in their taxes.
\@startsection{subsection}{2}{0mm}{-1\baselineskip}{0.7\baselineskip}{\normalfont\normalsize\bfseries}{Reduced-form and first-stage estimates}
Table \ref{table:RFandFS} below shows the AS and WAS estimates
of the reduced-form (Panel A) and first-stage (Panel B) effects of taxes on
gasoline quantities and prices. We control for lagged
prices $D_{t-1}$, to ensure that the resulting IV is robust to
heterogeneous effects over time, as explained in Section \ref{subsec:IV}.
Estimates are computed with cross-fitting, using 10 splits. We use a polynomial of order 1 in
$(Z_{t-1},D_{t-1})$ to estimate $E(\Delta Y_t|Z_{t-1},D_{t-1}, S_t=0)$,
$E(\Delta D_t|Z_{t-1},D_{t-1},S_t=0)$, $E[S/\Delta D|Z_{t-1},D_{t-1}]$,
$P(S_{+,t}=1|Z_{t-1},D_{t-1})$, $P(S_{-,t}=1|Z_{t-1},D_{t-1})$, and
$P(S_{t}=0|Z_{t-1},D_{t-1})$.
10-folds cross-validation selects a polynomial of order one
for all conditional expectations, except for $E(\Delta D_t|Z_{t-1},D_{t-1},S_t=0)$ where it selects a polynomial of order two. Thus, polynomials of order one are in line with those selected by cross validation.\footnote{In a previous version of this paper \citep{de2022difference}, we had shown results very similar to those below with polynomials of order two and without cross fitting. With polynomials of order two and cross fitting, estimates become noisy: due to the small sample size, $\widehat{P}(S_{t}=0|Z_{t-1},D_{t-1})$ is close to zero for some $(i,t)$.} Standard errors
clustered at the state level, computed following \eqref{eq:ASIFT>=3} and
\eqref{eq:WASIFT>=3} in the web appendix, are shown between parentheses. Estimations use 1632 (48$\times$ 35) first-difference observations: 7 periods
are excluded as they do not have stayers. The last line of
each panel shows the p-value of a test that AS=WAS.
\begin{table}[H]
\begin{center}
\begin{threeparttable}
\caption{Effects of gasoline tax on quantities and prices}
\label{table:RFandFS}
\begin{tabular}{lcc}
\toprule
{
\begin{tabular}{lc}
\midrule[.5pt] \multicolumn{2}{c}{Panel A: Reduced-form effect of taxes on quantities}\\\midrule[.5pt]
AS & -0.0043 (0.0027) \\
WAS & -0.0036 (0.0010)\\
p.value & 0.7527\\
\midrule[.5pt] \multicolumn{2}{c}{Panel B: First-stage effect of taxes on prices}\\\midrule[.5pt]
AS & 0.0038 (0.0024)\\
WAS & 0.0058 (0.0009) \\
p.value & 0.2921 \\\midrule[.5pt]
Observations & 1,632 \\
\end{tabular}
}
\\
\bottomrule
\end{tabular}
\end{threeparttable}
\end{center}
\footnotesize
Notes: Panel A (resp. B) shows the AS and WAS estimates of the reduced-form (resp. first-stage) effect of taxes on gasoline quantities (resp. prices). All estimates control for the lag of prices. All conditional expectations are estimated using a polynomial of order 1 in $(Z_{t-1},D_{t-1})$, with
cross-fitting, using ten splits.
Standard errors clustered at the state level are shown next to the estimates, between parentheses. The last line of each panel shows the p-value of a test that the AS and WAS effects are equal.
\end{table}
In Panel A, the AS estimate indicates that increasing gasoline tax
by 1 cent decreases quantities consumed by 0.43 percent on average for the
switchers. That effect is insignificant. The WAS estimate indicates that increasing gasoline tax
by 1 cent decreases quantities consumed by 0.36 percent, an effect that is significant. In Panel B, the AS
estimate of the first-stage effect is positive and insignificant. The WAS estimate is significant, and indicates that if gasoline tax increases by 1 cent on average, prices increase by 0.58 percent. In both panels, the AS estimates are close to, and
not significantly different from, the WAS estimates, but they are insignificant because they are
markedly less precise: their standard errors are almost three times larger, as predicted by Proposition
\ref{prop:variance_comparison}. Therefore, in what follows we focus on the WAS estimates.
\@startsection{subsection}{2}{0mm}{-1\baselineskip}{0.7\baselineskip}{\normalfont\normalsize\bfseries}{Placebo analysis}
Table \ref{table:placeboRFandFS} below shows ``pre-trends'' or ``placebo'' WAS estimates
of the reduced-form and first-stage effects. The placebos are analogous to the actual estimates, but they replace $\Delta Y_t$ by $\Delta Y_{t-1}$, and they restrict the sample to states whose taxes did not change between $t-2$ and $t-1$. This reduces the sample size by around a third. The placebos
are small and insignificant, both for quantities and prices. This shows that before switchers change
their gasoline taxes, switchers' and stayers' consumption and
prices do not follow detectably different evolutions. In Table \ref{table:RFandFS_placebosample} of the web appendix we reestimate the WAS in the placebo subsample. In that subsample, the WAS remains valid if the first lag of taxes affect current gasoline
prices and quantities, as discussed in Section \ref{subsec:Tgeq3}. The estimates therein are close to those in
Table \ref{table:RFandFS}.
\begin{table}[H]
\begin{center}
\begin{threeparttable}
\caption{Placebo WAS effects of gasoline tax on quantities and prices}
\label{table:placeboRFandFS}
\begin{tabular}{lcc}
\toprule
{
\begin{tabular}{lc}
Placebo reduced-form effect of taxes on quantities & 0.0011 (0.0015) \\
Placebo first-stage effect of taxes on prices & 0.0011 (0.0015)\\
\midrule[1pt]
Observations & 1,059\\
\end{tabular}
}
\\
\bottomrule
\end{tabular}
\end{threeparttable}
\end{center}
\footnotesize
Notes: The table shows placebo WAS estimates of the reduced-form and first-stage effects of taxes on quantities and prices. The estimates and their standard errors are computed as the actual estimates, replacing $\Delta Y_t$ by $\Delta Y_{t-1}$, and restricting the sample to states whose taxes did not change between $t-2$ and $t-1$.
\end{table}
\@startsection{subsection}{2}{0mm}{-1\baselineskip}{0.7\baselineskip}{\normalfont\normalsize\bfseries}{Comparison with estimates robust to unrestricted dynamic effects of taxes}
While results seem robust to allowing for dynamic effects up to one lag, one may want to allow for unrestricted dynamic effects. For that purpose, we re-estimate the effect of current taxes on consumption and prices, using the estimators of \cite{de2020difference}. Those estimators are applicable to continuous treatments, and they are robust to dynamic effects up to any lag. Results are shown in Table \ref{table:RFandFS_unrestricted_dyn} of the web appendix. While estimates are of the same sign as those in Table \ref{table:RFandFS}, only 215 $(g,t)$ cells are used in the estimation.
With unrestricted dynamic effects, instead of comparing states whose taxes change/do not change from $t-1$ to $t$, DID estimators need to compare states whose taxes change for the first time at $t$ to states whose taxes have not changed yet at $t$. This reduces the estimation sample to 46 first-time-switcher cells, to which the estimated effects apply, and 169 never-switcher cells used as the control group. Accordingly, estimated effects are insignificant: the standard error of the estimated effect on consumption is 13 times larger than that of the WAS estimate in Table \ref{table:RFandFS},\footnote{The target parameters in \cite{de2020difference} are generalizations of the WAS to models allowing for unrestricted dynamic effects.} while the standard error of the estimated effect on prices is more than nine times larger. Thus, in this application allowing for unrestricted dynamic effects comes with large costs in terms of external validity and statistical precision. As the data is at the yearly level, it may be reasonable to rule out effects of taxes in previous years on current consumption and prices, at least for taxes that prevailed two years or more ago.
\@startsection{subsection}{2}{0mm}{-1\baselineskip}{0.7\baselineskip}{\normalfont\normalsize\bfseries}{IV estimate}
Table \ref{table:IV} shows an IV-WAS estimate
of the price-elasticity of gasoline consumption. As the instrument's first
stage is not very strong and the sample only has 48 observations,
asymptotic approximations may not be reliable for inference. In line with
that conjecture, the bootstrap distribution of the
estimator in Table \ref{table:IV} is non-normal, with some outliers.
Therefore, for inference we use a percentile bootstrap confidence interval, clustered at the
state level. Reassuringly, that confidence interval has close to
nominal coverage in simulations tailored to our application.\footnote{Here is
the DGP used in our simulations. We estimate TWFE regressions of $Y_{i,t}$ on
state and year fixed effects and $Z_{i,t}$, and of $D_{i,t}$ on state and
year fixed effects and $Z_{i,t}$. We let
$\hat{\gamma}^Y_i+\hat{\lambda}^Y_t+\hat{\beta}^YZ_{i,t}+\epsilon^Y_{i,t}$
and
$\hat{\gamma}^D_i+\hat{\lambda}^D_t+\hat{\beta}^DZ_{i,t}+\epsilon^D_{i,t}$
denote the resulting regression decompositions. In each simulation, the
simulated instrument is just the actual instrument, while the simulated
outcomes and treatments are respectively equal to
$Y^s_{i,t}=\hat{\gamma}^Y_i+\hat{\lambda}^Y_t+\hat{\beta}^YZ_{i,t}+\epsilon^{Y,s}_{i,t},$
and
$D^s_{i,t}=\hat{\gamma}^D_i+\hat{\lambda}^D_t+\hat{\beta}^DZ_{i,t}+\epsilon^{D,s}_{i,t},$
where the vector of simulated residuals $(\epsilon^{Y,s}_{i,1},...,
\epsilon^{Y,s}_{i,T},\epsilon^{D,s}_{i,1},..., \epsilon^{D,s}_{i,T})$ is
drawn at random and with replacement from the estimated vectors of residuals
$((\epsilon^{Y}_{i',1},..., \epsilon^{Y}_{i',T},\epsilon^{D}_{i',1},...,
\epsilon^{D}_{i',T}))_{i'\in \{1,...,n\}}$. Thus, the first-stage and
reduced-form effects, the correlation between the reduced-form and
first-stage residuals, and the residuals' serial correlation are the same as
in the sample.} The IV-WAS estimate is negative and significant. It is more negative than, but not significantly different
from, estimates reported in recent literature \citep[][]{coglianese2017anticipation}.
\begin{table}
\begin{center}
\caption{Price-elasticity of gasoline consumption}
\label{table:IV}
\begin{tabular}{lcc}
\toprule
{
\begin{tabular}{lc}
\midrule[.5pt]
IV-WAS & -0.6556 \\
& [-1.1375, -.2526] \\
\midrule[.5pt]
Observations & 1,632 \\
\end{tabular}
}
\\
\bottomrule
\end{tabular}
\end{center}
\footnotesize
Notes: The table shows an IV-WAS estimate of the price-elasticity of gasoline consumption, controlling for the lag of prices. All conditional expectations are estimated using a polynomial of order 1 in $(Z_{t-1},D_{t-1})$, with
cross-fitting, using ten splits.
A percentile-bootstrap confidence interval, computed with 500 bootstrap replications clustered at the state level, is shown below the estimate.
\end{table}
\@startsection{section}{2}{0mm}{-1\baselineskip}{1\baselineskip}{\normalfont\large\bfseries}{Conclusion}
We propose new, doubly-robust difference-in-difference (DID) estimators for treatments continuously distributed at every time period. We assume that between pairs of consecutive periods, the treatment of some units, the switchers, changes, while the treatment of other units, the stayers, does not change. We propose a parallel-trends assumption on the outcome evolution of switchers and stayers with the same baseline treatment. Under that assumption, two target parameters can be estimated. Our first target is the average slope of switchers' period-two potential outcome function, from their period-one to their period-two treatment, referred to as the AS. Our second target is a weighted average of switchers' slopes, where switchers receive a weight proportional to the absolute value of their treatment change, referred to as the WAS. The AS and WAS serve different purposes. When it comes to estimation, the WAS has two advantages. First, it can be estimated at the parametric rate even if units can experience an arbitrarily small treatment change. Second, under some conditions, the asymptotic variance of the WAS estimator is strictly lower than that of the AS. In our application, we use US-state-level panel data to estimate the effect of gasoline taxes on gasoline consumption. The standard error of the WAS is almost three times smaller than that of the AS, and the two estimates are close.
\medskip
We also consider the instrumental-variable case, as there are instances where units experiencing/not experiencing a treatment change are unlikely to be on parallel trends, but one has at hand an instrument such that units experiencing/not experiencing an instrument change are more likely to be on parallel trends. Then, we propose widely applicable IV-DID estimators.
\medskip
Throughout, we assume that there are some stayers, namely units whose treatment does not change between consecutive periods. \cite{de2023nostayers} discuss the extension of the results in this paper to applications without stayers.
\newpage
\bibliography{biblio}
\newpage