EconBase
← Back to paper

Difference-in-Differences Estimators When No Unit Remains Untreated

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

89,466 characters · 0 sections · 70 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Difference-in-Differences Estimators When No Unit Remains Untreated

bibunit\begin{abstract} We consider treatment-effect estimation with a two-periods panel, where units are untreated at period one, and receive strictly positive doses at period two. First, we consider designs with some quasi-untreated units, with a period-two dose local to zero. We show that under a parallel-trends assumption, a weighted average of slopes of units’ potential outcomes is identified by a difference-in-difference estimand using quasi-untreated units as the control group. We leverage results from the regression-discontinuity-design literature to propose a nonparametric estimator. Then, we propose estimators for designs without quasi-untreated units. Finally, we propose a test of the homogeneous-effect assumption underlying two-way-fixed-effects regressions. \end{abstract} \@startsection{section}{2}{0mm}{-1\baselineskip}{1\baselineskip}{\normalfont}{Introduction} We consider treatment-effect estimation in designs in which no unit is treated initially, and then units simultaneously receive heterogeneous and strictly positive treatment doses. We refer to such designs as heterogeneous adoption designs (HADs). To simplify the exposition, in most of the paper we focus on the case with two time periods, where units are untreated at period one and treated at period two, but the extension to applications with several time periods is straightforward, see Section (ref) of the Appendix. HADs are common. They arise when a policy is implemented universally, but exposure varies. For instance, nationwide increases in the minimum wage affect industries or local areas differentially, depending on their prior fraction of workers below the new minimum wage dustmann_reallocation_2022. Similarly, China's accession to the WTO eliminated uncertainty on US-China trade tariffs, but affected US industries differentially depending on their prior levels of tariffs' uncertainty pierce2016surprisingly. Similarly, the creation of Medicare Part D affected drugs differentially, depending on their Medicare market share duggan2010. Such designs also arise when an innovation is introduced, with an exposure rate that varies over space enikolopov2011,Chaisemartin2011fuzzy. In these examples, no unit remains untreated: all industries or local areas have some workers with a wage below the new nationwide minimum wage, tariffs' uncertainty was reduced in all US industries when China joined the WTO, etc. Even when there are untreated units, researchers may prefer to discard them from their analysis, for fear that they differ too much from treated ones. To estimate the treatment's effect in those designs, a common strategy is to run a two-way fixed effects (TWFE) regression. One regresses $Y_{g,t}$, the outcome of unit $g$ at period $t$, on unit fixed effects, an indicator for period two, and $D_{g,t}$, the treatment dose of $g$ at $t$. As there are only two time periods and $D_{g,1}=0$, the coefficient on $D_{g,t}$ is equal to that on $D_{g,2}$ in a regression of $\Delta Y_g$ on a constant and $D_{g,2}$, where $\Delta Y_g=Y_{g,2}-Y_{g,1}$ denotes $g$'s outcome change from period one to two. Proposition S1 of dcDH2020 shows that in HADs, TWFE regressions fail to identify a convex combination of unit-specific effects, and could be misleading if effects vary across units. This paper proposes alternative, nonparametric estimators, as well as a test of the assumptions underlying the TWFE estimator. \@startsection{subsection}{2}{0mm}{-0.8\baselineskip}{0.6\baselineskip}{\normalfont}*{Nonparametric estimators} In the first part of the paper, our target parameter is a weighted average of slopes (WAS) of units' potential outcomes, between a treatment of zero and their actual treatment. That parameter can, for instance, be useful to do a cost-benefit analysis comparing units' actual treatments in period two to the counterfactual where they would have remained untreated. First, we consider designs where $0$ belongs to the support of the period-two treatment, meaning that there is a quasi-untreated group (QUG), that has a period-two treatment local to zero. Then, under a parallel-trends assumption, WAS is identified by an estimand comparing the outcome evolutions of the whole population and of the QUG. We leverage results from the regression discontinuity designs (RDDs) and nonparametric estimation literature Imbens2012,calonico2014,calonico2018effect, to propose: (i) an optimal bandwidth defining the QUG in a data-driven way, (ii) an estimator relying on a local-linear regression, and (iii) a robust confidence interval accounting for the estimator's first-order bias. Our estimator converges at the standard univariate nonparametric rate. It is computed by the did_had Stata and R packages. Second, we consider designs without a QUG. Then, we start by showing that under our parallel-trends assumption, the identification region for WAS is the entire real line. In view of this impossibility result, we consider two routes. First, we consider an additional assumption, which restricts the difference between the treatment effects of the full population and of the QUG, under which the sign of WAS is identified by an estimand using the least-treated units as the control group. Second, we consider another additional assumption, which requires that the treatment effect is the same in the full population and in the QUG, under which a parameter related to, but different from, WAS is identified. Third, as our identifying assumptions and estimators depend on whether there is a QUG, we propose a nonparametric and tuning-parameter-free test of that null hypothesis. \@startsection{subsection}{2}{0mm}{-0.8\baselineskip}{0.6\baselineskip}{\normalfont}*{Test of the assumptions underlying the TWFE estimator} Taking stock, in designs with a QUG we have proposed an estimator that converges at the univariate nonparametric rate. In view of its slow rate of convergence, it may have limited power to detect the treatment's effect when the sample size is not very large. In designs without a QUG, an additional limitation is that our estimator restricts treatment-effect heterogeneity. This motivates us to reconsider the usual TWFE estimator in the second part of the paper. As is now well-known, that estimator relies on strong assumptions: it estimates the treatment's effect if on top of parallel trends, one also assumes that treatment effects are mean-independent of the treatment variable. Those two assumptions imply that \begin{equation} E(\Delta Y_g|D_{g,2})=\beta_0+\beta_{fe} D_{g,2}, \end{equation} i.e. the conditional mean of the outcome's evolution should be linear in the period-two dose. Under the parallel-trends assumption, there is actually an “if and only if” between (ref) and the mean-independent-effects assumption, in designs where we have a QUG. If the data contains another pre-treatment period $t=0$, one can conduct a pre-trends test to assess the plausibility of the parallel-trends assumption. This motivates the following procedure in designs with a QUG: run a pre-trends test of the parallel-trends assumption, and run a test of (ref); if neither test is rejected, use the TWFE estimator. It follows from chaisemartin2024 that if parallel-trends and (ref) hold, conditional on not rejecting the two pre-tests, confidence intervals for $\beta_{fe}$ cannot undercover. Testing (ref) is straightforward when $D_{g,2}$ can take a finite number of values. When $D_{g,2}$ can take an infinite number of values, we use a test proposed by stute1997nonparametric and stute1998bootstrap. That test is nonparametric, tuning-parameter free, consistent, and it has power against local alternatives. We implemented it in the stute_test Stata and R packages. \@startsection{subsection}{2}{0mm}{-0.8\baselineskip}{0.6\baselineskip}{\normalfont}*{Applications} We first revisit garrett_tax_2020. In the 2002 Job Creation and Worker Assistance Act, the US government introduced a bonus depreciation regulation, whereby firms can deduct around 30 percent of the purchase price of a new investment from their taxable income when the investment is made, yielding a decrease in the present value of investment costs. The authors define a county-level treatment capturing exposure to that reform. 12 counties out of 2,954 are untreated, but the number of untreated counties is too low to use them only as the control group, so we also use quasi-untreated counties. Our nonparametric WAS estimate indicates a positive and significant treatment effect on employment, around twice as large as the paper's TWFE estimate, and we find little evidence of differential pre-trends between untreated and quasi-untreated counties and the whole population. Thus, the paper's result seems robust to allowing for heterogeneous effects. We then revisit pierce2016surprisingly. Since 1980, US imports from China have been subject to the low Normal Trade Relations (NTR) tariff rates reserved to WTO members. However, those rates required uncertain and politically contentious annual renewals. Without renewal, US tariffs on Chinese imports would have spiked to higher non-NTR rates. When China joined the WTO in 2001, the US granted it Permanent NTR (PNTR). This eliminated a potential tariff spike, equal to the difference between the non-NTR and NTR tariff. This so-called NTR-gap treatment varies substantially across industries and is strictly positive for all of them. We start by testing whether there are quasi-untreated industries, and find that the test is not rejected. Our nonparametric WAS estimates are noisy, and most of them are insignificant. However, we find that after accounting for industry-specific linear trends, pre-trends tests of the parallel-trends assumption underlying TWFE estimators are not rejected, and the test of (ref) is also not rejected. Then, TWFE estimators with industry-specific linear trends may be reliable in this application. They yield negative and marginally significant effects of eliminating a potential China-tariffs spike on US employment, that are smaller than in pierce2016surprisingly. \@startsection{subsection}{2}{0mm}{-0.8\baselineskip}{0.6\baselineskip}{\normalfont}*{Related Literature and organization of the paper} Difference-in-difference (DID) estimators relying on the standard parallel-trends assumption have been proposed in designs where some units stay untreated at period two deChaisemartin15b,dcDH2020,de2020difference,callaway2021difference. However, those estimators cannot be used in HADs. Few papers discuss designs without an untreated group. An early exception is fricke2017identification, who shows that under the standard parallel-trends assumption, a DID comparing a more-treated to a less-treated group identifies the difference between the treatment's effect in the two groups blundell2009alternative,deChaisemartin15b,callaway2021difference. If one is ready to assume that more- and less-treated groups have the same effect of receiving the lowest dose, then fricke2017identification shows that the more-versus-less-treated DID identifies the effect of receiving the higher rather than the lower dose, for the group that received the higher dose callaway2021difference. However, unlike parallel-trends, the plausibility of this homogeneous-effect assumption cannot be assessed via a pre-trends test, which may be why it is often unpalatable to applied researchers. We make three contributions to that literature. First, we show that in designs with a QUG, one can avoid making an homogeneous-effect assumption, by using a DID with the QUG as the control group, which only relies on a parallel-trends assumption whose plausibility can be assessed by pre-trends tests. Second, when there is no QUG, we exhibit “minimal” homogeneous effect assumptions under which a DID estimand using the least-treated as the control group identifies a causal effect or its sign. Third, we propose a test of the homogeneous-effect assumption. Another related paper is haxhiu2024continuous, who also consider designs without an untreated group. They assume that the period-two dose is discrete\footnote{They sketch an extension to the continuous case, which requires discretizing the continuous dose.} and that the treatment does not have an effect below an unknown “minimum effective” dose. They propose to estimate that dose and then estimate the treatment's effect using units with a treatment below the minimum effective dose as the control group. With respect to their proposal, our nonparametric estimators do not assume that the period-two dose is discrete or that the treatment has no effect below a minimum effective dose. Moreover, by relying on results from the RDD and nonparametric estimation literature we can propose analytic confidence intervals with proven guarantees, while they propose a bootstrap for inference but do not show its validity. Finally, the first part of our paper relies on results from the literature on RDDs and nonparametric estimation Imbens2012,calonico2014,calonico2018effect. Our contribution is to build upon that literature to propose a non-parametric DID estimator in designs without untreated units. The second part of our paper relies on results from the literature on functional-form specification tests stute1997nonparametric,stute1998bootstrap. There, our contribution is to show that the linearity of $E(\Delta Y_g|D_{g,2})$ is a testable implication of the homogeneous and linear effect assumption underlying TWFE estimators, and to further show that under parallel trends, in designs with a QUG there is an “if and only if” between the two conditions. The paper is organized as follows. In Section (ref), we introduce the set-up, the notation and our main assumptions. In Section (ref), we introduce our nonparametric estimators. Section (ref) considers the TWFE estimator. In Section (ref), we revisit two empirical applications. The proof of our first result appears in the text, other proofs are collected in the appendix. \@startsection{section}{2}{0mm}{-1\baselineskip}{1\baselineskip}{\normalfont}{Setup, assumptions, and examples} \@startsection{subsection}{2}{0mm}{-0.8\baselineskip}{0.6\baselineskip}{\normalfont}{Setup} \paragraph{Panel data.} We are interested in estimating the effect of a treatment on an outcome, using a panel data set, with $G$ units and two time periods. In Appendix (ref), we discuss how results generalize to designs with more time periods. Units could be aggregate entities, like regions or industries. Henceforth, units are indexed by $g$ and time periods by $t$. \paragraph{Treatment and potential outcomes.} Let $D_{g,t}$ denote the treatment of unit $g$ at $t$, with $D_{g,t}\ge 0$. The potential outcome of unit $g$ at $t$ under treatment $d$ is $Y_{g,t}(d)$, and the observed outcome is $Y_{g,t}:=Y_{g,t}(D_{g,t})$. Our potential outcome notation rules out anticipatory effects: units' outcome at period one does not depend on their period-two treatment. Our notation also rules out dynamic or carry-over effects: units' outcome at period two does not depend on their past treatments. In the designs we consider, units are all untreated at period one, so the no carry-over effects assumption is not of essence, we just impose it to simplify notation. On the other hand, the no-anticipation assumption is of essence. \paragraph{Notation and convention.} Hereafter, we let $\text{Supp}(A)$ denote the support of the random variable $A$, namely the smallest closed set $C$ such that $P(A\in C )=1$. Throughout the paper, whenever a limit is introduced in an assumption, it is implicitly assumed that this limit exists. $\Delta$ denotes the first-difference operator. \@startsection{subsection}{2}{0mm}{-0.8\baselineskip}{0.6\baselineskip}{\normalfont}{Assumptions} We consider the three following assumptions. \begin{hyp} (i.i.d. sample) $(Y_{g,1},Y_{g,2},D_{g,1},D_{g,2})_{g=1,...,G}$ are i.i.d. \end{hyp} As units are i.i.d., we drop the $g$ subscript below, except when we introduce estimators. We also let $\underline{d}:=\inf \text{Supp}(D_2)$ be the infimum of the support of the period-two treatment. \begin{hyp} (The least treated and the full population are on parallel trends)\\ $\lim_{d\downarrow \underline{d}}E\left[\Delta Y(0)|D_2\le d\right]=E\left[\Delta Y(0)\right]$. \end{hyp} Assumption (ref) requires that the least-treated group experiences the same average outcome evolution without treatment as the overall population. If the data contains another pre-treatment period $t=0$, we can assess the plausibility of Assumption (ref), by computing a pre-trends estimator comparing the average of $Y_1-Y_0=Y_1(0)-Y_0(0)$ in the overall population and among the least-treated group, see below for further details. When $D_2$ is discrete, we simply have $\lim_{d\downarrow \underline{d}}E\left[\Delta Y(0)|D_2\le d\right]=E\left[\Delta Y(0)|D_2=\underline{d}\right]$. Thus, with a binary treatment, Assumption (ref) reduces to the standard parallel-trends assumption: it requires that untreated and treated units have the same average outcome evolutions without treatment. \begin{hyp} (Uniform continuity of potential outcomes) There exists $d_0>\underline{d}$ such that for all $\varepsilon \in (0,\infty)$, there exists $\delta>0$ such that for all $(d,d')\in [0,d_0]^2$, $|d-d'|\le \delta$ implies that $|Y_2(d)-Y_2(d')|\le \varepsilon$. \end{hyp} Assumption (ref) imposes almost-sure uniform continuity of $d\mapsto Y_2(d)$. It holds for instance if $d\mapsto Y_2(d)$ is Lipschitz, namely $|Y_2(d)-Y_2(d')|\le M|d-d'|$ for some $M>0$, but it can also hold if $d\mapsto Y_2(d)$ has unbounded slopes. On the other hand, Assumption (ref) implies that $d\mapsto Y_2(d)$ is almost surely continuous at $0$. This rules out extensive margin effects, where infinitesimally small treatment doses already have a non-zero effect caetano2015test. \@startsection{subsection}{2}{0mm}{-0.8\baselineskip}{0.6\baselineskip}{\normalfont}{Heterogeneous adoption designs} We consider heterogeneous adoption designs (HADs), where all units are untreated at period one,\footnote{Our results also apply to designs where units all receive the same non-zero treatment dose $d_1\ne 0$ at period one. What is key is that all units receive the same period-one dose.} and receive heterogeneous, strictly positive treatment doses at period two. \begin{des} (HAD) $D_{1}=0$, $D_{2}>0$, and $V(D_2)>0$. \end{des} Some of our results apply to a subset of Design (ref), namely designs with a QUG. \setcounter{des}{0} \begin{des}{\bf '} (HAD with a quasi-untreated group) The conditions in Design (ref) hold and $\underline{d}=0$. \end{des} The condition in Design (ref)' holds when there are units whose period-two treatment is “very close” to zero: for any $\delta>0$, $P(0<D_2<\delta)>0$. For instance, it holds if $D_2$ is continuously distributed on $\mathbb{R}_+$ with a continuous density that is strictly positive at 0. \paragraph{Survey of HADs.} We conducted two searches. First, we used the Google Scholar advanced search with keywords “treatment intensity”, “twoway fixed effects”, and “differences in differences”, among articles published by the American Economic Review and the American Economic Journal: Applied Economics. Second, we conducted a specific search targeting papers relying on the so-called “fraction-treated” design, which is popular to estimate the effect of the minimum wage. Then, the treatment variable is defined as the fraction of individuals whose wage is below the new minimum wage level, and the design leverages variation in that treatment across firms, industries, or regions, thus corresponding exactly to an heterogeneous adoption design. We started from haanwinckel_does_2025, a recent paper investigating the fraction-treated design, and looked at the papers cited therein that use that design. Then, we looked at the papers cited by those papers that also use that design to estimate the effect of the minimum wage. Overall we found 10 papers with an HAD, or where the untreated group accounts for less than 1% of the sample, thus implying that using it only as the control group would likely yield a noisy estimator. Five papers come from our first search, and five come from our second search. For six papers, we could verify that the estimation sample of a TWFE regression shown in the paper does not have an untreated group. For two papers, we could verify that a TWFE regression has an estimation sample with less than 1% of untreated units. Two papers do not report their treatment's minimal value and the dataset is not publicly available, but they are unlikely to have an untreated group based on the nature of the treatment so we included them. Overall, HADs seem common, all the more so as our searches are not exhaustive: they did not retrieve several of the examples mentioned in the introduction. \@startsection{section}{2}{0mm}{-1\baselineskip}{1\baselineskip}{\normalfont}{Nonparametric estimators} \@startsection{subsection}{2}{0mm}{-0.8\baselineskip}{0.6\baselineskip}{\normalfont}{Designs with a quasi-untreated group} Throughout this section, we assume that $\underline{d}=0$: we are in a design with a QUG. \@startsection{subsubsection}{3}{0mm}{-0.6\baselineskip}{0.4\baselineskip}{\normalfont}{Target parameter} \paragraph{Actual-versus-no-treatment slopes.} As $D_2>0$, let $$\text{TE}_2:=\frac{Y_2(D_2)-Y_2(0)}{D_2}$$ denote the slopes of units' potential outcome functions between $0$ and their actual treatments, which we refer to as the actual-versus-no-treatment slope. In this paper, our target parameters are averages of the slopes $\text{TE}_2$. Instead, one may prefer to estimate the mapping $$d\mapsto E\left(\frac{Y_2(d)-Y_2(0)}{d}\right),$$ for instance to assess if the treatment has increasing, decreasing, or constant returns. Here is the reason why we instead focus on averages of $\text{TE}_2$. Estimating $\text{TE}_2$ only requires estimating $Y_2(0)$, which, as we will see, can be achieved under Assumption (ref), a parallel-trends assumption whose plausibility can be assessed via pre-trends tests. Instead, for any $0<d\ne D_2$, estimating $(Y_2(d)-Y_2(0))/d$ also requires estimating $Y_2(d)$. As $Y_2(d)$ is not observed at $t=2$ and at prior periods, estimating it requires making assumptions whose plausibility cannot be assessed via pre-trends tests. \paragraph{Averages of slopes.} $\text{TE}_2$ applies to only one unit and cannot be consistently estimated under Assumption (ref). We therefore turn attention to averages of the $\text{TE}_2$ slopes, that can be consistently estimated. Let \begin{align*} AS:=&E\left[TE_2\right]\\ WAS:=& E\left[\frac{D_2}{E[D_2]}TE_2\right]. \end{align*} $\text{AS}$ is the Average Slope of treated units. This parameter generalizes the well-known average treatment effect on the treated parameter to our setting with a non-binary treatment. $\text{WAS}$ is a weighted average of treated units' slopes, where units with a larger period-two treatment receive more weight. chaisemartin2022continuous put forward an economic and a statistical argument as to why WAS, while seemingly less natural than AS, may be a relevant target. First, WAS may actually be the relevant quantity to consider in a cost-benefit analysis comparing the actual treatments $D_2$ to a counterfactual policy where all units would have remained untreated at period two. Assume that the outcome is expressed in monetary units, and that treatment is costly, with a cost linear in dose, uniform across units, and known to the analyst: the cost of giving $d$ units is $c\times d$ for some known $c$. Then, $D_2$ is beneficial relative to $0$ if and only if $E(Y_2(D_2)-c_2D_2)>E(Y_2(0))$, namely if and only if WAS$>c:$ comparing WAS to $c$ is sufficient to evaluate if changing the treatment from $0$ to $D_2$ was beneficial. Second, estimating $\text{AS}$ may sometimes be more difficult than estimating $\text{WAS}$. When there are quasi-untreated units with a $D_2$ close to zero, the denominator of $\text{TE}_2$ is close to zero for those units. Then, estimators of those units' slopes may suffer from a small-denominator problem, which could substantially increase their variance, thus making it impossible to estimate AS at the standard $\sqrt{G}-$rate graham2012identification,sasaki2021slow. On the other hand, \begin{equation} \text{WAS}=E\left[\frac{D_2}{E[D_2]}\frac{Y_2(D_2)-Y_2(0)}{D_2}\right]=\frac{E[Y_2(D_2)-Y_2(0)]}{E[D_2]}, \end{equation} so estimators of $\text{WAS}$ are not affected by a small-denominator problem, even if there is a QUG. In this section, $\text{WAS}$ is our target parameter. \paragraph{Estimating a conditional average slope function?} An alternative goal could be to estimate $d_2\mapsto E\left[\text{TE}_2|D_2=d_2\right]$, the conditional average-slope function. However, as noted by fricke2017identification and callaway2021difference, variations in $E\left[\text{TE}_2|D_2=d_2\right]$ across values of $d_2$ conflate a dose-response relationship, of economic interest, with a selection bias which is often not of economic interest. Note that under a parallel-trends assumption on the untreated potential outcome, as we consider here, estimators of $d_2\mapsto E\left[\text{TE}_2|D_2=d_2\right]$ proposed previously callaway2021difference cannot be used in the settings we consider, because there are no untreated units. \@startsection{subsubsection}{3}{0mm}{-0.6\baselineskip}{0.4\baselineskip}{\normalfont}{Identification of $\text{WAS}$} \begin{thm} Suppose that we are in Design (ref)' and Assumptions (ref) and (ref) hold. Then, \begin{equation} \text{WAS}=\frac{E[\Delta Y]-\lim_{d\downarrow 0} E\left[\Delta Y|D_2\le d\right]}{E[D_2]}. \end{equation} \end{thm} Theorem (ref) shows that with a QUG, $\text{WAS}$ is identified by an estimand comparing the average outcome evolution to that of the QUG, scaled by the average treatment. \paragraph{Proof of Theorem (ref).} Fix $\varepsilon>0$. By Assumption (ref), $\exists \delta\in (0, d_0)$ such that $\forall d\le \delta$, \begin{align} \left|E[Y_2(D_2)-Y_2(0)|D_2\le d]\right| & \le E\left[\left|Y_2(D_2)-Y_2(0)\right| \big| D_2\le d\right] \le \varepsilon. \end{align} Because $\varepsilon>0$ was arbitrary, we obtain $\lim_{d\downarrow 0} E[Y_2(D_2)-Y_2(0) | D_2\le d]=0$. Since $$E[\Delta Y|D_2\le d] = E[\Delta Y(0)|D_2\le d] + E[Y_2(D_2)-Y_2(0)| D_2\le d],$$ we get, by Assumption (ref), that $\lim_{d\downarrow 0} E\left[\Delta Y|D_2\le d\right]=E[\Delta Y(0)]$. Moreover, \begin{align*} E[\Delta Y]-\lim_{d\downarrow 0} E\left[\Delta Y|D_2\le d\right]=&E[Y_2(D_2)-Y_1(0)]-E[Y_2(0)-Y_1(0)]\nonumber\\ =&E[Y_2(D_2)-Y_2(0)]. \end{align*} The result follows from the previous display and (ref) \textbf{QED.} \@startsection{subsubsection}{3}{0mm}{-0.6\baselineskip}{0.4\baselineskip}{\normalfont}{Estimation of $\text{WAS}$} \paragraph{Estimators' definition.} To estimate $\lim_{d\downarrow 0} E\left[\Delta Y|D_2\le d\right]$, we rely on a local linear regression, as in regression discontinuity designs (RDDs) and more generally in nonparametric estimation. We define the following estimators, indexed by a bandwidth $h$: $$\widehat{\beta}^{\text{np}}_h:=\frac{\frac{1}{G}\sum_{g=1}^G \Delta Y_g - \widehat{\mu}_h}{\frac{1}{G}\sum_{g=1}^G D_{g,2}},$$ with $\widehat{\mu}_h$ the intercept in the local linear regression of $\Delta Y_g$ on $D_{g,2}$, weighting observations by $k(D_{g,2}/h)/h$, for a kernel function $k$ and a bandwidth $h>0$. \paragraph{Regularity conditions underlying the estimator.} Let $m(d):=E[\Delta Y|D_2=d]$ and $\sigma^2(d):=V(\Delta Y|D_{2}=d)$. One can derive the asymptotic behavior of $\widehat{\beta}^{\text{np}}_h$ under the conditions below: \begin{hyp}[Regularity conditions] There exists a neighborhood $\mathcal{V}$ of 0 such that: \begin{enumerate} • The cumulative distribution function of $D_{2}$ is differentiable on $\mathcal{V}$, with derivative denoted by $f_{D_2}$. Moreover, $\lim_{d\downarrow 0} f_{D_2}(d)>0$. • $m$, defined on $\text{Supp}(D_{2})$, is twice differentiable on $\mathcal{V}$ and $\lim_{d\downarrow 0} m''(d)$ exists. • $\sigma^2(.)$, defined on $\text{Supp}(D_{2})$, is continuous on $\mathcal{V}$ and $\lim_{d\downarrow 0} \sigma^2(d)>0$. • $k$ is bounded and has bounded support. • As $G\to\infty$, the bandwidth $h_G$ satisfies $h_G\to 0$ and $Gh_G\to \infty$. \end{enumerate} \end{hyp} Assumption (ref) is inspired from commonly-made assumptions in the RDD literature: it imposes regularity conditions on the distribution of $D_{2}$, the function $m$, the kernel function and the bandwidth. Hereafter, we let $f_{D_2}(0):=\lim_{d\downarrow 0} f_{D_2}(d)$, and we define $m(0)$, $m''(0)$, and $\sigma^2(0)$ similarly.\footnote{One can show that if $\lim_{d\downarrow 0} m''(d)$ exists, then $\lim_{d\downarrow 0} m(d)$ also exists.} One difference with the RDD literature is that while there, the cut-off point lies in the interior of the support of the running variable, here $0$ lies on the boundary of the support of $D_2$. Then, the assumption that $f_{D_2}(0)>0$ is not innocuous. Below, we show in simulations based both on synthetic and on real data, that the confidence interval for WAS we propose can still have close to nominal coverage when $f_{D_2}(0)=0$. Proposing estimators for the case where $f_{D_2}(0)=0$ is an interesting avenue for future research, that goes beyond the scope of this paper. \paragraph{Estimators' asymptotic distribution.} Let $\kappa_k:=\int_0^\infty t^k k(t)dt$ for $k\in\mathbb N$ and \begin{align*} k^*(t) & := \frac{\kappa_2 - \kappa_1 t}{\kappa_0 \kappa_2 -\kappa_1^2} k(t), \\ C & := \frac{\kappa_2^2 - \kappa_1 \kappa_3}{\kappa_0 \kappa_2 -\kappa_1^2}. \end{align*} Since $\sum_{g=1}^G \Delta Y_g/G$ and $\sum_{g=1}^G D_{g,2}/G$ are root-$G$ consistent, their randomness is negligible compared to that of $\widehat{\mu}_h$. Thus, $$G^{2/5}\left(\widehat{\beta}^{\text{np}}_{h_G} - \text{WAS} \right) = G^{2/5}\frac{\widehat{\mu}_{h_G} - m(0)}{E[D_{2}]} + o_P(1).$$ Then, following the exact same reasoning as in, say, Imbens2012, we obtain that under Assumptions (ref)-(ref), \begin{equation} \sqrt{Gh_G}\left(\widehat{\beta}^{\text{np}}_{h_G} - \text{WAS} - h_G^{2}\frac{C m”(0)}{2E[D_2]} \right) \stackrel{d}{\longrightarrow} \mathcal{N}\left(0, \frac{\sigma^2(0) \int_0^\infty k^*(u)^2du}{E[D_2]^2 f_{D_2}(0)}\right). \end{equation} The fastest rate of convergence is obtained with $G^{1/5}h_G \to c>0$, in which case \begin{equation} G^{2/5}\left(\widehat{\beta}^{\text{np}}_{h_G} - \text{WAS} \right) \stackrel{d}{\longrightarrow} \mathcal{N}\left(\frac{c^2C m”(0)}{2E[D_2]} , \frac{\sigma^2(0) \int_0^\infty k^*(u)^2du}{c E[D_2]^2 f_{D_2}(0)}\right). \end{equation} \paragraph{Optimal bandwidth and robust confidence interval.} Based on (ref), one can derive a so-called optimal bandwidth, which, as in RDDs Imbens2012, minimizes the asymptotic mean squared error of $\widehat{\beta}^{\text{np}}_{h_G}$. Then, inference on $\text{WAS}$ is not straightforward, because the asymptotic distribution of $$\sqrt{Gh_G^*}(\widehat{\beta}^{\text{np}}_{h_G^*}-\text{WAS})$$ has a first-order bias that needs to be accounted for. However, the approach for local-polynomial regressions in calonico2018effect readily applies to our set up, in particular because it can be used to estimate a conditional expectation function at a boundary point. We rely on their results and software implementation calonico2019nprobust to: \begin{enumerate} • estimate an optimal bandwidth $\widehat{h}^*_G$; • compute $\widehat{\mu}_{\widehat{h}^*_G}$; • compute $\widehat{M}_{\widehat{h}_G^*}$, an estimator of $\widehat{\mu}_{\widehat{h}^*_G}$'s first-order bias; • compute $\widehat{V}_{\widehat{h}_G^*}$, an estimator of the variance of $\widehat{\mu}_{\widehat{h}^*_G}-\widehat{M}_{\widehat{h}_G^*}$. \end{enumerate} With those inputs, we can simply define our estimator as \begin{equation} \widehat{\beta}^{\text{np}}_{\widehat{h}_G^*}:=\frac{\frac{1}{G}\sum_{g=1}^G \Delta Y_g - \widehat{\mu}_{\widehat{h}^*_G}}{\frac{1}{G}\sum_{g=1}^G D_{g,2}}, \end{equation} and its bias-corrected confidence interval as \begin{equation} \left[\widehat{\beta}^{\text{np}}_{\widehat{h}_G^*}+\frac{\widehat{M}_{\widehat{h}_G^*}}{\frac{1}{G}\sum_{g=1}^G D_{g,2}} \pm\frac{q_{1-\alpha/2}\sqrt{\widehat{V}_{\widehat{h}_G^*}/(G\widehat{h}_G^*)}}{\frac{1}{G}\sum_{g=1}^G D_{g,2}}\right], \end{equation} where $q_x$ denotes the quantile of order $x$ of a standard normal. We refer to calonico2018effect for conditions ensuring the asymptotic validity of this confidence interval. \@startsection{subsubsection}{3}{0mm}{-0.6\baselineskip}{0.4\baselineskip}{\normalfont}{Extensions} \paragraph{Pre-trend and event-study estimators.} If the data contains another pre-treatment period $t=0$ where units are all untreated, we can define a pre-trends estimator mimicking $\widehat{\beta}^{\text{np}}_{\widehat{h}_G^*}$, where we just replace $\Delta Y_g$ by $ Y_{g,1}-Y_{g,0}$. One can show that if the analogue of Assumption (ref) from period 0 to 1 holds, this pre-trends estimator converges to zero. Therefore, if the pre-trends estimator is significantly different from zero, one rejects the analogue of Assumption (ref) from period 0 to 1, thus casting doubt on Assumption (ref). If the data contains another post-treatment period $t=3$, one may be interested in estimating the treatment effect at period $3$ rather than at period $2$, to assess the effect of being exposed to treatment for two periods. Then, one can just define an event-study estimator mimicking $\widehat{\beta}^{\text{np}}_{\widehat{h}_G^*}$, where one replaces $\Delta Y_g$ by $ Y_{g,3}-Y_{g,1}$. See Appendix (ref) for further details on the extension to multiple periods. \paragraph{Computation.} The \texttt{did_had} Stata did_hadStata and R did_hadR commands compute $\widehat{\beta}^{\text{np}}_{\widehat{h}_G^*}$ and its bias-corrected confidence interval. When more than one pre-treatment period is available, the commands can compute pre-trend estimators. With more than one post-treatment period, the commands can compute event-study estimators. \texttt{did_had} heavily relies on the \texttt{nprobust} package of calonico2019nprobust, which should be cited, together with calonico2018effect, whenever \texttt{did_had} is used. \texttt{did_had} uses the same default choices as \texttt{nprobust} for the kernel (Epanechnikov) and bandwidth (MSE-optimal bandwidth). \@startsection{subsubsection}{3}{0mm}{-0.6\baselineskip}{0.4\baselineskip}{\normalfont}{Simulations} Table (ref) below shows the results from a simulation study, where we assess the coverage rate of the bias-corrected confidence interval (BCCI) based on (ref) and computed by \texttt{did_had}. We consider three DGPs. In DGP 1, $D_2$ follows a uniform distribution on $[0,1]$, $\Delta Y(0)$ follows a standard normal independent of $D_2$, and $\Delta Y_{2}(D_2) = D_2 + D_2^2 + \Delta Y(0)$, thus implying that $\text{WAS}=5/3$. Assumptions (ref), (ref), and (ref) hold in this DGP, so the BCCI should have close to nominal coverage when the sample size is large enough. When $G=100$, $\hat{\beta}^{np}_{\hat{h}^*_G}$ is slightly upward biased, and the 95% level BCCI slightly undercovers, with a coverage rate of 89%. Its coverage rate increases to 93% when $G=500$, and to 95% when $G=2500$. DGP 2 is similar to DGP 1, except that now $D_2$ is drawn from a Beta distribution $B(2,2)$, thus implying that WAS$=8/5$. The rationale for DGP 2 is to test coverage in a setting where Assumption (ref) fails because $f_{D_2}(0) = 0$. With a coverage rate of 90%, the 95% level BCCI slightly undercovers when $G=100$, but its coverage almost reaches 95% when $G=2500$. Thus our BCCI can still have close to nominal coverage when $f_{D_2}(0) = 0$. Finally, in DGP 3 we draw $D_2$ without replacement from the empirical distribution of $(D_{g,2})_{g\in \{1,...,G\}}$ in pierce2016surprisingly, we independently draw without replacement $\Delta Y(0)$ from the empirical distribution of $(Y_{g,2}-Y_{g,1})_{g\in \{1,...,G\}}$, and we let $\Delta Y_{2}(D_2) =\Delta Y(0)$, so that WAS$=0$. Again, Assumption (ref) fails because $f_{D_2}(0) = 0$. When $G=100$, the 95% level BCCI has a coverage rate of 92%. Its coverage rate increases when the sample size increases, and reaches almost 95% when $G=2500$. Overall, these results suggest that inference based on (ref) should be reliable for moderately large sample sizes. \begin{table}[H] \caption{Simulation Results} \aboverulesep=0ex\belowrulesep=0ex \begin{tabular}{lcccccc} \toprule & \multicolumn{2}{c}{DGP 1} & \multicolumn{2}{c}{DGP 2} & \multicolumn{2}{c}{DGP 3} \\ \cmidrule(lr){2-3}\cmidrule(lr){4-5}\cmidrule(lr){6-7} & (1) & (2) & (1) & (2) & (1) & (2) \\ \midrule G = 100& 1.69& 0.89& 1.65& 0.90& -0.0013& 0.92 \\ G = 500& 1.70& 0.93& 1.63& 0.90& -0.0005& 0.94\\ G = 2500& 1.68& 0.95& 1.63& 0.94& 0.0002& 0.94 \\ \bottomrule \end{tabular} \\ \begin{minipage}\textwidth {\textit{Notes:} Column (1) shows $\hat{\beta}^{np}_{\hat{h}^*_G}$, computed following Equation (ref). Column shows (2) the coverage of the bias-corrected 95%-level CI, computed following Equation (ref). The three DGPs are described in the text. 2,000 simulations were conducted per DGP/sample size.} \end{minipage} \end{table} \@startsection{subsection}{2}{0mm}{-0.8\baselineskip}{0.6\baselineskip}{\normalfont}{Designs without a quasi-untreated group} \@startsection{subsubsection}{3}{0mm}{-0.6\baselineskip}{0.4\baselineskip}{\normalfont}{Identification problem} Without a QUG, we face an identification issue: since all units receive non-negligible treatment doses at period two, the outcome evolution of all units is the sum of the counterfactual evolution $\Delta Y(0)$ and of the treatment effect $D_2\text{TE}_2$, thus preventing us from identifying $E(\Delta Y(0))$. The following proposition formalizes this, by showing that without a QUG, we can rationalize any value of WAS under Assumptions (ref)-(ref). \begin{prop} Suppose that we are in Design (ref), $\underline{d}>0$ and Assumptions (ref)-(ref) hold. Then, the identification set for WAS is $\mathbb R$. \end{prop} In view of this impossibility result, we consider two restrictions on the heterogeneity of the treatment effect between the least treated and the overall population, under which comparing the outcome evolutions of those two groups identifies or at least partially identifies some causal effects. We conclude this section by discussing other routes for identification. \@startsection{subsubsection}{3}{0mm}{-0.6\baselineskip}{0.4\baselineskip}{\normalfont}{Identification of the sign of WAS} \paragraph{Restricting the ratio between the treatment effect of the least treated and the WAS.} Below, we use the convention that $0/0=0$. \begin{hyp} $$\frac{\lim_{d\downarrow \underline{d}} E(\text{TE}_2|D_2\le d)}{\text{WAS}}< \frac{E(D_2)}{\underline{d}}.$$ \end{hyp} First, assume that WAS $\ne 0$. Then, as $\frac{E(D_2)}{\underline{d}}>0$, Assumption (ref) mechanically holds when the average of TE$_2$ among the least treated is of a different sign than the WAS. When the two effects are of the same sign, Assumption (ref) requires that the least treated have an average response per dose of treatment less than $E(D_2)/\underline{d}$ times larger than the WAS. This condition may be plausible when $\underline{d}$ is much lower than $E(D_2)$. Note that while Assumption (ref) is probably not very restrictive when $\underline{d}$ is much lower than $E(D_2)$, unlike Assumption (ref) its plausibility cannot be assessed via a pre-trends test. Finally, in the knife-edge case where WAS $=0$, Assumption (ref) requires that $\lim_{d\downarrow \underline{d}} E(\text{TE}_2|D_2\le d)\le 0$. \paragraph{Sufficient condition for Assumption (ref) to hold.} The following condition is sufficient for Assumption (ref) to hold: for all $(d,d')\in \text{Supp}(D_2)^2$, \begin{equation} 0\leq \frac{E(\text{TE}_2|D_2=d)}{E(\text{TE}_2|D_2=d')}< \frac{E(D_2)}{\underline{d}}. \end{equation} (ref) restricts the heterogeneity of the conditional average slope function, by assuming that its sign never changes, and that no subgroup $D_2=d$ has an average response per dose more than $\frac{E(D_2)}{\underline{d}}$ times larger than that of another subgroup $D_2=d'$. An alternative to Assumption (ref) could for instance be to assume that $$0\leq \frac{E(\text{TE}_2|D_2=d)}{E(\text{TE}_2|D_2=d')}<K,$$ for a user-specified choice of $K$. Theorem (ref) below implies that under this alternative approach, the breakdown value of $K$ until which the sign of the WAS is identified is at least equal to $\frac{E(D_2)}{\underline{d}}$. \paragraph{Identification of the sign of WAS.} \begin{thm} Suppose that we are in Design (ref) and Assumptions (ref), (ref), and (ref) hold. Then, \begin{equation} \text{WAS}\geq 0 \Leftrightarrow \frac{E[\Delta Y]-\lim_{d\downarrow \underline{d}} E\left[\Delta Y |D_2\le d\right]}{E[D_2-\underline{d}]}\geq 0. \end{equation} \end{thm} Note that the estimand identifying the sign of WAS in Theorem (ref) is similar to that identifying WAS in Theorem (ref), except that we replace $D_2$ by $D_2-\underline{d}$ and $0$ by $\underline{d}$. \@startsection{subsubsection}{3}{0mm}{-0.6\baselineskip}{0.4\baselineskip}{\normalfont}{Identification of an alternative target parameter} Let \begin{align*} \text{WAS}_{\underline{d}}:=& E\left[\frac{D_2-\underline{d}}{E[D_2-\underline{d}]}\text{TE}_{2, \underline{d}}\right], \end{align*} where\footnote{$\text{TE}_{2, \underline{d}}$ may be undefined if $P(D_2=\underline{d})>0$. We can define it arbitrarily in this case since it is multiplied by 0 anyway in WAS$_{\underline{d}}$.} $$\text{TE}_{2, \underline{d}}:=\frac{Y_2(D_2)-Y_2(\underline{d})}{D_2-\underline{d}}.$$ $\text{WAS}_{\underline{d}}$ is similar to $\text{WAS}$, except that it is a weighted average of the actual-versus-lowest-treatment slopes $\text{TE}_{2, \underline{d}}$, instead of the actual-versus-no-treatment slopes $\text{TE}_2$. In the same way that WAS is the relevant quantity to consider in a cost-benefit analysis comparing $D_2$ to a counterfactual policy where all units would have remained untreated at period two, $\text{WAS}_{\underline{d}}$ is the relevant quantity to consider in a cost-benefit analysis comparing $D_2$ to a counterfactual where all units would have received the lowest dose at period two. However, we view this second cost-benefit analysis as less policy relevant than the first: in HADs, a planner may be able to control whether one can access treatment, as indicated by the fact that no one is treated at period one. On the other hand, the planner may not be able to control units' treatment choices once they have access, as indicated by the fact that doses are heterogeneous at period two. Thus, a counterfactual where all units receive the lowest dose may not be something that the planner can implement. Still, $\text{WAS}_{\underline{d}}$ retains a “weakly causal” interpretation as a convex combination of slopes blandhol2022tsls.\footnote{Under Assumptions (ref)-(ref) alone, one can show the analogue of Proposition (ref) for $\text{WAS}_{\underline{d}}$: the identification set for that parameter is the entire real line, see the end of the proof of Proposition (ref) for more details.} WAS$_{\underline{d}}$ is identified under the following assumption. \begin{hyp} (The least-treated group and the full population have the same effect of receiving the lowest dose) $\lim_{d\downarrow \underline{d}} E\left[Y_2(\underline{d})- Y_2(0)|D_2\le d\right]=E\left[Y_2(\underline{d})- Y_2(0)\right].$ \end{hyp} Similarly to Assumption (ref), Assumption (ref) restricts treatment effect heterogeneity, by requiring that the least treated and the full population have the same effect of receiving the lowest dose. Again, one cannot assess Assumption (ref)'s plausibility via a pre-trends test. \begin{thm} Suppose that we are in Design (ref) and Assumptions (ref), (ref), and (ref) hold. Then, \begin{equation} \text{WAS}_{\underline{d}}=\frac{E[\Delta Y]-\lim_{d\downarrow \underline{d}} E\left[\Delta Y |D_2\le d\right]}{E[D_2-\underline{d}]}. \end{equation} \end{thm} \@startsection{subsubsection}{3}{0mm}{-0.6\baselineskip}{0.4\baselineskip}{\normalfont}{Estimation} Irrespective of whether one works under Assumption (ref) or (ref), the estimation target is $$\frac{E[\Delta Y]-\lim_{d\downarrow \underline{d}} E\left[\Delta Y |D_2\le d\right]}{E[D_2-\underline{d}]}.$$ In designs with no untreated group but with a QUG, the distribution of $D_2$ has to be continuous, at least in a neighborhood of zero. In designs without a QUG, the distribution of $D_2$ may have a mass point at $\underline{d}$. Then, the estimation target reduces to $$\frac{E[\Delta Y]-E\left[\Delta Y |D_2=\underline{d}\right]}{E[D_2-\underline{d}]} = \frac{E[\Delta Y|D_2>\underline{d}]-E\left[\Delta Y |D_2=\underline{d}\right]}{E[D_2|D_2>\underline{d}]-E[D_2|D_2=\underline{d}]},$$ which can straightforwardly be estimated replacing expectations by sample averages, or by running a 2SLS regression of $\Delta Y$ on $D_2$, using $1\{D_2>\underline{d}\}$ as the instrument for $D_2$. When instead $D_2$ is continuously distributed in a neighborhood of $\underline{d}$, to estimate $$\frac{E[\Delta Y]-\lim_{d\downarrow \underline{d}} E\left[\Delta Y |D_2\le d\right]}{E[D_2-\underline{d}]}$$ we can follow the same steps as in Section (ref), replacing $0$ by $\underline{d}$ and $D_2$ by $D_2-\underline{d}$ in the estimators' definition. The only difference is that $\underline{d}$ must be estimated. But since the estimator requires the density of $D_2$ to be positive at $\underline{d}$ (now, Point 1 of Assumption (ref) needs to hold at $\underline{d}$ instead of $0$), $\min_g D_g$ converges to $\underline{d}$ at rate $G$, which is much faster than the nonparametric rate of convergence for $\widehat{\beta}^{\text{np}}_{h_G}$ (see Section (ref) above). Thus, the randomness associated with $\min_g D_g$ is asymptotically negligible. \@startsection{subsubsection}{3}{0mm}{-0.6\baselineskip}{0.4\baselineskip}{\normalfont}{Alternative routes to identification} While there are alternative routes to identification than those presented in Theorems (ref) and (ref), all those we can envision also amount to restricting treatment effect heterogeneity. In a previous version of this paper dechaisemartin2024twowayfixedeffectsdifferencesindifferencesv4, we had derived sharp bounds assuming that $|E(\text{TE}_2|D_2)|$ has a known bound. This approach implicitly restricts treatment effect heterogeneity, and it will often yield wide and uninformative bounds. We had also proposed a parametric approach assuming a known functional form for $E(\text{TE}_2|D_2)$. This approach also implicitly restricts treatment effect heterogeneity, in a perhaps less transparent and more parametric manner than Assumptions (ref) and (ref). Another alternative route would be to leverage results for RD designs with a discrete running variable kolesar2018inference. There as well, we do not observe any unit close to the threshold, which is similar to not observing any unit with a dose close to zero in our setting. However, importing those estimators into our context would require that the analyst specifies a bound on the second derivative of $E(\Delta Y|D_2)$. As $E(\Delta Y|D_2=d_2)=E(\Delta Y(0))+d_2E(\text{TE}_2|D_2=d_2)$, that second derivative involves $\partial E(\text{TE}_2|D_2=d_2)/\partial d_2$, so bounding it also implicitly restricts treatment effect heterogeneity, in a perhaps less interpretable manner than in Assumptions (ref) and (ref). \@startsection{subsection}{2}{0mm}{-0.8\baselineskip}{0.6\baselineskip}{\normalfont}{Testing the null that there is a QUG} Our identifying assumptions and our estimators depend on whether or not there is a QUG. Therefore, we now propose a test of the null hypothesis that $\underline{d}=0$, against $\underline{d}>0$. We consider the following simple and tuning-parameter-free test, of nominal level $\alpha$. The test statistic is $T = D_{2,(1)}/(D_{2,(2)}-D_{2,(1)})$, where $D_{2,(1)}\le ... \le D_{2,(G)}$ denotes the order statistic of $(D_{2,g})_{g=1,...,G}$. The critical region is $W_\alpha:=\{T>1/\alpha - 1\}$. Intuitively, we reject the null if the distance between $D_{2,(1)}$ and 0 is more than $1/\alpha - 1$ times larger than the distance between $D_{2,(2)}$ and $D_{2,(1)}$. We show that this test has asymptotic size equal to $\alpha$ and nontrivial local power on a broad class of cdfs for $D_2$. Specifically, let $\mathcal{D}$ denote the set of cdfs on the real line, let $\overline{d}>\underline{d}\ge 0$, $m>0$ and $K>0$, and let \begin{align*} \mathcal{F}^{\underline{d}, \overline{d}}_{m,K} := & \; \bigg\{F\in \mathcal{D}: F \text{ is differentiable on } [\underline{d}, \overline{d}], \; F(\underline{d}) =0, \, F'(d)\ge m\; \forall d\in [\underline{d}, \overline{d}] \\ & \; \text{ and } |F'(d) -F'(d_1)|\le K |d-d_1| \; \forall (d_1,d)\in [\underline{d}, \overline{d}]^2\bigg\}. \end{align*} This set includes differentiable cdfs whose support has infimum equal to $\underline{d}$ and with density bounded from below in a neighborhood of $\underline{d}$. We focus on cdfs with positive density in the neighborhood of $\underline{d}$ because this condition is required in Assumption (ref). Hereafter, we denote probabilities by $P_F$ instead of $P$, to emphasize their dependence on the cdf of $D_2$. \begin{thm} Fix $\alpha\in (0,1)$, $\overline{d}>\underline{d}> 0$, $m>0$, and $K>0$. We have: \begin{enumerate} • (Asymptotic size control) $\lim\sup_{G\to\infty} \sup_{F\in \mathcal{F}^{0,\overline{d}}_{m,K}} P_F(W_\alpha) = \alpha$. • (Uniform consistency) $\lim\inf_{G\to\infty} \inf_{F\in \mathcal{F}^{\underline{d},\overline{d}}_{m,K}} P_F(W_\alpha) = 1$. • (Local power) $\forall (\underline{d}_G)_{G\ge 1}:\lim\inf G \underline{d}_G >0$, $\lim\inf_{G\to\infty} \inf_{F\in \mathcal{F}^{\underline{d}_G,\overline{d}}_{m,K}} P_F(W_\alpha) >\alpha.$ \end{enumerate} \end{thm} Point 1 of Theorem (ref) establishes the asymptotic validity of the test. Point 2 shows that the test is consistent against fixed alternatives. Finally, Point 3 shows that the test has power against local alternatives converging to the null at rate $G$. We conjecture that under the null that $\underline{d}=0$, drawing inference on WAS only conditional on not rejecting the pre-test above does not distort inference. The reason behind this conjecture is asymptotic independence of extreme-order statistics and sample averages, see e.g. Theorem 2.4 in li2024limit. Also, $T$ above only depends on the extremes $(D_{2,(1)}, D_{2,(2)})$ and calonico2014 show that up to a remainder term, $\widehat{\mu}_{\widehat{h}^*_G}-\widehat{M}_{\widehat{h}_G^*}$ is equal to the sample mean of some appropriate i.i.d. variables. However, the distribution of these variables depends on $G$, and we are not aware of a generalization of Theorem 2.4 in li2024limit to triangular arrays. \@startsection{section}{2}{0mm}{-1\baselineskip}{1\baselineskip}{\normalfont}{TWFE estimator} \paragraph{Motivation and outline.} In designs with a QUG, we have proposed a $G^{2/5}$-consistent treatment-effect estimator, and $G^{2/5}$-consistent pre-trend estimators that one can use to assess the plausibility of the parallel-trends assumption underlying the estimator. Yet, in view of their slow rate of convergence, when the sample size is not very large those estimators may have limited power to detect the treatment's effect and differential trends. In designs without a QUG, an additional limitation of the estimators we have proposed is that they restrict treatment-effect heterogeneity. This leads us to reconsider the usual TWFE estimator in this section. While, as is now well-known, that estimator relies on strong assumptions, we show that those assumptions are testable. When the corresponding tests are not rejected, the TWFE estimator may be an appealing alternative to the estimators proposed in the previous section. In particular, its variance might be lower. \paragraph{Definition of the TWFE estimator.} Let $\widehat{\beta}_{fe}$ denote the coefficient on $D_{g,2}$ in a regression of $\Delta Y_g$ on a constant and $D_{g,2}$. $\widehat{\beta}_{fe}$ is equal to the coefficient on $D_{g,t}$ in a regression of $Y_{g,t}$ on unit fixed effects, an indicator for period $2$, and $D_{g,t}$. Let $$\beta_{fe}:=\frac{E[(D_{2}-E(D_{2}))\Delta Y]}{E[(D_{2}-E(D_{2}))D_{2}]}$$ denote the probability limit of $\widehat{\beta}_{fe}$ when $G\rightarrow +\infty$ under Assumption (ref) and if $E[D_2^2]<\infty$ and $E[\Delta Y^2]<\infty$. \@startsection{subsection}{2}{0mm}{-0.8\baselineskip}{0.6\baselineskip}{\normalfont}{Conditions under which the TWFE estimator is consistent for AS.} \begin{hyp} (Parallel trends across all treatment doses) There is a real number $\mu_0$ such that $E\left[\Delta Y(0)|D_2\right]=\mu_0$. \end{hyp} Assumption (ref) requires that units receiving different treatment doses at period two would all have experienced parallel trends without treatment. This assumption, which is similar to a strong exogeneity assumption in panel data models, is stronger than Assumption (ref). Yet, imposing Assumption (ref) rather than Assumption (ref) is not sufficient to restore identification of the WAS when there is no QUG: one can show that the impossibility result in Proposition (ref) still holds when one replaces Assumption (ref) by Assumption (ref). Relatedly, Assumption (ref) is not sufficient to ensure that DIDs comparing strongly and weakly treated units identify a causal effect: to obtain that result one needs to assume that the effect of receiving a low dose is the same in the two groups fricke2017identification. \begin{hyp} (Homogeneous and linear effect) $E(\text{TE}_2|D_2)=\text{AS}.$ \end{hyp} The following set of conditions is sufficient for Assumption (ref) to hold: \begin{align} &E\left(\frac{Y_2(d)-Y_2(0)}{d}\middle|D_2=d\right)=E\left(\frac{Y_2(d)-Y_2(0)}{d}\right),\\ &E\left(\frac{Y_2(d)-Y_2(0)}{d}\right)=\text{AS}, \end{align} for almost all $d\in\text{Supp}(D_2)$. (ref) is an homogeneous effect assumption, which assumes that the effect of moving treatment from 0 to $d$ is the same for groups of units with different treatment doses. (ref) is a linear effect assumption. To ease exposition, we refer to Assumption (ref) as an homogeneous and linear effect assumption, though (ref) and (ref) are sufficient but not necessary for Assumption (ref). Under Assumption (ref), one can show that \begin{equation} \beta_{fe}=E\left(\frac{(D_2-E(D_2))D_2}{E\left((D_2-E(D_2))D_2\right)}E(\text{TE}_2|D_2)\right). \end{equation} Equation (ref) is not a new result: it is essentially an asymptotic version of the decomposition in Proposition S1 of dcDH2020, who also coined the HAD terminology. It says that $\beta_{fe}$ is a weighted sum of the conditional average slopes (CAS) $E(\text{TE}_2|D_2=d)$, across all values that $D_2$ can take, where $E(\text{TE}_2|D_2=d)$ receives a weight proportional to $(d-E(D_2))d$. Therefore, under Assumption (ref), $\widehat{\beta}_{fe}$ is not consistent for $\text{AS}$ and it does not even converge to a convex combination of CAS: some weights in (ref) are necessarily negative, since $P(0<D_2<E(D_2))>0$. If one further assumes that Assumption (ref) holds, it follows from (ref) that $\text{AS}=\beta_{fe}:$ the TWFE estimator is consistent for AS. Of course, under Assumption (ref) AS and WAS are equal, so the TWFE estimator is also consistent for WAS. \@startsection{subsection}{2}{0mm}{-0.8\baselineskip}{0.6\baselineskip}{\normalfont}{Assumptions (ref) and (ref) are testable: they imply that $E(\Delta Y|D_2)$ is linear.} Let $\beta_0=E(\Delta Y)-\beta_{fe}E(D_2)$ denote the intercept of the TWFE regression. \begin{thm} \begin{enumerate} • In Design (ref), if Assumptions (ref) and (ref) hold, then $E(\Delta Y|D_2)=\beta_0+\beta_{fe}D_2$. • In Design (ref)', if Assumptions (ref) and (ref) hold, $E(\Delta Y|D_2)=\beta_0+\beta_{fe}D_2$ implies that Assumption (ref) holds and $\beta_{fe}=$WAS. • In Design (ref), if Assumptions (ref) and (ref) hold and $E(Y_2(\underline{d})-Y_2(0)|D_2)=\delta_0$, $E(\Delta Y|D_2)=\beta_0+\beta_{fe}D_2$ implies that $\beta_{fe}=\text{WAS}_{\underline{d}}.$ \end{enumerate} \end{thm} \paragraph{When a test that $E(\Delta Y|D_2)$ is linear is rejected, we recommend not using $\widehat{\beta}_{fe}$.} Point 1 of Theorem (ref) shows that if Assumptions (ref) and (ref) hold, $E(\Delta Y|D_2)$ is linear. By contraposition, if $E(\Delta Y|D_2)$ is not linear, then Assumptions (ref) and (ref) cannot hold. Then, we recommend not using $\widehat{\beta}_{fe}$. A caveat of that recommendation is that Assumptions (ref) and (ref) are not the weakest conditions under which $\widehat{\beta}_{fe}$ is consistent for AS. For instance, $\widehat{\beta}_{fe}$ is consistent for AS if \begin{equation} \text{Cov}(\Delta Y(0),D_2)=0 \end{equation} \begin{equation} \text{Cov}\left(\frac{(D_2-E(D_2))D_2}{E\left((D_2-E(D_2))D_2\right)},E(\text{TE}_2|D_2)\right)=0, \end{equation} two conditions that are respectively weaker than Assumption (ref) and (ref). Yet, we are not aware of a test of (ref) and (ref).\footnote{If the data contains another pre-period $t=0$ where units are all untreated, as in period $t=1$, one can run a pre-trends test of (ref), for instance by regressing $Y_{1} - Y_{0}$ on $D_{2}$, because $Y_{1} - Y_{0}=Y_{1}(0) - Y_{0}(0)$ is an outcome evolution without treatment. On the other hand, we are not aware of a test or placebo test that researchers can use to assess the plausibility of (ref).} This reduces the appeal of invoking (ref) and (ref) to justify $\beta_{fe}$: tests or placebo tests of identifying assumptions have become a hard-to-bypass standard to substantiate the credibility of one's findings in observational studies imbens2001estimating,imbens2024lalonde. \paragraph{When linearity of $E(\Delta Y|D_2)$ is not rejected, one may use $\widehat{\beta}_{fe}$ if a pre-trends test of Assumption (ref) and a test that there is a QUG are also not rejected.} If the data contains another pre-period $t=0$ where units are all untreated, as in period $t=1$, one can run a pre-trends test of Assumption (ref), by assessing if $Y_{1} - Y_{0}$ is mean-independent of $D_{2}$, because $Y_{1} - Y_{0}=Y_{1}(0) - Y_{0}(0)$ is an outcome evolution without treatment. Then, in designs with a QUG, Point 2 of Theorem (ref) shows that under Assumptions (ref) and (ref), there is an “if and only if” relationship between the homogeneous-and-linear-effect condition in Assumption (ref) and the linearity of $E(\Delta Y|D_2)$. This leads us to propose the following estimation rule: i) test the null that there is a QUG, using the test in Section (ref); ii) conduct of pre-trends test of Assumption (ref); iii) test that $E(\Delta Y|D_2)$ is linear; iv) if none of these tests is rejected, one may use $\widehat{\beta}_{fe}$ to estimate the treatment's effect. Below, we discuss the consequences of these pre-tests on inference on AS. \paragraph{Without a QUG, if $E(\Delta Y|D_2)$ is linear, a pre-trends test of Assumption (ref) is not rejected, and one is ready to assume $E(Y_2(\underline{d})-Y_2(0)|D_2)=\delta_0$, one may use $\widehat{\beta}_{fe}$.} In designs without a QUG, Point 3 of Theorem (ref) shows that under Assumptions (ref) and (ref) and if one is ready to assume constant effects of the lowest dose ($E(Y_2(\underline{d})-Y_2(0)|D_2)=\delta_0$), if $E(\Delta Y|D_2)$ is linear then $\beta_{fe}=\text{WAS}_{\underline{d}}$. However, in this case the key condition $E(Y_2(\underline{d})-Y_2(0)|D_2)=\delta_0$ cannot be tested. \@startsection{subsection}{2}{0mm}{-0.8\baselineskip}{0.6\baselineskip}{\normalfont}{Testing that $E(\Delta Y|D_2)$ is linear and $Y_{1} - Y_{0}$ is mean-independent of $D_{2}$.} \paragraph{Nonparametric and tuning-parameter-free tests when $D_2$ takes a finite number of values.} Assume that $D_2$ takes $K$ values. If $K=2$, we necessarily have $E(\Delta Y|D_2)=\beta_0+\beta_{fe}D_2$, so there is no room for testability. If $K>2$, to test that $E(\Delta Y|D_2)$ is linear one can just regress $\Delta Y_g$ on a constant, $D_{g,2}$, $D_{g,2}^2$, ... $D_{g,2}^{K-1}$, and test that the coefficients on $D_{g,2}^2$, ..., $D_{g,2}^{K-1}$ are all zero. One can proceed similarly to test that $Y_{1} - Y_{0}$ is mean-independent of $D_{2}$, except that one tests that the coefficients on $D_{g,2}$, ..., $D_{g,2}^{K-1}$ are all zero, and now there is room for testability even when $K=2$. \paragraph{Nonparametric and tuning-parameter-free tests when $D_2$ is continuous.} When $D_2$ is continuous, we rely on stute1997nonparametric to test the linearity of $E(\Delta Y|D_2)$. Under the null hypothesis that $E(\Delta Y|D_2)$ is linear, then $(\widehat{\varepsilon}_{\text{lin},g})_{g=1,...,G}$, the residuals of the linear regression of $\Delta Y_g$ on $D_{g,2}$, should not be correlated with any function of $D_2$. Then, consider the so-called cusum process of the residuals: $$c_G(d):=G^{-1/2} \sum_{g=1}^G \mathds{1}\left\{D_{g,2}\le d\right\} \widehat{\varepsilon}_{\text{lin},g}.$$ stute1997nonparametric shows that under the null hypothesis, $c_G$, as a process indexed by $d$, converges to a Gaussian process. On the other hand, under the alternative, $c_G(d)$ tends to infinity for some $d$. Then, one can consider Kolmogorov-Smirnov or Cram\'er-von Mises test statistics based on $c_G(d)$. We use a Cram\'er-von Mises test statistic: $$S := \frac{1}{G} \sum_{g=1}^G c^2_G(D_{g,2}).$$ To help build intuition, notice that sorting the data by $D_{g,2}$ and denoting by $(g)$ the resulting indexation, one has that $$S = \sum_{g = 1}^G \left(\frac{g}{G}\right)^2 \left(\dfrac{1}{g} \sum_{h = 1}^g \hat{\varepsilon}_{lin,(h)} \right)^2.$$ Now, $(1/g) \sum_{h = 1}^g \hat{\varepsilon}_{lin,(h)} \approx E(\varepsilon_\text{lin}|D_2\leq D_{2,g})$, and the null hypothesis that $E(\varepsilon_\text{lin}|D_2)=0$ holds if and only if $E(\varepsilon_\text{lin}|D_2\leq d)=0$ for all $d$ in the support of $D_2$. The limiting distribution of $S$ under the null is complicated, but stute1998bootstrap show that one can approximate it using the wild bootstrap. Specifically, consider i.i.d. random variables $(\eta_g)_{g=1,...,G}$ with $E[\eta_g]=0$, $E[\eta_g^2]=E[\eta_g^3]=1$.\footnote{In practice, we use the standard two-point distribution: $\eta_g=(1+\sqrt{5})/2$ with probability $(\sqrt{5}-1)/(2\sqrt{5})$, $\eta_g=(1-\sqrt{5})/2$ otherwise.} Then, let $\widehat{\varepsilon}^*_{\text{lin},g}:=\widehat{\varepsilon}_{\text{lin},g}\eta_g$ and $$\Delta Y^*_g = \widehat{\beta}_{0} + \Delta D_g \widehat{\beta}_{fe} + \widehat{\varepsilon}^*_{\text{lin},g}.$$ Then, we compute $S^*$, the bootstrap counterpart of $S$ based on the sample $(D_g, \Delta Y^*_g)_{g=1,...,G}$. One can also use the Stute test to test that $Y_{1} - Y_{0}$ is mean-independent of $D_{2}$: in the definition of the test statistics, one just replaces the residuals from a regression of $Y_2-Y_1$ on a constant and $D_2$ by the residuals from a regression of $Y_1-Y_0$ on a constant. \paragraph{Properties of the Stute test.} stute1997nonparametric and stute1998bootstrap show that the test has asymptotically correct size, is consistent under any fixed alternative, and has non-trivial power against local alternatives converging towards the null at the $1/G^{1/2}$ rate. \paragraph{Pre-testing and inference on AS.} We recommend pre-testing that $E[\Delta Y|D_2]$ is linear and that $E[Y_1-Y_0|D_2]=E[Y_1-Y_0]$ before estimating AS with the TWFE estimator. Theorem (ref) and the results in chaisemartin2024 imply that if, indeed, $E[\Delta Y|D_2]$ is linear and $E[Y_1-Y_0|D_2]=E[Y_1-Y_0]$, and Assumptions (ref) and (ref) hold, inference on AS conditional on accepting the two tests is still valid.\footnote{Also, the result in li2024limit discussed in Section (ref) implies that the test that there is a QUG is asymptotically independent of the TWFE estimator, so inference on AS conditional on accepting this test is also valid. We conjecture that inference on AS conditional on accepting the three tests is valid as well, but proving this would require proving independence between extreme-order statistics and the empirical process involved in the Stute test, and we are not aware of such a result in the literature.} \paragraph{Computation of the Stute test.} The \texttt{stute_test} Stata stute_testStata and R stute_testR commands compute the test. The test's p-value computation relies on the wild bootstrap, which is demanding in terms of computational power. To reduce computation time, the Stata and R commands implementing the test use a vectorization of the test statistic, see Online Appendix (ref) for further details. With this method, the test runs very quickly with moderate sample sizes. For instance, it takes less than one second with $G=5,000$. It still runs in less than two minutes with $G=50,000$. However, the vectorization requires specifying a $G\times G$ matrix, so starting at around $G=100,000$ the test does not run anymore on standard computers, which cannot allocate enough memory to store such a large matrix. For large datasets, we develop another test that does not rely on the bootstrap, and whose computation time remains below one minute, even for datasets as large as $G=50,000,000$. That test is computed by the \texttt{yatchew_test} Stata and R commands. See online Appendix (ref) for further details. \@startsection{section}{2}{0mm}{-1\baselineskip}{1\baselineskip}{\normalfont}{Applications} \@startsection{subsection}{2}{0mm}{-0.8\baselineskip}{0.6\baselineskip}{\normalfont}{Effect of bonus depreciation on employment.} \paragraph{Research question and data.} In the 2002 Job Creation and Worker Assistance Act, the US government introduced a bonus depreciation regulation, whereby firms can deduct around 30 percent of the purchase price of a new investment from their taxable income in the year the investment is made, yielding a decrease in the present value of investment costs. That decrease is larger for longer-lived assets: as those assets depreciate more slowly, an immediate tax deduction is more valuable in present value as otherwise deductions would have been realized further in the future. Based on this insight, garrett_tax_2020 generate a county-level exposure measure $D_g$, equal to county $g$'s labor force share that works in industries relying on long-lived assets. Our sample is the data used to produce Panel A of Figure 2 in the paper, a panel with employment at the county-industry-year level from 1997 to 2012. As exposure is at the county-year level, to fit in the setting considered in this paper we aggregate employment at the county-year level, thus yielding a balanced panel with 2,954 counties observed over 16 years. $Y_{g,t}$ is the growth of employment in county $g$ from 2001 to $t$, and $D_{g,t}=D_g 1\{t\geq 2002\}$. \paragraph{There are untreated and quasi-untreated counties in this application.} 12 counties are such that $D_g=0$, so the null hypothesis $\underline{d}=0$ is verified. However, the number of untreated counties is too low to use them only as the control group. One could also drop those 12 counties for fear that they differ too much from the overall population. We keep them, but dropping them yields results similar to those below. After dropping untreated counties, our test of the null hypothesis $\underline{d}=0$ is not rejected: $D_{2,(1)}=0.044$, $D_{2,(2)}=0.069$, and $T=1.77$, thus yielding a p-value of 0.361. Therefore, there are both untreated and quasi-untreated counties in this application. \paragraph{Nonparametric event-study estimators indicate a positive effect of the treatment on employment.} Figure (ref) below shows nonparametric event-study and pre-trend estimators, and their bias-corrected confidence intervals (BCCIs). Those estimators and BCCIs are computed as described in Sections (ref) and (ref). For instance, at $t=2002$ (resp. $t=2003$, $t=2004$...), the event-study estimator is computed as $\widehat{\beta}^{\text{np}}_{\widehat{h}_G^*}$, defining the outcome variable as $Y_{g,2002}-Y_{g,2001}$ (resp. $Y_{g,2003}-Y_{g,2001}$, $Y_{g,2004}-Y_{g,2001}$...). Pre-trend estimators are small. They are marginally significant and positive in 2000 and in 1999, thus suggesting that prior to the reform, employment was increasing less on average than in the least-treated counties. Assuming that such differential trends would have continued after the reform, then event-study estimators are downward biased. Starting in 2004, event-study estimators are positive and either significant or marginally insignificant, thus suggesting a positive effect of the reform on employment. A caveat is that we can only test for parallel trends for four years before the reform, while the long-run event-study estimators require parallel trends to hold for up to eleven years. Our event-study estimates are mostly similar to, but up to almost twice as large as the TWFE estimates shown on Panel A of Figure 2 in garrett_tax_2020. Thus, even if both estimates differ quantitatively, the conclusion that the treatment increases employment seems robust to allowing for heterogeneous effects across counties. \@startsection{subsection}{2}{0mm}{-0.8\baselineskip}{0.6\baselineskip}{\normalfont}{Effect on US employment of eliminating a potential tariff spike on China} \paragraph{Research question and data.} In 2001 the United States (US) granted Permanent Normal Trade Relations (PNTR) to China, thus eliminating the possibility of a tariff spike on Chinese imports. The treatment in this application is the magnitude of the potential tariffs' spike eliminated by the reform, which varies from 2 to 64 percentage points across industries, with a mean and standard deviation respectively equal to 30 and 14 percentage points. pierce2016surprisingly estimate its effect on US employment. Their data is proprietary, except that used to produce their Table 3. Our sample is the data used to estimate the regression in Column (3) therein,\footnote{Column (2) is a placebo looking at the effect of the treatment on employment in Europe, while Column (1) is a triple difference comparing the effect of the treatment in the US and in Europe.} namely a panel of 103 US industries from 1997 to 2002 and from 2004 to 2005,\footnote{As noted by pierce2016surprisingly, 2003 data is missing for all US industries in the UNIDO dataset used to produce the table. While the UNIDO dataset downloaded by the authors had 104 industries, the version we downloaded in 2023 has 103 industries, presumably due to some industry regrouping.} where $Y_{g,t}$ is the log employment of industry $g$ in year $t$ and $D_{g,t}$ is the potential tariff spike eliminated by the PNTR reform for industry $g$ in year $t$, which is by definition equal to zero for $t<2001$. \begin{figure}[H] \caption{Effects of Bonus Depreciation on Employment } \begin{center} \begin{minipage}\textwidth {\textit{Notes:} Figure (ref) shows nonparametric event-study estimators of the effect of bonus depreciation employment, computed following Equation (ref), using a 1997 to 2012 panel of 2,954 US counties constructed from the data of garrett_tax_2020. The figure also shows pre-trend estimators, as well as estimators' bias-corrected 95% level confidence intervals, computed following Equation (ref).} \end{minipage} \end{center} \end{figure} \paragraph{No group is untreated, but our test that there is a QUG is not rejected.} $D_{g,t}>0$ for $t\geq 2001$, so we follow Section (ref) to test the null that $\underline{d}=0$, against $\underline{d}>0$. $D_{2,(1)}=0.020$, $D_{2,(2)}=0.024$, and $T =6.150$. Therefore, the null is not rejected (p-value$=0.140$). \paragraph{Nonparametric event-study estimators are insignificant.} The blue line of Figure (ref) below shows nonparametric event-study and pre-trend estimators, and their BCCIs, computed as described in Sections (ref) and (ref). For instance, at $t=2001$ (resp. $t=2002$, $t=2004$...), the event-study estimator is computed as $\widehat{\beta}^{\text{np}}_{\widehat{h}_G^*}$, defining the outcome variable as $Y_{g,2001}-Y_{g,2000}$ (resp. $Y_{g,2002}-Y_{g,2000}$, $Y_{g,2004}-Y_{g,2000}$...). The 2001, 2004, and 2005 event-study estimators are insignificant, and only the 2002 estimator is significantly negative: using nonparametric estimators, there is not very strong evidence that eliminating a potential tariff spike on Chinese imports reduces US employment. Note that in view of the small sample size ($G=103$), the nonparametric estimators lack power in this application, and they cannot rule out large effects, especially in 2004 and 2005. Those estimators probably have more power when computed on the authors' proprietary dataset, which has 315 industries. This dataset is built from the Longitudinal Business Database, which can only be accessed from a US Federal Statistical Research Data Center. \paragraph{TWFE event-study estimators are significant, but so are TWFE pre-trend estimators.} The red line of Figure (ref) shows regressions of $Y_{g,t}-Y_{g,2000}$ on $D_{g,2001}$, for $t\in \{1997, 1998, 1999, 2001, 2002, 2004, 2005 \}$. As $G=103$ is not large, we follow the recommendations from imbens2016robust and use HC2 standard errors with the DOF adjustment recommended by bell2002bias to obtain more reliable confidence intervals. $\widehat{\beta}_{fe,t}$ is small and insignificant in $t=$2001, before becoming large and significant in $t=$2002, and even larger in $t=$2004 and $t=$2005.\footnote{The event-study TWFE estimators are all less negative than that in Table 3 Column (3) of the paper. This is because our regressions are not weighted. With weighting, some of our coefficients become more negative than the coefficient in the paper.} However, all pre-trend estimators are also statistically significant, though their magnitude is smaller than that of the event-study estimators. Thus, Assumption (ref) might be violated, though it does not seem that violations of Assumption (ref) can fully account for the event-study effects.\footnote{Those findings are at odds with those from Figure 2 in pierce2016surprisingly. Therein, the authors compute the same pre-trend estimators as we do, on the proprietary dataset they use for most of their analysis, where industries are defined at a more disaggregated level than in our data, and they do not find statistically significant pre-trends. Their regressions are weighted unlike ours, but if we weight our regressions by industries' 1997 employment, pre-trends tests are still rejected. It seems that while disaggregated treatments are uncorrelated with industries' pre-trends, the aggregated variables are correlated, a version of the so-called “ecological inference problem”.} \paragraph{TWFE pre-trend estimators are no longer significant when industry-specific linear trends are controlled for, and some TWFE event-study estimators remain marginally significant.} On the red line of Figure (ref), the pre-trend estimators increase as we look at employment evolutions over a longer horizon. Then, the violation of Assumption (ref) might be due to industry-specific linear trends correlated with industries' treatments, and Assumption (ref) might be plausible when linear trends are controlled for. Accordingly, we replace Assumption (ref) by \begin{equation} E(Y_{g,t}(0)-Y_{g,2000}(0)-(t-2000)\times (Y_{g,2000}(0)-Y_{g,1999}(0))|D_{g,2001})=\mu_t. \end{equation} $Y_{g,2000}(0)-Y_{g,1999}(0)$ captures industry $g$'s linear trend without treatment. Then, (ref) requires that $Y_{g,t}(0)-Y_{g,2000}(0)-(t-2000)\times (Y_{g,2000}(0)-Y_{g,1999}(0))$, industries' deviations from their linear trend, be mean-independent from the NTR-gap treatment. Under this assumption, TWFE event-study estimators can be obtained by regressing, for $t\in \{2001,2002,2004,2005\}$, $Y_{g,t}-Y_{g,2000}-(t-2000)\times (Y_{g,2000}-Y_{g,1999})$ on $D_{g,2001}$. Similarly, to assess the plausibility of (ref), one can either compute TWFE pre-trend estimators, by regressing $Y_{g,t}-Y_{g,1999}-(t-1999)\times (Y_{g,2000}-Y_{g,1999})$ on $D_{g,2001}$ for $t\in \{1998, 1997\}$, or one can run a joint Stute test that \begin{equation} E(Y_{g,t}-Y_{g,1999}-(t-1999)\times (Y_{g,2000}-Y_{g,1999})|D_{g,2001})=\mu_t \forall t\in \{1998, 1997\}. \end{equation} The green line of Figure (ref) shows TWFE estimators with linear trends. Pre-trend estimators in 1997 and 1998 are small and insignificant,\footnote{With linear trends, the pre-trends estimator is mechanically equal to zero in 1999.} and the p-value of the joint Stute test of (ref) is equal to $0.51$. This lends support to (ref). TWFE event-study estimators are smaller and less significant with than without linear trends, but the estimated effect in 2004 is still significant at the 5% level, and that in 2002 is significant at the 10% level. \begin{figure}[t] \caption{Effects on US jobs of eliminating a potential China tariff spike} \begin{minipage}\textwidth {\textit{Notes:} The blue line shows nonparametric event-study estimators of the effect US employment of eliminating potential tariff spikes on Chinese imports, computed following Equation (ref), using a 1997 to 2002 and 2004 to 2005 panel of 103 US industries used by pierce2016surprisingly. The blue line also shows pre-trend estimators, as well as estimators' bias-corrected 95% level confidence intervals, computed following Equation (ref). The red line shows event-study and pre-trend estimators computed using TWFE regressions. The green line shows event-study and pre-trend estimators from TWFE regressions with industry-specific linear trends.} \end{minipage} \end{figure} \paragraph{However, if one only assumes (ref), event-study TWFE estimators with linear trends do not estimate a convex combinations of effects.} We follow (ref),\footnote{This decomposition applies to the regression of $Y_{g,t}-Y_{g,1999}-(t-1999)\times (Y_{g,2000}-Y_{g,1999})$ on $D_{g,2}$.} and estimate the weights attached to the event-study TWFE coefficients with linear trends, using the \texttt{twowayfeweights} Stata package de2019twowayfeweights. They estimate a weighted sum of the effects of the treatment in the 103 industries, where 62 estimated weights are strictly positive, 41 are strictly negative, and the negative weights sum to -0.32. Then, if one only assumes (ref), TWFE estimators with linear trends are far from estimating a convex combination of effects, and could be biased if effects vary across industries. \paragraph{The Stute test of the homogeneous and linear effect assumption is not rejected, thus lending some support to TWFE estimators with linear trends.} To test if heterogeneous effects could bias the TWFE estimators with linear trends, we run a joint Stute test of the following null: $$E(Y_{g,t}-Y_{g,2000}-(t-2000)\times (Y_{g,2000}-Y_{g,1999})|D_{g,t})=\beta_{0,t}+\beta_{fe,t}D_{g,t} ~\forall t\in \{2001,2002,2004,2005\}$$ The test is not rejected (p-value=0.40). In view of Point 2 of Theorem (ref), as there seems to be a QUG in this application and pre-trends tests of (ref) are not rejected, the fact that the Stute test is not rejected lends some support to the homogeneous and linear effect condition in Assumption (ref), and therefore to TWFE estimators with linear trends. However, with only 103 observations and four years of data, it could also be that the Stute test lacks power to detect heterogeneous effects, so this conclusion remains tentative. \@startsection{section}{2}{0mm}{-1\baselineskip}{1\baselineskip}{\normalfont}{Conclusion} We consider treatment-effect estimation in designs in which no unit is treated initially, and then units simultaneously receive heterogeneous and strictly positive treatment doses. We show that in designs with a quasi-untreated group, under a parallel-trends assumption a weighted average of slopes of units’ potential outcomes is identified by a difference-in-difference estimand using the quasi-untreated group as the control group. We leverage results from the regression-discontinuity-design literature to propose a nonparametric estimator. Then, we propose estimators for designs without a quasi-untreated group. Finally, we propose a test of the homogeneous-effect assumption underlying commonly-used two-way-fixed-effects regressions. We use our results to revisit garrett_tax_2020 and pierce2016surprisingly. In garrett_tax_2020, our nonparametric estimators are close to the results that the authors had obtained with TWFE regressions, thus showing that those results are robust to allowing for heterogeneous effects. In pierce2016surprisingly our nonparametric estimators are too noisy to be informative, but our test of the homogeneous-effect assumption underlying two-way-fixed-effects regressions is not rejected. This suggests that those regressions might be reliable in this application, though it could also be that our test lacks power. \putbib