EconBase
← Back to paper

Dynamic Regression Discontinuity: An Event-Study Approach

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

117,423 characters · 24 sections · 122 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Dynamic Regression Discontinuity: An Event-Study Approach

abstractI propose a novel argument to identify economically interpretable intertemporal treatment effects in dynamic regression discontinuity designs (RDDs). Specifically, I develop a dynamic potential outcomes model and reformulate two assumptions from the difference-in-differences literature---no anticipation and common trends---to attain point identification of cutoff-specific impulse responses. The estimand of each target parameter can be expressed as the sum of two static RDD contrasts, thereby allowing for nonparametric estimation and inference with standard local polynomial methods. I also propose a nonparametric approach to aggregate treatment effects across calendar time and treatment paths, leveraging a limited path independence restriction to reduce the dimensionality of the parameter space. I apply this method to estimate the dynamic effects of school district expenditure authorizations on housing prices in Wisconsin.

\setcounter{page}{0}\thispagestyle{empty} \baselineskip1.47\baselineskip

Introduction

Regression discontinuity designs (RDDs) are commonly used to identify causal parameters when the probability of exposure to a treatment changes discontinuously at a known deterministic threshold. In a seminal contribution, cfr2010 extended the canonical design to multi-period settings in which treatment assignment is sequential and exhibits path dependence. Dynamic RDDs have proven popular in the empirical public finance literature that exploits local referenda to estimate the effect of increased government expenditure on various outcomes (darolia2013, hongzimmer2016, martorelletal2016, abottetal2020, baron2022, rohlinetal2022, baronetal2024, bilaschon2024).

Despite their growing relevance in empirical research, dynamic RDDs are methodologically understudied. This is especially true in relation to the challenge posed by the identification of economically interpretable long-term causal parameters. Because units are repeatedly exposed to treatment assignment and the outcome at any point in time may reflect the contribution of past and future treatment states, disentangling effects associated with sequential RD settings is inherently difficult. In recent work, hsushen2024 addresses this challenge in a potential outcomes framework that allows for flexible patterns of path dependence and heterogeneity in treatment effects. To identify the intertemporal effects of treatment exposure in a given period, they impose a selection-on-observables assumption that exploits future realizations of the running variable and, potentially, additional covariates.

In this paper, I consider a dynamic model in which potential outcomes at any point in time depend on past, contemporaneous, and future treatment states (robins1986). I subsequently propose an identification argument that does not incorporate external observables and formalizes a connection between dynamic RDDs and event-study designs---one that has been implicitly recognized in recent empirical work (bilaschon2024). Specifically, I reformulate two assumptions from the difference-in-differences literature---no anticipation and common trends---and show that they are sufficient to point identify interpretable causal parameters. The resulting, time-specific estimands can be expressed as sums of two standard RD outcome contrasts and can be aggregated, for interpretability and exposition, using a set of nonnegative weights that add to one. Finally, recognizing that the number of estimands grows exponentially with the number of time periods in the model, I reduce the dimensionality of the parameter space with a limited path independence assumption. This restriction takes advantage of the sparse and cyclical nature of treatment assignment that is typically true of empirical applications of dynamic RDDs. A more detailed comparison between the approach developed in this paper and that of hsushen2024 is provided in Section (ref).

For nonparametric estimation and statistical inference, I use local polynomial methods that are widely applied in empirical research. Specifically, I estimate time-specific parameters using local linear regression and aggregate them with convex sample weights (as in callawaysantanna2021). In the process, I incorporate standard bias correction techniques to address the distortion that arises when estimating cutoff-specific parameters with observations away from the cutoff. To quantify statistical uncertainty, I follow the method developed by cct2014 and construct nonparametric confidence intervals that account for the noise introduced by first-stage bias correction. Additionally, I derive bandwidth formulas that minimize the estimator's mean squared error, following standard optimality criteria (ik2012).

I implement the proposed method in the most common empirical application of dynamic RDDs: referenda through which school districts and other local jurisdictions authorize additional expenditures and bypass state-imposed fiscal policy constraints. I focus on Wisconsin, where such referenda are routinely held to finance both operational and capital spending (baron2022). The outcome of interest is housing prices in the years following referendum approval, a margin that has long been studied in public finance as a measure of how potential homebuyers value the tradeoff between improved public services and higher property taxes (see rossyinger1999 for a review of the earlier literature).

The approach developed in this paper is similar in spirit to recently proposed solutions to extrapolate effects away from the threshold in static designs only leveraging features of the joint distribution of the outcome, the running variable, and one or more cutoffs (donglewbel2015, bertanhaimbens2020, cattaneoetal2021). I view this approach as offering several advantages. First, by construction, it is applicable to settings in which external information may not be available or fail to satisfy prerequisites for validity, such as observed characteristics co-determined by the treatment of interest. This concern may be especially salient in dynamic settings in which units are repeatedly subject to treatment assignment\footnote{Statisticians and econometricians have long cautioned against adjusting for post-treatment observed confounders. See, for example, rosenbaum1984, pearl1995, rosenbaum2002, wooldridge2005, mostlyharmlsess, rubin2009, wooldridge2010, masteringmetrics, imbensrubin2015, gelman2021regression.}. Second, my proposed approach is based on well-understood identification assumptions with testable suggestive implications. Third, the method is easy to implement because the resulting estimators can be expressed as sums of standard regression discontinuity contrasts.

{\sloppy This paper contributes to three related strands of the econometric literature. First, it proposes a novel approach to point identify intertemporal treatment effects in dynamic RDDs. Second, it adds to the broad literature on the identification of economically interpretable dynamic effects under treatment effect heterogeneity (dcdh2020, goodmanbacon2021, callawaysantanna2021, sunabraham2021, dcdh_intertemp2024, callawaybaconsantanna2024, dcdhvazquez2024) and path dependence in treatment assignment (heckmanhumphriesveramendi2016). Third, it contributes to the growing literature on the causal interpretation of standard estimands in the presence of multiple treatments (dcdh_multi2023, pgphullkolesar2022) or multiple instrumental variables (mogstadtorgowalters2021, magnetorgohole2025) when treatment effects are allowed to be stochastic. }

The remainder of the paper is organized as follows. In Section (ref), I introduce the main identification challenge and discuss previous approaches to address it. In Section (ref), I develop a simple two-period regression discontinuity design and illustrate the core of my identification strategy. In Section (ref), I generalize the model to a setting with multiple time periods. In Section (ref), I propose an approach to aggregate time-specific estimands and reduce the dimensionality of the parameter space. In Section (ref), I illustrate how to use standard local polynomial methods to estimate the target parameters and construct nonparametric confidence intervals around them. In Section (ref), I conduct a Monte Carlo simulation to confirm that my proposed approach correctly recovers the parameters of interest. In Section (ref), I apply my method to estimate the average effects of school district expenditure authorizations on housing prices in Wisconsin. Section (ref) concludes.

Background

In standard regression discontinuity designs, a treatment is assigned based on whether an observed running variable exceeds a known deterministic threshold. For example, lee2008 exploits U.S. House elections to estimate the average effect of a party’s candidate narrowly winning an election on the party’s subsequent victory in future elections for the same seat. Dynamic RDDs extend this framework by distinguishing between current and future treatment states, which depend on whether the running variable crosses the cutoff in each period.

In the context of lee2008, the outcome of a House election in year $t$ determines political representation in both $t$ and $t+1$, allowing for the estimation of short-term treatment effects\footnote{This example assumes that House elections occur at the beginning of a calendar year and that outcomes are observed at the end of it.}. Estimating the average effect of winning an election in year $t$ on outcomes in years $t+2$ and beyond introduces an additional challenge. Because House elections occur every two years, a party that wins a seat in year $t$ may lose it in year $t+2$, and a party that loses in year $t$ may regain control. This dynamic turnover complicates the identification of interpretable long-term causal parameters. The key challenge lies in disentangling the long-term effects of the year-$t$ election outcome from the short-term effects of the year-$t+2$ election outcome. More broadly, the recurrence of elections introduces dynamic fuzziness, as future treatment states may not align with initial treatment assignments. Figure (ref) below provides a stylized representation of the main identification challenge in dynamic RDDs.

figure[figure omitted — 1,656 chars of source]

Since the seminal contribution by cfr2010, dynamic regression discontinuity designs have been employed almost exclusively in the field of local public finance. Numerous empirical studies leverage school district referenda on property tax hikes to estimate the effects of increased school expenditures on outcomes such as housing prices and student test scores. Similar to the stylized example above, dynamic fuzziness arises because school districts hold repeated referenda to authorize additional spending. A ballot measure that is initially rejected may be resubmitted and later approved. As a result, a researcher might observe that a referendum was rejected at time $t$ but approved at time $t+1$. If a study uses the year-$t$ referendum outcome to estimate the long-term effects of extra spending at time $t$, labeling the district as untreated at time $t$ overlooks its transition into a treated state at time $t+1$, potentially conflating the effects of treatment received in successive periods.

Most empirical studies employing a dynamic regression discontinuity design adopt the so called “one-step” approach proposed by cfr2010. This method relies on the following dynamic two-way fixed effects specification:

align[align omitted — 196 chars of source]

where $D$ is a binary treatment indicator, $Q$ is an indicator for whether a referendum is held, and $p_g \left ( \cdot \right )$ represents a $g$th-degree polynomial of the approval vote share $R$. This specification is typically estimated without bandwidth restrictions, allowing researchers to flexibly control for both current and past treatment states, as well as for current and past realizations of the running variable (darolia2013, hongzimmer2016, martorelletal2016, baronetal2024).

More recently, empirical researchers have proposed modifications to the “one-step” approach introduced by cfr2010. abottetal2020 and baron2022 extend the estimating equation in (ref) by including lead treatment indicators, allowing $s$ to take negative values. Similarly, bilaschon2024 incorporates leads but estimates the resulting linear regression in a “stacked” sample, where each component corresponds to a cohort-specific subsample. Here, a cohort is defined as a calendar year in which referenda occur\footnote{bilaschon2024 restricts each cohort to include only treated jurisdictions and untreated jurisdictions that held a referendum in the same year but did not approve any ballot measure in the subsequent ten years.}. The authors interact each regressor in specification (ref) with cohort indicators, effectively estimating the model for each cohort separately. Departing from cfr2010, hsushen2024 develops a dynamic potential outcomes model that accounts for heterogeneity in treatment effects and path dependence in treatment assignments. To identify interpretable long-term effects, the authors impose a selection-on-observables assumption that exploits future approval vote shares and, possibly, additional covariates. More details on their framework are provided in the following section.

A Two-Period Regression Discontinuity Design

In this section, I present the core of my identification strategy within a simple two-period model. Sections (ref) and (ref) extend this framework to accommodate an arbitrary number of time periods.

Model Setup

Consider a dynamic setting in which time periods are indexed by $t \in \left \{ 1, 2 \right \}$. Let $D_t \in \left \{ 0,1 \right \}$ denote a binary treatment that is assigned at the beginning of period $t$. The treatment is determined by whether the realization of an observed running variable $R_t \in \mathbb{R}$ exceeds a known deterministic constant $c \in \mathbb{R}$:

align[align omitted — 63 chars of source]

Let $Y_t \in \mathbb{R}$ be an outcome that is observed at the end of period $t$. In the canonical static setting, potential outcomes depend on a single treatment state, and the observed outcome satisfies the switching equation $Y_t = \sum_{d \in \left \{ 0,1 \right \}} \mathbb{I} \left [ D_t = d \right ] Y_t \left ( d \right )$. In a two-period model with repeated treatment assignment, the potential outcome in period 1 may reflect the treatment state in period 2 due to anticipation and, symmetrically, the potential outcome in period 2 may be affected by the treatment state in period 1 due to path dependence. Thus, following robins1986, it is natural to model potential outcomes as depending on the path of past, contemporaneous, and future treatment states at any point in time. Specifically, let the period-$t$ potential outcome be $Y_t \left ( d_1, d_2 \right )$, where $d_1, d_2 \in \left \{ 0,1 \right \}$. Without further assumptions, the evolution of treatment states is unrestricted. In fact, a unit may be repeatedly assigned the treatment, never receive it, or be exposed to it in only one period. As in hsushen2024, I allow for path dependence in treatment assignment by modeling potential treatments as functions of the history of treatment states. Specifically, let $D_1$ and $D_2 \left ( d_1 \right )$ denote the first- and second-period potential treatments, respectively\footnote{A more flexible setting could allow for anticipation in treatment assignment by letting the potential treatment in period 1 depend on the treatment state in period 2, i.e., $D_1 \left ( d_2 \right )$. For simplicity, I abstract from this scenario and allow for anticipation only in potential outcomes.}. In the canonical application of dynamic RDDs to local referenda, it is useful to interpret potential treatments as characterizing jurisdiction “types”. For example, a school district in which voters repeatedly reject a bond for school construction ($D_1 = 0$ and $D_2 \left ( 0 \right ) = D_2 \left ( 1 \right ) = 0$) is likely more fiscally conservative than one in which increased public expenditure is approved at the first round of elections ($D_1 = 1$) or at the second round ($D_1 = 0$ and $D_2 \left ( 0 \right ) = 1$). In this two-period model, the observed outcome in period $t$ is related to potential outcomes as

align[align omitted — 195 chars of source]

Target Parameters

Having outlined features of the dynamic potential outcomes model, it is natural to define target parameters. As in static RDDs, all causal parameters of interest are local to the nonstochastic threshold above which the treatment is assigned (hahnetal2001). Without loss, the analysis will focus on target parameters implied by the first-period discontinuity. Letting $\tau \in \left \{ 0,1 \right \}$ denote the number of leads, consider a class of threshold-specific Average Treatment Effects (ATEs),

align[align omitted — 181 chars of source]

where $d_2, d_2' \in \left \{ 0,1 \right \}$. Each of these average effects is constructed as the average contrast between a potential outcome that is treated in the first period and a potential outcome that is not. The choice of a second-period treatment state naturally gives rise to causal parameters with a different interpretation.

Since cfr2010, the empirical public finance literature that studies the effects of increased public expenditure authorized by local referenda has focused on a class of Average Direct Treatment Effects that sets $d_2 = d_2' = 0$,

align[align omitted — 166 chars of source]

Hereafter, I will refer to $\mathrm{ADTE}_{1,0} \left ( c \right )$ and $\mathrm{ADTE}_{1,1} \left ( c \right )$ as the instantaneous and cumulative Average Direct Treatment Effects, respectively. These target parameters can be interpreted as tracing out the impulse response of the outcome to the first-period discontinuity. Economically, they correspond to the thought experiment of measuring the cutoff-specific average effect of first-period treatment assignment in the counterfactual scenario in which exposure to the treatment is set to zero thereafter.

What Do Standard RD Estimands Identify?

In what follows, I will show that standard RD estimands do not, in general, identify ADTEs and will subsequently specify assumptions to achieve this goal. First, the (sharp) RD estimand implied by the first-period discontinuity is defined as

align[align omitted — 187 chars of source]

hahnetal2001 shows that continuity in average potential outcomes at the cutoff is sufficient to point identify the Average Treatment Effect at the threshold. It is natural to generalize this assumption to a dynamic setting.

ass[Continuity at the Cutoff] For any $t \in \left \{ 1,2 \right \}$ and $d_1, d_2 \in \left \{ 0,1 \right \}$, \begin{align*} \mathbb{E} \left [ D_{2} \left ( d_1 \right ) | R_1 = r \right ] \quad and \quad \mathbb{E} \left [ Y_{t} \left ( d_1, d_2 \right ) | R_1 = r, D_{2} \left ( d_1 \right ) \right ] \end{align*} are continuous at $r=c$.

Assumption (ref) generalizes the restriction that agents do not sort nonrandomly around the threshold. First, the probability of potential treatment assignment in the second period evolves smoothly at the first-period cutoff. Importantly, this statement does not imply that the probability of observing $D_2 = 1$ is unaffected by the first-period treatment state. Instead, it restricts the probability of any jurisdiction “type” -- as pinned down by the potential treatment $D_2 \left ( d_1 \right )$ -- not to vary abruptly at the first-period cutoff. Analogously, Assumption (ref) rules out a scenario in which agents systematically choose either side of the first-period cutoff based on observed and unobserved determinants of the outcome, within each jurisdiction “type” and for each possible path of treatment states. Under this assumption, it is immediate to derive the following identification argument.

prop[Standard RD Estimand] Suppose that Assumption (ref) holds. Then \begin{align*} \mathrm{RD}_t \left ( c \right ) & = \sum_{d_2 \in \left \{ 0,1 \right \}} \mathbb{E} \left [ Y_t \left ( 1,d_2 \right ) | R_1 = c, D_2 \left ( 1 \right ) = d_2 \right ] \times \mathbb{P} \left ( D_2 \left ( 1 \right ) = d_2 | R_1 = c \right ) \\ & - \sum_{d_2 \in \left \{ 0,1 \right \}} \mathbb{E} \left [ Y_t \left ( 0,d_2 \right ) | R_1 = c, D_2 \left ( 0 \right ) = d_2 \right ] \times \mathbb{P} \left ( D_2 \left ( 0 \right ) = d_2 | R_1 = c \right ) \end{align*}

Without further assumptions, the sharp regression discontinuity estimand identifies the contrast between two weighted sums of potential outcomes over second-period treatment states, where weights are nonnegative and add up to one. As discussed in cfr2010 and similarly in heckmanhumphriesveramendi2016, this target parameter may be interpreted as the “intent-to-treat” implied by the first-period discontinuity. Because treatment assignment exhibits path dependence and outcomes reflect past and future treatment states, the ITT is generally hard to interpret. By way of exposition, consider the extreme scenario in which $\mathbb{P} \left ( D_2 \left ( 1 \right ) = 1 | R_1 = c \right ) = 1$ and $\mathbb{P} \left ( D_2 \left ( 0 \right ) = 0 | R_1 = c \right ) = 1$. In this case, $\mathrm{RD}_{2} \left ( c \right )$ identifies $\mathbb{E} \left [ Y_2 \left ( 1,1 \right ) - Y_2 \left ( 0,0 \right ) | R_1 = c \right ]$. However, the treatment state in period 1 perfectly predicts the treatment state in period 2, making it infeasible to disentangle the cumulative effect of the first-period discontinuity, namely $\mathrm{ADTE}_{1,1} \left ( c \right )$, from the instantaneous effect of the second-period discontinuity. Formally,

align[align omitted — 333 chars of source]

Estimating the “intent-to-treat” from Proposition (ref) may suffice should a researcher be interested in measuring the reduced-form (or “total”) effect of a policy initiated by the first-period discontinuity, regardless of the subsequent path of treatment states. However, the composition of a long-term causal parameter is intrinsically of value in several empirical settings. For instance, does increased funding for local schools lead to sustained and permanent gains in student achievement or the effect of extra expenditure decays over time? To answer this and similar policy-relevant questions, a researcher may prefer to impose additional restrictions and estimate Average Direct Treatment Effects that trace out the impulse response of an outcome to the first-period round of treatment assignment.

Motivated by this negative result, in the next two sections I specify additional sufficient restrictions to identify the instantaneous and cumulative Average Direct Treatment Effects implied by the first-period discontinuity.

Identification of the Instantaneous ADTE

The goal of this section is to provide an identification argument for the instantaneous Average Direct Treatment Effect, $\mathbb{E} \left [ Y_1 \left ( 1,0 \right ) | R_1 = c \right ] - \mathbb{E} \left [ Y_1 \left ( 0,0 \right ) | R_1 = c \right ]$. By an application of the Law of Iterated Expectations, the first term can be expressed as

align[align omitted — 572 chars of source]

Notice that both probabilities are identified under Assumption (ref). In fact, for $d_2 \in \left \{ 0,1 \right \}$,

align[align omitted — 167 chars of source]

Similarly, the first conditional expectation in equation (ref) is identified as

align[align omitted — 179 chars of source]

On the other hand, $\mathbb{E} \left [ Y_1 \left ( 1,0 \right ) | R_1 = c, D_2 \left ( 1 \right ) = 1 \right ]$ is counterfactual. Intuitively, this term captures the average treated-untreated potential outcome among units who would receive the treatment in $t=2$ after receiving it in $t=1$. To identify this counterfactual expectation, it is sufficient to assume that, on average and within bins implied by potential treatment states, the first-period outcome is independent of the second-period treatment state.

ass[No Anticipation] For any $d_1, d_2 \in \left \{ 0,1 \right \}$, \begin{align*} \mathbb{E} \left [ Y_{1} \left ( d_1, 0 \right ) | R_1 = c, D_{2} \left ( d_1 \right ) = d_2 \right ] = \mathbb{E} \left [ Y_{1} \left ( d_1, 1 \right ) | R_1 = c, D_{2} \left ( d_1 \right ) = d_2 \right ] \end{align*}

Assumption (ref) is a cutoff-specific version of the canonical no anticipation assumption required to point identify the Average Treatment Effect on the Treated in difference-in-differences designs (abbringvanderberg2003). It restricts current potential outcomes not to reflect (expectations over) future treatment states. While this statement may appear to rule out the applicability of this argument to forward-looking outcome variables, it should be noted that the lack of anticipatory behavior must hold only at the first-period cutoff. In the application of dynamic RDDs to the effect of increased school spending on student outcomes, Assumption (ref) entails that average test scores after a narrowly rejected bond do not reflect the possibility that a bond proposition may be approved or rejected in a subsequent referendum. Under Assumptions (ref) and (ref), the counterfactual conditional mean is identified as

align[align omitted — 179 chars of source]

Thus, the left-hand side of $\textsc{ADTE}_{1,0} (c)$ can be expressed as a function of the following observables:

align[align omitted — 476 chars of source]

where the last equality follows from an application of the Law of Iterated Expectations. By a symmetric argument, the right-hand side of $\textsc{ADTE}_{1,0} (c)$ is point identified, too. Combining results yields the following proposition.

prop[Identification of the Instantaneous ADTE] \sloppy Suppose that Assumptions (ref) and (ref) hold. Then \begin{align*} ADTE_{1,0} (c) = \lim_{r \downarrow c} \mathbb{E} \left [ Y_1 | R_1 = r \right ] - \lim_{r \uparrow c} \mathbb{E} \left [ Y_1 | R_1 = r \right ] \equiv \mathrm{RD}_1 \left ( c \right ) \end{align*}

This result follows from considering Proposition (ref) in the case in which Assumption (ref) holds. Intuitively, the lack of anticipation implies that the second-period potential treatment state is irrelevant for the first-period outcome. Setting $d_2 = 0$ yields the identification result in Proposition (ref). To summarize, continuity and no anticipation are sufficient for the static regression discontinuity estimand to identify an interpretable causal parameter in this model.

Identification of the Cumulative ADTE

Consider, instead, the more interesting goal of identifying the cumulative Average Direct Treatment Effect, $\mathbb{E} \left [ Y_2 \left ( 1,0 \right ) | R_1 = c \right ] - \mathbb{E} \left [ Y_2 \left ( 0,0 \right ) | R_1 = c \right ]$. Mirroring the argument in the previous section, an application of the Law of Iterated Expectations implies that the first term can be expressed as

align[align omitted — 572 chars of source]

As shown in equation (ref), both conditional probabilities are identified under Assumption (ref). In addition, the first conditional expectation in equation (ref) is identified as

align[align omitted — 179 chars of source]

On the other hand, $\mathbb{E} \left [ Y_2 \left ( 1,0 \right ) | R_1 = c, D_2 \left ( 1 \right ) = 1 \right ]$ is counterfactual because $Y_2 \left ( 1,0 \right )$ is unobserved among units who would sort into the treated arm in $t=2$ after being treated in $t=1$. Differently from the previous section, Assumption (ref) alone is not sufficient to identify the counterfactual mean because the outcome is measured in $t=2$. However, point identification of the target parameter may be achieved by also assuming that the expected growth in untreated-at-2 potential outcomes is independent of the potential second-period treatment state at the cutoff.

ass[Common Trends] For $d_1 \in \left \{ 0,1 \right \}$, \begin{align*} \mathbb{E} \left [ \Delta_{1,2} \left ( d_1, 0 \right ) | R_1 = c, D_{2} \left ( d_1 \right ) = 0 \right ] = \mathbb{E} \left [ \Delta_{1,2} \left ( d_1, 0 \right ) | R_1 = c, D_{2} \left ( d_1 \right ) = 1 \right ] \end{align*} where $\Delta_{1,2} \left ( d_1, 0 \right ) \equiv Y_{2} \left ( d_1,0 \right ) - Y_1 \left ( d_1, 0 \right )$.

Assumption (ref) is a cutoff-specific version of the common trends assumption required to point identify the Average Treatment Effect on the Treated in difference-in-differences designs (abadie2005). As in that setting, groups (“treatment” and “control”) are implied by paths of potential treatment states. Differently from that setting, the treatment is available also in the first period, implying that there are four, not two, groups or “types”. As a consequence, the parallel trends restriction also applies to units who are potentially treated in the first period. In the classic application of dynamic RDDs to local public finance, Assumption (ref) entails that, at the cutoff implied by the first-period referendum, the expected growth in potential test scores absent a second referendum approval is independent of a jurisdiction's “type”, i.e., it is independent of whether a jurisdiction would approve or reject a school bond referendum in the second period. Under Assumptions (ref), (ref), and (ref), the counterfactual conditional mean is identified as

\begingroup

align[align omitted — 264 chars of source]

\endgroup As in the canonical difference-in-differences design, the counterfactual term is identified as the sum of a “level” and a “growth” term. In this case, the former measures the average observed first-period outcome among units who later sort into the treated arm. The latter captures the average observed outcome growth among units who sort into the untreated arm in the second period. Thus, the left-hand side of $\mathrm{ADTE}_{1,1} \left ( c \right )$ can be expressed as a function of the following observables:

align[align omitted — 510 chars of source]

By a symmetric argument, the right-hand side of $\textsc{ADTE}_{1,1} (c)$ is point identified, too. Combining results yields the following proposition.

prop[Identification of the Cumulative ADTE] Suppose that Assumptions (ref), (ref), and (ref) hold. Then \begin{align*} ADTE_{1,1} (c) & = \lim_{r \downarrow c} \mathbb{E} \left [ Y_1 | R_1 = r \right ] - \lim_{r \uparrow c} \mathbb{E} \left [ Y_1 | R_1 = r \right ] \\ & + \lim_{r \downarrow c} \mathbb{E} \left [ Y_2 - Y_1 | R_1 = r, D_2 =0 \right ] - \lim_{r \uparrow c} \mathbb{E} \left [ Y_2 - Y_1 | R_1 = r, D_2 =0 \right ] \end{align*}

This proposition entails that an interpretable long-term causal parameter determined by the first-period discontinuity is identified by the sum of two standard sharp Regression Discontinuity estimands, namely the RD estimand that uses $Y_1$ as an outcome and the RD estimand that uses $Y_2 - Y_1$ as an outcome in the subpopulation implied by $D_2 = 0$.

Relation to hsushen2024

The fundamental difference between this paper and hsushen2024 lies in the assumption used to identify the counterfactual expectation $\mathbb{E} \left [ Y_2 \left ( 1,0 \right ) | R_1 = c, D_2 \left ( 1 \right ) = 1 \right ]$. hsushen2024 introduces a potential indicator $Q_2 \left ( d_1 \right )$, which denotes whether a round of treatment assignment would occur in the second period if the first-period treatment state were $d_1 \in \left \{ 0,1 \right \}$. When $Q_2 \left ( d_1 \right ) = 1$, a corresponding potential second-period running variable $R_2 \left ( d_1 \right )$ is defined. Their key identifying assumption states that, if a second-period treatment assignment occurs, the potential outcome $Y_2 \left ( 1,0 \right )$ is mean independent of $R_2 \left ( d_1 \right )$ in a neighborhood of the first-period cutoff. This condition is further generalized to incorporate covariates $X$, such that, for any $r \in \mathrm{supp} \left ( R_2 \right )$,

align[align omitted — 266 chars of source]

This assumption has the appealing feature of restricting comparisons to units for which treatment assignment occurs in the second period as well. However, it may be vulnerable to violations if unobserved determinants of the second-period outcome also influence the second-period vote share. For instance, in the canonical application examining the effect of school expenditure referenda on student test scores, unobserved parental valuation of education may drive both support for a second-round referendum and student performance—e.g., because more involved parents are more likely to vote in favor of school funding (poterba1997) and also assist their children academically (guryanhurstkearney2008). Notably, such a violation can occur even if the unobservable is time-invariant and the first-round approval margin is idiosyncratically close to the cutoff.

Another distinction between the identification strategy developed in this paper and that in hsushen2024 pertains to the role of anticipation in the potential outcomes framework. The dynamic model presented here allows potential outcomes at any given time to depend on past, contemporaneous, and future treatment states. In a setting with two periods, this implies that potential outcomes in the first period are denoted as $Y_1 \left ( d_1, d_2 \right )$, while those in the second period are given by $Y_2 \left ( d_1, d_2 \right )$. By contrast, the setup considered by hsushen2024 restricts the sequence of treatment states to end in the period in which the outcome is realized, with first-period outcomes modeled as $Y_1 \left ( d_1 \right )$ and second-period outcomes as $Y_2 \left ( d_1, d_2 \right )$. In empirical applications where the outcome of interest may respond to anticipated future treatments---such as settings involving housing prices or other forward-looking variables commonly analyzed in public finance (e.g., cfr2010 and bilaschon2024)---it may be useful to explicitly state the role of anticipation in the underlying model.

Finally, the approach developed in this paper is related in spirit to event-study-type specifications commonly employed by empirical researchers in dynamic RDD applications. The inclusion of leads in the estimating equation (ref) by abottetal2020 and baron2022, together with the cohort-by-cohort identification strategy based on “clean controls” proposed by bilaschon2024, suggests that a conceptual link between dynamic RDDs and difference-in-differences designs has already been implicitly recognized in the literature. This paper formalizes that connection.

Generalization to Multiple Leads

It is natural to generalize previous arguments to identify long-term causal parameters in a dynamic potential outcomes model with more than two periods. In this section, I extend the setup and identifying assumptions of the two-period model to a setting with an arbitrary number of time periods indexed by $t \in \left \{ 1, 2, \dots, \overline{t} \right \}$.

First, potential outcomes at any point in time reflect the path of past, contemporaneous, and future treatment states. Specifically, for $\left \{ d_j \right \}_{j=1}^{\overline{t}}$ with $d_j \in \left \{ 0,1 \right \}$ for all $j$, the period-$t$ potential outcome is denoted as $Y_t \left ( d_1, \dots, d_{\overline{t}} \right )$. Second, to allow for path dependence in treatment assignment, the period-$t$ potential treatment is modeled as $D_t \left ( d_1, \dots, d_{t-1} \right )$. As a generalization of equation (ref), the relationship between potential and observed outcomes can thus be characterized as

align[align omitted — 345 chars of source]

To keep notation concise, let $P: \left \{ 0,1 \right \}^{t} \to \left \{ 0,1 \right \}$ be a function mapping a path of potential treatment states to an indicator that takes the value one if a unit's treatment-taking behavior for a given $d_1 \in \left \{ 0,1 \right \}$ is described by that path, i.e.,

align[align omitted — 229 chars of source]

Therefore, a compact formulation of equation (ref) is

align[align omitted — 270 chars of source]

The target parameter of interest is a class of Average Direct Treatment Effects implied by the first-period discontinuity. Letting $0_{\overline{t}-1}$ denote a vector of $\overline{t}-1$ zeros, the $\tau$-period-ahead ADTE at the first-period cutoff is defined as

align[align omitted — 225 chars of source]

It is convenient to interpret this class of target parameters as tracing out the impulse response function of the outcome $Y$ in a dynamic model in which the initial “shock” is determined by the threshold-crossing rule $D_1 \equiv \mathbb{I} \left [ R_1 \geq c \right ]$. Because treatment assignment exhibits path dependence, the sequence of states $0_{\overline{t}-1}$ may not be observed. As a matter of fact, it is counterfactual in most empirical applications. For instance, local jurisdictions in which voters reject a referendum for increased public spending may submit the same proposition again and have it approved in a subsequent round of elections. In this scenario, a standard regression discontinuity estimand will fail to identify an ADTE or a causal parameter with a clear economic interpretation. To formalize this argument, the following assumption generalizes the continuity restriction naturally embedded in the standard RDD model.

ass[Continuity at the Cutoff] For any $t \in \left \{ 1, \dots, \overline{t} \right \}$, any $\left \{ d_j \right \}_{j=1}^{\overline{t}}$ with $d_j \in \left \{ 0,1 \right \}$ for all $j$, and any $\left \{ d'_j \right \}_{j=2}^{\overline{t}}$ with $d'_j \in \left \{ 0,1 \right \}$ for all $j$, \begin{align*} \mathbb{E} \left [ P \left ( d_1, \dots, d_{\overline{t}} \right ) | R_1 = r \right ] \quad and \quad \mathbb{E} \left [ Y_{t} \left ( d_1, d_2, \dots, d_{\overline{t}} \right ) | R_1 = r, P \left ( d_1, d'_2, \dots, d'_{\overline{t}} \right ) \right ] \end{align*} are continuous at $r=c$.

Assumption (ref) generalizes Assumption (ref) to a model with an arbitrary number of time periods. Intuitively, restricting $\mathbb{E} \left [ P \left ( d_1, \dots, d_{\overline{t}} \right ) | R_1 = r \right ]$ to be continuous at the first-period cutoff entails that agents must not systematically sort around the threshold to make one treatment path more likely than another. Similarly, the continuity of $\mathbb{E} \big [ Y_{t} \left ( d_1, d_2, \dots, d_{\overline{t}} \right ) | R_1 = r, P \left ( d_1, d'_2, \dots, d'_{\overline{t}} \right ) \big ]$ implies that observed and unobserved determinants of the outcome evolve smoothly at the cutoff, within each possible “type” defined by a path of potential treatment states. Under Assumption (ref) and without further restrictions, it is immediate to show that a naive regression discontinuity estimand identifies a causal parameter with an undesirable interpretation.

prop[Standard RD Estimand] Suppose that Assumption (ref) holds. Then the estimand $\mathrm{RD}_t \left ( c \right ) \equiv \lim_{r \downarrow c} \mathbb{E} \left [ Y_t | R_1 = r \right ] - \lim_{r \uparrow c} \mathbb{E} \left [ Y_t | R_1 = r \right ]$ identifies \begin{small} \begin{align*} & \sum_{\left ( d_2, \dots, d_{\overline{t}} \right ) \in \left \{ 0,1 \right \}^{\overline{t}-1}} \mathbb{E} \left [ Y_t \left ( 1,d_2, \dots, d_{\overline{t}} \right ) | R_1 = c, P \left ( 1, d_2, \dots, d_{\overline{t}} \right ) = 1 \right ] \times \mathbb{E} \left [ P \left ( 1, d_2, \dots, d_{\overline{t}} \right ) | R_1 = c \right ] \\ - \ & \sum_{\left ( d_2, \dots, d_{\overline{t}} \right ) \in \left \{ 0,1 \right \}^{\overline{t}-1}} \mathbb{E} \left [ Y_t \left ( 0, d_2, \dots, d_{\overline{t}} \right ) | R_1 = c, P \left ( 0, d_2, \dots, d_{\overline{t}} \right ) = 1 \right ] \times \mathbb{E} \left [ P \left ( 0, d_2, \dots, d_{\overline{t}} \right ) | R_1 = c \right ] \end{align*} \end{small}

Since $P \left ( d_1, \dots, d_{\overline{t}} \right )$ is a Bernoulli random variable, each $\mathbb{E} \left [ P \left ( d_1, \dots, d_{\overline{t}} \right ) | R_1 = c \right ]$ is a conditional probability and $\sum_{\left ( d_2, \dots, d_{\overline{t}} \right ) \in \left \{ 0,1 \right \}^{\overline{t}-1}} \mathbb{E} \left [ P \left ( d_1, d_2, \dots, d_{\overline{t}} \right ) | R_1 = c \right ]$ is equal to one for $d_1 \in \left \{ 0,1 \right \}$. Thus, the causal parameter identified by a standard RD estimand under Assumption (ref) can be viewed as the contrast between two weighted sums of potential outcome means, where only the first-period treatment state is held fixed and each subsequent treatment path is weighted according to its probability at the focal threshold.

In the remainder of this section, I first provide an identification argument for the instantaneous ADTE and then consider long-term parameters. To identify the cutoff-specific effect of the first-period discontinuity on the outcome at the end of the same period, it is sufficient to assume that the outcome of interest does not reflect agents' knowledge of or expectations over future treatment states.

ass[No Anticipation] For any $t < \overline{t}$, any $\left \{ d_j \right \}_{j=1}^{\overline{t}}$ with $d_j \in \left \{ 0,1 \right \}$ for all $j$, any $\left \{ d'_j \right \}_{j=t+1}^{\overline{t}}$ with $d'_j \in \left \{ 0,1 \right \}$ for all $j$, and any $\left \{ d''_j \right \}_{j=t+1}^{\overline{t}}$ with $d''_j \in \left \{ 0,1 \right \}$ for all $j$, \begin{align*} & \mathbb{E} \left [ Y_{t} \left ( d_1, \dots, d_t, d_{t+1}, \dots, d_{\overline{t}} \right ) \big | R_1 = c, P \left ( d_1, \dots, d_t, d”_{t+1}, \dots, d”_{\overline{t}} \right ) \right ] \\ & = \mathbb{E} \left [ Y_{t} \left ( d_1, \dots, d_t, d'_{t+1}, \dots, d'_{\overline{t}} \right ) \big | R_1 = c, P \left ( d_1, \dots, d_t, d”_{t+1}, \dots, d”_{\overline{t}} \right ) \right ] \end{align*}

Starting from Proposition (ref), it is easy to show that that the instantaneous ADTE is point identified under Assumption (ref). In fact, letting $t=1$, the no anticipation restriction entails that the treatment path subsequent to the first period can be equivalently replaced with a sequence of $t-1$ zeros. Then, repeated applications of the Law of Iterated Expectations and Assumption (ref) yield the following proposition.

prop[Identification of the Instantaneous ADTE] \sloppy Suppose that Assumptions (ref) and (ref) hold. Then \begin{align*} ADTE_{1,0} (c) = \lim_{r \downarrow c} \mathbb{E} \left [ Y_1 | R_1 = r \right ] - \lim_{r \uparrow c} \mathbb{E} \left [ Y_1 | R_1 = r \right ] \equiv \mathrm{RD}_1 \left ( c \right ) \end{align*}

This result generalizes Proposition (ref) to a dynamic potential outcomes model with an arbitrary number of periods. In words, the instantaneous regression discontinuity estimand identifies an impulse response parameter provided that the outcome is not forward-looking at the first-period cutoff. However, by placing restrictions only on the conditional expectation of the first-period outcome, the no anticipation assumption is not sufficient to trace out the full impulse response function. To identify long-term Average Direct Treatment Effects, it is sufficient to generalize the common trends assumption introduced in the previous section.

ass[Common Trends] For any $t \in \left \{ 2, \dots, \overline{t} \right \}$ and $\left \{ d_j \right \}_{j=1}^{\overline{t}}$ with $d_j \in \left \{ 0,1 \right \}$, \begin{align*} & \mathbb{E} \left [ \Delta_{1,t} \left ( d_1, 0_{\overline{t}-1} \right ) | R_1 = c, P \left ( d_1, 0_{\overline{t}-1} \right ) \right ] = \mathbb{E} \left [ \Delta_{1,t} \left ( d_1, 0_{\overline{t}-1} \right ) | R_1 = c, P \left ( d_1, d_2, \dots, d_{\overline{t}} \right ) \right ] \end{align*} with $\Delta_{1,t} \left ( d_1, 0_{\overline{t}-1} \right ) \equiv Y_{t} \left ( d_1, 0_{\overline{t}-1} \right ) - Y_1 \left ( d_1, 0_{\overline{t}-1} \right )$.

For any initial potential treatment state $d_1$, $\Delta_{1,t} \left ( d_1, 0_{\overline{t}-1} \right )$ is a random variable that measures the level shift experienced by the never-treated potential outcome between the first and a subsequent period. Assumption (ref) restricts the expectation of this random variable, conditional on $R_1 = c$, not to depend on the potential treatment path after $d_1$. In the canonical application of dynamic RDDs to the effect of local school bonds on student test scores, the path indicator $P \left ( d_1, d_2, \dots, d_{\overline{t}} \right )$ collapses the heterogeneity in jurisdiction “types” into a sequence of potential treatment states. Thus, Assumption (ref) restricts the expected change in test scores after the first period, absent any subsequent referendum approval, not to depend on whether a jurisdiction is more or less likely to approve said referenda. This local “parallel trends” assumption, jointly with continuity and no anticipation, are sufficient to point identify all Average Direct Treatment Effects.

prop[Identification of the Cumulative ADTE] Suppose that Assumptions (ref), (ref), and (ref) hold. Then, for $\tau \in \left \{ 1, \dots, \overline{t}-1 \right \}$, \begin{small} \begin{align*} \mathrm{ADTE}_{1,\tau} \left ( c \right ) & = \lim_{r \downarrow c} \mathbb{E} \left [ Y_{1} | R_1 = r \right ] - \lim_{r \uparrow c} \mathbb{E} \left [ Y_{1} | R_1 = r \right ] \\ & + \lim_{r \downarrow c} \mathbb{E} \left [ Y_{1+\tau} - Y_{1} \bigg | R_1 = r, \bigcap_{s=2}^{1+\tau} \left \{ D_{s} = 0 \right \} \right ] - \lim_{r \uparrow c} \mathbb{E} \left [ Y_{1+\tau} - Y_{1} \bigg | R_1 = r, \bigcap_{s=2}^{1+\tau} \left \{ D_{s} = 0 \right \} \right ] \end{align*} \end{small}

Proposition (ref) states that continuity, no anticipation, and common trends are sufficient for any long-term Average Direct Treatment Effect to be point identified as the sum of two regression discontinuity estimands implied by the first-period cutoff. The first estimand measures the average outcome contrast at the cutoff immediately after the first round of treatment assignment. On the other hand, the second term compares the average outcome growth on the two sides of the cutoff among agents who are not exposed to the treatment forever after.

Aggregation of Time-Specific Parameters

The identification arguments presented so far have focused on causal parameters implied by the first-period discontinuity. In practice, the focal round of treatment assignment may not occur at $t=1$. Moreover, agents may be exposed to multiple rounds of treatment assignment at different points in time.

To allow for the focal discontinuity to occur in an arbitrary period $t$, it is convenient to introduce a random variable $H_{t-1} \in \left \{ 0,1 \right \}^{t-1}$ that summarizes the history of treatment states up to and excluding period $t$. Without further restrictions, past treatment states may affect the probability distribution of potential outcomes when the treatment is assigned. Thus, for a causal parameter to be interpretable, it must contrast agents who have been exposed to an identical path of treatment states. Letting $t$ denote the focal discontinuity period, the Average Direct Treatment Effect at the period-$t$ cutoff subsequent to the treatment history $h_{t-1} \in \mathrm{supp} \left ( H_{t-1} \right )$ is defined as

align[align omitted — 248 chars of source]

for any $\tau \in \left \{ 0, 1, \dots, \overline{t} - t \right \}$\footnote{It is not necessary for the conditioning set to include the event $H_{t-1} = h_{t-1}$. However, an empirical researcher is unlikely to be interested in the average effect of a treatment subsequent to a treatment history $h_{t-1}$ in the subpopulation implied by a different treatment history $h'_{t-1}$. For practical purposes, it is natural for one treatment history to both be an argument of potential outcomes and belong to the conditioning set.}. Importantly, potential outcomes differ solely by the period-$t$ treatment state, implying that this class of impulse response parameters generalizes equation (ref).

Because each $\mathrm{ADTE}_{t,\tau} \left ( h_{t-1}, c \right )$ is specific to a history of treatment states, focal period of treatment assignment, and number of leads, it is natural to think of $\mathrm{ADTE}_{t,\tau} \left ( h_{t-1}, c \right )$ as a building block to construct aggregate policy-relevant parameters. This interpretation is similar in spirit to the one put forward by callawaysantanna2021, which proposes several approaches to aggregate cohort- and time-specific causal parameters in the context of difference-in-differences designs with staggered adoption of an absorbing treatment. Analogously, bojinovetal2021 aggregates unit-time-specific causal parameters in sequentially randomized experiments. In dynamic RDDs, a natural first step is to aggregate target parameters over histories of treatment states, for given focal period of treatment assignment and number of leads. The resulting parameter is defined as

align[align omitted — 224 chars of source]

$\mathrm{ADTE}_{t,\tau} \left ( c \right )$ integrates Average Direct Treatment Effects over the probability distribution of histories up to the focal round of treatment assignment. It can thus be interpreted as the average impulse response of the outcome to the period-$t$ discontinuity, $\tau$ periods after the initial shock.

In addition, an empirical researcher may be interested in aggregating time-specific parameters into a single summary measure of the $\tau$-period-ahead effect of the treatment. In this case, assuming a balanced panel, each term of the aggregate impulse response function is defined as

align[align omitted — 184 chars of source]

Dimension Reduction via Limited Path Independence

Assumptions (ref), (ref), and (ref) are easily generalized to achieve the point identification of each possible $\mathrm{ADTE}_{t,\tau} \left ( h_{t-1}, c \right )$. In observational settings, however, it is generally infeasible to estimate one target parameter for each possible treatment history, focal calendar time, and number of leads. For instance, in the canonical application of dynamic RDDs to the effects of local referenda, units are infrequently subject to treatment assignment because jurisdictions typically hold new elections only once funds approved by a previously approved proposition are close to expiration. As a consequence, only a limited number of local governments participates in treatment assignment every year and any given jurisdiction does so in a cyclical fashion. In addition, a narrowly rejected referendum is generally submitted to voters again only once or twice briefly after the initial attempt, implying that the identification challenge posed by repeated treatment assignment is practically salient only in a narrow time window. To summarize, in the local public finance application of dynamic RDDs, treatment assignment is sparse and cyclical.

Motivated by this fact, it is natural to impose additional restrictions to limit the number of target parameters at stake and reduce the dimensionality of the problem. In a similar fashion to hsushen2024, I assume that all causal parameters of interest share a weak form of limited path independence. Specifically, for any given focal round of treatment assignment, I restrict the Average Direct Treatment Effect not to depend on the history of treatment states provided that units are not exposed to the treatment for a specific number of periods prior to the focal round of treatment assignment.

ass[Limited Path Independence] Let $k \in \mathbb{N}_{+}$ be a deterministic constant. For any $h_{t-k-1}, h'_{t-k-1} \in \mathrm{supp} \left ( H_{t-k-1} \right )$ such that $h_{t-1} = \left ( h_{t-k-1},0_k \right )$ and $h'_{t-1} = \left ( h'_{t-k-1},0_k \right )$, \begin{align*} & \mathbb{E} \left [ Y_{t} \left ( h_{t-1},1,0_{\overline{t}-t} \right ) - Y_{t} \left ( h_{t-1},0,0_{\overline{t}-t} \right ) | H_{t-1} = h_{t-1}, R_t = c \right ] \\ = \ & \mathbb{E} \left [ Y_{t} \left ( h'_{t-1},1,0_{\overline{t}-t} \right ) - Y_{t} \left ( h'_{t-1},0,0_{\overline{t}-t} \right ) | H_{t-1} = h'_{t-1}, R_t = c \right ] \end{align*} for all $t > k$. Thus, $\mathrm{ADTE}_{t,\tau} \left ( h_{t-1}, c \right ) = \mathrm{ADTE}_{t,\tau} \left ( h'_{t-1}, c \right )$ for all $\tau \in \left \{ 0,1, \dots, \overline{t} - t \right \}$.

The constant $k$ is to be determined by the researcher depending on the empirical application\footnote{In a similar spirit, callawaysantanna2021 considers a limited treatment anticipation assumption, based on which the researcher must specify the number of time periods the outcome variable may reflect agents' knowledge of (or expectations over) future exposure to the treatment.}. Clearly, the choice of a larger $k$ entails a more stringent requirement for the limited path independence assumption to be satisfied. On the other hand, a smaller $k$ restricts ADTEs not to be determined in time periods further antecedent to the focal round of treatment assignment.

The practical implication of this Markov-type assumption is that, for the purpose of estimating a target parameter, units that share a common “recent” past can be pooled, thereby improving statistical inference. Another consequence of Assumption (ref) is that the plausibility of the local common trends restriction, namely Assumption (ref), may be tested as in difference-in-differences and event-study designs. Because all units are untreated for $k$ periods before the focal calendar time period, a researcher may test whether the conditional mean of the outcome exhibits pre-trends at the focal cutoff.

For ease of exposition, consider a four-period model ($\overline{t}=4$) in which all units are untreated until $t=3$, when a round of treatment assignment occurs. For $\left ( t,\tau \right ) = \left ( 3,1 \right )$, let the target parameter be

align[align omitted — 190 chars of source]

This is the cutoff-specific, one-period-ahead average effect of the treatment assigned in $t=3$, conditional on a never-treated path in the first two periods. Because, by construction, $D_1 = D_2 = 0$, the history of treatment states $H_2 = 0_2$ is observed. To identify $\mathrm{ADTE}_{3,1} \left ( 0_2, c \right )$, Assumption (ref) restricts $\mathbb{E} \left [ Y_4 \left ( 0_2,d_1,0 \right ) - Y_3 \left ( 0_2,d_1,0 \right ) | H_2 = 0_2, R_3 = c, P \left ( 0_2, d_1, 0 \right ) = 1 \right ]$ to be equal to $\mathbb{E} \left [ Y_4 \left ( 0_2,d_1,0 \right ) - Y_3 \left ( 0_2,d_1,0 \right ) | H_2 = 0_2, R_3 = c, P \left ( 0_2, d_1, 1 \right ) = 1 \right ]$ for both $d_1 = 0$ and $d_1 = 1$. Leveraging the intuition that the path indicator random variable $P$ describes a unit's treatment-taking behavior and thus their “type”, it is natural to expect that a similar common trends assumption hold in time periods preceding the focal round of treatment assignment. In other words, if a unit's “type” is time-invariant, an empirical researcher may plausibly conjecture that

align[align omitted — 355 chars of source]

for $d_1 \in \left \{ 0,1 \right \}$. If the outcome variable is non-anticipatory and Assumption (ref) holds, these conditional expectations are identified for both the upper and lower limit of the running variable in $t=3$. Specifically,

align[align omitted — 195 chars of source]

and symmetrically for the lower limit.

These equalities naturally give rise to testable restrictions aimed at assessing the plausibility of Assumption (ref). Specifically, an empirical researcher may provide a more compelling argument in favor of the local common trends restriction by failing to reject the equality of average differences in outcome trends at the cutoff in all available “pre-periods”. In the canonical difference-in-differences design, it is common practice to test the equality of pre-period outcome trends across groups implied by whether units are assigned to the “treatment” or “control” arm. In a dynamic RDD, groups are implicitly determined by the treatment path subsequent to the focal time period. Unless more restrictions are placed on the evolution of treatment states, the number of groups or “types” will in general exceed two, thereby increasing the number of testable equalities.

Under Assumption (ref), the focal time period is preceded by $k$ periods for which $D_{t-1} = \dots = D_{t-k} = 0$. Then, based on Assumption (ref), the null hypothesis in equation (ref) may be generalized to any sequence of future treatment states $d_{t+1}, \dots, d_{\overline{t}} \in \left \{ 0,1 \right \}$ and any pair of periods $u$ and $v$ such that $t-k \leq u < v \leq t$. Specifically,

align[align omitted — 431 chars of source]

with $\Delta_{u,v} \left ( h_{t-k-1}, 0_k, d_t, 0_{\overline{t}-t} \right ) \equiv Y_v \left ( h_{t-k-1}, 0_k, d_t, 0_{\overline{t}-t} \right ) - Y_u \left ( h_{t-k-1}, 0_k, d_t, 0_{\overline{t}-t} \right )$. Because units are untreated in the $k$ periods preceding $t$, both conditional expectations are identified using the upper or lower limit of $R_t$:

align[align omitted — 360 chars of source]

and the equality of lower limits follows from a symmetric argument. Appendix (ref) develops statistical theory to conduct tests of hypotheses pertaining to the plausibility of the common trends assumption.

Estimation and Statistical Inference

In this section, I illustrate that standard local polynomial methods are applicable to the estimation of Average Direct Treatment Effects in a dynamic RDD. Hereafter, I do not explicitly condition on the history of treatment states $H_{t-1} = h_{t-1}$ to keep notation concise. Then, by an application of the Law Iterated Expectations and the properties of limits, the target estimand from Proposition (ref) can be equivalently expressed as

align[align omitted — 754 chars of source]

In this expression, the parameter of interest is a nonlinear function of cutoff-specific conditional means, all of which can be nonparametrically estimated in the full sample, rather than in subsamples implied by realizations of $\left \{ D_{t+s} \right \}_{s=1}^{\tau}$. This feature will prove convenient for the purpose of constructing robust nonparametric confidence intervals. To keep notation compact, I also define the following random variables:

align[align omitted — 208 chars of source]

For any random variable $A$, let $\mu_{A} \left ( r \right ) \equiv \mathbb{E} \left [ A | R = r \right ]$ and $\sigma^2_{AA} \left ( r \right ) \equiv \mathbb{V} \left [ A | R = r \right ]$. Furthermore, let $\mu_{A+} \equiv \lim_{r \downarrow c} \mu_{A} \left ( r \right )$ and $\mu_{A-} \equiv \lim_{r \uparrow c} \mu_{A} \left ( r \right )$. Then the parameter of interest can be written as

align[align omitted — 103 chars of source]

Prior to discussing statistical inference on $\theta$, Assumption (ref) outlines a number of necessary conditions pertaining to moments and features of the running variable. These requirements match those listed in Assumptions 1 and 3 of cct2014.

assLet $\delta \in \mathbb{N}$. For $\kappa_0 \in \mathbb{R}_{++}$, the following conditions hold in a neighborhood of the cutoff $\left ( c-\kappa_0, c+\kappa_0 \right )$. \begin{enumerate}[(a)] • The density $f_R \left ( r \right )$ is continuous and bounded away from zero. • $\mathbb{E} \left [ Y^4 | R=r \right ]$ is bounded. • $\mu_{Y} \left ( r \right )$, $\mu_{W} \left ( r \right )$, $\mu_{D} \left ( r \right )$ are $\delta$ times continuously differentiable. • $\sigma^2_{YY} \left ( r \right )$, $\sigma^2_{WW} \left ( r \right )$, $\sigma^2_{DD} \left ( r \right )$ are continuous and bounded away from zero. \end{enumerate}

Consider a random sample $\left \{ \left [ R_{i}, Y_{i}, W_{i}, D_{i} \right ]' \right \}_{i=1}^{n}$. Given a positive bandwidth $h_n > 0$, let $\mathcal{R} \left ( h_n \right ) = \left [ c-h_n, c+h_n \right ]$ be a discontinuity window implied by realizations of the running variable around the first-period cutoff. Moreover, let $\mathcal{R}^{-} \left ( h_n \right ) = \left [ c-h_n, c \right )$ and $\mathcal{R}^{+} \left ( h_n \right ) = \left [ c, c+h_n \right ]$ indicate the left and right discontinuity windows, respectively. Throughout this section, I employ standard local polynomial estimators to approximate the value of several conditional means on either side of the cutoff. Specifically, for any random variable $A$, I consider local linear estimators defined as

align[align omitted — 702 chars of source]

where $k_h \left ( u \right ) \equiv k \left ( \frac{u}{h} \right ) \slash h$, with $k \left ( \cdot \right )$ denoting a kernel function. Mirroring Assumption 2 in cct2014, Assumption (ref) below reports minimal conditions for the validity of $k \left ( \cdot \right )$.

assFor $\kappa \in \mathbb{R}_{++}$, the kernel function $k \left ( \cdot \right ): \left [ 0,\kappa \right ] \to \mathbb{R}$ is bounded and nonnegative, zero outside its support, and positive and continuous on $\left ( 0, \kappa \right )$.

In the notation of equations (ref) and (ref), the subscript $p$ in $\widehat{\mu}^{(\nu)}_{+,p}$ indicates the chosen polynomial degree in the regression specification, while the superscript $\nu$ denotes the order of the derivative of interest. For example, in standard sharp designs, the estimated coefficient of interest is $\widehat{\mu}^{(0)}_{+,1} \left ( h_n \right ) - \widehat{\mu}^{(0)}_{-,1} \left ( h_n \right )$, namely the difference between two intercepts in a local linear regression specification. In this paper's dynamic design, the estimator of interest replaces the conditional means in equation (ref) with their sample counterparts. As a matter of fact, a natural estimator for $\theta$ is

align[align omitted — 385 chars of source]

As in any regression discontinuity design, the estimator in equation (ref) is a function of the width of the discontinuity window $h_n$. The choice of a larger bandwidth may lead to more precise estimates, but also increase their bias relative to the true intercept at the cutoff. Concerns about bias are particularly salient when adopting the prevalent data-driven approach to bandwidth selection, namely choosing one that minimizes the mean squared error of the estimator (ik2012). A natural solution is to first estimate the bias using a pilot bandwidth $b_n$ and then subtract the estimated bias from the point estimate of interest. cct2014 shows how to account for the statistical uncertainty stemming from the first step and construct robust nonparametric confidence intervals for canonical estimands in sharp and fuzzy regression discontinuity and kink designs.

In this section, I leverage the same approach to construct robust nonparametric confidence intervals for the target estimand in equation (ref). The choice of notation and the structure of proofs largely mirror those in cct2014supp. Similarly to standard estimators in fuzzy regression discontinuity and kink designs, the estimator in equation (ref) is nonlinear. Thus, to obtain the bias expression in equation (ref), I proceed as in cct2014 and linearize $\widehat{\theta} \left ( h_n \right )$ with a second-order Taylor expansion. The resulting estimator is defined as $\widetilde{\theta} \left ( h_n \right ) \equiv \widetilde{\eta}_{+} \left ( h_n \right ) - \widetilde{\eta}_{-} \left ( h_n \right )$, with

align[align omitted — 559 chars of source]

Letting $\widehat{\textsc{B}} \left ( h_n, b_n \right )$ denote the estimated bias of $\widetilde{\theta} \left ( h_n \right )$, the bias-corrected estimator of $\theta$ is

align[align omitted — 170 chars of source]

with

align[align omitted — 923 chars of source]

where $\mathcal{B}_{+} \left ( h_n \right )$ and $\mathcal{B}_{-} \left ( h_n \right )$ are observed and asymptotically bounded quantities defined in Appendix (ref). Proposition (ref) below establishes the asymptotic distribution of the bias-corrected estimator $\widehat{\theta}^{\textsc{bc}} \left ( h_n, b_n \right )$.

prop[Asymptotic Distribution] Suppose that Assumptions (ref) and (ref) hold. Let $\mu_{D+} \neq 0$ and $\mu_{D-} \neq 0$. If $n \min \left \{ h^5_n, b^5_n \right \} \max \left \{ h^2_n, b^2_n \right \} \to 0$ and $n \min \left \{ h_n, b_n \right \} \to \infty$, \begin{align*} T^{rbc} \left ( h_n , b_n \right ) \equiv \frac{\widehat{\theta}^{bc} \left ( h_n, b_n \right ) - \theta}{\sqrt{V^{bc} \left ( h_n, b_n \right )}} \stackrel{d}{\to} \mathcal{N} \left ( 0,1 \right ) \end{align*} provided that $h_n \to 0$ and $\kappa b_n < \kappa_0$. The exact expression for $\textsc{V}^{\textsc{bc}} \left ( h_n, b_n \right )$ is provided in Appendix (ref).

Bandwidth Selection

I adopt the solution proposed by cct2014 by computing both an optimal main bandwidth $h$ and an optimal pilot bandwidth $b$. The main bandwidth is used to estimate the parameter of interest, while the pilot bandwidth is employed to estimate its bias.

Following ik2012, I define optimality in terms of minimizing the mean squared error (MSE) of the estimator. Two distinct MSE criteria are considered depending on the bandwidth. First, the main bandwidth minimizes the MSE of the linearized estimator $\widetilde{\theta} \left ( h_n \right )$, which consists of a linear combination of intercepts in local linear regression specifications. Second, the pilot bandwidth minimizes the MSE of an alternative linearized estimator $\widetilde{\theta}_{2,2} \left ( h_n \right )$, in which the estimated intercepts are replaced by second-order derivatives in local quadratic specifications and, analogously, population means are replaced by the corresponding second-order derivatives\footnote{In general, $\widetilde{\theta}_{\nu,p} \left ( h_n \right )$ denotes the linearized estimator involving $\nu$th-order derivatives in local polynomial specifications of order $p$. Thus, the main linearized estimator $\widetilde{\theta} \left ( h_n \right )$ corresponds to $\widetilde{\theta}_{0,1} \left ( h_n \right )$.}. The following proposition provides expression for both MSEs and derives the corresponding optimal bandwidths.

propSuppose Assumptions (ref) and (ref) hold with $\delta \geq 3$. Let $h_n \to 0$ and $n h_n \to \infty$. The mean squared error of the linearized estimator $\widetilde{\theta} \left ( h_n \right )$ is \begin{align*} MSE \left ( \widetilde{\theta} \left ( h_n \right ) \right ) & = h^4_{n} \left ( \dot{B}^2_{h} + o_p \left ( 1 \right ) \right ) + \frac{1}{n h_n} \left ( \dot{V}_{h} + o_p \left ( 1 \right ) \right ) \end{align*} and the MSE-optimal main bandwidth is \begin{align*} h_{mse} = \left ( \frac{\dot{V}_{h}}{4 \dot{B}^2_{h}} \right )^{1 \slash 5} n^{-1 \slash 5} \end{align*} The mean squared error of the linearized estimator $\widetilde{\theta}_{2,2} \left ( h_n \right )$ is \begin{align*} \text{MSE} \left ( \widetilde{\theta}_{2,2} \left ( h_n \right ) \right ) & = h^{2}_{n} \left ( \dot{\textsc{B}}^2_{b} + o_p \left ( 1 \right ) \right ) + \frac{1}{n h_n^{5}} \left ( \dot{\textsc{V}}_{b} + o_p \left ( 1 \right ) \right ) \end{align*} and the MSE-optimal pilot bandwidth is \begin{align*} b_{\textsc{mse}} = \left ( \frac{5 \dot{\textsc{V}}_{b}}{2 \dot{\textsc{B}}^2_{b}} \right )^{1 \slash 7} n^{-1 \slash 7} \end{align*} Exact expressions for $\dot{\textsc{B}}_{h}$, $\dot{\textsc{V}}_{h}$, $\dot{\textsc{B}}_{b}$, and $\dot{\textsc{V}}_{b}$ are provided in Appendix (ref).

It is important to note that the MSE-optimal bandwidths specified in Proposition (ref) are asymptotic and, hence, not directly implementable. In practice, I follow the steps outlined in Section S.2.6 of cct2014supp and employ analogous direct plug-in bandwidth selectors.

Standard Errors

The variance of the bias-corrected estimator $\textsc{V}^{\textsc{bc}} \left ( h_n, b_n \right )$ in Proposition (ref) is asymptotic. To compute standard errors and conduct statistical inference on the parameter of interest, it is natural to replace unknown population moments with their sample analogues. First, note that the linearized estimator in equations (ref) and (ref) is defined as a linear combination of estimated intercepts, where the constant coefficients are unknown population means. I replace these means with their corresponding estimates computed within each discontinuity half-window. Second, population variances and covariances are unknown. For any pair of random variables $A$ and $B$, define these second-order moments as $\sigma^2_{AB+} \equiv \lim_{r \downarrow c} \mathbb{C} \left [ A,B | R=r \right ]$ and $\sigma^2_{AB-} \equiv \lim_{r \uparrow c} \mathbb{C} \left [ A,B | R=r \right ]$. To estimate these moments, I follow the approach in abadieimbens2006 and cct2014, employing a nearest-neighbor method that computes each sample covariance using a fixed number of observations. Specifically, letting $j^*$ denote this tuning parameter,

align[align omitted — 534 chars of source]

where $\mathcal{J}^{+}_{i}$ and $\mathcal{J}^{-}_{i}$ denote the sets of the $j^*$ closest units to unit $i$ in the right and left discontinuity half-windows, respectively. Once all of the first- and second-order moments are estimated, the resulting estimated variance $\widehat{\textsc{V}}^{\textsc{bc}} \left ( h_n, b_n \right )$ can be used to compute the standard error of $\widehat{\theta}^{\textsc{bc}} \left ( h_n, b_n \right )$ and construct confidence intervals for the target parameter.

Simulation

In this section, I apply the methods developed in the preceding sections to simulated data. I employ a data generating process similar to that presented by bilaschon2023\footnote{Specifically, the authors conduct a Monte Carlo simulation in Appendix B.}, which models both school districts' decisions to hold local referenda and the subsequent evolution of an outcome of interest based on referendum results. The primary objective of this exercise is to demonstrate that the estimator developed in this paper effectively recovers the parameter of interest.

Data Generating Process

Let $i \in \left \{ 1, \dots, n \right \}$ and $t \in \left \{ 1, \dots, \overline{t} \right \}$ index schools districts and calendar years, respectively. Let $Q_{it}$ be a Bernoulli random variable equal to one if district $i$ holds a referendum in year $t$. Denote by $R_{it}$ the approval vote margin and define $D_{it} \equiv \mathbb{I} \left [ R_{it} \geq 0 \right ]$ as the indicator for referendum approval. In each period, a district's decision to hold a referendum is influenced, among other factors, by whether referenda were recently approved or rejected. Specifically,

align[align omitted — 233 chars of source]

where $A^q_i$ is a district-specific intercept, $B^q_t$ is a time-specific intercept, $U^q_{it}$ is an idiosyncratic shock, and $\overline{q}$ is a deterministic threshold that must be exceeded for a referendum to be proposed. If a referendum is held, the approval vote margin is determined according to a similar statistical model:

align[align omitted — 258 chars of source]

where $A^r_i$ and $B^r_t$ are district- and time-specific intercepts, respectively, and $U^r_{it}$ is an idiosyncratic shock. This specification allows the vote margin to flexibly depend on whether similar ballot measures were recently proposed and, if so, whether they were approved or rejected. The outcome of interest is generated according to

align[align omitted — 107 chars of source]

where $V_{it}$ is an idiosyncratic shock and $R_i \equiv \left ( \sum_{g=1}^{\overline{t}} Q_{ig} \right )^{-1} \sum_{g=1}^{\overline{t}} Q_{ig} R_{ig}$ represents the average approval vote margin for district $i$ over the period of interest. Importantly, the stochastic treatment effect $\theta_{igt}$ is determined by

align[align omitted — 109 chars of source]

This formulation allows the effect of referendum approval to vary across districts, calendar times, and cohorts (with a cohort defined as the calendar time at which a referendum is approved). Here, $\theta_{0g}$ is a cohort-specific random variable capturing the permanent effect of a referendum approved in calendar time $g$, and $\theta_{1g}$ is a cohort-specific random variable that measures the incremental effect per unit of time elapsed since approval, meaning that this additional effect is proportional to the relative time $t-g$. Finally, $\phi$ is a deterministic coefficient that quantifies the extent to which the treatment effect varies with the magnitude of the approval vote margin.

Parameterization

The distributional assumptions and parameter values largely mirror those in bilaschon2023. In the referendum equation (ref), the unit intercept is assumed to be $A^q_i \sim \mathcal{N} \left ( 0, \sigma^2_q \right )$, the time intercept is $B^q_t \sim \mathcal{N} \left ( 0, \sigma^2_q \right )$, and the idiosyncratic shock is $U^q_{it} \sim \mathcal{N} \left ( 0, \sigma^2_{uq} \right )$. Similarly, in the vote margin equation (ref), $A^r_i \sim \mathcal{N} \left ( 0, \sigma^2_r \right )$, $B^r_t \sim \mathcal{N} \left ( 0, \sigma^2_r \right )$, and $U^r_{it} \sim \mathcal{N} \left ( 0, \sigma^2_{ur} \right )$. In both cases, draws are independently distributed. In the outcome equation (ref), the idiosyncratic shock is given by $V_{it} \sim \mathcal{N} \left ( 0, \sigma^2_v \right )$. In the treatment effect equation (ref), the permanent effect is modeled as $\theta_{0g} \sim \mathcal{N} \left ( \overline{\theta}_{0} , \sigma^2_0 \right )$ and the additional, time-proportional effect is $\theta_{1g} \sim \mathcal{N} \left ( \overline{\theta}_{1} , \sigma^2_1 \right )$. I simulate data for $n = 30,000$ school districts over a period of $\overline{t} = 30$ years. To fix initial conditions, I set $Q_{i0} \sim \text{Binomial} \left ( 1, 0.1 \right )$ and, conditional on $Q_{i0} = 1$, $R_{i0} \sim \mathcal{N} \left ( \mu_r,\sigma^2_r \right )$. The chosen parameter values are as follows: $\overline{q} = 1.7$; $\sigma^2_q = 0.5$; $\sigma^2_{uq} = 1$; $\sigma^2_r = 0.01$; $\sigma^2_{ur} = 0.02$; $\mu_r = 0.05$; $\sigma^2_v = 0.2$; $\eta = 0$; $\phi = 0.1$; $\overline{\theta}_{0} = 0.4$; $\overline{\theta}_{1} = -0.1$; $\sigma^2_{0} = 0.08^2$; $\sigma^2_{1} = 0.02^2$. In addition, I cap the incremental effect per unit of time at $t-g = 4$, so that $\theta_1 \left ( t - g \right ) = 4 \times \theta_1$ for any $t-g > 4$. Finally, I set $\overline{s} = 3$ and $\left ( \rho^q_1, \rho^q_2, \rho^q_3 \right ) = \left ( -0.9, -0.5, -0.1 \right )$, $\left ( \delta^q_1, \delta^q_2, \delta^q_3 \right ) = \left ( 0.5, 0.3, 0.1 \right )$, $\left ( \rho^r_1, \rho^r_2, \rho^r_3 \right ) = \left ( -0.05, -0.03, -0.01 \right )$, $\left ( \delta^r_1, \delta^r_2, \delta^r_3 \right ) = \left ( 0.05, 0.03, 0.01 \right )$, and $\left ( \gamma_1, \gamma_2, \gamma_3 \right ) = \left ( 0.05, 0.03, 0.01 \right )$.

Target Parameters

By definition, $\theta_{igt}$ represents the unit-specific effect of a referendum approved in period $g$ on the outcome observed in period $t$. Equivalently, letting $\tau = t-g$ denote the number of periods since approval, $\theta_{igt}$ can be written as $\theta_{ig\tau}$. The average effect of referendum approval for cohort $g$ at relative time $\tau$ is then

align[align omitted — 156 chars of source]

where the expectation is taken over the distribution of $R_{ig}$ across units in cohort $g$. As in any regression discontinuity design, the running variable must be evaluated at the cutoff, implying that the target parameter of interest is $\theta_{0g} + \tau \theta_{1g}$. Using the notation introduced earlier, this corresponds to $\mathrm{ADTE}_{g,\tau} \left ( 0 \right )$, i.e., the $\tau$-period-ahead Average Direct Treatment Effect when the focal discontinuity occurs in period $g$. To summarize effects across cohorts, I aggregate the cohort-specific parameters for each relative time $\tau$ following equation (ref):

align[align omitted — 168 chars of source]

where the weight $\omega_{g \tau} \equiv \mathbb{P} \left ( Q_{ig}=1 \right ) \slash \sum_{\ell=1}^{\overline{t}-\tau} \mathbb{P} \left ( Q_{i\ell}=1 \right )$ is the probability that a referendum occurs in period $g$, conditional on a referendum occurring at any time within the window $g \leq \overline{t} - \tau$. This weighting ensures that the aggregate effect appropriately reflects the distribution of referenda over time (callawaysantanna2021).

The remainder of this section presents a simulation exercise that evaluates how accurately the estimated $\mathrm{ADTE}_{\tau} \left ( 0 \right )$ recovers the true average effect at relative time $\tau$, defined as

align[align omitted — 164 chars of source]

where the expectation is taken over the distribution of $\theta_{0g}$ and $\theta_{1g}$ across cohorts.

Results

To reduce the number of estimands, I leverage Assumption (ref) by defining a cohort $g$ as the set of jurisdictions that held a referendum in period $g$ and did not approve a ballot measure in any of the preceding $k=3$ periods, i.e., $Q_{ig} = 1$ and $D_{i,g-1} = D_{i,g-2} = D_{i,g-3} = 0$. Then, for each cohort and for relative times $\tau \in \left \{ 1, 2, 3, 4, 5 \right \}$, I estimate $\mathrm{ADTE}_{g,\tau} \left ( 0 \right )$ using a local linear specification within the discontinuity window implied by the MSE-optimal bandwidth, as detailed in Section (ref). I then aggregate cohort-specific point estimates by replacing the weights in equation (ref) with their sample counterparts in the simulated data. To estimate the instantaneous causal parameter $\mathrm{ADTE}_{g,0} \left ( 0 \right )$ and to conduct a placebo analysis on pre-period outcomes ($\tau \in \left \{ -3, -2, -1 \right \}$), I use standard (i.e., static) sharp RD estimators. Figure (ref) below reports the true DGP-implied treatment effects alongside the corresponding estimates, all of which are averaged over 250 replications.

figure[figure omitted — 1,079 chars of source]

Figure (ref) confirms that the estimated ADTEs accurately capture the dynamic effects of referendum approval on the outcome. As expected, outcomes measured in negative relative time periods remain unaffected by the approval of referenda in period $t=g$. In Section (ref), I suggest that pre-period outcomes can serve as a useful check on the plausibility of the local common trends assumption. For example, for any cohort $g$, consider the first difference $\Delta_{g-3,g-1} \equiv Y_{g-1}- Y_{g-3}$. By construction, this difference reflects untreated potential outcomes. It is therefore natural to test whether, near the cutoff at $\tau = 0$, the difference $\Delta_{g-3,g-1}$ is mean independent of the treatment path following the focal referendum. For instance, one might compare the $D_{g+1} = D_{g+2} = D_{g+3} = D_{g+4} = D_{g+5} = 0$ path with any alternative path involving at least one treated period:

align[align omitted — 303 chars of source]

and similarly for the lower limit. I conduct a test of the null hypothesis implied by the equality in (ref) and its symmetric counterpart for $r \uparrow 0$. Following the procedure described in Appendix (ref), I construct a test statistic for these two equalities and compare its value in each simulated sample to the appropriate quantile of a $\chi^2 \left ( 2 \right )$-distributed random variable. Across 250 replications, the null hypothesis is not rejected at the conventional 0.05 significance level in approximately 92 percent of cases.

Empirical Application: School District Expenditure Authorization and Housing Prices in Wisconsin

In this section, I apply the proposed method to estimate the intertemporal effects of changes in school district expenditures on housing prices in Wisconsin. Due to property tax limits imposed by state governments, local jurisdictions often face constraints in their fiscal policy decisions lincolntaxlimits. To secure additional funding, school districts may hold referenda seeking voter authorization for property tax increases designated to finance specific expenditure initiatives.

Housing prices are chosen as the outcome variable for two primary reasons. First, research in local public finance has long recognized that the extent to which property tax changes are capitalized into housing prices is an informative measure of the efficiency of local public goods provision (bickerdike1902, marshall1948, oates1969, brueckner1982, cushing1984, barrowrouse2004, figliolucas2004, cfr2010, bilaschon2024). Evidence of positive capitalization following an expenditure increase is typically interpreted as an indication that prospective homebuyers value the associated improvement in public service provision more than the change in the tax burden required to fund it. Second, many capital investment projects approved through school referenda unfold over multiple years, and their long-term implications may not be fully internalized by potential buyers at the time of the vote. This feature underscores the importance of identifying interpretable dynamic treatment effects over extended horizons.

Data

I focus my application on the state of Wisconsin, where school districts routinely hold referenda to authorize both operational and capital expenditures (baron2022). The Wisconsin Department of Public Instruction collects and publishes comprehensive data on all such local referenda held in the state since 1990 (wiscreferenda). This dataset includes, among other variables, information on the approval vote share, which I use as the running variable in the dynamic RDD.

To construct the outcome of interest, I follow an approach similar to that of bilaschon2024. Specifically, I rely on a repeat-sales house price index developed by contatlarson2024, which covers all Census tracts located within Core-Based Statistical Areas\footnote{The term “Core-Based Statistical Area” refers collectively to both Metropolitan Statistical Areas and Micropolitan Statistical Areas (censusglossary).} in the United States from 1989 to 2021. The index is normalized to 100 in 1989 for all tracts, allowing for within-tract temporal comparisons but not cross-sectional ones. To allow for level comparisons across school districts, I incorporate data on the average value of owner-occupied single-family homes at the Census tract level, as reported in the U.S. Census Bureau’s 2000 Decennial Census\footnote{The collection of this variable was discontinued beginning with the 2010 Decennial Census.}. For each tract, I compute a calibration factor as the ratio of the 2000 Census home value to the 2000 value of the house price index from contatlarson2024, and apply this factor to the full time series of the index. The resulting measure of housing prices allows for both cross-sectional and intertemporal comparisons. Next, I compute the centroid of each Census tract and assign it to the corresponding elementary, secondary, or unified school district based on the 2010 TIGER/Line shapefiles provided by the U.S. Census Bureau\footnote{I use 2010 tract and school district boundaries because the house price index constructed by contatlarson2024 is based on 2010 Census tracts.} (censustiger). Finally, for each district, I calculate a population-weighted average of housing prices across its constituent Census tracts\footnote{Although I compute population-weighted averages, this choice is not consequential, as Census tracts are designed to contain approximately 4,000 inhabitants (censusglossary).}. This yields the outcome variable used in the dynamic RDD.

The matched sample includes 1,001 referenda\footnote{To construct a balanced panel in event time, I restrict the sample to referenda with outcomes observed in all relative years from –5 to 5, where relative year is defined with respect to the year of the referendum.}, of which 57.6 percent were approved. The average approval vote share margin is 2.05 percentage points, with an average of 9.86 percentage points among approved referenda and –8.57 percentage points among those that were rejected. Figure (ref) displays a histogram of the approval vote share margin. To assess the validity of the design, I test for discontinuities in the density of the running variable at the cutoff using the local polynomial density estimators developed by cattaneojanssonma2020. The null hypothesis of equal densities on either side of the cutoff is not rejected ($p$-value = 0.80), suggesting that manipulation of the running variable around the threshold is unlikely to be a concern in this setting.

figure[figure omitted — 741 chars of source]

Results

As in Section (ref), I leverage Assumption (ref) by defining a cohort as the set of jurisdictions that held a referendum in a given year and did not approve a ballot measure in any of the preceding three years. Given the limited sample size, estimating a separate parameter for each cohort is not feasible. To address this constraint, I partition the sample into groups of adjacent years, ensuring that each group contains a roughly similar number of observations. The results presented in this section are based on a specification with six groups, each comprising approximately 150 observations. The findings are broadly robust to alternative groupings\footnote{Pooling observations across adjacent cohorts implicitly imposes the restriction $\mathrm{ADTE}_{t,\tau} = \mathrm{ADTE}_{t',\tau}$ for all periods $t$ and $t'$ within a group.}. To quantify statistical uncertainty, I compute standard errors using the nearest-neighbor method described in Section (ref), setting the tuning parameter to $j^* = 3$\footnote{This choice corresponds to the default setting in the rdrobust package (rdrobust2014, rdrobust2015, rdrobust2017).}. Cohort-specific standard errors are then aggregated using the Delta method.

table[table omitted — 2,494 chars of source]

Table (ref) presents the results. Panel A reports cohort-specific estimates, while Panel B shows aggregated parameters. The latter indicate that, on average, the approval of school expenditure measures leads to cumulative increases in housing prices of approximately 12 percent after five years. These estimates are consistent with recent findings in the literature (bilaschon2024), although they also highlight substantial heterogeneity across cohorts. In particular, the aggregate effects appear to be primarily driven by the 1999--2001 and 2002-2005 year groups.

figure[figure omitted — 1,873 chars of source]

To assess the validity of the identification strategy, I conduct a placebo analysis by estimating treatment effects on housing prices in periods prior to the referenda. These estimates rely on standard (i.e., static) local linear regression discontinuity estimators. Figure (ref) presents event-study-style plots that show average direct treatment effects for each relative year, based on two specifications: one that aggregates cohort-specific estimates, and another that pools observations across all cohorts. The two approaches yield comparable results in both magnitude and precision. Reassuringly, the estimated effects in pre-treatment periods are close to zero and not statistically significant, suggesting no systematic differences in housing prices between jurisdictions that later approved versus rejected referenda.

table[table omitted — 2,451 chars of source]

As a final validation exercise, I assess the plausibility of the local common trends assumption, as discussed in Section (ref). By construction, outcomes are untreated in the three periods preceding each referendum. If the common trends assumption holds, an empirical researcher may be willing to conjecture that, at the cutoff of the running variable, trends in untreated outcomes prior to the referendum are similar across jurisdictions that are subsequently treated and those that remain untreated after the focal period. To test this suggestive implication, I examine all possible pairs of relative years $(u, v) \in \{ (3,2), (2,1), (3,1) \}$ and nonparametrically estimate the difference in conditional means of $Y_{g-v} - Y_{g-u}$ between two subpopulations: jurisdictions for which $D_{g+1} = D_{g+2} = D_{g+3} = D_{g+4} = D_{g+5} = 0$, and its complement. These differences correspond to contrasts in pre-referendum outcome trends, evaluated at the cutoff. Table (ref) reports the estimated differences for each pair of years, both using aggregated cohort-specific estimates and a specification that pools observations across all cohorts. In all cases, the estimated differences are not statistically distinguishable from zero at the 0.05 significance level. For each pair $(u, v)$, I also conduct an $F$-test of the null hypothesis that the upper and lower limits of the conditional mean differences are jointly equal to zero. The null is not rejected in any case, suggesting that the local common trends assumption might be plausible in this setting.

Conclusion

This paper proposed a novel argument to point identify economically interpretable intertemporal treatment effects in dynamic regression discontinuity designs. I developed a dynamic potential outcomes model and reformulated two assumptions from the difference-in-differences literature---no anticipation and common trends---to point identify impulse response-like treatment effects at the cutoff. The estimand of each target parameter can be expressed as the sum of two static RD outcome contrasts, thereby allowing for nonparametric estimation and inference with standard local polynomial methods. I also proposed a nonparametric approach to aggregate treatment effects across calendar time and treatment paths, relying on a limited path independence restriction to reduce the dimensionality of the parameter space. I applied the proposed method to study the dynamic effects of school district expenditure referenda on housing prices in Wisconsin. The results indicate that, on average, referendum approval leads to a cumulative increase in housing prices of approximately 12 percent within five years.

spacing{1}