Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
103,921 characters · 10 sections · 75 citation commands
Regression Discontinuity Design with Many Thresholds
{
\footnotetext[1]{ Gilbert F. Schaefer Assistant Professor, Department of Economics, University of Notre Dame; Address: 3060 Jenkins Nanovic Halls, Notre Dame, IN 46556. Email: [email removed]. Website: www.nd.edu/$\sim$mbertanh. } }
Keywords: Regression Discontinuity, Multiple Cutoffs, Average Treatment Effect, Peer-effects
JEL Classification: C14, C21, C52, I21.
Applications of regression discontinuity design (RDD) have become increasingly popular in economics since the late 1990s (black1999better, angrist1999, and van2002). One of RDD's main advantages is identification of a local causal effect under minimal functional form assumptions. More recently, with increasing availability of richer data sets, there have been many applications with multiple cutoffs and treatments (for example, black2007, egger2010, delamata2012, pop2013going). Existing one-cutoff RDD methods applied to each individual cutoff produce many local effects that are estimated using only a few observations near each cutoff. Researchers often prefer one takeaway summary effect that is more precisely estimated by pooling all the data. The meaning of a summary effect crucially depends on heterogeneity assumptions and weights imposed on the different local effects.
Applied studies with multiple cutoffs often normalize all cutoffs to zero and use the one-cutoff estimator. This normalization procedure estimates an average of local treatment effects weighted by the relative density of individuals near each of the cutoffs (cattaneo2016keele, Proposition 3). Such an average effect would be a meaningful summary measure only in two cases: (i) local treatment effects are all identical and the weighting scheme does not matter; or (ii) local treatment effects are heterogeneous but the researcher is only interested in the average effect on the individuals near the existing cutoffs. However, researchers are often interested in combining observed data with assumptions weaker than (i) to make inferences on counterfactual scenarios more general than (ii).\footnote{ In a RDD setting with multiple cutoffs and treatments, it is unreasonable to expect that different local treatment effects are always identical. For example, pop2013going find that the impact of going to a better high school on academic achievement is heterogeneous across students with different ability levels. Another example is delamata2012, who finds that the eligibility for Medicaid benefits decreases the probability of having private health insurance more strongly for lower income individuals. Although I allow for heterogeneous effects across cutoffs, counterfactual analysis requires a pooling and a policy invariance assumption (Section (ref)). }
This paper proposes a novel estimation procedure for average treatment effects (ATE). These ATEs are more valuable summary measures than the average effect estimated by the normalization procedure described above for two reasons. First, the researcher explicitly chooses the counterfactual distribution of the ATE, and this distribution may include individuals at or between existing cutoffs. Second, the researcher does not need to assume any specific functional form for the heterogeneity of treatment effects across different cutoffs. As an example of an application, suppose we are interested in estimating the effect of Medicaid benefits on health care utilization. Medicaid eligibility is triggered by income cutoffs that vary across states. Existing one-cutoff RDD methods identify the average effect on individuals with income equal to the income cutoffs. However, most interesting policy questions require the average effect over the entire range of income values in the data.
The framework for RDD with many thresholds is introduced here using a simple example based on the work of pop2013going, PU from now on. Using a wealth of variation of cutoffs from high school assignments in Romania, PU provide rigorous evidence of the impacts of attending a better school on students' academic performance. The economic logic of this application is briefly summarized as follows. A central planner assigns students to high schools based on their scores from a placement test. High schools have limited capacities and are ranked by their qualities. The central planner ranks students by their scores and assigns each of them to the best school available. Each student $i$ submits her score $X_i$ (forcing variable) to the central planner who, based on the entire distribution of scores, determines a minimum test score $c_j$ (cutoff) for admission to each high school $j$. The quality of high school $j$ is denoted $d_j$ (treatment dose).
The RDD assignment is assumed sharp for now. That is, students attend the best high school available to them based on their score and the cutoffs that apply to them. As the test score crosses an admission threshold $c_j$, the quality of the school the student attends changes from $d_{j-1}$ to $d_j$. Local average effects are denoted by $\mathbb{E}[Y_i(d_j)-Y_i(d_{j-1})|X_i=c_j] = \beta(c_j,d_{j-1},d_j)$, where $Y_i(d)$ is the potential academic achievement student $i$ has if attending a high school of quality $d$, and $\beta(c,d,d')$ is the treatment effect function. Heterogeneity of local effects comes from values of cutoffs and treatment doses that change across the different cutoffs. PU give a particularly illustrative application, because it exhibits sufficient variation in cutoff and treatment doses to generate ATEs with substantially greater economic relevance than the typical average based on normalizing all of the cutoffs to zero.
Numerous other examples of RDD with multiple cutoffs and treatments exist in different fields of economics. For instance, egger2010 study the effect of the size of city government councils on municipal expenditures, where council size is determined by population cutoffs. delamata2012 estimates the effects of Medicaid benefits on health care utilization, where Medicaid eligibility is triggered by income cutoffs that vary across states. agarwal2016 and degiorgi2017 look at multiple cutoffs on credit scores, used by banks to make credit decisions. Education economics also provides a variety of applications. angrist1999 and hoxby2000 use class size rules to estimate the impact of class size on student achievement. hoxby2000 utilizes variation in cutoff values from specific school district class size rules. Several researchers exploit different school starting dates to estimate the impact of educational attainment on various outcomes, for example, dobkin2010school, and mccrary2011. duflo2011 analyze school cohorts that are split into low and high-achieving classes based on test scores, where each school has its own cutoff score. garibaldi2012 look at different income cutoffs that determine tuition subsidies to study the impact of tuition payment on the probability of late graduation from university. In short, despite many applications with variation in cutoffs and treatment doses, a lack of theory on how to combine observations from all cutoffs impedes our ability to estimate economically-relevant average effects.
Whether local effects can be combined into an average effect depends on how comparable the researcher believes these effects are. The comparability of local treatment effects essentially depends on the heterogeneity of treatment doses and on the heterogeneity of the treatment effect function $\beta(c,d,d')$. This paper considers two types of assumptions regarding these two aspects of heterogeneity. The first heterogeneity assumption says that treatment doses are credibly quantifiable by some variable $d$. For example, PU find behavioral evidence that average student performance at each school is a good summary measure for school quality. Another example is the case of a single treatment being triggered by varying cutoffs, as when each state has its own income threshold for Medicaid coverage. The second heterogeneity assumption specifies a parametric functional form for $\beta(c,d,d')$ guided by economic theory or a priori knowledge of the researcher. For example, in a class size application like Hoxby's hoxby2000, a functional form based on Lazear's lazear2001 model of achievement can be derived as a function of class size. Another example is given by bajari2017 who present a principal-agent model to study how insurers reimburse hospitals. The marginal reimbursement rate is discontinuous on health expenditures.
This paper proposes a consistent and asymptotically normal estimator for the ATE of a counterfactual distribution of treatment assignments specified by the researcher. A counterfactual policy scenario specifies the distribution of $(c,d,d')$, and the ATE is the integral of $\beta(c,d,d')$ weighted by such a distribution. The ability to predict effects of counterfactual policies depends crucially on assuming that the distribution of potential outcomes $Y_i(d)$ does not depend on the initial schedule of cutoff-dose values. This policy invariance assumption, along with the first heterogeneity assumption, allows the researcher to choose counterfactual distributions with support more general than the discrete set of cutoff-dose values observed in the data.
The estimator proposed in this paper approximates the ATE integral by averaging estimates of $\beta(c,d,d')$ at existing cutoffs using a proper weighting scheme. Under the first heterogeneity assumption with $\beta(c,d,d')$ non-parametric, the proposed ATE estimator is shown to be consistent and asymptotically normal. This result is novel, because estimation of the non-parametric function $\beta(c,d,d')$ is only possible at deterministic points of the domain, and that creates an additional source of bias. Asymptotic normality requires both the number of observations and cutoffs to grow to infinity, and I provide sufficient conditions on their rate of growth. I demonstrate that the minimax rate of ATE estimation in this setting is root-$n$, and that the proposed estimator attains the minimax optimal rate for a specific choice of tuning parameters. This extends the previous literature on minimax optimality of non-parametric estimation of regression functions at a boundary point to estimation of averages of these regression functions.
Many applications of RDD with multiple cutoffs are, in fact, fuzzy rather than sharp. In the high school assignment example, a student may choose to attend a high school other than the school she is originally eligible to attend. Multiple treatments result in multiple compliance behaviors, and one-cutoff identification results do not apply. Building on classic definitions of compliance behaviors (imbens1997), I define compliance groups in terms of changes in treatment eligibility and receipt. “Ever-compliers” are those whose treatment received changes if and only if it changes to the treatment dose for which they become eligible for. I assume that individuals never change into a treatment dose different from the dose of eligibility, a “no-defiance” condition. In the high school example, if the test score of a student currently in school B increases so as to grant her access to school A, no-defiance implies she either chooses to attend school A or stay at school B, and that she is not triggered to attend some other school C.
This paper shows that even local identification in fuzzy RDD with finite multiple treatments is impossible unless the class of treatment effect functions of ever-compliers is restricted to a finite-dimensional class. Important empirical analyses of fuzzy RDD with multiple treatments include those of angrist1999, chen2008vanderklaauw, and hoekstra2009effect; nevertheless, this is the first paper to define compliance and study causal identification in a general framework for multi-cutoff fuzzy RDD. This framework lays out conditions for the interpretation of two-stage least squares (2SLS) estimates in applications of multi-cutoff fuzzy RDD, a common practice in applied work. The second heterogeneity assumption states that the treatment effect function is of a parametric class. This assumption allows for consistent and asymptotically normal estimation of ATEs on ever-compliers. It also results in efficiency gains, because observations are optimally combined across cutoffs to minimize the mean squared error (MSE) of the ATE estimator.
The rapid growth in the number of applications of RDD in economics in the late 1990s was accompanied by substantial theoretical contributions for inference in the one-cutoff case. Identification and estimation in the sharp and fuzzy cases were formalized by hahn2001id. fan1996 and porter2003 demonstrated low-order bias and rate optimality of the local polynomial estimator. Recent theoretical contributions have addressed the optimal bandwidth choice (imbens2012optimal), alternative asymptotic approximations with better finite sample properties (cattaneo2014calonico), quantile treatment effects (frandsen2012), kink treatment effects (dong_jumpkink), and the difficulty of uniform inference (bertanha2019).
The contribution of this paper is more closely related to the study of treatment effect extrapolation of angrist2004, bertanha_imbens, donglewbel2015, angrist2013away, and rokkanen2014jmp. These last two authors use observations on additional covariates. They restrict the relationship between the heterogeneity of treatment effects after conditioning on these covariates to obtain identification away from the cutoff. This paper differs from these other contributions, because the variation of multiple cutoffs and doses identify ATEs over distributions of individuals both between and at cutoffs, without additional covariates.
The remainder of this paper is organized as follows. Section (ref) presents the notation and lays out basic assumptions. Section (ref) describes the ATE estimator for the sharp case and proves asymptotic normality. It is divided into two sub-sections. Section (ref) treats ATEs of discrete counterfactual distributions, which is a straightforward generalization of one-cutoff RDD. Section (ref) is novel; it studies ATEs of continuous counterfactual distributions under the first heterogeneity assumption. Section (ref) analyzes the fuzzy case. Appendix (ref) contains all proofs. Supplemental Appendix (ref) collects auxiliary results to the proofs in Appendix (ref).\footnote{Appendix (ref) is available online at www.nd.edu/$\sim$mbertanh.}
This section sets up the framework for RDD with multiple cutoffs. There are $P$ sub-populations of individuals indexed by $p=1,\ldots, P$. An example of a sub-population may be a town-year in the high school application, or a state in the Medicaid example. Each individual $i$ in sub-population $p$ is fully characterized by a vector of random variables $(X_{i,p},U_{i,p})$ drawn iid across $i$ from each sub-population. The forcing variable $X_{i,p}$ is a scalar score that governs eligibility for treatment, and it lives in a compact interval $ \mathcal{X}=[\underline{ \mathcal{X}},\overline{ \mathcal{X}}]$; $U_{i,p}$ is a vector of unobserved heterogeneity. Individual $(i,p)$ receives a treatment dose $D_{i,p}$ from a set of possible treatments $\mathcal{D}$. The outcome variable $Y_{i,p}$ is determined by a function $\mathbb{Y}$ of the individual characteristics and treatment,
I start with the simpler sharp RDD setting and defer the fuzzy RDD case to Section (ref). In the sharp case, the treatment received by the individual is a deterministic function of the forcing variable. For an individual with forcing variable $X_{i,p}$ close to a cutoff $c$, the treatment dose is $d$ if $X_{i,p}<c$, or $d'$ if $X_{i,p}\geq c$. hahn2001id demonstrate that continuity of the conditional mean of outcomes is sufficient to identify average causal effects for individuals local to the cutoff $c$.
Lemma (ref) generalizes to the case of multiple cutoffs and treatments under the assumption of continuity of $\mathbb{E}[ \mathbb{Y}(X_{i,p},d,U_{i,p}) | X_{i,p}=x]$ as a function of $x$ for every $d\in \mathcal{D}$. Many cutoffs arise because data sets may have many sub-populations with few cutoffs (e.g. Medicaid benefit with one cutoff per state, many states); or few sub-populations with many cutoffs (e.g. Romanian high schools with one town and many schools). The ability to exploit variation in cutoff-dose values relies on the following pooling assumption. \ExecuteMetaData[assumptions.tex]{assu00A}
Assumption (ref) does not restrict average outcomes to be the same across different sub-populations. It is less restrictive than common specifications for pooling data in applied work, for example, time-trends and sub-population fixed effects. The pooling assumption says that individuals with the same forcing variable that undergo the same change in treatment have the same average response across different sub-populations. The rest of the paper builds on Assumption (ref), and it becomes irrelevant to distinguish sub-populations. Thus, I drop the subscript $p$ and focus on the case of one population with multiple cutoffs.
The cutoffs are ordered such that $c_{1}< c_{2} < \ldots < c_{K}$. Sharp RDD means that an individual with forcing variable $X_i$ is deterministically assigned to a treatment dose $D_i=D(X_i)$ according to the following rule:
where $c_{0}=\underline{ \mathcal{X}}$, and $c_{K+1}=\overline{ \mathcal{X}}$. Each cutoff is characterized by three variables: the scalar threshold $c_{j}$; the treatment dose $d_{j-1}$ the individual receives if $c_{j-1}\leq X_i <c_{j}$; and the treatment dose $d_{j}$ the individual receives if $c_{j}\leq X_i < c_{j+1}$. Let $\boldsymbol{\mathrm{c}}_j = (c_{j},d_{j-1},d_{j})$. The schedule of cutoffs and treatment doses is given by the non-random set $\mathcal{C}_K=\left\{ \boldsymbol{\mathrm{c}}_{j} \right\}_{j=1 }^K$. The richness of set $\mathcal{C}_K$ increases as the researcher collects more data.\footnote{ The validity of the RDD depends crucially on exogeneity of cutoffs and no manipulation of the forcing variable $X$ by individuals. See mccrary2008 for a test of forcing variable manipulation. bajari2017 present a modified RDD estimator that is consistent under forcing variable manipulation in a class of structural models.}
The data generating process is summarized as follows. Values for the forcing variable $X_i$ and heterogeneity $U_i$ are drawn iid $i=1,...,n$ from a joint distribution. Given $D(x)$, these $n$ individuals are assigned to different treatment doses $D_i=D(X_i)$. The observed outcome is determined by $Y_i = \mathbb{Y}(X_{i},D_i,U_{i})$. The econometrician observes the schedule of cutoffs and treatment doses $D(x)$ and $(Y_i,X_i,D_i)$ for $i=1,...,n$. Following Rubin's model of potential outcomes, let $Y_i(d) = \mathbb{Y}(X_{i},d,U_{i})$, and assume continuity of $\mathbb{E}[Y_i(d)|X_i=x]$ for every $d\in \mathcal{D}$. A simple extension of Lemma (ref) identifies average effects at every cutoff $\boldsymbol{\mathrm{c}} \in \mathcal{C}_K$,
Data with multiple cutoff-dose values allow the researcher to learn the causal effect of a variety of dose changes applied to individuals at various levels of the forcing variables. This fact opens the possibility of using observed data to estimate the effect of new policy changes. The individual response function $\mathbb{Y}$ may well depend on the initial assignment of treatments $D_i$, and it could potentially change under counterfactual policies. Unless such dependence is restricted, it becomes impossible to use existing data to infer the effect of new policies. The remainder of this paper relies on the following policy-invariance assumption. \ExecuteMetaData[assumptions.tex]{assu00AB}
The methods of this paper leverage RDD variation in cutoff-dose values to make inferences on average effects of policy changes. A policy change is a counterfactual distribution of changes in treatment doses that are randomly applied to individuals, conditional on the forcing variable. An individual $i$ is assigned to a change in treatment dose from $D_i^*$ to $D_i^{**}$, where the distribution of $(D_i^* , D_i^{**})$ is independent of $U_i$ after conditioning on $X_i$. Under Assumption (ref), the average causal effect of such an experiment is
where the last equality uses the definition of $\beta$ in Equation (ref). The average effect $\mu$ equals an average of the $\beta$ function over the counterfactual distribution of $( X_{i}, D_i^{*}, D_i^{**} )$. The inference methods of this paper first identify $\beta$ from RDD with many cutoffs, then identify the average of $\beta$ under a counterfactual distribution pre-specified by the researcher. In a similar setting, cattaneo2016keele study identification under conditions equivalent to Assumptions (ref) and (ref) (respectively, their Assumptions 5a and 5b).
The definition of $\mu$ captures both the direct effect of changing $D$, and the composition effect of a change in the distribution of $D$ conditional on $X$. To investigate the direct effects of $D$, rothe2012 proposes methods for inference on partial policy effects that preserve the distribution of ranks of $(D,X)$ unchanged, thus controlling for composition effects. Although not the focus of this paper, Rothe's methods may be combined with the RDD identification strategy to study partial policy effects.
This section investigates estimation and inference of averages of the non-parametric function $\beta$ under sharp RDD with many cutoffs. First, I treat the case of qualitative treatment doses. This is a straightforward extension of single-cutoff RDDs which identify ATEs of discrete counterfactual distributions, with support contained in $\mathcal{C}_K$. Second, I treat the case of quantitative treatment doses, that is, the first heterogeneity assumption. Substantial variation in cutoff-dose values allows for novel methods that estimate ATEs with support more general than $\mathcal{C}_K$.
Consider applications of RDD where the treatment dose variable has a qualitative nature, and is not credibly summarized by a real-valued metric. For example, hastings2013 study the assignment of students into different degree programs in universities in Chile. There are multiple cutoffs on a test score, but different cutoffs switch students to completely different programs, e.g. physics, engineering, economics, etc. This limits the ability to combine local effects across cutoffs, which restricts ATEs to counterfactual distributions with discrete support contained in $\mathcal{C}_K$. In this section, it is not possible to identify effects of policies that places weight on cutoff-dose combinations $(c,d,d')$ that are not in $\mathcal{C}_K$.
The focus is on discrete counterfactual distributions with probability mass function $\omega^d (\boldsymbol{\mathrm{c}})$ where $\omega_j^d = \omega^d(\boldsymbol{\mathrm{c}}_j)$ for every $j$. For example, in the high school assignment application, a new policy may reallocate students with test scores marginally across the existing cutoffs. The weight $\omega_j^d$ represents the probability mass of students with test score equal to $c_{j}$ that undergo a change in school quality from $d_{j-1}$ to $ d_{j}$ in the reallocation policy.
The parameter of interest is the average effect on these students, which is a weighted average of local effects at the existing cutoffs:
Identification follows from Equation (ref).\footnote{The common practice of normalizing all cutoffs to zero and estimating only one effect produces an estimator consistent for $\mu^d$ with weights $\omega_j^d = f(c_j) / \sum_l f(c_l)$ where $f$ is the probability density function of $X$.} Estimation is conducted in two steps. The first step uses local polynomial regressions (LPR) near each cutoff $c_{j}$ to non-parametrically estimate
The researcher chooses a bandwidth parameter $h_{1j}>0$ for each cutoff, a kernel density function $k(.)$, and the order of the polynomial regression $\rho_1 \in \mathbb{Z}_{+}$. A polynomial in $X$ is fitted on each side of the cutoff, and the estimator $\hat B_{j}$ is the difference between the intercepts of these two polynomial regressions:
where
and $\boldsymbol{\mathrm{b}}=(b_1,\ldots,b_{\rho_1})$. The estimator $\widehat{B}_j$ uses observations with $X_i$ in the estimation window $[c_j - h_{1j}, c_j + h_{1j}]$. The choice of bandwidths may allow the windows to overlap at consecutive cutoffs. However, it must be the case that $c_j + h_{1j}<c_{j+1}$ and $c_j \leq c_{j+1}-h_{j+1}$ for $j=1,\ldots,K-1$. This ensures that $Y_i=Y_i(d_j)$ for $X_i \in [c_j, c_j + h_{1j}]$, and $Y_i=Y_i(d_{j-1})$ for $X_i \in [c_j - h_{1j}, c_j)$.\footnote{This is the first-step estimation procedure for one sub-population with $K$ cutoffs. In many settings, the data have many sub-populations $p =1, \ldots, P$ with one or more cutoffs $j=1, \ldots, K(p)$ in each sub-population. In that case, the researcher first estimates $\widehat{B}_{j,p}$ for every $j$ in each sub-population $p$. Then, Assumption (ref) allows for pooling of $\widehat{B}_{j,p}$ across $p$ in the second step.}
In the second step, the researcher averages out $\hat B_{j}$ to obtain the estimator $\widehat{\mu}^d$:
For the case of one cutoff, hahn2001id and porter2003 derive the asymptotic normal distribution of the LPR estimator $\hat B_{j}$. I build on their arguments to derive the asymptotic distribution of $\widehat{\mu}^d$ under the assumptions listed below.
\ExecuteMetaData[assumptions.tex]{assu00B}
\ExecuteMetaData[assumptions.tex]{assu00C}
\ExecuteMetaData[assumptions.tex]{assu00D}
The variance of $\widehat{\mu}^{d}$ is consistently estimated by
The squared residuals $\widehat{\varepsilon}_i^2$ are computed by a nearest-neighbor matching estimator, as suggested by cattaneo2014calonico (CCT from now on):
and $\ell(i,l)$ is the index of the $l$-th closest $X$ to $X_i$ that lies within the same cutoffs $c_j$ and $c_{j+1}$ that $X_i$ does. CCT's Theorem A3 demonstrates that $\widehat{\mathcal{V}}_n^d / \mathcal{V}_n^d \overset{p}{\to} 1$ in the case of one cutoff, and a straightforward generalization yields the same conclusion for a finite number of cutoffs. If the bandwidth choices are such that the standardized bias term $\left(\mathcal{V}_n^d \right)^{-1/2} \mathcal{B}_n^d$ differs from zero asymptotically, then inference must be done using a bias-corrected estimator. A practical way of doing bias correction is to increase the order of the polynomial from $\rho_1$ to $\rho_1+1$ and compute $\widehat{\mu}^{d'}$ and $\widehat{\mathcal{V}}_n^{d'}$ using the same bandwidth choices as $\widehat{\mu}^{d}$ and $\widehat{\mathcal{V}}_n^{d}$. It follows that $(\widehat{\mathcal{V}}_n^{d'})^{-1/2} (\widehat{\mu}^{d'} - \mu^d )\overset{d}{\to} N(0,1)$.
The multi-cutoff setup of Theorem (ref) allows for choices of bandwidths that produce overlapping estimation windows in finite samples. For example, if $c_1 + h_{11}>c_2-h_{12}$, the estimator $\widehat{B}_2$ uses some of the same observations that the estimator $\widehat{B}_{1}$ does. In theory, a finite number of cutoffs with shrinking bandwidths leads to non-overlapping estimation windows in large samples. As a consequence, the asymptotic variance of $\sqrt{n \overline{h}_1} (\widehat{\mu}^d - \mathcal{B}_n^d - \mu^d) $ may not approximate its finite-sample variance well in case of overlap. Instead, the variance term in (ref) takes into account overlap because its formula is constructed based on the finite-sample variance.
In practice, implementation of $\ha\mu^{d}$ requires the researcher to choose bandwidths $h_{1j}>0$, the polynomial order $\rho_1 \in \mathbb{Z}_+$, and a kernel density function $k(\cdot)$. In the one-cutoff case, common choices in applied work include the edge kernel $k(u)=\mathbb{I}\{|u| \leq 1 \} (1-u)$, local linear regression $\rho_1=1$, and a bandwidth choice that minimizes the mean squared error (MSE) of estimation. Recent work by imbens2012optimal (IK from now on) provides a practical data-driven rule for choosing the bandwidth in the case of one cutoff. With multiple cutoffs, an interesting aspect of the optimal bandwidth problem is the variance reduction from overlapping estimation windows.\footnote{Following Equation (ref), $COV(\widehat{B}_j,\widehat{B}_{j+1}) =COV( \widehat{a}_j^+ -\widehat{a}_j^- ,\widehat{a}_{j+1}^+ - \widehat{a}_{j+1}^-) =COV( \widehat{a}_j^+ ,- \widehat{a}_{j+1}^-)<0$ because $\widehat{a}_j^+$ and $\widehat{a}_{j+1}^-$ use some of the same observations in the case of overlap.} A formal investigation on optimal bandwidths in the multi-cutoff case is deferred to future work.
A simple recommendation to implement Theorem (ref) is to use the IK bandwidth based on local linear regressions with the edge kernel applied to the sub-sample pertaining to each cutoff. These bandwidths produce asymptotic bias, and valid inference must use a bias-corrected estimator and its variance. Use local quadratic regressions ($\rho_1=2)$ with the edge kernel and the same bandwidths as before to compute the consistent bias-corrected estimator $\widehat{\mu}^{d'}$ and its variance $\widehat{\mathcal{V}}_n^{d'}$. calonico2018optimal propose shrinking MSE-optimal bandwidths as a rule of thumb to improve finite sample coverage of confidence intervals. As means of a robustness check, the researcher may shrink the IK bandwidths by multiplying them by $n^{-1/20}$, and examine the resulting confidence intervals (Section 4.1, calonico2018optimal).
The first heterogeneity assumption allows the researcher to identify counterfactual ATEs with support more general than $\mathcal{C}_K$. An empirical application satisfies the first heterogeneity assumption if the treatment dose is credibly quantifiable in a real-valued variable $d$. For example, in the high school assignment of PU, the treatment dose is a quality measure for each school. Possible measures of school quality include the average test score of peers, the average number of teachers, or funding per student. An infinite amount of data gives rise to a countably-infinite set of cutoff-dose values $\mathcal{C}_{\infty}$. In terms of the high school assignment example, a large number of towns and years produce substantial variation in cutoff-dose values. Define $\mathcal{C}$ to be the convex hull of $\mathcal{C}_{\infty}$. If variation in cutoff-dose values is sufficiently rich, then ATEs with counterfactual distributions supported in $\mathcal{C}$ are identified (Lemma (ref)).
I focus on scalar treatment doses $d$ and counterfactual distributions with continuous probability density function $\omega^c (\boldsymbol{\mathrm{c}})$. Minor changes to the setup can accommodate multivariate $d$ and discrete or mixed counterfactual distributions. The ATE is defined as
The researcher may impose further heterogeneity restrictions to reduce the dimension of $\beta(\boldsymbol{\mathrm{c}})$ and increase the set of possible counterfactual distributions. For instance, linear returns to school quality say that $\beta(c,d,d')$ depends on $(c,d'-d)$ instead of $(c,d,d')$. This implies that $\beta(\boldsymbol{\mathrm{c}})=\phi(c)(d'-d)$ for a smooth function $\phi(c)$, and changes the dimension of set $\mathcal{C}_K$. See Figure (ref) in Section (ref) for an empirical illustration. Medicaid coverage is an example of binary treatment that is triggered by various income cutoffs across states. In the case of binary treatment, the treatment effect function depends only on the cutoff value, that is, $\beta(c,d,d')= \phi(c)$. Identification of averages of $\beta(\boldsymbol{\mathrm{c}})$ requires identification of averages of $\phi(c)$, which relies on infinitely many cutoff values that cover a compact interval on the real line. For example, such variation identifies the average effect of giving Medicaid benefits to an entire neighborhood of individuals within the range of income cutoffs seen in the data.\footnote{ For the Medicaid example, delamata2012 has many income cutoffs that differ by state, age, and year. De La Mata's Table I suggests variation between US\$ 21,394 and US\$36,988. Other examples of rich variation in cutoff values include: (i) agarwal2016 who have 714 credit-score cutoffs distributed between 620 and 800 (see their Figure II(E)); and (ii) hastings2013 who have at least 1,100 cutoffs on admission scores varying between 529.15 and 695.84 (refer to their online appendix's Table A.I.I). Although Angrist and Lavy (1999) have few cutoff values, the pattern of their Figure I suggests variation in dose changes across grades and schools. In such cases, non-parametric identification of $\beta$ is possible for a range of dose changes at a few cutoff values. }
The parameter $\mu^c$ is estimated in two steps. The first step is identical to the procedure described in Equations (ref)-(ref). That is, LPRs produce estimates $\widehat{B}_{j}$, $j=1,\ldots,K$. The second step computes a weighted average of the first-step estimates, using specially designed weights $\{ \Delta_j\}_{j=1}^{K}$ that I call “correction weights,”
Unlike the intuition of the discrete case, the correction weight $\Delta_j$ is not necessarily equal or proportional to $\omega^c_j = \omega^c(\boldsymbol{\mathrm{c}}_j)$. An analytical expression for $\Delta_j$ is given below in Equation (ref), and constructed as follows. The correction weight $\Delta_j$ is the contribution of estimate $\widehat{B}_j$ to the integral $\int_{\mathcal{C}} \omega^c(\boldsymbol{\mathrm{c}}) \widehat{\beta}(\boldsymbol{\mathrm{c}}) ~d(\boldsymbol{\mathrm{c}})$, where $\widehat{\beta}(\boldsymbol{\mathrm{c}})$ is a non-parametric estimate of $\beta(\boldsymbol{\mathrm{c}})$. A weighted regression of $\widehat{B}_j$ on polynomial functions of $\boldsymbol{\mathrm{c}}_j$ centered at $\boldsymbol{\mathrm{c}}$ produces the estimate $\widehat{\beta}(\boldsymbol{\mathrm{c}})$. The researcher specifies the order of the polynomials $\rho_2 \in \mathbb{Z}_+$, and a bandwidth $h_2 >0$ that defines an estimation neighborhood around $\boldsymbol{\mathrm{c}} \in \mathcal{C}$. The estimate $\widehat \beta(\boldsymbol{\mathrm{c}})$ is the intercept of the following weighted least squares regression:
The formula for $\Delta_j$ comes from integrating $\omega^c(\boldsymbol{\mathrm{c}}) \widehat{\beta}({\boldsymbol{\mathrm{c}}})$:
where the third equality uses the Cramer rule, and $\boldsymbol{\mathrm{E}}_{\boldsymbol{0} \leftarrow e_j}(\boldsymbol{\mathrm{c}})$ is a $K \times J$ matrix equal to $\boldsymbol{\mathrm{E}}(\boldsymbol{\mathrm{c}})$ except for the first column, which is replaced by the $K \times 1$ vector $e_j$. The vector $e_j$ has one in its $j$-th entry and zero otherwise.
The main contribution of this paper concerns inference on $\mu^c$ where $\beta(\boldsymbol{\mathrm{c}})$ is estimated non-parametrically and then averaged across cutoffs. This is not the first paper to study estimation of averages of non-parametric functions; for example, see newey1994partial. The novelty here is that the non-parametric estimation step only occurs at $K$ fixed boundary points $\boldsymbol{\mathrm{c}}_j$. A necessary condition for consistency of $\hat \mu^{c}$ is an “infill type of asymptotics,” that is, $K$ grows large with the sample size $n$, and $\mathcal{C}_K$ becomes dense in its convex hull $\mathcal{C}$. Assumption (ref) makes the dependence of $K$, $h_{1j}$, $h_2$, and $c_{j}$ on $n$ explicit with a subscript. The main text omits the subscript $n$ whenever possible to simplify notation. \ExecuteMetaData[assumptions.tex]{assu00E}
For large $K$, cutoff-dose values must be uniformly distributed on the domain $\mathcal{C}$ such that $\boldsymbol{\mathrm{E}}(\boldsymbol{\mathrm{c}}/h_2)' \boldsymbol{\mathrm{\Omega}}(\boldsymbol{\mathrm{x}};h_2) \boldsymbol{\mathrm{E}}(\boldsymbol{\mathrm{c}}/h_2)$ is invertible and of magnitude $Kh_2^3$, that is, $K$ times the volume of every $h_2$-neighborhood of $\boldsymbol{\mathrm{c}}$, for every $\boldsymbol{\mathrm{c}}$ in $\mathcal{C}$. These conditions are satisfied in a variety of examples of triangular arrays of points. In Section (ref) of the supplemental appendix, these conditions are verified for one example of a triangular array. Asymptotic normality also relies on additional smoothness conditions on the moments of the data. \ExecuteMetaData[assumptions.tex]{assu00F}
Theorem (ref) states the rate conditions under which the estimator $\widehat{\mu}^c$ has an asymptotic normal distribution. Estimation of the ATE consists of approximating the integral of the treatment effect function by a weighted sum of the values of such function at a finite number of points in its domain. The approximation error converges to zero as the number of points grows large. Function evaluations $B_j$ are estimated by $\widehat{B}_j$. The correction weights guarantee that the integral approximation error converges to zero faster than the estimation error.
A consistent estimator for ${\mathcal{V}}^c_n$ is
where $\widehat{\varepsilon}_i^2$ is computed using Equation (ref). Lemma (ref) in the supplemental appendix's Section (ref) demonstrates that $\widehat{\mathcal{V}}_n^c / \mathcal{V}_n^c \overset{p}{\to} 1$ under the condition that $(K\underline{h}_1)^{-1}=O(1)$. If the bandwidth choices are such that the standardized bias term $\left(\mathcal{V}_n^c \right)^{-1/2} \left(\mathcal{B}_{1n}^c + \mathcal{B}_{2n}^c \right)$ differs from zero asymptotically, then inference must be done using a bias-corrected estimator. A practical way of performing bias correction is to increase the order of the polynomials from $(\rho_1,\rho_2)$ to $(\rho_1+1, \rho_2+1)$, and to compute $\widehat{\mu}^{c'}$ and $\widehat{\mathcal{V}}_n^{c'}$ using the same bandwidth choices as $\widehat{\mu}^{c}$ and $\widehat{\mathcal{V}}_n^{c}$. It follows that $(\widehat{\mathcal{V}}_n^{c'})^{-1/2} (\widehat{\mu}^{c'} - \mu^c )\overset{d}{\to} N(0,1)$.
Convexity of $\mathcal{C}$, along with the asymptotic behavior of the schedule of cutoff-doses (Assumption (ref)), is crucial for the numerical integration error to vanish sufficiently quickly, as required by Theorem (ref). Continuity of $\omega^c(\boldsymbol{\mathrm{c}})$ implies that the boundary of $\mathcal{C}$ has zero probability under the counterfactual distribution. Therefore, the convergence rate of $\widehat{\mu}^c$ is not affected by the value of $\omega^c(\boldsymbol{\mathrm{c}})$ over the boundary of $\mathcal{C}$. In finite samples, local polynomial estimates of $\widehat{\beta}(\boldsymbol{\mathrm{c}})$ may be noisy for values of $\boldsymbol{\mathrm{c}}$ at the boundary of the convex-hull of $\mathcal{C}_K$. Researchers should take that into account when specifying the support of the counterfactual distribution $\omega^c(\boldsymbol{\mathrm{c}})$.
A simple example illustrates the three rate conditions of Theorem (ref). Suppose $h_{1j}=n^{-\lambda_1}$ for all $j$, $h_2=n^{-\lambda_2}$, and $K=n^{\theta}$. The first-step estimation uses local-linear regression ($\rho_1=1$), and the second step, local cubic regression ($\rho_2=3$). The first rate condition says the first-step bandwidths have to converge to zero fast enough to control the asymptotic bias. That is, $\left( K n \overline{h}_1 \right)^{1/2} \overline{h}_1^{\rho_1 + 1} = O(1)$; in terms of the example, this condition becomes $\lambda_1 \geq (1+\theta)/(3+2 \rho_1)$. The second rate condition restricts how fast the number of cutoffs grows with $n$. It cannot grow too fast to ensure having enough observations around the cutoffs for uniform consistency of first-step estimates. The second condition has two parts: (a) ${ K^{1/2} \log n}/{ \left( n \overline{h}_1 \right)^{1/2} } =o(1)$ $\Leftrightarrow$ $\lambda_1 < 1 - \theta$; and (b) $K\overline{h}_1=O(1)$ $\Leftrightarrow$ $\lambda_1 \geq \theta$. The third rate condition limits how slowly $K$ grows, relative to the sample size, to ensure that the integral approximation error vanishes faster than the estimation variance. Part (a) of the third condition says $\left( K n \overline{h}_1 \right)^{1/2} h_2^{ \rho_2+1 } =O(1)$ $\Leftrightarrow$ $\lambda_1 \geq 1+\theta - 2 \lambda_2 (\rho_2 + 1)$; part (b) is $1 / K h_2^{3} =O(1)$ $\Leftrightarrow$ $\lambda_2 \leq \theta/3$.
Figure (ref) illustrates these conditions and the feasible set for bandwidth choices (shaded area).\footnote{ Section (ref) in the supplemental appendix gives an example of a schedule of cutoff-dose values that satisfies Assumption (ref) for feasible choices of $(h_1, h_2)$ in this example. } Panel (a) shows the conditions in terms of $(\lambda_1,\theta)$ assuming $\lambda_2=\theta/3$ so that $h_2=K^{-1/3}$, which satisfies part (b) of the third condition. Panel (b) depicts the same conditions in terms of $(\lambda_1,\lambda_2)$, assuming $\theta=0.4$. The feasible set is well-defined as long as $K$ grows no faster than $\sqrt{n}$, that is, $\theta < 0.5$. In addition, $\rho_2 \geq 3$ because line 3(a) has to be below line 2(a). The maximum rate of convergence of the estimator is $\sqrt{n}$, and it is reached along the dashed line 2(b).
Implementation of Theorem (ref) requires the researcher to choose $\rho_1 \in \mathbb{Z}_+$, $h_{1j}>0 ~\forall j$, $\rho_2 \in \mathbb{Z}_+$, $h_2>0$, and $k(\cdot)$. A theory of optimal choice of these tuning parameters is beyond the goals of this paper. Optimal choice of bandwidths is an interesting topic for future research, because optimality in the multi-cutoff case would account for: (i) the interaction between first and second-stage bandwidths; (ii) the variance reduction from overlapping estimation windows at consecutive cutoffs; and (iii) the recent advances of robust bias-corrected inference and coverage-error optimal bandwidths by calonico2018optimal.
The IK bandwidth formula may produce first-step bandwidths with an incorrect rate of convergence. For example, if $\rho_1=1$, these bandwidths converge to zero at $n^{-0.2}$, which is not fast enough if $\rho_2=3$ (Figure (ref)). A simple way to correct for this is to adjust the bandwidths by multiplying them by $n^{0.2-\lambda_1}$ for $\lambda_1 \geq \theta$, so that their rate becomes $n^{-\lambda_1}$. Conditions 2(a) and (b) imply that $\theta$ is never bigger than $0.5$ regardless of $\rho_1$, $\rho_2$, and $\lambda_2$. Thus, the smallest value for $\lambda_1$ consistent with these restrictions is $0.5$. The same idea applies to the coverage-error optimal bandwidths by calonico2018optimal, which converge to zero at rate $n^{-0.25}$, and need to be adjusted.
In certain cases, the $\beta$ function may depend on less than the three arguments $(c,d,d')$. For example, in the Medicaid application, the treatment is binary and $\beta$ is only a function of $c$. This is a particular case of the theory in this section. The only rate condition that changes is condition 3(b). It becomes $1 / K h_2 =O(1)$, or $\lambda_2 \leq \theta $ in terms of Figure (ref). A non-empty feasible set of bandwidth choices requires $\rho_2 \geq 1$, as opposed to $\rho_2 \geq 3$ in the general case.
A simple recommendation to implement Theorem (ref) is to use the edge kernel, first-step rate-adjusted IK bandwidths for each cutoff, and a second-step bandwidth $h_2$ that minimizes the MSE of estimation. First, use observations pertaining to each cutoff $j$, compute the IK bandwidth $h_{1j}^{ik}$ for sharp RD and local-linear regression; adjust the rate of the bandwidths so that $h_{1j}=h_{1j}^{ik} \times n^{-0.3}$. Second, create a grid of possible values for $h_2$. For each value on the grid, compute $\widehat{\mu}^c(h_2)$ using the edge kernel, the choices of $h_{1j}$ given above, $\rho_1=1$, and $\rho_2=3$ (or $\rho_2 = 1$ in the binary treatment case). Similarly, compute $\widehat{\mu}^{c'}(h_2)$ using the edge kernel, the choices of $h_{1j}$ given above, $\rho_1=2$, and $\rho_2=4$ (or $\rho_2 = 2$ in the binary treatment case). Use Equation (ref) to estimate the variance of $\widehat{\mu}^c(h_2)$ and call it $\widehat{\mathcal{V}}_n^c(h_2)$. Evaluate the approximated MSE of $\widehat{\mu}^{c}(h_2)$ by $(\widehat{\mu}^{c}(h_2) - \widehat{\mu}^{c'}(h_2))^2 + \widehat{\mathcal{V}}_n^c(h_2)$. Choose the bandwidth value on the grid that minimizes the MSE and call it $h_2^*$. The bias-corrected estimate is $\widehat{\mu}^{c'}(h_2^*)$, and its variance estimate is $\widehat{\mathcal{V}}_n^{c'}(h_2^*)$.
It may not be immediately clear that root-$n$ is the fastest estimation rate achievable in a setting where both $K$ and $n$ grow large. The double asymptotic setting is conceptually different from the usual asymptotic setting where only $n \to \infty$ and non-parametric averages are estimable at root-$n$. Estimation rates depend not only on bandwidth choices, but also on how fast $K$ grows, relative to $n$. Similar examples in econometrics include panels with a large number of observations and time periods, and asymptotics with many instruments. The following theorem demonstrates that the minimax optimal rate of estimation of $\mu^c$ is indeed root-$n$, as long as first-step bandwidths converge to zero at $1/K$ rate.
Equation $\ref{theo_srd_kinf_minimax_a}$ shows that no estimator converges faster than $\sqrt{n}$ uniformly over $\mathcal{P}$. Equation $\ref{theo_srd_kinf_minimax_b}$ says the estimator proposed in Theorem (ref) converges at root-$n$ uniformly over $\mathcal{P}$ as long as first-step bandwidths converge to zero at $1/K$ rate. Therefore, root-$n$ is the minimax optimal rate of convergence in the non-parametric estimation of ATE in RDD with many thresholds. Authors have previously analyzed minimax optimality of non-parametric estimators of a regression function at a boundary point, for example, cheng1997 and sun2005. Theorem (ref) is novel because it combines boundary points to estimate averages of non-parametric regression functions.
This section relaxes the sharp assignment mechanism of previous sections and studies the fuzzy RDD case. The analysis focuses on multiple cutoffs, but $K$ is finite as opposed to approaching infinity as in Section (ref). This makes the exercise more tractable, because the number of compliance cases grows super-exponentially with the number of cutoffs. In contrast to the sharp case, non-parametric identification of local effects in the fuzzy case is impossible. As a result, inference methods in this section rely on a second heterogeneity assumption, namely, the treatment effect function is assumed parametric. Section (ref), in the supplemental appendix, provides practical guidelines to compute an MSE-optimal ATE estimator, and demonstrates asymptotic normality.
In the sharp RDD case, all individuals with forcing variable equal to $x$ receive the same treatment $D(x)$ (Equation (ref)). In the fuzzy RDD case, many of these individuals may receive treatments different from $D(x)$. In the high school assignment example, students may choose to go to a school that is not the best school for which they are eligible. For instance, a student may want to attend the same high school as a certain friend or sibling. Another example is given by garibaldi2012. In their study, a schedule of tuition subsidies applies to most students at Bocconi University, but the university reserves the right to grant certain students different subsidies after reassessing their ability to pay.\footnote{The source of fuzziness varies across applications. One example is the case where the assignment of individuals into different treatments is made through a matching mechanism, and the econometrician does not observe all the individual characteristics used in the matching algorithm. This is the reason why the RDD of PU is fuzzy: based on the entire distribution of test scores and preferences, the central planner ranks students by their test scores and assigns each one to her preferred school among schools with vacancies. }
The fuzzy RDD case is modeled in terms of a potential treatment assignment framework. A potential treatment assignment function $\mathcal{U}:\mathcal{X} \to \mathcal{D}$ describes the treatment received for every value of the forcing variable $x \in \mathcal{X}$. For simplicity, these functions are assumed to belong to the following class:
Sharp RDD is the particular case where the individual potential treatment assignment function $\mathcal{U}_i$ is the same for every individual $i$, that is, $\mathcal{U}_i(x)=D(x) ~ \forall i$ with $D(x)$ defined in Equation (ref). In the fuzzy case, $\mathcal{U}_i$ is sampled iid from a distribution of functions with support in $\mathcal{U}^*$. Potential treatment functions $\mathcal{U}_i (x)$ are unobserved, but the treatments received are observed and given by
Using classic definitions of compliance behaviors (imbens1997), three types of compliance groups are defined in terms of changes in treatment eligibility. “Never-changers” are those whose treatment received never changes when eligibility changes. The treatment received by “ever-compliers” or “ever-defiers” changes at least once when eligibility changes. Ever-compliers are those whose treatment received changes if and only if it changes to the treatment dose for which they become eligible. Ever-defiers change to a treatment dose different from the one for which they become eligible. In the case of one cutoff and two treatments, the definition of ever-complier (ever-defier) is equivalent to the classic definition of complier (defier) of imbens2008regression.
The three compliance groups are measurable events that partition the population of individuals with $\mathbf{G}_{nc}$ denoting never-changers, $\mathbf{G}_{ec}$ ever-compliers, and $\mathbf{G}_{ed}$ ever-defiers.\footnote{ These definitions allow for non-monotonic treatment schedules; for example, the average class-size varies non-monotonically across cutoffs on enrollment (angrist1999). Table (ref) in Section (ref) of the supplemental appendix illustrates these definitions of compliance groups using a simple example with 3 treatments and 2 cutoffs. }
where $\emptyset$ denotes empty set.
In the high school assignment case, an example of a never-changer is a student who strongly prefers the high school with the lowest admission cutoff and attends that high school even if she is admitted to better schools. An example of an ever-complier is a student who attends the best school into which she is admitted, or a student who chooses the best school among the nearby schools. Suppose a student has rational preferences and is never indifferent. Assume her choice set is equal to those schools with admission cutoffs that are less than or equal to her test score. Then, such a student is never an ever-defier. In other words, as her test score increases, a new school is added to her choice-set of schools; she either chooses to go to the new school for which she becomes eligible, or she stays at the school which she preferred prior to the increase in her choice-set. Thus, it seems natural to rule out “ever-defiers” in this and other applications.
Never-changers do not produce changes in treatments, so there is no identification on them. For ever-compliers, there are multiple possible changes in treatment at a given cutoff, and ever-compliers may differ in terms of the treatments they comply with. For example, the student who is willing to attend the best school possible complies with all changes in treatment eligibility. On the other hand, the student who is willing to attend the best possible school within a certain distance from home only complies with some of the changes in treatment eligibility. Therefore, besides no-defiance, identification also requires the heterogeneity of ever-compliers to be restricted.
Assumption (ref) generalizes the sufficient conditions for identification on compliers in the one-cutoff case (hahn2001id and dong_alternative). In addition, it restricts the heterogeneity on ever-compliers.
\ExecuteMetaData[assumptions.tex]{assu00H}
A fuzzy assignment produces several different treatment changes at each cutoff, even after ruling out ever-defiers. The researcher only observes one aggregate change in $Y_i$ at each cutoff, but there are several treatment effects on ever-compliers to be identified at that cutoff. Theorem (ref) below shows that identification of these effects is not possible without further restricting the class of functions $\beta_{ec}(\boldsymbol{\mathrm{c}})$. Economic theory or a priori knowledge guides the choice of a functional form that credibly summarizes the heterogeneity of treatment effects. For example, the principal-agent model of bajari2017 yields a functional form to study reimbursement of hospitals by insurers. The second heterogeneity assumption (Assumption (ref)) restricts the treatment effect function on ever-compliers to a finite-dimensional vector space of functions. \ExecuteMetaData[assumptions.tex]{assu00G}
In this case, the ATE on ever-compliers is a linear combination of the true parameter vector $\boldsymbol{\mathrm{\theta}}_0^{ec}$. For a counterfactual distribution $F$ chosen by the researcher,
Theorem (ref) shows that the observed change in average outcome at a given cutoff is a weighted average of treatment effects on ever-compliers who switch from various doses into the dose of eligibility at that cutoff. Assumption (ref) and variation in cutoff characteristics are sufficient conditions for identification. Conversely, identification on ever-compliers implies that $\beta_{ec}(\boldsymbol{\mathrm{c}})$ belongs to a finite-dimensional class of functions.
Theorem (ref) reveals the requirement of stronger functional form assumptions on $\beta_{ec}(\boldsymbol{\mathrm{c}})$ even for identification of local effects in the fuzzy case with a finite number of multiple cutoffs. For example, identification is not possible when $\widetilde{\mathcal{H}}$ is the class of all smooth functions studied in the non-parametric case of Section (ref). The result is striking because non-parametric identification of local effects is possible both in the sharp case with a finite number of cutoffs and in the fuzzy case with a single cutoff. It is likely possible to obtain non-parametric identification of $\beta_{ec}(\boldsymbol{\mathrm{c}})$ under a large variation of cutoff-dose values. The function $\beta_{ec}(\boldsymbol{\mathrm{c}})$ may be approximated by a sequence of parametric functions from Assumption (ref), where $q$ grows to infinity more slowly than $K$, so to keep $dim \mathcal{G} \leq K$ as $K\to \infty$. In this paper, the number of cutoffs is kept finite for simplicity, and the case with large $K$ is deferred to future work.
Theorem (ref) also clarifies the interpretation of two-stage least squares (2SLS) estimates in applications of fuzzy RD with multiple cutoffs, a common practice in applied work. The practice consists of using $D(X_i)$ as an instrument for $D_i$ in the regression of $Y_i$ on a constant, $D_i$, and $X_i$. See angrist2008mostly for a discussion. In the single-cutoff case, both the non-parametric RD estimator and 2SLS applied to a neighborhood of the cutoff are consistent to the average treatment effect on compliers (hahn2001id). To my knowledge, such an equivalence has never been studied in the multiple-cutoff case. Nevertheless, many important applications have multiple-fuzzy cutoffs and use 2SLS; for example, angrist1999, chen2008vanderklaauw, and hoekstra2009effect. The 2SLS estimator is consistent for a data-driven weighted average of treatment effects on ever-compliers as long as a sufficiently flexible specification is used; for example, cutoff fixed-effects or varying slopes. The economic meaning of the 2SLS estimands depends crucially on the choice of such a weighting scheme. Unless a parametric functional form is imposed on $\beta_{ec}(\boldsymbol{\mathrm{c}})$, or there is large variation in cutoff-doses, only a data-driven weighted average of $\beta_{ec}(\boldsymbol{\mathrm{c}})$ is identified. In other words, if $\beta_{ec}(\boldsymbol{\mathrm{c}})$ is non-parametric and there are only a few cutoffs, the researcher does not have control over the weighting scheme, and 2SLS estimates don't have a clear interpretation.
Theorem (ref) leads to a two-step estimation procedure for $\boldsymbol{\mathrm{\theta}}^{ec}_0$ and $\mu^{ec}$. The mechanics are similar to the previous sections, so I omit the details from the main text for brevity. In the first step, the researcher estimates the jump discontinuity of the vector $[Y_i ~~ \boldsymbol{\mathrm{\mathcal{W}}}(X_i,D_i)' ]'$ using LPRs at each cutoff to obtain $[ \widehat{B}_j ~~ \widehat{\widetilde{W}}_j']'$. In the second step, a regression of $\widehat{B}_j$ on $\widehat{\widetilde{\boldsymbol{\mathrm{W}}}}_j'$ obtains $\widehat{\boldsymbol{\mathrm{\theta}}}^{ec}$. The ATE estimator is $\widehat{\mu}^{ec} = \boldsymbol{\mathrm{Z}}(F) \widehat{\boldsymbol{\mathrm{\theta}}}^{ec}$. Estimation precision varies across cutoffs, and the parametric form of $\beta_{ec}$ allows us to optimally combine different cutoffs to minimize the MSE of $\widehat{\boldsymbol{\mathrm{\theta}}}^{ec}$. The researcher can simply re-weight the second-step regression by the inverse of the MSE matrix of the first-step estimators. Section (ref) in the supplemental appendix delineates the estimation and inference procedures of $\boldsymbol{\mathrm{\theta}}^{ec}_0$ and $\mu^{ec}$ with practical steps.
In this section, Monte Carlo simulations illustrate the finite sample behavior of the ATE estimator proposed in Section (ref). The analysis considers estimation precision and coverage of confidence intervals for different choices of tuning parameters and a non-linear specification for $\beta$. As predicted by Theorem (ref), an incorrect choice of the second-step polynomial degree leads to severe bias and extremely poor coverage of confidence intervals. Moreover, first-step bandwidths that imply overlapping estimation windows produce lower MSE than cases with no overlap, regardless of other tuning parameters.
The DGP draws $n$ iid observations of $(X_i,\varepsilon_i)$ where $X_i$ is uniformly distributed over $[0,1]$, $\varepsilon_i$ is normally distributed with zero mean and unit variance, and these variables are independent of each other. There are $K$ cutoffs $c_j = j/(K+1)$, $j=1, \ldots, K$, on the unit interval $[0,1]$. The number of cutoffs is $K=\lfloor n^{0.4} \rfloor$, where $\lfloor a \rfloor$ denotes the largest integer smaller than or equal to $a$. An individual with forcing variable $X_i$ receives a treatment dose equal to $D(X_i)$ as in Equation (ref). The dose increases by one unit at each cutoff, starting at $d_0 = 1$ and ending at $d_K=K+1$. The outcome variable is $Y_i = \phi(X_i) D(X_i) + \varepsilon_i$ where $\phi(X_i) = 15 X_i^3 + 7.5 X_i^2 - 18.75 X_i +2.125$. This implies that $\beta(\boldsymbol{\mathrm{c}})=\phi(c)(d'-d)$, which falls into the binary treatment case (see discussion on page (ref)). Consider a counterfactual policy that uniformly increases treatment doses by one unit. The ATE parameter $\mu$ is the integral of $\phi(c)$ over $c \in [0,1]$, which equals $-1$ in this case.
Estimation follows the procedure suggested in Section (ref). For given bandwidth choices $h_1$ and $h_2$, I compare the ATE estimator $\widehat{\mu}$ that uses $\rho_1=\rho_2=1$, to the bias-corrected ATE estimator $\widehat{\mu}^{bc}$ that uses $\rho_1=\rho_2=2$. To emphasize the importance of the second step, I also compute a naive ATE estimator that simply averages the first-step estimates. The naive and bias-corrected naive estimators, respectively $\widetilde{\mu}$ and $\widetilde{\mu}^{bc}$, are constructed as $\widehat{\mu}$ and $\widehat{\mu}^{bc}$ except for the tuning parameters in the second step. Both naive estimators use $\rho_2=0$ and $h_2=\infty$. To examine the effect of overlapping estimation windows in the first step, I compare estimators for two choices of $h_1$. The first choice is the largest possible bandwidth $h_{1}=1/(K+1)$, which leads to maximum overlap. The second choice is the largest possible bandwidth with no overlap, that is, $h_{1}=0.5/(K+1)$. Finally, I study the effects of ten different choices for the second-step bandwidth, $h_2 \in \{3/(K+1), \ldots, 12/(K+1)\}$. All choices of tuning parameters satisfy the rate conditions of Theorem (ref) and produce a convergence rate of root-$n$ for $\widehat{\mu}$ and $\widehat{\mu}^{bc}$. The Monte Carlo experiment simulates 10,000 draws of an iid sample with $n \in \{ 1789, 10120, 27886, 57244,100000\} $ and $K \in \{20, 40, 60, 80, 100\}$, respectively. Section (ref) in the supplemental appendix repeats the experiment with data-driven bandwidth choices, following the bandwidth rules proposed on page (ref) (Section (ref)).
The bias and variance of all estimators converge to zero as the sample size increases, regardless of the choice of $h_1$ (Table (ref)). The bias-correction of $\widehat{\mu}^{bc}$ eliminates almost all the bias of $\widehat{\mu}$, at the cost of a higher variance. The naive estimator $\widetilde{\mu}$ oversmooths the second step beyond the conditions of Theorem (ref). As a result, the bias of $\widetilde{\mu}$ is substantially larger than that of $\widehat{\mu}$. Simply correcting for bias in the first step does not solve the problem, as the difference in bias between $\widetilde{\mu}$ and $\widetilde{\mu}^{bc}$ is small. First-step bandwidths that produce overlap (Table (ref), rows 1-5) yield approximately the same bias, but substantially smaller variance, compared to first-step bandwidths that produce no overlap (Table (ref), rows 6-10).
Next, I study how the choice of $h_2$ affects precision of $(\widehat{\mu}, \widehat{\mu}^{bc})$ for a fixed choice of $h_1=1/(K+1)$ (Table (ref)). The smallest value for $h_2$ is $3/(K+1)$. This defines a second-step estimation window with at least three cutoffs to ensure invertibility of matrices in the regressions. The bias of $\widehat{\mu}$ is substantially smaller when $h_2$ is set to its smallest value. All other measures are practically unaffected across different $h_2$.
The significant bias of the naive ATE estimators $\widetilde{\mu}$ and $\widetilde{\mu}^{bc}$ decreases the coverage of 95% confidence intervals as the sample increases (Table (ref)). The naive estimators oversmooth in the second step, and Theorem (ref) implies the bias grows faster than root-$n$. For each of the four estimators, the confidence intervals equal the estimator plus or minus $1.96$ times its standard error. The variance of estimators are obtained as described in Equation (ref). The bias-corrected ATE estimator $\widehat{\mu}^{bc}$ produces confidence intervals with correct coverage for all samples sizes. Although $\widehat{\mu}$ yields intervals with average length smaller than $\widehat{\mu}^{bc}$, the bias of $\widehat{\mu}$ leads to a slightly lower coverage.
In this section, the methods proposed in this paper are illustrated using the data from PU on high school assignments in Romania.\footnote{The data set is available online in the supplemental materials of PU on the website of the American Economic Review.} Many policy questions demand an ATE of a continuous counterfactual distribution of treatments, and this section provides an example of such a policy question. The estimators designed for the sharp RDD case are consistent for “Intent-to-Treat” (ITT) average effects when applied to the fuzzy data of PU. In this application, the ITT effect measures the impact of being assigned to a better school but not necessarily attending it. The parametric methods of Section (ref) yield noticeable efficiency gains in the estimation of the ATE on ever-compliers. Treatment effects for ever-compliers reveal a heterogeneity pattern unlike the heterogeneity of ITT effects.
The administrative data from Romania cover 3 cohorts of 9th grade students for the years 2001, 2002, and 2003, with a total of 334,137 observations. The essential elements of the high school assignment in Romania are described below. The assignment to high school is nationally centralized by the Ministry of Education. At the end of grade 8, students submit a transition score and a complete ranking of preferences for high schools. The transition score is an average of the student's performance on a national exam taken in grade 8 and the student's grade point average during grades 5-8. The Ministry of Education ranks students by their transition score and no other criteria. The mechanism assigns the student ranked first to her most preferred school, the student ranked second to her most preferred school, etc. Students cannot decline their assignment, and they have incentives to truthfully reveal their preference rankings.
The observed variables are the town and year of student $i$, the transition score $X_i$, the school the student is assigned to, and the student's score on the “baccalaureate exam.” This is an exam taken at the end of high school, and the grade on the exam is the outcome variable $Y_i$. The quality of school $j$ (treatment dose $d_j$) is measured by the average transition score of the students attending that school $j$. The cutoff $c_j$ for admission into a school $j$ is equal to the minimum transition score among the students that are assigned to that school $j$. The student's preferences in high schools are not observed in the data, which makes the RDD fuzzy. For example, a student may have a score greater than the cutoff for the best school in her town, but still be assigned to a different school because of her personal preferences. For a transition score $X_i$, the treatment dose of eligibility $D(X_i)$ is equal to the largest $d$ among those schools with admission cutoff $c$ less than $X_i$. The treatment dose received $D_i$ coincides with the treatment dose of eligibility $D(X_i)$ for 40% of the students in the sample. Thus, the assignment is fuzzy, and causal inference beyond ITT effects requires the methods of Section (ref). Following PU, I drop observations with missing values for $Y_i$. I also drop cutoffs without enough observations around them to carry out the matrix inversions of the local polynomial regressions. The dropping of cutoffs leaves the empirical distribution of outcomes, forcing variable, cutoffs, and treatment doses practically unchanged. The estimation sample has 588 cutoffs with a total of 179,995 individuals from 769 schools in 121 towns and 3 years. The variation of cutoff and dose values is displayed in Figure (ref).
Non-parametric identification of $\beta(\boldsymbol{\mathrm{c}})$ is limited to the set $\mathcal{C}$, which is the convex-hull of $\mathcal{C}_{\infty}$. The set $\mathcal{C}_{\infty}$ is not entirely observed, and the researcher relies on the observation of $\mathcal{C}_K$ (Figure (ref)(a)). For the sake of simplicity, I restrict $\beta$ to be a function of dose changes ($u=d'-d$) instead of doses before and after ($d$ and $d'$). The restriction greatly simplifies the visualization and estimation of $\beta(c,d,d')$, because it implies that $\beta(c,d,d')=\phi(c)(d'-d)$, where $\phi$ is a continuously differentiable function. Figure (ref)(b) illustrates the variation of cutoff and dose-change values and defines the limits on identification of policy counterfactuals. For example, it is not possible to identify the effects of randomly assigning students with grades between 8 and 9 to a change in treatment dose of 2. The support of such counterfactual distribution falls outside the observed variation of cutoff and dose-change values. On the other hand, it is possible to identify the ATE of randomly assigning students with grades between 6.5 and 8.5 to dose increases between 0 and 1.5.
The following policy question illustrates the ATE estimator proposed in this paper. Suppose a new charter school is constructed in one of the towns in Romania. The new charter school has more autonomy and better management than traditional public schools, and admitted students experience an increase in school quality as if they were admitted to a school with better peers. More specifically, the policy counterfactual is to give a 0.5 increase in peer-quality to a uniform distribution of scores between 6.5 and 8.5. The ATE parameter is defined as
I follow the estimation procedure suggested in Section (ref) and take into account the restriction $\beta(c,d,d')=\phi(c)(d'-d)$. As in the binary treatment case, the restriction lowers the polynomial degree requirement on the second-step estimation to $\rho_2=1$. See Figure (ref) and the discussion that follows it. The grid for $h_2$ has 32 equally-spaced points between 0.1 and 3.6, respectively, the smallest bandwidth for which the estimator is computable, and the maximum distance between two different cutoffs. The MSE-optimal bandwidth choice is $h_2^* = 1.837$. The new charter school has a bias-corrected ATE of $0.2964$ with standard error of $0.1296$, and it is statistically significant at 5%. Figure (ref)(a) plots $0.5 \widehat{\phi}(c)$, that is, the effect of a $0.5$ increase in the treatment dose for various levels of $c$. The graph reveals heterogeneous marginal effects of ability on returns to school quality. Heterogeneity of treatment effects is a priori unknown, and the ATE estimator proposed in this paper is consistent for $\mu$ regardless of the shape of $\phi(c)$. This highlights the empirical relevance of Theorem (ref) and the importance of the second-step estimation. In other words, the common strategy of normalizing all cutoffs to zero and estimating one discontinuity using the pooled data is not consistent for $\mu$ when $\phi(c)$ has such heterogeneity.
Estimation of treatment effects on ever-compliers requires a parametric functional form on $\beta_{ec}$ (Theorem (ref)). I assume $\beta_{ec}(c,d,d') = \theta_1 (d'-d) + \theta_2 c (d'-d) +\theta_3 c^2 (d'-d)+ \theta_4 c^3 (d'-d)$ and carry out the iterative estimation procedure described in the supplemental appendix's Section (ref). The algorithm achieves convergence of $\theta$s within 30 iterations. The iterated bias-corrected ATE on ever-compliers equals $0.107$ with standard error of $0.0109$. The precision is substantially greater than the non-parametric case. Figure (ref)(b) displays the treatment effect function on ever-compliers for a dose change of $0.5$. Compared to ITT effects in Figure (ref)(a), the return of better schooling on ever-compliers is also positive, but much less heterogeneous across ability levels.
Difficulty in gathering experimental data in many fields within the social sciences makes quasi-experimental techniques such as RDD extremely important to evaluate policies and social programs. RDD has been used in a wide range of applications in economics since the late 1990s. More recently, there has been an increasing number of applications with one forcing variable and multiple cutoffs, assigning individuals to heterogeneous treatments. The demand for multi-cutoff RDD methods is constantly growing, as richer data sets become ever more available.
This paper states conditions under which multiple RDD effects are combined to infer ATE over the entire range of cutoff values. The proposed estimator is consistent and asymptotically normal for ATEs over the entire support of variation in cutoffs and treatment doses. Asymptotic results are derived under a large number of observations and cutoffs in the sharp case of non-parametric treatment effect functions. Sufficient conditions on the rate of growth of the number of cutoffs, relative to the number of observations, are given. These rate conditions determine the feasible choice set of tuning parameters. This paper also shows that non-parametric identification in fuzzy RDD with multiple cutoffs is impossible unless the treatment effect function is finite-dimensional, or there is large variation of cutoff-dose values. A parametric specification provides an MSE-optimal ATE estimator for the fuzzy case that is consistent and asymptotically normal.
The relevance of the ATE estimators proposed in this paper is illustrated with the data of pop2013going on high school assignment in Romania. Of interest is the effect of high school quality on academic performance of students. I find strong evidence of non-linearities in the returns to better schooling, as a function of students' ability level. Monte Carlo simulations demonstrate that such non-linearities severely bias a naive average of local effects that does not use the correction weighting scheme proposed in this paper. Applying the fuzzy RDD methods to the Romanian data reveals causal effects on ever-compliers that are smaller and less heterogeneous than ITT effects.
The proposed estimator converges at the minimax optimal rate of root-$n$, as long as first-step bandwidths converge to zero at $1/K$ rate. It would be interesting to learn about efficiency properties of the ATE estimator. Theoretical tools commonly employed to derive efficiency lower-bounds may not be immediately applicable to the setting of this paper. These tools are designed for regular estimators, and for data drawn from a population where the parameter of interest is identified. In contrast, the fixed-cutoff RDD design relies on an “identification at infinity argument”, and I wonder about the sufficient conditions that would obtain regularity of the ATE estimator. A possibility for future work is to the generalize the uniform convergence tools from this paper to arrive at such conditions.
I am indebted to Han Hong, Caroline Hoxby, and Guido Imbens for invaluable advice. The paper also benefited from feedback received from seminar participants at Stanford, Boston University, Cambridge, Iowa, Notre Dame, CORE-UcLouvain, UCSD, UCDavis, FGV-EESP, FGV-EPGE, Insper, PUC-Rio, Toulouse, UIUC, and at various conferences. I thank Tim Bresnahan, Arun Chandrasekhar, Michael Dinerstein, Ivan Fernandez-Val, Ivan Korolev, Michael Leung, Huiyu Li, Jessie Li, Petra Moser, Stephen Terry, Xiaowei Yu, and anonymous referees for suggestions and comments. I gratefully acknowledge the financial support received from the B.F. Haley and E.S. Shaw Fellowship at SIEPR-Stanford, CORE-UcLouvain, ISLA-Notre Dame, and while visiting the Kenneth C. Griffin Department of Economics at the University of Chicago.