Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
67,725 characters · 14 sections · 69 citation commands
Extrapolating Treatment Effects in Multi-Cutoff Regression Discontinuity Designs
\setcounter{page}{0}\thispagestyle{empty}
Keywords: causal inference, regression discontinuity, extrapolation.
\doublespacing {5pt} {5pt}
The regression discontinuity (RD) design is one of the most credible strategies for estimating causal treatment effects in non-experimental settings. In an RD design, units receive a score (or running variable), and a treatment is assigned based on whether the score exceeds a known cutoff value: units with scores above the cutoff are assigned to the treatment condition, and units with scores below the cutoff are assigned to the control condition. This treatment assignment rule creates a discontinuity in the probability of receiving treatment which, under the assumption that units' average characteristics do not change abruptly at the cutoff, offers a way to learn about the causal treatment effect by comparing units barely above and barely below the cutoff. Despite the popularity and widespread use of RD designs, the evidence they provide has an important limitation: the RD causal effect is only identified for the very specific subset of the population whose scores are “just” above or below the cutoff, and is not necessarily informative or representative of what the treatment effect would be for units whose scores are far from the RD cutoff. Thus, by its very nature, the RD parameter is local and has limited external validity.
We empirically illustrate the advantages and limitations of RD designs employing a recent study of the ACCES (Acceso con Calidad a la Educaci\'on Superior) program by MelguizoSanchezVelasco2016-WD. ACCES is a subsidized loan program in Colombia, administered by the Colombian Institute for Educational Loans and Studies Abroad (ICETEX), that provides tuition credits to underprivileged populations for various post-secondary education programs such as technical, technical-professional, and university degrees. In order to be eligible for an ACCES credit, students must be admitted to a qualifying higher education program, have good credit standing and, if soliciting the credit in the first or second semester of the higher education program, achieve a minimum score on a high school exit exam known as SABER 11. In other words, to obtain ACCES funding students must have an exam score above a known cutoff. Students who are just below the exam cutoff are deemed ineligible, and therefore are not offered financial assistance. This discontinuity in program eligibility based on the exam score leads to a RD design: MelguizoSanchezVelasco2016-WD found that students just above the threshold in SABER 11 test scores were significantly more likely to enroll in a wide variety of post-secondary education programs. The evidence from the original study is limited to the population of students around the cutoff. This standard causal RD treatment effect is informative in its own right but, in the absence of additional assumptions, it cannot be used to understand the effects of the policy for students whose test scores are outside the immediate neighborhood of the cutoff. Treatment effects away from the cutoff are useful for a variety of purposes, ranging from answering purely substantive questions to addressing practically important policy making decisions such as whether to roll-out the program or not.
We propose a novel approach for estimating RD causal treatment effects away from the cutoff that determines treatment assignment. Our extrapolation approach is design-based as it exploits the presence of multiple RD cutoffs across different subpopulations to construct valid counterfactual extrapolations of the expected outcome of interest, given different scores levels, in the absence of treatment assignment. In sum, our approach imputes the average outcome in the absence of treatment of a treated subpopulation exposed to a given cutoff, using the average outcome of another subpopulation exposed to a higher cutoff. Assuming that the difference between these two average outcomes is constant as a function of the score, this imputation identifies causal treatment effects at score values higher than the lower cutoff.
The rest of the article is organized as follows. The next section presents further details on the operation of the ACCES program, discusses the particular program design features that we use for the extrapolation of RD effects, and presents the intuitive idea behind our approach. In that section, we also discuss related literature on RD extrapolation as well as on estimation and inference. Section (ref) presents the main methodological framework and extrapolation results for the case of the “Sharp” RD design, which assumes perfect compliance with treatment assignment (or a focus on an intention-to-treat parameter). Section (ref) applies our results to extrapolate the effect of the ACCES program on educational outcomes, while Section (ref) illustrates our methods using simulated data. Section (ref) presents an extension to the “Fuzzy” RD design, which allows for imperfect compliance. Section (ref) concludes. The supplemental appendix contains additional results, including further extensions and generalizations of our extrapolation methods.
The SABER 11 exam that serves as the basis for eligibility to the ACCES program is a national exam administered by the Colombian Institute for the Promotion of Postsecondary Education (ICFES), an institute within Colombia's National Ministry of Education. This exam may be taken in the fall or spring semester each year, and has a common core of mandatory questions in seven subjects---chemistry, physics, biology, social sciences, philosophy, mathematics, and language. To sort students according to their performance in the exam, ICFES creates an index based on the difference between (i) a weighted average of the standardized grades obtained by the student in each common core subject, and (ii) the within-student standard deviation across the standardized grades in the common core subjects. This index is commonly referred to as the SABER 11 score.
Each semester of every year, ICFES calculates the 1,000-quantiles of the SABER 11 score among all students who took the exam that semester, and assigns a score between 1 and 1,000 to each student according to their position in the distribution---we refer to these scores as the SABER 11 position scores. Thus, the students in that year and semester whose scores are in the top 0.1% are assigned a value of 1 (first position), the students whose scores are between the top 0.1% and 0.2% are assigned a value of 2 (second position), etc., and the students whose scores are in the bottom 0.1% are assigned a value of 1,000 (the last position). Every year, the position scores are created separately for each semester, and then pooled. MelguizoSanchezVelasco2016-WD provide further details on the Colombian education system and the ACCES program.
In this sharp RD design, the running variable is the SABER 11 position score, denoted by $X_i$ for each unit $i$ in the sample, and the treatment of interest is receiving approval of the ACCES credit. Between $2000$ and $2008$, the cutoff to qualify for an ACCES credit was $850$ in all Colombian departments (the largest subnational administrative unit in Colombia, equivalent to U.S. states). To be eligible for the program, a student must have a SABER 11 position score at or below the 850 cutoff.
In the canonical RD design, a single cutoff is used to decide which units are treated. As we noted above, eligibility for the ACCES program between $2000$ and $2008$ followed this template, since the cutoff was $850$ for all students. However, in many RD designs, the same treatment is given to all units based on whether the RD score exceeds a cutoff, but different units are exposed to different cutoffs. This contrasts with the assignment rule in the standard RD design, in which all units face the same cutoff value. RD designs with multiple cutoffs, which we call Multi-Cutoff RD designs, are fairly common and have specific properties Cattaneo-Keele-Titiunik-VazquezBare_2016_JOP.
In $2009$, ICFES changed the program eligibility rule, and started employing different cutoffs across years and departments. Consequently, after $2009$, ACCES eligibility follows a Multi-Cutoff RD design: the treatment is the same throughout Colombia---all students above the cutoff receive the same financial credits for educational spending---but the cutoff that determines treatment assignment varies widely by department and changes each year, so that different sets of students face different cutoffs. This design feature is at the core of our approach for extrapolation of RD treatment effects.
Multi-Cutoff RD designs are often analyzed as if they had a single cutoff. For example, in the original analysis, MelguizoSanchezVelasco2016-WD redefined the RD running variable as distance to the cutoff, and analyzed all observations together using a common cutoff equal to zero. In fact, this normalizing-and-pooling approach Cattaneo-Keele-Titiunik-VazquezBare_2016_JOP, which essentially ignores or “averages over” the multi-cutoff features of the design, is widespread in empirical work employing RD designs. See the supplemental appendix for a sample of recent papers that analyze RD designs with multiple cutoffs across various disciplines.
We first present some initial empirical results using the normalizing-and-pooling approach as a benchmark for later analyses. The outcome we analyze is an indicator for whether the student enrolls in a higher education program, one of several outcomes considered in the original study by MelguizoSanchezVelasco2016-WD. In order to maintain the standard definition of RD assignment as having a score above the cutoff, we multiply the SABER 11 position score by $-1$. We focus on the intention-to-treat effect of program eligibility on higher education enrollment, which gives a Sharp RD design. We discuss an extension to Fuzzy RD designs in Section (ref). We focus our analysis on the population of students exposed to two different cutoffs, $-850$ and $-571$.
For our main analysis, we employ statistical methods for RD designs based on recent methodological developments in Calonico-Cattaneo-Titiunik_2014_ECMA,Calonico-Cattaneo-Titiunik_2015_JASA, Calonico-Cattaneo-Farrell_2018_JASA,Calonico-Cattaneo-Farrell_2020_ECTJ,Calonico-Cattaneo-Farrell_2020_wp, Calonico-Cattaneo-Farrell-Titiunik_2019_RESTAT, and references therein. In particular, point estimators are constructed using mean squared error (MSE) optimal bandwidth selection, and confidence intervals are formed using robust bias correction (RBC). We provide details on estimation and inference in Section (ref).
Figure (ref) reports two RD plots of the data, reporting linear and quadratic global polynomial approximations. Using RBC local polynomial inference, Table (ref) reports that the pooled RD estimate of the ACCES program treatment effect on expected higher education enrollment is $0.125$, with corresponding $95\%$ RBC confidence interval $[0.012 , 0.219]$. These results indicate that, in our sample, students who barely qualify for the ACCES program based on their SABER 11 score are 12.5 percentage points more likely to enroll in a higher education program than students who are barely ineligible for the program. These results are consistent with the original positive effects of ACCES eligibility on higher education enrollment rates reported in MelguizoSanchezVelasco2016-WD. However, the pooled RD estimate only pertains to a limited set of ACCES applicants: those whose scores are barely above or below one of the cutoffs.
Cattaneo-Keele-Titiunik-VazquezBare_2016_JOP show that the pooled RD estimand is a weighted average of cutoff-specific RD treatment effects for subpopulations facing different cutoffs. The empirical results for the pooled and cutoff-specific estimates can be seen in the upper panel of Table (ref). In our sample, the pooled estimate of 0.125 is a (linear in large samples) combination of two cutoff-specific RD estimates, one for units facing the low cutoff $-850$ and one for units facing the high cutoff $-571$. We provide a detailed analysis of these estimates in Section (ref). These cutoff-specific estimates are not directly comparable, as these magnitudes correspond not only to different values of the running variable but also to different subpopulations. We discuss next how the availability of multiple cutoffs can be exploited to learn about treatment effects far from the cutoff in the context of the ACCES policy intervention.
Our key contribution is to exploit the presence of multiple RD cutoffs to extrapolate the standard RD average treatment effects (at each cutoff) to students whose SABER 11 scores are away from the cutoff actually used to determine program eligibility. Our method relies on a simple idea: when different units are exposed to different cutoffs, different units with the same value of the score may be assigned to different treatment conditions, relaxing the strict lack of overlap between treated and control scores that is characteristic of the single-cutoff RD design.
For example, consider the simplest Multi-Cutoff RD design with two cutoffs, $\mathcal{l}$ and $\mathcal{h}$, with $\mathcal{l} < \mathcal{h}$, where we wish to estimate the average treatment effect at a point $\bar{x} \in (\mathcal{l},\mathcal{h})$. Units exposed to $\mathcal{l}$ receive the treatment according to $\mathbbm{1}(X_i\ge \mathcal{l})$, where $X_i$ is unit's $i$ score and $\mathbbm{1}(.)$ is the indicator function, so they are all treated at $X=\bar{x}$. However, the same design contains units who receive the treatment according to $\mathbbm{1}(X_i\ge \mathcal{h})$, so they are controls at both $X=\bar{x}$ and $X=\mathcal{l}$. Our idea is to compare the observable difference in the control groups at the low cutoff $\mathcal{l}$, and assume that the same difference in control groups occurs at the interior point $\bar{x}$. This allows us to identify the average treatment effect for all score values between the cutoffs $\mathcal{l}$ and $\mathcal{h}$.
Our identifying idea is analogous to the “parallel trends” assumption in difference-in-difference designs Abadie_2005_RES, but over a continuous dimension---that is, over the values of the continuous score variable $X_i$.
We contribute to the causal inference and program evaluation literatures Imbens-Rubin_2015_Book,Abadie-Cattaneo_2018_ARE and, more specifically, to the methodological literature on RD designs. See Imbens-Lemieux_2008_JoE, Cattaneo-Titiunik-VazquezBare_2017_JPAM, Cattaneo-Titiunik-VazquezBare_2020_Sage and Cattaneo-Idrobo-Titiunik_2019_Book,Cattaneo-Idrobo-Titiunik_2020_Book, for literature reviews, background references, and practical introductions.
Our paper adds to the recent literature on RD treatment effect extrapolation methods, a nonparametric identification problem in causal inference. This strand of the literature can be classified into two groups: strategies assuming the availability of external information, and strategies based only on information from within the research design. Approaches based on external information include Mealli-Rampichini_2012_JRSSA, WingCook2013-JPAM, Rokkanen2015-wp, and AngristRokkanen2015-JASA. The first two papers rely on a pre-intervention measure of the outcome variable, which they use to impute the treated-control differences of the post-intervention outcome above the cutoff. Rokkanen2015-wp assumes that multiple measures of the running variable are available, and all measures capture the same latent factor; identification relies on the assumption that the potential outcomes are conditionally independent of the available measurements given the latent factor. AngristRokkanen2015-JASA rely on pre-intervention covariates, assuming that the running variable is ignorable conditional on the covariates over the whole range of extrapolation. All these approaches assume the availability of external information that is not part of the original RD design.
In contrast, the extrapolation approaches in Dong-Lewbel_2015_ReStat and Bertanha-Imbens_2020_JBES require only the score and outcome in the standard (single-cutoff) RD design. Dong-Lewbel_2015_ReStat assume mild smoothness conditions to identify the derivatives of the average treatment effect with respect to the score, which allows for a local extrapolation of the standard RD treatment effect to score values marginally above the cutoff. Bertanha-Imbens_2020_JBES exploit variation in treatment assignment generated by imperfect treatment compliance imposing independence between potential outcomes and compliance types to extrapolate a single-cutoff fuzzy RD treatment effect (i.e., a local average treatment effect at the cutoff) away from the cutoff. Our paper also belongs to this second type, as it relies on within-design information, using only the score and outcome in the Multi-Cutoff RD design.
Cattaneo-Keele-Titiunik-VazquezBare_2016_JOP introduced the causal Multi-Cutoff RD framework, which we employ herein, and studied the properties of normalizing-and-pooling estimation and inference in that setting. Building on that paper, Bertanha_2020_JoE discusses estimation and inference of an average treatment effect across multi-cutoffs, assuming away cutoff-specific treatment effect heterogeneity. Neither of these papers addressed the topic of RD treatment effect extrapolation across different levels of the score variable, which is the main goal and innovation of the present paper.
All the papers mentioned above focus on extrapolation of RD treatment effects away from the cutoff by relying on continuity-based methods for identification, estimation and inference, which are implemented using local polynomial regression Fan-Gijbels_1996_Book. As pointed out by a reviewer, an alternative approach to analyzing RD designs is to employ the local randomization framework introduced by Cattaneo-Frandsen-Titiunik_2015_JCI. This framework has been later used in the context of geographic RD designs Keele-Titiunik-Zubizarreta_2015_JRSSA, principal stratification Li-Mattei-Mealli_2015_AoAS, and kink RD designs Ganong-Jager_2018_JASA, among other settings. More recently, this alternative RD framework was expanded to allow for finite-sample falsification testing and for local regression adjustments Cattaneo-Titiunik-VazquezBare_2017_JPAM. See also Sekhon-Titiunik_2016_ObsStud,Sekhon-Titiunik_2017_AIE for further conceptual discussions, Cattaneo-Titiunik-VazquezBare_2020_Sage for another review, and Cattaneo-Idrobo-Titiunik_2020_Book for a practical introduction.
Local randomization RD methods implicitly give extrapolation within the neighborhood where local randomization is assumed to hold because they assume a parametric (usually constant) treatment effect model as a function of the score. However, those methods cannot aid in extrapolating RD treatment effects beyond such neighborhood without additional assumptions, which is precisely our goal. Since local randomization methods explicitly view RD designs as local randomized experiments, we can summarize the key conceptual distinction between that literature and our paper as follows: available local randomization methods for RD designs have only internal validity (i.e., within the local randomization neighborhood), while our proposed method seeks to achieve external validity (i.e., outside the local randomization neighborhood), which we achieve by exploiting the presence of multiple cutoffs (akin to multiple local experiments) together with an additional identifying assumption within the continuity-based approach to RD designs (i.e., parallel control regression functions across cutoffs).
Our core extrapolation idea can be developed within the local randomization framework, albeit under considerably stronger assumptions. To conserve space, in the supplemental appendix, we discuss multi-cutoff extrapolation of RD treatment effects using local randomization ideas, and develop randomization-based estimation and inference methods Rosenbaum_2010_Book,Imbens-Rubin_2015_Book. We also empirically illustrate these methods in the supplemental appendix.
We assume $(Y_i,X_i,C_i,D_i)$, $i=1,2,\dots,n$, is an observed random sample, where $Y_i$ is the outcome of interest, $X_i$ is the score (or running variable), $C_i$ is the cutoff indicator, and $D_i$ is a treatment status indicator. We assume the score has a continuous positive density $f_X(x)$ on the support $\mathcal X$. Unlike the canonical RD design where the cutoff is a fixed scalar, in the Multi-Cutoff RD design the cutoff faced by unit $i$ is the random variable $C_i$ taking values in a set $\mathcal C\subset \mathcal X$. For simplicity, we consider two cutoffs: $\mathcal{C} = \{\mathcal{l},\mathcal{h}\}$, with $\mathcal{l}<\mathcal{h}$ and $\mathcal{l},\mathcal{h} \in \mathcal X$. Extensions to more than two cutoffs and to geographic and multi-score RD designs are conceptually straightforward, and hence discussed in the supplemental appendix.
The conditional density of the score at each cutoff is $f_{X|C}(x|c)$, $c\in\mathcal C$. In sharp RD designs treatment assignment and status are identical, and hence $D_i=\mathbbm{1}(X_i\ge C_i)$. Section (ref) discusses an extension to fuzzy RD designs. Finally, we let $Y_{i}(1)$ and $Y_{i}(0)$ denote the potential outcomes of unit $i$ under treatment and control, respectively, and $Y_i=D_iY_{i}(1)+(1-D_i)Y_{i}(0)$ is the observed outcome.
The potential outcome regression functions are $\mu_{d,c}(x) = \mathbbm{E}[Y_{i}(d)|X_i=x,C_i=c]$, for $d=0,1$. We express all parameters of interest in terms of the “response” function
This function measures the treatment effect for the subpopulation exposed to cutoff $c$ when the running variable takes the value $x$. For a fixed cutoff $c$, it records how the treatment effect for the subpopulation exposed to this cutoff varies with the running variable. As such, it captures a key quantity of interest when extrapolating the RD treatment effect. The usual parameter of interest in the standard (single-cutoff) RD design is a particular case of $\tau_c(x)$ when cutoff and score coincide:
It is well known that, via continuity assumptions, the function $\tau_c(x)$ is nonparametrically identifiable at the single point $x=c$. Our approach exploits the presence of multiple cutoffs to identify this function at other points on a portion of the support of the score variable.
Figure (ref) contains a graphical representation of our extrapolation approach for Multi-Cutoff RD designs. In the plot, there are two populations, one exposed to a low cutoff $\mathcal{l}$, and another exposed to a high cutoff $\mathcal{h}$. The RD effects for each subpopulation are, respectively, $\tau_\mathcal{l}(\mathcal{l})$ and $\tau_\mathcal{h}(\mathcal{h})$. We seek to learn about the effects of the treatment at points other than the particular cutoff to which units were exposed, such as the point $\bar{x}$ in Figure (ref). Below, we develop a framework for the identification of $\tau_\mathcal{l}(x)$ for $\mathcal{l}<x\leq\mathcal{h}$ so that we can assess what would have been the average treatment effect for the subpopulation exposed to the cutoff $\mathcal{l}$ at score values above $\ell$ (illustrated by the effect $\tau_\mathcal{l}(\bar{x})$ in Figure (ref) for the intermediate point $X_i=\bar{x}$).
In our framework, the multiple cutoffs define different subpopulations. In some cases, the cutoff to which a unit is exposed depends only on characteristics of the units, such as when the cutoffs are cumulative and increase as the score falls in increasingly higher ranges. In other cases, the cutoff depends on external features, such as when different cutoffs are used in different geographic regions or time periods. This means that, in our framework, the cutoff $C_i$ acts as an index for different subpopulation “types”, capturing both observed and unobserved characteristics of the units.
Given the subpopulations defined by the cutoff values actually used in the Multi-Cutoff RD design, we consider the effect that the treatment would have had for those subpopulations had the units had a higher score value than observed. This is why, in our notation, the index for the cutoff value is fixed, and the index for the score is allowed to vary and is the main argument of the regression functions. This conveys the idea that the subpopulations are defined by the multiple cutoffs actually employed, and our exercise focuses on studying the treatment effect at different score values for those pre-defined subpopulations. For example, this setting covers RD designs with a common running variable but with cutoffs varying by regions, schools, firms, or some other group-type variable. Our method is not appropriate to extrapolate to populations outside those defined by the Multi-Cutoff RD design.
The main challenge to the identification of extrapolated treatment effects in the single-cutoff (sharp) RD design is the lack of observed control outcomes for score values above the cutoff. In the Multi-Cutoff RD design, we still face this challenge for a given subpopulation, but we have other subpopulations exposed to higher cutoff values that, under some assumptions, can aid in solving the missing data problem and identify average treatment effects. Before turning to the formal derivations, we illustrate the idea graphically.
Figure (ref) illustrates the regression functions for the populations exposed to cutoffs $\mathcal{l}$ and $\mathcal{h}$, with the function $\mu_{1,\mathcal{h}}(x)$ omitted for simplicity. We seek an estimate of $\tau_\mathcal{l}(\bar{x})$, the average effect of the treatment at the point $\bar{x} \in (\mathcal{l},\mathcal{h})$ for the subpopulation exposed to the lower cutoff $\mathcal{l}$. In the figure, this parameter is represented by the segment $\overline{ab}$. The main identification challenge is that we only observe the point $a$, which corresponds to $\mu_{1,\mathcal{l}}(\bar{x})$, the treated regression function for the population exposed to $\mathcal{l}$, but we fail to observe its control counterpart $\mu_{0,\mathcal{l}}(\bar{x})$ (point $b$), because all units exposed to cutoff $\mathcal{l}$ are treated at any $x > \mathcal{l}$. We use the control group of the population exposed to the higher cutoff, $\mathcal{h}$, to infer what would have been the control response at $\bar{x}$ of units exposed to the lower cutoff $\mathcal{l}$. At the point $X_i = \bar{x}$, the control response of the population exposed to $\mathcal{h}$ is $\mu_{0,\mathcal{h}}(\bar{x})$, which is represented by the point $c$ in Figure (ref). Since all units in this subpopulation are untreated at $\bar{x}$, the point $c$ is identified by the average observed outcomes of the control units in the subpopulation $\mathcal{h}$ at $\bar{x}$.
Of course, units facing different cutoffs may differ in both observed and unobserved ways. Thus, there is generally no reason to expect that the average control outcome of the population facing cutoff $\mathcal{h}$ will be a good approximation to the average control outcome of the population facing cutoff $\mathcal{l}$. This is captured in Figure (ref) by the fact that $\mu_{0,\mathcal{l}}(\bar{x})\equiv b \neq c \equiv \mu_{0,\mathcal{h}}(\bar{x})$. This difference in untreated potential outcomes for units facing different cutoffs can be interpreted as a bias driven by differences in observed and unobserved characteristics of the different subpopulations, analogous to “site selection” bias in multiple randomized experiments. We formalize this idea with the following definition.
Table (ref) defines the parameters associated with the corresponding segments in Figure (ref). The parameter of interest, $\tau_\mathcal{l}(\bar{x})$, is unobservable because we fail to observe $\mu_{0,\mathcal{l}}(\bar{x})$. If we replaced $\mu_{0,\mathcal{l}}(\bar{x})$ with $\mu_{0,\mathcal{h}}(\bar{x})$, we would be able to estimate the distance $\overline{ac}$. This distance, which is observable, is the sum of the parameter of interest, $\tau_\mathcal{l}(\bar{x})$, plus the bias $B(\bar{x},c,c')$ that arises from using the control group in the $\mathcal{h}$ subpopulation instead of the control group in the $\mathcal{l}$ subpopulation. Graphically, $\overline{ac} = \overline{ab} + \overline{bc}$. Since we focus on the two-cutoff case, we denote the bias by $B(\bar{x})$ to simplify the notation.
We use the distance between the control groups facing the two different cutoffs at a point where both are observable, to approximate the unobservable distance between them at $\bar{x}$---that is, to approximate the bias $B(\bar{x})$. As shown in the figure, at $\mathcal{l}$, all units facing cutoff $\mathcal{h}$ are controls and all units facing cutoff $\mathcal{l}$ are treated. But under standard RD assumptions, we can identify $\mu_{0,\mathcal{l}}(\mathcal{l})$ using the observations in the $\mathcal{l}$ subpopulation whose scores are just below $\mathcal{l}$. Thus, the bias term $B(\mathcal{l})$, captured in the distance $\overline{ed}$, is estimable from the data.
Graphically, we can identify the extrapolation parameter $\tau_\mathcal{l}(\bar{x})$ assuming that the observed difference between the control functions $\mu_{0,\mathcal{l}}(\cdot)$ and $\mu_{0,\mathcal{h}}(\cdot)$ at $\mathcal{l}$ is constant for all values of the score:
We now formalize this intuitive result employing standard continuity assumptions on the relevant regression functions. We make the following assumptions.
The observed outcome regression functions are $\mu_c(x) = \mathbbm{E}[Y_i|X_i=x,C_i=c]$, for $c\in\mathcal C= \{\mathcal{l}, \mathcal{h}\}$, and note that by standard RD arguments $\mu_{0,c}(c) = \lim_{\varepsilon \uparrow 0} \mu_{c}(c+\varepsilon)$ and $\mu_{1,c}(c) = \lim_{\varepsilon \downarrow 0} \mu_{c}(c+\varepsilon)$. Furthermore, $\mu_{0,\mathcal{h}}(x) = \mu_{\mathcal{h}}(x)$ and $\mu_{1,\mathcal{l}}(x) = \mu_{\mathcal{l}}(x)$ for all $x \in (\mathcal{l}, \mathcal{h})$.
Our main extrapolation assumption requires that the bias not be a function of the score, which is analogous to the parallel trends assumption in the difference-in-differences design.
While technically our identification result only needs this condition to hold at $x=\bar{x}$, in practice it may be hard to argue that the equality between biases holds at a single point. Combining the constant bias assumption with the continuity-based identification of the conditional expectation functions allows us to express the unobservable bias for an interior point, $\bar{x} \in (\mathcal{l},\mathcal{h})$, as a function of estimable quantities. The bias at the low cutoff $\mathcal{l}$ can be written as
Under Assumption (ref), we have \[\mu_{0,\mathcal{l}}(\bar{x}) = \mu_{\mathcal{h}}(\bar{x}) + B(\mathcal{l}), \qquad \bar{x} \in (\mathcal{l},\mathcal{h}),\] that is, the average control response for the $\mathcal{l}$ subpopulation at the interior point $\bar{x}$ is equal to the average observed response for the $\mathcal{h}$ subpopulation at the same point, plus the difference in the average control responses between both subpopulations at the low cutoff $\mathcal{l}$. This leads to our main identification result.
This result can be extended to hold for $\bar{x}\in(\mathcal{l},\mathcal{h}]$ by using side limits appropriately. In Section (ref), we discuss two approaches to provide empirical support for the constant bias assumption. We extend our result to Fuzzy RD designs in Section (ref), and allow for non-parallel control regression functions and pre-intervention covariate-adjustment in the supplemental appendix.
While we develop our core idea for extrapolation from “left to right”, that is, from a low cutoff to higher values of the score, it follows from the discussion above that the same ideas could be developed for extrapolation from “right to left”. Mathematically, the problem is symmetric and hence both extrapolations are equally viable. However, conceptually, there is an important asymmetry. Theorem (ref) requires the regression functions for control units to be parallel over the extrapolation region (Assumption (ref)), while a version of this theorem for “right to left” extrapolation would require that the regression functions for treated units be parallel. These two identifying assumptions are not symmetric because the latter effectively imposes a constant treatment effect assumption across cutoffs (for different values of the score), while the former does not because it pertains to control units only.
We estimate all (identifiable) conditional expectations $\mu_{d,c}(x)=\mathbbm{E}[Y_i(d)|X_i=x,C_i=c]$ using nonparametric local polynomial methods, employing second-generation MSE-optimal bandwidth selectors and robust bias correction inference methods. See Calonico-Cattaneo-Titiunik_2014_ECMA, Calonico-Cattaneo-Farrell_2018_JASA,Calonico-Cattaneo-Farrell_2020_ECTJ,Calonico-Cattaneo-Farrell_2020_wp, and Calonico-Cattaneo-Farrell-Titiunik_2019_RESTAT for more methodological details, and Calonico-Cattaneo-Farrell-Titiunik_2017_Stata and Calonico-Cattaneo-Farrell_2019_JSS for software implementation. See also Hyytinen-etal_2018_QE, Ganong-Jager_2018_JASA and Dong-Lee-Gou_2020_wp for some recent applications and empirical testing of those methods.
To be more precise, a generic local polynomial estimator is $\hat{\mu}_{d,c}(x)=\mathbf{e}_0'\hat{\boldsymbol{\beta}}_{d,c}(x)$, where \[\hat{\boldsymbol{\beta}}_{d,c}(x) = \operatorname*{argmin}_{\mathbf{b}\in\mathbbm{R}^{p+1}}\sum_{i=1}^n(Y_i-\mathbf{r}_p(X_i-x)'\mathbf{b})^2K\left(\frac{X_i-x}{h}\right)\mathbbm{1}(C_i=c)\mathbbm{1}(D_i=d),\] $\mathbf{e}_0$ is a vector with a one in the first position and zeros in the rest, $\mathbf{r}_p(\cdot)$ is a polynomial basis of order $p$, $K(\cdot)$ is a kernel function, and $h$ a bandwidth. For implementation, we set $p=1$ (local-linear), $K$ to be the triangular kernel, $h$ to be a MSE-optimal bandwidth selector, unless otherwise noted. Then, given the two cutoffs $\mathcal{l}$ and $\mathcal{h}$ and an extrapolation point $\bar{x}\in(\mathcal{l},\mathcal{h}]$, the extrapolated treatment effect at $\bar{x}$ for the subpopulation facing cutoff $\mathcal{l}$ is estimated as \[\hat{\tau}_\mathcal{l}(\bar{x}) =\hat{\mu}_{1,\mathcal{l}}(\bar{x})-\hat{\mu}_{0,\mathcal{h}}(\bar{x})-\hat{\mu}_{0,\mathcal{l}}(\mathcal{l})+\hat{\mu}_{0,\mathcal{h}}(\mathcal{l}). \]
The estimator $\hat{\tau}_\mathcal{l}(\bar{x})$ is a linear combination of nonparametric local polynomial estimators at boundary and at interior points depending on the choice of $\bar{x}$ and data availability. Hence, optimal bandwidth selection and robust bias-corrected inference can be implemented using the methods and software mentioned above. By construction, $\hat{\mu}_{d,\mathcal{l}}(\cdot)$ and $\hat{\mu}_{0,\mathcal{h}}(\cdot)$ are independent because the observations used for estimation come from different subpopulations. Similarly, $\hat{\mu}_{0,\mathcal{l}}(\cdot)$ and $\hat{\mu}_{1,\mathcal{l}}(\cdot)$ are independent since the first term is estimated using control units whereas the second term uses treated units. On the other hand, in finite samples, $\hat{\mu}_{0,\mathcal{h}}(\mathcal{l})$ and $\hat{\mu}_{0,\mathcal{h}}(\bar{x})$ can be correlated if the bandwidths used for estimation overlap (or, alternatively, if $\mathcal{l}$ and $\bar{x}$ are close enough), in which case we account for such correlation in our inference results. More precisely, $\mathbbm{V}[\hat{\tau}_\mathcal{l}(\bar{x})|\mathbf{X}] =\mathbbm{V}[\hat{\mu}_{1,\mathcal{l}}(\bar{x})|\mathbf{X}] +\mathbbm{V}[\hat{\mu}_{0,\mathcal{h}}(\bar{x})|\mathbf{X}] +\mathbbm{V}[\hat{\mu}_{0,\mathcal{l}}(\mathcal{l})|\mathbf{X}] +\mathbbm{V}[\hat{\mu}_{0,\mathcal{h}}(\mathcal{l})|\mathbf{X}] -2\mathbbm{C}\mathrm{ov}(\hat{\mu}_{0,\mathcal{h}}(\mathcal{l}),\hat{\mu}_{0,\mathcal{h}}(\bar{x})|\mathbf{X})$, where $\mathbf{X}=(X_1,X_2,\dots,X_n)'$.
Precise regularity conditions for large sample validity of our estimation and inference methods can be found in the references given above. The replication files contain details on practical implementation.
Assessing the validity of our extrapolation strategy should be a key component of empirical work using these methods. In general, while the assumption of constant bias is not testable, this assumption can be tested indirectly via falsification. While a falsification test cannot demonstrate that an assumption holds, it can provide persuasive evidence that an assumption is implausible. We now discuss two strategies for falsification tests to probe the credibility of the constant bias assumption that is at the center of our extrapolation approach.
The first falsification approach relies on a global polynomial regression. We test globally whether the conditional expectation functions of the two control groups are parallel below the lowest cutoff. One way to implement this idea, given the two cutoff points $\mathcal{l} < \mathcal{h}$, is to test $\boldsymbol{\delta}=\mathbf{0}$ based on the regression model \[Y_i = \alpha + \beta\mathbbm{1}(C_i=\mathcal{h}) + \mathbf{r}_p(X_i)'\boldsymbol{\gamma} + \mathbbm{1}(C_i=\mathcal{h})\mathbf{r}_p(X_i)'\boldsymbol{\delta} + u_i, \qquad \mathbbm{E}[u_i|X_i,C_i]=0, \] only for units with $X_i<\mathcal{l}$. In words, we employ a $p$-th order global polynomial model to estimate the two regression functions $\mathbbm{E}[Y_i|X_i=x,X_i<\mathcal{l},C_i=\mathcal{l}]$ and $\mathbbm{E}[Y_i|X_i=x,X_i<\mathcal{l},C_i=\mathcal{h}]$, separately, and construct a hypothesis test for whether they are equal up to a vertical shift (i.e., the null hypothesis is $\mathsf{H}_0:\boldsymbol{\delta}=\mathbf{0}$). This approach is valid under standard regularity conditions for parametric least squares regression. This approach could also be justified from a nonparametric series approximation perspective, under additional regularity conditions.
The second falsification approach employs nonparametric local polynomial methods. We test for equality of the derivatives of the conditional expectation functions for values $x<\mathcal{l}$. Specifically, we test for $\mu^{(1)}_{\mathcal{l}}(x)=\mu^{(1)}_\mathcal{h}(x)$ for all $x<\mathcal{l}$, where $\mu^{(1)}_{\mathcal{l}}(x)$ and $\mu^{(1)}_{\mathcal{h}}(x)$ denote the derivatives of $\mathbbm{E}[Y_i|X_i=x,X_i<\mathcal{l},C_i=\mathcal{l}]$ and $\mathbbm{E}[Y_i|X_i=x,X_i<\mathcal{l},C_i=\mathcal{h}]$, respectively. This test can be implemented using several evaluation points, or using a summary statistic such as the supremum. Validity of this approach is also justified using nonparametric estimation and inference results in the literature, under regularity conditions.
We use our proposed methods to investigate the external validity of the ACCES program RD effects. As mentioned above, our sample has observations exposed to two cutoffs, $\mathcal{l} = -850$ and $\mathcal{h} = -571$. We begin by extrapolating the effect to the point $\bar{x} = -650$; our focus is thus the effect of eligibility for ACCES on whether the student enrolls in a higher education program for the subpopulation exposed to cutoff 850 when their SABER 11 score is $650$.
As described in Section (ref), the observations facing cutoff $\mathcal{l}$ correspond to years 2000 to 2008, whereas observations facing cutoff $\mathcal{h}$ correspond to years 2009 and 2010. Our identification assumption allows these two groups to differ in observable and unobservable characteristics so long as the difference between the conditional expectations of their control potential outcomes is constant as a function of the running variable. In addition, our approach relies on the assumption that the underlying population does not change over time (which is implicit in our notation). We offer empirical support for these assumptions in two ways. First, we implement the tests discussed in Section (ref) to assess the plausibility of Assumption (ref). In addition, Section SA-2 in the Supplemental Appendix shows that our results remain qualitatively unchanged when restricting the empirical analysis to the period 2007-2010, which reduces the (potential) heterogeneity of the underlying populations over time.
We begin by assessing the validity of our constant bias assumption with the methods described in Section (ref). The results can be seen in Tables (ref) and (ref). Specifically, Table (ref) reports results employing global polynomial regression, which does not reject the null hypothesis of parallel trends. Figure (ref) offers a graphical illustration. Table (ref) shows the results for the local polynomial approach, which again does not reject the null hypothesis. Additionally, Figure (ref) plots the difference in derivatives (solid line) between groups estimated nonparametrically at ten evaluation points below $\mathcal{l}$, along with pointwise robust bias-corrected confidence intervals (dashed lines). The figure reveals that the difference in derivatives is not significantly different from zero.
As discussed in Section (ref) and Table (ref), the pooled RD estimated effect is 0.125 with a RBC confidence interval of $[0.012,0.219]$. The single-cutoff effect at $-850$ is $0.137$ with $95\%$ RBC confidence interval of $[0.036,0.232]$, and the effect at $-571$ is somewhat higher at $0.169$, with $95\%$ RBC confidence interval of $[-0.039 , 0.428]$. These estimates based on single-cutoffs are illustrated in Figures (ref)(\subref{subfig:lowC}) and (ref)(\subref{subfig:highC}), respectively.
In finite samples, the pooled estimate may not be a weighted average of the cutoff-specific estimates as it contains an additional term that depends on the bandwidth used for estimation and small sample discrepancies between the estimated slopes for each group. This is evident in Table (ref), where the pooled estimate does not lie between the cutoff specific estimates. This additional term vanishes as the sample size grows and the bandwidths converge to zero, yielding the result in Cattaneo-Keele-Titiunik-VazquezBare_2016_JOP. To provide further evidence on the overall effect of the program, we also estimated a weighted average of cutoff-specific effects using estimated weights. This average effect equals 0.156 with a RBC confidence interval of $[0.025,0.314]$. Since this estimate is a proper weighted average of cutoff-specific effects, it may give a more accurate assessment of the overall effect of the program.
The extrapolation results are illustrated in Figure (ref)(\subref{subfig:extrapol}) and reported in the last two panels of Table (ref). At the $-650$ cutoff, the treated population exposed to cutoff $-850$ has an enrollment rate of $0.756$, while the control population exposed to cutoff $-571$ has a rate of $0.706$. This naive comparison, however, is likely biased due to unobservable differences between both subpopulations. The bias, which is estimated at the low cutoff $-850$, is $-0.142$, showing that the control population exposed to the $-850$ cutoff has lower enrollment rates at that point than the population exposed to the high cutoff $-571$ (0.525 versus 0.667). The extrapolated effect in the last row corrects the naive comparison according to Theorem (ref). The resulting extrapolated effect is $0.756 - (0.706 - 0.142) = 0.191$ with RBC confidence interval of $[ 0.080, 0.336]$.
The choice of the point $-650$ is simply for illustration purposes, and indeed considering a set of evaluation points for extrapolation can give a much more complete picture of the impact of the program away from the cutoff point. In Figures (ref) and (ref), we conduct this analysis by estimating the extrapolated effect at 14 equidistant points between $-840$ and $-580$. The effects are statistically significant, ranging from around $0.14$ to $0.25$.
We report results from a simulation study aimed to assess the performance of the local polynomial methods described in Section (ref). We construct $\mu_{0,\mathcal{h}}(x)$ as a fourth-order polynomial where the coefficients are calibrated using the data from our empirical application, and $\mu_{0,\ell}(x) = \mu_{0,\mathcal{h}}(x) + \Delta$. Based on our empirical findings, we set $\Delta=-0.14$ and an extrapolated treatment effect of $\tau_\ell(\bar{x})=0.19$. We consider three sample sizes: $N=1,000$ (“small N”), $N=2,000$ (“moderate N”), and $N=5,000$ (“large N”). To assess the effect of unbalanced sample sizes across evaluation points/cutoffs, our simulation model ensures that some evaluation points/cutoffs have fewer observations than others. In particular, the available sample size to estimate $\mu_\ell(\ell)$ is always less than a third of the sample size available to estimate $\mu_\mathcal{h}(\bar{x})$. We provide all details in the supplemental appendix to conserve space.
The results are shown in Table (ref). The robust bias-corrected $95\%$ confidence interval for $\tau_\ell(\bar{x})$ has an empirical coverage rate of around 91 percent in the “small N” case. This is because one of the parameters, $\mu_\ell(\ell)$, is estimated using very few observations. The empirical coverage rate increases slightly to 92 percent in the “moderate N” case, and to about 94 percent in the “large N” case. In sum, in our Monte Carlo experiment, we find that local polynomial methods can yield estimators with little bias and RBC confidence intervals with accurate coverage rates for RD extrapolation.
The main idea underlying our extrapolation methods can be extended in several directions that may be useful in other applications. We briefly discuss an extension to Fuzzy RD designs employing a continuity-based approach. In the supplemental appendix we discuss other extensions: covariate adjustments (i.e., ignorable cutoff bias), score adjustments (i.e., polynomial-in-score cutoff bias), many multiple cutoffs, and multiple scores and geographic RD designs.
In the Fuzzy RD design, treatment compliance is imperfect, which is common in empirical applications. For simplicity, we focus on the case of one-sided (treatment) non-compliance: units assigned to the control group comply with their assignment but units assigned to treatment status may not. This case is relevant for a wide array of empirical applications in which program administrators are able to successfully exclude units from the treatment, but cannot force units to actually comply with it.
We employ the Fuzzy Multi-Cutoff RD framework of Cattaneo-Keele-Titiunik-VazquezBare_2016_JOP, which builds on the canonical framework of AngristImbensRubin1996-JASA. Let $D_i(x,c)$ be the binary treatment indicator and $\underline{x}\leq\bar{x}$. We define compliers as units with $D_i(\underline{x},c)<D_i(\bar{x},c)$, always-takers as units with $D_i(\underline{x},c)=D_i(\bar{x},c)=1$, never-takers as units with $D_i(\underline{x},c)=D_i(\bar{x},c)=0$, and defiers as units with $D_i(\underline{x},c)>D_i(\bar{x},c)$. We assume the following conditions:
The conditions are standard in the fuzzy RD literature and used to identify the local average treatment effect (LATE), which is the treatment effect for units that comply with the RD assignment. The following result shows how to recover a LATE-type extrapolation parameter in this fuzzy RD setting.
The left-hand side can be interpreted as an “adjusted” Wald estimand, where the adjustment allows for extrapolation away from the cutoff point $\mathcal{l}$. More precisely, this theorem shows that under one-sided (treatment) noncompliance we can recover the average extrapolated effect on compliers by dividing the adjusted intention-to-treat parameter by the proportion of compliers.
We introduced a new framework for the extrapolation of RD treatment effects when the RD design has multiple cutoffs. Our approach relies on the assumption that the average outcome difference between control groups exposed to different cutoffs is constant over a chosen extrapolation region. Our method does not require any information external to the design, and can be used whenever two or more cutoffs are used to assign the treatment for different subpopulations, which is a very common feature in many RD applications. Our main extrapolation idea can also be used in settings with more than two cutoffs, multi-scores RD designs PapayWillettMurnane2011-JoE,ReardonRobinson2012-JREE, and geographic RD designs Keele-Titiunik_2015_PA. In addition, our main idea can be extended to the RD local randomization framework introduced by Cattaneo-Frandsen-Titiunik_2015_JCI and Cattaneo-Titiunik-VazquezBare_2017_JPAM. These additional results are reported in the supplemental appendix for brevity.
\onehalfspacing