EconBase
← Back to paper

On Extrapolation of Treatment Effects in Multiple-Cutoff Regression Discontinuity Designs

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

70,624 characters · 29 sections · 88 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

On Extrapolation of Treatment Effects in Multiple-Cutoff Regression Discontinuity Designs

\doublespacing

abstractWe investigate how to learn treatment effects away from the cutoff in multiple-cutoff regression discontinuity designs. Using a microeconomic model, we demonstrate that the parallel-trend type assumption proposed in the literature is justified when cutoff positions are assigned as if randomly and the running variable is non-manipulable (e.g., parental income). However, when the running variable is partially manipulable (e.g., test scores), extrapolations based on that assumption can be biased. As a complementary strategy, we propose a novel partial identification approach based on empirically motivated assumptions. We also develop a uniform inference procedure and provide two empirical illustrations.

{Keywords: Decision model, external validity, partial identification, regression discontinuity designs}

{JEL Classification: C14, C21, D00, D84}

Introduction

Regression discontinuity (RD) designs are among the most credible quasi-experimental methods for identifying causal effects. In the RD framework, treatment status changes discontinuously at a known threshold, allowing for identification of the treatment effect at the cutoff, provided a mild smoothness assumption holds Hahn_etal:2001. This feature provides RD designs with strong internal validity.

However, this internal validity stems from the local nature of the RD designs, and hence, their external validity often remains uncertain. This is a primary limitation of RD designs, as the treatment effect estimated at the single cutoff is not necessarily the only parameter of interest. Researchers are often interested in the treatment effects away from the cutoff to guarantee a certain external validity of the local estimate Cerulli_etal:2017. Furthermore, in some applications, policymakers may wish to understand the treatment effect at specific points or regions other than the current cutoff point---e.g., when considering moving the current cutoff to a different level Dong_Lewbel:2015. Standard RD designs, however, provide limited information for addressing such questions, and the extrapolation of RD treatment effects remains a “crucial open question" Abadie_Cattaneo:2018.

To address the issue of weaker external validity in RD designs, an additional source of information is required. A promising avenue is to leverage the presence of multiple cutoffs, a scenario frequently encountered in empirical work (Bertanha:2020, Cattaneo_etal2016jop). For example, eligibility thresholds for scholarships often vary by student background, such as gender, race, or cohort.

Although such multiple cutoffs have often been normalized in empirical studies, Cattaneo_etal:2021JASA_extrapolating recently proposed a novel strategy to identify the treatment effects away from the cutoff point by effectively utilizing these multiple cutoffs. Their identification strategy is as follows: Let $\mu_{0,l}(x)$ and $\mu_{1,l}(x)$ represent the conditional expectation functions under controlled and treated status, respectively, for a group with cutoff $l$ (see Figure (ref)). In standard RD, we typically identify and estimate $\mu_{1,l}(l)-\mu_{0,l}(l)$. The fundamental challenge in extrapolation is that, while $\mu_{1,l}(x)$ can be observed for $x > l$, $\mu_{0,l}(x)$ cannot. However, when another group with a higher cutoff $h(>l)$ is present, their regression function under control status, $\mu_{0,h}(x)$, can be observed for $x\in(l,h)$. Cattaneo_etal:2021JASA_extrapolating identifies the treatment effects for group $l$ at $\bar{x}\in(l,h)$ as $\mu_{1,l}(\bar{x}) - \mu_{0,h}(\bar{x}) + \{\mu_{0,h}(l) - \mu_{0,l}(l)\}$---referred to as “Extrapolated Effect" in Figure (ref)---by introducing a “parallel trend"-type condition: that $\mu_{0,h}(x) - \mu_{0,l}(x)$ is constant over $(l, h)$.

figure[figure omitted — 170 chars of source]

This parallel trend type assumption, termed the constant bias assumption, is undoubtedly useful for drawing additional policy insights from RD studies. Considering the popularity of difference-in-differences (DID) analysis in empirical economics, this extrapolation strategy has the potential to be widely used. At the same time, it appears somewhat ad hoc, and its empirical motivation and validity have not been fully explored. Consequently, researchers lack clear guidance on how to assess its plausibility, which may have restricted the potential to conduct extrapolation analysis relying on this assumption.

Our first goal is to investigate the plausibility of the constant bias assumption---when it holds and when it fails---thereby clarifying its scope of applicability. To do this, building on Fudenberg_Levine:2022AEJmicro, we develop an economic model that links multiple-cutoff RD designs to the decision-making behavior of rational agents.

The key implications of our analysis are twofold. First, the constant bias assumption is likely to hold when the distribution of the unobserved characteristics is similar across groups and the running variable is entirely non-manipulable by agents (e.g., parental income level). In other words, when the running variable is non-manipulable, the assumption is plausible if the cutoff position is assigned as if at random. This insight will be helpful for empirical researchers, as it links the validity of the constant bias assumption---that is, a functional form assumption---to a familiar random assignment-like condition often invoked in the causal inference literature. For example, if the income threshold for financial aid is revised from the previous year but unobserved characteristics are assumed to remain stable across cohorts, then the constant bias assumption among these two cohorts will plausibly hold.

Unfortunately, however, our model suggests that when the running variable is partially manipulable in the sense of McCrary:2008 (e.g., test scores), this assumption may fail even when the only difference between groups is the cutoff positions. The intuition behind this result is simple: if agents optimally choose their level of effort---which affects both the running variable and the outcome---then a shift in the cutoff position can induce a change in effort. This, in turn, alters the distribution of both the running variable (e.g., test scores) and the outcome variable (e.g., initial wage). These distributional shifts may undermine the plausibility of the constant bias assumption. Consequently, extrapolated estimates may be biased, even when the groups are otherwise identical.

To provide a practical solution for extrapolation when the constant bias assumption may fail, we introduce an alternative framework that is particularly well-suited to multi-cutoff RD designs and broadly applicable, especially in educational settings, one of the most common applications of RD designs. Our strategy relies on a set of qualitative assumptions that are often easier to assess in practice. Specifically, our identification approach leverages a commonly employed monotonicity assumption for $\mu_{0,c}(x),c\in\{l,h\}$, along with a dominance assumption, $\mu_{0,h}(x)\geq\mu_{0,l}(x)$. As noted in Babii_Kumar:2023, the monotonicity assumption is often reasonable in RD applications. The dominance assumption is also plausible in several multi-cutoff RD settings, particularly when cutoffs are designed to reduce inequities or reflect pre-existing differences between groups. For example, consider an affirmative action scenario in which scholarship thresholds are relaxed for students from disadvantaged backgrounds---e.g., those from high-poverty regions (Melguizo_etal:2016). In such cases, it is reasonable to assume that the average outcome function for disadvantaged students is lower than that of more advantaged students.

Under these assumptions, we derive partial identification results that remain valid even when the constant bias assumption fails. These bounds offer a practical and robust alternative, complementing the point identification result of Cattaneo_etal:2021JASA_extrapolating. In addition, we develop estimation and uniform inference procedures tailored to these bounds. The empirical relevance of our approach is illustrated through applications to two empirical examples.

Plan of the Article

The remainder of this article is organized as follows. The rest of this section reviews the related literature. In Section (ref), we investigate the applicability and limitations of the constant bias assumption of Cattaneo_etal:2021JASA_extrapolating using an economic model. Section (ref) proposes an alternative identification strategy under the sharp RD design. Empirical illustrations are provided in Section (ref). Section (ref) concludes this article. Proofs are collected in Appendix (ref). Online Appendix---containing an extension to the one-sided fuzzy RD case, other omitted theoretical results, and simulation studies---is available. The R code for replicating the empirical analysis and implementing the proposed methods is also provided.

Related Literature

The methodological RD literature has substantially expanded over the past few decades. In the standard single-cutoff RD setting, the theoretical foundations have been developed for nonparametric identification Hahn_etal:2001, point estimation Imbens_Kaluanaraman:2011, robust bias-corrected inference procedure Calonico_etal:2014, calonico2018effect, calonico2020optimal, falsification tests Arai_etal:2022, Bugni_Canay:2021, Canay_Kamat:2017, Cattaneo_etal:2020, Fusejima_etal:2024, McCrary:2008, among other important extensions. For a comprehensive recent review of these and related developments, see Cattaneo_Titiunik:2022.

While these works focus on the internal validity of RD designs, recent studies have begun to address concerns regarding external validity. Angrist_Rokkanen:2015jasa proposes an extrapolation strategy applicable when the potential outcomes and the running variable are mean-independent conditional on some covariates. Bertanha_Imbens:2020jbes considers the case where the potential outcomes and compliance types are independent conditional on the running variable. Dong_Lewbel:2015 examines marginal extrapolation under mild smoothness conditions. Mehta:2019 derives bounds on the average treatment effect by assuming that the policymaker knows the treatment effect and sets the cutoff optimally. Deaner_Kwon:2025 proposes an extrapolation strategy that is applicable when an additional covariate satisfying the so-called comonotonicity assumption is available.

Cattaneo_etal:2021JASA_extrapolating explicitly leverages multiple cutoffs to extrapolate treatment effects, relying on a constant bias assumption. This approach serves as the foundation for the discussion in the present paper. Specifically, we connect their constant bias assumption to an economic model to clarify its applicability, and we further propose a complementary strategy by introducing a set of different assumptions based on an empirical motivation. Relatedly, Sun:2023 relaxes the constant bias assumption by introducing a bias bounding constant, following Manski_Pepper:2018. Her approach is also useful when the constant bias assumption seems implausible in multi-cutoff RD settings. One crucial difference between her strategy and ours is that the former requires researchers to specify a theoretical smoothness bounding constant, which can be a non-trivial task in practice, as emphasized in Cattaneo_Titiunik:2022. In contrast, our identification result does not require researchers to specify such constants. Instead, it relies only on a set of qualitative assumptions, whose plausibility is often easier to assess in practice.

This paper is related to the literature on microeconomic analysis of econometric methods. Marx_etal:2024JPEmicro analyzes the validity of the parallel trend assumption in the DID framework by embedding it within a dynamic choice model of rational agents. Motivated by this work, our paper interprets the constant bias assumption of Cattaneo_etal:2021JASA_extrapolating through the lens of individual decision-making in an RD environment, aiming to better understand when this assumption is empirically plausible. To this aim, we build on the work by Fudenberg_Levine:2022AEJmicro, which originally examines how agents’ learning behavior influences causal estimands, showing that such effects are neutral in an RD design.

Finally, our work also contributes to the partial identification literature (e.g., Manski:1997). In the context of program evaluation, numerous bounding strategies have been developed. Manski_Pepper:2018 and Rambachan_Roth:2023, for example, consider bounds on DID effects that are still valid when the parallel trend assumption is not satisfied. Within the RD context, Gerard_etal:2020 provides bounds that are robust to manipulation of the running variable.

The Constant Bias Assumption

Econometric Setup

We begin by introducing the notion of constant bias assumption using econometric terminology. For simplicity, we focus on a two-cutoff sharp RD design.

Let $Y_i$ denote the outcome variable, $X_i$ the running variable, $C_i\in\{l,h\}\,(l<h)$ the cutoff, and $D_i\in\{0,1\}$ be a treatment indicator, which takes one if $i$ is treated. Treatment assignment follows the rule: $D_i = \mathbf{1}\{X_i \geq C_i\}$. Let $Y_i(d)$ denote the potential outcome under treatment status $D_i=d\in\{0,1\}$. Define the conditional expectations and treatment effect functions as

align*[align* omitted — 151 chars of source]

In typical RD settings, we identify $\tau_c(c)$ under a mild continuity assumption at $c$ Hahn_etal:2001:

assumption[Continuity] $\mu_{d,c}(x)$ is continuous at $x=c$.

Yet, researchers are often interested in treatment effects at other points, either to assess the external validity of the local RD estimate or to learn about subpopulations far from the cutoff (as discussed in Section (ref)). These parameters, however, are generally not identified under the standard continuity assumption alone.

To overcome this limitation, Cattaneo_etal:2021JASA_extrapolating propose the following assumption:

assumption[Constant Bias] Let $B(x) = \mu_{0,h}(x) - \mu_{0,l}(x)$. It holds that $B(l) = B(x)$ for all $x\in(l,h)$.

Under Assumptions (ref) and (ref), Cattaneo_etal:2021JASA_extrapolating provides the following identification result:

align[align omitted — 161 chars of source]

where the CB stands for “Constant Bias."

The continuity assumption is a weak requirement. Hence, the key identifying condition is Assumption (ref), namely the constant bias assumption. When, then, will this constant bias assumption be justified? Clarifying the conditions under which this assumption holds is crucial for understanding the scope and limitations of Cattaneo_etal:2021JASA_extrapolating's (Cattaneo_etal:2021JASA_extrapolating) result. When valid, this assumption provides a powerful extrapolation strategy, potentially yielding richer policy implications. However, their strategy can lead to a biased estimate if the regression functions of the two groups exhibit a different pattern.

To investigate the plausibility of this assumption, we turn in the next subsection to an economic model of a rational agent's behavior in a multi-cutoff RD environment, and examine its implications for Assumption (ref).

A Decision Model of the Multi-Cutoff RD Environment

Economic Setup

To facilitate intuitive understanding, we describe the model using an educational setting---one of the most common applications---although the framework will apply broadly to other settings.\footnote{Depending on the empirical context, the model developed below may not be directly applicable. However, we believe that our analysis nonetheless provides potentially valuable insight for such cases. In particular, in settings that can be formulated as some agent’s decision problem, a researcher could develop and analyze a variant of our model to examine the validity of the constant bias assumption in their context as well.}

Building on Fudenberg_Levine:2022AEJmicro, we assume that agents decide how much costly effort $e_i$ to invest, influencing their short-term outcome $S_i$ and future outcome $Y_i(0)$ under the control state. For example, in a typical educational setting, student $i$ inputs their study effort $e_i$ to achieve a higher test score $S_i$ and higher future earnings $Y_i(0)$. In reality, the realized outcomes of $S_i$ and $Y_i(0)$ are not solely determined by the effort and are also affected by stochastic shocks. We model this as follows:

align*[align* omitted — 129 chars of source]

where $y(\cdot)$ and $s(\cdot)$ are structural functions, $\eta_i^y$ and $\eta_i^s$ are zero-mean stochastic errors that are independent of other factors, and $\gamma$ represents a group-level difference in $Y_i(0)$. While our formulation is inspired by Fudenberg_Levine:2022AEJmicro, we depart from their framework in two ways that are particularly relevant for the questions we study: we do not assume linearity of $y(\cdot)$ and $s(\cdot)$ in $e_i$, and we allow the running variable to differ from $S_i$.

Effort incurs a cost $K(e_i, \epsilon_i)$, which also depends on the innate ability $\epsilon_i$, known to the agent $i$. Thus, the cost of a given effort level varies across individuals.

Agents anticipate that they can receive an additional benefit $\tilde{\tau}_i$ in the future if their running variable $X_i$---which may or may not coincide with $S_i$---exceeds a predetermined threshold $C_i$, known to the individual. This $\tilde{\tau}_i$ represents the agent's subjective belief about the benefit from the treatment, which may differ from the actual one.

In this setup, we assume an agent decides the amount of effort to maximize expected utility:

align*[align* omitted — 243 chars of source]

where $v(\cdot)$ represents the utility derived from the short-run outcome $S_i = s(e_i) + \eta_i^s$, $\beta$ denotes the discount factor, and the expectation is taken with respect to $\eta_i^s$ and $\eta_i^y$. This is simplified as

align[align omitted — 187 chars of source]

where $u\left(s(e_i)\right)=\E{v\left(s(e_i) + \eta_i^s\right)}$.

Decision Problem and the Running Variable Manipulability

The optimization problem in Equation (ref) reveals that the agent's decision depends on $\P{X_i \geq C_i}$---that is, the probability of crossing the threshold. The agent interprets this probability in a different way depending on whether the running variable $X_i$ is influenced by their effort. Consider two examples:

itemize• If financial aid eligibility is based on family wealth level, which is fixed and unaffected by student effort. In this case, $\P{X_i \geq C_i}$ is either 0 or 1---fully deterministic. • If aid eligibility is determined by a test score, $X_i = S_i$, then the probability of crossing the threshold depends on effort via $s(e_i)$.

Following McCrary:2008, we say that the running variable is partially manipulable when it is under the agent's control but also affected by an idiosyncratic element, i.e., $X_i=S_i=s(e_i)+\eta^s_i$ in our formulation.\footnote{Note that the so-called score manipulation problem, which undermines the validity of standard RD designs, does not arise when the running variable is only partially manipulable (Lee:2008, McCrary:2008). We also note that this article does not consider the case in which the running variable is completely manipulable, as such manipulability invalidates the identification of RD designs.} In contrast, we say that the running variable is non-manipulable when it is not a function of the agent's effort.

Thus, agents face one of two decision problems depending on whether the running variable is manipulable:

subequations\begin{align} &\max_{e_i}\bigg[ u\left(s(e_i)\right) - K(e_i, \epsilon_i) + \beta y(e_i) \bigg] if X_i is non-manipulable,\\ &\max_{e_i}\bigg[ u\left(s(e_i)\right) - K(e_i, \epsilon_i) + \beta\big\{ y(e_i) +\tilde{\tau}_i \P{s(e_i) + \eta_i^s \geq C_i} \big\} \bigg] if X_i=S_i. \end{align}

We will write the optimal level of effort by $e_i^*$.

Main Results

We are now in a position to analyze the constant bias assumption within our model environment.

Non-Manipulable Running Variable Case

We begin with the case where the running variable is non-manipulable, corresponding to equation (ref). An immediate implication from (ref) is that the optimal effort level $e_i^*$ does not depend on the cutoff $C_i$. This means that differences across groups can only arise from differences in the distribution of ability $\epsilon_i$ (through the cost function) and from the group-level shift $\gamma$ in outcomes. This leads to the following result:

propSuppose that the running variable is non-manipulable. If the distribution of $\epsilon_i$ is identical between those with $C_i=l$ and those with $C_i=h$, Assumption (ref) holds.

Statistically, Proposition (ref) says that the constant bias assumption holds if $\epsilon_i \perp\!\!\!\!\perp C_i$, that is, when cutoff assignment is as good as random. This interpretation is appealing for empirical researchers, as it links the validity of the constant bias assumption to a familiar random assignment-like condition often invoked in the causal inference literature. It also implies that a conditional version of the assumption aligns with the logic of unconfoundedness, making such an extension a natural one (see Remark (ref) below).

In this view, the proposition offers a useful reference point for assessing whether the constant bias assumption is reasonable in practical applications. If an economist believes that the cutoff positions are set “exogenously" and that groups are comparable in terms of unobserved characteristics, then the constant bias assumption is likely to hold in settings with a non-manipulable running variable. Consider, for example, a case in which the income threshold for financial aid is revised from the previous year---perhaps due to a change in budget constraints. Alternatively, imagine a newly introduced aid program that uses a wealth-based threshold. If unobserved characteristics are assumed to remain stable across cohorts, then the constant bias assumption may plausibly hold, and the extrapolation strategy of Cattaneo_etal:2021JASA_extrapolating can be applied. An empirical setting of this kind is found in Londono-Velez_etal:2020aejep. In particular, Figure 5 of their paper seems to illustrate a context where the constant bias assumption appears credible.

That said, the assumption can fail when the cutoff is correlated with unobserved group characteristics. For instance, if more generous cutoffs are systematically applied to disadvantaged or minority groups, then the equality in $\epsilon_i$ distributions may not hold, even when the running variable itself is non-manipulable. Hence, when applying Cattaneo_etal:2021JASA_extrapolating's (Cattaneo_etal:2021JASA_extrapolating) method, we recommend that researchers justify the assumption of similarity in unobservables across groups. Balance tests on observed pre-treatment covariates may offer indirect evidence. For example, comparing covariate distributions or formally testing for differences can help assess plausibility.

remark[Unconfoundedness] As discussed above, the constant bias assumption is likely to hold when cutoff assignment is as good as random, i.e., $C_i \perp\!\!\!\!\perp \epsilon_i$. While this assumption may not necessarily hold in all settings, researchers can still appeal to the constant bias assumption conditional on covariates, in a manner analogous to the unconfoundedness assumption in observational studies (e.g., imbens_rubin:2015causal). Specifically, even if $\epsilon_i \perp\!\!\!\!\perp C_i$ is questionable, one might instead assume $\epsilon_i \perp\!\!\!\!\perp C_i \mid \bm{Z}_i$, where $\bm{Z}_i$ is a vector of predetermined covariates not influenced by effort $e_i$. In such cases, the conditional constant bias assumption, briefly discussed in Cattaneo_etal:2021JASA_extrapolating, provides a powerful alternative identification strategy.

Partially Manipulable Running Variable Case

We now turn to the case where the running variable is partially manipulable, as characterized by equation (ref). To proceed, we impose some conditions.

assumption\begin{itemize} • $u\left(s(e_i)\right) + \beta\left\{y(e_i) + \tilde{\tau}_i\P{s(e_i) + \eta_i^s \geq C_i}\right\}$ is strictly concave in $e_i$, • $u\left(s(e_i)\right)$, $K(e_i, \epsilon_i)$, $y(e_i)$, $s(e_i)$, and $\P{s(e_i) + \eta_i^s \geq C_i}$ are continuously differentiable in $e_i$. • The distribution of the shock $\eta_i^s$ has a density function $f_{\eta^s}$. • $K(e_i, \epsilon_i)$ is convex in $e_i$, • The support of $\epsilon_i$ is identical across both groups. \end{itemize}
assumption$\tilde{\tau}_i = \tilde{\tau}$ for every $i$.

Assumption (ref) comprises standard regularity conditions. Assumptions (ref)(i) and (ii) are smoothness and (high-level) concavity assumptions that ensure analytical traceability. Assumption (ref)(iii) is also a weak smoothness assumption. Assumption (ref)(iv) requires convexity of the cost function, covering commonly used specifications such as the linear and quadratic cost functions. Assumption (ref)(v) imposes a mild support condition on ability.\footnote{To derive the sharpest theoretical prediction, we would need to rely on some parametric functional form assumptions on the structural functions. However, we hesitate to do so, as such implications could be driven more by auxiliary restrictions than by fundamental economic mechanisms. Instead, we maintain generality, while illustrating the intuition with numerical examples later in this section.}

Assumption (ref) rules out heterogeneity in agents’ beliefs about treatment benefits. While strong, the next result suggests that even under this homogeneity assumption, the constant bias assumption may not hold.

propSuppose that the running variable is partially manipulable and that Assumptions (ref)-(ref) are satisfied. Then the optimal effort $e_i^*$ does not depend on $C_i$ for any $\epsilon_i$ if and only if the density function $f_{\eta^s}$ is periodic with period $h-l$, i.e., $f_{\eta^s}(z)=f_{\eta^s}(z+h-l)$, on the interval $[l-\sup_{\epsilon}s(e^*(\epsilon)), h-\inf_{\epsilon}s(e^*(\epsilon))]$.

In most empirical settings, the condition made on the density function is questionable. We typically have that $l < \sup_{\epsilon}s(e^*(\epsilon))$ and $\inf_{\epsilon}s(e^*(\epsilon)) < h$, so the interval $[l-\sup_{\epsilon}s(e^*(\epsilon)), h-\inf_{\epsilon}s(e^*(\epsilon))]$ includes zero. Now, recalling that $f_{\eta^s}$ represents the distribution of idiosyncratic noise in the score, it is far more natural to assume that this density is unimodal and centered at zero, such as the Gaussian errors, rather than being periodic (see Figure (ref)). Therefore, the proposition states that the optimal effort generally does depend on the cutoff position $C_i$.

figure[figure omitted — 736 chars of source]

The dependence of the effort input $e^*_i$ on $C_i$ implies that the distribution of $S_i$ and $Y_i(0)$, which are functions of $e^*_i$, can differ between the two groups. Hence, in contrast to the case with a non-manipulable running variable, the validity of the constant bias assumption is not guaranteed, even under random assignment, identical distributions of $\epsilon_i$, and the strong homogeneity assumption $\tilde{\tau}_i = \tilde{\tau}$. Of course, the dependence of $e_i^*$ on $C_i$ does not rule out the possibility that the constant bias assumption holds; however, motivating this assumption in practice may be challenging, since its validity depends on the functional forms of the structural functions, which are not observable by economists. Furthermore, the example below illustrates that a deviation from constancy can be substantial:

exampleSuppose that $u\left(s(e_i)\right)=s(e_i) = 5\sqrt{e_i}$, $y(e_i) = 10\sqrt{e_i}$, and $K(e_i, \epsilon_i)=15(2-\epsilon_i)e_i$. The distribution of $\eta_i^s$ is triangle, i.e., $f_{\eta^s}(z) = (1 - |z|)_{+}$, which is made to obtain an explicit solution. The ability $\epsilon$ follows the uniform distribution, $\text{Uniform}(0,1)$. The cutoffs are $C_i\in\{2,3\}$. We set $\tilde{\tau}=1$, $\gamma=0$, and $\beta=1$. In this setup, the optimal effort can be computed as \begin{align} e_i^* = \begin{cases} \dfrac{1}{4(2-\epsilon_i)^2} & if \,\epsilon_i \leq \dfrac{4C_i -9}{2(C_i - 1)} or \,\dfrac{4C_i -1}{2(C_i + 1)} < \epsilon_i\\ \dfrac{(4+C_i)^2}{(17-6\epsilon_i)^2} & if \,\dfrac{6C_i -10}{3C_i} < \epsilon_i \leq \dfrac{4C_i -1}{2(C_i + 1)}\\ \dfrac{(4-C_i)^2}{(7-6\epsilon_i)^2} & if \,\dfrac{4C_i -9}{2(C_i-1)} < \epsilon_i \leq \dfrac{6C_i -10}{3C_i} \end{cases}, \end{align} which confirms the dependence of $e_i^*$ on $C_i$. Plugging in this optimal effort to $s(e_i)$ and $y(e_i)$, we can compute the conditional expectations $\E{Y_i(0) | X_i}$ for each group, which are shown in Figure (ref). The constant bias assumption is not satisfied, even though the two groups are similar in their ability and face the same decision problem except for the cutoff points.

As illustrated in the example above, even when all determinants except the cutoff are identical across groups, the resulting regression functions can differ when the running variable is partially manipulable.

Besides, the validity of the assumption becomes more unclear when the distribution of $\epsilon_i$ is supposed to differ, as illustrated in the next example:

exampleSuppose the same structural functions as Example (ref). We here assume that the group with $C_i=3$ is more advantaged in that the ability $\epsilon$ for the group with $C_i=3$ follows $\text{Uniform}(2/3,5/3)$. The optimal effort is determined by (ref). The regression functions are shown in Figure (ref). The constant bias assumption is, again, not satisfied.
figure[figure omitted — 619 chars of source]

These findings highlight a concern for extrapolating treatment effects under the constant bias assumption in settings where the running variable is partially manipulable. When the structural functions are unknown---as is typically the case---justifying constancy becomes empirically difficult.

remarkEven when the running variable partly reflects agent effort, such as a test score, the concerns discussed in this section may be mitigated depending on the design. For instance, if the introduction of a scholarship program or the implementation of multiple cutoffs is determined after the entrance exam has been administered, then all groups can be seen as facing the same decision problem at the time of their effort choice. In such cases, the distributional equivalence of $\epsilon_i$ implies the constant bias assumption.
comment\subsubsection{Summary of Implications for the Constant Bias Assumption} In standard RD settings, the type of running variable is not particularly crucial for identification, as long as it is not completely manipulable (Lee:2008, McCrary:2008). However, when extrapolating treatment effects under the constant bias assumption, the manipulability of the running variable becomes an important consideration. If the running variable is non-manipulable, then the extrapolated RD estimates remain valid, provided that the distribution of individual characteristics is similar across groups. By contrast, if the running variable is partially manipulable, this conclusion no longer holds in general. In such cases, even when group characteristics appear similar, extrapolated estimates may be biased due to endogenous responses to the cutoff. Moreover, in empirical applications where researchers cannot confidently assert the similarity of unobserved characteristics across groups, it becomes even more difficult to justify the constant bias assumption, regardless of the running variable's manipulability.

Alternative Identification Results

In the previous section, we demonstrated that the constant bias assumption may be violated in some empirical settings. As a result, the extrapolation formula in equation (ref), which relies on Assumption (ref), can yield biased estimates. This motivates the need for an alternative approach when the plausibility of the constant bias assumption is in doubt.

One potential strategy is to fully specify the structural functions and distributional assumptions and then estimate the model structurally to recover $\E{Y(0)|X}$.\footnote{See Todd_Wolpin:2023 for a review of approaches that integrate structural modeling with causal inference, particularly in the context of randomized controlled trials.} However, implementing such a strategy may require detailed individual-level data sufficient to identify underlying preference parameters and to serve as proxies for individual effort and belief---perhaps unavailable since many RD studies are observational and not designed to collect such granular information.

In light of these challenges, this section develops an alternative identification strategy that remains within a reduced-form framework but does not rely on the constant bias assumption. We begin by introducing a new set of assumptions and then derive identification results under these conditions. Our focus is on the two-cutoff case for expositional clarity, though the extension to multiple cutoffs is straightforward (see Remark (ref)). The extension to the one-sided fuzzy RD case is deferred to the Online Appendix.

Main Results

Assumptions and Identification Results

Our identification strategy is based on two empirically motivated assumptions. We begin with a commonly employed shape restriction:

assumption[Monotonicity] $\mu_{0,c}(x)$ is weakly increasing in $x\in(l,h)$ for $c\in\{l,h\}$.

This monotonicity assumption posits that the untreated potential outcome is a non-decreasing function of the running variable. Such a monotonicity assumption is standard in the partial identification literature (e.g., Manski:1997).

It is plausible in many empirical RD settings. As Babii_Kumar:2023 wrote, “[r]egression discontinuity designs encountered in empirical practice are frequently monotone." For example, when $X_i$ represents a test score and $Y_i(0)$ denotes future earnings in the absence of any treatment, it is natural to assume that $\mu_{0,c}(x)$ is increasing. A similar logic applies when $X_i$ is family income level, and higher-income families are associated with greater expected earnings due to inherited ability or increased investment in human capital (e.g., Bjorklund_etal:2006).

Second, we introduce an alternative restriction to the constant bias assumption, one that relates the untreated outcome functions across groups:

assumption[Dominance] $\mu_{0,l}(x) \leq \mu_{0,h}(x)$ holds on $x\in(l,h)$.

This dominance assumption assumes that the untreated conditional mean function of the lower cutoff group lies below that of the higher cutoff group. This is plausible in many multi-cutoff RD settings, especially when cutoffs are designed to reduce inequities or reflect pre-existing differences between groups.

For instance, scholarship thresholds are often relaxed for students from disadvantaged backgrounds---such as those from high-poverty regions---allowing them to qualify with lower test scores (Melguizo_etal:2016). Conversely, more academically prepared students may apply to competitive schools with higher cutoffs (Beuermann_etal:2022). In both examples, the cutoff reflects differences in group characteristics; that is, it is determined “endogenously." In such cases, the dominance assumption is more likely to hold.

Our main identification result is as follows:

theorem[Bounds on Extrapolated RD Effects] Under Assumptions (ref), (ref), and (ref), the treatment effect for group $C_i = l$ on $\bar{x}\in(l,h)$, $\tau_l(\bar{x})$, is bounded from below and above by \begin{align*} \rotatebox[origin=c]{180}{$\nabla$}_{l}(\bar{x}) = \mu_{1,l}(\bar{x}) - \mu_{0,h}(\bar{x}),\, and \, \nabla_{l}(\bar{x}) = \mu_{1,l}(\bar{x}) - \mu_{0,l}(l). \end{align*} These bounds $[\rotatebox[origin=c]{180}{$\nabla$}_{l}(\bar{x}), \nabla_{l}(\bar{x})]$ are pointwise sharp.
corollarySuppose the “reverse" of Assumptions (ref) and (ref) hold instead, that is, $\mu_{0,c}(x)$ is weakly decreasing and $\mu_{0,l}(x) \geq \mu_{0,h}(x)$ on $x\in(l,h)$. Then, the sharp bounds are given by $[\nabla_{l}(\bar{x}), \rotatebox[origin=c]{180}{$\nabla$}_{l}(\bar{x})]$.

The idea behind Theorem (ref) is illustrated in Figure (ref). By the monotonicity of $\mu_{0,l}(x)$, $\mu_{0,l}(\bar{x})$ can be bounded from below by $\mu_{0,l}(l)$. The dominance assumption ensures that it is bounded above by $\mu_{0,h}(\bar{x})$. Hence, we obtain that $\tau_l(\bar{x}) \in [\rotatebox[origin=c]{180}{$\nabla$}_{l}(\bar{x}), \nabla_{l}(\bar{x})]$. Monotonicity of $\mu_{0,h}(x)$ guarantees the sharpness of the bounds (see also Remark (ref) below).

figure[figure omitted — 165 chars of source]

The same reasoning applies across any points in $(l,h)$, leading to the following corollary:

corollaryTake an arbitrary closed interval $\mathcal{X}\subset(l,h)$. Then, under the same assumptions in Theorem (ref), the bounds $\rotatebox[origin=c]{180}{$\nabla$}_{l}(x)$ and $\nabla_{l}(x)$ are both attainable as a function over the interval $\mathcal{X}$, in the sense that $\rotatebox[origin=c]{180}{$\nabla$}_{l}(x)$ and $\nabla_{l}(x)$ are consistent with the observed data and maintained assumptions over $\mathcal{X}$.

This corollary states that $\rotatebox[origin=c]{180}{$\nabla$}_{l}(x)$ and $\nabla_{l}(x)$ provide the tightest lower and upper bounds on $\tau_l(x)$ as a function over a closed interval $\mathcal{X}$ within $(l,h)$. Unfortunately, the bounds $[\rotatebox[origin=c]{180}{$\nabla$}_{l}(x), \nabla_{l}(x)]$ are not uniformly sharp in general, that is, there exists a function $\delta(x)$ that is inconsistent with our maintained assumptions although $\delta(x) \in [\rotatebox[origin=c]{180}{$\nabla$}_{l}(x), \nabla_{l}(x)]$ over the interval $\mathcal{X}$. Such a counterexample can be easily constructed and is provided in the Online Appendix.

Nevertheless, we emphasize that both the lower and upper bounds are attainable and cannot be rejected as representing a true treatment effect function over $\mathcal{X}$. Thus, the practical implication that the true treatment effect function lies inbetween $\rotatebox[origin=c]{180}{$\nabla$}_{l}(x)$ and $\nabla_{l}(x)$ remains valid. In this view, the uniform (non)sharpness may be of limited practical consequence.

Remarks

We conclude this subsection with several important remarks:

remark[Flexibility of the Treatment Effect Function] Both of our main assumptions impose restrictions only on the control state, and no assumptions are made about the functional form under the treated status. Thus, the treatment effect function itself is left unrestricted, in line with Cattaneo_etal:2021JASA_extrapolating.
remark[(Non-)Sensitivity to Transformation] In RD studies, it is common practice to apply a monotonic transformation to the outcome variable---for example, log transformation of annual earnings as in Oreopoulos:2006. The constant bias assumption can be sensitive to such transformations. This type of sensitivity to transformations has been pointed out by Roth_SantAnna:2023 in the DID setting. Our identification conditions are invariant to monotonic transformations. This robustness makes the bounds especially valuable when the transformed outcome is of interest, but the plausibility of constant bias in the transformed space is unclear.
remark[Multiple Cutoff Points] Suppose there are $J+1(\geq3)$ cutoff points, denoted by ${c_0, c_1, \ldots, c_J}$. Without loss of generality, we focus on $\tau_{c_0}(x)$. A natural extension of the dominance assumption (Assumption (ref)) is a sequential dominance: $\mu_{0,c_0}(x) \leq \mu_{0,c_1}(x) \leq \cdots \leq \mu_{0,c_J}(x)$. Under this assumption, we can derive a sharp lower bound for $\tau_{c_0}(\bar{x})$ as $\mu_{1,c_0}(\bar{x}) - \mu_{0,c_K}(\bar{x})$ when $\bar{x}\in(c_{K-1}, c_K)$. The upper bound remains the same as in Theorem (ref). Hence, using the notation from the main text, the bounds over the interval $(l, h)$ are (weakly) tightened when there exists an “intermediate” group $m$ satisfying $l < m < h$. Estimation and inference proceed analogously to those described in the following section.
remark[Effect of Changing Threshold] $\tau_l(\bar{x})$ should be understood as the treatment effect in an environment where agents make decisions under cutoff $l$. Consequently, when an economist is interested in the effect of shifting the threshold from $l$ to $\bar{x}$, an additional assumption is required. To interpret $\tau_l(\bar{x})$ in this context, a policy invariance assumption, akin to the local policy invariance assumption in Dong_Lewbel:2015, becomes necessary. Under such a condition, the derived bounds characterize the effect of adjusting the threshold to $\bar{x}$.
remark[Limitation of the Multi-Cutoff RD Designs] Our bounds do not provide any information about $\tau_l(x)$ on $x<l$, which may also be of interest. This limitation mirrors the challenge discussed by Cattaneo_etal:2021JASA_extrapolating. Exploring identification strategies in this region is an important area for future work. We note, however, that the marginal extrapolation strategy proposed by Dong_Lewbel:2015 remains applicable in this setting.
remark[Testability and Falsification of Assumptions] Under Assumption (ref), the dominance assumption is refuted if $\lim_{x\uparrow l}\mu_{0,l}(x) > \lim_{x\uparrow l}\mu_{0,h}(x)$, which is directly testable. In general, the monotonicity of $\mu_{0,l}$ is not testable, while the monotonicity of $\mu_{0,h}$ is directly testable by, for example, Chetverikov:2019's (Chetverikov:2019) procedure. Technically, the falsification of the monotonicity of $\mu_{0,h}$ affects only the assertion of pointwise sharpness or the attainability of the lower bound, and does not immediately invalidate the bounds themselves. However, in many applications, the rejection of the monotonicity of $\mu_{0,h}$ may serve as indirect evidence for the rejection of the monotonicity of $\mu_{0,l}$, suggesting the potential invalidity of the upper bound. In such cases, researchers may focus on the lower bound $\rotatebox[origin=c]{180}{$\nabla$}(x)$, which requires only Assumption (ref). While this bound is only partially informative, it still offers valuable insight into the external validity of RD estimates.
remark[Sensitivity Analysis] In empirical studies, it is common practice to report layered estimates---such as point or partially identified intervals nested within one another---to assess the identifying power of different sets of assumptions and examine the sensitivity of the results (e.g., Kreider_etal:2012). The bounds obtained in Theorem (ref) well align with this purpose. Under Assumptions (ref) and (ref), it follows that $\tau_{l, \text{CB}}(x)\in[\rotatebox[origin=c]{180}{$\nabla$}_{l}(x), \nabla_{l}(x)]$. Thus, researchers can use our bounds to assess the identification power of the constant bias assumption, given that the shape restrictions hold. However, this nesting implies that the bounds cannot be used to test the constant bias assumption itself.

Estimation and Inference

Local Linear Estimation

To construct the bounds estimates, we need to estimate $\mu_{1,l}(x)$, $\mu_{0,h}(x)$, and $\mu_{0,l}(l)$. These quantities are all consistently estimable by standard nonparametric regression techniques. Following the recent RD literature, we employ the local linear smoothing with mean-squared error (MSE) optimal bandwidth selector (see Fan_Gijbels:1996 for a comprehensive review). Specifically, we estimate each function ${\mu}_{d,c}(x)$ by $\widehat{\mu}_{d,c}(x) \coloneqq (1,0) \widehat{\beta}_{d,c}(x)$, where

align[align omitted — 241 chars of source]

where $K$ is the kernel function and $b$ is the bandwidth. Note that we only use the observations with $D_i=d$ and $C_i=c$ to estimate $\mu_{d,c}(x)$. Note also that $K$ and $b$ can differ in each estimation. One can estimate the bounds by $\widehat{\rotatebox[origin=c]{180}{$\nabla$}}_{l}(x) = \widehat{\mu}_{1,l}(x) - \widehat{\mu}_{0,h}(x)$ and $\widehat{\nabla}_{l}(x) = \widehat{\mu}_{1,l}(x) - \widehat{\mu}_{0,l}(l)$.

Pointwise Inference

In some applications, researchers may have a specific point of interest, $\bar{x}$. In this case, pointwise uncertainty quantification will be useful.

To address the bias introduced by kernel smoothing, we employ the robust bias-corrected inference procedure developed by Calonico_etal:2014, calonico2018effect, calonico2022coverage. This method corrects for the leading-order smoothing bias and adjusts the variance induced by this bias correction, thereby enabling valid inference under MSE-optimal bandwidth selection.

Let $B_{d,c}(x)$ denote the asymptotic smoothing bias of $\widehat{\mu}_{d,c}(x)$ due to the local linear regression (ref) and $\widehat{B}_{d,c}(x)$ be its estimator, typically computed via local quadratic regression using a bandwidth $b_\text{bias}$. A common and practical choice for $b_\text{bias}$ is to set $b_\text{bias}=b$ (calonico2018effect; Calonico_stata:2019), which we adopt hereafter.

Define the bias corrected estimator $\widehat{\mu}_{d,c}^{\text{BC}}(x)=\widehat{\mu}_{d,c}(x)-\widehat{B}_{d,c}(x)$, and let $\mathcal{V}_{d,c}(x) =\V{\widehat{\mu}_{d,c}(x)-\widehat{B}_{d,c}(x) | X_1,\ldots,X_n}$, with $\widehat{\mathcal{V}}_{d,c}(x)$ as its estimator. Under standard smoothness and regularity conditions (see Calonico_etal:2014, calonico2018effect), we have that

align*[align* omitted — 145 chars of source]

where $\to_d$ denotes the convergence in distribution. Put $\widehat{\rotatebox[origin=c]{180}{$\nabla$}}_{l}^{\text{BC}}(\bar{x}) = \widehat{\mu}_{1,l}^{\text{BC}}(\bar{x}) - \widehat{\mu}_{0,h}^{\text{BC}}(\bar{x})$ and $\widehat{\nabla}_{l}^{\text{BC}}(\bar{x}) = \widehat{\mu}_{1,l}^{\text{BC}}(\bar{x}) - \widehat{\mu}_{0,l}^{\text{BC}}(l)$. Then, since $\widehat{\mu}_{1,l}^{\text{BC}}(\bar{x})$, $\widehat{\mu}_{0,h}^{\text{BC}}(\bar{x})$, and $\widehat{\mu}_{0,l}^{\text{BC}}(l)$ are independent by the assumption of random sampling, $\widehat{\rotatebox[origin=c]{180}{$\nabla$}}_{l}^{\text{BC}}(\bar{x})$ and $\widehat{\nabla}_{l}^{\text{BC}}(\bar{x})$ are asymptotically normal with the asymptotic variances $\widehat{\mathcal{V}}_{L} = \widehat{\mathcal{V}}_{1,l}(\bar{x}) + \widehat{\mathcal{V}}_{0,h}(\bar{x})$, and $\widehat{\mathcal{V}}_{U} = \widehat{\mathcal{V}}_{1,l}(\bar{x}) + \widehat{\mathcal{V}}_{0,l}(l)$. Using them, we can draw confidence intervals (CIs) for each lower and upper bound in the usual manner. One can also construct the CI that covers $[\rotatebox[origin=c]{180}{$\nabla$}_{l}(\bar{x}), \nabla_{l}(\bar{x})]$, or the CI for $\tau_l(\bar{x})$ using the method of Imbens_Manski:2004 and Stoye:2009.

Uniform Confidence Band

In other applications, researchers are interested in a range of values for the running variable rather than a specific point. In such cases, constructing a uniform confidence band over the region of interest provides a useful tool for inference. This subsection proposes a procedure for uniform inference based on the multiplier bootstrap, following the ideas of Fan_etal:2022 and imai2025. In the Online Appendix, we establish the asymptotic validity of the proposed procedure by combining the results of cck2013, cck_anti, cck14.

Let $\mathcal{I} \subset (l,h)$ be a closed interval of interest. We proceed with the following steps:

itemize• Obtain $\widehat{\mu}_{1,l}^{\text{BC}}(x)$ and $\widehat{\mu}_{0,h}^{\text{BC}}(x)$ using local linear regressions with integrated MSE (IMSE) optimal bandwidths $b_{1,l}$ and $b_{0,h}$ (e.g., Calonico_stata:2019). Construct $\widehat{\mu}_{0,l}^{\text{BC}}(l)$ similarly but with an MSE-optimal bandwidth $b_{0,l}$. For all components, we apply bias correction using the same bandwidths for both the main and bias estimation: $b_{\text{bias},d,c} = b_{d,c}$, and we employ the same kernel function throughout. • Choose a large number of bootstrap replications $M$ (e.g., $M=1000$). For each $m=1,\ldots,M$, draw an i.i.d. random variable $\{\xi_i^m\}_{i=1}^{n}$ from Mammen:1993's (Mammen:1993) two-point distribution.\footnote{Mammen:1993's two-point distribution is defined as $\xi_i = (1-\sqrt{5})/2$ with probability $(1+\sqrt{5})/(2\sqrt{5})$ and $\xi_i = (\sqrt{5}+1)/2$ with probability $(\sqrt{5}-1)/(2\sqrt{5})$. Theoretically, the Gaussian multiplier can be used, but in small samples, the Gaussian multiplier may occasionally cause the matrix $\bm{R}^\top \text{diag}((\xi_i^m +1)K((X_i-x)/b_{d,c})) \bm{R}$ to become (nearly) singular due to the possible negativity of the weights $\xi_i^m+1$. To avoid this, we recommend using Mammen:1993's weights, guaranteeing $\xi_i^m+1>0$, and the computation becomes stabler.} Compute the local quadratic regression estimators, $\widehat{\mu}_{d,c}^{\star m}(x) = (1,0,0)\widehat{\beta}_{d,c}^{\star m}(x)$, where $\widehat{\beta}_{d,c}^{\star m}(x)$'s are defined as \begin{align*} \mathop{\rm arg\,min}\limits_{(b_0,b_1,b_2)^\top\in\mathbb{R}^3} \sum_{i: D_i=d, C_i=c} (\xi_i^m+1)\left\{Y_i - b_0 - b_1(X_i - x) - b_2(X_i - x)^2\right\}^2 K\left(\frac{X_i - x}{b_{d,c}}\right). \end{align*} Note that the bandwidths $b_{d,c}$ are the same as the ones used in step 1 in every iteration. Define $\widehat{\rotatebox[origin=c]{180}{$\nabla$}}_{l}^{\star m}(x) = \widehat{\mu}_{1,l}^{\star m}(x) - \widehat{\mu}_{0,h}^{\star m}(x)$ and $\widehat{\nabla}_{l}^{\star m}(x) = \widehat{\mu}_{1,l}^{\star m}(x) - \widehat{\mu}_{0,l}^{\star m}(l)$. • For each replication $m$, calculate the studentized maximum deviations: \begin{align*} S^{\star}_L(m) = \sup_{x\in\mathcal{I}} \frac{\widehat{\rotatebox[origin=c]{180}{$\nabla$}}_{l}^{\star m}(x) - \widehat{\rotatebox[origin=c]{180}{$\nabla$}}_{l}^{BC}(x)}{\widehat{\mathcal{V}}^{1/2}_L(x)}\,\, and \,\, S^{\star}_U(m) = \sup_{x\in\mathcal{I}} \frac{\widehat{\nabla}_{l}^{\star m}(x) - \widehat{\nabla}_{l}^{BC}(x)}{\widehat{\mathcal{V}}^{1/2}_U(x)}. \end{align*} In practice, the supremum is approximated by the maximum over some fine grid points. Given a confidence level $1-\alpha$, compute the critical values \begin{align*} c^\star_{V}(1-\alpha/2) \coloneqq the (1-\alpha/2) -quantile of \{S^{\star}_V(m): m=1,\ldots,M\},\,\,V\in\{L,U\}. \end{align*} • Construct a confidence band \begin{align*} \widehat{\mathcal{C}}(x) = \left[ \widehat{\rotatebox[origin=c]{180}{$\nabla$}}_{l}^{\text{BC}}(x) - c^\star_L(1-\alpha/2) \widehat{\mathcal{V}}_L^{-1/2}(x), \widehat{\nabla}_{l}^{\text{BC}}(x) + c^\star_U(1-\alpha/2) \widehat{\mathcal{V}}_U^{-1/2}(x) \right]. \end{align*} Then, under assumptions made in Appendix (ref), it holds that \begin{align*} \lim_{n\to\infty}\P{\left[{\rotatebox[origin=c]{180}{$\nabla$}}_{l}(x), {\nabla}_{l}(x)\right] \subseteq \widehat{\mathcal{C}}(x) \text{ for all } x\in\mathcal{I}} \geq 1-\alpha. \end{align*} Trivially, it also holds that $\lim_{n\to\infty}\P{\tau_l(x) \in \widehat{\mathcal{C}}(x) \text{ for all } x\in\mathcal{I}} \geq 1-\alpha$.

Empirical Illustrations

In this section, we present two empirical applications to demonstrate the potential usefulness of the bounds derived in the previous section.

SPP Program (Non-Manipulable Running Variable Case)

We begin with a financial aid program setting with a non-manipulable running variable, originally studied by Londono-Velez_etal:2020aejep. Our primary focus is on the intent-to-treat (ITT) effects; we defer issues related to the incomplete compliance to the Online Appendix.

Empirical Context

We investigate the effect of Ser Pilo Paga (SPP), a financial aid program introduced in Colombia in 2014. Eligibility for full scholarship loans from SPP is determined by a combination of merit- and need-based criteria, using national test scores and a family wealth index as running variables.

Londono-Velez_etal:2020aejep exploited these eligibility rules to estimate standard RD effects on higher education enrollment. Following their setup, we focus on merit-eligible students—those who meet the minimum test score threshold---and estimate what they term “frontier-specific" RD effects on enrollment rates. Notably, the merit-based cutoff is constant across all students, while the need-based threshold varies across geographic areas. This creates a multi-cutoff RD setting.

In this setup, the running variable is the wealth index, which ranges from $0$ (poorer) to $100$ (wealthier). For consistency with the stylized RD settings, we multiply the wealth index by $-1$, so that larger values correspond to poorer households, with $-100$ representing the wealthiest. Under this transformation, students are eligible for aid if they come from households whose index exceeds a given threshold (i.e., sufficiently poor households). The cutoff for students from metropolitan areas is $-57.21$, while that for rural areas is $-40.75$, resulting in a multi-cutoff RD setting.

Discussion of Assumptions

In our analysis, the outcome variable is the enrollment rate, and the running variable is the family wealth level. Given this context, the “reverse" versions of Assumptions (ref) and (ref) appear plausible in our context. The monotonicity assumption posits that $\mu_{0,c}$ is decreasing, meaning that the probability of enrollment declines as family wealth decreases. This aligns with economic intuition and empirical patterns. The dominance assumption (in reverse) posits that the regression function for students from rural areas (who face a higher cutoff) lies below that for students from metropolitan areas. This is a natural assumption, as students in rural areas may be more disadvantaged in terms of access to educational resources and may also receive less parental support or guidance regarding the benefits of pursuing higher education.

In the present context, where the running variable is non-manipulable, concerns about the constant-bias assumption may be less severe than in the partially manipulable case, although certain issues remain---for instance, differences in access to infrastructure or social support between urban and rural students. Focusing exclusively on merit-eligible students, however, may help ensure more comparable unobserved characteristics across groups, akin to conditioning on covariates.

In the next subsection, we compute the bounds derived in Corollary (ref) and evaluate potential deviations from constancy by comparing them to the extrapolated RD effects under the constant bias assumption.

Results and Implications

figure[figure omitted — 489 chars of source]

We consider the closed interval $\mathcal{I}=[-55.0,-42.5]\subset(-57.21, -40.75)$ in our analysis. Figure (ref) presents the estimated regression functions, and Figure (ref) displays the estimated bounds over $\mathcal{I}$, along with the extrapolated RD effect function under the constant bias assumption (blue dot-dashed line). The pink dashed line indicates the level of the treatment effect at the cutoff, $-57.21$.

Several empirically important findings emerge. First, the estimated bounds indicate positive treatment effects throughout the interval $\mathcal{I}$. These bounds are sufficiently narrow to draw meaningful policy implications, suggesting that the effect lies approximately between 0.20 and 0.25. Second, the horizontal line indicating the RD effect at the cutoff consistently falls within the (tight) bounds, implying that the hypothesis of constant average treatment effects is not rejected. These observations offer strong support for the external validity of the standard RD estimate. Furthermore, the finding that similarly sized effects persist across other points in $\mathcal{I}$ provides useful guidance for future policy adjustments, such as modifying eligibility thresholds.

We also confirm that the extrapolation under the constant bias assumption performs well, although the upper bound suggests that the true effect may be slightly larger. Overall, the tightness of the bounds indicates that any bias from deviations from the constant bias assumption is likely limited and not practically severe.

ACCES Program (Partially Manipulable Running Variable Case)

We now turn to a different context involving a financial aid program, in which the running variable is partially manipulable.

Empirical Context and Background

We revisit the empirical analysis of Cattaneo_etal:2021JASA_extrapolating. They investigated the extrapolated effect of the Acceso con Calidad a la Educación Superior (ACCES) program---a national merit-based financial aid initiative---on higher education enrollment among Colombian students, originally studied by Melguizo_etal:2016.

Eligibility for ACCES requires students to score above a specific threshold on the national high school exit exam (SABER 11). The score ranges from 1 (best) to 1000 (worst), and the eligibility cutoff was fixed at 850 prior to 2008. Beginning in 2009, however, region-specific cutoffs were introduced, resulting in a multi-cutoff RD setting. For consistency, we multiply the scores by $-1$, so that higher achievement corresponds to larger values.

Cattaneo_etal:2021JASA_extrapolating leveraged this design to estimate extrapolated RD effects on college enrollment, focusing on two cohorts: students who applied between 2000 and 2008 (cutoff $-850$) and those who applied between 2009 and 2010 in a region with a cutoff of $-571$. They reported that the extrapolated RD effect at $-650$ was $0.191$, larger than the local RD estimate at $-850$, which was $0.137$.

Discussion on Assumptions

The plausibility of the constant bias assumption in this context is debatable. First, the running variable is partially manipulable, suggesting that the constant bias assumption is not guaranteed even when two groups are similar (Section (ref)). Second, the 2009 reform introduced cutoffs in a progressive manner: regions with greater disadvantage received lower thresholds, while more advantaged areas were subject to higher ones. Notably, the region with a cutoff of $-571$ is among the most advantaged, with a very low share of students from low socioeconomic backgrounds Melguizo_etal:2016. As a result, the two groups under comparison may differ substantially in unobservable characteristics.

Another potential concern regarding the constant bias assumption arises from how it was assessed in Cattaneo_etal:2021JASA_extrapolating. In their analysis, the authors fitted separate quadratic regression functions on the left side of the lower cutoff at $-850$, and compared them. Building on this approach, we can also extrapolate the fitted quadratic model to the right of the cutoff, as depicted by the red dotted line in Figure (ref). For comparison, the extrapolation based on the constant bias assumption is shown as the pink dot-dashed line. The two curves diverge notably just above the cutoff, especially in their slopes, suggesting potential violations of the constant bias assumption. Furthermore, the parametric extrapolation implies a sizable negative treatment effect farther from the cutoff, yielding markedly different conclusions depending on the extrapolation method employed.

In this application, the monotonicity and dominance assumptions appear reasonable: students with higher SABER 11 scores are more likely to enroll in college, and those from more advantaged regions (i.e., $C_i = h$) are expected to have higher average enrollment rates than those from disadvantaged regions.

figure[figure omitted — 497 chars of source]

Results and Implications

Figure (ref) displays the estimation results over the interval $\mathcal{I}=[-840,-590]$. The bounds suggest that the ACCES program likely has a nonnegative effect overall, and the effects are roughly inbetween $[0.00,0.15]$ just above the cutoff and about $[0.10,0.25]$ far away from the cutoff.

While the bounds are admittedly wide, they nonetheless deliver meaningful information about treatment effects away from the cutoff. Importantly, they rule out large negative effects implied by the parametric extrapolation, which lie well outside the identified set. This provides reassurance that strong pessimistic conclusions are inconsistent with the maintained assumptions. At the same time, it remains possible that the true effect is smaller than what is implied by extrapolation under the constant bias assumption---for instance, closer in magnitude to the effect estimated at the cutoff. In the absence of additional institutional knowledge, this possibility should also be borne in mind when evaluating the ACCES program.

Conclusion

This paper explored when and how the treatment effect in RD designs can be extrapolated away from cutoff points in multi-cutoff settings. We began by examining the plausibility of the constant bias assumption proposed by Cattaneo_etal:2021JASA_extrapolating, interpreting it through the lens of rational agent behavior. We found that the assumption is indeed plausible when the two groups are composed of agents with comparable characteristics and when the running variable is non-manipulable. However, this justification may fail when the running variable is partially manipulable by the agent, potentially resulting in biased estimates.

To address this issue, we proposed an alternative identification strategy grounded in empirically motivated assumptions---monotonicity and dominance---which do not require the constant bias assumption. We derived sharp bounds on extrapolated treatment effects under these assumptions and established a uniform inference procedure. Our empirical applications illustrated the potential usefulness of these bounds.