Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
79,173 characters · 19 sections · 63 citation commands
Safe Policy Learning under Regression Discontinuity Designs with Multiple Cutoffs
\etocdepthtag.toc{mtchapter} \etocsettagdepth{mtchapter}{subsection} \etocsettagdepth{mtappendix}{none}
\singlespacing
The regression discontinuity (RD) design is widely used in causal inference and program evaluation with observational data thistlethwaite1960regression,lee:more:butl:04,eggers_hainmueller_2009,card2009does,lee2010regression. A primary reason for this popularity is that the RD design can provide credible causal inference without assuming that the treatment is unconfounded. Under the RD design, one can exploit a known deterministic treatment assignment mechanism and identify the average treatment effect at the treatment assignment cutoff. On the whole, the causal inference literature has focused on the development of rigorous estimation and inference methodology for this RD treatment effect calonico2014robust, calonico2018effect, cheng1997automatic, armstrong2018optimal, hahn2001identification, imbens2008regression, kolesar2018inference,imbens2019optimized.
While treatment effect estimation at the existing cutoff is useful, policy makers may be interested in improving the outcome by considering alternative treatment assignment cutoffs heckman2005structural. Despite the widespread use of the RD design, there has been limited research on policy learning under this design. We address this gap in the literature by developing a methodology for learning new and better cutoffs from observed data. Because the treatment assignment mechanism is deterministic under the RD design, learning a new policy requires extrapolation. Building on the robust optimization framework proposed by ben2021safe, we develop a safe policy learning methodology under the RD design that leverages the existence of multiple cutoffs. Specifically, we provide a theoretical safety guarantee that a new, learned policy will not yield a worse overall utility than the existing status quo policy.
In many applications of the RD designs, the treatment assignment cutoff varies across subpopulations based on certain criteria including geographical regions, time periods, and age groups chay2005central,klavsnja2017incumbency. We show how to exploit the existence of such multiple cutoffs to make the required extrapolation more credible. Specifically, we compare the conditional potential outcome functions across different sub-populations and assume that the cross-group difference is a smooth function of the running variable with a bounded slope in the extrapolation region. Based on this smoothness assumption, we separate the expected utility into point-identified and partially-identified components, and propose a doubly-robust estimator for the point-identifiable component. We then minimize the worst case utility loss relative to the status quo policy. The resulting policy, based on new treatment cutoffs, has a safety guarantee that they will not yield a worse overall utility than the existing cutoffs. We also establish the asymptotic regret bounds for the learned safe policy using semiparametric efficiency theory.
We apply the proposed methodology to a recent study of the ACCES loan program in Colombia melguizo2016credit. In this application, geographical departments in Colombia used different cutoffs, which are based on test scores, to determine the eligibility for financial aid to low-income students. We use the proposed methodology to learn a cutoff for each department that improves the overall enrollment in post-secondary institution throughout the country.
\paragraph{Related Work.} There exist few prior studies that consider policy learning under the RD design. As noted above, many existing studies focus on the identification, estimation, and inference of the average treatment effects at cutoffs. Recently, some scholars have developed extrapolation methods to improve the external validity of the RD design mealli2012evaluating,dong2015identifying,angrist2015wanna,cattaneo2020extrapolating,dowd2021donuts,bertanha2020external. For example, dong2015identifying consider how an infinitesimal change of the cutoff alters the RD treatment effect by estimating its derivative at the existing cutoff. eckles2020noise propose a different randomization-based framework that exploits measurement error in the running variable and examine how a change in the cutoff value affects the observed outcome. Although the methods proposed in these studies are related to ours, they do not explicitly address the problem of policy learning under the RD design.
An important exception is yata2021optimal, who focuses on choosing between two specific policies. This method, however, derives a finite-sample, exact minimax regret decision rule for choosing between the two policies, which may potentially results in a randomized decision. Additionally, their approach requires either knowledge of the distribution of the running variable or a fixed-design evaluation using the empirical measure to define the expected utility. In contrast, our approach seeks the optimal deterministic policy across a broad class of policies under a super-population framework.
Our proposed methodology leverages the existence of multiple cutoffs. In most applied studies with multi-cutoff RD designs, however, researchers simply pool sub-populations and normalize the running variable such that the standard single-cutoff RD methodology can be applied. An important exception is cattaneo2020extrapolating, who propose a methodology to extrapolate treatment effects under the RD design with two cutoffs. The authors rely on a parallel-trend assumption that the difference in the conditional expectation of potential outcome between the two groups is constant in the running variable. We generalize and weaken their assumption by allowing the the cross-group difference to change smoothly in the running variable. In addition, while cattaneo2020extrapolating focuses on the treatment effect estimation with two cutoffs, we focus on policy learning under multi-cutoff RD designs.
We also contribute to a growing literature on policy learning with observational data. We found no prior work in this literature that studies the RD design. Indeed, most existing work relies upon inverse probability weighting (IPW) swaminathan2015counterfactual,qian2011performance,zhao2012estimating,Kitagawa2018 or doubly-robust augmented IPW (AIPW) athey2021policy to estimate and optimize the overall expected utility. These weighting-based approaches require a sufficient degree of overlap between the baseline policy and all candidate policies within a policy class. jin2022policy consider a scenario of limited overlap and propose a pessimism-based weighted learning approach that only requires the overlap between the baseline and optimal policy.
In contrast, we address the policy learning problem with the complete lack of overlap that arises from the deterministic treatment assignment mechanism under RD designs. Specifically, we build on the robust optimization framework of ben2021safe to develop a safe policy learning methodology tailored to RD designs. While ben2021safe focused on simple extrapolation assumptions within a specific application, we incorporate the unique features of RD designs, such as cross-group smoothness and doubly robust construction, making it applicable to a wide range of real-world RD settings.
Related to our approach, some researchers also have used partial identification for policy evaluation and learning under different non-standard settings cui2021semiparametric,han2019optimal,khan2023off, while others have applied robust optimization to causal effect estimation and sensitivity analysis kallus2018confounding,pu2021estimating,gupta2020maximizing,bertsimas2023distributionally.
\paragraph{Organization of the paper.} The paper is organized as follows. Section (ref) defines the policy learning problem in RD designs with multiple cutoffs. Section (ref) presents the proposed methodological framework and our smoothness assumption for partial identification. Section (ref) discusses the empirical policy learning problem, where we propose a doubly-robust estimator. Section (ref) establishes asymptotic bounds on the utility loss of our learned policy relative to the baseline policy and the oracle optimal-in-class policy. In Section (ref), we apply the proposed methodology to learning new cutoffs for the ACCES loan program in Colombia. Lastly, Section (ref) concludes.
We begin by describing the policy learning problem under the general multi-cutoff RD design and introducing the notation used throughout the paper.
In our setting, for each unit $i=1,\ldots,n$, we observe a univariate running variable $X_i \in \mathcal{X} \subset \ensuremath{\mathbb{R}}$, a binary treatment assignment $W_i \in \{0,1\}$, and an outcome variable $Y_i \in \ensuremath{\mathbb{R}}$. We assume that the running variable $X$ has a continuous positive density $f_X$ on the support $\mathcal{X}$. Each unit $i$ belongs to one of $Q$ different subgroups, denoted by $G_i \in \mathcal{Q}:=\{1,\ldots,Q\}$, where each group $g$ has a corresponding cutoff value $\tilde{c}_g$. These existing cutoff values may differ between groups. For simplicity, we assume there are no ties, but note that our proposed methodology can still be applied in situations where tied cutoffs may exist. We also assume that these cutoffs have been sorted in an increasing order (i.e., $\tilde c_1 < \tilde c_2 <\ldots < \tilde{c}_Q$). Note that both $X$ and $G$ are pre-treatment variables, which are neither affected by the cutoffs nor the treatment variable.
Throughout the paper, we consider the sharp RD setting, where given a set of known cutoff values $\{\tilde{c}_{g}\}_{g=1}^Q$, the treatment assignment $W_i$ is completely determined by both the running variable $X_i$ and the group indicator $G_i$. Specifically, unit $i$ in group $g$ is assigned to the treatment condition when the value of the running variable $X_i$ exceeds a fixed, known treatment cutoff $\tilde{c}_g$, i.e., $$W_i = \sum_{g=1}^{Q} \mathds{1}(G_i =g)\mathds{1}(X_i \geq \tilde c_g)$$ where $\mathds{1}\left(\cdot\right)$ denoting the indicator function. Following the potential outcome framework neyman1923, rubin1974, we define $Y_i(1)$ and $Y_i(0)$ as the potential outcomes under the treatment and control conditions, respectively. The observed outcome then is given by $Y_{i}=W_{i} Y_{i}(1)+\left(1-W_{i}\right) Y_{i}(0)$. We further assume that the tuples $\{X_i, W_i,G_i, Y_i(1),Y_i(0)\}$ are drawn i.i.d from an overall target population $\mathcal{P}$. For notational convenience, we will sometimes drop the unit subscript $i$ from random variables.
Much of the RD design literature has focused on the evaluation of the conditional average treatment effect (CATE) at the cutoff. Under the multi-cutoff RD design, we define the group-specific CATE for each group $g$ at the cutoff value $\tilde{c}_g$,
The key identification assumption for $\tau_g(\tilde{c}_g)$ under the sharp RD design is the local continuity of the conditional expectation function of the potential outcome hahn2001identification.
Under this assumption, we can identify $\tau_g(\tilde{c}_g)$ via, $$\tau_g(\tilde{c}_g) \ = \ \lim _{x \downarrow \tilde{c}_g} \mathbb{E}[Y_i \mid X_i=x,G_i=g]-\lim _{x \uparrow \tilde{c}_g} \mathbb{E}[Y_i \mid X_i=x,G_i=g].$$
Under Assumption (ref) alone, however, the deterministic treatment assignment makes it impossible to identify the average treatment effect beyond the existing cutoff hahn2001identification,imbens2008regression. In the literature, researchers have proposed additional assumptions to extrapolate beyond the existing cutoffs. For example, under the single-cutoff RD setting, dowd2021donuts assume the bounded partial derivative of the conditional expectation function with respect to the running variable. Under the fuzzy RD setting, bertanha2020regression restricts heterogeneity across compliance classes. Lastly, under the multi-cutoff RD design, cattaneo2020extrapolating propose a “common-trend” assumption. Building on these recent advancements, we show how to learn new and improved cutoffs from observed data under multi-cutoff RD design.
We introduce our policy learning problem under the multi-cutoff RD design. Throughout, we assume that potential outcomes depend on cutoffs only through the treatment assignment, which is implicit in our notation dong2015identifying,bertanha2020regression. In other words, changing the cutoff value does not alter outcomes unless it leads to a change in the treatment condition. The assumption is violated, for example, if individuals change their behavior in response to a change in the cutoff value under the same treatment condition.
Under the multi-cutoff RD design, a policy $\pi$ is a deterministic function that maps the running variable and group indicator to a binary treatment assignment variable, $\pi:\mathcal{X} \times \mathcal{Q} \rightarrow \{0,1\}$. Let $u(y,w)$ denote the utility for outcome $y$ under treatment $w$. Thus, the utility for unit $i$ with treatment $w$ and potential outcome $Y_i(w)$ is given by $u(Y_i(w),w)$. Our goal is to find a new policy $\pi$ that has a high average value (expected utility) for the overall population. For the ease of exposition, we use the observed outcome as the utility, i.e., $u(y, w)=y$, but our discussion below readily extends to other utility functions. In Section (ref), we consider an adjusted utility function that takes into account the cost of treatment.
For any given policy $\pi$, we define its overall value as the expected utility in the target population $\mathcal{P}$:
where the expectation is taken with respect to the joint distribution of the running variable $X$, the group indicator $G$, and potential outcomes $Y(w)$. The goal of policy learning is to find the best possible policy within a pre-specified policy class $\Pi$, i.e., $$\pi^{*} \ = \ \underset{\pi \in \Pi}{\operatorname{argmax}}\ V(\pi).$$
Under the multi-cutoff RD design, we first consider the following class of (group-specific) linear threshold policies:
where each group-specific cutoff value $c_g$ corresponds to a candidate policy for group $g$. We note that the baseline cutoff value $\tilde{c}_g$ may be different from ${c}_g$, but both belong to this policy class. Also, by definition, each group's policy (cutoff) in $\Pi_{\text{Multi}}$ is determined independently of the others' choices, meaning no joint constraints are placed on $\{c_g\}_{g=1}^Q$.
Learning an optimal policy requires the evaluation of $V(\pi)$ for any given policy $\pi \in \Pi_{\text{Multi}}$. Under RD designs, however, the average potential outcomes under cutoffs that are different from the existing ones are not identifiable. Since the treatment assignment is deterministic rather than stochastic, the usual overlap assumption, i.e., $0 < \Pr(W_i = 1 \mid X_i = x) < 1 \ \text{for all } x \in \mathcal{X}$, is violated under the RD designs. As a result, we cannot use standard policy learning methodology based on the inverse probability-of-treatment weighting (IPW) or its variants zhao2012estimating, Kitagawa2018,athey2021policy. Therefore, to evaluate $V(\pi)$, we must extrapolate beyond the existing threshold $\tilde{c}_g$ and infer the average unobserved potential outcomes.
We now outline our basic framework of policy learning under the multi-cutoff RD designs. The proposed framework has three steps. We first decompose the expected utility for any given policy into identifiable and unidentifiable components. Second, we partially identify the unidentifiable component using a relaxed version of the “parallel-trends” assumption from the difference-in-differences literature rambachan2023more. Lastly, we propose a robust optimization approach that maximizes the worst case expected utility. We show that the resulting maximin value optimal policy has a safety guarantee; the value of deploying the learned policy is no worse than the baseline policy under the assumption used for partial identification.
We begin by defining the baseline policy that generates the observed data as $$\tilde\pi(x,g) \ := \ \mathds{1}(x\geq \tilde c_g).$$ Note that this baseline policy $\tilde\pi$ is contained in the linear-threshold policy class $\Pi_{\text{Multi}}$ defined in Equation (ref). By comparing the deviation of a given policy $\pi$ from the baseline policy $\tilde\pi$, we decompose the expected utility in Equation (ref) into point-identifiable and non-identifiable terms, denoted by $\mathcal{I}_{\text{iden}}$ and $\mathcal{I}_{\text{uniden}}$, respectively:
where $m(w,x,g)$ is the group-specific conditional potential outcome function.
The first term in Equation (ref) corresponds to the cases where the policy $\pi$ agrees with the baseline policy $\tilde\pi$. This component can be point-identified from the observable data by replacing the counterfactual outcome $Y(\pi(X,G))$ with the observed outcome $Y$. The second and third terms, however, are not point-identifiable without further assumptions; the counterfactual outcome cannot be observed when the policy under consideration $\pi$ disagrees with the baseline policy $\tilde\pi$.
We consider the following restricted set of group-specific linear threshold policies,
Like the unrestricted policy class $\Pi_{\text{Multi}}$ in Equation (ref), this policy class contains the baseline policy $\tilde{\pi}(x,g)$. In this policy class, however, the new cutoffs are found between the lowest and highest cutoffs, denoted by $\tilde{c}_1$ and $\tilde{c}_Q$, respectively. The reason for this restriction is that the observed data provide no information to leverage beyond this range. For example, above the highest cutoff $\tilde{c}_Q$, all units receive the treatment and no outcome is observed under the control condition.
Under this restricted policy class $\Pi_{\text{Overlap}}$, the unidentifiable term $\mathcal{I}_{\text{uniden}}$ can be further decomposed into the conditional mean function of the observed outcome, which is point-identifiable, and an unidentifiable “difference function,”
The difference function $d(w,x,g,g^\prime)$ represents the difference in the conditional mean function of potential outcomes between two groups at the same value of the running variable.
Note that if $d(w,x,g,g^\prime)=0$, then the conditional average treatment effect is identical between groups $g$ and $g^\prime$ given the running variable. This type of assumption is often invoked to combine treatment effect estimates from multiple studies stuart2011use,dahabreh2019generalizing, but it is typically not credible here because the running variable is the only conditioning variable. In Section (ref), we will instead partially identify the difference function by requiring a weaker assumption that the difference function is a smooth function of the running variable.
We now formally state our decomposition result, showing that we can use the group whose existing cutoff is closest to the counterfactual cutoff of another group for extrapolation.
Figure (ref) illustrates the decomposition using a three-group case and focusing on the outcome under the treatment. Suppose that we wish to learn a new cutoff for Group 3 whose existing cutoff is $\tilde{c}_3$. To do this, we must extrapolate the conditional mean outcome $m(1,x,3)$ for this group in the region below the existing cutoff (represented by the dashed blue line). To evaluate the counterfactual value of the conditional mean outcome at $x^{\prime\prime} \in [\tilde{c}_2, \tilde{c}_3]$ for Group 3, we choose Group 2 whose existing cutoff is the closest to $\tilde{c}_3$. We then decompose the counterfactual value $m(1,x^{\prime\prime},3)$ into the identifiable conditional mean for Group 2, i.e., $m(1,x^{\prime\prime},2)$, and the difference in the conditional mean function between Groups 3 and 2, i.e., $d(1,x^{\prime\prime},3,2)=m(1,x^{\prime\prime},3)-m(1,x^{\prime\prime},2)$. Similarly, to infer $m(1, x^{\prime},3)$, we use Group 1 as a reference group and decompose this counterfactual value into the identifiable conditional mean value for Group 1, i.e., $m(1, x^\prime, 1)$, and the difference between Groups 3 and 1, i.e., $d(1,x^\prime,3,1)$.
In sum, Proposition (ref) implies the following decomposition of the expected utility in Equation (ref):
For any given policy $\pi\in\Pi_{\text{Overlap}}$, when the baseline policy $\pi$ and a candidate policy $\tilde\pi$ agree, the expected utility equals the identifiable component $\mathcal{I}_{\text{iden}}$. When they disagree, the expected utility equals the sum of two components: the identifiable term $\Xi_{1}$, and the partially identifiable term $\Xi_{2}$.
Using the decomposition above, we propose a robust optimization approach to policy learning. Specifically, we first partially identify the difference function $\Xi_2$ and then use robust optimization to derive an optimal policy by maximizing the worst case value.
As illustrated in Figure (ref) and discussed above, the cross-group difference in the conditional outcome needed for extrapolation is not identifiable. For example, at $x^{\prime\prime}$, this difference function between Groups 3 and 2 cannot be point-identified. Yet, the same difference function is identifiable in the region above the existing cutoff of Group 3, i.e., $x \ge \tilde{c}_3$. We leverage this fact for extrapolation under the following assumption that the cross-group difference changes smoothly as a function of the running variable.
In the function class $\mathcal{F}$, the group index ${g^\ast}$ depends on the treatment condition. For example, in the case of the conditional outcome model under the treatment, ${g^\ast}$ represents the group with the higher cutoff, whereas under the control condition it corresponds to the group with the lower cutoff. Thus, $\tilde c_{g^\ast}$ equals the boundary point at which the difference model can be identified.
Assumption (ref) states that the cross-group heterogeneity is a smooth function of the running variable with a varying slope, which is assumed to be bounded when the difference function is unidentifiable. The assumption limits the strength of the interaction between the running variable and group indicator in the conditional expected potential outcome. Assumption (ref) generalizes and hence weakens the assumption of cattaneo2020extrapolating who consider an important special case with two groups. Specifically, they assume that the cross-group heterogeneity does not depend on the running variable at all, i.e., $\lambda_{w12}=0$ implying $d(w,x,1,2)$ is constant in $x$. Their assumption is analogous to the parallel trend assumption under the difference-in-differences design.
Note that for some values of the running variable, the difference function is identifiable as $\tilde d(x,g,g^\prime) \ := \mathbbm{E}\left[Y \mid X=x,G=g\right]- \mathbbm{E}\left[Y \mid X=x,G=g^\prime\right]$. Therefore, under Assumptions (ref), we can directly compute the set of potential models for the difference function by extrapolating from the closest boundary point $\tilde c_{g^\ast}$. This leads to the following restricted model class for the difference function,
with the point-wise upper and lower bounds given by
where $\mathcal{X}_{gw}=\{x \in \mathcal{X} \mid \tilde{\pi}(x,g)=w\}$ represents the set of values of the running variable where the baseline policy gives the treatment condition $w$ to group $g$. Here the limits are taken as the value of the running variable $x'$ approaches the baseline cutoff $\tilde{c}_{g^\ast}$ from within $\mathcal{X}_{gw} \cap \mathcal{X}_{g^\prime w} $. We also note that the point-wise bounds for $d(w,x,g,g^\prime)$ in Equation (ref) could potentially be improved through alternative constructions. For instance, one could use extrapolated bounds for other pairwise difference functions involving an additional group $g''$, and calculate the difference between the upper and lower bounds of those difference functions: $[B_{l}(w,x,g,g'')-B_{u}(w,x,g'',g^\prime),B_{u}(w,x,g,g'')-B_{l}(w,x,g'',g^\prime)]$. This is an alternative partial identification bound for $d(w,x,g,g^\prime)$. While further intersecting these bounds could yield tighter partial identification results, in practice this may require an iterative process to progressively tighten the bounds. We construct the bounds using the straightforward approach in Equation (ref), leaving an exploration of other approaches to future work.
Finally, we use robust optimization to find an optimal policy. We do so by maximizing the worst case value of the expected utility $V(\pi,d)$ in Equation (ref) over the restricted policy class $\Pi_{\text{Overlap}}$, with the difference function taking values in the ambiguity set $\mathcal{M}_d$:
Note that the value under the baseline policy $ V(\tilde{\pi}, d)$ is identifiable, and so maximizing the the minimum value is equivalent to minimizing the maximum loss relative to the baseline policy. Below we will focus on the maximin formulation of the problem. As ben2021safe show, the resulting policy $\pi^{\inf}$ has a safety guarantee that the overall expected utility of the learned policy is no less than that of the baseline policy. This safety property holds because $\pi^{\inf}$ maximizes the worst case value when it disagrees with the baseline policy.
In the previous section, we define the population optimal policy $\pi^{\inf}$ that maximizes the worst case value of the population expected utility $V(\pi)$ for any given policy $\pi\in\Pi_{\text{Overlap}}$. Next, we show how to construct this maximin policy using observed data. Recalling our decomposition in Proposition (ref), the optimization problem of Equation (ref) can be decomposed into the sum of three components, $$\pi^{\inf} \ = \ \underset{\pi \in \Pi_{\mathrm{Overlap}}}{\operatorname{argmax}}\ \mathcal{I}_{\text{iden}}+\Xi_{\text{1}}+\underset{d \in \mathcal{M}_d}{\operatorname{min}} \Xi_{2}$$ where the first two components, $\mathcal{I}_{\text{iden}}$ and $\Xi_{\text{1}}$, are point-identifiable, while the last term $\Xi_{2}$ is not. We introduce our proposed estimator by examining each of these three terms in turn.
Since the first term $\mathcal{I}_{\text{iden}}$ in Equation (ref) is identifiable, we use its sample analog for unbiased estimation:
The second term $\Xi_{1}$ in Equation (ref) is also nonparametrically identified, but this term contains an unknown nuisance component, corresponding to the conditional outcome regression for each group $g$, $\tilde{m}(X, g)$. We compute the efficient influence function for estimating $\Xi_1$, which yields the following equivalent representation,
where $e_g(x):=\mathbb{P}\left[G_i=g\mid X_i=x\right]$ is the the conditional probability of group membership $g$ given the value of the running variable $X_i = x$.
The above identification requires an assumption on the existence of overlap between groups in terms of the running variable.
Although Equation (ref) resembles the standard doubly robust construction of the average treatment effect, there is a key difference; the nuisance model $e_g(x)$ represents the conditional probability of group membership given the running variable, rather than the conditional probability of treatment assignment. This difference has important implications. First, our framework neither requires covariate overlap nor unconfoundedness for treatment assignment. In fact, we also do not require group membership to be unconfounded. Second, the nuisance model is a function of a single-dimensional running variable, making it substantially easier to estimate than typical propensity score models.
We adopt a fully nonparametric approach to estimating these nuisance components $\tilde m(\cdot,\cdot)$ and $e_g(\cdot)$ and use cross-fitting chernozhukov2016locally,athey2021policy. Since the running variable is one dimensional, many potential nonparametric estimators can be used. In our empirical analyses (Section (ref)), we fit a local linear regression for $\tilde m(\cdot,\cdot)$ available in the R-package nprobust and fit a multinomial logisitic regression for $e_g(\cdot)$ based on a single hidden layer neural network of the R package nnet.
We cross-fit our estimators, randomly splitting the data into $K$ disjoint folds, where each fold contains a subset of data from all $Q$ groups. For each fold $k$, we use the remaining $K-1$ folds to estimate the nuisance components and then compute the estimates on fold $k$. We denote these estimates by $\hat{\tilde m}^{-k[i]}(X_i,g)$ and $\hat e_g^{-k[i]}(X_i)$, where we use the superscript $-k[i]$ to denote the model fitted to all but the fold that contains the $i$th observation. Now, we can write the proposed cross-fitted doubly-robust estimator as,
This doubly robust estimator $\widehat\Xi_{\text{DR}}$ achieves semiparametric efficiency while only requiring mild sub-parametric convergence rate conditions of the two nuisance models $\tilde{m}(x,g)$ and $e_g(x)$ (see Assumption (ref)). In Section (ref), we show that this rate-double robustness property allows us to learn the identifiable component of the policy value at a faster rate than the naive outcome imputation method.
We now turn to estimating the worst-case value of $\Xi_{2}$ in Equation (ref). We first define the sample analog of $\Xi_{2}$ as follows:
Unfortunately, it is not possible to directly optimize $\widehat\Xi_2$ over the restricted model class $\mathcal{M}_d$ because a model in this class contains an unknown nuisance component ${\tilde{d}}(\cdot,g,g^\prime)$. This term appears in the point-wise upper and lower bounds, $ B_{\ell}(w,x,g,g^\prime)$ and $ B_{u}(w,x,g,g^\prime)$. In Appendix (ref), we propose an efficient two-stage doubly robust estimator $\hat{\tilde{d}}^{DR}(\cdot,g,g^\prime)$ that has a fast convergence rate. This model builds upon the DR-learner developed for estimating the conditional average treatment effect kennedy2020towards.
We construct an empirical model class $\widehat{\mathcal{M}}_d$ by plugging in the estimates in place of their true values given in Equation (ref). The empirical bounds, $\hat{B}_{\ell}(w,x,g,g^\prime)$ and $\hat{B}_{u}(w,x,g,g^\prime)$, for any value of $x$, are given by:
We discuss choosing values of the smoothness parameter $\lambda_{wgg^\prime}$ in Section (ref) below.
With these three components in place, we obtain an empirical optimal policy by maximizing the worst-case in-sample value across policies $\pi \in \Pi_{\text{Overlap}}$:
Smoothness restrictions are often used to obtain partial identification manski1997monotone,kim2018identification. Similar to these cases, Assumption (ref) requires a priori knowledge of the degree of smoothness, directly controlled by the corresponding smoothness parameter $\lambda_{wgg^\prime}$. A greater value of $\lambda_{wgg^\prime}$ will lead to a more conservative policy whose new thresholds are closer to the original ones. In principle, the choice of the smoothing parameter value $\lambda_{wgg^\prime}$ should be guided by subject-matter expertise. Here, we propose a default data-driven method based on the assumption that the least smooth area of the difference function occurs in the region of overlapping policies between the two groups rambachan2023more.
To choose the value of $\lambda_{wgg^\prime}$, we leverage our proposed two-stage doubly robust estimation procedure described in Appendix (ref). In short, we first construct a DR pseudo-outcome for the actual observed difference $\tilde d(\cdot, g, g')$. We then fit a local polynomial regression with the running variable as the only predictor. The fitted nonparametric model also yields the estimates of their first local derivatives $\tilde d^{(1)}(\cdot, g, g')$ based on the local polynomial regression method proposed by calonico2018effect,calonico2020coverage. We compute the absolute value of the estimated first-order derivative at a grid of points in the region of overlapping policies between the two groups, and take the maximum value as $\lambda_{wgg^\prime}$. In practice, as illustrated in Section (ref), we recommend considering a range of plausible values of $\lambda_{wgg^\prime}$, e.g., multiplying those estimates by a constant, e.g., 2, 4, and 8, and conducting a sensitivity analysis to illustrate the sensitivity of learned policies to the strength of the smoothness assumption (see rambachan2023more,kim2018identification,imbens2019optimized for a similar practical advice).
To better understand the statistical properties of the learned policy described above, we compute an asymptotic bound on the utility loss relative to the baseline policy and the oracle optimal policy within the policy class. Before presenting the results, we introduce additional assumptions necessary for characterizing the estimation error of the policy value and the regret bound. These assumptions are stronger than the minimal assumptions required for identification.
First, as in much of the existing policy learning literature zhou2018offline,nie2021learning,zhan2021policy, we assume that the potential outcomes are bounded.
This bounded assumption can be generalized to unbounded but sub-Gaussian random variables zhou2018offline,athey2021policy.
Second, we consider the following conditions on the convergence rates for the estimators of the nuisance components in $\widehat{\Xi}_{\text{DR}}$ (Equation (ref)). These conditions are satisfied for many semi-parametric or non-parametric estimators (see zhou2018offline for a brief summary).
Assumption (ref) allows us to obtain tight regret bounds for learning policies using the proof strategy of zhou2018offline athey2021policy. Such conditions have been extensively used in the semi-parametric estimation literature newey2018cross,farrell2015robust,chernozhukov2018double, which only requires the product error bound of the two nuisance components to be $o\left(n^{-1}\right)$. Under the multi-cutoff RD design, Assumption (ref) is especially reasonable since the running variable $X$ is one-dimensional, and so we do not suffer from the curse of dimensionality when non-parametrically estimating the nuisance functions. This contrasts with typical observational studies where one must adjust for many observed confounders.
Finally, we quantify the effect of using the estimated empirical bounds (i.e., using the plug-in estimates of ${\tilde{d}}(x,g,g^\prime)$) on the regret of the learned policy by invoking the following assumption that characterizes the estimation error rate of $\hat{\tilde{d}}(\cdot,g,g^\prime)$.
This assumption concerns the estimation error of the conditional cross-group difference function ${\tilde{d}}(\cdot,g,g^\prime)$ at the existing discrete cutoff values $x=\tilde{c}_{1},\ldots,\tilde{c}_{Q}$. In practice, this convergence rate $\rho_n$ will depend on the specific choice of the estimator $\hat{\tilde{d}}(\cdot,g,g^\prime)$. To maintain the generality of our results, we remain agnostic about the exact rate of $\rho_n^{-1}$ in our theoretical results, but we discuss special cases in Remark (ref) below.
From the definitions of the lower bounds in Equations (ref) and (ref), we can write the estimation error of the empirical (lower) bound $\hat{B}_\ell(w, x, g, g')$ in terms of the estimation error of $\hat{\tilde{d}}(\cdot,g,g^\prime)$,
Therefore, under Assumption (ref), the estimation error of the empirical (lower) bound is also governed by the rate $\rho_n^{-1}$: $$ \mathbb{E}\left[\frac{1}{n}\sum_{i=1}^{n} \max _{w \in \{0,1\},\ g,g^\prime\in \mathcal{Q}} \left | {B}_{\ell}(w,X_i,g,g^\prime)-\hat{B}_{\ell}(w,X_i,g,g^\prime)\right|\right]=O\left(\rho_n^{-1}\right).$$ Here, we focus on the lower bounds since the inner optimization problem of Equation (ref) depends only on the worst value of $\hat{\tilde{d}}(x,g,g^\prime)$, i.e, ${B}_{\ell}(w,x,g,g^\prime)$; however, this rate applies to the upper bound as well. In the subsequent sections, we will examine how this error affects the final performance of the learned policy.
Under the above assumptions, we compare the learned policy $\hat{\pi}$ to the baseline policy as well as the oracle policy that maximizes the expected utility. To simplify the statement of the main theorem, we first define $ {\Gamma}^{(1)}_{ig} $ and $ {\Gamma}^{(0)}_{ig} $ as the doubly-robust scores for Group $g$ used to extrapolate to the left (from the treated to the control) and right (from the control to the treated) of the current cutoff, respectively: {
}
To characterize the safety property of our learned policy $\hat{\pi}$, we first bound the utility loss relative to the baseline policy $\tilde \pi$. Here, we report the simplified version of this bound, and provide a more general version in Appendix (ref).
The theorem ensures, asymptotically, that the learned policy will not yield a worse overall utility than the baseline policy. The constant term $V^\ast_{g}$ in the bound arises from the proposed doubly robust estimator $\widehat\Xi_{\text{DR}}$ zhou2018offline,athey2021policy. In addition, the rate of convergence depends on the slower of the two estimation error rates: one is the standard $n^{-1/2}$ rate that comes from estimating ${\mathcal{I}}_{\text{iden}}$ and $\Xi_{\text{DR}}$, and the other is the estimation error $\rho_n^{-1}$ for the empirical model class $\widehat{\mathcal{M}}_d$. This term $\rho_n^{-1}$ typically exhibits a slower-than-parametric rate, depending on the accuracy of estimating the local difference ${\tilde{d}}(\cdot,g,g^\prime)$ at fixed points.
Compared to the optimal $O_p(n^{-1/2})$ regret bound achieved in the standard policy learning settings with point identification kallus2018confounding,Kitagawa2018,zhan2021policy,athey2021policy, the presence of this additional error term $\rho_n^{-1}$ is due to the partial identification under multi-cutoff RD designs. Such a penalty is generally unavoidable under partial identification settings kallus2018confounding,khan2023off. Although it is slower than $\sqrt{n}$, in RD contexts with a single running variable and smooth difference functions, $\rho_n^{-1}$ can typically achieve a good convergence rate.
Lastly, we compare our learned policy to the infeasible, oracle optimal policy $\pi^\ast$ in the same policy class, defined as: $$\pi^{\ast} \ := \ \underset{\pi \in \Pi_{\text{Overlap}}}{\operatorname{argmax}}\ V(\pi).$$ As before, we report a simpler version of the bound here, while its complete version is found in Appendix (ref).
Compared to Theorem (ref), there is an additional term that is proportional to the smoothness parameter $\lambda_{wgg^\prime}$ and the average distance to the nearest cutoff. This term measures the sample average of the length of the point-wise partial identification bound for $d(w,X_i,g,g^\prime)$ under Assumption (ref). While the robust optimization procedure ensures that our learned policy will be no worse than the baseline policy, asymptotically, Theorem (ref) shows that this safety guarantee comes at the cost of obtaining a potentially suboptimal policy. If the cross-group difference deviates substantially from a parallel trend, i.e., $\lambda_{wgg^\prime}>0$ is larger, then the interval will be wider, leading to a potentially larger optimality gap.
We note that Theorem (ref) assumes the smoothness parameters $\lambda_{wgg^\prime}$ are known a priori. The result does not account for the uncertainty that might arise due to the use of a data-driven method for estimating $\lambda_{wgg^\prime}$. In different contexts with partial identification, kallus2018confounding and khan2023off take an approach similar to the one discussed in Section (ref) by treating such hyper-parameters as known and conducting a sensitivity analysis.
In this section, we apply the proposed methodology to the ACCES (Access with Quality to Higher Education) program, which is a national-level subsidized loan program in Colombia. We present additional simulation results evaluating our approach in Appendix (ref). Starting in 2002, the ACCES program has provided financial aid to low-income students. Eligibility for the program is determined by a test score cutoff, which varies across different regions of Colombia and time periods. The original analysis by melguizo2016credit estimated a single RD treatment effect by pooling and normalizing all 33 cutoffs, while cattaneo2020extrapolating focused on two cutoffs from different time periods and extrapolated the RD treatment effect for each one. In contrast to these prior studies, we analyze all the cutoffs across different geographical regions at the same time and learn a new safe policy for each one of them by leveraging the existence of multiple cutoffs.
We begin by providing some background information about the ACCES program (see melguizo2016credit for more details). To be eligible for the ACCES program, a student must achieve a high score on the national high school exit exam, the SABER 11, administered by the Colombian Institute for the Promotion of Postsecondary Education (ICFES). Each semester of every year, once students take the SABER 11, the ICFES authorities compute a position score that ranges from 1 to 1,000, based on each student's position in terms of 1,000 quantiles. For example, the students who are assigned a position score of 1 should be in the top 0.1%, while those with the test scores in the bottom 0.1% receive a position score of 1000. The SABER 11 position scores are calculated for each semester, and then pooled for each year. Throughout our analysis, as in cattaneo2020extrapolating, we multiply the position score by $-1$ so that the values of the running variable above a cutoff lead to the program eligibility.
To select ACCES beneficiaries, the Colombian Institute for Educational Loans and Studies Abroad establishes an eligibility cutoff such that students whose position scores are higher than the cutoff becomes eligible for the ACCES program. The eligibility cutoff is defined separately for each department (or region) of Colombia, and has changed over time. Between 2002 and 2008, the cutoff was $-850$ for all Colombian departments. After 2008, however, the ICEFES used varying cutoffs across different years and departments.
We use data collected from all 33 departments in 2010, when each department used a different cutoff value. In our application, the running variable, despite being discrete, takes a large number of distinct values. In addition, we drop departments whose sample size is less than 200, which ensures the reliable use of local polynomial regression calonico2014robust,calonico2018effect. The final dataset includes observations from 23 different departments of Colombia. Following melguizo2016credit and cattaneo2020extrapolating, the outcome of interest is an indicator for whether the student enrolls in a postsecondary institution. Table (ref) reports the sample size, group-specific baseline cutoff $\tilde{c}_g$, and estimated group-specific RD treatment effect at the baseline cutoff for each department.
We estimate the empirical improvement of the expected utility, i.e., $\hat{V}(\pi)-\hat{V}(\tilde \pi)$, for each candidate policy in the policy class $\Pi_{\text{Overlap}}$, restricting the candidate cutoffs for each department to be within the range of the original lowest and highest cutoffs, i.e., $c_g \in [-828, -559]$. We select the value of the smoothness parameter $\lambda_{wgg^\prime}$ for each pairwise difference function $d(w,x,g,g^\prime)$ under Assumption (ref), by applying the procedure described in Section (ref). We first fit a local polynomial regression of each difference function $d(w,\cdot,g,g^\prime)$, and then take the maximal absolute value of the estimated local first derivative at a grid of points as $\lambda_{wgg^\prime}$. We also perform a sensitivity analysis by multipling the data-driven choice of $\lambda_{wgg^\prime}$ by a positive constant $M$ and seeing how the learned policy changes.
We first consider the stronger parallel-trend assumption that the cutoff differences are exactly constant, i.e., $\lambda_{wgg^\prime}=0$, allowing us to point-identify the counterfactual function $d(w,\cdot,g,g^\prime)$. As shown in the first column ($M=0$) in Figure (ref), for most departments, the learned cutoffs are lower than the baseline cutoffs (highlighted by purple). Thus, lowering the cutoff, which increases the number of eligible students, is likely to result in a higher overall enrollment rate. This is not surprising, as financial aid typically should help students attend post-secondary institutions.
However, the learned cutoffs in La Cuarjira, Cordoba, Norte De Santander, and Huila are much greater than the actual baseline cutoffs used. Referring to the summary results in Table (ref), only Cordoba showed a significant negative treatment effect at the cutoff, while the treatment effect of the other three departments were ambiguous. Therefore, we find that the parallel-trend assumption may be too aggressive and lead to unnecessary policy changes.
Next, we consider a weaker smoothness restriction using the data-driven estimate of $\lambda_{wgg^\prime}$, corresponding to the $M=1$ column in Figure (ref). These estimated values range from $9.2 \times 10^{-6}$ to $1.4\times 10^{-1}$ with a median of $5.8\times 10^{-3}$ for different possible values of $w,g$ and $g'$. Overall, we observe that the learned policies are less aggressive compared to the results obtained under the parallel-trend assumption. Although most departments still have reduced cutoffs, the magnitude of the changes is smaller. Notably, the policy changes are now more conservative for those departments that would have seen an increase in the cutoff under the “parallel-trend” assumption. Specifically, Cordoba shows the largest change, aligning with the significant negative treatment effect estimate at the cutoff for Cordoba (see Table (ref)).
We conduct a sensitivity analysis by varying the multiplicative factor $M$. As shown in Figure (ref), the learned policies are relatively robust to the choice of the smoothing parameter, although, as expected, a greater value of $\lambda_{wgg^\prime}$ results in a more conservative learned policy that is closer to the baseline policy.
Finally, we incorporate the cost of treatment into the utility function. Given that the outcome enrollment rate is binary, a general utility function for outcome $y$ under treatment $w$ can be expressed as $yu(1,w)+(1-y)u(0,w)=yu(w) + c(w)$. Here, $u(w):=u(1,w) - u(0,w)$ represents the utility change in the treatment condition, and $c(w):=u(0,w)$ can be interpreted as the baseline cost of implementing treatment $w$. We set the utility gain for enrollment to $1$ under with and without access---i.e. $u(1)=u(0)=1$--- and set the cost of not providing financial aid $c(0)$ to zero and the cost of providing it nto $C$. Thus, the utility function is of the form $u(y,w)=y-C\times w$.
We examine how changing the cost $C$ affects the learned policy. We consider a range of values for the cost, $C\in[0,1]$, to explore the trade-off between cost and utility. Figure (ref) shows the results. Changing the cost of offering a loan yields optimal robust cutoffs that affect the number of eligible students. The overall trend indicates that the learned policy will still lower the cutoffs relative to the baseline for most departments when the cost is low to moderate. When the cost is high relative to the utility gain of enrollment---i.e. when $C \geq 0.6$---the cutoffs start to meaningfully increase, leading to more students being excluded from the program relative to the baseline. Note that in the Distrito Capital, where the baseline cutoff is the highest, the learned cutoff will eventually reach a point where it remains unchanged, due to the restriction that the learned cutoffs cannot exceed the maximum baseline cutoff.
Despite the popularity of regression discontinuity (RD) designs for program evaluation with observational data, little attention has been given to the question of how to learn new and better treatment assignment cutoffs. In this paper, we propose a methodology for safe policy learning that guarantees that a new, learned policy performs no worse than the existing status quo. We partially identify the overall utility function and then find an optimal policy in the worst case via robust optimization. We stress, however, that extrapolation is required to learn new policies. Our approach leverages the presence of multiple cutoffs, and is based on a weak, non-parametric smoothness assumption about the level of cross-group heterogeneity.
There are several directions for future research. First, the partial identification assumption we employ involves smoothness parameters. While we have suggested ways that the data can inform these hyperparameters, developing rigorous selection criteria for this choice will be an important aspect of future work.
Second, we can generalize Assumption (ref) to the case where it holds only after conditioning on the other available covariates in RD design. Consequently, our proposed doubly robust estimator can be modified to include not only the single running variable, but also other observed covariates.
Third, we can incorporate other constraints such as fairness, risk factors, simplicity, or other functional form constraints into the policy class. For example, candidate treatment cutoffs may be restricted by a lower threshold $c_{\text{lower}}$, i.e., $ \Pi=\left\{ \pi(x,g) = \mathds{1}(x\geq c_g) \mid c_g \geq c_{\text{lower}} \right\}$, or federal policymakers want to learn a unified cutoff for all subpopulations, i.e. $ \Pi=\left\{ \pi(x,g) = \mathds{1}(x\geq c) \mid c \in [\tilde{c}_1,\tilde{c}_Q] \right\}$. In these cases, our robust optimization approach can be directly applied to learn the worst-case optimal policy in the class, although the safety guarantee relative to the status quo may be compromised because the baseline policy may not be included in the policy class. We could also incorporate hard budget constraints of the form $\ensuremath{\mathbb{E}}[\pi(X, G)] \leq \delta$, although there may be additional estimation issues regarding uncertainty about the constraint sun2021empirical.
Fourth, while our current methodology primarily addresses scenarios with continuous running variables, future work should extend the proposed framework to handle discrete running variables, which introduce partial identification challenges similar to those posed by extrapolation. As noted by kolesar2018inference and imbens2019optimized, explicit smoothness assumptions on the outcome or difference functions might help address these challenges.
Fifth, we can change the target population of interest and expand the notion of safety. In particular, it is of interest to establish a statistical safety guarantee for a specific group $g$ or other well-defined subpopulation so that a learned policy does not perform worse in this subpopulation than the status quo policy jia2023. This differs from the current guarantees that hold over the entire population on average, and so can lead to worse outcomes for certain subpopulations.
Finally, one may consider the use of safe policy learning in other study designs. For example, the tie-breaker design of li2022general combines treatment assignment based on both randomization and thresholding. Extending our methodological framework to this and other study designs is left to future work.