Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
68,958 characters · 13 sections · 42 citation commands
Regression Discontinuity Design under Self-selection
Regression discontinuity (RD) design is an important policy evaluation tool that has been widely used in empirical studies. Under the continuity assumptions Lee2008, the RD design gives rise to many testable restriction similar to a randomized control trial, and it allows for the identification of causal effects Hahn2001. Standard nonparametric tools like series expansion method or kernel regression method can be applied to estimate this quantity under this minimal assumption, see Imbens2008 and CattaneoRD2017.
An implication of the continuity assumption is that the distribution of the covariates conditional on the assignment variable is continuous at the cutoff. To valid this assumption, a variety of statistical tests and nonparametric inference procedures have been proposed, including Cattaneo2015 and Canay2018. However, this assumption may not hold in many applications. In particular, we consider the following two motivating examples.
Due to different distributions of the covariates at the cutoff in these examples, the standard RD estimand is no longer valid. We show in those cases the standard RD estimand can be decomposed into a direct treatment effect and an indirect treatment effect. The indirect effect is due to the unbalanced covariates near the cut-off. For example, the policy intervention may result more boys than girls to receive scholarship. If in general boys perform different from girls in SAT test, the difference in SAT score due to gender will also be accounted into the standard RD estimand as if the policy intervention is selecting on genders. However as no direct causal mechanism is assumed between the unbalanced covariates and running variable, the selection could be generated due to an unknown equilibrium, completely reversed or purely spurious. For example, in a GSP auction, the reservation score is designed to separate the bidders by quality and in the meanwhile bidders may self-select such that low quality bidders may have incentive to bid higher. As a result, policy changes to move reservation score can be risky if the equilibrium between reservation score and quality of bidders are not disentangled.
In this paper, we propose a new framework to address this problem by adjusting unbalanced covariates due to self-selection. Consider the following sharp RD setup: $T$ is a binary treatment variable, $Y(1)$ and $Y(0)$ are the potential outcomes under $T=1$ and $T=0$ and $X$ is the running variable so that the treatment is fully determined by $T = \bold{1}(X>c)$ for a known threshold $c$. In the classical RD framework, it is assumed that $\EE(Y(t)|X=x)$ is continuous at $x=c$ for $t=0, 1$; see Assumption (ref). This essentially assumes away self-selection based on both observed and unobserved covariates. To account for the self-selection effect, we assume that further covariate information can be collected. In particular, we define $Z(1)$ and $Z(0)$ as the potential covariates with or without treatment. By using the potential “outcome" formulation, we allow the distribution of the covariates on two sides of the threshold to be different, i.e., discontinuity of the covariate distribution. By controlling all unbalanced covariates, we assume that $\EE(Y(t)|X=x, Z(t)=z)$ is continuous in $x$ at the threshold; see Assumption (ref). This is the main assumption made in this paper.
In the classical RD framework, the standard causal parameter of interest is the marginal treatment effect $\mathbb{E}(Y(1) - Y(0)|X=c)$ at the cutoff. In the presence of self-selection, the continuity assumption on $\EE(Y(t)|X=x)$ may fail. Thus, the marginal treatment effect can be confounded by the discontinuity of the conditional mean of covariates at the cutoff, as seen in equation ((ref)). In this paper, we propose a class of estimands for RD design in the framework of weighted average treatment effect (WATE). We show that our estimands can tease out the effect of discontinuity of the conditional distribution of covariates through re-weighting the marginal treatment effects. For instance, under the constant treatment effect model with unbalanced covariates, our estimands reduce to the direct treatment effect. One special case of our estimands is FrolichHuber2018. Another special case can be interpreted as the “global" average treatment effect (ATE) $\mathbb{E}(Y(1) - Y(0))$ under the conditional independence assumption (CIA) as in Angrist2015. The CIA assumes that the treatment is mean independent of the running variable near the cutoff conditional on the covariates. Intuitively, after projecting the outcome variable onto a rich set of covariates (excluding running variable), the residual should not depend on the running variable and can be viewed as an experiment with randomly assigned treatment. We show that our method identifies ATE under CIA assumption, whereas the standard RD estimand remains a “local" treatment effect at the cutoff Lee2008.
We further provide the nonparametric identification for our estimands and propose nonparametric estimators based on the inverse propensity score weighted (IPW) approach, see HorvitzThompson1952 and AbadieImbens2016. However, notice that the treatment assignment is degenerate with respect to the running variable. We get around this problem by considering the marginal effect of the treatment and re-weighting on the other covariates first. The kernel method is applied to estimate the conditional mean function. The consistency and asymptotic normality of the proposed estimator are established. We further extend our method to the fuzzy RD design where the treatment compliance is imperfect. Similarly, we provide the nonparametric identification of the causal effect under the fuzzy RD design, and propose a nonparametric estimator. The proposed estimator is similar to that of the local average treatment effect (LATE) in a fractional format but with numerator being adjusted to incorporate additional selections.
This work is connected to the growing literature on the RD design with covariates. In particular, two recent papers provide insightful guidance on the subject. Matias2017 estimated the marginal treatment effect by a local linear regression with the linear-in-parameters specification for the covariates. The main advantage of their method is that the nonparametric estimation of $\EE(Y(t)|X=x, Z(t)=z)$ is avoided. In another paper, FrolichHuber2018 proposed a fully nonparametric estimator of the marginal treatment effect by estimating $\EE(Y(t)|X=x, Z(t)=z)$ nonparametrically. They allowed the conditional density of $Z(t)$ given $X$ to be discontinuous. Our work differs from the above papers by considering a different estimand that is less local under CIA and identifies the direct treatment effect under the constant treatment effect model. Unlike Matias2017, we do not require the continuity of the conditional mean of $Z(t)$ given $X$. Instead, our Assumption (ref) is similar to assumption 1 (iv) in FrolichHuber2018 under the sharp RD design. The proposed IPW estimator is also different from the above regression based estimators.
The rest of the paper is organized as follows. In Section (ref), we propose our new estimands in the form of weighted average treatment effects and establish its nonparametric identification. In Section (ref), we propose a set of weighted local linear estimators. The theoretical properties are established in Section (ref). In Section (ref), we conduct simulation studies and also apply the method to two empirical data sets: the U.S. House elections data in Lee2008 and a novel data set from Microsoft Bing on Generalized Second Price (GSP) auction. Finally, we consider the extension to the fuzzy RD case in Section (ref). The proofs are deferred to the appendix.
In the standard RD design setting, we observe $n$ i.i.d. random samples $\{Y_i, X_i, Z_i, T_i\}_{i = 1}^n$, where $Y_i$ is the outcome variable of interest for the $i$th sample, $T_i\in\{0,1\}$ is the binary treatment variable, $X_i\in\RR$ is the running variable and $Z_i\in\RR^p$ is the covariate. In the sharp RD design, the treatment $T_i$ is perfectly assigned through the running variable $X_i$ relative to a known cutoff $c$. For example, we have \[T_i = \bold{1}(X_i>c).\] Adopting a potential outcome framework, we can write the observed outcome variable $Y_i$ as \[Y_i = Y_i(0)\cdot (1-T_i)+Y_i(1)\cdot T_i,\] where $Y_i(0)$ and $Y_i(1)$ represent the potential outcomes without or with treatment. The average treatment effect is defined as $\mathbb{E}(Y_i(1) - Y_i(0))$. However, this estimand is not identifiable under the RD design as the treatment assignment $T_i$ is a deterministic function of $X_i$. In this framework, Hahn2001, Lee2008 and Cattaneo2015 showed that one can still identify the treatment effect at the cutoff \[\tau_{SRD} = \mathbb{E}(Y_i(1) - Y_i(0)|X_i=c), \] under the following continuity assumption.
This assumption implies that the conditional mean of the potential outcomes near the cutoff $x=c$ are similar. There is no discontinuity of the conditional mean functions at the cutoff. This assumption enables us to identify $\tau_{SRD}$ in the RD design. We refer to Hahn2001, Lee2008 and Cattaneo2015 for further discussion on this assumption.
Now let us consider the case that additional covariates $Z_i$ are observed. Denote \[Z_i = Z_i(0)\cdot (1-T_i)+Z_i(1)\cdot T_i,\] where $Z_i(0)$ and $Z_i(1)$ represent the potential covariates without or with treatment. In the presence of covariates $Z_i$, the causal parameter $\tau_{SRD}$ can be rewritten as
In a recent work, Matias2017 proposed a kernel based estimator of $\tau_{SRD}$ by accounting for the additional covariates $Z_i$. In addition to the continuity Assumption (ref), it is also assumed that $\EE(Z_i(1)|X_i=c) = \EE(Z_i(0)|X_i=c)$ for the consistency of the resulting kernel estimator, that is the potential covariates $Z_i(1)$ and $Z_i(0)$ have the same conditional mean at the cutoff $X_i=c$.
However, in some applications we may observe $\EE(Z_i(1)|X_i=c) \neq \EE(Z_i(0)|X_i=c)$, when self-selection based on the covariates exists. For example, consider the classical scholarship example. Students with SAT score higher than a threshold will receive scholarship. The treatment effect of interest is the effect of scholarship on the students' first semester GPAs. If the cutoff is pre-released, it might be possible that students with some common characteristics (i.e. gender) may study harder to pass the bar. This leads to an ex-ante selection based on the covariates. Thus, one may observe that the conditional mean functions of $Z_i$ given $X_i$ right below or above the threshold are different, i.e.,
The following simple lemma essentially says that the self-selection based on the covariates (i.e., eq (ref)) implies $\EE(Z_i(1)|X_i=c) \neq \EE(Z_i(0)|X_i=c)$.
The above lemma provides a convenient way to check whether $\EE(Z_i(1)|X_i=c) = \EE(Z_i(0)|X_i=c)$ holds in empirical studies. One may simply plot the observed covariates $Z_i$ against $X_i$ and examine whether there is a discontinuity of the trend around $x=c$. The method is applied in the real data analysis.
In the following, we investigate the consequence of $\EE(Z_i(1)|X_i=c) \neq \EE(Z_i(0)|X_i=c)$. To be specific, we consider the following constant treatment effect model
for $t\in\{0,1\}$, where $g(\cdot)$ is an arbitrary continuous function. By ((ref)), one can show that
The estimand $\tau_{SRD}$ can be decomposed into two terms. The first term $\tau$ represents the direct treatment effect after controlling the running variable $X_i$ and the covariates $Z_i(t)$. The second term in the right hand side of ((ref)) can be interpreted as the indirect effect of the policy due to the unbalanced covariates near the cutoff or self-selection, which is nonzero if $\gamma\neq 0$ and $\EE(Z_i(1)|X_i=c) \neq \EE(Z_i(0)|X_i=c)$. In many applications, the direct treatment effect $\tau$ is usually more meaningful and interpretable than $\tau_{SRD}$, as $\tau_{SRD}$ is confounded by the self-selection effect.
In this example, when the indirect effects in $\tau_{SRD}$ is assumed away by requiring the continuity of the conditional density of the covariates $Z_i$ at the cutoff $X_i = c$, the data around the cutoff can be viewed as a natural experiment and continuity on the covariates implies a balanced design for this experiment so that we can estimate a local average treatment effect. However when the self-selection exits, the experiment is no longer balanced and the average treatment effect $\tau_{SRD}$ will typically differ from the direct effect $\tau$.
As seen in ((ref)), if $\EE(Z_i(1)|X_i=c) \neq \EE(Z_i(0)|X_i=c)$ (i.e., the covariates are unbalanced at the cutoff) and $\gamma\neq 0$, $\tau_{SRD}$ can be different from the causal parameter of interest. To overcome this difficulty, we propose a new class of causal parameters, called the weighted average treatment effect (WATE), which are defined as
where $w_1(\cdot)$ and $w_0(\cdot)$ denote different choices of weights to form the estimand, and $f_{Z(1)|X}(\cdot |\cdot)$ and $f_{Z(0)|X}(\cdot |\cdot)$ are the conditional density of $Z(1)$ and $Z(0)$ given $X$. In order to interpret ((ref)) as the WATE, we require the following normalization condition for $w_1(\cdot)$ and $w_0(\cdot)$: $$ \int w_1(z)f_{Z(1)|X}(z|c)dz=\int w_0(z)f_{Z(0)|X}(z|c)dz=1. $$ In particular, by choosing appropriate $w_1(\cdot)$ and $w_0(\cdot)$, ((ref)) can be interpreted as the average of the difference of the conditional mean functions corresponding to a target population. To see this, we consider the following examples. Denote $$ \Delta(c,z)=\mathbb{E}(Y(1)|X=c, Z(1)=z) - \mathbb{E}(Y(0)|X=c, Z(0)=z). $$
A summary of the first three estimands is provided in Table (ref). In practice, which causal estimand in above examples to use should depend on the target population of interest and is often determined on a case-by-case basis. Indeed, our framework opens a door towards designing new causal parameters tailored to specific applications. For instance, similar to $\tau^{w2}_{SRD}$, one can define the average treatment effect over locally treated population i.e., $\int \Delta(c,z) f_{Z(1)|X}(z|c)dz$. Since the goal of the paper is to deal with unbalanced covariates, to fix the idea we will mainly focus on the first three examples.
In the following, we comment on two properties of our estimands $\tau^{w1}_{SRD}$, $\tau^{w2}_{SRD}$ and $\tau^{w3}_{SRD}$. First, under the constant treatment effect model ((ref)), direct calculation shows that $\Delta(c,z)=\tau$ and thus $\tau^{w1}_{SRD}=\tau^{w2}_{SRD}=\tau^{w3}_{SRD}$ equals to the direct treatment effect $\tau$ without any further assumption. In contrast, the classical RD estimand $\tau_{SRD}$ reduces to $\tau$ under the extra assumption that $\gamma= 0$ or $\EE(Z_i(1)|X_i=c) = \EE(Z_i(0)|X_i=c)$.
Second, our estimand $\tau^{w1}_{SRD}$ generalizes to the overall average treatment effect (ATE) under the conditional independence assumption (CIA) proposed by Angrist2015. Assume that $Z_i$ are the pre-treatment covariates, i.e, $Z_i(1)=Z_i(0)=Z_i$. The CIA is defined as
which says that the potential outcomes are mean independent of the running variable conditional on the covariates. By controlling a rich set of covariates, CIA seems to be a reasonable assumption as the link between the running variable and outcomes can be blocked Angrist2015. Since the CIA ((ref)) implies $\Delta(c,z)=\Delta(z)$, our estimand $\tau^{w1}_{SRD}$ reduces to $$ \tau^{w1}_{SRD}=\int \Delta(z) f_Z(z)dz=\mathbb{E}(Y_i(1) - Y_i(0)), $$ which is the overall ATE. In contrast, $$ \tau^{w2}_{SRD}=\tau^{w3}_{SRD}=\tau_{SRD}=\mathbb{E}(Y_i(1) - Y_i(0)|X_i=c) $$ remains a “local" treatment effect at the cutoff. It requires further conditions to generalize to the overall ATE. For instance, if the constant treatment assumption $\mathbb{E}(Y_i(1)|Z_i)-\mathbb{E}(Y_i(0)|Z_i)=a$ holds for some constant $a$, then $\tau_{SRD}$ reduces to $\mathbb{E}(Y_i(1) - Y_i(0))$. Thus, the new estimand $\tau^{w1}_{SRD}$ can represent a causal effect that is less local than the standard RD estimand $\tau_{SRD}$.
In this subsection, we study the nonparametric identification of $\tau^{w1}_{SRD}$, $\tau^{w2}_{SRD}$ and $\tau^{w3}_{SRD}$. Instead of Assumption (ref), we impose the following continuity assumption.
Intuitively, this assumption says, once all unbalanced covariates are controlled, there is no further discontinuity between the running variable and outcomes at the threshold. This assumption is similar to assumption 1 (iv) in FrolichHuber2018 under the sharp RD design. In addition, this assumption is weaker than CIA, as ((ref)) implies our Assumption (ref) when the covariates are pre-determined.
The following theorem shows that $\tau^{w1}_{SRD}$ is identifiable based on the distribution of the observed data under Assumption (ref). In addition, $\tau^{w2}_{SRD}$ and $\tau^{w3}_{SRD}$ are identifiable under some extra continuity assumptions.
In the causal inference literature, inverse propensity score weighting (IPW) is one of the most widely used tools to handle the unbalanced covariates in the treatment and control groups HorvitzThompson1952. However, the standard IPW method is not directly applicable because in the sharp RD design the treatment assignment is a deterministic function of the running variable and thus the propensity score is degenerate. In this section, we propose a class of nonparametric estimators of $\tau^{w1}_{SRD}$, $\tau^{w2}_{SRD}$ and $\tau^{w3}_{SRD}$ by modifying the inverse propensity score weighting approach.
To motivate our nonparametric estimator, we consider the following notation. Denote by $\pi_1(Z_i)$ a function of the covariate $Z_i$ to be chosen later, $K(\cdot)$ a symmetric kernel function and $h$ a bandwidth that shrinks to 0. The detailed conditions on the kernel function and bandwidth are deferred to the next section. Consider the following inverse weighted kernel estimator
for the estimand $\EE\{Y(1)w_1(Z(1))|X=c\}$, where $\pi_1(Z)$ plays the same role as the propensity score in the IPW method. Since in the RD design the propensity score function is degenerate, in the following we will show that the choice of $\pi_1(Z_i)$ differs from the standard propensity score model. The rationale is to choose $\pi_1(Z_i)$ so that the estimator ((ref)) is asymptotically unbiased,
Thus, it suffices to calculate the expectation of the kernel estimator as follows
where the first step follows from $Z_i=Z_i(1)$ under $T=1$ and the last step follows from the definition $T_i = \bold{1}(X_i>c)$. Thus, ((ref)) implies
where the last step follows from the symmetry of the kernel function $\int_{u>0} K(u)du=1/2$ and can be made rigorous given the regularity conditions specified in the next section. Comparing with $$ \EE\{Y_i(1)w_1(Z_i(1))|X_i=c\}=\int \mathbb{E}(Y_i(1)|X_i=c, Z_i(1)=z) w_1(z)f_{Z(1)|X}(z|c)dz, $$ we can see that ((ref)) holds provided
where $f_X(c)$ is the p.d.f of $X$ at $x=c$. Following from the same argument, one can show that $\pi_0(z)=\frac{f_X(c)}{2w_0(z)}$. In Table (ref), we give the detailed expression of $\pi_1(z)$ and $\pi_0(z)$ for the proposed three estimands $\tau^{w1}_{SRD}$, $\tau^{w2}_{SRD}$ and $\tau^{w3}_{SRD}$. Since the weight $\pi_1(z)$ depends on the unknown density functions, we propose to estimate those densities by the following kernel estimators
where $h_1$ and $h_2$ are bandwidth parameters. Note that the kernel estimator is known to suffer from the curse of dimensionality and is only applicable when the dimension of $Z_i$ is small. For the applications in which a large number of covariates can be collected, one may consider alternative parametric or semiparametric approaches for density estimation. In this work, we only focus on the above kernel estimators and leave the alternatives for future investigation.
Replacing the unknown density functions in $\pi_1(z)$ and $\pi_0(z)$ as shown in Table (ref) with the corresponding kernel estimators, we can obtain $\hat\pi_1(z)$ and $\hat\pi_0(z)$. While we can construct the final estimator by plugging $\hat\pi_1(z)$ into ((ref)), for practical use and theoretical analysis we recommend the local linear estimator, since it has smaller asymptotic bias and better finite sample behavior near the boundary fan1996local. Motivated by the formulation of the kernel estimator ((ref)), we propose the following weighted local linear (WLL) estimator
Thus, we can estimate the WATE ${\tau}_{SDR}^w$ by
In this section, we study the asymptotic properties of the proposed estimator. We focus on the local linear estimator ((ref)) for the estimand $\tau^{w1}_{SRD}$, due to the nice properties of $\tau^{w1}_{SRD}$ as explained in Section 2.2. The analysis of the estimators for the other two estimands $\tau^{w2}_{SRD}$ and $\tau^{w3}_{SRD}$ is similar, but may require different technical conditions.
Let $\delta$ denote a small positive constant. For notational simplicity, define $\cF^-$ as the class of functions of $x\in(c-\delta, c]$ and $z\in\cZ$ such that for any $f\in\cF^-$, $\frac{\partial^3}{\partial a\partial b\partial c}f(x,z)$ is continuous, where $a,b,c=\{x,z\}$. Here, the derivatives of $f(x,z)$ with respect to $x$ at $x=c$ (say $\frac{\partial^3}{\partial x^3}f(c,z)$) are interpreted as left derivatives. Similarly, we define $\cF^+$ as the class of functions of $x\in[c, c+\delta)$ and $z\in\cZ$ such that for any $f\in\cF^+$, $\frac{\partial^3}{\partial a\partial b\partial c}f(x,z)$ is continuous, where $a,b,c=\{x,z\}$. The derivatives at $x=c$ correspond to the right derivatives. For simplicity, we only consider the case that $\textrm{dim}(Z)=1$. The generalization to multivariate covariates follows from the similar argument.
Since the proposed method requires to estimate the unknown density functions $f_{X, Z(0)}(x,z)$ and $f_{X, Z(1)}(x,z)$, we assume they are sufficiently smooth. In addition, we also need the smoothness of $m_t(x,z)$ in order to study the local linear estimator ((ref)). This assumption implies Assumption (ref), which is required for identification purpose.
Assumption (ref) has four parts. The first and second parts are standard assumptions in kernel density estimation problems. Since our RD estimator applies local linear estimators at the boundary, the uniform kernel or triangular kernel has been shown to have good performance under such scenario Matias2017. The third part requires $f_{X,Z(1)}(x,z)$ and $f_{X,Z(0)}(x,z)$ to be bounded away from 0 so that the inverse weights in the local linear estimator can be well controlled. This is similar to IPW estimator which requires the propensity score to be bounded away from 0. The last part assumes that the (homoscedastic) noise has finite variance.
Denote $\alpha_t=\int \mathbb{E}(Y(t)|X=c, Z(t)=z) f_Z(z)dz$, and recall that $\tau^{w1}_{SRD}=\alpha_1-\alpha_0$. The main theorem in this section shows the rate of convergence of the local linear estimators $\hat{\alpha}_{1}$ and $\hat{\alpha}_{0}$ and their limiting distributions.
We note that from ((ref)) the optimal choice of $h$ is of order $O(n^{-1/6})$ and the corresponding convergence rate is $|\hat{\alpha}_{t} - \alpha_{t}|=O_p(n^{-1/3})$ due to the boundary effect. Specifically, the estimated weights $\hat\pi_1(z)$ and $\hat\pi_0(z)$ depend on the density estimators at the boundary. Since we only require $f_{X, Z(0)}(x,z)\in \cF^-$ and $f_{X, Z(1)}(x,z)\in \cF^+$ to be smooth from one side, the corresponding density estimators have a slower rate. So, the plug-in error becomes the dominant term when establishing the rate of $\hat{\alpha}_{t}$. However, if $Z_i$ are the pre-treatment covariates, i.e, $Z_i(1)=Z_i(0)=Z_i$, we can estimate the density $f_{X,Z}(c,z)$ by
The following corollary shows that in this case $\hat{\alpha}_{t}$ has an improved rate. With an optimal choice of the bandwidth parameters, we prove that $|\hat{\alpha}_{t} - \alpha_{t}|=O_p(n^{-2/5})$, that is the boundary effect for estimating $\alpha_t$ is automatically removed without applying any additional bias correction procedures.
In this first setting, consider the following data generating process: \[y_i(1) = 2+x_i+\beta z_i + \epsilon_i,\] \[y_i(0) = 1+x_i+\beta z_i + \epsilon_i,\] where $x_i$, $z_i$, and $\epsilon_i$ are generated independently from $N(0,1)$ distribution. The treatment $T_i$ is assigned at the cutoff 0: $T_i = \bold{1}(x_i>0)$. In this case, there is no discontinuity of the conditional distribution of $z_i$ given $x_i=0$. Thus, our estimaind $\tau^{w1}_{SRD}$ equals the standard RD estimand $\tau_{SRD}$ (both are equal to 1). We vary $\beta$ from 0 to 5 and compare weighted local linear (WLL) estimator with the standard RD estimator (Imbens2008) in terms of bias, variance, root-mean-squared error (MSE), coverage probability of 95% confidence intervals (Coverage) and its length (CI length). When implementing both methods, we set the bandwidth parameter for standard RD estimator using cross-validation and then use the same bandwidth for WLL. The results based on 500 simulations are shown in Table (ref). When $\beta=0$, there is no covariates involved in the outcome function. Standard RD estimator performs neck to neck with our estimator. When $\beta\neq 0$, the standard RD estimator performs slightly better in terms of bias, however, our estimator consistently has smaller variance and MSE.
Figure (ref) compares the MSE of the standard RD estimator with our WLL estimator across different bandwidth choices. The figure is consistent with corollary (ref) as our estimator is asymptotically more efficient through including additional covariates into the estimation. Moreover, the advantage of WLL estimator is greater when perform under-smoothing and the difference of the two estimators becomes smaller as the bandwidth increases.
In the second setting, we consider the following data generating process: \[y_i(1) = 3+x_i+z_i + \epsilon_{1i},\] \[y_i(0) = 1+x_i+z_i + \epsilon_{0i},\] where $x_i$ and $\epsilon_i$ are generated independently from $N(0,1)$ distribution, however, $z_i$ is generated from another independent $N(0,1)$ process with a discontinuity at $X>0$, i.e. $z_i = \gamma \cdot \bold{1}(x_i>0)+z_i^*$, where $z_i^* \sim N(0,1)$. The treatment $T_i$ is assigned at the cutoff 0: $T_i = \bold{1}(x_i>0)$. When $\gamma=0$, again there is no discontinuity of the conditional distribution of $z_i$ given $x_i=0$. Both our estimaind $\tau^{w1}_{SRD}$ and the standard RD estimand $\tau_{SRD}$ are equal to 2. However, as $\gamma$ differs from 0, the conditional distribution of $z_i$ given $x_i$ is discontinuous at $x_i=0$. Our estimand $\tau^{w1}_{SRD}$ is still equal to 2, which is the direct causal effect of interest. If we adopt the standard RD framework and ignore the discontinuity of the conditional distribution of $z_i$ given $x_i$, we would expect that the standard RD estimator is biased for estimating the direct causal effect of interest (which is 2 in this example). In the data generating process, we vary $\gamma$ from 0 to 1 and compare our estimator with the standard RD estimator. The results are shown in Table (ref). The standard RD estimator has large bias and very poor coverage probability when $\gamma$ is close to 1, which agrees with our expectation. In contrast, the proposed estimator has relatively small MSE and accurate coverage probabilities across different choices of the sample size $n$ and the parameter $\gamma$. In summary, our simulation studies confirm that one should apply the proposed framework to the RD study if there exists some potential discontinuity of the conditional distribution of the covariates given the running variable.
We further apply our method to study the “incumbency advantage" in the U.S. House elections as in Lee2008. The “incumbency advantage" states that current incumbent party in a district are more likely to win the next. The dataset contains 6560 observations on elections to the United States House of Representatives (1946-1998). We evaluate the probability a Democrat both running in and winning election $t+ 1$ as a function of the Democratic vote share margin of victory in election $t$. The other covariates we are considering includes the Democratic vote share at time $t-1$, Democrat winning at $t-1$, Democrat's political experience, opponent's political experience, Democrat's electoral experience, opponent's electoral experience. This is the same setting as studied in Lee2008 where he uses local 4th order polynomial RD estimates and includes these additional covariates as robustness check. We apply our method in the same five settings and results are reported in Table (ref) below:
As shown in Table (ref), there is no significant change on the estimated effect when using our method. The main reason is that the conditional densities of covariates are truly continuous at the cutoff in this data set. We expect that both standard RD estimator and our method are valid and therefore the results are very similar.
Next we apply our method to study the generalized second price auction (GSP) problem. GSP is an auction mechanism for multiple items and it has been used widely for the assignment of advertisement positions by internet search engine like google and bing. Let $n$ be the number of bidders, and let $b^1 \geq b^2 \geq \cdots\geq b^n$ be the bids from high to low. Denote by $v_{(1)}, v_{(2)}, \cdots, v_{(n)}$ the bidders' valuation associated with the rank of bids and $r^k$ the click through rate for the $k$th position. The $k$th bidder's payoff in a GSP is given as $(v_{(k)} - b^{k+1})r^k$.
An important metrics is the click through rate $r^k$ for the $k$th position. Bidders are interested in the potential growth in their search traffic by winning the auction. And furthermore, in a Vickrey-Clarke-Groves (VCG) auction, $r^k$ will determine the total cost for placing each bidder in the sponsored advertisement region. In real world, GSP is usually implemented through a reservation score. A search score is formed for every bidder based on their bid and other quality measures. When the search score is bigger than a pre-set reservation score, the bidder's link will be displayed in the sponsored area. Otherwise, they will be displayed after all the sponsored advertisements. The reservation score cut-off creates a natural regression discontinuity setting to evaluate $r^k$.
We study the Microsoft Bing search data from Oct 2nd to Oct 22nd in 2015 and estimate the effect of advertisement positions. We focus on a set of searches with first advertisement positions displayed but without third advertisement position displayed.\footnote{Microsoft Bing allows a maximum of 4 advertisement to be displayed at the time of study.} This allows us to analyze the effect of second advertisement positions by comparing click-abilities for bidders near the search score cutoff. Figure (ref) plots the connection between search score and click-ability for the second position in the sponsored advertisement area. Once the score passed the cut-off at 15.17, the customers' links will be placed at the sponsored advertisement area and a significant increase in the click traffic can be observed.
Consider the bid as a covariate. Figure (ref) plots the mean bidding price before and after the search score cut-off. The mean bids before the cut-off is higher than the mean bids after the cut-off, implying the discontinuity of the conditional distribution of the covariate. See also Figure (ref) for the conditional density of the covariate before and after the cut-off. Although a local envy free equilibrium exists when bidders are all bidding their true valuation Edelman2007, information asymmetry or bidder inertia may still lead to bidder selections. For example, active bidders may have the incentive to bid more aggressively to take advantage of the bidders with high inertia.
Table (ref) presents the results of our estimator and a polynomial RD estimator. The classic RD estimator may not be valid in this case due to the discontinuity in the covariates. It estimates that placing the advertisement on the second position can increase the click-ability by 1.91% and it is statistically significant. On the other hand, the proposed estimator delivers only 1.20%, 57% less than the RD estimator and it is not statistically significant at 5% level. The difference is mainly because low quality bidders with high willingness to pay have the incentive to bid higher to move their search scores pass the threshold. But when we match bidders with similar bids before and after the cut-off, the effect goes away.
Fuzzy RD design refers to the case of a RD design when the treatment compliance is imperfect. Although $T_i$ is no longer a deterministic function of $X_i$, there is still a discontinuity in $\mathbb{P}(T_i|X_i, Z_i)$. Hahn2001 showed the equivalence between the fuzzy RD design and local average treatment effect (LATE). Under independence and monotonicity assumption, the cutoff can be used as an instrument for the treatment status and the LATE can be interpreted as an intention-to-treat effect on the compliers: subjects take treatment as their assigned one.
Introducing self-selection under the Fuzzy RD design is the same as adding an additional endogenous variable. As a result, additional instrumental variable is required for identification. We consider a scenario when the self-selection is based on the cut-off instead of the treatment assignment mechanism. This setting allows us to deal with unobservables that affect the covariates. Let $\tau$ denote the type of individual: complier ($co$), always-taker ($at$) and never taker ($nt$). A post-selection may exist and we can write \[Z_i = \bold{1}(X_i>c)Z_i(1) + \bold{1}(X_i\leq c)Z_i(0).\] Notice that in sharp RD design, the self-selection based on the cut-off is the same as the selection on treatment assignment mechanism since the assignment mechanism is a deterministic function of $X_i$. We study the selection based on the cut-off in fuzzy RD design as it is a more practical design. For example, students may notice that with SAT score higher than a threshold may increase their chance to receive scholarship, but the true assignment mechanism (gender, family income, etc.) may not be disclosed to them.
To consider self-selection under this setting, we first need to modify the independence and monotonicity assumption to incorporate additional covariates.
Due to self-selection, $Z_i$ is discontinuous at $X_i = c$, thus the independence assumption needs to be modified to insure individual types are remain constant under self-selection. Monotonicity assumption is essential to the identification by restricting the direction of selection. We may only have compliers, always-takers, never-takers in the population. As a result, we can view the treatment as an instrument as in Imbens1994.
By allowing types of individuals to depend on $Z_i$, our framework allows selection on treatment compliance to vary across $Z_i$. For example, individuals with family incomes less than a threshold are more likely to be always takers for the scholarship. An extreme scenario when $Z_i$ is binary would be that always takers all have $Z_i=1$. In a similar spirit to the sharp RD design, we consider the following three estimands under our WATE framework:
where $\mathcal{N}_\epsilon$ is a symmetric $\epsilon$ neighborhood around $c$.
Theorem (ref) proves the identification results for the three WATEs under the fuzzy RD design. The proof for the first two results are straightforward. The third result relies on the CIA in addition. This is because the sharp RD estimand is to ignore the conditioning on $X_i$, however, the fuzzy RD has to condition on $X_i$ to cancel with the denominator. The CIA is to directly remove the conditioning on $X_i$ and thus is necessary for the results to hold. The local linear estimators of these three estimands can be developed following the argument in Section (ref). However, the technical details and the theoretical results are much more involved, and are beyond the scope of this work. Given the practical importance of the fuzzy RD design, we think this is an important future research problem to explore.
In this paper we study the RD problem when the conditional distribution of the covariates given running variables is discontinuous at the cutoff, for example, due to self-selection. The standard RD design is no longer valid as the continuity of potential outcomes assumption is violated. We show that casual effect can still be recovered when the covariates related to self-selection are observed. We thus propose a set of estimands under the framework of WATE and show that these estimands can be estimated using a class of weighted local linear estimators. We derive the theory for our estimators include consistency and asymptotic normality. We further compare our estimator with the standard RD estimator in simulation exercises to demonstrate its finite sample performance.
We apply our estimator to two empirical examples. First, we study the U.S. House elections data in Lee2008, which is the classical data for RD estimator. Our estimator and local polynomial RD estimator have similar performance. Second, we apply our method to evaluate the effect of a GSP auction. We show that the result from our estimator is different from standard RD estimator due to self-selection. In particular, the obtained effect of advertising by using our estimator is smaller and statistically insignificant when self-selection is taking into account.