Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
74,481 characters · 10 sections · 71 citation commands
Testing Continuity of a Density via g-order statistics in the Regression Discontinuity Design
\thispagestyle{empty}
KEYWORDS: Regression discontinuity design, $g$-ordered statistics, sign tests, continuity, density.
JEL classification codes: C12, C14.
\setcounter{page}{1}
The regression discontinuity design (RDD) has been extensively used in recent years to retrieve causal treatment effects - see lee/lemieux:10 and imbens/lemieux:08 for exhaustive surveys. The design is distinguished by its unique treatment assignment rule where individuals receive treatment when an observed covariate, known as the running variable, crosses a known cut-off. Such an assignment rule allows nonparametric identification of the average treatment effect (ATE) at the cut-off, provided that potential outcomes have continuous conditional expectations at the cut-off hahn/etal:01. The credibility of this identification strategy along with the abundance of such discontinuous rules have made RDD increasingly popular in empirical applications.
While the continuity assumption that is necessary for nonparametric identification of the ATE at the cut-off is fundamentally untestable, researchers routinely assess the plausibility of their RDD by exploiting two testable implications of a stronger identification assumption proposed by lee:08. We can describe the two implications as follows: (i) the treatment is locally randomized at the cut-off, which translates into the distribution of all observed baseline covariates being continuous at the cut-off; and (ii) individuals have imprecise control over the running variable, which translates into the density of the running variable being continuous at the cut-off. The practice of judging the reliability of RDD applications by assessing either of the two above stated implications (commonly referred to as manipulation, or falsification, or placebo tests) is ubiquitous in the empirical literature. Indeed, Table (ref) surveys RDD empirical papers in four leading applied economic journals during the period 2011-2015. Out of 62 papers, 43 of them include some form of manipulation, falsification, or placebo test.
This paper proposes an approximate sign test for the null hypothesis on the second testable implication, i.e., the density of the running variable is continuous at the cut-off.\footnote{It is important to emphasize that the null hypothesis we test in this paper is neither necessary nor sufficient for identification of the ATE at the cut-off; see Remark (ref).} The approximate sign test has a number of distinctive attractive properties relative to existing methods used to test our null hypothesis of interest. First, the test does not require consistent non-parametric estimators of densities and simply exploits the fact that a certain functional of order statistics of the data is approximately binomially distributed under the null hypothesis. Second, our test controls the limiting null rejection probability under fairly mild conditions that, in particular, do not require existence of derivatives of the density of the running variable.\footnote{We use the term null rejection probability as opposed to asymptotic size as a way to acknowledge that, for a given sample size, there always exists a heavily steep smooth function that is indistinguishable from a discontinuous one. Remark (ref) discusses this further and provides important references.} In addition, our test is valid in finite samples under stronger, yet plausible, conditions. Third, the asymptotic validity of our test holds under two alternative asymptotic frameworks; one in which the number $q$ of observations local to the cut-off is fixed as the sample size $n$ diverges to infinity, and one where $q$ diverges to infinity slowly as $n$ diverges to infinity. Importantly, both frameworks require similar and arguably mild conditions. Fourth, our test is simple to implement as it only involves computing order statistics, a constant critical value, and a single tuning parameter. This contrasts with existing alternatives that require local polynomial estimation of some order and either bias correction or under-smoothed bandwidth choices. Finally, we have developed a companion \verb+Stata+ package to facilitate the adoption of our test.\footnote{The \verb+Stata+ package \verb+rdcont+ can be downloaded from \url{http://sites.northwestern.edu/iac879/software/}.}
The construction of our test is based on the simple intuition that, when the density of the running variable is continuous at the cut-off, the fraction of units under treatment and control local to the cut-off should be roughly the same. This means that the number of treated units out of the $q$ observations closest to the cut-off, is approximately distributed as a binomial random variable with sample size $q$ and probability $\frac{1}{2}$. To formalize this intuition, we exploit and develop properties of the so-called $g$-order statistics kaufmann/reiss:92,reiss:89 and consider the two asymptotic frameworks mentioned earlier to capture the local behavior of the density at the cut-off. In the first asymptotic framework, $q$ is fixed as $n\to \infty$ to represent a finite sample situation where the effective number of observations used by the test is too small to credibly invoke approximations for “large” $q$. This may arise, for example, when the density is not so well behaved around the cut-off as illustrated in some of our simulations. This framework is similar to the one in canay/kamat:18, who in turn exploit results from canay/romano/shaikh:17. It is worth noting that the hypothesis we test, the test statistic, the critical value, and most of the formal arguments are different from those in canay/kamat:18 or canay/romano/shaikh:17. In the second asymptotic framework, $q$ diverges to infinity slowly as $n\to \infty$ to represent a finite sample situation where the effective number of observations used by the test is large enough to invoke approximations for “large” $q$. This framework is similar to the one in mccrary:08,otsu/etal:13,cattaneo/jansson/ma:17,armstrong/kolesar:19, among others, and is in line with more traditional asymptotic arguments in non-parametric tests.
From a technical standpoint, this paper has several contributions relative to the existing literature. To start, our results exhibit two important differences relative to canay/kamat:18 that go beyond the difference in the null hypotheses. First, we do not study our test as an approximate randomization test but rather as an approximate sign test. This not only requires different analytical tools, but also by-passes some of the challenges that would arise if we were to characterize our test as an approximate randomization test; see Remark (ref) for a discussion on this. In addition, our approach in turn facilitates the analysis for the second asymptotic framework in which $q\to \infty$. Second, we develop results on $g$-order statistics as important intermediate steps towards our main results. Some of them may be of independent interest; e.g., Theorem (ref). In addition, relative to the results in mccrary:08,otsu/etal:13,cattaneo/jansson/ma:17; our test does not involve consistent estimators of density functions to either side of the cut-off and does not require conditions involving existence of derivatives of the density of the running variable local to the cut-off. To the best of our knowledge, the formal asymptotic results we present are original to this paper.
It is relevant to note that similar binomial tests have been recently proposed in the RDD literature by cattaneo/Tit/VB:16,cattaneo/Tit/VB:17 and frandsen:17. As we explain in more detail in Remark (ref), there are important differences between these binomial tests and ours when it comes to the null hypothesis being tested, the formal arguments, and the practical implementation of the tests. cattaneo/Tit/VB:16,cattaneo/Tit/VB:17 rely on finite sample arguments to justify their test construction for the hypothesis of local randomization. frandsen:17 also relies on finite sample arguments to test the hypothesis of manipulation of a discretely distributed running variable. In contrast, we test the hypothesis that the density of the running variable is continuous at the cut-off. Our focus on this particular null hypothesis prevents us from invoking finite sample arguments at the level of generality we consider and leads us to study the asymptotic properties of the approximate sign test. Our analysis also guides how to choose $q$ in data-dependent way and this, in turn, leads to a distinctive implementation of the test that we propose.
The remainder of the paper is organized as follows. Section (ref) introduces the notation and describes the null hypothesis of interest. Section (ref) defines $g$-order statistics, formally describes the test we propose, and discusses all aspects related to its implementation including a data-dependent way of choosing $q$. Section (ref) presents the main formal results of the paper, dividing those results according to the two alternative asymptotic frameworks we employ. In Section (ref), we examine the relevance of our asymptotic analysis for finite samples via a simulation study. Finally, Section (ref) implements our test to reevaluate the validity of the design in lee:08 and Section (ref) concludes. The proofs of all results can be found in the Appendix.
Let $Y\in \mathbf R$ denote the observed outcome of interest for an individual or unit in the population and $A\in \{0,1\}$ denote an indicator for whether the unit is treated or not. Further denote by $Y(1)$ the potential outcome of the unit if treated and by $Y(0)$ the potential outcome if not treated. As usual, the observed outcome and potential outcomes are related to treatment assignment by the relationship
The treatment assignment in the (sharp) RDD follows a discontinuous rule,
where $Z\in \mathcal Z\equiv \operatorname*{supp}(Z)$ is an observed scalar random variable known as the running variable and $\bar{z}$ is the known threshold or cut-off value. For convenience we normalize $\bar{z}=0$, which is without loss of generality as we can always redefine $Z$ as $Z-\bar{z}$. This treatment assignment rule allows us to identify the average treatment effect (ATE) at the cut-off; i.e.,
In particular, hahn/etal:01 establish that identification of the ATE at the cut-off relies on the discontinuous treatment assignment rule and the assumption that
Reliability of the RDD thus depends on whether the mean outcome for units marginally below the cut-off identifies the true counterfactual for those marginally above the cut-off.
The continuity assumption in (ref) is arguably weak, but fundamentally untestable. In practice, researchers routinely employ two specification checks in RDD that, in turn, are testable implications of a stronger sufficient condition proposed by lee:08. The first check involves testing whether the distribution of pre-determined characteristics (conditional on the running variable) is continuous at the cut-off. See shen/zhang:16 and canay/kamat:18 for a recent treatment of this problem. The second check involves testing the continuity of the density of the running variable at the cut-off, an idea proposed by mccrary:08. This second check is particularly attractive in settings where pre-determined characteristics are not available or where these characteristics are likely to be unrelated to the outcome of interest. Formally, we can state the hypothesis testing problem for the second check as
where $f_Z^{+}(0)$ and $f_Z^{-}(0)$ are the one-sided limits of the probability density function of $Z$, i.e.,
In RDD empirical studies, the aforementioned specification checks are often implemented (with different levels of formality) and referred to as falsification, manipulation, or placebo tests (see Table (ref) for a survey).
In this paper we consider an approximate sign test for the null hypothesis of continuity in the density of the running variable $Z$ at the cut-off $\bar{z}=0$, i.e., (ref). This test has three attractive features compared to existing approaches mccrary:08,otsu/etal:13,cattaneo/jansson/ma:17. First, it does not require commonly imposed smoothness conditions on the density of $Z$, as it does not involve non-parametric estimation of such a density. Second, it exhibits finite sample validity under certain (stronger) easy to interpret conditions. Finally, it involves a single tuning parameter cattaneo/jansson/ma:17 as opposed to multiple ones in mccrary:08. We discuss these features further in Section (ref).
Let $P$ be the distribution of $Z$ and $Z^{(n)}=\{Z_i:1\le i\le n\}$ be a random sample of $n$ i.i.d.\ observations from $P$. Let $q$ be a small (relative to $n$) positive integer and $g:\mathcal Z \to \mathbf R$ be a measurable function such that $g(Z)$ has a continuous distribution function. For any $z,z'\in \mathcal Z$ define $\le_{g}$ as
The ordering defined by $\le_{g}$ is called a $g$-ordering on $\mathcal Z$. The $g$-order statistics $Z_{g,(i)}$ corresponding to $Z^{(n)}$ are defined as the values satisfying
see, e.g., reiss:89 and kaufmann/reiss:92.
To construct our test statistic, we use the sign of the $q$ values of $\{Z_i:1\le i\le n\}$ that are induced by the $q$ smallest values of $\{g(Z_i)=|Z_i|:1\le i\le n\}$. That is, for $Z_{g,(1)},\dots,Z_{g,(q)}$, let
and
The test statistic of our test only depends on the data via $S_n$ and is defined as
In order to describe the critical value of our test it is convenient to recall that the cumulative distribution function (CDF) of a binomial random variable with $q$ trials and probability of success $\frac{1}{2}$ is given by
where $\lfloor x \rfloor$ is the largest integer not exceeding $x$. Using this notation the critical value for a significance level $\alpha\in (0,1)$ is given by
where $b_q(\alpha)$ is the unique value in $\{0,1,\ldots ,\lfloor \frac{q}{2}\rfloor \} $ satisfying
The test we propose is then given by
where
Intuitively, the test $\phi(S_n)$ exploits the fact that, under the null hypothesis in (ref), the distribution of the treatment assignment should be locally the same to either side of the cut-off. That is, local to the cut-off, the treatment assignment behaves as purely randomized under the null hypothesis, so the fraction of units under treatment and control should be similar.
Given $q$, the implementation of our test proceeds in the following five steps.
In this section we discuss the practical considerations involved in the implementation of our test, highlighting how we addressed these considerations in the companion \verb+Stata+ package. The only tuning parameter of our test is the number $q$ of observations closest to the cut-off. We propose a data-dependent way to choose $q$ that combines a rule of thumb with a local optimization. We call this data-dependent rule the “informed rule of thumb” and its computation requires the two steps described below. For the sake of clarity, in this section we do not use the normalization $\bar{z}=0$. Additional computational details are presented in Appendix (ref).
In Section (ref) we consider the asymptotic framework where $q$ diverges as $n\to\infty$. Under Assumption (ref) and $H_0$ in (ref), we show in that section that the value of $q$ that sets the worst case asymptotic bias equal to the standard deviation is given by
where $f_Z(\bar{z})$ equals $f_Z^{+}(\bar{z})=f_Z^{-}(\bar{z})$ under $H_0$, and $C_P$ is the Lipschitz constant in Assumption (ref)(i'). Since the results in Theorem (ref) also require $q^{3/2}/n\to 0$, we propose to start with an initial rule of thumb where $f_Z(\bar{z})$ and $C_P$ are computed under the assumption $Z\sim N(\mu, \sigma^2)$ and the rate is set to $n^{1/2}$. This leads to
where we used that $C_P = |\phi'_{\mu,\sigma}(\mu+\sigma)| = \frac{1}{\sigma}\phi_{\mu,\sigma}(\mu+\sigma)$ when $Z\sim N(\mu,\sigma^2)$, and $\phi_{\mu,\sigma}(\cdot)$ and $\phi'_{\mu,\sigma}(\cdot)$ denote the density of $N(\mu,\sigma^2)$ and its derivative. This initial rule of thumb is location and scale invariant and, by definition, is inversely related to the asymptotic bias of the test statistic in the asymptotic framework of Section (ref). In turn, the constant multiplying $n^{1/2}$ in (ref) is fairly intuitive. First, it captures the idea that a steeper density at the cut-off should be associated with a smaller value of $q$. Intuitively, the steeper the density, the more it resembles a density that is discontinuous (Figure (ref).(c) illustrates this in Section (ref)). Since the maximum slope is determined by the Lipschitz constant, the rule is inversely proportional to that. Second, it also captures the idea that $q$ should be small if the cut-off is a point of low density. Intuitively, when $f_Z(\bar z)$ is low, the $q$ closest observations to $\bar{z}$ are likely to be “far” from $\bar z$ (Figure (ref).(a) with $\mu=-2$ illustrates this in Section (ref)). One could alternatively replace the normality assumption with a non-parametric estimator of $f_Z(\bar{z})$ but it is unfortunately impossible to choose $C_P$ adaptively for testing (ref) low:97,armstrong/kolesar:18. Since any data-dependent rule for $q$ will require a reference for $C_P$, we prefer to prioritize its simplicity and use normality for both $f_Z(\bar{z})$ and $C_P$.
The second step involves a local maximization of the asymptotic null rejection probability of the non-randomized version of the test. In particular, based on our results, we propose
where $\Psi_q(\cdot)$ is the CDF defined in (ref), $b_q(\alpha)$ is defined in (ref), and $\mathcal N(q_{\rm rot})$ is a discrete neighborhood defined in (ref) in Appendix (ref). This step helps the performance of the non-randomized version of the test (see Remark (ref)) as $ \Psi_q(b_q(\alpha)-1)$ is non-monotonic in $q$ (see Figure (ref)) and so optimizing locally to $q_{\rm rot}$ over $\mathcal N(q_{\rm rot})$ prevents choosing a value of $q$ with a low value of $\Psi_q(b_q(\alpha)-1)$. In practice, we replace $(\mu,\sigma)$ with sample analogs to obtain the feasible informed rule of thumb $\hat q_{\rm irot}$.
In this section we derive the asymptotic properties of the test in (ref) using two alternative asymptotic frameworks. The first one requires $q$ to be fixed as $n\to \infty$, and represents a finite sample situation where the effective number of observations used by the test is too small to credibly invoke approximations for “large” $q$. The second framework requires $q\to \infty$ slowly as $n\to \infty$, and represents a finite sample situation where the effective number of observations used by the test is large enough to invoke approximations for “large” $q$.
There are three main features of our results that are worth highlighting: (i) our test exhibits similar properties under both asymptotic frameworks, (ii) the implementation of the test does not depend on which asymptotic framework one has in mind, and (iii) all formal results require similar, and arguably mild, conditions. We start by introducing these conditions.
Assumptions (ref)(i) and (ref)(i') each impose different degrees of smoothness on the density of $Z$ local to the cut-off $\bar{z}=0$. Indeed, Assumption (ref)(i') strengthens Assumption (ref)(i) by replacing the requirement of left- and right-continuity at the cut-off with its Lipschitz version. In the formal results that follow, we use Assumption (ref)(i) in the asymptotic framework where $q$ is fixed as $n\to \infty$ and Assumption (ref)(i') in the asymptotic framework where $q\to \infty$ as $n\to \infty$. Both assumptions allow for the distribution of $Z$ to be discontinuous outside of a neighborhood of the cut-off.\footnote{In Appendix (ref) we also allow for situations with a mass point at the cut-off, i.e., $P\{Z=0\}>0$.} More importantly, they do not require the density of $Z$ to be differentiable anywhere. This is in contrast to mccrary:08, who requires three continuous and bounded derivatives of the density of $Z$ (everywhere except possibly at $\bar{z}=0$), and cattaneo/jansson/ma:17 and otsu/etal:13, who require the density of $Z$ to be twice continuously differentiable local to the cut-off (in the case of a local-quadratic approximation). Assumption (ref)(ii) rules out a situation where $f^{-}_{Z}(0)=f^{+}_{Z}(0)=0$, which is implicitly assumed away in mccrary:08 and otsu/etal:13, and is weaker than assuming a positive density of $Z$ in a neighborhood of the cut-off as in cattaneo/jansson/ma:17. In Section (ref) we explore the sensitivity of our results to violations of these conditions.
In this section we present two main results. The first result, Theorem (ref), describes the asymptotic properties of $S_n$ in (ref) when $q$ is fixed as $n\to\infty$. This result about $g$-order statistics with $g(\cdot)=|\cdot|$ represents an important milestone in proving the asymptotic validity of our test. The second result, Theorem (ref), exploits Theorem (ref) to show that the test in (ref) controls the limiting rejection probability under the null hypothesis.
Theorem (ref), although fairly intuitive, does not follow from standard arguments. First, the random variables $\{A_{g,(j)}:1\le j\le q\}$ are indicators of $g$-order statistics so, in general, they are neither independent nor identically distributed. Second, applying results from the literature on $g$-order statistics kaufmann/reiss:92 requires $g(Z)=|Z|$ to have a continuous distribution function everywhere on its domain. Under Assumption (ref)(i) this is only true in $[0,\delta)$, and mass points are allowed outside of $[0,\delta)$. In the proof of Theorem (ref) we use a smoothing transformation of $Z$ as an intermediate step and then accommodate the results in kaufmann/reiss:92 to reach the desired conclusion.
The following result, which heavily relies on Theorem (ref), is the main result of this section and characterizes the asymptotic properties of the test $\phi(S_n)$ in (ref).
Theorem (ref) shows that $\phi(S_n)$ behaves asymptotically, as $n\to \infty$, as the two-sided sign test in an experiment where one observes $S\sim {\rm Bi}(q,\pi)$ and wishes to test the hypotheses $H_0:\pi=\frac{1}{2}$ versus $H_1:\pi\ne \frac{1}{2}$. For this reason, we refer to $\phi(S_n)$ as an approximate sign test.
In this section we study the properties of $\phi(S_n)$ in (ref) in an asymptotic framework where $q$ diverges to infinity as $n\to \infty$. This asymptotic framework is in line with traditional non-parametric arguments and so our results depend on the assumed smoothness of the density of $Z$ and the rate at which $q$ is allowed to grow. Importantly, the results in this section follow from Assumption (ref)(i')-(ii) and so, accounting for the differences between Assumptions (ref)(i) and (ref)(i'), the result below shows that the asymptotic properties of the approximate sign test under both asymptotic frameworks require similar, and arguably mild, conditions.
Theorem (ref), although fairly intuitive again, does not follow from standard arguments. In particular, given that the random variables $\{A_{g,(j)}:1\le j\le q\}$ are neither independent nor identically distributed, the result does not follow from a simple application of the central limit theorem. We instead adapt kaufmann/reiss:92 and prove the result using first principles and the normal approximation to the binomial distribution.
Given the result in Theorem (ref), we can provide some insight on the properties of the data-dependent rule for choosing $q$ that we describe in Section (ref). Specifically, we focus on providing interpretation to $q_{\rm rot}$ in (ref), as $q_{\rm irot}$ in (ref) is a modification of $q_{\rm rot}$ to improve the performance of the non-randomized version of the test. Under $H_0$ in (ref) and Assumption (ref)(i')-(ii), the results in armstrong/kolesar:19 imply that
where $\zeta_n \overset{d}{\to} \zeta \sim N(0,1)$ and $B_{n,q}$ is a standardized bias term satisfying
with $f_Z(0)$ equals $f_{Z}^{+}(0)=f_{Z}^{-}(0)$ under $H_0$. Denote by $t^{\ast}$ the right-hand side of (ref) and note that this can be interpreted as the worst (in absolute value) ratio of bias to standard deviation (sd) of the left-hand side of (ref). We can then solve for $q$ to obtain
This derivation shows that the requirement $\frac{q^{3/2}}{n}\to 0$ is analogous to under-smoothing as this is the rate condition that removes the worst-case asymptotic bias.\footnote{A previous version of this paper did not include Assumption (ref)(i') and the requirement $q^{3/2}/n\to 0$, which is required to control the asymptotic bias term. We thank Tim Armstrong for pointing this out to us.} This immediately gives two alternative interpretations to the data-dependent rule $q_{\rm rot}$ in (ref) armstrong/kolesar:19. In order to describe these two interpretations, note that by (ref) and (ref) we obtain that $q_{\rm rot}=q^*$ whenever
Assume for a moment that the rule-of-thumb assumption of normality is correct (which means that the term within brackets in (ref) equals 1). Then, $q_{\rm rot}$ is equivalent to $q^{\ast}$ for a worst ratio of bias to sd $t^{\ast}$ given by
This implies the size of $\phi(S_n)$ for $\alpha=5\%$ and $n=5,000$ would approximately be $P\{|\zeta+0.12|>z_{\alpha/2}\} =5.16\%$. In this sense, $q_{\rm rot}$ makes the size distortion of the bias negligible when $n=5,000$. Next, suppose that the rule-of-thumb assumption of normality over-estimates the ratio $f^2_Z(0)/C_P$. In other words, suppose that the term within brackets in (ref) equals a constant $a>1$. In this case, $q_{\rm rot}$ would be equivalent to $q^{\ast}$ for a worst ratio of bias to sd $t^{\ast}$ given by
This implies that the size of $\phi(S_n)$ for $\alpha=5\%$ and $n=5,000$ would approximately be $P\{|\zeta+0.36|>z_{\alpha/2}\} =6.38\%$. When $f_Z(0)=\phi_{\mu,\sigma}(0)$, this means that even if the true Lipschitz constant $C_P$ is three times larger than the one imposed by normality, $\phi(S_n)$ would still exhibit mild over-rejection under the null hypothesis. The price we pay for this robustness under the null hypothesis (in terms of performance and mild requirements) is possibly a lower power under the alternative hypothesis, a feature that we explore in the simulations of Section (ref).
In this section we examine the finite-sample performance of the test in (ref) with a simulation study. Instead of just presenting designs where this test excels relative to competing ones, we present an array of data generating processes that hopefully illustrate its relative strengths and weaknesses. The data for the study are simulated as i.i.d.\ samples from the following designs.
Design 1 in Figure (ref)(a) is the canonical normal case and, by Remark (ref), our test is expected to control size in finite samples when $\mu=0$ but not when $\mu\in \{-2,-1\}$. Indeed, $\mu=-2$ is a challenging case due to the low probability of getting observations to the right of the cut-off. Design 2 in Figure (ref)(b) is taken from canay/kamat:18. Design 3 in Figure (ref)(c) is a parametrization of the taxable income density in saez:10. This design exhibits a spike (almost a kink) to the left of the cut-off which is essentially a violation of the smoothness assumptions required by mccrary:08 and cattaneo/jansson/ma:17. It also exhibits a steep density at the cut-off, which also makes it a difficult case in general. Similar to Design 3, Design 4 in Figure (ref)(d) also illustrates the difficulty in distinguishing a discontinuity from a very steep slope; see low:97, kamat:17, armstrong/kolesar:18, and bertanha/moreira:2019 for a formal discussion. Here we can study the sensitivity to the slope by changing the value of $\kappa$. Design 5 in Figure (ref)(e) requires $\delta$ in Assumption (ref)(a) to be such that $\delta<\kappa$ in order for our approximations to be accurate, but as opposed to Design 4, it is locally symmetric around the cut-off. As $\kappa$ gets smaller, we expect our test to perform worse if $q$ is not chosen carefully. Finally, Design 6 in Figure (ref)(f) draws data i.i.d.\ from the non-parametric density estimate of the running variable in lee:08, i.e., $Z$ is the difference in vote shares between Democrats and Republicans.
We consider sample sizes $n \in \{1,000; 5,000 \}$, a nominal level of $\alpha=10\%$, and perform $10,000$ Monte Carlo repetitions. Designs 1 to 6 satisfy the null hypothesis in (ref). We additionally consider the same models under the alternative hypothesis by randomly changing the sign of observations in the interval $z\in [0,0.1]$ with probability $\Pr=0.2-2z$. We report results for the following tests.
Tables (ref) and (ref) report rejection probabilities under the null and alternative hypotheses for the six designs we consider and for sample sizes of $n=1,000$ and $n=5,000$, respectively. We start by discussing the results under the null hypothesis. AS-NR delivers rejection probabilities under the null hypothesis closer to the nominal level than those delivered by McC and CJM in most of the designs. The two empirically motivated designs (Designs 3 and 6) illustrate this feature clearly. Designs 4 and 5 also show big differences in performance, both in cases where AS-NR delivers rejection rates equal to the nominal level (Design 5) and McC and CJM severely over-reject; as well as in cases where all tests over-reject (Design 4, $\kappa=0.05$) but AS-NR is relatively closer to the nominal level. A particularly difficult case for AS-NR is Design 1 with $\mu=-2$, where the probability of getting observations to the right of the cut-off is below $2\%$. This showcases the satisfactory performance of our data-dependent rule $\hat q_{\rm irot}$, which takes the lowest value in that particular design. Tables (ref) and (ref) also show negligible differences between the randomized (AS-R) and non-randomized (AS-NR) versions of our test, consistent with our discussion in Remark (ref).
To describe the performance of the different tests under the alternative hypothesis, we focus on designs where the rejection probability under the null hypothesis is close to the nominal level for all tests: Design 1, Design 2, Design 4 with $\kappa = 0.25$, and Design 6. In those cases, we see that AS-NR has competitive power, and can sometimes even be the test with the highest rejection probability under the alternative hypothesis. For $n=1,000$, AS-NR delivers the highest rejection probability under the alternative hypothesis in Design 1 for all values of $\mu$ and Design 4 with $\kappa=0.25$. In the rest of the cases under consideration, McC exhibits the highest power and is followed by AS-NR. The results for $n=5,000$ are qualitatively similar, with a few exceptions. McC has the highest rejection probability in Design 1 with $\mu=-1$, and CJM are have the second highest rejection probability in Design 2 with $\lambda = \frac{1}{3}.$
Table (ref) shows the mean values of $\hat q_{\rm irot}$ across simulations for all designs and sample sizes. As described in Section (ref), $\hat q_{\rm irot}$ takes into account both the slope and the magnitude of the density at the cut-off. As a result, $\hat q_{\rm irot}$ is relatively high in designs with flat density at the cut-off and high $f_Z(0)$ (e.g., Design 1 with $\mu=0$) and relatively low in designs with steep slopes or low $f_{Z}(0)$ (e.g., Design 1 with $\mu=-2$ or Design 2 with $\lambda=1$). Table (ref) also reports the average number of observations in $[-h_{n,L},h_{n,R}]$ for McC and CJM, where $h_{n,L}$ and $h_{n,R}$ are the left and right bandwidths used to estimate $f^{-}_Z(0)$ and $f^{+}_Z(0)$, respectively (in the case of McC, $h_{n,L}=h_{n,R}$). In comparison, AS-NR uses significantly fewer observations than either McC or CJM, a feature that may support the the asymptotic framework in Section (ref). Finally, and to gain further insight on the sensitivity of our test to the choice of $q$, Figure (ref) displays the rejection probabilities of AS-NR and AS-R as a function of $q$ in two types of designs. In the top row we illustrate two designs where the rejection probability is mostly insensitive to the choice of $q$ (Design 1 with $\mu=0$ and Design 6). These are designs where the density is rather flat around the cut-off so increasing $q$ does not deteriorate the performance of our test. In the bottom row we illustrate two designs where the rejection probability is highly sensitive to the choice of $q$ (Design 1 with $\mu=-1$ and Design 3). These are designs that feature a steep density at the cut-off (also low in Design 1) so increasing $q$ very quickly deteriorates the performance of the test under the null hypothesis. The data-dependent rule $\hat q_{\rm irot}$ is displayed in each case with a vertical dashed line and seems to be doing a good job at choosing relatively smaller values in the sensitive cases.
We conclude this section with two final remarks. First, one could compare the results in Tables (ref) and (ref) for a fix value of $q$ to appreciate the results in Section (ref). For example, taking $q=75$, the rejection probability in Design 1 with $\mu=-2$ and Design 3 are $99.8$ and $40.1$, respectively, when $n=1,000$. The same numbers when $n=5,000$ are $26$ and $8.4$, respectively, which are closer to the nominal level as predicted by our results. Second, at the request of a referee, the results for $n=1,000,000$ and $\alpha=1\%$ are available upon request. Notably, AS-NR with the data-dependent rule $\hat q_{\rm irot}$ delivers rejection probabilities under $H_0$ equal to $\alpha$ across all designs when $n=1,000,000$ whereas McC and/or CJM still significantly over-reject for Designs 3, 5, and 6.
In this section we briefly reevaluate the validity of the design in lee:08. Lee studies the benefits of incumbency on electoral outcomes using a discontinuity constructed with the insight that the party with the majority wins. Specifically, the running variable $Z$ is the difference in vote shares between Democrats and Republicans in a house election; see Figure (ref)(f) for a graphical illustration of the density of $Z$. The assignment rule then takes a cut-off value of zero that determines the treatment of incumbency to the Democratic candidate, which is used to study their outcome in the next election. The total number of observations is 6,559 with 2,740 below the cut-off. The dataset is publicly available at \url{http://economics.mit.edu/faculty/angrist/data1/mhe}.
Lee assesses the credibility of the design in this application by inspecting discontinuities in means of the baseline covariates, but mentions in footnote 19 the possibility of using the test proposed by mccrary:08. Here, we frame the validity of the design in terms of the hypothesis in (ref) and use the approximate sign test we describe in Section (ref), using $\hat q_{\rm irot}$ as our default choice for the number of observations $q$. This test delivers a $p$-value of $0.55$ for $S_n = 73$ out of $\hat q_{\rm irot}=138$ observations. The null hypothesis of continuity of the density is therefore not rejected.
This paper presents an approximate sign test based on $g$-order statistics for testing the continuity of a density at a point in RDD. We study the properties of this test under two asymptotic frameworks; one in which the number $q$ of observations local to the cut-off is fixed as the sample size $n$ diverges to infinity, and one in which $q$ diverges to infinity slowly as $n$ diverges to infinity. We show that the test has limiting rejection probability under the null hypothesis not exceeding the nominal level in both asymptotic frameworks under similar and arguably mild conditions. More importantly, our test is easy to implement, asymptotically valid under weaker conditions than those used by competing methods, exhibits finite sample validity under stronger conditions than those needed for its asymptotic validity, and delivers competitive power in simulations.
A final aspect we would like to highlight of our test is its simplicity. The test only requires to count the number of non-negative observations out of the $q$ observations closest to the cut-off (this is all that is required to compute the p-value in (ref)), and does not involve kernels, local polynomials, bias correction, or bandwidth choices. Importantly, we have developed the \verb rdcont \verb+Stata+ package that allows for effortless implementation of the test we propose in this paper.