Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
101,579 characters · 23 sections · 12 citation commands
Optimal Model Selection in RDD and Related Settings Using Placebo Zones
\newgeometry{left=1.25in,right=1.25in,bottom=1.5in,top=1.5in}
Policy rules frequently create discontinuities in exposure to policies and programs. Regression discontinuity design (RDD) has become a key tool for empirical researchers in these settings \shortcite<see e.g.>[for overviews]{imbens_lemieux_2008,lee_lemieux_2010,cattaneo_idrobo_titiunik_2020}. In the canonical sharp RDD case, the treatment $T$ changes discontinuously from $T=0$ to $T=1$ at some threshold along the running variable $X$. Setting $x=0$ as that threshold, the goal is to estimate the change in the outcome $Y$ at $x=0$:
$\tau(x)$ is commonly estimated by local polynomial regression. Researchers select some neighbourhood of observations around $x=0$ (the bandwidth) where $E[Y_{T=0}|X=x]$ and $E[Y_{T=1}|X=x]$ are expected to meet the continuity assumption hahn_etal_2001 and estimate the jump in $Y$ while flexibly controlling for $X$ above and below $x=0$.
RDD is appealing because it facilitates estimation of causal effects under relatively weak assumptions. Moreover, the assumptions for RDD have simple, testable implications \shortcite<see e.g.>{mccrary_2008,cattaneo_jansson_ma_2019}. The ability to visualize RDDs in simple plots of the running variable and outcome calonico_cattaneo_titiunik_2015 also gives an appealing air of transparency to this approach. A number of related estimators extend the basic RDD. The regression kink design (RKD) identifies causal effects by exploiting discontinuous changes in the slope of the running variable under similar assumptions to RDD card_etal_2015. Fuzzy RDD and RKD deal with situations where only the probability of treatment changes at the threshold, or treatment is a continuous variable. \citeA{dong_2018} suggests a regression probability jump and kink design (RPJKD) for settings where there is both a discontinuous `jump' and/or `kink'. In many related settings it is also possible to fit a global polynomial through the running variable and instrument the treatment using binned means (cohort-IV), as in \citeA{angrist_lavy_1999}.\footnote{The cohort-IV approach can also be used in related situations where there is no clear discontinuity, and yet treatment is a non-smooth function of the running variable. See for example \citeA{imben_van_der_klaaw_1995}, \citeA{bound_turner_2002} and \shortciteA{cousley_etal_2017}.}
While theoretically appealing, when using discontinuity designs researchers face a daunting challenge in selecting a preferred estimator. The choice of bandwidth involves a difficult trade-off between bias and variance. Researchers must also choose what order of polynomial to use, what kernel to use, whether to include covariates, and in some applications what discontinuity model to estimate (e.g. in situations where there is a `jump' and a `kink'). Sometimes it is also useful to adopt different polynomial orders on the left and the right of the threshold, or different bandwidths. Consequently, researchers will typically have thousands of potential estimators to select from, and there is no widely accepted standard for making this choice. In a given application, estimates may vary widely depending on the choices the researcher makes.
Various solutions to model selection have been suggested, but these typically focus on one decision and fix other important decisions. Optimal bandwidth selection has received a lot of attention. In economics, least squares cross validation and plug-in approaches have dominated imbens_lemieux_2008. Cross validation methods typically select a bandwidth to minimize the mean squared error of the local polynomial fit. \citeA{ludwig_miller_2005,ludwig_miller_2007} discuss an alternative approach that minimizes error at the boundary, although ultimately reject this method for their application. Early plug-in approaches also focused on the polynomial function's fit, for example the rule-of-thumb approach discussed in \citeA{fan_gijbels_1996} \cite<see also>{lee_lemieux_2010}. In an influential paper, \citeA{imbens_kalyanaraman_2012} (IK) argue that instead of focusing on the global fit of the polynomial function, the bandwidth should minimize the asymptotic mean squared error of the treatment effect (boundary) estimator \cite<see also>{ludwig_miller_2005}. They derive a plug-in algorithm that selects the optimal bandwidth to achieve this. \shortciteA{calonico_cattaneo_titiunik_2014} (CCT) derive a bias correction to IK's method to improve confidence interval estimation. The IK/CCT approach is now popular in applied work.
\shortciteA{card_etal_2017}, however, caution against using the IK/CCT approach as a default and demonstrate through simulations that it may not perform best for a given application.\footnote{\shortciteA{card_etal_2017} suggest that the regularization term used in the IK/CCT plug-in approach may be overly punitive to large bandwidths in practice. The regularization term is used to account for the fact that the curvature parameters for the polynomial fits -- which are parameters themselves in the plug-in formula -- are unknown and must be estimated from the data.} Further, none of these approaches deals with the simultaneous modelling choices researchers need to make. For example, the optimal bandwidth will almost always depend on the polynomial order.\footnote{\citeA{hall_racine_2015} propose a leave-one-out cross validation approach that jointly selects the bandwidth and polynomial order. Their cross validation approach is subject to the issues discussed in IK.} There has been less theoretical development on the question of polynomial choice. \citeA{gelman_imbens_2019} argue that researchers should generally use local linear or quadratic regressions because higher order terms can induce undesirable effects on the estimates.\footnote{\citeA{gelman_imbens_2019} point out that higher order terms can have the practical effect of giving disproportional weighting to certain observations, are generally not selected on the basis of optimizing the objective of boundary estimation, and can lead to misleading inference.} \shortciteA{pei_etal_2020} are less critical of higher order terms and suggest that, conditional on a given bandwidth and other modelling choices, researchers should calculate the implied asymptotic mean squared error for the boundary estimator (similar in spirit to IK/CCT for bandwidth selection). They also show, through a review of recent literature, that most researchers simply default to using local linear estimation.
To understand how researchers are dealing with the challenges of model selection in discontinuity designs, we conducted a review of papers published in leading journals for applied economics research in 2019.\footnote{See \citeA{kettlewell_siminski_2020} Appendix D for further details.} Of the 26 papers we identified, 12 gave no formal rationale for their preferred bandwidth. Of those that did motivate their choice, 13 used IK/CCT; however, many of these merely used the method to `guide' their choice (e.g. by noting that the IK/CCT bandwidth was similar to whatever bandwidth they ultimately used). Almost all studies conducted some kind of sensitivity testing by varying the bandwidth. Only five studies provided any justification for their chosen order of polynomial. Most studies (18) used local linear regression and typically added higher order terms as a robustness check.\footnote{Other common robustness checks included adding covariates, different kernels, `donuts' around the threshold and falsification tests using placebo cut-off points.}
Overall, we surmise that there is no consensus among applied researchers about how to select a preferred model in discontinuity settings. In many cases, researchers seem to be selecting a baseline model either arbitrarily or based on possible defaults like local linear regression. The focus away from emphasizing a preferred model and towards sensitivity analysis may be problematic in certain applications. Further, if there is no clear preferred model, there is also no clarity around the confidence interval for the treatment effect. More generally, the emphasis on robustness tests means that we may be `setting the bar too high' for what constitutes credible evidence in RDD and related contexts.
In this paper we propose a new method for model selection with broad application. Our method allows researchers to select an optimal combination of bandwidth, polynomial, and any other choice parameters he/she wants to consider. It can also be used to choose between competing models (e.g. RDD versus cohort-IV) in certain settings and can accommodate nonlinear models. It relies on using observations of the running variable away from the discontinuity (the placebo zone) as a training ground to assess the performance of candidate models where a `pseudo-treatment' effect is known to be zero. The estimator that minimizes the preferred performance criterion (e.g. lowest root mean squared error) across all pseudo-treatments is then selected as the `best' specification for estimating the actual treatment effect. Our approach is applicable in settings where the point of discontinuity can be reasonably thought of as being randomly chosen from the domain of the running variable.\footnote{While we focus on applications of linear models in the paper, our method naturally extends to nonlinear models like logit and ordered logit \cite<see>[for an alternate IK-type approach to bandwidth selection for categorical dependent variables]{xu_2017}. In nonlinear models the marginal effect for the treatment indicator when $x = 0$ could be used as the target parameter for RMSE minimization, for example.}
We are not the first to recognize the value in placebo zone data. \citeA{imbens_lemieux_2008} suggest testing for jumps at specific psuedo-thresholds as a general test for specification error, a common practice in applied work. \citeA{wing_cook_2013} use the placebo zone to create a kind of differences-in-differences structure, which they argue can improve precision and allow one to learn something about the treatment effect away from the threshold. \citeA{gelman_imbens_2019} use results from the distributions of placebo estimates to inform general advice about higher order terms in RDD studies. Closest to our work is \citeA{ganong_jager_2018}, who suggest a randomization inference approach to hypothesis testing based on the distribution of pseudo-treatment effect estimates (we propose extensions to this procedure). We extend all of this work by using the placebo zone for ex ante model selection, which to the best of our knowledge is a new idea.
Our approach also has parallels with studies that use estimates from randomized controlled trials (RCTs) to assess RDD estimators. A prominent example is \shortciteA{hyytinen_etal_2018} who use one such RCT-RDD pair, and conclude that CCT bias-corrected estimators perform well in that application.\footnote{See \shortciteA{chaplin_etal_2018} for a review and meta-analysis of similar studies.} In this literature, as in our approach, the assessment rests on knowing the true parameter that the RDD estimator targets. In that literature, the target is the estimate generated by an RCT. In our case, the target estimate is zero, since there is no actual treatment. Our approach builds on the RCT-RDD approach in three important ways. First, rather than making a single comparison of RDD to RCT estimates, our approach assesses each candidate estimator’s performance repeatedly -- at hundreds or thousands of placebo thresholds throughout the placebo zone. Second, these comparisons serve to inform the choice of estimator to apply within the same context, to estimate the effect of a real treatment using the same data, in a range of the running variable that borders the placebo zone. There is no reason to believe that the best-performing estimator will perform best in other unrelated contexts where the DGP may be completely different, or with other sources of data. Thirdly, the target parameter in the RCT-RDD literature is subject to sampling bias, whereas the placebo-zone target of zero is known with certainty.
Finally, our approach positions point-estimation as a principal goal of model selection in RDD. As emphasised by \citeA{cattaneo_vazquez-bare_2016}, the choice of method for (point) estimation should be seen as distinct from the task of inference, or constructing a valid confidence interval. Our approach is primarily focussed on point estimation, rather than inference. In contrast, a key motivation for using the CCT approach, which has become increasingly popular in applied research, has been that it provides theoretically robust coverage even when there is misspecifcation. Recent work has also been explicit in making confidence interval estimation the primary objective \cite<e.g.>{armstrong_kolesar_2018,armstrong_kolesar_2020}. While inference is important, point estimation is also important.
Whilst we are primarily interested in point estimation, we also suggest using the placebo zone to assess coverage of associated confidence intervals, and propose a randomization inference procedure which again draws on estimates in the placebo zone. We also assess coverage for our approach in a range of simulations, and obtain favourable results. If desired, one can always use our approach to obtain point estimates and then use other approaches for inference, such as those promoted in the papers above.
We begin by outlining sufficient conditions under which our selection algorithm is `asymptotically optimal', which is when it converges on selecting the estimator with the lowest mean squared error amongst all candidate estimators as the number of placebo zone repetitions goes to infinity.
We then conduct Monte Carlo simulations using a number of DGPs. Some of these DGPs are stylized, and others are based on well-known RDD applications. Many of these DGPs depart greatly from the conditions we discuss in our proof of optimality. Our approach has lower RMSE than CCT's estimator in almost every simulation, and always lower RMSE than both CCT and IK estimators with DGPs derived from real data.
We then demonstrate our approach with a novel evaluation of a policy designed to reduce motor vehicle accidents (MVAs) for young drivers. The policy requires that learner drivers meet a minimum supervised driving hours (MSDH) mandate before being able to drive independently; a common requirement in jurisdictions using graduated driver licencing systems.\footnote{Countries implementing graduated driver licensing include Australia, Canada, New Zealand and the U.S.} We are among the first to causally evaluate the effect of MSDH on MVAs.
In New South Wales, Australia, policy rules created two discontinuities whereby young drivers needed to complete either 0, 50 or 120 MSDH depending on their birth cohort and date of obtaining license. There are apparent discontinuities in both the level and slope of the first-stage relationship between treatment and date of birth. We could estimate a global polynomial model like \citeA{angrist_lavy_1999} (cohort-IV), RKD, RDD, or RPJKD, and it is a priori unclear which approach we should adopt. More importantly, within each model type, we need to make important functional form and bandwidth choices. In this setting, there are also good reasons to consider models with both asymmetric bandwidths and asymmetric polynomial orders. Institutional details prevent long bandwidths on the left (but not on the right) of the threshold. Institutional details also result in complete compliance on the right, but strong non-linearity on the left, of the threshold. In total we consider almost 10,000 different estimators considering model type, functional form and bandwidth.
Somewhat surprisingly, our `best' estimator is a month-of-birth cohort-IV with linear trend. A mixed order RPJKD also performs well, and indeed performs best when asymmetric bandwidths are allowed. Strikingly, the root mean squared error is about five times greater across the placebo zone if we use the bandwidths suggested by CCT rather than the best performing RDD or RKD bandwidths. In a different application, \shortciteA{card_etal_2017} come to a similar conclusion, drawing on Monte Carlo simulations. When comparing across model types, our `best' estimator is a month-of-birth cohort-IV with linear trend. A mixed order RPJKD also performs well, and indeed performs best when asymmetric bandwidths are allowed.
We find that going from 0 to 50 MSDH lowers the probability of an MVA in the first year of independent driving by 1.4 percentage points (21%). This estimate is highly statistically significant and robust to the randomization inference procedure that we propose in this paper.
We also re-evaluate evidence on the effect of Head Start on child mortality ludwig_miller_2007 and the minimum legal drinking age on drinking behavior lindo_siminski_yerokhin_2016. For both applications, the best performing model is linear RDD with a relatively long bandwidth (much longer than in the original papers) and CCT estimators perform considerably worse than our best models in the placebo zone.
Our Stata command -pzms- implements the placebo zone model selection algorithm, and our proposed approach to randomization inference.\footnote{The program can be installed by typing “ssc install pzms” into the Stata command window.} We recommend that researchers consider using our approach whenever the researcher has access to a sufficiently wide placebo zone. Our companion Stata command -pzms_sim- can be used to easily compare the performance of our approach to CCT in simulations using DGPs tailored to any specific application.\footnote{pzms_sim can be downloaded at: \\ https://drive.google.com/file/d/1rtNn2McXKbrY_lQybpDhHujWN8bk7X_S/view.}
The remainder of the paper is structured as follows. In Section (ref) we outline sufficient conditions under which our method is asymptotically optimal. Section (ref) presents results from Monte Carlo simulations. In Section (ref) we illustrate in detail how to apply the placebo-zone approach in our evaluation of MSDH laws. Section (ref) discusses using the placebo zone for randomization inference. Section (ref) presents results for our MSDH evaluation which adopt the chosen estimators. Section (ref) presents a re-evaluation of Head Start and minimum legal drinking age studies. Section (ref) concludes and discusses practical considerations and recommendations for using the placebo zone approach.
Following standard sharp-RDD notation, consider a random sample ($Y_i(0)$, $Y_i(1)$, $X_i$), $i = 1, 2,\dots , n$, where $Y_i(0)$,$ Y_i(1)$ are potential outcomes, with and without treatment. Treatment ($T$) is determined by the forcing variable exceeding a threshold at $X = 0$ so that $T_i = \textbf{1}(X_i \geq 0)$. The observed sample is therefore ($Y_i$, $X_i$), where $Y_i = (1- T_i )Y_i(0) + T_i Y_i(1)$.
The parameter of interest is the average treatment effect at the threshold $\tau = E[Y_i(1) - Y_i(0)|X_i = 0]$.
There are many candidate estimators of $\tau$. Amongst the set of such candidates, assume the existence of a single estimator which has the lowest MSE($\hat{\tau}$) of all candidate estimators. For each candidate estimator, MSE($\hat{\tau}$) is not observed. Intuitively, our approach will be useful if the set of placebo estimates are informative of MSE($\hat{\tau}$) for each candidate specification.
In Appendix A, we prove that our model selection approach is `asymptotically optimal' under a set of sufficient conditions. By `asymptotically optimal', we mean that it converges on selecting the estimator with the lowest MSE amongst all candidate estimators of $\tau$, as the number of placebo zone repetitions, $m$, gets large, keeping constant the distance between consecutive placebo thresholds. Perhaps the most important of these conditions is that the conditional expectations of $Y_i(0)$ and $Y_i(1)$ are continuous and have zero fourth derivatives with respect to $X$. In other words, that the global DGP is cubic, or a lower order polynomial. The proof also requires homoskedasticy and uniformly distributed $X$.
While these sufficient conditions are restrictive, they are not intended to be realistic. Rather, it serves as a baseline case. More importantly, as we show with simulations in Section (ref), the approach continues to perform favourably under major violations of these assumptions, and with a finite number of placebo estimates.
In this section we present the results of Monte Carlo simulations, in which we illustrate the performance of our approach under various conditions, and compare this to popular bandwidth selection algorithms. While our approach can be used to choose between candidate estimators that vary on a range of dimensions, its main use is likely to be as an aid for choosing bandwidth and polynomial order, for RDD estimators. We therefore focus on these designs and choices in this section. First we consider a range of stylised DGPs, and then turn to realistic DGPs, based on well-known applications.
We commence with some simple DGPs. In each case the sample size is 900 observations. For observation $i$, the running variable $x$ is equal to $i - 100.5$. The running variable is therefore uniformly distributed across the range (-100, 800).\footnote{This structure is common for RDD applications with sample sizes of this order. In particular, this structure arises when a larger original data set has been collapsed into a smaller data set in which each observation represents the mean of $y$ within a particular range of $x$.} The outcome variable is given by
where 0.3 is the discontinuity at $x = 0$, and $\sigma^2 = 0.1^2$ (representing `large' error variance), or $0.03^2$ (`small' error variance).\footnote{The size of the discontinuity has no bearing on the results. 0.3 was chosen for presentational purposes.} $f(x)$ is either:
Linear: $f(x) = x/400$
Quadratic: $f(x) = (x/400)^2$
Cubic: $f(x) = (x/400)^3$
Sine: $f(x) = \sin(2\pi x/400)/2$
Cosine: $f(x) = \cos(2\pi x/400)/2$
Linear, quadratic and cubic DGPs were chosen because these are theoretically ideal conditions for our approach. The sine and cosine functions were chosen to represent DGPs that do not satisfy assumption (3) in Section (ref) (zero fourth order derivative) and are dissimilar near the treatment threshold to the placebo zone. More specifically, the 3rd derivative of the sine curve takes its maximum value at the threshold. In contrast, the 3rd derivative of the cosine curve takes its minimum value at the threshold.
The resulting ten DGPs are depicted in Figure C1. Each panel shows the deterministic component of $f(x)$ as well as a scatter plot with one entire simulated data set (the data generated in the first iteration of the simulation). The vertical lines at $x=400$ show where the placebo zone is truncated when we test the performance of each approach with a smaller placebo zone. Given that our approach relies on asymptotics for the number of placebo replications, we expect it to perform better when we use a longer placebo zone.
The results of the simulations are shown in Panels A-E of Table (ref). For each DGP, we show the RMSE of the estimated discontinuity across 1,000 iterations for our approach (labelled KS), as well as CCT and IK with polynomial order 1.\footnote{We focus on polynomial order 1 since this is the typical specification used with these methods. Moreover, we found that across all our specifications, CCT with order 1 always outperformed CCT with order 2 (results available on request). We also ran versions of CCT and IK without the regularisation term. The RMSEs from these versions were similar but usually slightly higher than those shown for the DGPs in Panels A-E. The exception is for the cosine function, where the IK RMSE was much higher without the regularisation term.} We also show the average `optimal' bandwidth across iterations, as selected by each approach. For our approach, we also show the share of iterations in which the `optimal' estimator is linear (as opposed to quadratic). For our approach, we set the maximum bandwidth to 300 for the `long' placebo zone trials, and 200 for the `short' placebo zone trials. In each candidate model, the bandwidths are set to be symmetrical, until reaching 100 units (which is equal to the full support on the left side of the discontinuity). At higher (right side) bandwidths, the left bandwidth is fixed at 100 units. We discuss the choice of maximum bandwidth further in the conclusion.
A main feature of Table (ref) is that KS attains a lower RMSE than CCT in almost every version of the simulation. The CCT bandwidth (and its performance) are relatively insensitive to most of the parameters of the simulations: $f(x)$, the error variance, and the size of the placebo zone.
The relative performance of KS is best in the linear, quadratic and cubic DGPs, especially with large SEs and large bandwidth. Here, the RMSE is always considerably lower for KS than CCT, and in all but one case lower than IK. In all of these simulations, our algorithm selects linear functions in a large majority (over 91%) of the iterations.
For the sine DGP, the results are more mixed. The IK performs best in 3 of the 4 versions. KS also performs quite well, and is actually wins when the placebo zone is short and error variance is low. For the Cosine function, IK performs best and CCT worst. KS performs reasonably well, especially when the error variance is large.
We conclude from this that KS outperforms the alternate approaches when the DGP is linear, quadratic or cubic, even when the placebo zone is not overly long. Even with very unstable DGPs, and relatively small placebo zones, our approach still performs reasonably well.
We now turn to DGPs which are designed to mimic realistic scenarios, drawing on three well known applications -- Head Start ludwig_miller_2007, political incumbency lee_2008, and Minimum Legal Drinking Age (MLDA) lindo_siminski_yerokhin_2016.\footnote{The MLDA context is one of the best known applications of RDD, beginning with \citeA{carpenter_dobkin_2009}. \citeA{carpenter_dobkin_2009} used restricted variables from the NHIS, which are not easily available. Instead, we draw on data from \shortciteA{lindo_siminski_yerokhin_2016}'s corresponding analysis for the Australian state of New South Wales.} In each case, we take the following approach. Using the original data from each application, we fit $f(x)$: a 5th order global polynomial through the support of the data, allowing a discontinuity and a kink at the threshold, which allows the treatment effect to be heterogeneous in a way that satisfies assumption (4) in Section (ref) (i.e., the average treatment effect has a zero second derivative with respect to $X$).\footnote{Other papers which have undertaken similar exercises have taken the approach of fitting 5th order polynomials on either side of the threshold. We do not believe this is appropriate for the present exercise. Fitting 5th order polynomials on either side results in discontinuities in every derivative of $y$ w.r.t. $x$. It can also result in a wildly unstable fit on either side of the threshold -- but a much smoother fit elsewhere \cite<see for example>[Figure 1, Model 2]{calonico_cattaneo_titiunik_2014}. This is only realistic under a violation of the usual assumption of “smoothness” of all other determinants of $y$, or when the ATE is a highly non-linear function of $X$. This is not appropriate for testing our approach, which relies on the assumption that data patterns away from the threshold can sometimes be informative about likely patterns near the threshold. Nevertheless, our approach still outperforms the other candidate approaches in a majority of the cases discussed here, even when a 5th order polynomial is fitted on either side (results available on request).} \footnote{ For the political incumbency application, we follow the precedent of IK by first dropping observations where the running variable $< -.99$ or $>.99$.} We then fit a beta distribution to the same data to summarize the distribution of the running variable.
For each iteration of the simulation, the sample size is set equal to the original sample. We randomly draw values of the running variable from the beta distribution. Finally, we set $y = f(x) + \epsilon$, where $\epsilon$ is normally distributed with zero mean and variance equal to the variance of the residuals from the regression in the first step.
The resulting DGPs are depicted in Figure C2. Each panel shows $f(x)$ and a scatter plot with the full data set generated in the first iteration of the simulations.
For the Head Start application, in each iteration we allocate each placebo zone observation randomly into one of two groups. This is to address the fact that the density is approximately twice as large in the placebo zone as the treatment zone. We expand on this in Section (ref).
The results of these simulations are shown in Panels F-H of Table (ref). The key result is that KS outperforms the other approaches for each application, while CCT is consistently last.\footnote{We also ran versions of CCT and IK without the regularisation term (available on request). The RMSEs from these versions were slightly lower than those shown for the DGPs in Panels E-G, but never enough to change the RMSE rankings of the three approaches.} We conclude from this that the KS approach outperforms the other approaches on `realistic' DGPs based on well known applications.
The primary objective of our approach is optimal point estimation. However, it is worth recognizing that an important contribution of CCT is to pioneer `robust confidence intervals' for RDD designs, since conventional standard errors do not guarantee correct coverage rates. We therefore report coverage results from our simulations in Table C1, where confidence intervals for the KS method are constructed using conventional asymptotic standard errors clustered at units of the running variable. We also report results for our method using a novel randomization inference approach based on the placebo estimates \cite<building on>{ganong_jager_2018}, which we outline in Section (ref).
The results can be summarized as follows. When our approach is combined with the conventional inference procedure, coverage is close to 95% for each of the 'realistic' DGPs, with smaller average confidence intervals than other methods, especially CCT. Similarly, for each of the stylized DGPs except sine, our approach with conventional inference also achieves coverage close to 95%, often closer to 95% than CCT. Our confidence intervals are also markedly shorter than CCT (sometimes less than half the length). The only DGP where conventional inference does not perform well with our approach is the sine DGP. The randomization inference procedure also performs well, with similar coverage to CCT for most DGPs, and better coverage in the majority of the DGPs considered. Overall, our approach generally performs better or at least similarly on coverage using either conventional inference or randomization inference than CCT, while also having much shorter confidence intervals.
We now demonstrate our method in detail, in the context of a novel application -- estimating the effectiveness of learners' permit policy changes in New South Wales.\footnote{Our study received ethics approval from the UTS Human Research Ethics Committee (Application number ETH17-1547).} Online Appendix B describes institutional details, including the two policy changes and key features of the data. The policy reforms provide exogenous variation in the probability of being subject to MSDH, as a function of date of birth (DOB).What makes this application so interesting is there are many classes of estimators (e.g., RDD, RKD, cohort IV) that can be used to estimate treatment effects in this setting, as we will show. We use our method to compare the performance of estimators within each class, as well as between classes.
We begin by describing the many credible candidate models which could be applied to estimate the effect of the policy changes. We then describe the `placebo zone'. This is a set of 2,556 consecutive DOBs (from 1 July 1984 to 30 June 1991). Within this zone, there is no reason to suspect any systematic relationship between DOB and the outcome variable (MVAs within 1 year of obtaining a provisional license). It therefore provides an opportunity for testing the performance of candidate models in estimating the true treatment effect within this zone (which is zero). Next, we summarize the performance of the candidate models within this zone.\footnote{In the working paper version of this paper, we also consider the implications of various types of treatment effect heterogeneity which we impose into the placebo zone data kettlewell_siminski_2020.}
The first-stage relationship between DOB and holding a `new' learner's permit following the 2000 reform (0 to 50 MSDH) is shown in Figure (ref) (see Appendix Figure C3 for the 2007 reform, and Figure C4 for scatter plots showing the reduced form for both reforms). Both panels draw on the same underlying data, differing only in the bin-size used in the plots. Panel A uses a `small' bin-size of 2 days, while Panel B uses a `large' bin size of 30 days.
Figure (ref) shows complete compliance to the right of 1 July 1984.\footnote{While there appears to be very minor non-compliance, this is due to imprecision around DOB, as described in Section B.4. We drop observations where we are uncertain about treatment status in our regression analysis.} It was not possible for anyone born on this date or after to hold an `old' learner's permit due to administrative rules. The pattern on the left side is more complicated. Both panels show a monotonic upward, non-linear pattern. Panel B suggests the presence of a discontinuity at the threshold. In contrast, Panel A suggests no discontinuity, but a kink, caused by a very steep rise on the left side of the threshold.
This figure illustrates that many different estimators could potentially be used to estimate the effect of the reform. Candidate estimators could exploit the apparent kink, or the approximate discontinuity, or both. Or, they could instead employ a between-cohort-IV strategy. Each approach could be implemented using various alternate functional form assumptions (i.e. orders of polynomial, which need not be the same on each side of the threshold). Finally, one can choose between many bandwidths.
We first consider a total of 4,634 alternate candidate specifications, each with symmetrical bandwidth around the threshold. This consists of 14 different models, estimated using each possible bandwidth in the range of 35 to 365 days. In principle, we could consider larger bandwidths as well. This is prevented by practical considerations in our application. People born before 1 July 1983 were eligible for driver's licenses which differed in other important ways. Therefore we need an estimator which does not use data on people born before that date, hence making 365 days the largest feasible bandwidth.
Denoting outcome (i.e. MVA 1-year indicator) for person $i$ by $Y_i$, DOB by $X_i$ (centred at zero around 1 July 1984), treatment (obtained learner's permit after policy change) by $T_i$ and an indicator for DOB $\geq$ 1 July 1984 (1991) by $D_i$, the first 11 candidate models are fuzzy RDD, RPJKD and RKD estimators. Each of these can be treated as instrumental variable models, with the structural equation given by Eq. (ref) and first-stage given by Eq. (ref). Full details on the estimation equations are in Table C2.\footnote{We only consider a uniform kernel in our application, although it would be straightforward to vary the kernel along with other modelling dimensions. In practice, kernel choice typically has little influence on the estimates lee_lemieux_2010.}
Model 1 is a conventional (fully-interacted) linear RDD. Model 2 is an RDD model with a linear fit on the right side of the threshold, and a quadratic on the left. This is motivated by the first-stage relationship in Figure (ref), characterized by a clearly nonlinear relationship on the left, and perfect linearity on the right. We refer to this as a `mixed polynomial' specification. Model 3 is a conventional (fully-interacted) quadratic RDD.
The next four candidate models exploit both the discontinuity and the kink for identification. These are RPJKD estimators of the following form. Model 4 is a conventional (fully-interacted) linear RPJKD. Model 5 is a quadratic RPJKD, in which the quadratic term is not interacted with the threshold indicator. Model 6 is an RPJKD model with a linear fit on the right side of the threshold, and a quadratic on the left. Model 7 is a fully-interacted quadratic RPJKD.
Four more candidate models adopt conventional regression kink designs. Model 8 is a conventional (fully-interacted) linear RKD. Model 9 is a quadratic RKD, in which the quadratic term is not interacted with the threshold indicator. Model 10 is an RKD model with a linear fit on the right side of the threshold, and a quadratic on the left. Model 11 is a fully-interacted quadratic RKD.
The remaining three candidate models are month-of-birth cohort-IV models, which exploit between-cohort variation in the probability of `treatment'. Denoting month-of-birth fixed effects by $\theta_m$, for these models, the first-stage becomes:
Model 12 assumes a linear secular relationship between DOB and the outcome variable. Model 13 assumes a quadratic secular relationship between DOB and the outcome variable. Model 14 assumes a cubic secular relationship between DOB and the outcome variable.
The placebo zone is the set of DOBs between 1 July 1984 and 30 June 1991, inclusive. There were no apparent major licencing policy changes which were likely to have affected MVAs in a way that depends on DOB within this zone. Figure (ref) Panel A shows the MVA rate by month of birth within this zone (in 30 day bins), with a lowess fit. Generally, the pattern is relatively smooth, with a slight downward trend, apart from perhaps the first 5 months.
Within this zone, we create placebo treatments in a way that mimics the true treatment selection process. For example, in the first placebo, persons are deemed treated if they obtained their license on or after 1 July 2001. The first-stage relationship between DOB and this placebo treatment is shown (in 2-day bins) in Figure (ref) Panel B, with a 365 day bandwidth around the DOB threshold of 1 July 1985. This relationship closely resembles the true treatment profile around the 1 July 1984 DOB, which we show in Figure (ref). Similar patterns are found for the other placebo DOB thresholds in this zone.
After collapsing to DOB-level (and weighting by cell-size), we estimate the placebo treatment effect (which we know to be zero and constant across entities) using each of the 4,634 candidate models.\footnote{The results are almost identical when uncollapsed microdata are used instead, but estimation is much faster with collapsed data.} We repeat this for all 1,826 placebo treatment thresholds, and summarize the performance of each candidate model.
Table (ref) summarizes the performance of each candidate model. It would not be practical to report on the performance of all 4,634 candidates. Instead we show only the results for the bandwidth which yields the lowest root mean squared error (RMSE) for each model type. The first clear feature of this table is that for every model considered, large bandwidths (365 days in all but one case) yield the smallest RMSEs, compared with smaller bandwidths. Secondly, most models have appropriate coverage rates.
In our application, four models stand out with the lowest RMSE. The best performing model (RMSE = 0.0051) is `Model 12' -- the month-of-birth cohort-IV model with a linear trend. This is closely followed by `Model 6' (RMSE = 0.0052) -- the RPJKD with mixed polynomial fit (quadratic on the left and linear on the right). Next are the RKD with mixed-polynomials (RMSE = 0.0057) and the linear RPJKD (RMSE = 0.0060). All four have similarly good coverage (at least 93.6%), and small average bias (0.001 at most).
We also consider a model-averaging approach. It is defined as the weighted average of the estimates from the 14 candidate models (each with full 365 day bandwidth). The weights are set to the inverse of the MSE of each candidate model. The performance of this weighted average estimator is also shown in Table (ref). Whilst its performance is good, its RMSE is higher than Models 12 and 6. The coverage of this estimator is not shown as its variance has not been derived.
The final four rows of Table (ref) summarize the performance of four estimators proposed by CCT, and implemented using Stata’s -rdrobust- command. These are conventional, and bias-corrected estimates using RDD and RKD, respectively.\footnote{The results shown are for models estimated on collapsed microdata. As with the other estimators considered, the results with collapsed (DOB) data are very similar.} The `optimal' bandwidths for these estimators are determined within rdrobust, rather than the placebo zone procedure that we adopt for the other estimators.\footnote{More precisely, the CCT bandwidths shown in Table (ref) are the average of bandwidths selected by rdrobust through the placebo zone.} As seen in the table, these bandwidths are much smaller than the others. The key result, however, is that the performance of these estimators, as measured by RMSE, is worse than any of the other candidate models, and an order of magnitude worse than the best performing candidate models. This is consistent with the findings of Card et al. (2017)'s RKD Monte Carlo simulations and our results in Section (ref).
In every model tested on the placebo zone thus far, we have followed conventional practice and imposed the same bandwidth on the left and right sides of the threshold. Here we explore whether model performance can be improved by allowing for asymmetric bandwidths.
In particular, we have so far capped the bandwidth at 365 days on each of the thresholds. This is motivated by practical constraints in our application. Any more than 365 days to the left of the 1 July 1984 threshold would take us into territory where other important policy changes were implemented in a way that relates systematically with DOB. But we do not have the same issue on the right side of the threshold. Similarly, for the 1991 threshold, we have no constraints in the left side, though data constraints prevent us from considering bandwidths greater than 365 days on the right.\footnote{The constraint is due to the fact that drivers who obtain their P1 license after age 25 are dropped from the sample because they are not required to meet the MSDH requirement (see Appendix B). We cannot impose this constraint consistently on the RHS of the 2007 reform because the end-date for our license data mean we do not always observe whether people got their P1 license by age 25.}
We now repeat the placebo-zone model selection procedure for all 14 candidate models using two similar procedures.
The results for Version 1 of this exercise are summarized in Table (ref). It shows that performance is improved considerably for every model by allowing larger bandwidths on the right. In some cases, RMSE is reduced by more than 50%. The optimal right-side bandwidth varies considerably, from 550 up to the 730 day limit. Model 6 is the best performing model, with an optimal RHS bandwidth of 550 days. This is the best performing estimator amongst all candidates for estimating the effect of the 2000 reform. The weighted-average estimator, shown in the lowest row, does just as well as Model 6. Models 12, 4 and 10 continue to perform well.
The results for Version 2 of this exercise are summarized in Table (ref). They are similar to those of the previous exercise -- Models 12, 10 and 6 continue to perform well. Optimal bandwidths vary, but are generally considerably larger than the baseline exercise. Model 12 has the lowest RMSE, with an optimal LHS bandwidth of 560 days. This is the single best performing specification amongst all candidates for estimating the effect of the 2007 reform.
Our main interest in this paper is estimation. However the placebo zone may also be useful for inference. As already shown, the placebo zone can be used to assess coverage of confidence intervals stemming from standard approaches. In this section, we propose some new approaches to randomization inference.
In our main application, the placebo zone consists of 1,826 overlapping data windows, corresponding to 1,826 separate placebo estimates for each (symmetric) estimator. One can use the distribution of these estimates for alternative approaches to inference -- randomization inference, in the spirit of \citeA{ganong_jager_2018}. We discuss two alternative inference approaches.
Randomization inference considers the estimate of interest alongside the distribution of placebo estimates. Consider an estimate which lies outside of the range of the placebo estimates. If these 1,826 placebo estimates were independent, as per \citeA{ganong_jager_2018}, one would conclude that the two-sided p-value $< 2/1826 = 0.0011$. For an estimate lying inside the range of placebo estimates, $p = 2*\min(i/1826,(1826-i)/1826)$, where $i$ is the rank of the estimate alongside the 1,826 placebo estimates.
However, in our application, the 1,826 placebo estimates are not independent. Indeed they are strongly serially correlated. This is almost certainly the case in other applications as well, given the rolling data window used for consecutive placebo estimates. In our Model 12, the serial correlation of placebo estimates = 0.9895. This equates to an effective sample size (ESS) of just 10 independent observations.\footnote{This calculation draws on Eq. 5 in \citeA{zwiers_storch_1995}.} A more appropriate two-sided p-value for estimates lying outside the placebo zone is $p < 2/ESS$. In our case, for Model 12, $p < 2/10 = 0.2$.
A much more powerful approach to randomization inference is available if one invokes an assumption that the placebo estimates are drawn from a normal distribution. Under this approach, the variance of the placebo estimates is informative of the variance of the treatment effect estimate. This approach to randomization inference facilitates not only p-values but also leads to straightforward calculation of standard errors, t-statistics and confidence intervals. In effect, the variance of placebo estimates is assumed to equal the variance of the treatment effect estimator.
As in the first approach above, it is important to account for serial correlation of the placebo estimates. Such serial correlation reduces confidence in the estimated variance of the population distribution. This can be easily accounted for with a degrees of freedom adjustment. The test statistic is hence assumed to follow a t-distribution, with degrees of freedom equal to the ESS of placebo estimates minus 1.
This is our preferred approach to randomization inference. Our simulations suggest that it performs well in comparison to alternate approaches such as CCT and IK. See Section (ref) and Table C1.
In an earlier version of this paper we discuss how one can also draw on the placebo estimates for bias correction, both for estimation and inference kettlewell_siminski_2020.
In our own application, the placebo zone is contiguous. But in other applications it may not be, particularly if placebo data are available on `both sides' of the real threshold. In such cases, it is not obvious how to determine the `effective sample size', since the estimates on either side of the `gap' may be correlated, but less so than estimates from immediately adjacent thresholds. One approach is to bound the effective sample size. The lower bound essentially ignores this discontinuity in serial correlation, and derives the ESS using a weighted average of the autocorrelations within each contiguous segment. The upper bound treats estimates between each segment as distinct and so the total ESS is the sum of the ESS in each segment.
In our application, this approach would produce tight bounds. To illustrate, if our placebo zone was not contiguous but instead consisted of two equally sized zones on either side of the treatment threshold, the lower bound for the ESS for our best symmetric bandwidth estimator would be 10 and the upper bound would be 11.
The main estimation results are presented in Table (ref). Panel A shows the estimated effects of the 2000 reform and Panel B shows the estimated effects of the 2007 reform. Each panel shows results from five separate estimators -- one in each column.
Column (1) shows results from the `best' estimator. This is the estimator with the lowest RMSE of all candidate models evaluated on the placebo zone. For the 2000 reform, this is `Model 6', with a bandwidth of 365 days on the left and 550 on the right. For the 2007 reform, this is `Model 12' with a bandwidth of 560 days on the left and 365 days on the right.\footnote{Recall that we are constrained to a maximum bandwidth of 365 days on the left of the threshold for the 2000 reform, and a maximum bandwidth of 365 days on the right of the threshold for the 2007 reform. Model 6 is a RPJKD model with quadratic polynomials to the left and linear polynomial to the right of the threshold in each stage. Model 12 is a month-of-birth cohort-IV model, controlling for a linear secular relationship between DOB and the outcome variable.}
These `best' estimates suggest that the first reform had a strong impact on reducing MVAs, while the second reform did not. The first reform is estimated to have reduced the crash rate by -0.014, a reduction of 21% relative to the predicted value at the threshold for the untreated. The conventional p-value associated with this estimate is 0.0005. We have no reason to be sceptical about the validity of this p-value, since this estimator was found to have good coverage in the placebo zone trial, as well as an estimated bias that is close to zero. Nevertheless, we also show alternate p-values, based on the distribution of placebo estimates, as discussed in the previous section. This p-value is larger (0.0096), though still strongly significant. The alternate p-value is larger, even though the alternate standard error is smaller, due to the small number of degrees of freedom in this approach to inference.\footnote{Just 6 degrees of freedom are used for this estimate. This is equal to the `effective sample size' of placebo estimates calculated in the placebo zone minus 1, taking into account the very strong serial correlation of those estimates (0.9915).} For the 2007 reform, the alternate p-value is also slightly higher than the conventional p-value, but remains very far from any conventional threshold of statistical significance.
Column (2) shows results from the best symmetric estimator -- which is Model 12 with a bandwidth of 365 days on each side. For the 2000 reform, all of the key parameters from this model are similar to those in Column 1. The standard error and both p-values are all slightly larger, but qualitatively the same as in Column (1). The estimate for the 2007 reform is close to zero.
Columns (3), (4) and (5) show the results from the best symmetric RPJKD, RKD and RDD estimators, respectively. In each case, the maximum feasible bandwidth (365 days) is used, consistent with the outcomes of the placebo zone trials. Again, the qualitative conclusions are the same, with strongly significant negative effects of the 2000 reform, and approximately zero for the second reform.
Appendix D delves deeper into the effects of the 2000 reform. We find that: (i) delaying of obtaining a license is at most only a small factor in the treatment effects that we have estimated; (ii) there is no MVA reduction in the following year (12-24 months) after obtaining a license; (iii) the results carry over to only more serious MVAs involving injury; and (iv) our main estimates are similar for males and females.
In Appendix E we undertake a back-of-the envelope cost-benefit analysis using our main estimates. Our estimates imply an average social gain of \$2,300 per person due to the 50 MSDH reform. If we take the conservative view that on average people would complete 20 hours supervised in the absence of the reform, then this would constitute a net social improvement provided that supervisors' and learners' combined cost of obtaining hours is less than \$46 per hour. Since we find no evidence the 120 MSDH reform improved safety, we cannot rule out nil social benefits for that reform.
In this section we re-evaluate evidence for two `well-known' applications suited to our method; Head Start and MLDA. These applications feature relatively large placebo zones, which makes them suitable for our method.
Since \citeA{ludwig_miller_2007}'s RDD analysis of the Head Start program, the data from this study have been used widely for illustrative purposes in the RDD methodological literature, including papers by \shortciteA{calonico_cattaneo_titiunik_2014}, \shortciteA{cattaneo_titiunik_vazquez_2017}, \shortciteA{ganong_jager_2018} and \shortciteA{calonico_et_al_2019}.
The unit of analysis is the county. Treatment is eligibility for technical assistance to develop Head Start funding applications. Eligibility is tied to a sharp, arbitrary cut-off in the county-level poverty rate at 59.198 percentage points. The outcome variable is child mortality from Head Start-relevant causes.
Earlier papers have used a range of methodological approaches to estimate the same discontinuity. \citeA{ludwig_miller_2007} prefer local-linear regressions with a triangular kernel. Citing a lack of consensus on bandwidth selection, they show results using bandwidths of 9, 18 and 36 percentage points, as well as from regular linear and quadratic specifications. \shortciteA{calonico_et_al_2019} use the CCT bandwidth-selection algorithm, which yields bandwidths of 6.81 and 6.98, varying by the use of covariates.
We consider 10 separate estimators, each with a range of alternate bandwidths. These were chosen to examine questions of functional form (linear versus quadratic) kernel (uniform versus triangular), weights (unweighted or population-weighted), and covariates (include or exclude). A priori, weighted estimates are likely to be more precise, since residual variance is likely inversely proportional to population size, and population varies greatly (Mean = 38,964; Standard Deviation = 117,460), ranging from 224 to 2,664,438. We also show four estimates using CCT models. Models 11 and 12 are unweighted conventional and bias-corrected estimates with CCT bandwidths. Models 13 and 14 are corresponding weighted estimates.
We face two challenges for adopting our approach in this context. The first is a relatively small range of the forcing variable within the placebo zone. The placebo zone has a range of 47 percentage points (spanning 15.2 to 59.198 percentage points). When models with relatively large bandwidths are trialled, the effective sample size of the resulting placebo estimates is small. The second challenge is a considerably larger density (about 2.1 times larger) in the placebo zone than in the treatment zone \shortcite<see>[Figure A1]{cattaneo_titiunik_vazquez_2017}. The results of trials within such a high density zone may not be relevant for choosing models to adopt in a low density zone. We address both of these challenges by splitting the placebo zone sample into two independent groups.\footnote{We split the placebo zone observations into two groups because the placebo zone density is 2.1 times greater than the treatment zone density. This approach can be generalized for other contexts where the density is uneven. Practitioners may split the placebo zone into $g$ groups, where $g$ = round(placebo zone density / treatment zone density). It is not clear however if our approach is useful for situations where the treatment zone density is markedly greater than the placebo zone density.} We randomly allocated each county into one of these groups.\footnote{When these random allocations are repeated, the results are generally very similar. The RMSEs for the unweighted specifications are most sensitive to these repetitions, but they seem to always exceed the RMSEs for corresponding weighted specifications, usually by a large factor.} This solves the second challenge, since the resulting density is very similar to that of the treatment zone. It also helps with the first challenge, since the effective sample size of placebo estimates is approximately doubled.
The results of these placebo zone trials are shown in Table (ref). For every model considered, the optimal bandwidths are either the maximum (15 percentage points), or close to it. This is considerably larger than the bandwidths in \shortciteA{calonico_et_al_2019}.\footnote{The bandwidths in \shortciteA{calonico_et_al_2019} are in turn similar to the average CCT-selected bandwidth within the placebo zone, which are shown in the last four rows of Table (ref).} Since we are unable to test larger bandwidths, these should be seen as lower bounds for each optimal bandwidth. The results suggest that for this application, the population weights are very helpful -- reducing the RMSE by around 28% in the linear model. The table also shows that covariates do not help, in fact they increase RMSE slightly. This is perhaps unsurprising, since the set of covariates is not rich and does not account for much residual variation.\footnote{The covariates are: percentage of black and urban population, levels and percentages of population in three age groups (children aged 3 to 5, children aged 14 to 17, and adults older than 25) as well as total population. We do not include `total population' as a covariate whenever we use it as a weight instead.} The results suggest that models with a triangular kernel do worse than a regular rectangular kernel, and that a linear polynomial is preferred to higher orders. The CCT estimators (which are characterized by small bandwidths) perform poorly, but not to the same extent as they do in our main application.
The best performing estimator is the weighted linear RDD, with no controls, and with full bandwidth. The estimated discontinuity using this estimator in the treatment zone is shown in Column (1) of Table (ref). The estimate is statistically significant, consistent with \citeA{ludwig_miller_2007} and with \shortciteA{calonico_et_al_2019}. But the estimate is also considerably smaller than that of \shortciteA{calonico_et_al_2019}. The randomization inference procedure generates a similar standard error, but a larger p-value. This is primarily because the effective sample size of placebo estimates is small (7).
\afterpage{
}
We now illustrate our approach with discontinuities in drinking behaviour at the MLDA. The MLDA context is one of the best known applications of RDD, beginning with \citeA{carpenter_dobkin_2009}. It is featured in econometric textbook treatments of RDD, such as \citeA{angrist_pischke_2015}.
We draw on data from \shortciteA{lindo_siminski_yerokhin_2016}'s analysis for the Australian state of New South Wales. We use the same three self-reported drinking outcomes as Lindo et al.: `Ever drinks', `Drinks regularly' and `Proportion of Days Drinks'. And we use the same data: waves 1-11 of the HILDA survey.
Following \citeA{carpenter_dobkin_2009}, Lindo et al. show estimates from linear specifications with bandwidths up to two years of age. These are centred around the 18th birthday MLDA threshold. Here, we consider the performance of a range of specifications -- linear and quadratic, with and without weights, as well as CCT estimators. We consider a much wider bandwidth range, from three months to five years on the right side (in 90 day increments), with the left side capped at three years. The 3-year cap on the left reflects the limit of data availability in the treatment zone, since all respondents were aged 15 years and over.
Table (ref) shows results from the placebo zone trials, for which the placebo zone consists of 18-30 year old respondents.\footnote{The results are qualitatively similar when a longer placebo zone is used (eg. 18-40 years, or 18-50 years) or if the maximum bandwidth is changed (e.g. 10 years, or 20 years). These are available on request.} In many respects, the results are consistent across outcome variables used, and indeed consistent with the earlier applications we have shown: (i) long bandwidths are optimal for each estimator -- much larger than those selected by CCT’s procedure; (ii) linear RDD yields the lowest RMSEs; (iii) the CCT estimator does poorly, with or without bias adjustment.
Table (ref) shows the estimated discontinuities at the MLDA, using the placebo-zone-optimal models we have identified. These are each linear RDD models, with a bandwidth of three years on the left, and between 4.68 and 4.93 years on the right, as per Table (ref). These results are directly comparable to those in Lindo et al.’s Figure 3. Each of the point estimates is similar to Lindo et al.’s 2-year bandwidth estimates.
Regression Discontinuity Design and related estimators are amongst the most important tools of empirical economics. When using such estimators, however, applied researchers are typically faced with choosing between hundreds or thousands of candidate specifications. The large number of candidates is due to the numerous dimensions by which these estimators can vary -- bandwidth, functional form, kernel, covariates are some of these dimensions, and these need not be the same on either side of the threshold. Various guidelines have been developed for model selection, but these generally only address one of these dimensions, whilst keeping others constant. In practice, contemporary applied work in leading economics journals still relies more on robustness testing than on model selection algorithms. Many such papers provide no explicit justification for model specification.
We have outlined a new approach for model selection which allows the performance of all candidate models to be assessed. The approach is conceptually straightforward. Each candidate model is assessed on its performance in estimating treatment effects in a placebo zone of the running variable – where the true effect is known to be zero. The RMSE of the resulting placebo estimates is the summary statistic by which each estimator is judged.
We have proven that this new approach is asymptotically optimal under restrictive conditions. More importantly, our simulations show that the approach also performs well under very different conditions, including with DGPs based on well-known applications. Our approach has potential to be useful for model-selection in a wide range of applications. We have demonstrated its use with three such applications within the paper. Researchers can implement the approach using our Stata command -pzms-. We recommend that researchers consider our approach whenever the available range of the running variable is relatively wide. One can compare the likely performance of our approach to other approaches by conducting a simulation exercise based on the data from any given application. This is quick and easy to do using our companion Stata command -pzms_sim-, which implements simulations following the steps we discuss in Section (ref). However the approach should not be seen as a completely automated procedure for unproblematically choosing an objectively best specification. In this section we discuss some complications and suggestions for using the approach judiciously.
Our method is relatively straightforward to apply for testing a set of candidate sharp-RDD models. For designs that rely on a first stage (e.g., fuzzy-RDD and RKD), a first stage relationship between the treatment variable and running variable may not exist in the placebo zone.\footnote{Our own application is unusual since the placebo treatments have a natural definition -- as a function of date obtained learners license.} In these settings, one option is to use our approach to choose an estimator for the reduced-form discontinuity (or kink) only -- that is, the discontinuity (or kink) of the outcome variable at the threshold. Such an approach may be useful in cases where it can be reasonably assumed that the first-stage discontinuity (or kink) estimate is relatively insensitive to the specification used. This is often the case in practice -- see for example \citeA{Abdulkadiroglu_etal_2014}.\footnote{This is often the case in practice, sometimes because the first-stage results from policy rules that function in a nearly deterministic way. Of the fuzzy-RDD papers we identified in Table (ref), the majority arguably fall in this category in our view.}
In any given application of our proposed method, the analyst must choose a maximum bandwidth for the set of candidate models. This choice will depend on the specific constraints of the application. In principle, one would like to consider all possible bandwidths, but this is not practical. If the chosen maximum bandwidth is too large, the number of thresholds within the placebo zone will be too small for the procedure to be informative about model performance.\footnote{Larger bandwidths are also likely to yield higher serial correlation in placebo estimates.}
When natural constraints do not occur, we suggest that researchers undertake some data examination before choosing a maximum bandwidth. A key consideration, for a given model type, is the relationship between treatment effect estimates and bandwidth. In Figure C7 we plot this for our applications in Section (ref). For Head Start, the estimated treatment effect changes markedly up to a bandwidth of around 15, which indicates worth in setting a maximum at least equal to this value. Beyond 15, the effect size is more stable and (arguably) the discrepancy between estimates with bandwidths between 15 and 30 are not economically important. Since our estimator selects 15 as the preferred bandwidth when we set this as the maximum, there is little value in choosing a higher bandwidth and this seems like a sensible choice (if the solution had been interior, we would suggest increasing the maximum and reassessing). For MLDA outcomes, the estimates are fairly stable after five years.\footnote{We also suggest researchers consider the ESS. The ESS of the placebo estimates for the Head Start applications is quite small (5), due to the relatively small placebo zone. A larger maximum bandwidth would decrease the ESS even further. Importantly, however, even with this small ESS, our monte carlo simulations suggest that our procedure still outperforms popular alternative approaches (Table (ref), Panel E). The small ESS perhaps has greater implications for our alternate inference procedure, by which the estimate is only marginally significant.}
Our approach is perhaps most useful for model selection within (rather than between) a class of estimators. For example, consider the large set of candidate RDD estimators for a given application. Our approach assesses performance of such models with different bandwidths and different polynomial orders. Each of those candidate estimators has the same target parameter, and so comparing performance is relatively unproblematic.
Comparing performance between classes of models is more problematic, because they often estimate different parameters. Fuzzy-RDD models estimate LATEs, while RKDs estimate MTEs, RPJKDs estimate a weighted average of a LATE and a MTE (under additional assumptions of local MTE stability), while cohort-IV estimates a weighted average of a different set of LATEs. Our approach can be used to compare performance between such models. But this can only be done unproblematically if one is willing to assume that selection into treatment is unrelated to potential gains from that treatment. In our own application, this may be a reasonable assumption. It is less reasonable in many other applications.
More generally, researchers adopting our approach should carefully consider the implications of potential treatment effect heterogeneity. To be clear, placebo treatment effects in the raw data are precisely zero. This implies that model performance is assessed in a constant-treatment-effect context. This may be informative for model selection in more general contexts. But a more nuanced approach is to explore the implications for model performance if treatment effect heterogeneity is imposed into the placebo zone. In \citeA{kettlewell_siminski_2020} (Section 4.6) we demonstrate how one might go about this using our own MSDH application as an example.