EconBase
← Back to paper

Visual Inference and Graphical Representation in Regression Discontinuity Designs

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

103,511 characters · 16 sections · 109 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

\pagestyle{plain}

\thispagestyle{empty}

\setcounter{footnote}{0}

titlepage\begin{center} \\ {VISUAL INFERENCE AND GRAPHICAL REPRESENTATION IN }\\ {REGRESSION DISCONTINUITY DESIGNS\footnote[1]{We are grateful for the insightful and constructive comments from two co-editors, Larry Katz and Andrei Shleifer, and four anonymous reviewers. We have also benefited from discussions with Alberto Abadie, Sahara Byrne, Colin Camerer, Matias Cattaneo, Damon Clark, Geoff Fisher, Paul Goldsmith-Pinkham, Nathan Grawe, Jessica Hullman, David Lee, Lars Lefgren, Thomas Lemieux, Pauline Leung, Jia Li, Adam Loy, Alex Mas, Doug Miller, Ted O'Donoghue, Bitsy Perlman, Steve Pischke, Jonathan Roth, Jesse Rothstein, Rocio Titiunik, Cindy Xiong, Stephanie Wang, Andrea Weber, and Xiaoyang Ye, as well as participants of various seminars and conferences. Lexin Cai, Matt Comey, Michael Daly, Rebecca Jackson, Motasem Kalaji, Xingyue Li, Fiona Qiu, and Tatiana Velasco provided research assistance, and we thank Brad Turner, Mary Ross, and Patti Tracey for providing logistical support. We are indebted to our friends for testing the experiment and to colleagues for participating in the study. We gratefully acknowledge financial support from the Cornell Institute of Social Sciences and the Princeton Industrial Relations Section. The authors have no relevant or material financial interests that relate to the research of this paper. This study is registered in the Open Science Framework and the AEA RCT Registry with ID AEARCTR-0004331. Any opinions and conclusions expressed herein are those of the authors and do not reflect the views of the U.S.\ Census Bureau.}}\\ \end{center} \begin{singlespace} \begin{center} {Christina Korting, University of Delaware}\\ {Carl Lieberman, U.S. Census Bureau}\\ {Jordan Matsudaira, Columbia University}\\ {Zhuan Pei \setcounter{daggerfootnote}{\value{footnote}} {\fnsymbol{footnote}} \footnote[2]{Corresponding author. Associate Professor, Department of Economics and Jeb E. Brooks School of Public Policy, Cornell University, Ithaca, NY 14853, USA; [email removed]} \setcounter{footnote}{\value{daggerfootnote}} {\arabic{footnote}} , Cornell University and IZA}\\ {Yi Shen, University of Waterloo}\\ \end{center} \end{singlespace} \begin{center} {January 2023} \end{center} \begin{abstract} \begin{singlespace} Despite the widespread use of graphs in empirical research, little is known about readers' ability to process the statistical information they are meant to convey (“visual inference”). We study visual inference within the context of regression discontinuity (RD) designs by measuring how accurately readers identify discontinuities in graphs produced from data generating processes calibrated on 11 published papers from leading economics journals. First, we assess the effects of different graphical representation methods on visual inference using randomized experiments. We find that bin widths and fit lines have the largest impacts on whether participants correctly perceive the presence or absence of a discontinuity. Our experimental results allow us to make evidence-based recommendations to practitioners, and we suggest using small bins with no fit lines as a starting point to construct RD graphs. Second, we compare visual inference on graphs constructed using our preferred method with widely used econometric inference procedures. We find that visual inference achieves similar or lower type I error (false positive) rates and complements econometric inference. Key Words: Graphical Methods; Visual Inference; Regression Discontinuity Design; Expert Prediction; Scientific Communication JEL Code: A11, C10, C40 \end{singlespace} \end{abstract}

\setcounter{footnote}{0}

{3pt}

{3pt}

{3pt}

{3pt}

Introduction

center[center omitted — 135 chars of source]

Graphical analysis is increasingly prevalent in empirical research, a phenomenon currie2020technology call the “graphical revolution.” Effective use of graphs conveys a large set of statistical information at once and improves research transparency (andrews2020transparency). However, there are different ways to construct a graph with the same data, and the particular construction an analyst chooses has the potential to mislead readers (schwartzstein2021using). To understand the best use of graphical evidence, it is important to study readers' ability to process information from graphs---which we term “visual statistical inference” or visual inference per Majumderetal2013---as well as the sensitivity of visual inference to choices in graph construction. To date, little is known about visual inference for commonly presented graphs in empirical research designs.

We begin to fill this knowledge gap and study visual inference in the regression discontinuity design (RDD or RD design). The popularity of RDD in the modern causal inference toolkit, which began in economics with AngristandLavy1999, makes it an important setting in which to study visual inference. Standard practices in applying RDD today perhaps best embody the spirit of Watson's quote above, with graphs playing a central role in the presentation of findings. The key RD graph plots the bivariate relationship between outcome variable $Y$ and running variable $X$ and is meant to display a discontinuity (or lack thereof) in the underlying conditional expectation function (CEF) as $X$ crosses a policy threshold. Influential practitioner guides by ImbensandLemieux2008, LeeandLemieux2010, and cattaneo2019practical recommend creating this graph by dividing $X$ into bins, computing the average of $Y$ within each bin, and generating a scatter plot of these $Y$-averages against the midpoints of the bins.

We assess the performance of visual inference by studying whether people presented with this graph can accurately extract the embedded statistical information, where our main criterion is the correct identification of the existence or absence of a discontinuity at the policy threshold. Our project has two major components. In the first, we build on work pioneered by eells1926relative and refined by ClevelandandMcGill1984b by conducting a series of randomized experiments to examine how different graphical parameters affect visual inference in RD. We present participants recruited through the Cornell University Johnson College's Business Simulation Lab with RD graphs produced from data generating processes (DGPs) based on microdata from 11 published papers that we randomly select from a list of 110 empirical studies from top economics journals. For each graph, we ask participants to identify the existence or absence of a discontinuity. We randomize respondents into different treatment arms and show participants within each arm graphs produced with particular graphical parameters such as small bin widths and evenly spaced bins.

There is limited research on how to choose these parameters in practice. For example, Calonicoetal2015 propose two popular data-driven bin width selectors: one that minimizes the integrated mean squared error (IMSE) of the bin averages, resulting in fewer, larger bins, and another that mimics the variability of the underlying data (mimicking variance or MV), which leads to more, smaller bins. While both proposals set a graphical parameter to satisfy an econometric criterion, practitioners are left with little basis to choose between them. Moreover, a host of other choices over graphical parameters remains with minimal guidance from the literature, such as including smoothed regression lines in the binned scatter plot, adding a vertical line to indicate the policy threshold, and choosing the axis scales.

By comparing the rates at which respondents correctly classify discontinuities across treatment arms, we can assess the advantages and disadvantages of different graphical parameters. We find that certain graphical parameters such as bin width and smoothed regression lines create important tradeoffs between type I errors (identifying a discontinuity when there is none, i.e., a false positive) and type II errors (identifying the absence of a discontinuity when there is one, i.e., a false negative). Relative to MV (small) bins, using IMSE (large) bins tends to increase type I error rates but decrease type II error rates. Similarly, imposing fit lines may also increase type I error rates, echoing the concerns of ct2021nber,cattaneo2021regression.

To translate our findings to recommendations on graphical practices in RDDs, we empirically implement a decision theoretic framework that incorporates classification accuracy as a metric to compare graphical methods. The method that uses MV bins with no fit lines consistently performs well relative to IMSE bins and/or imposing fit lines. Bin spacing (equally spaced versus quantile-spaced), axis scaling, and the presence of a vertical line indicating the policy threshold do not appear to matter, implying that researchers can adhere to reasonable personal preferences. Our recommendation is robust to alternative decision theoretic framework formulations.

Because only non-experts participate in our randomized experiments on the effects of graphical methods, one may be concerned that our results are less relevant to academic audiences. To assess whether our findings generalize across experience levels, we also recruit experts from a pool of seminar attendees and affiliates of the National Bureau of Economic Research (NBER) and the Institute of Labor Economics (IZA) to participate in our study. Although our expert sample is not large enough to conduct the same randomized experiments, we can compare the non-expert and expert results by using the subset of non-experts who saw the same graphs as the experts. We find that the two groups perform comparably.

In the spirit of DellaVignaandPope2018, we also test whether experts are able to predict the graphical parameters that result in the highest rate of visual inference success by non-experts. We find that experts only partly anticipate the aforementioned effects of bin widths and fit lines.

As a second major component of the project, we compare the performances of visual inference and econometric inference. For visual inference, we use results from the sample of experts who viewed graphs constructed with the best-performing technique from our experiments. For econometric inference, we apply three influential methods by ImbensKalyanaraman2012, Calonicoetal2014, and ArmstrongandKolesar2017 (henceforth IK, CCT, and AK, respectively) and conduct hypothesis testing at the 5% (asymptotic) level. We find that visual inference achieves a type I error rate that, at just below 8%, is lower than the IK and CCT procedures (the CCT type I error rate is not significantly higher), but the two econometric procedures enjoy considerably lower type II error rates. Visual inference performs very similarly to the AK procedure, a remarkable result given the minimax optimality property of AK.

We also find that visual and econometric inferences appear to be complementary. First, we examine the joint distribution of visual and econometric tests: while they commit similar type II errors, there does not appear to be a strong association in their type I errors. Second, we assess the performance of a combined visual and econometric inference. One simple way of combining the two inferences mirrors an approach in which a researcher believes a discontinuity exists if and only if a formal test rejects the null hypothesis of no effect and she sees a discontinuity in the RD graph. We find that the combined IK and visual inference performs similarly to the AK procedure, which may help explain the enduring credibility of the RD design despite formal inference issues in earlier RD papers.

Finally, we ask experts to estimate the discontinuity magnitude when they classify a discontinuity, and we compare the accuracy of their estimates to that of econometric methods. On this front, econometric methods tend to do better. For example, the simple local linear IK estimator yields lower mean squared errors than experts across all 11 DGPs, shedding light on the limits of visual inference.

This paper connects a diverse set of literatures and makes the following contributions. First, we begin to fill an important gap in our understanding of graphical evidence by evaluating visual inference and graphical representation practices in a widely used quasi-experimental research design. Our endeavor draws from three strands of the statistics literature that study the choice of graphical parameters (e.g., Calonicoetal2015, li2020essential), their effects on visual inference (e.g., ClevelandandMcGill1984b), and the evaluation of visual inference through comparison with econometric inference (e.g., Majumderetal2013). Our paradigm can be applied to other important areas, as discussed in Section (ref).

Second, to guide our study design and to help interpret our empirical results, we propose a general conceptual framework, which may extend to future studies of visual inference in other contexts. In particular, we can interpret the average type I or II error rate we use as an estimate of the probability that a randomly sampled reader commits such an error when viewing a graph generated from a randomly chosen DGP. We show that these error probabilities are key inputs in a standard decision theoretic framework, which helps inform best graphical practices. We also discuss alternative decision theoretic formulations and make connections to the recent literature on scientific communication (e.g., andrews2020model).

Third, we add to the literature on expert judgments (e.g., CamererandJohnson1997) and expert forecasts of research results (e.g., sanders2015just). Our finding that experts only partly anticipate our experimental results underscores the value of empirically evaluating visual inference and providing evidence-based guidance on graphical methods.

We introduce the conceptual framework in Section (ref), describe the design of our experiments and studies in Section (ref), present results in Section (ref), and conclude in Section (ref). For readers in a hurry, the key takeaway results are in Figures (ref) and (ref) with corresponding discussions in Sections (ref) and (ref).

Conceptual Framework

In this section, we propose a conceptual framework for evaluating visual inference to guide our study design and aid in the interpretation of our empirical results. More specifically, we show how to aggregate visual inference performances across subjects, who may reach different conclusions even when viewing the same graph, and how to meaningfully interpret the parameter to which our aggregate measure corresponds. This performance measure helps to inform best graphical practices.

Although RD graphs may serve other purposes, we view their most important function as accurately conveying discontinuity existence and magnitude at the policy threshold. According to LeeandLemieux2010, other purposes of RD graphs include i) helping to assess regression specifications and ii) allowing for the inspection of discontinuities away from the policy cutoff. But ultimately, these other functions are also motivated by inference on the discontinuity at the policy threshold: i) can be viewed as reconciling visual and econometric inferences thereof and ii) informs the reader, under implicit global homogeneity assumptions, whether to believe the existence of a discontinuity at the policy threshold.

We focus on binary classifications of a graph and treat type I and type II errors as the main performance measures for visual inference. The conceptual framework easily generalizes to assessing visual estimates of the discontinuity magnitude, which we elicit from experts. A person commits a type I error in RD visual inference if she classifies a continuous graph as having a discontinuity and a type II error if she classifies a discontinuous graph as continuous.

To define our aggregate measures of visual inference performance, we introduce the following notations. First, the vector $\gamma$ denotes a combination of graphical parameters (see wilkinson2013grammar for an extensive list). We study five parameters in this paper (bin width, bin spacing, axis scaling, polynomial fit lines, and a vertical line at the policy threshold), and each of the five entries of $\gamma$ represents the value of a particular parameter.\footnote{In our experiments, we randomly assign each participant to view only graphs generated with a certain fixed value of $\gamma$. Ideally, one could run a large experiment with a full factorial design to test all combinations of the graphical values outlined above, but resource constraints force us to test a subset of the graphical parameter space via a sequence of studies as described in Section (ref).}

The combination $(g,d)$ denotes the probability model underlying an RD dataset. $g$ encompasses four elements: i) the distribution of the running variable $X$; ii) the conditional expectation function $E[\tilde{Y}|X=x]$ which is continuous at the policy threshold $x=0$; iii) the distribution of the error term $u$ where $\tilde{Y}\equiv E[\tilde{Y}|X=x]+u$; and iv) the sample size $N$. Intuitively, $g$ specifies everything in the probability model except for the discontinuity, including the shape of the conditional expectation function. The discontinuity then results from shifting the right arm of the smooth function $E[\tilde{Y}|X=x]$ by some discontinuity level $d$, that is, $Y=\tilde{Y}+d\cdot1_{[X\geqslant0]}$ (which implies that $E[Y|X=x]=E[\tilde{Y}|X=x]+d\cdot1_{[X\geqslant0]}$; we provide a graphical illustration in Section (ref)). We note that the variable $Y$ can represent the outcome, baseline covariates, or treatment take-up, so this framework applies to all graphs typically included in RD studies, including those from a fuzzy design. Typically, the $(g,d)$ combination is jointly referred to as the “data generating process,” but we separate the discontinuity level $d$ and call $g$ the DGP for ease of exposition below.

We think of each (bivariate) RD dataset as a realization from the probability model $(g,d)$, which we denote by $W$, or $W(g,d)$ if we want to emphasize the underlying probability model. Implementing a graphical procedure with parameters $\gamma$ on dataset $W$ results in an RD graph $(\gamma,W)$, which we denote by $T$ or $T(\gamma,g,d)$. Alternatively, we can think of $T$ as a realization from $(\gamma,g,d)$ and refer to $(\gamma,g,d)$ as the graph generating process (GGP).

When presented with the same RD graph, readers may draw different visual inferences. For example, some readers may be more skilled than others at classifying a discontinuity because they have received more training in statistics, have more experience with RD graphs, or otherwise have superior ability. We use $\phi$ to capture these human characteristics that affect graph perception.

The probability that a reader with characteristics $\phi$ reports that a discontinuity exists in RD graph $T(\gamma,g,d)$ is denoted by $\tilde{p}(T(\gamma,g,d),\phi)$. From casual observation, we know that the same reader may be influenced by idiosyncratic elements not encapsulated in $\phi$ and classify the same graph differently on different days. The probability formulation $\tilde{p}$ allows these factors to affect visual inference.

We now define the type I and type II error probabilities we use to gauge reader performance. First, averaging $\tilde{p}$ over both data realizations $W$ and reader characteristics $\phi$ leads to the quantity \[ p(\gamma,g,d)\equiv E_{W,\phi}[\tilde{p}(T(\gamma,g,d),\phi)]. \] This is the probability that a randomly chosen reader reports a discontinuity in a graph randomly generated from the GGP $(\gamma,g,d)$. A high value of $p$ indicates a high classification error probability when the true discontinuity $d$ is zero (type I error), but a low classification error probability when $d$ is nonzero (type II error). Formally, the DGP-specific or $g$-specific type I and type II error probabilities for graphical parameters $\gamma$ are defined as:

align*[align* omitted — 173 chars of source]

Conceptually, we can further average $p(\gamma,g,d)$ over the space of DGPs, $\mathcal{G}$ (we discuss $\mathcal{G}$ after Assumption (ref) below) to arrive at the overall discontinuity classification probability for $\gamma$: \[ \bar{p}(\gamma,d)\equiv E_{g\in\mathcal{G}}[p(\gamma,g,d)]. \] Correspondingly, the overall type I and type II error probabilities are defined as

align*[align* omitted — 161 chars of source]

Consistent with the definitions in casella2021statistical, we call $p(\gamma,g,d)$ and $\bar{p}(\gamma,d)$ power functions as functions of $d$.

In this paper, we design experiments to estimate the type I and type II error probabilities as defined above. For each GGP $(\gamma,g,d)$, we generate $M$ different realized graphs and present each to a random participant. That is, participant $i$ is shown one RD graph denoted by $T_i(\gamma,g,d)$, where $i$ takes on values in the set $\{1,...,M\}$, and is asked to assess the presence of a discontinuity. Let the binary variable $R_{i}(T_i(\gamma,g,d))$ denote participant $i$'s discontinuity classification, which equals one if the participant reports a discontinuity at the policy threshold. Under random sampling, the following assumption holds:

assumptionFor a given GGP $(\gamma,g,d)$, the $R_{i}(T_{i}(\gamma,g,d))$'s are i.i.d.\ with $E[R_{i}(T_{i}(\gamma,g,d))]=p(\gamma,g,d).$

A natural estimator for $p(\gamma,g,d)$ is the sample average of discontinuity classifications: \[ \hat{p}(\gamma,g,d)=\frac{1}{M}\sum_{i}R_{i}(T_{i}(\gamma,g,d)). \] Proposition (ref) in Online Appendix (ref) states the distribution of $\hat{p}(\gamma,g,d)$ and shows the estimator to be unbiased and consistent as $M\to\infty$ for $p(\gamma,g,d)$ under Assumption (ref).

To estimate the overall probability $\bar{p}(\gamma,d)$, we need to sample from the DGP space $\mathcal{G}$, which we formally define in Online Appendix (ref). While the infinite dimensionality of $\mathcal{G}$ makes it difficult to characterize the distribution of DGPs, we think of the data used in empirical RD research as realizations when sampling from $\mathcal{G}$ according to this distribution. To that end, we can specify $J$ DGPs that approximate data from existing research and present graphs generated with discontinuity $d$ for each DGP $g_{j}$ ($j=1,...,J$) to a distinct group of $M$ participants for a total of $M\cdot J$ participants and visual discontinuity classifications.

assumptionThe DGP $g_{j}$'s are randomly sampled from $\mathcal{G}$.

A natural estimator for $\bar{p}(\gamma,d)$ is \[ \hat{\bar{p}}(\gamma,d)\equiv\frac{1}{J}\sum_{j}\hat{p}(\gamma,g_{j},d)=\frac{1}{M\cdot J}\sum_{i,j}R_{i}(T_{i}(\gamma,g_{j},d)), \] the average of discontinuity classifications across the $M\cdot J$ classifications. Proposition (ref) in Online Appendix (ref) states the distribution of $\hat{\bar{p}}(\gamma,d)$ and shows the estimator to be unbiased and consistent (as $J\to\infty$) for $\bar{p}(\gamma,d)$ under Assumptions (ref) and (ref) (given that $J=11$ in our experiments, consistency here is a conceptual statement implying that were we to incorporate DGPs from more RD studies, our estimators would be closer in probability to the population parameters of interest). We henceforth refer to $\hat{p}(\gamma,g,0)$ as the DGP-specific or $g$-specific type I error rate, and $\hat{\bar{p}}(\gamma,0)$ as the average type I error rate (or simply the type I error rate). For a particular $d\neq0$, we refer to $1-\hat{p}(\gamma,g,d)$ as the DGP-specific or $g$-specific type II error rate, and $1-\hat{\bar{p}}(\gamma,d)$ as the average type II error rate (or simply the type II error rate) at discontinuity $d$.

We can also define the type I and type II error probabilities of an econometric inference procedure based on a discontinuity estimator, $\hat{\theta}$, but with two adjustments. First, $\gamma$ is no longer an argument in these expressions because we directly implement $\hat{\theta}$ on microdata $W$. Second, we need to specify the level of the testing procedure, which we set to 5%, the prevailing standard in empirical studies. Because the definitions of these probabilities and their estimators are similar to the quantities defined above, we omit them here.

In subsequent sections, we empirically trace out $\hat{\bar{p}}(\gamma,d)$ as functions of $d$, which concisely summarize the type I and type II error probabilities of a graphical method. For brevity, we also use the term “power functions,” as opposed to “estimated power functions,” to refer to their empirical estimates. We study visual inference by comparing its power functions $\hat{\bar{p}}(\gamma,d)$ across $\gamma$ and against the corresponding power functions of various econometric inference procedures. We discuss the calculation of the standard errors on the differences between the visual and econometric power functions when we present our empirical results in Section (ref), as its details depend on the design of our experiments.

The estimated power functions help shed light on best graphical practices. While the “optimal” graphical method depends in part on the optimality criterion we specify, the power functions provide key inputs for some common criteria that are well suited to this context. A simple criterion, as suggested by a referee, is how close the type I error rate is to 5 percent, the conventional threshold in econometric analysis. A second criterion quantifies the costs of type I and type II errors and uses the power function to consider explicitly the tradeoff between the two types of errors. We also consider a third criterion, which is adapted from andrews2020model; instead of using the power function, this criterion relies on readers' confidence in their discontinuity classification, which we can proxy using participants' bonus scheme choice as discussed in Section (ref) and Online Appendix (ref). A full discussion of the second and third criteria require formal decision-theoretic frameworks, and we leave it to Online Appendix (ref).

In summary, we have defined the type I and type II error rates for visual inference. The type I error rate is the fraction of continuous graphs participants incorrectly classify as having a discontinuity, and the type II error rate is the fraction of discontinuous graphs incorrectly classified as being continuous. Our framework allows us to interpret these rates as unbiased and consistent estimates of the probabilities of type I and type II errors a randomly chosen person commits when classifying a graph generated from a representative DGP. These probabilities also help inform best graphical practices.

Description of Experiments and Studies

Graphical Parameters Tested

We test the effects of bin width, bin spacing, parametric fit lines, vertical lines at the policy threshold, and $y$-axis scaling. We discuss each of these treatments in detail below and provide graphical illustrations in Figure (ref).

The most studied graphical parameter in RD is the width of each bin in the binned scatter plot. The first class of bin width selection algorithms comes from LeeandLemieux2010: start with some number of bins, double that number, test whether the additional bins fit the data significantly better, and repeat until the test fails to reject the null hypothesis. Calonicoetal2015 propose two bin width selection algorithms based on different econometric criteria. The first, which is more in line with the convention of the nonparametric regression literature, is the bin selector that minimizes the IMSE of the bin-average estimator of the CEF, where the resulting number of bins increases with the sample size $N$ at the rate $N^{1/3}$. For their second bin selector, the MV selector, Calonicoetal2015 state that they “choose the number of bins so that the binned sample means have an asymptotic (integrated) variability approximately equal to the amount of variability of the raw data.” The resulting number of bins increases with the sample size more quickly, at the rate $N/\log(N)^{2}$. The IMSE-optimal bin width therefore selects fewer bins, and hence has larger bin widths, than the MV algorithm (we describe these algorithms further in Online Appendix (ref)). In addition to these two algorithms, Calonicoetal2015 provide an interpretation for any given number of bins as the output of a weighted IMSE-optimal algorithm. Applying the LeeandLemieux2010 algorithms to our datasets leads to bin numbers that tend to be between the IMSE and MV selectors and closer to those of the IMSE. Thus, we restrict our analysis to the visual inference properties of the IMSE and MV bin selectors.

Although the prevailing approach is to adopt evenly spaced bins, this method has drawbacks in that the resulting bins may contain vastly different numbers of observations, or even none at all.\footnote{In a literature review we conduct for current practices, 98% of the more than 100 studies we compile use evenly spaced bins. Our review includes RDD studies as well as studies that apply the regression kink design---RK design or RKD.} This can happen when the distribution of the running variable is far from uniform. As a remedy, Calonicoetal2015 also propose quantile-spaced bins where each bin contains (approximately) the same number of observations. Both spacings support IMSE and MV bin selectors, and we test each of these combinations.

Following suggestions by ImbensandLemieux2008 and LeeandLemieux2010, RD graphs frequently feature parametric fit lines and a vertical line at the treatment threshold. Both papers suggest fit lines improve “visual clarity” by approximating the conditional expectation functions, and the default in the popular rdplot command by Calonicoetal2015 uses piecewise global quartic regressions on each side of the policy threshold. Of the 11 RDD papers on which we calibrate our DGPs, 10 include fit lines. For the six of these papers that generate fit lines using polynomial regressions, we use the same polynomial order as in the source graph, and in the remaining cases, our team unanimously decides on the fit that best matches the original fit line or the data. We could also use a formal data-driven approach to select the polynomial order, but different criteria from LeeandLemieux2010 lead to conflicting recommendations (for example, for our first DGP, the Akaike information criterion selects a fifth-order polynomial, while the Bayesian information criteria and $F$-test they describe choose a zeroth-order polynomial).

The presence of a vertical line is designed to visually separate observations above and below the cutoff. We test all combinations of the fit line and vertical line treatments except for including the fit lines and excluding the vertical line, which is used infrequently in practice (our literature review shows that fewer than 10% of papers use this combination).

The motivation for rescaling graph axes comes from Clevelandetal1982, who note that correlations on scatter plots seem stronger when scales are increased. We use two axis scaling options in our experiments. First is the default output returned by Stata 14. Second, we double that range by recording the range of the $y$-variable from the default graph and increasing the bounds by 50% of the original range in each direction, resulting in a graph where the data are condensed along the vertical axis. We do not manipulate the scale of the $x$-axis because our survey of the literature suggests that this is not a common adjustment. A related decision researchers encounter is the range of the running variable to use in producing graphs: should they use the entire dataset or only a subsample close to the policy threshold? We do not test this margin of adjustment in our experiments due to the difficulty in generalizing the findings from such an exercise. Suppose we find that selecting 50% of the observations closest to the threshold improves visual inference, should researchers “chop” the sample they are already planning to use? And after doing so, will it be beneficial to chop again? One could argue for testing the effect of using the full sample versus the subsample falling within the IK or CCT bandwidth, but these bandwidths themselves depend on the full sample---the first step in bandwidth calculation is a (semi-)global regression---and it is not even clear how we should define the “full” sample: some of the replication data used for our DGP calibration are already subsets of a larger sample.

There are other graphical parameters we do not test in our experiments. One is plotting confidence bands around the binned averages or fit lines. However, confidence intervals are too complex to explain to the non-experts in our short tutorial without potentially affecting the way participants think about the classification task itself, and therefore we do not experimentally test their effects on visual inference. Because the moderate size of our expert pool makes it unsuited to randomized experiments, we refrain from testing other graphical parameters.

Creation of Simulated Datasets and Graphs

We specify data generating processes based on the actual data used in published research. We randomly sample 11 from a total of 110 empirical RD papers published in the American Economic Review, American Economic Journals, Econometrica, Journal of Business and Economic Statistics, Journal of Political Economy, Quarterly Journal of Economics, Review of Economic Studies, and Review of Economics and Statistics between 1999 and 2017 that have replication data available to create our DGPs. We refer to these DGPs as DGP1-DGP11.

The calibration of each DGP $g$ entails the specification of its four components: the distribution of the running variable $X$, the continuous conditional expectation function $E[\tilde{Y}|X=x]$, the distribution of the error term $u$, and the sample size $N$.\footnote{Out of the 11 studies, two plot residualized outcomes instead of the original $Y$ to adjust for covariates, conceptually consistent with the covariate-adjusted RD estimation by Calonicoetal2019regression and the ideas of AngristandRokkanen2015wanna.} We use the empirical distribution of the running variable from each of the 11 papers but normalize it to lie in $[-1,1]$ by dividing $X$ of each data point by the maximum $|X|$ on its side of the cutoff, with zero representing the location of that policy cutoff. We also remove the most extreme observations where $|X|>0.99$, following ImbensKalyanaraman2012 and Calonicoetal2014. For two papers which feature semi-discrete running variables, we add small amounts of normal noise to the running variable to match the regularity conditions from Calonicoetal2015. To create a continuous CEF, we fit global piecewise quintics (still following ImbensKalyanaraman2012 and Calonicoetal2014) and vertically shift the right arm. We specify the distribution of the error term $u$ as i.i.d.\ normal with mean zero and standard deviation $\sigma$, which we set as the root mean squared error (RMSE) of the piecewise quintic regression. We use the same number of observations as the original paper minus any observations removed while trimming the data. Plots of the resulting CEFs before we vertically shift their right arm to make them continuous are in Figure (ref), and Figure (ref) illustrates the construction process. We describe the DGP creation process in full detail in Online Appendix (ref).

Because the outcomes from the 11 papers are measured in different units, we need to standardize the discontinuity levels and choose to specify discontinuity levels $d$ as multiples of $\sigma$. Alternatively, we could specify $d$ as multiples of the overall standard deviation of the outcome variable net of the discontinuity, a measure that also captures the variation due to the conditional expectation function besides the error term. Using this alternative measure turns out not to make a difference: the variance of the error term, $\sigma^2$, dominates the variance of $E[Y|X]$ in all of our DGPs, with the ratio of the two ranging from 8 to 690.

As a multiple of $\sigma$, $d$ takes on 11 values: $0,\pm0.1944\sigma,\pm0.324\sigma,\pm0.54\sigma,\pm0.9\sigma,\pm1.5\sigma$. We choose the upper bound $|d|=1.5\sigma$ based on our own visual judgment: it represents the point at which we expect every reasonable person to say a graph from any of our 11 DGPs features a discontinuity. The nonzero magnitudes of $d$ are equally spaced on the log scale. We use this scale rather than a linear scale to generate more graphs with smaller discontinuities, which are harder to detect, to better capture the shape of the power functions.

Our discontinuity magnitudes are similarly distributed to those observed in the main outcome graphs in our literature review. The average absolute value of the discontinuity $t$-statistics in our datasets from piecewise quintic regressions is 5.0 with a standard deviation of 5.3, compared to the observed mean of 3.9 with a standard deviation of 6.4. If we instead compare the distributions of the absolute value of the discontinuity divided by the control magnitude (the left intercept of the CEF), the means are similar, 1.3 in our datasets and 1.9 in the field, while our standard deviations are somewhat smaller at 2.2 compared to 7.7.

As argued in Section (ref), we want our DGPs to be representative. Although we select the papers randomly, we also need to evaluate how well our DGPs approximate the actual data from the respective studies. To do this, we adapt the lineup protocol from buja2009statistical and Majumderetal2013, which uses visual inference to conduct hypothesis testing. In our case, we test the null hypothesis that the original datasets come from the calibrated DGPs. Specifically, we present one graph of the original data randomly placed among 19 graphs from datasets drawn from the corresponding DGP. The goal is to identify the true dataset by choosing the graph that least resembles the others. If the viewer does not select the original graph, then we cannot reject the null hypothesis. Under the null hypothesis, the probability of identifying the graph produced from the original data (or the type I error probability) among the 19 simulated datasets is 5% ($1/20$) for a single reader. For our lineup protocol, each graph is a binned scatter plot using the MV bin selector. We present two examples in Figure (ref).

Based on visual testing among the authors, we cannot identify the graph from the true data for eight out of our 11 DGPs, which supports the idea that our DGPs approximate the original datasets well. For three DGPs, however, there is an obvious difference, as exemplified by DGP3 in the right panel of Figure (ref). All three “fail” seemingly because of the misspecification of the variance structure of the error term $u$. Recall that we specify $u$ as being i.i.d.\ across observations and homoskedastic. But in the right panel of Figure (ref), for example, the running variable is time, and there is positive serial correlation in the outcome. As a consequence, the outcome variability in the binned scatter plot is understated when $u$ is assumed to be i.i.d. Nevertheless, we adhere to the i.i.d.\ specification because it is standard in Monte Carlo exercises to evaluate RD estimators and inference procedures.

Another caveat of our DGP specification is that using global quintic regressions can lead to overfitting, the same issue that has brought forth the warning by gelman2019high against using high-order global polynomial regressions to estimate RD treatment effects (see Peietal2018local for related discussions on the order of local polynomial regressions). We acknowledge this potential drawback of using quintics as some of our graphs indeed feature high variation in the tails. That said, the lineup protocol we adapt offers a novel and transparent method to evaluate our DGP specifications, and we find our inability to distinguish the real data from those drawn from one of our DGPs in a majority of cases reassuring. To further assuage the concerns regarding our DGP specification, we carry out a supplemental phase of experiments to gauge the sensitivity of visual inference to alternative DGP specifications. In Online Appendix (ref), we demonstrate the remarkable robustness of our results to using local linear estimates as an alternative to model the CEF (and allowing for heteroskedasticity), which is much less likely to overfit.

Non-Expert Experiments

In our randomized experiments, we present non-expert participants with binned scatter plots made from our DGPs and ask them to classify the graphs as having a discontinuity or not. We conduct five phases of computer-based experiments online through the Cornell University Johnson College's Business Simulation Lab. Our subject pool consists of current and former Cornell students, Cornell staff, and non-student local residents with an expressed interest in focus groups or surveys. Although these educated laypeople are not the primary audience for academic research, RD graphs are sufficiently transparent that they are featured in popular media articles in publications such as The New York Times, The Washington Post, and The Atlantic Dynarski2014,Sides2015,Rosen2012, suggesting the participants in our sample should be capable of interpreting the graphs.

Before the experiment, participants watch a video tutorial explaining how the graphs are constructed.\footnote{The video tutorial is available at \url{https://storage.googleapis.com/rd-video-turial/rd_video_tutorial.mp4}.} We do not instruct participants on how to make their decisions, e.g., whether only to look at points near the cutoff or mentally to trace out the CEF. The video contains an attention check with a corresponding question later in the experiment to ensure that subjects are attentive to the instructions. After the video, participants complete a series of interactive example tasks and receive feedback on their answers. As part of the instructions, we explicitly tell participants that all, some, or none of the 11 graphs they classify may feature a discontinuity.

In each phase of the experiment, we present participants with a series of RD graphs using data generated as described in Section (ref). Participants see two graphs with zero discontinuities, one each of $\pm0.1944\sigma,\pm0.324\sigma,\pm0.54\sigma,\pm0.9\sigma$, and one of either $1.5\sigma$ or $-1.5\sigma$. Participants see one graph from all 11 DGPs in a randomized order. We have up to 88 participants per treatment arm, and every graph we generate is seen by only one participant. For each graph, we ask participants whether they believe there is a discontinuity at $x=0$.

Because running an experiment with $2^{5}=32$ treatment arms is infeasible with our resources, we conduct our experiment in phases, testing only a few treatments in each phase. Table (ref) details the timeline of the experiments and lists the graphical parameters we test and hold fixed for various experimental phases. In phase 1, we test both bin width and axis scaling options. In phase 2, we test bin widths and bin spacings. In phase 3, we test imposing fit lines and a vertical line at the treatment threshold. Based on the results from these three phases, in which only bin widths and fit lines have major impacts, phase 4 tests all four combinations of those two treatments together.

Participants receive a base pay of \$3 for being in the experiment. To stimulate participant engagement and elicit participants' confidence in their response, participants can choose for each graph they classify a bonus that is either based on a monetary wager which pays 40 cents if their judgment is correct but nothing otherwise or a fixed payment of 20 cents irrespective of their classification. In Online Appendix (ref), we explore the implication of a participant's bonus choice and discuss how we use it---in addition to the type I and type II error rates---to evaluate graphical methods as mentioned in Section (ref).

We do not give participants real-time feedback on the accuracy of their responses. Instead, we report total earnings and the final tally of correct classifications at the very end of the experiment after a short exit survey soliciting demographic information and comments.

The experiments are programmed in oTree chen2016otree and pre-registered at the AEA RCT registry aearct and the Center for Open Science's OSF platform osf. The study takes participants approximately 15 minutes to complete.

Expert Study

In addition to our non-expert experiment, we conduct a study with researchers in economics and related fields who work on topics that often employ RDDs. We collect data at three technical social science seminars and online by contacting randomly selected members of the NBER in applied microeconomic fields (aging, children, development, education, health, health care, industrial organization, labor, and public) and IZA fellows and affiliates. After removing six responses from participants who completed the survey more than once, did not provide a valid email address for payment, or were not part of our recruited sample, we are left with 143 expert responses.

This expert study allows us to answer two questions. First, how do classification accuracy and the impacts of graphical techniques differ between experts and non-experts? And second, can experts correctly predict which graphing options perform best for our non-expert sample? This second question speaks to experts' ability to predict which visualization choices are best suited for interpretation by a lay audience. Because the success of a graphical technique ultimately lies in the reader's correct perception of graphs using it, it is important to understand whether experts' intuition regarding the relative advantages and drawbacks of alternative representation choices aligns with the evidence we find in practice. In related work on experts' ability to predict non-expert performance, DellaVignaandPope2018 find that economic experts are better than non-experts at estimating the effect of alternative incentive schemes on performance in a real effort task, but perform similarly to non-experts in terms of a simple ranking of incentive schemes.

Our expert study consists of two parts. The first is similar in structure to the non-expert experiment. Participants see a series of RD graphs and are asked to classify them by whether they have a discontinuity. To assess the accuracy of point estimates in addition to binary classifications of discontinuities, we also ask participants for an estimate of the discontinuity magnitude whenever they report a discontinuity. Due to sample size limitations, we do not randomize graphical treatments in the expert study, and all participants see graphs with equally spaced bins, no fit lines, default axis scaling, and a vertical line at the treatment threshold. All expert graphs use small bins, except for one seminar where participants see large bins. Four randomly selected participants receive a base payment of \$450 plus a bonus payment of \$50 per correct discontinuity classification. The bonus payment does not depend on the accuracy of the magnitude estimate.

The second part of the expert study asks about experts' preferences and their beliefs regarding non-expert performance across alternative graphical parameters. We present experts with three discontinuity magnitudes: $0$, $0.54\sigma$, and $1.5\sigma$. At each magnitude, we present four graphs, one for each combination of bin width and fit lines, in a random order using the same underlying data from the DGP where visual inference performs most closely to the average across all 11 DGPs. At each magnitude, we ask the experts to indicate which of the four treatment options they prefer and which they believe perform best and worst in our non-expert sample. We evaluate the experts' predictions about non-expert performances using phase 4 of the non-expert experiment, which tests these four treatment permutations simultaneously.

Results

Non-Expert Experiment Results and Graphical Method Recommendation

For each combination of graphical parameters in all phases of the experiment, we compute power functions based on participants' classifications of whether graphs feature a discontinuity. Using notation from Section (ref), a DGP-specific power function represents the estimates $\hat{p}(\gamma,g,d)$ for graphical parameters $\gamma$ and DGP $g$ across different levels of discontinuity $d$. An overall power function represents the estimates $\hat{\bar{p}}(\gamma,d)$.

The intercept of a power function indicates the type I error rate as defined in Section (ref). At all other discontinuity magnitudes, the power function represents the proportion of graphs with discontinuities that participants classify correctly, which can be interpreted as one minus the type II error rate when the DGP is chosen uniformly randomly from the 11 possibilities. A desirable inference method has a small intercept before quickly rising to achieve high power. Plots of the overall power functions for each phase are shown in Figure (ref), and Figure (ref) plots the corresponding DGP-specific power functions. The $x$-axis in these graphs is the magnitude of the discontinuity divided by the DGP-specific $\sigma$. This normalization facilitates aggregation and comparisons across DGPs, which then have six identical discontinuity magnitudes: 0, 0.1944, 0.324, 0.54, 0.9, and 1.5. Estimated effects on visual inference are in Tables (ref)-(ref). Because we effectively adopt stratified randomization in the design of our experiments as described in Section (ref), where each of the 11 strata is determined by the DGPs seen for every discontinuity magnitude, we obtain these estimates by regressing the participants' responses on treatment indicators and stratum fixed effects.

Phase 1 of the experiment tests the four combinations of the bin width treatments (large IMSE-optimal bins and small MV bins) with the $y$-axis scaling treatments (the default in Stata 14 and double that range). Comparing power functions, large bins have a significantly higher type I error rates relative to small bins. While both small bin treatments lead to a type I error rate of approximately 5%, the large bins have a type I error rate of around 20% to 25%. The large bins have a type II error advantage over the small bins, which is related to their higher type I error rate, but the power functions converge as the discontinuity magnitude increases. In contrast to bin widths, axis scaling has little effect on participant perception. Based on this result, we use Stata's default scaling for all subsequent phases.

In phase 2, we again test the two bin width treatments, this time interacted with the two bin spacing treatments: even spacing and quantile spacing. Note that large bins and small bins with even spacing appear in both phases 1 and 2. With this design, we can gauge the stability of visual inference across different samples from the non-expert population, and it is encouraging to see the results for these two treatments being virtually identical across phases. Comparing the power functions for these repeated treatments with their new quantile-spaced versions, we see that evenly-spaced and quantile-spaced bins perform very similarly. We conclude that bin spacing has a small or null effect on visual inference for the DGPs we test.

Phase 3 tests three treatments: the inclusion of a vertical line at the treatment threshold with and without polynomial fit lines and the omission of both the vertical line and fit lines. We find that the vertical line at the cutoff makes little difference in perception. Fit lines, on the other hand, appear to increase type I error rates in this phase, in line with a common concern that they may be overly suggestive of discontinuities.

Jointly, these three phases of experiments suggest that the presence of fit lines and the bin width choice have the largest impact on visual perceptions of discontinuities. We therefore base our analysis of expert preferences and expert predictions about non-expert performance on the interaction of these two treatments and run a final phase of experiments, phase 4, directly comparing the four possible treatment combinations. Interestingly, while the effects of bin width choice are once again robust across phases, the effects of fit lines are more muted in this phase. In particular, the treatment with small bins and fit lines has a type I error rate of only 0.052 in phase 4 but 0.175 in phase 3. This finding suggests that we cannot conclude that fit lines unequivocally result in an increase in type I errors, but that they do add uncertainty to visual inference.

We provide additional results in Online Appendices (ref) and (ref). Online Appendix (ref) includes evidence for the balance of covariates across treatment arms and the measurement of predictive power of demographic and DGP characteristic variables on visual inference performance. Online Appendix (ref) examines how large the $t$-statistics need to be for readers to detect a discontinuity visually.

Based on our experimental results, we recommend the graphical method that uses small bins, no fit lines, even spacing, default $y$-axis scaling, and a vertical line at the policy threshold as a sensible default for generating RD graphs. It has an associated type I error rate close to 5 percent in all four phases of the experiment. As we document in Online Appendix (ref), it also performs well under the two additional criteria mentioned in Section (ref). Of the five graphical parameters, bin spacing, $y$-axis scaling, and the presence of the vertical line do not appear to matter much, allowing researchers to use reasonable discretion. The other two parameters are much more important, and the use of small bins and no fit lines appears to be key for good visual inference performance.

Finally, we emphasize that our recommendation is not intended as a doctrine that practitioners must abide by. We are limited to the 11 DGPs we test. In addition, as described in Online Appendix (ref), there is an ad hoc element to the construction of small bins. In fact, we use quantile-spaced large bins in the regression kink design experiments in a previous working paper kortingetal2020wp, where the sample sizes are much larger than for our RD DGPs. In this regard, we view our recommended graphical method as a reasonable starting point based on the best evidence we have. An equally important takeaway is the value in documenting the robustness of graphical evidence given our finding of divergent visual inferences under commonly used methods.

Expert Study Results

We show most expert participants (95 out of 143) graphs generated with our preferred method as discussed above. Figure (ref) plots the expert power functions against those of the non-experts who saw the same graphs. When comparing expert and non-expert performances, we use solid (hollow) markers to indicate that the point is (not) statistically significantly different from the reference curve, and present plots of the corresponding 95% confidence intervals (created here with the large sample approximation described at the end of Online Appendix (ref) and by assuming independence between the experts and non-experts) in Figure (ref). The two groups perform similarly, with experts having a slightly higher type I error rate (approximately 8% to the non-expert 5%) and a slightly lower type II error rate. The only statistically significant differences are for the experts' marginally lower type II error rates at the 0.1944 and 0.324 discontinuities.

In addition to the aforementioned treatment, we show experts in one seminar pool (48 out of 143) graphs using the large bins and no fit lines treatment. The two groups again perform similarly, and the expert and non-expert power functions are not statistically significantly different anywhere. Both groups have type I error rates well above their corresponding small bin rates. Like non-experts, experts do worse when viewing graphs constructed with large bins.

Expert Preferences and Predicting Non-Expert Performance

We present experts' preferences and their beliefs about non-expert performance across the four considered treatments in Figure (ref). When asked about graphing options for the main graph of a paper that conveys the treatment effect, most experts report preferring small bins, usually with fit lines. These results hold at all three discontinuity magnitudes considered, including zero. Experts' predictions about the most effective treatments for non-experts tend to mirror their preferences. By a large margin, experts believe small bins with fit lines to be the most efficacious treatment for non-experts at all discontinuity magnitudes. Conversely, most experts view large bins without fit lines least favorably in the context of non-expert performance.

Comparing the expert predictions to our experimental data from phase 4, we find substantial discordance for the effects of bin width choice on non-expert classification accuracy. The best- and worst-performing treatments at each discontinuity magnitude have + and - signs, respectively, in Figure (ref). The actual power functions are shown in Figure (ref) (Figure (ref) shows the power functions based only on DGP9, the DGP used in the example graphs shown to experts in the second part of the expert study). While a majority of experts correctly identifies the bin width treatment with lowest type I error rates (i.e., most experts prefer small bins at the zero discontinuity level, either with or without fit lines), there is also significant expert support for the large bin with fit lines treatment, even when there is no discontinuity, which exhibits the greatest type I error rate in our sample. In addition, experts fail to predict the type I vs type II error tradeoff presented by the bin width choice: most experts expect large bins to perform worst even at large discontinuities, while we find this treatment arm has the lowest type II error rates in those cases. Although in the actual power functions, the effects of bin width are much more pronounced than the effect of fit lines, we find more expert disagreement regarding non-expert performance on bin width than on fit lines. We also find expert predictions to be similar whether their own visual inference performance is above or below the median.

Visual versus Econometric Inference

Here we compare the performances of visual inference from our small bin expert sample with various econometric RD procedures. We present both the overall power functions and the difference between visual and econometric inferences for each econometric procedure. For a fair comparison, we base the estimators' power calculations on their rejection decisions over the same set of datasets underlying the graphs seen by the experts. That is, the estimators “see” the same data as the experts, preventing differences driven by variation in sampling from the same DGP.

As a benchmark, our first estimator comes from a correctly specified model: a global piecewise quintic regression with homoskedastic standard errors. The power function for the corresponding 5% test compared with human performance is presented in the left panel of Figure (ref). We again use solid (hollow) markers to indicate that the difference to the comparison power function is (not) statistically significant, and provide plots of the differences in Figure (ref). We additionally include in Table (ref) type I and II error rates (the latter at both \(\lvert d \rvert = 0.324\sigma\) and averaged across nonzero discontinuities) for visual and econometric inferences.

Next, we implement the IK, CCT, and AK inference procedures, again plotting the corresponding power functions in the left panel of Figure (ref). All three procedures build upon local linear regressions but take different approaches to conducting inference. We introduce them here briefly before discussing the performance of each. We provide a more detailed review of the methods in Online Appendix (ref).

Strictly speaking, ImbensKalyanaraman2012 do not study inference but propose an MSE-optimal bandwidth selector, the IK bandwidth. The IK inference procedure we refer to is the “conventional” (terminology from Calonicoetal2014) inference procedure practitioners typically implement in conjunction with the IK bandwidth in which the asymptotic bias is ignored. As seen in the left panel of Figure (ref), the IK procedure achieves even lower type II error rates than the piecewise quintic estimator, and is significantly better than visual inference at detecting discontinuities up to $1.5\sigma$. But with this advantage in type II error rate comes a significant disadvantage in type I error rate. When there is truly no discontinuity, the estimator still rejects the null hypothesis in 22.6% of datasets.

Unlike IK, the CCT procedure by Calonicoetal2014 directly estimates the asymptotic bias, centers the confidence interval at the bias-corrected estimate, and adjusts the width of the confidence interval to account for the uncertainty in the bias estimate. CCT also generalize IK and propose a new class of MSE-optimal bandwidth selectors, which are implemented in the Stata package rdrobust along with the inference procedure.\footnote{While the CCT bandwidth remains the default and modal choice of rdrobust, new work by Calonicoetal2020optimal proposes inference-optimal RD bandwidth selectors. As seen in Figure (ref), using these bandwidths reduces the excess type I error rate relative to visual inference.} Note that although CCT's type I error rate is approximately 12.5%, as seen in the left panel of Figure (ref), this could be a small sample problem. As mentioned in Section (ref), we effectively have 88 datasets modulo the discontinuity level for the expert study. In a separate Monte Carlo simulation with 1,000 draws, the type I error rate is in line with that of the experts at 7.0%. However, it is possible that the type I error rate of visual inference also decreases over graphs based on these alternative datasets---the realization of the disturbance term can impact both econometric and visual inference---and therefore we keep the comparison based on the datasets the experts saw. The CCT inference procedure achieves lower type II error rates, enjoying a significant 15 to 20 percentage point advantage over expert visual inference at intermediate discontinuity levels.\footnote{One may be concerned that visual inference's lower type I error rate is a consequence of our experimental design where the majority (nine out of 11) of graphs feature a discontinuity. If subjects speculate that only about half the graphs feature a discontinuity, our type-I-error-rate result is biased in favor of visual inference. As mentioned in Section (ref), we explicitly tell participants that “all, some or none of the 11 graphs you see in this survey may feature a discontinuity,” and we present evidence against this bias in Online Appendix (ref), where we test dynamic visual inference.}

Finally, the AK inference procedure by ArmstrongandKolesar2017 adapts donoho1994statistical and produces asymptotically valid and minimax (near-)optimal confidence intervals over a class of conditional expectation functions with a bound---loosely speaking---on the second derivative magnitudes just above and just below the cutoff. This bound informs the worst-case bias of a local linear estimator over the class of functions, and AK corrects for this bias in their procedure (unlike CCT, AK's bias correction is nonrandom). Using the default rule-of-thumb method in the RDHonest R package to estimate the tuning parameter, the AK inference has a type I error rate of approximately 6%, and the power function is very close to that of the experts.\footnote{ArmstrongKolesar2017 propose analogous confidence intervals that maintain coverage and enjoy minimax optimality over a H\"{o}lder class of functions, which is determined by a global, as opposed to local, bound on the second derivative of the CEF. Though not presented here, we find the corresponding power functions to be similar. Like ArmstrongandKolesar2017,ArmstrongKolesar2017, imbens2019optimized also adapt the idea of donoho1994statistical. They propose an RD estimator through numerical optimization that is minimax mean-squared-error optimal over CEFs with a global second derivative bound. Because the corresponding inference procedure performs similarly to ArmstrongandKolesar2017 in simulations by Peietal2018local in their 2018 working paper version and can be computationally demanding, we do not implement the imbens2019optimized procedure here.} Given the minimax optimality of AK, the comparable performance of visual inference is remarkable. However, it is worth emphasizing that although they have approximately the same average type I error rate, AK offers a theoretical guarantee to control the (asymptotic) type I error rate for all DGPs in the Taylor class while visual inference does not.

The IK and CCT inference procedures at the 5% level exhibit higher type I and lower type II error rates than visual inference. We can circumvent this tradeoff by adjusting their type I error rate to the level of visual inference: we search for alternative critical $t$-values such that the resulting type I error rate of the econometric inference procedure is equal to that of visual inference and then use those critical values to conduct inference. For IK, this critical $t$-value is 2.46, and it is 2.28 for CCT. We present the results for these type-I-error-rate-adjusted inference procedures in the right panel of Figure (ref), with the differences between these procedures and expert visual inference in Figure (ref). Despite the extent of their differences in type I error rates relative to visual inference, both econometric procedures' type II error rates only increase by around 5-10 percentage points from this adjustment and are still significantly lower than those of visual inference at moderate discontinuities. However, we again caution that this result is specific to the sample of the 11 DGPs we consider; the type-I-error-rate adjusted IK and CCT procedures are not guaranteed to have correct coverage over the broader Taylor class as AK does.

We provide two sets of additional results in Online Appendix (ref). First, to investigate the mechanisms that underlie the econometric methods' performances, we impose our knowledge of the DGP when implementing them. We find that the driving force of IK's high type I error rate appears to be the noisy estimates that feed into the optimal bandwidth formula. When we use the theoretical optimal bandwidth, the corresponding type I error rate is lower and statistically indistinguishable from visual inference's.

Second, we venture beyond binary classifications of a discontinuity and compare the accuracy of discontinuity point estimates across visual and econometric approaches. We find that a simple method like IK attains a lower RMSE than visual inference for each of the 11 DGPs. We conclude that visual inference is less competitive at estimating discontinuity magnitudes than at identifying their existence.

The Complementarity of Visual and Econometric Inferences

Thus far, we have studied the “marginal” power functions for visual and econometric inferences. In this section, we use the joint distribution of visual and econometric discontinuity tests to explore their complementarity and evaluate the performance of a simple combined visual-econometric inference procedure.

First, we examine the joint distribution of visual and econometric inferences to see whether they tend to agree on the same data. For each discontinuity magnitude, we characterize the joint distribution of visual and econometric classifications in the form of a two-by-two contingency table. We conduct Fisher's exact test for independence and present one-sided $p$-values in Table (ref) by discontinuity magnitude and for each econometric method. We report one-sided $p$-values because two-sided $p$-values are method-dependent due to the ambiguity in classifying contingency tables as extreme in the opposite direction (see, e.g., Agresti1992).\footnote{In principle, we could also report correlations between inferences, but they may be hard to interpret. Because our classification variables are binary, the maximum of the correlation measure depends on the marginal distributions of the classifications. For example, if the probabilities of rejecting no discontinuity are different between two inference procedures, then their correlation is strictly less than 1. Because the (marginal) classification probabilities vary across discontinuity levels and by methods, it is difficult to compare the correlation measures.}

Two patterns emerge from Table (ref). The first row shows no strong support for an association between visual and econometric classifications when the true discontinuity is zero. But when the true discontinuity is nonzero (rows two through six), there appears to be strong evidence in support of association (though not reported in Table (ref), all correlations are positive in these cases). In other words, type II errors by experts are predictive of type II errors by various econometric inference methods, but this is not true for type I errors, which highlights the complementarity of visual and econometric inferences.

Second, we illustrate this complementarity more concretely by studying the performance of a particular combined visual-econometric inference procedure. It infers a discontinuity if and only if both the visual and econometric procedures reject the null hypothesis of no discontinuity. As a referee points out, many researchers may already use this procedure informally when reading or writing RD papers.

We plot the resulting power functions in the top panel of Figure (ref) along with the power function of the AK procedure for comparison. Because of the lack of dependence between visual and econometric classifications when $d=0$, the combined inferences achieve lower type I error rates than either of the individual inference types. On the other hand, the same mechanism pushes type II error rates higher, but the positive associations between the classifications when $d\neq0$ help limit their increase.

In fact, the power function of the combined visual-IK inference procedure is fairly close to that of AK. The bottom panel of Figure (ref) presents the difference between the two. Despite the limitations of the IK procedure, with a type I error rate above 20% as shown before, the IK-expert hybrid has a type I error rate of 2.6% while not performing statistically significantly differently from AK at any of the nonzero discontinuity levels. This finding helps to explain the enduring credibility of RDDs despite potential issues with the econometric inference method used prior to CCT and AK. It also suggests that the de facto type I error probability may be lower than the nominal level if researchers informally combine different statistical evidence instead of relying on a single econometric inference result, a point that deserves attention in future research.

Conclusion

This paper studies visual inference and graphical representation in RD designs via crowdsourcing. Through a series of experiments and studies that recruit both non-expert and expert participants, we provide answers to two sets of questions. First, how do graphical representation techniques affect visual inference and which technique should practitioners use? And second, when presented with well constructed graphs, how does visual inference perform compared to common econometric inference procedures in RDDs?

To answer the first set of questions, we experimentally assess how five graphical parameters impact visual inference accuracy. We find that generating graphs with the Calonicoetal2015 IMSE (large) bin selector leads to higher type I error rates but lower type II error rates relative to their MV (small) bin selector. Imposing fit lines can have a similar effect as using large bins, confirming the worries by ct2021nber,cattaneo2021regression. We recommend the graphical method of using small bins and no fit lines as a sensible starting point in practice, which we show performs well under several criteria. Bin spacing, a vertical line at the policy threshold, and $y$-axis scaling have little effect, implying that researchers can adhere to reasonable preferences.

For the second set of questions, we find that visual inference performs competitively on graphs constructed with the recommended method. It achieves a lower type I error rate than econometric inference at the 5% level based on the ImbensKalyanaraman2012 and Calonicoetal2014 methods (the difference between the visual and CCT type I error rates is not statistically significant), though the two econometric inference procedures offer considerable type-II-error advantages. The performance of visual inference is very similar to that based on the procedure suggested by ArmstrongandKolesar2017. Furthermore, visual and econometric inferences appear to be complementary. Through the analysis of the joint distribution of visual and econometric tests we find that, while they commit similar type II errors, there does not appear to be a strong association in their type I errors.

Our study is subject to several important limitations. The first is the restricted set of parameters we are able to test experimentally. As mentioned earlier, we do not impose fit lines with confidence intervals in our graphs, which researchers sometimes do, due to the difficulty in explaining it to non-expert participants. We also do not vary the size and color of the dots. However, these choices may impact inference due to their effect on visual attention and visual complexity as suggested by the literature on the psychological and neurological mechanisms underlying the processing of (visual) information hegarty2010thinking,kriz2007top,rosenholtz2007measuring,wolfe2004attributes. We leave these investigations to future work.

Second, our results are based on a specific set of DGPs. For example, while the bin spacing choice---equally spaced or quantile spaced---appears immaterial in our experiments, it could be important when the distribution of the running variable is farther from uniform than in our DGPs. On the other hand, the number of DGPs used in Monte Carlo simulations that lead to methodological recommendations is often far lower than our 11, and those DGPs sometimes bear no resemblance to real-world data. In addition, we test the validity of our simulated datasets by adapting the lineup protocol from Majumderetal2013 to assess the degree to which our DGPs approximate the original data, and we document the robustness of our experimental results to alternative DGP specifications.

And third, the mechanism of RD visual inference remains elusive. In a previous working paper kortingetal2020wp, we reported the results from an eyetracking study, in which we sought to identify eyegaze patterns (e.g., “visual bandwidths”) that robustly predict visual inference success. Had predictive patterns emerged from the eyetracking study, we would have followed up with additional experiments, in which we instruct a random subset of the participants to focus their visual attention according to our finding. But we were not able to identify predictive ocular patterns and could only conclude that the processing of visual signals, as opposed to where in a graph participants looked, drove visual inference success. A next step toward better understanding the mechanism is to systematically study the types of DGPs for which visual inference performs well and poorly.

These limitations notwithstanding, our study answers the call by leek2015statistics to provide empirical evidence on best practices in data analysis, and our approach can find applications in other important areas. We have conducted analogous experiments to study visual inference and graphical representations in RKDs using DGPs based on CardLeePei2009, Cardetal2015, and Cardetal2015ECMA (GanongJager2018 also discuss RK visual inference, albeit informally). Interested readers can consult our previous working paper kortingetal2020wp for our nuanced findings. Another related topic to study follows the recent work by Cattaneoetal2019, who, among other contributions, propose econometric tests of linearity and monotonicity based on binned scatter plots, which are motivated by studies such as Chettyetal2011 and Chettyetal2014. One could assess the impact of the graphical parameters on reader perception and compare visual and econometric linearity/monotonicity tests. A third related topic is structural breaks in time series econometrics, in which graphs serve practically the same purpose as those in RDD. Finally, within time series econometrics, studying visual inference for unit root/stationarity analysis may also be promising.\footnote{In recent work, ShenWirjanto2019 propose a new framework for stationarity tests, which formalizes the intuition that a visual characteristic of stationary time series is the infinite recurrence of “simple events” asymptotically.} In an influential textbook, StockWatson3E conduct an augmented Dickey-Fuller (ADF) test for the presence of a unit root in U.S. inflation. Upon finding that the test rejects a unit root at the 10% level but not at the 5% level, StockWatson3E write “The ADF statistics paint a rather ambiguous picture\ldots Clearly, inflation in {[}the figure{]} exhibits long-run swings, consistent with the stochastic trend model.\textquotedblright In this case, StockWatson3E apply visual unit root inference when the test statistic is marginal, which raises the question: can we leverage our eyes to begin with?

singlespace

Tables

table[table omitted — 3,368 chars of source]
table[table omitted — 2,511 chars of source]
table[table omitted — 2,317 chars of source]

Figures

figure[figure omitted — 1,601 chars of source]
figure[figure omitted — 596 chars of source]
figure[figure omitted — 977 chars of source]
landscape\begin{figure}[H] \caption{Lineup Protocol Graph Examples: DGP10 and DGP3} \begin{minipage}[t]{0.8\columnwidth} {\scriptsizeNotes: One of the 20 graphs for each lineup protocol is produced from the real data. The other 19 are produced from simulated data drawn from the DGP calibrated to the real data. For DGP10 on the left, the graph produced from data used in the original paper is in row $-3\cdot2+7$ and column $\sqrt{4+5}-2$ (using conventional matrix index notation), while the remaining graphs are generated from our specified DGP (we follow Majumderetal2013 and use simple arithmetic to indicate the graph location, so that readers do not accidentally see the answer before reaching their own conclusion). For DGP3 on the right, the graph made with the original data is in row 3 and column 2.} \end{minipage} \end{figure} \begin{figure}[H] \caption{Power Functions by Experimental Phase} \begin{minipage}[t]{0.85\columnwidth} {\scriptsizeNotes: Plotted are empirical power functions from the four non-expert experiment phases. As described in Table (ref), we test the effects of two graphical parameters in each phase while holding fixed the other three parameters. The legend underneath each graph indicates the combination of graphical parameters each phase tests. The power functions are defined in Section (ref). The discontinuity magnitude on the $x$-axis is specified as a multiple of the error standard deviation. The $y$-axis represents the share of respondents classifying a graph as having a discontinuity at the policy threshold. Large bins corresponds to the Calonicoetal2015 bin width selector that minimizes the integrated mean squared error of the bin-average estimators of the conditional expectation function; Small bins corresponds to the Calonicoetal2015 bin width selector that aims to approximate the variability of the underlying data; Quantile spacing indicates that bins were spaced by quantiles rather than evenly spaced; Fit line indicates the presence of parametric fit lines; Vertical line indicates the presence of a vertical line at the policy threshold; Normal scale corresponds to the default $y$-axis scaling using Stata 14; \textit{Large scale} doubles that default range.} \end{minipage} \end{figure}
figure[figure omitted — 1,216 chars of source]
figure[figure omitted — 1,198 chars of source]
landscape\begin{figure}[H] \caption{Expert Visual vs Econometric Inference} \begin{minipage}[t]{0.75\columnwidth} {\scriptsizeNotes: PQ uses a correctly specified regression model with global piecewise quintics above and below the treatment threshold and assuming homoskedasticity. IK is based on a local linear estimator using the IK bandwidth. CCT Robust is the default RDD inference procedure from CCT's }\scriptsizerdrobust{\scriptsize. AK uses the }\scriptsizeRDHonest{\scriptsize procedure with the rule-of-thumb bound on each DGP's second derivative. Markers in the figure are shown as solid, matching the legend, whenever the econometric inference procedure performs statistically significantly differently at the 5% level from expert visual inference at the same discontinuity magnitude. Markers in the figure are shown as the same shape but hollow instead whenever the econometric inference procedure does not perform statistically significantly differently at the 5% level from expert visual inference at the same discontinuity magnitude. Expert visual inference is shown for the case of small bins, corresponding to the Calonicoetal2015 bin width selector that aims to approximate the variability of the underlying data.} \end{minipage} \end{figure}
figure[figure omitted — 2,241 chars of source]

\pagenumbering{arabic}