Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
64,343 characters · 13 sections · 39 citation commands
Manipulation testing based on Benford's Law for discrete scores
Regression discontinuity design (RDD) is a widely utilized treatment effect estimator and is considered one of the most credible techniques. Like any other estimator, its validity hinges on a set of identifying assumptions. A crucial assumption is the absence of sorting or manipulation, which posits that units do not have perfect control over the running variable (or score) used to determine treatment. This assumption ensures that the assignment to treatment is effectively “as if random” around the cutoff point. In some empirical applications, the presence or absence of manipulation can be readily assessed based on the study design. For example, when the running variable is temporal, such as individuals' or firms' age, and the cutoff is defined as a specific point in time (e.g., age for eligibility for the treatment), manipulation is unlikely if the events influencing the outcome occurred prior to the cutoff. A case like this is reported by mellace2023short, who analyze a policy fostering young innovative firms in Italy that can be considered as eligible based on their date of birth. In many other cases, manipulation is a potential concern, and its likelihood must be empirically assessed. \\ When perfect manipulation is not possible, random change would place roughly the same number of units on either side of the cutoff, leading to a continuous probability density function when the score is continuously distributed. Along this line, the seminal work by mccrary2008manipulation introduced the idea to estimate twice the density of the units around the cutoff once at the left-hand-side (LHS) and once at the right-hand-side (RHS) and to perform an equality test. From a statistical point of view, the idea was first implemented by a nonparametric local-polynomial estimator, which introduces additional tuning parameters to pre-bin the data. Successively, the methodology was enhanced by otsu2013estimation, who used boundary-corrected kernels, and more recently cattaneo2020simple developed a set of manipulation tests based on local-polynomial density estimator, which does not require pre-binning of the data. \\ The null hypothesis of these McCrary-type tests is that there is “no manipulation" of the density at the cutoff, formally stated as continuity of the density functions for control and treatment units at the cutoff. It follows that failing to reject the null implies no statistical evidence of manipulation and offers evidence supporting the validity of the RDD cattaneo2019practical.\\ The papers mentioned are of paramount relevance in the RDD environment. However, they offer a method for assessing manipulation only when the score has a continuous distribution. In most of the applications, assuming the continuity of the running variable is hard and sometimes inadequate. By contrast, in empirical cases a discrete setting is standard, where the score takes a countable and finite number of realizations. It follows that one cannot use the concept of continuity and must refer to the weaker concept of asymmetry around the cutoff. \\ In this vein, an equivalent version of the McCrary-type test for discrete running variables is given by frandsen2017party, who follows the same idea, counting units on the two sides of the cutoff. This approach is reported also in cattaneo2024practical, where the authors present the density of the running variable by discussing a falsification test counting the number of observations above and below the cutoff (see Section 2.3.2). However, a counting procedure does not guarantee a difference between the two sides of the cutoff. One can reproduce the same behavior of the score above and below the cutoff, even if the number of occurrences varies. \\ This paper elaborates on this point. Specifically, we propose a new type of test for manipulation of discrete scores -- i.e., scores taking a finite number of values -- that is realistically grounded in the cutoff-based symmetry of the density instead of unreliable concepts of continuity and counting procedures. Specifically, we adopt an approach based on a natural law, notably the so-called Benford's Law (BL). As we will see, such a law is particularly suitable for our purposes. One of the main advantage of our strategy is to avoid discretionary choices on the part of the researcher. In particular, when carrying out a McCrary-type test following the procedure suggested by cattaneo2018manipulation, the researcher selects the local polynomial order to construct the density estimators, the local polynomial order used to construct the bias-corrected density estimators, the density estimator method and the kernel function. In contrast, BL provides us with predetermined thresholds of the support of the running variable within which we carry out the comparison between RHS and LHS of the cutoff. Furthermore, BL predetermines the density values for these thresholds, thus avoiding estimation issues. Importantly, BL is not an ad-hoc choice. Such a probabilistic law is a natural property of the digits of the elements of empirical datasets; thus, it is a universal regularity of the data that removes possible arbitrariness in the analytical approach. For an overview of BL, we refer to berger2015. The main characteristics of this statistical law are appreciated mainly when BL does not hold. Indeed, this case is often associated with data manipulations and artificial intervention to modify the elements of the dataset. Literature provides clear evidence on the matter arezzo2023benford, ausloos2016, ausloos2021, mir2014, cerioli2019, todter2009. For this reason, BL is particularly appropriate in our context. However, this natural law tends to fail in some cases that are inherently associated with the constitutive properties of the dataset considered. Specifically, the dataset should be of large cardinality, and data should not be theoretically bounded in a finite range. Importantly, as we will see, this paper overcomes these limitations by referring merely to the statistical definition of the BL as a discrete probability distribution. Indeed, we avoid the constraint of referring exclusively to the digits of the data, hence corroborating the universality of the approach followed. It follows that another advantage of our procedure is that the bandwidth (BW) within which carrying out the test is automatically suggested by BL, avoiding once again discretionary choices of the optimum data-driven method to be chosen. \\ The procedure presents two phases. In the first phase, we provide a new BL-based optimization criterion for computing an optimal data-driven BW around the cutoff; this criterion is of the iterative type. In the second phase, we propose two new methods for assessing the asymmetry of the running variable at the cutoff. Briefly, the procedure for assessing the asymmetry is as follows. We split the distribution of the running variable within the optimal BW into two parts -- above and below the cutoff. Then, we adopt a twofold approach. In the first approach, we identify the thresholds on the $x$-axis such that the running variable obeys the BL probability distribution. Such an identification is implemented above and below the cutoff, after a suitable rescaling of these two parts of the density function such that their integral on the two sides sums up to unity. Then, these thresholds are statistically compared to determine whether the distribution exhibits symmetry properties when observed before and after the cutoff. In the second approach, we use the thresholds found above for deriving the ones having the same distance from the cutoff and compare the empirical probabilities through a MAD test. These approaches provide a comprehensive view of all aspects related to the asymmetry in the case of discrete running variables. \\
The paper proceeds as follows. Section (ref) introduces BL and its properties, along with its applications in economics. Section (ref) explains the iterative procedure for deriving the optimal BW. Section (ref) sets out the testing methodologies. Section (ref) contains extensive simulations, while two applications to relevant empirical cases is offered in Section (ref). Finally, Section (ref) presents some concluding remarks.
The BL belongs to the natural law category, arising from observations of real-world phenomena. Indeed, it was first observed by Simon Newcomb in 1881 necomb1881note, who noticed the pages of logarithm tables near the beginning of the book were more worn. Frank Benford independently rediscovered and popularized the law in 1938 benford1938law. BL, also known as the First-Digit Law, describes the surprising non-uniform distribution of leading digits in many real-world datasets. Contrary to the intuitive expectation of equal probability for each digit $1, \dots, 9$, BL predicts a highly skewed distribution, with smaller digits appearing more frequently than larger ones. Specifically, a set of numbers is said to satisfy BL if the leading digit $j$, for $j \in \{1, \dots,9 \}$, occurs with probability
The quantity $B_j$ is proportional to the ($\log$ of the) space between $j$ and $j+1$. It follows that this is the distribution expected if the log of the numbers are uniformly and randomly distributed. Since the interval $\left[\log_{10}(1), \log_{10}(2)\right]$ is wider than the interval $\left[\log_{10}(8), \log_{10}(9)\right]$ (around $0.30$ and $0.05$, respectively) a randomly chosen value from a uniform distribution on the log scale is more likely to fall within the wider interval, making numbers starting with $1$ more likely than those starting with $9$. Notice also that (ref) still applies also in other $\log$ basis. \\ Some key contributing factors help explaining this phenomenon, such as: $(i)$ scale invariance: the law remains largely unaffected by changes in units of measurement, suggesting its origin lies in underlying multiplicative processes. $(ii)$ Invariance under power laws: many natural and social phenomena exhibit power-law distributions (e.g., city sizes, earthquake magnitudes, income distribution). These lead to skewed leading digit distributions. $(iii)$ Multiplicative growth processes: economic and financial data often involve multiplicative growth (e.g., compound interest, economic growth), which naturally generate data conforming to BL. \\ The law has had disparate applications, such as genomics data, election data, and criminal cases, to the extent that in the US, evidence based on BL has been admitted in criminal cases at all levels of the court system (for applications of BL, see berger2015 and the references therein). Also in economics, the law has already raised attention. varian26benford first noticed that the law could be used to detect possible fraud in lists of socio-economic data submitted in support of public planning decisions. Based on the assumption that people who fabricate figures tend to distribute their digits fairly uniformly, a simple comparison of first-digit frequency distribution from the data with the expected distribution according to BL ought to show up any anomalous results. A similar reasoning has been applied to accounting data arezzo2023benford, pricing patterns el2005price, scientific data validation diekmann2007not, macroeconomic data report rauch2011fact, and network analysis tovsic2021use. In this article, we follow the intuition by varian26benford and propose a manipulation test based on BL in a RDD context. \\ The BL accuracy has been documented by cai2020surprising, but despite its attractiveness, it does not apply universally. Short datasets, artificially generated data, and data with specific constraints may not exhibit the expected digit distribution. However, in our approach, we do not refer to digits, but rather to the thresholds generated by the BL and the ensuing distribution, which allow to sidestep these limitations.
Due to obvious endogeneity problems, the testing procedure must be executed within an optimal BW. In our case, the optimality criterion should be refined consistently with the context we are working with. So that, in what follows, we proceed to derive a new data-driven optimality criterion based on the BL. Notably, the optimal BW is found through an iterative procedure that solves an optimization problem for determining BL-based thresholds for the running variable.
The running variable $R_i$ is the quantity of interest. We denote the probability density function of the running variable $R_i$ by $f:\mathbb{R} \to \mathbb{R}$ and the cutoff by $C \in \mathbb{R}$. We deal with the potential asymmetry of $f$ above and below the cutoff, so that $f$ exhibits different behaviors to the left and to the right of $C$.
It is important to notice that the number of available data points might affect the approximation of Benford's distribution with the thresholds in $\mathcal{X}_{n}^{(-)}$ and $ \mathcal{X}_{n}^{(+)}$, with $n=1, \dots, N$. In fact, as the number of data points increases, the computed thresholds yield a more accurate approximation of Benford's distribution. Conversely, the approximation deteriorates as the number of steps $n$ increases because the number of available observations tends to decrease. \\ On the one hand, this consideration seems to suggest that a limited number of iterations may be desirable to have a good approximation. On the other hand, we need to be as close as possible to the cutoff $C$ to reduce bias and to check for the asymmetry of the density function $f$. Therefore, we have a counteracting force pushing towards a high number of steps $n$. Substantially, there exists a trade-off between these two competing forces.\\ To describe the way the optimal BW is selected, we need an empirical dataset. We assume that this dataset has observed values of the running variable given by the set $$ \mathcal{R}= \{r_1, \dots, r_K\}. $$ We assume that this empirical dataset has the function $f$ used in the steps above as its empirical probability density function. Using the theoretical thresholds suggested by BL, $\mathcal{X}_{n}^{(-)}$ and $ \mathcal{X}_{n}^{(+)}$, we compute the empirical distribution, i.e., the histograms, by setting
where $(a,b)$ is as in Formula (ref), $j=1, \dots, 9$, $n=1, \dots, N$ and
In words, $\tilde{B}^{(-)}_{n,j}$ and $\tilde{B}^{(+)}_{n,j}$ are the share of data points in the interval $[a,b)$ suggested by the BL. The idea to optimally select the BW consists in finding $n^\star \in \{1, \dots, N\}$ such that the couple $(X^{(-)}_{n^\star,1}, X^{(+)}_{n^\star,1})$ leads to a compliance of the empirical probability distributions in (ref) with the BL according to a distance criterion. In our context, we exploit the MAD, that is particularly suitable in the BL context cerqueti2021data. Specifically, $$ MAD_n^{(-)}= \frac{1}{9} \sum_{j=1}^9|\tilde{B}^{(-)}_{n,j}-B_j| \,\,\,\,\, \text{ and } \,\,\,\,\, MAD_n^{(+)}=\frac{1}{9} \sum_{j=1}^9|\tilde{B}^{(+)}_{n,j}-B_j|. $$ The $MAD$ criterion can be regarded as the mean of the absolute deviations between the empirical and the theoretical BL distribution, i.e., a summary statistic of the distance between the observed distribution and the one predicted by the BL. According to the arguments above, the $MAD$ tends to increase as $n$ increases, but we need $n$ sufficiently large to compute many well-accurate differences. Therefore, we can introduce an acceptance level $\alpha$ for $MAD$ and define the $\alpha$-optimal BW $n_\alpha^\star$ by setting
The $\alpha$-optimal couple is $(X^{(-)}_{n_\alpha^\star,1}, X^{(+)}_{n_\alpha^\star,1})$. For an easy notation, we set hereafter $X^{(-)}_{\alpha,\star}:=X^{(-)}_{n_\alpha^\star,1}$ and $ X^{(+)}_{\alpha,\star}:=X^{(+)}_{n_\alpha^\star,1}$\\
Notice that the trivial case where no $n^\star \in \{1, \dots, N\}$ exists such that the pair $(X^{(-)}_{n^\star,1}, X^{(+)}_{n^\star,1})$ complies with BL typically arises when the dataset contains a very limited number of distinct values. Our procedure excludes these uninformative instances. Figure (ref) provides a visual representation of the procedure shown in Section (ref) for the case of $g^{(+)}(R_i)$ and assuming that $n^\star=3$.
Comparing the thresholds in $\mathcal{X}^{(-)}$ and $\mathcal{X}^{(+)}$ is a way of assessing the symmetry features of the running variable. This is exactly the point where our approach differs from those based on density comparison; we use threshold values. \\ Operationally, we follow two different perspectives to face the symmetry problem, recalling that the methods proposed are applied within the BW found in the previous section. For the sake of simplicity, we refer to $X^{(-)}_{n,j}$ and $X^{(+)}_{n,j}$ simply as $X^{(-)}_{j}$ and $X^{(+)}_{j}$, respectively, for each $j=1, \dots, 9$.
In the case of a symmetric density function $f$, one has that
that can be rewritten as
Reasonably, the condition of perfect equivalence between all the thresholds in $\mathcal{X}^{(-)}$ and $\mathcal{X}^{(+)}$ as in (ref) might be too restrictive, and defining a tolerance level may lead to subjective choices. To circumvent the problem we convert the discrepancies in (ref) into probabilities, whence the possibility to use the MAD and its associated critical values.
We first set a threshold $\alpha>0$ -- which is the closely acceptance threshold for MAD by Nigrini -- set $X^{(-)}_9:=X^{(-)}_{\alpha,\star}$ and $X^{(+)}_9:=X^{(+)}_{\alpha,\star}$ and define $\Gamma=\max \{C-X^{(-)}_9, X^{(+)}_9-C \}$ that will be used to normalize the thresholds as: $$ \tilde{X}^{(-)}_j= \frac{C-{X}^{(-)}_j}{\Gamma},\,\,\,\,\,\,\tilde{X}^{(+)}_j= \frac{{X}^{(+)}_j-C}{\Gamma}, \qquad \forall \, j=1, \dots, 9. $$ Now, the thresholds $\tilde{X}^{(-)}$'s and $\tilde{X}^{(+)}$'s are probabilities that are collected into two sets, $\tilde{\mathcal{X}}^{(-)}$ and $\tilde{\mathcal{X}}^{(+)}$, respectively. At this point we are in the position to apply a distance measure, $D$, between the probabilities in $\tilde{\mathcal{X}}^{(-)}$ and $\tilde{\mathcal{X}}^{(+)}$, as follows:
In this second set up, we assume that the endogenous thresholds $\mathcal{X}^{(-)}$ and $\mathcal{X}^{(+)}$ can be conveniently used to define other thresholds to be used for the assessment of the asymmetry in the running variable. We distinguish the cases of the right and left tail of the distribution, which are respectively the sets $g^{(+)}(R_i)$ and $g^{(-)}(R_i)$. Also in this case, we consider a threshold $\lambda>0$ and set $X^{(-)}_9:=X^{(-)}_{\lambda,\star}$ and $X^{(+)}_9:=X^{(+)}_{\lambda,\star}$. \\ Specifically, we introduce the thresholds $Y^{(+)}_k$ and $Y^{(-)}_k$ such that they are symmetric values of the $X$'s with respect to the cutoff $C$, i.e.: $$ C-Y^{(-)}_k=X^{(+)}_k-C,\,\,\,\,\,C-X^{(-)}_k=Y^{(+)}_k-C, $$ for each $k=1, \dots, 9$. We collect the $Y$'s into two sets $\mathcal{Y}^{(-)}=\{Y^{(-)}_1, \dots, Y^{(-)}_9\} \subset \mathbb{R}$ and $\mathcal{Y}^{(+)}=\{Y^{(+)}_1, \dots, Y^{(+)}_9\} \subset \mathbb{R}$. We consider
and, analogously,
where $\gamma$ is defined in (ref) and the $p^{(+)}$'s and $p^{(-)}$'s form two probability distributions. Then, we apply the MAD criterion again to measure the deviation between the BL-type distributions in (ref) and the ones in (ref) and (ref). Specifically,
Accordingly, we classify the BL-type distributions as symmetric whenever $MAD^{(+)}<\lambda$ and $MAD^{(-)}<\lambda$ and as asymmetric whenever at least one of these inequalities is violated. If $MAD^{(+)}$ and $MAD^{(-)}$ signal symmetry we conclude in favour of no manipulation. \\ In principle, it may happen that only one of the two MADs signals divergence, which still implies asymmetry. This apparent contradiction actually provides a further source of information. Ultimately, the final judgment rests upon the specific economic case under investigation. One can, for instance, consider the economic incentive that units have to place their score on the RHS or the LHS of $C$. Think of a policy in which units scoring above $C$ receive a desirable treatment. In this case, if $MAD^{(+)}<\lambda$ and $MAD^{(-)}\geq \lambda$, this suggests manipulation with no economic effects. Conversely, manipulation engenders a substantial effect when $MAD^{(+)} \geq \lambda$ and $MAD^{(-)}<\lambda$. Mutatis mutandis, very similar reasoning applies when units have an incentive to place their score below the cutoff. This is a remarkable innovation with respect to the McCrary-type tests. \\
Figure (ref) visualizes the two testing procedures. The first method (shown in red) compares the thresholds $X_1^{(+)}$ and $X_1^{(-)}$, where the area under the curve between $X_1^{(+)}$ and $C$ (denoted $B_1$) is equal to the area under the curve between $C$ and $X_1^{(-)}$. The second method compares $X_1^{(+)}$ and $Y_1^{(+)}$, which are equidistant from $C$, but the areas under the curve between these points and $C$ are different.
To gain an understanding of how the proposed testing methodology performs, we conducted a simulation exercise. Specifically, we started by generating a perfectly symmetric standard Normal (0,1) and we assumed that the cutoff is $C=0$. We then set a new value of the cutoff, $C+\beta$, with $\beta>0$ and computed the BW extremes (UB and LB). This step generates a data-driven asymmetric interval around $C=0$, which is then used in the subsequent step. Considering the true cutoff value $C=0$, in such an uneven BW, we compute the testing indices to check whether they could capture the asymmetry and signal its direction. One question that may arise at this point is: why not look at the symmetry indices when $C=\beta$? The result in this case would be obvious, the indices clearly point to asymmetry. However, we want to verify the potential of the indices to signal asymmetry even with a symmetric distribution, by considering an asymmetric interval that makes the distribution asymmetric de facto. This is a rather finer exercise. The following pseudo-code provides a schematic breakdown of the simulation exercise for $Q=1,000$ repetitions, and Figure (ref) illustrates the process visually. Notice that the threshold $\alpha$ is used only for the identification of BW. Moreover, we denote by $N$ the highest number of iterations to compute $n_\alpha^\star$, defining the optimal BW.
Since the cutoff has been shifted rightward, we expect $|LB|<UB$, implying that the interval on the LHS of the cutoff point is shorter than the corresponding interval on the RHS, thereby endogenously generating an asymmetric BW around zero. In principle, the exercise can be repeated for different values of $\beta$ as long as $LB<0$; namely, for sufficiently small shifts. \\ Given this setup, we expect $MAD^{(-)}$ to point toward symmetry, and conversely for $MAD^{(+)}$. This follows from the fact that the values of $X_{i}^{(+)}$ are expected to be larger than $X_{i}^{(-)}$ (in absolute terms), so that when reflecting $X_{i}^{(+)}$'s values on the LHS (i.e., when generating $Y_i^{(+)}$), some or even all may exceed $|LB|$. In the presence of a marked asymmetry of the BW, it can even happen that $Y_1^{(+)}$ is to the left of LB, $|Y_1^{(+)}|>|LB|$. In such a situation, all the probability mass on the LHS is contained within $LB$ and the cutoff, and the value of $MAD^{(+)}$ will be high.\footnote{If this situation applies $MAD^{(+)}=\frac{1}{9}\left[(1-0.301)+(1-0.301)\right]=0.155333$.} \\ Summing up: we expect the following three results from the simulation exercise: (i) $|LB|<UB$, (ii) $MAD^{(-)}$ taking low values (iii) $MAD^{(+)}$ taking high values. \\ Therefore, in accordance with the discussion below Formula (ref), we can introduce the concept of `partial symmetry' or `left(right)-symmetry.' This apparent oxymoron describes a situation in which only one side of the distribution conforms to BL. Such a finding can be quite informative depending on the specific economic context under scrutiny. For instance, consider a scenario where units scoring above the cutoff are entitled to a benefit. We would expect significantly more units on the RHS, resulting in $UB \gg |LB|$ and a disproportionately high $MAD^{(+)}$.
This exercise can be viewed as an in vitro experiment aimed at studying the behavior of the test proposed by thoroughly isolating the causes of potential success or failure. To enhance the readability of the simulation results, and without loss of generality, we adopted a single critical value for each metric.
The threshold $\lambda$ used for the case of Formula (ref) is selected in accordance with nigrini. Indeed, in close analogy to the popular use of critical $p$-values at $1\%$, $5\%$ and $10\%$ for rejection of the null hypothesis, nigrini identifies some variation ranges of the value of MAD to have close/acceptable/marginal conformity with BL. Notably, the threshold values for $\alpha$ are $0.006$, $0.012$ and $0.015$ so that MADs below $0.006$ identify close conformity, values between $0.006$ and $0.012$ acceptable conformity and between $0.012$ and $0.015$ marginal conformity.\footnote{See cerqueti2021data Table 2 for a correspondence between the MAD thresholds and other criteria elaborated in the literature.} While alternative critical values can be chosen, adopting the threshold established by nigrini provides an additional strength to our approach.
Since only $0.5\%$ of the cases for the $MAD^{(-)}$ fall between the close and marginal acceptance thresholds (i.e., between $0.006$ and $0.015$), we used the close conformity threshold in our experiments. Similarly, for $MAD^{(+)}$, even if this proportion is slightly higher. \\ Very briefly, the results can be summarized in the following three points:
As an initial result, the average value of the ratio $\text{LB}/\text{UB}$ is around $94\%$ and the mean $MAD^{(-)}$ value of $0.0008$ falls well below the critical threshold for close conformity. In $98.2\%$ of the cases $|LB| < UB$. \\ A breakdown of the results by BW length is reported in Tables (ref) and (ref) for $MAD^{(-)}$ and $MAD^{(+)}$, respectively. Table (ref) reveals that in all cases where $|LB| < UB$, $MAD^{(-)}$ indicates symmetry (lower-right quadrant). Furthermore, the $MAD^{(-)}$ remains robust even when interval proportions shift -- specifically when $|LB| \geq UB$. Even under these conditions, the measure correctly identifies symmetry in $1.2\%$ of cases (upper-right quadrant), bringing the total symmetry detection to $99.4\%$.
Regarding $MAD^{(+)}$, a converse pattern is observed. Its mean value of $0.0254$ stands above the threshold for close conformity. Table (ref) provides a cross-tabulation akin to that of Table (ref). The most significant finding is in the lower-left quadrant: whenever $|LB| < UB$, $MAD^{(+)}$ indicates asymmetry in $95.30\%$ of cases, with the remaining $4.7\%$ pointing to close conformity.
To evaluate the quality of the results, additional reference values beyond Nigrini's thresholds have recently been proposed by cano2025divergence, who tabulated the critical values of the MAD test. In particular, $0.3216$, $0.3492$, and $0.4046$ correspond to the $90$th, $95$th, and $99$th percentiles, respectively. However, the MAD score exhibits a mechanical negative dependence on the number of observations $M$, leading to different critical values for each value of $M$. This dependence can be accounted for by multiplying the MAD by $\sqrt{M}$. Equivalently, one can divide the percentiles by $\sqrt{M}$ barney2016moderating, de2023study, nigrini2015persistent. The principal drawback to this approach is that, with a sufficiently large $M$, even small departures from the tabulated MAD will result in a finding of no conformity; even more oddly, a large MAD with a small $M$ is more likely to lead to a conclusion of conformity than a low MAD with a larger $M$. In order not to leave this approach unexplored, we have counted how many times the MAD values are below $\frac{P_{95}}{\sqrt{M}}$ -- being $P_{95}$ the $95$th percentile -- finding that $MAD^{(-)}$ does not reject conformity to BL $100\%$ of the time, and $MAD^{(+)}$ rejects it $99.88\%$ of the time. \\
A further complementary measure of non-conformity to BL is provided by the Equivalent Contamination Proportion (ECP). This represents the proportion of non-BL conforming observations within an otherwise BL-conforming sample that would result in an expected MAD value equal to the one observed in the actual data, see cano2025divergence and cano2025much. This measure presents the advantage of being independent of sample size and offers consistent results across different divergence statistics. When applied to our simulation this methodology returns an ECP equal to $0.0\%$ and $40.84\%$ for $MAD^{(-)}$ and $MAD^{(+)}$, respectively. \\ As far as the tolerance threshold approach is considered, eq. (ref), we get $D(\tilde{\mathcal{X}}^{(-)}, \tilde{\mathcal{X}}^{(+)})=0.0737$.
Finally, for the sake of completeness, we have supplemented the simulation results with McCrary-type tests. A comparison between the two methodologies is not straightforward, as the McCrary-type test only allows for symmetric BW. This implies that repeating the experiment on the perfectly symmetric generated distribution would be a redundant exercise, resulting in the acceptance of the null hypothesis of symmetry $100\%$ of the time with the highest possible p-value. \\ Consequently, to ensure an informative comparison, we proceeded as follows: first, we allowed the rddensity routine to endogenously compute the symmetric BW around the cutoff $C=0$. Second, we rescaled the BW by applying the average proportions obtained from the simulation (i.e., $LB/UB$). Third, we replicated the McCrary test within this interval. The result is a non-rejection of the null hypothesis of symmetry ($p\text{-value} = 0.850$). \\ As a further robustness check, we repeated the test using a narrower BW, determined by the average $LB$ and $UB$ from the simulation, and obtained a qualitatively unaltered result ($p\text{-value} = 0.986$). \\ Overall, these results highlight the complementary nature of the two approaches. While the McCrary-type test provides a useful baseline by confirming global symmetry, our proposed methodology offers additional layers of information by successfully identifying specific rejections of right-symmetry. This suggests that our test serves as a more granular diagnostic tool, capable of detecting nuanced distributional asymmetries that remain overlooked by traditional tests.
We illustrate our methodology using two popular datasets that have already been used for similar purposes. In the first example we used the dataset from meyersson2014islamic who studied the causal effect of Islamic political representation winning municipal elections in Turkey on women's highest educational degree. The score in this RDD is the margin of victory of the largest Islamic party in the municipality. The second example is taken from londono2020upstream, in which the authors examine how financial aid influences post-secondary enrollment, college choice, and student composition. Their analysis is based on a large-scale program in Colombia that provides high-achieving, low-income students with access to high-quality colleges. Eligibility for the program requires students to exceed a cutoff in a national standardized high school exit exam (SABER) and to fall below a threshold in the household wealth index (SISBEN). The empirical analysis focuses on the discontinuity generated by the SISBEN cutoff, treating the achievement requirement as an additional eligibility condition. The choice of these datasets is motivated by evidence of continuity of the running variable around the cutoff, as documented by the McCrary test (see the original papers). However, the distributions of the variables considered are asymmetric, highlighting a key difference between the present paper and the existing contributions.
Figure (ref) plots the empirical density of the normalized running variable which ranges from $-1$ up to $0.9905$ with an average value of $-0.281$ and a median around $-0.314$.
The iterative procedure for BW selection has been repeated ten times ($N=10$) and the MADs are reported in Table (ref).
For the values to the RHS of the cutoff, the iterate 4 identifies both $MAD^{(-)}$ and $MAD^{(+)}$ in close conformity to the BL, with $MAD^{(-)}_4=0.0023<0.006$ and $MAD^{(+)}_4=0.0059<0.006$. It follows that the optimal BW is: $\left[-X_{4,9}, X_{4,9}\right]$, corresponding to $\left[-0.1138, 0.0274\right]$ . \\ We can now proceed to apply the two types of tests. In the tolerance threshold approach, we set $M=0.1138$ and $C=0$ so that \\
Accordingly, $D(\tilde{\mathcal{X}}^{(-)}, \tilde{\mathcal{X}}^{(+)})$ is computed as:
In the MAD approach, we set $Y_j^{(+)}=X_{6,j}^{(-)}$ and $Y_j^{(-)}=X_{6,j}^{(+)}$ for $j=1, \ldots, 9$, so that equations (ref) and (ref) can be operationalized as the sum of the score values within the intervals defined by $Y_j^{(+)}$ and $Y_j^{(-)}$, respectively. This yields $MAD^{(+)}=0.0502$ and $MAD^{(-)}=0.0254$. Interestingly, $MAD^{(+)} > MAD^{(-)}$, indicating that the right-asymmetry is greater than the left and that the eligibility is concentrated on the RHS. When these indices are rescaled to account for the sample size, we obtain $\sqrt{M} \cdot MAD^{(+)}=0.4166$ and $\sqrt{M}\cdot MAD^{(-)}=0.4760$, both of which are slightly higher than the $99$th percentile $P_{99}$ threshold (0.4046). Finally, the ECP indicates that the tabulated MAD values correspond to a BL-conforming distribution, with $68.87\%$ of corrupted data on the RHS and $37.42\%$ on the LHS.
Figure (ref) plots the empirical density of the centered running variable for the londono2020upstream case. Specifically, the distance to the eligibility cutoffs in terms of the households' wealth index, the SISBEN score. It is immediately apparent that the distribution of the index is slightly left-skewed. The index ranges from $-45.85$ to $56.40$, with a mean of $17.96$ and a median of $19.12$.
The results of the iterative procedure for the BW selection are reported in Table (ref). In this case, we have $MAD_6^{(+)}=0.0016$ as being the last consistent with close conformity for positive values of the score, and the corresponding $MAD_6^{(-)}=0.0034$. Thus, we have the boundaries of the BW $\left[-X_{6,9}^{(-)}, X_{6,9}^{(+)}\right]$ corresponding to $\left[-0.92 , 1.91\right]$. Mutatis mutandis, going through the same steps as for the Meyersson case, we obtain $D(\tilde{\mathcal{X}}^{(-)}, \tilde{\mathcal{X}}^{(+)})$ = 0.2254, $MAD^{(+)}=0.0154$ and $MAD^{(-)}= 0.0257$.
It is worth noticing that in this case eligible individuals are located on the LHS of the cutoff and that we computed a $MAD^{(-)} > MAD^{(+)}$, indicating that the left-asymmetry is greater than the right. In addition, $\sqrt{M}\cdot MAD^{(+)}=0.4354$ and $\sqrt{M} \cdot MAD^{(-)}=0.5401$, while the ECP is $22.05\%$ and $39.25\%$ for right and left-asymmetry, respectively.
The paper proposes a new methodology for the identification of possible data manipulations in a RDD context. Based on the BL, we first propose an iterative procedure to select a BW within which to carry out the tests. The test is broken down into two mechanisms. The first compares threshold values of the running variable while keeping the density constant, and the second compares densities while keeping the threshold values constant. This methodology can be considered complementary to the popular McCrary-type tests and offers some clear advantages. First, it does not require researcher prespecified parameter settings. Second, it overcomes the limitations of the law itself. Third, the BW within which to carry out the test is automatically suggested by BL, again avoiding discretionary choices regarding the optimum data-driven method to be chosen. The deviation of only one of the two MADs from the critical values may suggest further information on the asymmetry properties of the manipulation. In this case, we suggest considering only the MAD on the side where units are incentivized to place their score. Finally, two empirical applications suggest that our procedure is more prudential than a McCrary-type test in detecting asymmetries; therefore, one cannot refrain from employing it as a necessary complement. For instance, if the canonical McCrary test suggests manipulation, examining the different values of $MAD^{(+)}$ and $MAD^{(-)}$ may be highly informative about the source of the deviation. If this deviation originates solely from the side where units have no incentive to place their score, a more lenient final verdict may be appropriate-even if asymmetry is still present. Similarly, our proposed test can support empirical decisions when the traditional benchmark suggests rejecting the null at the $10\%$ level but not at $5\%$. In sum, while the McCrary test remains the industry standard, its reliance on symmetric bandwidths inherently limits its diagnostic power, occasionally failing to reject the null even when underlying asymmetries are present. In contrast, our approach offers a more nuanced decomposition by explicitly distinguishing between left- and right-symmetry. By uncovering a rejection of one-sided symmetry where traditional frameworks remain uninformative, this new methodology provides researchers with a more granular lens through which to evaluate density behavior, ensuring a more robust and comprehensive validation of the identifying assumptions in RDD settings. \\ At the same time, some limitations of the proposed approach call for further investigation. In particular, additional research is needed to establish clearer operational thresholds for accepting or rejecting the null of symmetry as a function of the observed MAD. A key challenge in this respect is the dependence of the MAD on sample size, which complicates the definition of universal or easily interpretable critical regions. Relatedly, future work could aim at deriving or approximating critical values for alternative summary measures, such as the ECP, in order to provide complementary and potentially more scale-invariant decision criteria. Addressing these issues would further enhance the interpretability and practical applicability of the proposed testing framework.