Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
111,002 characters · 13 sections · 112 citation commands
Unconditional Randomization Tests for Interference
JEL Classification: C12, C18, C52.\newline Keywords: Causal inference, Non-sharp null hypothesis, Dense network.
\baselineskip=18.0pt \thispagestyle{empty} \setcounter{page}{0}
When one person receives new information, they often share it with friends. Similarly, when an urban policy is implemented in a neighborhood, its effects can ripple into distant areas. In such situations, the outcome for one unit may depend not only on its own treatment, but also on the treatment assigned to others—a phenomenon known as interference. Understanding the extent of interference is crucial: it refines causal inference and model specification,\footnote{See, for example, Sacerdote2001, Miguel2004, Angrist2014, paluck2016, Jayachandran2017, Rajkumar2022, and wang2024.} and it guides efficient resource allocation, especially when interventions are costly SarahB2018, Eliana2020.
However, testing for interference presents significant econometric challenges, particularly because complex clustering patterns can render large-sample approximations intractable morgan2021, blattman2021. Even in randomized experiments, valid inference may require assumptions beyond simple random assignment aronow2012, pollmann2023causal. As a result, recent studies bond2012, blattman2021 have turned to randomization-based approaches, such as the Fisher Randomization Test (FRT), to detect interference. Yet FRTs are generally valid only for testing the sharp null hypothesis of no treatment effect, which assumes that all potential outcomes remain unchanged under any assignment—an assumption known as imputability Rosenbaum2007, hudgens2008, athey2018. In network settings, this sharp null excludes both direct treatment effects and any form of interference. Therefore, when the FRT rejects the sharp null, it may indicate the presence of direct effects, interference, or both, making it impossible to distinguish between these possibilities.
In this paper, I introduce the pairwise imputation-based randomization test (PIRT): an unconditional, design-based framework for detecting and analyzing interference in experimental settings. PIRT treats potential outcomes as fixed and considers random treatment assignment as the sole source of randomness abadie2020, Abadie.\footnote{As blattman2021 note, design-based approaches and randomization inference are particularly well-suited for network contexts with unknown spillover effects.} The core hypothesis tests whether interference exists beyond a specified distance by comparing units’ potential outcomes when they are separated by more than a threshold distance \(\epsilon_s\) from treated units. The method repeatedly reassigns treatments while holding outcomes fixed, computes targeted test statistics, and derives a p-value; a sufficiently small p-value provides evidence against the null. PIRT makes minimal assumptions about network structure—making it suitable even for dense or complex networks—and relies solely on random assignment, ensuring validity without further assumptions.
More generally, I define partially sharp null hypotheses, in which only a subset of potential outcomes is assumed known across treatment assignments zhang2023. Testing for interference is a special case of this framework, as it requires isolating direct treatment effects while assessing whether a unit’s outcome depends on others’ treatment statuses.\footnote{If one is concerned about pre-testing, the proposed method can also be directly used for causal inference, just like a traditional randomization test.} Overall, PIRT is a non-parametric, finite-sample valid, and easily implementable method for testing any partially sharp null hypothesis.\footnote{It is finite-sample exact, meaning the probability of a false rejection in finite samples does not exceed the user-prescribed nominal rate Guillaum2024.}
Testing partially sharp null hypotheses with PIRT involves addressing two key technical challenges. First, only a subset of potential outcomes is “imputable”—that is, their values can be inferred from observed data under the partially sharp null hypothesis. For example, under the hypothesis of no peer effects on non-treated units in a social network, outcomes can be imputed only for non-treated units; the outcomes of treated units remain unknown. Second, the set of units with imputable outcomes changes with each treatment assignment, since the identity of non-treated units varies across assignments. Together, these challenges complicate the direct application of traditional randomization inference methods and underscore the need for specialized approaches.
To address the first challenge, I propose a class of test statistics termed pairwise imputable statistics, each defined as a function of two treatment assignment vectors. The first assignment determines how units are grouped or compared, while the second identifies those units for which outcomes are imputable. These statistics closely resemble conventional test statistics (as in Imbens_Rubin_2015), but are restricted to the subset of imputable units specified by the partially sharp null under both assignments. Despite this restriction, pairwise imputable statistics can accommodate a wide range of commonly used test statistics. For example, a difference-in-means estimator might compare imputable individuals who have treated friends to those who do not: here, the first assignment defines the groupings, paralleling conventional test statistics, while the second determines which individuals are included in the calculation.
To tackle the second challenge, I establish a novel connection between randomization tests for non-sharp null hypotheses and the cross-sectional conformal prediction literature vovk2018a, barber2021predictive, guan2023conformal. I construct PIRT p-values via pairwise comparisons of two pairwise imputable statistics. In the first, the randomized assignment determines which units are imputable, while the observed assignment defines the groupings and comparisons; in the second, the observed assignment selects the imputable units, with groupings and comparisons determined by the randomized assignment. The validity of this procedure relies on the symmetry of these pairwise comparisons, which is analogous to the conformal lemma of guan2023conformal.
To illustrate PIRT’s applicability, I apply the method to a large-scale experiment by blattman2021, which evaluated a policing strategy that concentrated resources on high-crime “hotspots” in Bogot\'a, Colombia, using street segments as units. I use PIRT to test for interference—such as crime displacement or deterrence—in nearby neighborhoods.\footnote{This analysis assumes that interactions occur through neighboring units, resulting in potential spillover effects.} While the original authors report significant displacement effects of increased police patrols on property crime but not on violent crime, my PIRT-based analysis yields different conclusions. Specifically, PIRT detects a marginally significant displacement effect on violent crime at the 10% level, and an insignificant effect on property crime—contrary to the findings of blattman2021.\footnote{I also propose a sequential testing procedure that automatically controls the family-wise error rate (FWER) when defining the “neighborhood” of interference.} These results have important implications for welfare analysis and may reshape our understanding of criminal behavior, particularly if more severe violent crime warrants stricter interventions.
A simulation study calibrated to this dataset further demonstrates the strong empirical performance of PIRT compared to existing methods. Specifically, I test for displacement effects, in which interference causes outcomes to “spill over” to neighboring units. At the \( \alpha \) rejection level, PIRT successfully controls type I error rates, maintaining robustness even under worst-case scenarios. In contrast, classical FRT may over-reject under partially sharp null hypotheses. In terms of power, PIRT at the \( \alpha \) level outperforms competing alternatives—an especially important advantage in network analysis, where data collection is costly and interference effects are often subtle Taylor2018, breza2020. Nonetheless, there is a trade-off between ease of implementation and conservatism under the null, as PIRT may be conservative in some settings.
\paragraph{Literature Review} This paper contributes to three main strands of literature. First, it advances the study of network analysis. Following the seminal work of Manski1993, a substantial body of research has developed model-based approaches that rely on parametric assumptions Sacerdote2001, bowers2013, Toulis2013, graham2017, dePaula2018. These approaches typically require the imposition of specific network structures and must contend with the high dimensionality inherent in modeling network interactions. In contrast, my method adopts a nonparametric framework that exploits the null hypothesis to reduce dimensionality, thereby relaxing restrictive parametric assumptions and allowing for greater flexibility in capturing network effects.
Second, this paper contributes to the literature on design-based causal inference under interference. Two principal frameworks have been developed for this setting: the Fisherian and Neymanian perspectives Li2018. The Neymanian approach focuses on randomization-based unbiased estimation and variance calculation hudgens2008, Aronow2017, pollmann2023causal, typically relying on asymptotic normal approximations and often requiring assumptions such as network sparsity or local interference.\footnote{See also basse2018, viviano2022, wang2023designbased, VAZQUEZBARE2023, Leung2020, Leung2022, and Bayati2024.}
In contrast, this paper adopts the Fisherian perspective, focusing on the detection of causal effects through finite-sample valid, randomization-based tests dufour2003, lehmann2006, rosenbaum2020design. Recognizing the limitations of classical Fisher Randomization Tests (FRTs) in the presence of interference, recent literature has developed conditional randomization tests (CRTs), which restrict inference to a subset of units and assignments where the null hypothesis is sharp.\footnote{See, for example, aronow2012, athey2018, basse2019b, Puelz2022, Zhang2021, basse2024, and hoshino2023.} However, existing CRTs are often tailored to specific interference structures, such as clustered interference basse2019b, basse2024, limiting their generalizability. Moreover, designing effective conditioning events that are both powerful and feasible is challenging, frequently resulting in substantial power loss Puelz2022. Implementation of CRTs in settings with general interference can also be computationally intensive.
This paper builds on these foundations by introducing an alternative approach that applies to a broad class of interference structures, is straightforward to implement, and remains valid even when construction of informative conditioning events is difficult or infeasible. Confidence intervals for causal parameters of interest can subsequently be obtained by inverting these tests. \footnote{Randomization-based methods can also be integrated with model-based frameworks, such as the linear-in-means model Manski1993, to increase power or extend applicability beyond randomized experiments while preserving test validity ding2021, basse2024, Borusyak2023.}
Finally, this paper contributes to the literature by extending randomization testing to hypotheses that are not fully sharp. While the primary focus is on partially sharp null hypotheses defined via distance measures, the underlying principles of PIRT appear broadly generalizable, potentially extending beyond network contexts. Since neyman1935 highlighted that FRTs are limited to sharp null hypotheses, subsequent work has developed strategies for testing weak or non-sharp nulls azeem2025. For instance, Ding2016, Li2016, and ding2020 investigate randomization tests for the null of no average treatment effect, while caughey2023 considers bounded null hypotheses. Zhang2021 introduces CRTs for partially sharp nulls, employing approaches similar to those in athey2018 and Puelz2022 for time-staggered adoption designs. To the best of my knowledge, PIRT is the first unconditional randomization testing method that accommodates partially sharp null hypotheses, thus broadening the scope of randomization inference in experimental and observational studies.
The remainder of the paper is organized as follows. Section (ref) introduces the general framework, notation, and the null hypothesis of interest. Section (ref) details the PIRT procedure, including the construction of pairwise imputable statistics and the corresponding p-value based on pairwise comparisons. Section (ref) applies the method to a large-scale policing experiment in Bogot\'a, Colombia, while Section (ref) presents results from a Monte Carlo study calibrated to this setting. Section (ref) concludes. Additional empirical and theoretical results, as well as proofs, are provided in the appendix.
Consider $N$ units indexed by $i \in \{1, 2, \dots, N\}$, connected via an undirected network observed by the researcher. My goal is to study the extent of interference between units, as determined by factors such as distance, adjacency, and connection strength. These relationships are summarized by an $N \times N$ proximity matrix $G$. The $(i,j)$-th entry, $G_{i,j} \geq 0$, represents a “distance measure” between units $i$ and $j$. This measure can be either continuous or discrete, depending on the application. I normalize $G_{i,i} = 0$ for all $i$, and assume $G_{i,j} > 0$ for all $i \neq j$. The precise definition of “distance” is context-specific:\footnote{Researchers may also define distance in product space, particularly for firms selling differentiated products. In this case, units represent products, and $G_{i,j}$ could denote the Euclidean distance in a multi-dimensional space of product characteristics, as in pollmann2023causal. Such measures are useful for defining market boundaries, e.g., when merger authorities assess whether two products belong to the same relevant market.}
In this paper, I focus on experimental settings in which treatment assignment is random and follows a known probability distribution $P$, where $P(d) = \Pr(D = d)$ denotes the probability that the treatment assignment vector $D$ equals $d$. Let $X$ represent observed pre-treatment characteristics (e.g., age, gender), which can be used to control for unit heterogeneity. However, I do not attempt to estimate their direct effects on the outcome. The probability distribution $P$ may or may not depend on covariates $X$: in complete or cluster randomization, it does not depend on $X$, while in stratified or matched-pair designs, it does.
I adopt the potential outcomes framework with a binary treatment assignment vector $D = (D_1, \dots, D_N) \sim P$, where $D \in \{0,1\}^N$ and $D_i \in \{0,1\}$ denotes whether unit $i$ is treated. Let $Y(d) = (Y_1(d), \dots, Y_N(d)) \in \mathbb{R}^N$ denote the vector of potential outcomes under treatment assignment $d$, where the potential outcome for unit $i$ is $Y_i(d) = Y_i(d_1, \dots, d_N)$. This notation allows unit $i$'s potential outcome to depend on the treatment assignments of all units, thereby relaxing the classic Stable Unit Treatment Value Assumption (SUTVA) of cox1958 and accommodating settings with spatial or network interference. Throughout, I assume that the proximity matrix $G$ is unaffected by treatment assignment.
The following variables are observed: (1) the realized vector of treatments for all units, denoted by $D^{obs}$; (2) the realized outcomes for all units, denoted by $Y^{obs} \equiv Y(D^{obs}) = (Y_1(D^{obs}), \dots, Y_N(D^{obs}))$; (3) the proximity matrix $G$; (4) the covariates $X$; and (5) the treatment assignment probability distribution $P$. I adopt a design-based inference approach, treating $D$ as random, while $G$, $X$, $P$, and the unknown potential outcome schedule $Y(\cdot)$ are considered fixed throughout. For notational simplicity, these elements will not be treated as arguments of functions in the remainder of the paper.
To illustrate these notations, consider the following running example.
\paragraph{Running Example.}
Consider four street segments, labeled $i_1$, $i_2$, $i_3$, and $i_4$, where two segments are considered adjacent if they are directly connected, as depicted in Figure (ref). Units $i_1$ and $i_2$ are connected, forming one area, while units $i_3$ and $i_4$ are connected, forming another area. For simplicity, I set the distance between units within the same area to $1$. In practice, distance measures should be determined by economic intuition, and the distance between units in different areas could be arbitrarily large. However, for the sake of this example, I set it to $2$.
Suppose the outcome of interest, $Y$, is the total number of crimes recorded over a year. In this example, exactly one unit receives a randomly assigned treatment to increase policing, so $P(d) = 1/4$ for each possible assignment. Let the observed treatment vector be $D^{obs} = (1,0,0,0)$, and the observed outcomes be $Y^{obs} = (2,4,3,2)$.
Table (ref) illustrates the potential outcome schedule under the design-based framework for all assignments that have positive probability. The first row corresponds to the observed dataset. Although all potential outcomes are fixed values, only those under the observed treatment assignment are actually observed. In general, since potential outcomes can depend on the treatment assignments across all units, there could theoretically be up to $2^N$ potential outcomes.
\setcounter{example}{0}
The term “partially sharp null” was first introduced by zhang2023. I begin by providing a formal definition of the partially sharp null hypothesis.
The partially sharp null hypothesis reduces dimensionality by restricting potential outcomes to vary only across certain subsets of assignments. The set $ \mathcal{D}_i $ can vary across units and is always a strict subset of $ \{0,1\}^N $, thus offering greater flexibility than the sharp null hypothesis, which corresponds to the case where $ \mathcal{D}_i = \{0,1\}^N $ for all $ i $. For instance, researchers can specify $ \mathcal{D}_i $ based on an exposure mapping---a function linking treatment assignments to exposure levels---to test outcome constancy within each exposure level, especially when concerned about potential misspecification hoshino2023.
More generally, researchers can define alternative forms of $ \mathcal{D}_i $ that reflect specific hypotheses and research contexts, including cases where the null hypothesis is expressed as the intersection of multiple $ \mathcal{D}_i $ sets owusu2023, Puelz2022. Appendix (ref) discusses extensions of the current framework to these more general and complex hypotheses. Although the method introduced in this paper is applicable to any partially sharp null hypothesis, I specifically focus on cases where $ \mathcal{D}_i $ is defined based on a distance measure.
This definition involves two key concepts: $ \mathcal{D}_i (\epsilon_s) $ and the interval $ (\epsilon_s, \infty) $, both of which are specific to unit $ i $. The distance interval assignment set $ \mathcal{D}_i (\epsilon_s) $ maps a distance $ \epsilon_s $ to a set of treatment assignments where unit $ i $ is at least a distance $ \epsilon_s $ away from any treated units. For any $ \epsilon_s \geq 0 $, since $ G_{i,i} = 0 $, it follows that $ 1\{G_{i,i} \leq \epsilon_s\} = 1 $, implying that unit $ i $ is untreated ($ d_i = 0 $) for any assignment $ d \in \mathcal{D}_i(\epsilon_s) $. Specifically, when $ \epsilon_s = 0 $, all $ G_{i,j} $ for $ i \neq j $ are positive, which ensures that $ 1\{G_{i,j} \leq \epsilon_s\} = 0 $. As a result, there is no restriction on the treatment status of other units $ d_j $, and $ \mathcal{D}_i(0) $ includes all treatment assignments $ d $ where $ d_i = 0 $, while allowing others to be treated.\footnote{For any $ \epsilon_s < 0 $, since $ G_{i,j} \geq 0 $ for all $ i,j $, we have $ 1\{G_{i,j} \leq \epsilon_s\} = 0 $, meaning that $ \mathcal{D}_i (\epsilon_s) = \{0,1\}^N $, where all treatment assignments are included.}
The distance interval assignment set $ \mathcal{D}_i(a)/\mathcal{D}_i(b) $ corresponds to treatment assignments where unit $ i $ is within the distance interval $ (a,b] $. For any treatment assignment $ d $, the set $ \{i: d \in \mathcal{D}_i(a)/\mathcal{D}_i(b)\} $ contains all units that fall within the distance interval $ (a, b] $ relative to treated units.
Using the concept of distance interval assignment sets, I now define the partially sharp null hypothesis of interference based on distance.
Under $ \mathcal{D}_i(\epsilon_s) $, all units within $ \epsilon_s $ distance of unit $ i $, as well as unit $ i $ itself, are not treated. Hence, this hypothesis asserts that no interference occurs beyond distance $ \epsilon_s $, meaning the potential outcomes for unit $ i $ remain unchanged for any treatment assignment where unit $ i $ is at least a distance $ \epsilon_s $ away from all treated units. Under this null hypothesis, the potential outcomes for unit $ i $ can be imputed for treatment assignment vectors that satisfy this distance condition, allowing for a partial imputation of outcomes. The interpretation of distance here is context-specific and depends on the nature of the interference in the particular application.\footnote{The paper focuses on constant distance across untreated units, but the same procedure can be implemented when allowing unit-level distance and a joint test for interference for both treated units and untreated units.}
The null hypothesis defined in Definition (ref) facilitates the assessment of whether, and to what extent, interference is present within a network. In many empirical settings, researchers seek to determine whether interference effects extend beyond a particular distance $ \epsilon_s $. When $ \epsilon_s > 0 $, this framework can be leveraged to delineate the neighborhood within which interference is operative, or to identify an appropriate comparison group for subsequent estimation.
\paragraph{Comparison to the Traditional t-Test.} The traditional t-test compares units situated at various distances from treated units, but it is subject to two fundamental limitations. First, the distances between units and treated units are not random, even under randomized treatment assignment, which can induce bias absent further assumptions aronow2012, pollmann2023causal. Second, standard large-sample approximations are complicated by the presence of complex clustering patterns morgan2021, blattman2021. By contrast, the partially sharp null hypothesis articulated in Definition (ref) evaluates potential outcomes for the same unit under all treatment assignments in which it is farther than $ \epsilon_s $ from any treated unit. This approach requires only random assignment of treatment and circumvents biases that may arise from comparing outcomes across potentially non-comparable units.
\paragraph{Running Example Continued.}
Suppose researchers wish to test for the existence of spillover effects using the partially sharp null hypothesis from Definition (ref) with $ \epsilon_s = 0 $: \[ H^{0}_0: Y_i(d) = Y_i(d^\prime) \text{ for all } i \in \{1, \dots , N\}, \text{ and any } d, d^\prime \in \{0,1\}^N \text{ such that } d_i = d^\prime_i = 0. \]
Throughout the paper, I use the above $ H^0_0 $ for illustration in the running example. This hypothesis implies that the potential outcome for any untreated unit $ i $ remains unchanged regardless of the treatment assignments of other units. The potential outcome schedule under $ H^0_0 $ is displayed in Table (ref).
As shown in Table (ref), the null hypothesis $ H^0_0 $ allows us to impute many of the previously missing potential outcomes. For example, since we observe the outcome when unit $ i_2 $ is not treated, we can impute other outcomes as long as $ i_2 $ remains untreated. Consequently, the outcome for $ i_2 $ when either unit $ i_3 $ or $ i_4 $ is treated is also 4.
However, as illustrated by Table (ref), the potential outcome schedule under $H^{0}_0$ still contains missing values, complicating the use of FRT zhang2023. This highlights two key technical challenges for implementing randomization tests in the presence of interference.
\paragraph{Traditional Test Statistics.}
In practice, researchers often specify a distance $\epsilon_c$ to calibrate test power when interference diminishes with distance to treated units. For instance, in a spatial setting, $\epsilon_c$ is often set to a distance beyond which interference is considered negligible; for cluster interference, $\epsilon_c$ may be chosen to exceed the maximum distance within a cluster, so that no interference is expected across clusters. The purpose of this threshold is to separate units likely to be impacted by interference from those that can serve as clean controls.
A natural test statistic compares units within the distance interval $(\epsilon_s, \epsilon_c]$ to the treated group, while using units in the interval $(\epsilon_c, \infty)$ as a pure control group. When researchers lack prior knowledge to specify $\epsilon_c$, Appendix (ref) proposes a sequential testing procedure to help select an appropriate threshold. Even if $\epsilon_c$ is misspecified and does not provide a perfectly clean control group, the proposed procedure remains valid, though it may reduce test power basse2024.
For example, consider the difference-in-means estimator using the control distance $\epsilon_c$: \[ T(Y(D^{obs}), D) = \underbrace{\bar{Y}(D^{obs})_{\{i: D \in \mathcal{D}_i(\epsilon_s) \setminus \mathcal{D}_i(\epsilon_c)\}}}_{\text{Mean of neighbor group}} - \underbrace{\bar{Y}(D^{obs})_{\{i: D \in \mathcal{D}_i(\epsilon_c)\}}}_{\text{Mean of control group}}, \] where for sets $A_i \subset \{0,1\}^N$, we define \[ \bar{Y}(D^{obs})_{\{i: D \in A_i\}} = \frac{\sum_{i=1}^N 1\{D \in A_i\} Y_{i}(D^{obs})}{\sum_{i=1}^N 1\{D \in A_i\}}. \] In particular, $A_i = \mathcal{D}_i(\epsilon_s) \setminus \mathcal{D}_i(\epsilon_c)$ corresponds to the distance interval $(\epsilon_s,\,\epsilon_c]$, while $A_i = \mathcal{D}_i(\epsilon_c)$ corresponds to $(\epsilon_c,\,\infty)$. The difference-in-means estimator is widely used in the literature (see, e.g., basse2019b; Puelz2022).
\paragraph{Running Example Continued.}
For the remainder of the running example, I set $\epsilon_c = 1$. Thus, there are two relevant distance intervals for the difference-in-means estimator: $(0, 1]$ and $(1, \infty)$. Figure (ref) illustrates how these intervals change with different treatment assignments.
Applying traditional test statistics, such as the difference-in-means estimator, can be problematic when some potential outcomes remain unknown under $H^{0}_0$. Although the first row can be computed as $4-(3+2)/2=1.5$, Table (ref) shows that test statistics under non-observed treatment assignments still involve missing values. This occurs because randomization requires knowledge of all $ Y_i(d) $ values for the relevant assignment. This renders FRT inapplicable under the partially sharp null hypothesis and highlights two specific challenges that persist in more general settings:
First, only a subset of potential outcomes can be observed or imputed. For example, under $ H^{0}_0 $, if unit $ i_2 $ is treated, the hypothesis provides no information about the potential outcomes of unit $ i_1 $, leaving the potential outcomes for both $ i_1 $ and $ i_2 $ missing.
Second, the set of units with imputable outcomes depends on the treatment assignment. For instance, if unit $ i_3 $ is treated instead, the missing values now belong to $ i_1 $ and $ i_3 $, differing from other assignments.
The remainder of the paper develops new methods to address these challenges and enable valid inference in the presence of interference.
For simplicity, I begin by fixing $\epsilon_s$ and $\epsilon_c$, deferring discussion of their selection until the end of this section. For each treatment assignment $d$, I formally define the set of units imputable under $H^{\epsilon_s}_0$.
The set of imputable units is the subset of units for which imputation is possible, corresponding to those in the distance interval $(\epsilon_s, \infty)$ under the partially sharp null hypothesis $H^{\epsilon_s}_0$. This concept shares a similar spirit with the “super focal units” in owusu2023: given the observed treatment $D^{obs}$, the set $\mathbb{I}(D^{obs})$ includes all units with an imputable observed outcome. Units outside this set provide no additional information because their observed outcomes cannot be imputed to other treatment assignments under the partially sharp null. For example, if $\epsilon_s = 0$, then $\mathcal{D}_i(\epsilon_s)$ includes all assignments $d$ where $d_i = 0$, meaning $\mathbb{I}(D^{obs})$ consists of all units not treated under $D^{obs}$.
\paragraph{Running Example Continued.}
Under $H^0_0$, as illustrated in Figure (ref), when unit $i_1$ is treated, units $i_2$ to $i_4$ belong to the imputable set. Similarly, when unit $i_2$ is treated, units $i_1$, $i_3$, and $i_4$ are imputable. This setting corresponds to a special case where all untreated units are imputable. In more general settings, however, the composition of the imputable set depends on the value of $\epsilon_s$ specified in the null hypothesis.
As demonstrated in Figure (ref), in general $\mathbb{I}(d) \neq \mathbb{I}(d')$ for different assignments $d$ and $d'$. For instance, when testing for spillover effects among friends in a network, the set of friends whose outcomes are imputable will change across treatment assignments, reflecting the underlying social structure and the value of $\epsilon_s$.
In practice, $\mathbb{I}(D^{obs})$ could sometimes be empty, depending on the network structure and the specific partially sharp null hypothesis. If no units meet the required criteria (i.e., \( \mathbb{I}(D^{obs}) \) is empty), one approach is to reject the null hypothesis \( \alpha \) percent of the time, in line with the desired significance level. This ensures control of the test's size, even in cases where the imputable set is empty. However, to achieve power in such cases, additional data or a different study design may be necessary. See Appendix (ref) for further discussion.
The set of imputable units can also be defined under the sharp null hypothesis, though in this case, $\mathbb{I}(d) = \{1, \dots, N\}$ for any assignment $d$, meaning all units are imputable under the sharp null. Therefore, there has been less focus on the imputable units set in the randomization tests literature.
To address the first technical challenge—the presence of missing potential outcomes—I construct a class of test statistics that remain valid by relying only on the observed, imputable outcomes.
The set $\mathbb{I}(d) \cap \mathbb{I}(d')$ in Definition (ref) closely parallels the set $H$ in Definition 1 of zhang2023, capturing units whose potential outcomes are jointly imputable under both assignments. At first glance, the pairwise imputability requirement may appear restrictive. However, it is flexible enough to encompass commonly used test statistics with slight modification. For example, the classic difference in means can be written as \[ T(Y(D^{\mathrm{obs}}), D, D^{\mathrm{obs}}) = \underbrace{\bar{Y}_{\mathbb{I}(D^{\mathrm{obs}})}(D^{\mathrm{obs}})_{\{i: D \in \mathcal{D}_i(\epsilon_s) \setminus \mathcal{D}_i(\epsilon_c)\}}}_{\text{Mean of \textit{imputable} neighbor}} \\ -\ \underbrace{\bar{Y}_{\mathbb{I}(D^{\mathrm{obs}})}(D^{\mathrm{obs}})_{\{i: D \in \mathcal{D}_i(\epsilon_c)\}}}_{\text{Mean of \textit{imputable} control}}. \] where, for any collection of sets $A_i \subset \{0,1\}^N$, \[ \bar{Y}_{\mathbb{I}(D^{\mathrm{obs}})}(D^{\mathrm{obs}})_{\{i: D \in A_i\}} = \frac{\sum_{i \in \mathbb{I}(D^{\mathrm{obs}})} 1\{D \in A_i\} Y_{i}(D^{\mathrm{obs}})} {\sum_{i \in \mathbb{I}(D^{\mathrm{obs}})} 1\{D \in A_i\}}\,. \]
This construction generalizes the classical sharp null setting: when all potential outcomes are observed (i.e., under the sharp null hypothesis), $\mathbb{I}(d) \cap \mathbb{I}(d') = \{1, \ldots, N\}$ for all $d, d'$, and any standard test statistic---such as the difference in means---falls within this framework Imbens_Rubin_2015. In particular, when all units are imputable (i.e., $\mathbb{I}(D^{\mathrm{obs}}) = \{1, \ldots, N\}$), the above reduces to the standard difference in means, with groupings determined by the treatment assignment $D$.
If, for a given assignment, no unit in $\mathbb{I}(D^{\mathrm{obs}})$ falls into one of the specified intervals, the corresponding sample mean is undefined. In such cases, I set $T = \max(Y^{\mathrm{obs}}) - \min(Y^{\mathrm{obs}})$, ensuring the test remains valid, albeit conservative.\footnote{Any constant exceeding the observed statistic would suffice.} In practice, this situation did not arise in my empirical application and is unlikely in bipartite experiments with moderate $\epsilon_s$. For further discussion, see Appendix (ref).\footnote{Alternatively, to maximize power, conditional randomization testing can be used to exclude problematic assignments zhang2023.}
\paragraph{Running Example Continued.} Consider the test statistic \[ T(Y(D^{\mathrm{obs}}), D, D^{\mathrm{obs}}) = \bar{Y}_{\mathbb{I}(D^{\mathrm{obs}})}(D^{\mathrm{obs}})_{\{i: D \in \mathcal{D}_i(0)\setminus\mathcal{D}_i(1)\}} - \bar{Y}_{\mathbb{I}(D^{\mathrm{obs}})}(D^{\mathrm{obs}})_{\{i: D \in \mathcal{D}_i(1)\}}. \]
Table (ref) presents the corresponding values for the first and second terms of the test statistic, while Figure (ref) provides a visual representation of how we determine the imputable neighbor units and imputable control units.
As illustrated in Figure (ref), $D^{\mathrm{obs}}$ refers to the scenario where unit $i_1$ is treated, so the set of imputable units remains the same across different potential assignments $D$. However, the potential assignment $D$ itself can change, thereby altering which units belong to the neighborhood and control sets. When $D = D^{\mathrm{obs}}$ and unit $i_1$ is treated, the first term $\bar{Y}_{\mathbb{I}(D^{\mathrm{obs}})}(D^{\mathrm{obs}})_{\{i: D \in \mathcal{D}_i(0)\setminus\mathcal{D}_i(1)\}}$ corresponds to the outcome of $i_2$, while the second term $\bar{Y}_{\mathbb{I}(D^{\mathrm{obs}})}(D^{\mathrm{obs}})_{\{i: D \in \mathcal{D}_i(1)\}}$ is the mean outcome of $i_3$ and $i_4$. When unit $i_2$ is treated, there are no imputable units in the neighborhood set; in this case, we define $T = \max(Y^{\mathrm{obs}}) -\min(Y^{\mathrm{obs}}) = 2$ to ensure the test's validity.
The proposed test statistic satisfies Definition (ref) because only units in the intersection $\mathbb{I}(D^{\mathrm{obs}})\cap\mathbb{I}(D)$ are used to construct it. For example, if $D^{\mathrm{obs}}$ has $i_1$ treated and $D$ has $i_3$ treated, then $\mathbb{I}(D^{\mathrm{obs}}) = \{i_2, i_3, i_4\}$ and $\mathbb{I}(D) = \{i_1, i_2, i_4\}$, so their intersection is $\{i_2, i_4\}$. As shown in the third row of Table (ref), the test statistic in this case depends only on the outcomes of units $i_2$ and $i_4$.
While the method remains valid without covariate adjustments, incorporating them may improve the test's power in practice ding2021. See Appendix (ref) for a discussion on incorporating covariates. Moreover, since the proposed method is finite-sample valid, researchers can conduct subgroup analyses when different patterns of interference are expected across covariates.
Following Definition (ref) of pairwise imputable statistics, I can derive a property that allows calculation of test statistics using only observed information:
The proof is provided in Appendix (ref). By Proposition (ref), letting $d = D$ and $d^\prime = D^{\mathrm{obs}}$, we have $T(Y(D), D, D^{\mathrm{obs}}) = T(Y(D^{\mathrm{obs}}), D, D^{\mathrm{obs}})$ under the null $H^{\epsilon_s}_0$, ensuring the construction of the counterfactual test statistic.
In this paper, I focus on the unconditional randomization test framework, which is defined as follows:
The key feature of the unconditional randomization test is that the probability of rejection, $\phi(D^{\mathrm{obs}})$, is computed by randomizing the treatment assignment according to the same probability distribution $P$ that governs the original assignment. This stands in contrast to methods in the existing literature, such as athey2018, where the rejection function is based on randomizing the treatment assignment within a conditional probability space, conditioning on certain events. One example is the simple randomization test, which uses pairwise imputable statistics and constructs p-values analogously to the classic Fisher Randomization Test (FRT).
\paragraph{Running Example Continued.}
Using the pairwise imputable statistics $T(Y(D^{\mathrm{obs}}), D, D^{\mathrm{obs}})$ and following Table (ref), we can construct Table (ref) with the test statistics for each assignment.
Based on Table (ref) and following Definition (ref), the p-value is $2/4$. However, one might question whether this procedure guarantees finite-sample validity---specifically, whether it satisfies the condition $E_P(\phi(D^{\mathrm{obs}})) \leq \alpha$ under the null hypothesis.
\paragraph{Investigating Finite-Sample Validity.}
Although pairwise imputable statistics are used, naively constructing the p-value as defined in the classic FRT does not guarantee the test's validity. For the test to be valid, the following condition must hold under the partially sharp null hypothesis: \[ T(Y(D^{\mathrm{obs}}), D, D^{\mathrm{obs}}) \stackrel{d}{=} T(Y(D^{\mathrm{obs}}), D^{\mathrm{obs}}, D^{\mathrm{obs}}), \] where $ \stackrel{d}{=} $ indicates equality in distribution. The distribution on the left-hand side (LHS) is with respect to $ D $, while the distribution on the right-hand side (RHS) is with respect to $ D^{\mathrm{obs}} $.
By Proposition (ref), under the null hypothesis, we also have: \[ T(Y(D^{\mathrm{obs}}), D, D^{\mathrm{obs}}) \stackrel{\mathclap{\normalfont\mbox{$H_0$}}}{=} T(Y(D), D, D^{\mathrm{obs}}). \]
\[ T(Y(D^{\mathrm{obs}}), D^{\mathrm{obs}}, D^{\mathrm{obs}}) \stackrel{d}{=} T(Y(D), D, D). \]
Thus, for the test to maintain validity, we require: \[ T(Y(D), D, D^{\mathrm{obs}}) \stackrel{d}{=} T(Y(D), D, D). \]
This condition is not guaranteed under the partially sharp null hypothesis because $ \mathbb{I}(D^{\mathrm{obs}}) \neq \mathbb{I}(D) $ in general. Different treatment assignments $ D $ result in different sets of imputable units, leading to variability in $ \mathbb{I}(D) $. This is a key technical challenge. In the special case of testing the sharp null hypothesis, where $ \mathbb{I}(D^{\mathrm{obs}}) = \{1, \dots, N\} = \mathbb{I}(D) $, the validity trivially holds.
To address the challenges posed by varying imputable unit sets, previous literature proposes a remedy through the design of a conditioning event, consisting of a fixed subset of imputable units (referred to as focal units) and a fixed subset of assignments (focal assignments). CRTs are then performed by conducting FRTs within this conditioning event. While this approach has been influential, it also presents certain practical considerations.
First, as zhang2023 highlighted, there exists a trade-off between the sizes of focal units and focal assignments: expanding the subset of treatment assignments typically requires a reduction in the subset of experimental units. This trade-off may result in less information being utilized within the conditioning events, which can impact the power of the test. Second, constructing the conditioning event introduces an additional layer of computational complexity. This naturally raises the question of whether unconditional randomization testing remains valid in finite samples.
Whereas previous approaches ensure the validity of randomization testing by carefully designing a fixed subset of units, my method takes a different route. It avoids fixing the subset of units during implementation and instead achieves valid testing through a carefully constructed p-value calculation, thereby ensuring finite-sample validity without relying on conditioning events.
Formally, I refer to any randomization test with p-values constructed through this pairwise comparison method as a “PIRT” (Pairwise Imputation-based Randomization Test).
See Appendix (ref) for the proof. The validity result follows from Proposition (ref) and the conformal lemma in the conformal prediction literature guan2023conformal. Specifically, under the null, $T(Y(D^{obs}), D^{obs}, D) = T(Y(D), D^{obs}, D)$, which coincides with the randomized test statistic $T(Y(D^{obs}), D, D^{obs})$ when swapping the roles of $D$ and $D^{obs}$. Thus, the pairwise comparison is symmetric, which allows the application of the conformal lemma to complete the proof.
Theorem (ref) provides a worst-case validity guarantee, analogous to those in cross-conformal prediction and jackknife+ methods, due to certain pathological cases vovk2018a, barber2021predictive, guan2023conformal. As in that literature, the test empirically achieves size control at level $\alpha/2$, as demonstrated in Section (ref), but this property cannot be established theoretically due to the existence of pathological examples.\footnote{See Appendix (ref) for a more conservative minimization-based PIRT, which achieves theoretical size control with a rejection threshold of $\alpha$.}
Even with a large sample, an unbiased estimator of the p-value can be computed using Algorithm (ref), which calculates the p-value as the average over $1 + R$ draws, where $r=0$ corresponds to $d = D^{\mathrm{obs}}$. See Appendix (ref) for a detailed discussion.
\paragraph{Running Example Continued.}
Using the difference-in-mean estimator as before, \[ T(Y(D^{\mathrm{obs}}), D^{\mathrm{obs}}, D) = \bar{Y}_{\mathbb{I}(D)}(D^{\mathrm{obs}})_{\{i: D^{\mathrm{obs}} \in \mathcal{D}_i(0)/\mathcal{D}_i(1)\}} - \bar{Y}_{\mathbb{I}(D)}(D^{\mathrm{obs}})_{\{i: D^{\mathrm{obs}} \in \mathcal{D}_i(1)\}}. \]
As shown in Figure (ref) and Table (ref), for each treatment assignment $D$, the test statistic is calculated as the mean value of $i_2$ (excluding missing values) minus the mean value of $i_3$ and $i_4$ (excluding missing values). Based on Tables (ref) and (ref), I can construct Table (ref), where each row represents the values used to compare and construct the p-value for each $(D^{\mathrm{obs}}, D)$ pair.
Similar to guan2023conformal, for non-directional tests, the absolute value of the test statistic can be used. For directional tests, the statistic can be applied to test for positive effects, while the negation of the statistic can be used to test for negative effects.
\paragraph{Comparison to the CRTs}
When testing under a partially sharp null hypothesis, the p-values constructed in Definition (ref) align closely with those from CRTs if we interpret $\mathbb{I}(D^{\mathrm{obs}})$ as a focal unit set and $\{0,1\}^N$ as a focal assignment set. The pair $(\mathbb{I}(D^{\mathrm{obs}}), \{0,1\}^N)$ represents a broader conditioning event than the traditional conditioning sets used in CRTs. Depending on whether the additional potential units and assignments contribute meaningful information, this broader conditioning may or may not yield higher statistical power.
In settings where a conditioning event can be specified over all imputable units in $\mathbb{I}(D^{\mathrm{obs}})$, as demonstrated by basse2024, CRTs with a well-defined focal assignment set may allow for more targeted comparisons, potentially increasing power. Nevertheless, in scenarios where constructing a suitable conditioning event is infeasible or would result in only a limited number of focal units and assignments, the PIRT framework may provide a more practical alternative. More generally, when including all assignments from $\{0,1\}^N$ is suboptimal, combining elements of PIRTs and CRTs may further improve power by focusing on more relevant test statistics and selected assignments lehmann2006, Hennessy2015ACR. Exploring this integration presents a promising direction for future research aimed at optimizing power through the flexibility of both PIRTs and CRTs.
\paragraph{Selection of $\epsilon_s$ and $\epsilon_c$}
In 2016, a large-scale experiment was conducted in Bogot\'a, Colombia, as described by blattman2021. The study covered 136{,}984 street segments, of which 1{,}919 were identified as crime hotspots. Among these hotspots, 756 were randomly assigned to a treatment involving increased daily police patrolling—from 92 to 169 minutes per day over eight months. The experiment also included a secondary intervention aimed at enhancing municipal services, though this is peripheral to the primary focus of my empirical application. The key outcome of interest is the number of crimes per street segment, including both property crimes and violent crimes (such as assault, rape, and murder).
Figure (ref) shows the distribution of hotspots, with many located in close proximity to one another. While only 756 street segments received the treatment, every segment potentially experienced spillover effects, resulting in a dense network of possible interactions. This complexity complicates the use of cluster-robust standard errors for addressing unit correlation.
In evaluating the total welfare impact of the policy, it is essential to determine whether interference occurred following treatment assignment—such as crime displacement or deterrence in neighboring areas. To this end, I address three key considerations: (1) whether interference exists; (2) if present, whether it manifests as displacement or deterrence; and (3) the distance at which this interference is effective. Given the complexity of modeling correlations among units in such a dense network, testing a partially sharp null hypothesis, as proposed by blattman2021 and Puelz2022, is particularly relevant.
Table (ref) presents descriptive statistics for the number of crimes observed during the intervention period. The t-statistics from t-tests comparing each pair of columns reveal two salient findings: first, treated hotspots experienced significantly fewer crimes; and second, non-hotspot areas reported progressively fewer crimes the farther they were from any treated unit. However, the exceptionally high values of the t-statistics warrant caution in interpreting these results as evidence of a displacement effect. As discussed previously, standard errors may be underestimated, and units at different distances from treated areas may not be directly comparable. Both factors could contribute to the elevated t-statistics observed in the table.
The original study estimated a negative treatment effect and conducted inference using Fisher Randomization Tests (FRTs) under a sharp null hypothesis of no effect. blattman2021 reported no significant displacement effect for violent crimes and a marginally significant displacement effect for property crimes.\footnote{The treatment effect for violent crime was significant, but property crime effects were insignificant.} However, as previously discussed, p-values from t-tests may not adequately capture the extent of interference, and employing FRTs to test partially sharp null hypotheses may not be valid in this context. Consequently, it is important to consider how these conclusions might differ if a valid testing approach is employed.
To compare the power and validity of competing inference methods under spatial interference, I conduct a simulation study designed to preselect the most promising approach for the main analysis. The simulation sample consists of 1,000 units, comprising 20 hotspots and 7 randomly treated units, closely matching the proportions observed in the original Bogot\'a study. I focus on two distance thresholds, using $(\epsilon_0, \epsilon_1, \epsilon_2) = (0, 0.1, 0.2)$.
To approximate the Bogot\'a context, I calibrate the schedule of potential outcomes using gamma distributions that match the observed mean and variance of total crimes. A negative treatment effect of $-1$ is imposed, ensuring non-negative crime counts for treated units. Additionally, a decreasing displacement effect is introduced via a positive parameter $\tau$: this models spatial spillovers such that the indirect effect of treatment decays with distance from the treated unit, reflecting the empirical pattern observed in the data and representing the primary focus of the analysis.
The partially sharp null hypothesis for $k = 0$ and $1$ is given by \[ H^{\epsilon_k}_0: Y_i(d) = Y_i(d^\prime) \quad \text{for all } i \in \{1, \dots, N\}, \text{ and any } d, d^\prime \in \mathcal{D}_i(\epsilon_k). \]
In the analysis, I compare four methods: (1) the classic FRT, applied under the sharp null hypothesis of no effect, as in blattman2021 for inference on spillover effects; (2) the biclique Conditional Randomization Test (CRT) proposed by Puelz2022, which serves as a benchmark due to its strong power in simulations under general interference; (3) the PIRT with rejection at the $\alpha/2$ level, ensuring validity under the worst-case scenario; and (4) the PIRT with rejection at the nominal $\alpha$ level.
Two main criteria guide the choice of test. First, under no spillover effect ($\tau = 0$), the partially sharp null hypothesis should be rejected no more than 5% of the time, maintaining type I error control. Second, when a spillover effect is present ($\tau > 0$), the test should maximize rejection of the null, i.e., maximize power. To assess power, I consider 50 equally spaced values of $\tau$ between 0 and 1, conducting 2,000 simulations for each value, and compute the average rejection rate for each method. Due to the computational challenge posed by the biclique CRT, I first draw $5{,}000$ assignments to approximate the full space of potential treatment assignments, then construct the conditioning event.\footnote{The current code still requires over four hours to construct the conditioning events for the simulations.} See Appendix (ref) of the Supplemental Material for details.
Figure (ref) (left panel) shows that the FRT over-rejects the true partially sharp null when $\tau = 0$, consistent with athey2018's observation that testing the sharp null of no effect is invalid for partially sharp nulls. In my simulation, with only seven treated units (0.7% of the total), the FRT rejection rate is around 10%. Notably, the unadjusted PIRT at level $\alpha$ maintains good size control, indicating that the $\alpha/2$ threshold is mainly a worst-case guarantee and can be conservative, with rejection rates below 5%. The biclique CRT is also valid, with a rejection rate near 5%.
Regarding power, I exclude the FRT from comparison due to its invalidity. The unadjusted PIRT ($\alpha$) demonstrates the best performance, outperforming other methods across all effect sizes $\tau$. Among methods with theoretical size control, the PIRT with $\alpha/2$ rejection is optimal, though it slightly lags the biclique CRT at small $\tau$. Despite its validity, the biclique CRT's power increases slowly as the spillover effect grows: the rejection rate remains below 90% even at $\tau=1$.
The right panel of Figure (ref) shows a contrasting pattern. First, all methods—including the FRT—maintain validity under the null. This may reflect that hotspots rarely fall into exposure levels $(0.1, 0.2]$ or $(0.2, \infty)$, so despite a negative treatment effect, its impact on the test statistic is minimal. As with $H^0_0$, both PIRT and biclique CRT exhibit rejection rates near 5%, while the PIRT at $\alpha/2$ remains conservative.
Second, all methods show considerably lower power than for $H^0_0$. This is mainly because only 60% of units are relevant to the partially sharp null, and the spillover effect is halved ($0.5\tau$). Nonetheless, the PIRT method still achieves appreciable power when $\tau$ is large, outperforming alternatives, especially with the unadjusted $\alpha$ rejection level. Interestingly, the FRT now under-rejects, with almost no power for any $\tau$. This occurs because the FRT's $p$-value remains large unless the observed test statistic exceeds most under randomization. However, units in $(0, 0.1]$ under the observed assignment contribute to test statistics for other assignments, and these units, subject to spillover $\tau$, cause observed statistics for $(0.1, 0.2]$ and $(0.2, \infty)$ to remain low even at high $\tau$. As a result, the FRT $p$-value remains large, explaining the lack of power.
Taken together, these results illustrate how the FRT, when applied to partially sharp nulls, may yield either over- or under-rejection, depending on the scenario. Although the biclique CRT displays power, its increase is slower than with PIRT, likely due to the challenge of finding optimal conditioning events under spatial interference. The biclique method's performance may improve with advanced computing resources, as it depends on parameter choices in biclique finding. The main advantage of my method is computational simplicity; further research could expand power comparisons among approaches.
Overall, the results favor PIRT methods, especially the unadjusted PIRT. Accordingly, I apply PIRTs to replicate the results of blattman2021, using the non-absolute difference-in-means estimator.
Consider the experimental setting described in blattman2021, where the observed treatment assignment is denoted by $D^{\mathrm{obs}}$. Following their approach, let $Y$ denote the number of crimes, and let $S(D^{\mathrm{obs}})$ be an indicator for whether a unit is within 125 meters of any treated unit under the observed assignment. This proximity indicator is determined by $D^{\mathrm{obs}}$ and would change under a different assignment $D$. Although additional covariates may be included in the regression, the primary test statistic is the coefficient from regressing $Y$ on $S(D)$.
To implement the PIRT for testing the partially sharp null hypothesis in practice, proceed as follows:
Steps 3 and 4 construct the pairwise imputable statistics, relying exclusively on units untreated under both the observed and randomized assignments.
For hypothesis testing regarding displacement effects, the PIRT computes the $p$-value as the fraction of reassignments $D$ for which $\beta' \geq \beta$. The null hypothesis of no displacement effect is rejected if this $p$-value is less than or equal to $\alpha/2$. Simulation results indicate that using $\alpha$ as the rejection threshold is also empirically valid. The method can be readily adapted to two-sided tests by comparing $|\beta'| \geq |\beta|$, or to tests for deterrence effects by comparing $-\beta' \geq -\beta$.
I apply the proposed method to the publicly available dataset from blattman2021, which provides street-level treatment assignments and distance intervals at thresholds of 125, 250, and 500 meters. The dataset also includes $1,000$ pseudo-randomized treatment assignments and their associated distance intervals, as used in the original study for randomization inference. However, the absence of precise longitude and latitude data for street segments precludes extending randomization testing beyond these $1,000$ assignments.
Given the high frequency of zero outcomes, I consider both an indicator for any crime occurrence and the raw number of crimes as outcome variables. Table (ref) reports results for these outcomes across different distance thresholds.
In the main analysis, I implement the PIRT using the difference-in-means estimator as the test statistic. A key advantage of this framework is its validity regardless of the chosen statistic. Nonetheless, researchers may incorporate covariates to improve power or examine heterogeneity; see Appendix (ref) for further discussion and empirical results with covariate adjustment.
Table (ref) reveals evidence of significant displacement effects for violent crime, but not for property crime. This finding stands in contrast to the original study, which found no significant displacement for violent crime. After correcting for multiple hypothesis testing using Algorithm (ref), the PIRT detects a significant short-range displacement effect (within 125 meters) at the 10% level when using the difference-in-means estimator, and potentially at the 5% level if rejecting directly at level $\alpha$, as supported by simulation evidence. For property crime, no spillover effects are detected at any distance.
Further, for violent crime, there is evidence of additional spillover effects beyond 250 meters, with an unadjusted p-value of $0.045$ for the $(250\text{m}, \infty)$ interval when using the number of crimes as the outcome in PIRT. This suggests the existence of two types of offenders: high-risk offenders who relocate farther away from the intervention site, and lower-risk offenders who are displaced to nearby areas.
To the best of my knowledge, this is the first causal evidence of a displacement effect extending to more distant areas rather than proximate neighborhoods.\footnote{This finding suggests that the distance interval $(500m, \infty)$ may serve as a more appropriate control group than the $(250m, \infty)$ interval used by blattman2021.} However, after adjusting for multiple hypothesis testing using Algorithm (ref), these results are no longer statistically significant. Applied researchers should interpret these findings with caution in future studies.
These results not only highlight the general applicability of the PIRT method but also provide suggestive evidence for policy implications and potential criminal motives in Bogot\'a, following the insights of blattman2021. From a policy perspective, it remains unclear whether reallocating state resources to these hotspots has led to an overall reduction in crime. Further investigation is needed to identify the specific locations most affected by displacement so that those areas can be directly targeted.
Regarding criminal motives in Bogot\'a, a possible explanation—consistent with standard economic models of crime—is that violent crime in the city's hotspots is not purely expressive, as suggested by blattman2021. Instead, some violent crimes, such as contract killings, may be driven by generally mobile criminal rents. By increasing the risk of detection, intensive policing deters criminals from committing crimes in specific locations, but the crimes themselves may relocate rather than be entirely prevented. In contrast, property crimes—which are often instrumental and linked to immobile criminal rents—appear to be deterred without causing further spillover effects. As blattman2021 noted, violent crimes are often considered more severe than property crimes, making displacement effects an essential consideration when evaluating the overall welfare impact of policy interventions.
This paper introduces a testing framework for detecting interference in network settings. The proposed tests offer computational simplicity over previous methods while retaining strong power and size properties, making them highly applicable for empirical research.
Theoretically, I formalize unconditional randomization testing and PIRT, addressing two primary challenges in testing partially sharp null hypotheses: only a subset of potential outcomes is imputable, and the set of units with imputable potential outcomes varies across treatment assignments. PIRT addresses the first challenge by employing pairwise imputable statistics and the second by constructing p-values through pairwise comparisons. I prove that PIRT maintains size control, and I propose a sequential testing procedure to estimate the “neighborhood” of interference, ensuring control over the FWER.
Beyond network settings, PIRT may have broader applicability. For instance, Zhang2021 shows that partially sharp null hypotheses are relevant in time-staggered designs. This opens promising avenues for future research, including extending the framework to quasi-experimental settings and observational studies. In quasi-experimental designs, developing a unified framework that can be applied to time-staggered adoption, regression discontinuity, and network settings would be highly valuable Borusyak2023, morgan2021. For observational studies, incorporating propensity score weighting to create pseudo-random treatments and conducting sensitivity analyses would be crucial, as noted by rosenbaum2020design.
While simulations suggest that PIRT performs favorably compared to CRTs, their power properties remain unexplored. Insights from studies such as Puelz2022 on CRT power and wen2023residual on the near-minimax optimality of minimization-based p-values suggest that further investigation into the power of PIRT could yield valuable insights. Additionally, power may increase when PIRT is combined with CRTs in specific settings, making the construction of an optimal testing framework for interference an important direction for future research.