EconBase
← Back to paper

Randomization Tests in Randomized Saturation Designs

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

64,190 characters · 21 sections · 20 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Randomization Tests in Randomized Saturation Designs

\def\spacingset#1 \spacingset{1}

\onehalfspacing

abstractRandomized saturation designs are widely used to study spillover effects in clustered populations. In these designs, clusters are first assigned to treatment saturation levels, and units are then randomized within clusters according to the assigned saturation. This paper develops randomization tests for such experiments under several null hypotheses that arise naturally in spillover analysis. For a fixed pair of saturation levels, we first study two individual-level hypotheses: a partially sharp null of no spillover effect for every untreated unit and a bounded null that restricts individual spillover effects by a prespecified constant. Both hypotheses can be tested using a common conditional randomization framework, with finite-sample validity obtained by combining the same focal-unit relabeling distribution with null-specific statistics. We then study weak average-spillover nulls and show that, although these nulls do not yield finite-sample exact conditional tests, studentized relabeling statistics deliver asymptotically valid randomization-based inference. Finally, for multiple ordered saturation levels, we develop a finite-sample valid unconditional pairwise-imputation test for global monotonicity of spillover effects. Simulations and an application to the Zomba Cash Transfer experiment illustrate the finite-sample behavior and practical implementation of the methods.

Keywords: causal inference; conditional randomization test; interference; randomized saturation design; spillover effects; studentization.

\hypersetup{pageanchor=false} \thispagestyle{empty} \hypersetup{pageanchor=true} \setcounter{page}{1}

\spacingset{1.7}

Introduction

Randomized saturation designs are a central tool for studying spillovers in clustered populations. In a typical design, clusters are first randomized to different treatment saturation levels, and then units within each cluster are randomized to treatment subject to the assigned saturation. The first stage creates variation in the intensity of exposure to treated peers, while the second stage separates own treatment from peer treatment intensity. Such designs have been used in a wide range of applications and have generated an active literature on estimation and inference under partial interference hudgens2008,toulis2013,Basse2018,Basse2019,Imai2021,liutwostage.

Despite this flexibility, valid inference in randomized saturation designs remains challenging for two related reasons. First, the number of clusters is often too small to justify large-sample approximations, and cluster sizes may be highly heterogeneous (see Table (ref)). Fisherian randomization tests Fisher1935Design are attractive in this setting because they rely only on the known assignment mechanism and can deliver exact finite-sample \(p\)-values under sharp null hypotheses imbens2015causal. Second, however, many hypotheses of substantive interest concern spillover effects or total effects, and these hypotheses are typically only partially sharp on the full assignment space under interference. They restrict selected exposure conditions and therefore do not impute the full schedule of missing potential outcomes zhong2024unconditional. As a result, a naive randomization test that resamples the full assignment vector need not generate a valid reference distribution. This motivates conditioning on focal units and restricting the reference distribution to assignments for which the relevant potential outcomes are imputable Athey2018,Basse2019.

This paper develops randomization tests for randomized saturation designs under several spillover null hypotheses. For a fixed pair of saturation levels \(s\) and \(s'\), we first study two individual-level nulls. The first is a partially sharp null requiring \(Y_i(0,s)=Y_i(0,s')\) for every unit \(i\), where \(Y_i(0,s)\) denotes unit \(i\)'s potential outcome under control treatment status \(0\) and cluster saturation level \(s\). The second is a bounded null requiring \(Y_i(0,s)-Y_i(0,s')\le \delta\) for every unit \(i\), where \(\delta\) is a researcher-chosen constant. For these two nulls, conditioning on untreated focal units in clusters whose realized saturation is either \(s\) or \(s'\) converts the problem into a cluster-level relabeling problem. Under the partially sharp null, the focal outcomes are invariant across relabelings. Under the bounded null, the equality boundary provides a least-favorable imputation for the one-sided alternative. Both tests therefore have finite-sample conditional validity.

We then consider weak average-spillover nulls. These nulls do not require every individual spillover effect to satisfy a unit-level restriction; instead, they set a prespecified weighted average of spillover effects equal to zero. Leading examples include the unit-average spillover effect and the equally weighted cluster-average spillover effect. Such weak nulls are scientifically less restrictive than the individual-level equality and bounded nulls, but they do not impute the missing focal potential outcomes. Consequently, the conditional relabeling distribution is not an exact finite-sample null distribution. We show that, when paired with an appropriately studentized cluster-level statistic, the same relabeling device yields an asymptotically valid randomization-based test under regularity conditions for the corresponding weighted cluster-level array.

We also study hypotheses involving more than two saturation levels. In many applications the object of interest is not a single pairwise contrast, but the entire shape of the spillover response as the saturation level changes. For example, researchers may want to test whether untreated outcomes are monotone in the saturation level, or whether the data reveal nonmonotonicity. For this global monotone null, pairwise conditional tests would require a multiple-testing or intersection procedure. We instead develop an unconditional randomization test based on pairwise imputation zhong2024unconditional. The test compares the observed assignment with each reassignment in the design support, using only units that are untreated under both assignments and whose exposures lie in the ordered set covered by the null. This construction yields finite-sample size control for the global monotone null and extends naturally to ordered-exposure nulls beyond the randomized saturation design.

The paper makes three contributions. First, it develops finite-sample conditional randomization tests for partially sharp and bounded spillover nulls in multi-saturation two-stage designs. Existing work on randomized saturation designs has primarily studied estimation and large-sample inference for direct and spillover effects hudgens2008,Basse2018,Imai2021,Gonzalo-two-stage,liutwostage, while the conditional randomization-test literature has emphasized nonsharp nulls under interference more generally Aronow2012,Athey2018,Basse2019,puelz2021,basse2024,liu2026randomizationinferencetwosidedmarket,liu2026randomizationtestsswitchbackexperiments. Closest to our setting, Basse2019 study a two-stage design with a binary first stage and exactly one treated unit in each treated cluster. We extend this line of work to multi-saturation designs with multiple treated units per cluster, which are common in empirical applications but create a richer assignment space and a more complicated partial-imputation problem.

Second, the paper contributes to inference for weak null hypotheses on average spillover effects. Weak nulls restrict only average treatment or spillover effects, so exact finite-sample Fisherian validity is generally unavailable without additional structure chung2013,diciccio2017robust,ZHAO2021278,Wu2021,toulis2025asymptotic. In randomized saturation designs, weak spillover nulls are especially natural because researchers often want to test whether average outcomes for untreated units differ across saturation levels. We show that the conditional relabeling construction can be paired with studentized cluster-level statistics to obtain asymptotically valid tests for such average spillover nulls. This connects the Fisherian randomization-test perspective with Neymanian large-sample inference in saturation designs.

Third, the paper extends randomization inference in randomized saturation designs beyond equality nulls. Recent work has emphasized that scientifically meaningful null hypotheses need not require exact equality of potential outcomes, but may instead impose bounded, ordered, or monotonic restrictions Caughey2023,huang2025randomization. For bounded spillover nulls, we show how conditional randomization tests can be adapted to allow nonzero but bounded spillover effects. For monotone spillover nulls across ordered saturation levels, we develop a randomization test that targets the global monotonicity restriction directly. Relative to approaches based on local or pairwise ordered-exposure comparisons, our procedure tests the global null in a single randomization-based framework and applies beyond randomized saturation designs to more general ordered-exposure settings.

The rest of the paper is organized as follows. Section (ref) introduces the randomized saturation design, feasible untreated exposures, and the null hypotheses. Section (ref) develops finite-sample conditional tests for the partially sharp and bounded two-saturation nulls. Section (ref) presents the weak average null as an asymptotic extension. Section (ref) studies the multiple-saturation monotone null and presents the pairwise-imputation test. Section (ref) gives the empirical illustration. Section (ref) concludes. The appendices contain the formal conditioning construction, proofs, and implementation details.

Setup and Notation

We work under a finite-population framework, following Basse2018,Basse2019. There are \(K\) clusters indexed by \(j=1,\ldots,K\). Cluster \(j\) contains the unit set \(\mathcal I_j\), with \(|\mathcal I_j|=n_j\), and the total number of units is \(N=\sum_{j=1}^K n_j\). Let \([i]\) denote the cluster containing unit \(i\). Potential outcomes are treated as fixed throughout; randomness comes only from the known assignment mechanism.

Randomized saturation design

The assignment proceeds in two stages: clusters are first assigned to treatment saturation labels, and units are then randomized within clusters conditional on the assigned label. Let \[ \mathcal S=\{s_0,s_1,\ldots,s_L\}\subset[0,1] \] be the set of cluster-level saturation labels, where \(s_0=0\) denotes zero treatment saturation. The label \(s_0\) is a pure control label only when no other intervention components are present. Fix integers \(K_0,\ldots,K_L\) satisfying \[ \sum_{\ell=0}^L K_\ell=K. \] These counts are fixed by design, with \(K_\ell\) denoting the number of clusters assigned to saturation label \(s_\ell\).

Let \(A_j\in\mathcal S\) denote the saturation assigned to cluster \(j\), and write \(A=(A_1,\ldots,A_K)\). In the first stage, \(A\) is assigned by complete randomization over \[ \mathcal A = \Bigl\{ a\in\mathcal S^K: \#\{j:a_j=s_\ell\}=K_\ell,\ \ell=0,\ldots,L \Bigr\}, \] so that \[ \Pr(A=a) = \binom{K}{K_0,\ldots,K_L}^{-1}\mathbf 1\{a\in\mathcal A\}. \]

For each cluster \(j\) and saturation label \(s\in\mathcal S\), let \[ m_j(s)\in\{0,1,\ldots,n_j\} \] denote the prespecified number of treated units if \(A_j=s\), with \(m_j(0)=0\). When \(s n_j\) is an integer, a natural choice is \(m_j(s)=s n_j\); otherwise \(m_j(s)\) may be determined by any rule fixed ex ante, such as \(\lfloor s n_j\rfloor\) or \(\lceil s n_j\rceil\). When the labels are interpreted as ordered treatment intensities, we additionally require \[ q<q' \quad\Longrightarrow\quad m_j(q)\le m_j(q') \quad\text{for every cluster }j. \] Without this restriction, monotonicity in the label remains mathematically meaningful, but it need not coincide with monotonicity in the number of treated peers.

Let \(D_i\in\{0,1\}\) denote the treatment indicator for unit \(i\), and write \(D=(D_i)_{i=1}^N\). Conditional on \(A=a\), the second stage randomizes independently across clusters, selecting exactly \(m_j(a_j)\) treated units uniformly without replacement within cluster \(j\). Define \[ \mathcal D_j(s) = \left\{ d_{\mathcal I_j}\in\{0,1\}^{\mathcal I_j}: \sum_{i\in\mathcal I_j} d_i=m_j(s) \right\}. \] Then \[ \Pr(D=d\mid A=a) = \prod_{j=1}^K \binom{n_j}{m_j(a_j)}^{-1} \mathbf 1\{d_{\mathcal I_j}\in\mathcal D_j(a_j)\}. \] The full assignment is \(Z=(A,D)\), with support \[ \mathcal Z = \left\{ (a,d): a\in\mathcal A, d_{\mathcal I_j}\in\mathcal D_j(a_j)\ \text{for all }j \right\}. \] The designs studied in Basse2018,Basse2019,liutwostage arise as the special case \(\mathcal S=\{0,\pi_2\}\) for some \(\pi_2\in(0,1)\).

Potential outcomes and feasible untreated exposures

For each unit \(i\) and feasible assignment \(z\in\mathcal Z\), let \(Y_i(z)\) denote the potential outcome of unit \(i\) under assignment \(z\). The observed outcome is \[ Y_i^{\mathrm{obs}}=Y_i(Z^{\mathrm{obs}}), \qquad i=1,\ldots,N. \]

We impose the standard homogeneous partial-interference restriction for randomized saturation designs hudgens2008,Basse2018,Basse2019,Forastiere2021,Imai2021,Gonzalo-two-stage,liutwostage.

assumption[Homogeneous partial interference] For any unit \(i\) and any two feasible assignments \(z=(a,d),z'=(a',d')\in\mathcal Z\), \[ Y_i(z)=Y_i(z') \quad\text{whenever}\quad d_i=d_i' \quad\text{and}\quad \sum_{k\in\mathcal I_{[i]}}d_k = \sum_{k\in\mathcal I_{[i]}}d_k'. \]

Assumption (ref) has two components. First, it rules out interference across clusters: assignments outside unit \(i\)'s cluster do not affect \(i\)'s outcome. Second, within a cluster, the outcome depends on the assignment only through unit \(i\)'s own treatment status and the number of treated units in the cluster. Conditional on these two quantities, the identities of the treated peers are irrelevant.

Because the design fixes the number of treated units in cluster \(j\) at \(m_j(s)\) whenever \(A_j=s\), Assumption (ref) allows us to index potential outcomes by own treatment status and cluster saturation whenever that exposure is feasible. For cluster \(j\), define the feasible exposure set \[ \mathcal E_j = \{(0,s):m_j(s)<n_j\} \cup \{(1,s):m_j(s)>0\}. \] For \(i\in\mathcal I_j\), the shorthand \(Y_i(r,s)\) is used only for \((r,s)\in\mathcal E_j\). The observed outcome can therefore be written as \[ Y_i^{\mathrm{obs}} = Y_i\!\left(D_i^{\mathrm{obs}},A_{[i]}^{\mathrm{obs}}\right), \] where the realized exposure is necessarily feasible.

The spillover tests in this paper compare untreated exposures. To avoid off-support untreated potential outcomes in the main theory, define the common untreated-support set \[ \mathcal S_0 = \{s\in\mathcal S:m_j(s)<n_j\ \text{for every }j=1,\ldots,K\}. \] All pairwise untreated contrasts in Sections (ref)--(ref) use saturation levels \(s,s'\in\mathcal S_0\). The monotone null in Section (ref) is stated over ordered subsets \(\mathcal S_M\subseteq\mathcal S_0\). Thus a full-saturation label with \(m_j(s)=n_j\) for some cluster is part of the assignment support, but it is not part of an untreated-spillover null unless the null is modified to use a cluster-specific feasible domain.

Assumption (ref) is a substantive exposure-mapping assumption, not a consequence of random assignment. It combines a form of partial interference, which rules out cross-cluster effects, with a form of stratified interference, which reduces within-cluster interference to the cluster saturation. Thus, the randomization justifies the assignment probabilities used below, but the interpretation of the resulting tests depends on whether the exposure mapping captures the relevant channels of interference. If the exposure mapping is misspecified, a rejection may reflect either a violation of the stated exposure-defined null or a failure of the exposure mapping itself. This is the standard interpretational issue that arises in randomization inference under exposure mappings basse2024.

Individual-level nulls for a pair of saturation levels

Fix two distinct saturation levels \(s,s'\in\mathcal S_0\). Throughout this subsection, the target comparison is the untreated spillover contrast \[ Y_i(0,s)-Y_i(0,s'). \] The sign convention is arbitrary but fixed: a positive value means that the untreated potential outcome is larger at saturation \(s\) than at saturation \(s'\).

The first null is the individual-level no-spillover null

equation[equation omitted — 120 chars of source]

This null is partially sharp on the full assignment space. It links two untreated exposures but does not determine treated outcomes or outcomes under other saturation levels. It becomes sharp for a suitable set of untreated focal units after the conditioning step in Section (ref).

The second individual-level null is a bounded spillover null. For a prespecified constant \(\delta\in\mathbb R\), define

equation[equation omitted — 139 chars of source]

This null is useful when the researcher wants to test whether the spillover effect exceeds a substantively meaningful threshold. The equality version

equation[equation omitted — 144 chars of source]

will serve as the least favorable configuration for the finite-sample bounded-null test.

Conditional Tests for Two Saturation Levels

This section develops the finite-sample results for pairwise nulls. For a fixed pair \(s,s'\in\mathcal S_0\), the construction has two steps. First, we select untreated focal units from clusters whose observed saturation is either \(s\) or \(s'\). Second, we generate the reference distribution by relabeling the two saturation labels across those clusters while preserving the observed numbers of \(s\)- and \(s'\)-clusters. Under the partially sharp null, this conditioning makes the focal outcomes imputable. Under the bounded null, the equality boundary gives a least-favorable imputation. The weak average nulls do not share this finite-sample imputability property and are treated separately in Section (ref).

Focal units and relabeling distribution

Fix two distinct saturation levels \(s,s'\in\mathcal S_0\). The comparison uses only clusters whose observed saturation is one of these two levels. Let \[ J^{\mathrm{obs}} := \{j:A_j^{\mathrm{obs}}\in\{s,s'\}\} \] be this set of contrasted clusters. For each cluster \(j\), choose an integer \(k_j\) before observing the assignment such that

equation[equation omitted — 84 chars of source]

Because \(s,s'\in\mathcal S_0\), the right-hand side is positive for every cluster. Condition (ref) guarantees that cluster \(j\) has at least \(k_j\) untreated units available whether its saturation label is \(s\) or \(s'\). After observing the assignment, for each \(j\in J^{\mathrm{obs}}\), sample \[ U_j\subseteq \{i\in\mathcal I_j:D_i^{\mathrm{obs}}=0\}, \qquad |U_j|=k_j, \] uniformly without replacement, independently across contrasted clusters. Let \[ U=\bigcup_{j\in J^{\mathrm{obs}}}U_j \] be the focal unit set. Thus the test uses exactly \(k_j\) observed untreated focal units from each contrasted cluster and uses no units from clusters with other observed saturation labels.

The reference distribution is obtained by permuting the two saturation labels \(s\) and \(s'\) across the contrasted clusters. Let \[ A_J^{\mathrm{obs}}=(A_j^{\mathrm{obs}}:j\in J^{\mathrm{obs}}), \] and define the observed margins \[ K_s:=\sum_{j\in J^{\mathrm{obs}}}\mathbf 1\{A_j^{\mathrm{obs}}=s\}, \qquad K_{s'}:=\sum_{j\in J^{\mathrm{obs}}}\mathbf 1\{A_j^{\mathrm{obs}}=s'\}. \] The relabeling space is

equation[equation omitted — 177 chars of source]

The relabeling distribution is uniform over \(\mathcal A_J^{\mathrm{obs}}\). Equivalently, a reference draw selects exactly \(K_s\) of the contrasted clusters to carry label \(s\), assigns label \(s'\) to the remaining \(K_{s'}\) contrasted clusters, and leaves all other clusters outside the comparison. This is a cluster-level relabeling distribution: focal outcomes are not moved across units or clusters, and the unit-level treatment assignment is not re-randomized in the implementation.

For each contrasted cluster, define the observed focal-cluster mean

equation[equation omitted — 147 chars of source]

Under any relabeling \(a_J\in\mathcal A_J^{\mathrm{obs}}\), the selected focal units remain untreated, and only their cluster saturation label changes from the observed label to the relabeled value. Thus the focal outcomes are the relevant observed quantities for comparing the untreated exposures \((0,s)\) and \((0,s')\). Under the partially sharp null (ref), these focal outcomes are invariant to the relabeling. For the bounded null, they are combined with the least-favorable imputation described below.

The appendix gives the formal conditional-randomization construction under the framework of Basse2019 that justifies the uniform cluster-level relabeling distribution. The main text only uses the resulting focal-unit and relabeling objects because those are the objects needed to implement the test.

Statistics for the finite-sample nulls

Throughout, write \(J=J^{\mathrm{obs}}\), and let \[ K_q=\sum_{j\in J}\mathbf 1\{A_j^{\mathrm{obs}}=q\}, \qquad q\in\{s,s'\}. \] The relabeling space \(\mathcal A_J^{\mathrm{obs}}\) preserves these margins. We use upper-tail statistics, so large values provide evidence that untreated outcomes are larger at saturation \(s\) than at saturation \(s'\).

For the partially sharp null \(H_{0,\mathrm{PS}}^{s,s'}\), the conditioning step makes the focal outcomes imputable over \(\mathcal A_J^{\mathrm{obs}}\). Hence any focal statistic can be used. Our default choice is the difference in focal-cluster means,

equation[equation omitted — 216 chars of source]

For the bounded null \(H_{\delta,\mathrm B}^{s,s'}\), we use the equality boundary \(H_{\delta,\mathrm E}^{s,s'}\) in (ref) only to construct the least-favorable imputation. Under this equality boundary, the focal-cluster mean under saturation \(s'\) is imputed as \[ \widetilde Y_{j,\delta}^{s'} := \widetilde Y_j^{\mathrm{obs}} - \delta\,\mathbf 1\{A_j^{\mathrm{obs}}=s\}, \] and the corresponding imputed focal-cluster mean under saturation \(s\) is \[ \widetilde Y_{j,\delta}^{s} = \widetilde Y_{j,\delta}^{s'}+\delta . \] Thus, for a relabeling \(a_J\), the imputed observed focal-cluster outcome is \[ \widetilde Y_{j,\delta}(a_j) = \widetilde Y_{j,\delta}^{s'} + \delta\,\mathbf 1\{a_j=s\}. \] The literal imputed difference-in-means statistic is therefore \[ T_{U,\delta}^{\mathrm{full}}(a_J) = \frac{1}{K_s} \sum_{j\in J} \left(\widetilde Y_{j,\delta}^{s'}+\delta\right)\mathbf 1\{a_j=s\} - \frac{1}{K_{s'}} \sum_{j\in J} \widetilde Y_{j,\delta}^{s'}\mathbf 1\{a_j=s'\}. \] Because every relabeling in \(\mathcal A_J^{\mathrm{obs}}\) has exactly \(K_s\) clusters labeled \(s\), this statistic can be written as \[ T_{U,\delta}^{\mathrm{full}}(a_J) = \delta+T_{U,\delta}^{\mathrm B}(a_J), \] where

equation[equation omitted — 227 chars of source]

The additive constant \(\delta\) is common to all relabelings and therefore does not affect the randomization \(p\)-value. We can therefore compute the bounded-null test using the simpler \(s'\)-anchored statistic in (ref). At the observed relabeling, \[ T_{U,\delta}^{\mathrm B}(A_J^{\mathrm{obs}}) = \hat\tau_U(A_J^{\mathrm{obs}})-\delta, \] whereas \[ T_{U,\delta}^{\mathrm{full}}(A_J^{\mathrm{obs}}) = \hat\tau_U(A_J^{\mathrm{obs}}). \] Thus the \(s'\)-anchored formulation does not ignore saturation \(s\); it simply removes the common shift \(\delta\) from the full imputed statistic. The same construction may be applied to focal-count weighted differences in means, which are often convenient in applications with unequal focal-set sizes, provided the statistic is monotone in the same least-favorable direction.

For any statistic \(T_U(a_J)\) defined on \(\mathcal A_J^{\mathrm{obs}}\), the corresponding one-sided conditional randomization \(p\)-value is \[ p_T = \frac{1}{|\mathcal A_J^{\mathrm{obs}}|} \sum_{a_J\in\mathcal A_J^{\mathrm{obs}}} \mathbf 1\!\left\{ T_U(a_J)\ge T_U(A_J^{\mathrm{obs}}) \right\}. \] For the partially sharp and bounded nulls, we use \(T_U=\hat\tau_U\) and \(T_U=T_{U,\delta}^{\mathrm B}\), respectively.

The conditional randomization test

The following procedure summarizes the finite-sample conditional testing algorithm. The permutation step is at the cluster level: we permute the saturation labels of the contrasted clusters and keep the focal units and their observed outcomes fixed.

testprocedure[Two-saturation conditional randomization test] Fix two saturation levels \(s\neq s'\) in \(\mathcal S_0\), focal-set sizes \(\{k_j\}\) satisfying (ref), and a statistic \(T_U(a_J)\) defined on \(\mathcal A_J^{\mathrm{obs}}\). \begin{enumerate}[(1)] • Select focal controls. For each \(j\in J^{\mathrm{obs}}\), sample \(k_j\) units uniformly without replacement from \[ \{i\in\mathcal I_j:D_i^{\mathrm{obs}}=0\}. \] Let \(U_j\) be the selected set in cluster \(j\), and let \(U=\cup_{j\in J^{\mathrm{obs}}}U_j\). • Construct the relabeling space. Form \(\mathcal A_J^{\mathrm{obs}}\) as in (ref). • Evaluate the statistic. For each \(a_J\in\mathcal A_J^{\mathrm{obs}}\), compute \(T_U(a_J)\). For the partially sharp null, one may use (ref) or any other focal statistic. For the bounded null, use the shifted statistic (ref). • Compute the conditional \(p\)-value. For a one-sided alternative with large values of the statistic unfavorable to the null, set \begin{equation} p_T = \frac{1}{|\mathcal A_J^{\mathrm{obs}}|} \sum_{a_J\in\mathcal A_J^{\mathrm{obs}}} \mathbf 1\left\{ T_U(a_J)\ge T_U(A_J^{\mathrm{obs}}) \right\}. \end{equation} \end{enumerate}

The test can be implemented by enumerating \(\mathcal A_J^{\mathrm{obs}}\) or by drawing Monte Carlo relabelings uniformly from \(\mathcal A_J^{\mathrm{obs}}\). For a Monte Carlo implementation of the finite-sample tests, the usual plus-one correction can be used to preserve conservativeness of the simulated relabeling \(p\)-value.

theorem[Finite-sample conditional validity for individual-level two-saturation nulls] Fix \(s\neq s'\) in \(\mathcal S_0\), and suppose the randomized saturation design in Section (ref) is used with focal-set sizes satisfying (ref). Let \(U\) be the focal set generated in Procedure (ref). Then the following statements hold. \begin{enumerate}[(a)] • Partially sharp null. Under \(H_{0,\mathrm{PS}}^{s,s'}\) in (ref), the permutation test in Procedure (ref) is finite-sample valid for any statistic based only on the focal outcomes and relabeling vector. That is, for every \(\alpha\in[0,1]\), \[ \mathbb P\{p_T\le \alpha\mid U\}\le \alpha. \] • Bounded null. Under the bounded null \(H_{\delta,\mathrm B}^{s,s'}\) in (ref), the permutation test in Procedure (ref) with the shifted statistic \(T_{U,\delta}^{\mathrm B}\) in (ref) is finite-sample valid: \[ \mathbb P\{p_T\le \alpha\mid U\}\le \alpha \qquad \text{for every } \alpha\in[0,1]. \] The same conclusion holds for any one-sided statistic satisfying the same least-favorable monotonicity property as \(T_{U,\delta}^{\mathrm B}\). \end{enumerate}

Theorem (ref) is the finite-sample core of the paper. The focal-unit selection and the cluster-level relabeling distribution are common to the two tests, but the statistic is tailored to the null. For the partially sharp null, the null itself makes the focal outcomes invariant to relabeling, so any focal statistic is valid. For the bounded null, the equality null \(H_{\delta,\mathrm E}^{s,s'}\) is least favorable for the one-sided alternative, so imputing under this equality yields a conservative finite-sample test of the inequality null.

remark[Choice of focal-set sizes] The integers \(k_j\) must be chosen independently of the realized labels \(A_j\in\{s,s'\}\). This ensures that the probability of selecting the observed focal set is the same under every admissible relabeling. In practice, a simple choice is \[ k_j=\min\{n_j-m_j(s),\,n_j-m_j(s')\}, \] which uses as many untreated focal units as possible while preserving symmetry across the two saturation labels. Smaller choices may be useful when computation is costly or when the analyst wants identical focal-set sizes across clusters.

Weak Average Nulls as an Asymptotic Extension

The preceding section relies on unit-level restrictions that either impute the focal potential outcomes or admit a least-favorable imputation, yielding finite-sample conditional tests. We now replace these restrictions with weaker zero-average restrictions. These nulls are scientifically less restrictive but do not determine the missing focal potential outcomes. Consequently, the conditional relabeling distribution is no longer an exact finite-sample null distribution; it is instead used to calibrate a studentized statistic with unconditional asymptotic size control.

Weak average nulls for a pair of saturation levels

The weak nulls replace the unit-level restrictions above with average restrictions. For cluster \(j\), define \[ \mu_{a,j} := \frac{1}{n_j}\sum_{i\in\mathcal I_j}Y_i(0,a), \qquad a\in\{s,s'\}, \] and \[ \tau_j^{s,s'} := \mu_{s,j}-\mu_{s',j} = \frac{1}{n_j}\sum_{i\in\mathcal I_j}\{Y_i(0,s)-Y_i(0,s')\}. \] Let \(\lambda_{1K},\ldots,\lambda_{KK}\) be nonnegative, nonrandom weights chosen by the researcher, with \[ \sum_{j=1}^K\lambda_{jK}=1. \] The weighted weak null is

equation[equation omitted — 158 chars of source]

Two special cases are useful. If \(\lambda_{jK}=1/K\), then (ref) is the equally weighted cluster-average weak null,

equation[equation omitted — 146 chars of source]

If \(\lambda_{jK}=n_j/N\), then (ref) is the unit-average weak null,

equation[equation omitted — 209 chars of source]

The two targets coincide when cluster sizes are constant.

Unlike the partially sharp and bounded nulls, the weighted weak null does not determine the missing focal potential outcomes. The relabeling distribution is therefore not an exact finite-sample null distribution in general. The remainder of this section treats weak-null inference as an asymptotic extension based on studentized relabeling statistics.

Studentized relabeling statistic and asymptotic validity

Fix \(s\neq s'\) in \(\mathcal S_0\), and use the focal set and relabeling space from Section (ref). Write \(J=J^{\mathrm{obs}}\). For the weighted weak null \(H_{0,\mathrm W}^{s,s'}(\lambda)\), define \[ w_{jK}:=K\lambda_{jK}, \qquad X_j^{\mathrm{obs}}:=w_{jK}\widetilde Y_j^{\mathrm{obs}}. \] The normalization \(w_{jK}=K\lambda_{jK}\) is convenient because \(K^{-1}\sum_{j=1}^K w_{jK}=1\).

For \(q\in\{s,s'\}\) and \(a_J\in\mathcal A_J^{\mathrm{obs}}\), define \[ \bar X_U(q;a_J) = \frac{1}{K_q} \sum_{j\in J} X_j^{\mathrm{obs}}\mathbf 1\{a_j=q\}, \] and, when \(K_q\ge 2\), \[ \hat S_{X,U}^2(q;a_J) = \frac{1}{K_q-1} \sum_{j\in J} \left\{X_j^{\mathrm{obs}}-\bar X_U(q;a_J)\right\}^2 \mathbf 1\{a_j=q\}. \] The Neyman variance estimator for the transformed focal-cluster outcomes is \[ \hat V_{X,U}^{\mathrm{Ney}}(a_J) = \frac{\hat S_{X,U}^2(s;a_J)}{K_s} + \frac{\hat S_{X,U}^2(s';a_J)}{K_{s'}}. \] The studentized statistic is

equation[equation omitted — 161 chars of source]

When the denominator is zero, we use a fixed deterministic convention, such as setting the statistic to zero. Assumption (ref) in Appendix (ref) rules out this degeneracy in the limit. The corresponding upper-tail relabeling \(p\)-value is defined by (ref) with \(T_U=T_{U,\lambda}^{\mathrm{Ney}}\); write this \(p\)-value as \(p_{T,\lambda}\).

The equally weighted cluster-average weak null corresponds to \(w_{jK}=1\), in which case \(T_{U,\lambda}^{\mathrm{Ney}}\) reduces to the unweighted studentized statistic. The unit-average weak null corresponds to \[ w_{jK}=\frac{n_j}{N/K}. \] Since multiplication of all transformed cluster outcomes by a common positive constant cancels after studentization, the unit-average version can equivalently be implemented by replacing each focal-cluster mean by \(n_j\widetilde Y_j^{\mathrm{obs}}\).

theorem[Unconditional asymptotic validity for weighted weak average nulls] Fix \(s\neq s'\) in \(\mathcal S_0\), and suppose the randomized saturation design in Section (ref) is used with focal-set sizes satisfying (ref). Let \(\lambda_{1K},\ldots,\lambda_{KK}\) be nonnegative, nonrandom weights summing to one, and let \(p_{T,\lambda}\) be the relabeling \(p\)-value computed from \(T_{U,\lambda}^{\mathrm{Ney}}\) in (ref). If the weighted weak null \(H_{0,\mathrm W}^{s,s'}(\lambda)\) in (ref) is true and certain regularity conditions defined in Assumption (ref) hold, then \[ \limsup_{K\to\infty} \mathbb P\{p_{T,\lambda}\le \alpha\} \le \alpha \qquad \text{for every } \alpha\in(0,1/2). \]

Theorem (ref) uses the same focal set and relabeling space as the finite-sample tests, but its validity statement is different. The randomization distribution used for calibration is conditional on the realized contrasted clusters and focal outcomes, whereas the size guarantee is asymptotic and unconditional over the contrasted cluster set, focal sampling, and saturation assignment. The distinction matters because the population weak null \(\sum_j\lambda_{jK}\tau_j^{s,s'}=0\) does not imply that the realized contrasted clusters have exactly zero weighted average spillover, nor does it determine the missing focal potential outcomes.

Multiple Saturation Levels and Monotone Nulls

The preceding sections study a fixed pair of saturation levels. This is the right object when the research question concerns a particular contrast. In other applications, the substantive question concerns the shape of the spillover response across several saturation levels. For example, a researcher may want to test whether untreated outcomes are monotone in treatment saturation over a specified range of saturation levels. This section develops a finite-sample valid unconditional randomization test for such monotone spillover nulls following zhong2024unconditional.

Monotone nulls over an ordered set of saturation levels

Let \[ \mathcal S_M=\{q_0,q_1,\ldots,q_M\}\subseteq \mathcal S_0, \qquad q_0<q_1<\cdots<q_M, \] be the ordered set of feasible untreated saturation levels over which the researcher wants to test monotonicity. The case \(\mathcal S_M=\mathcal S_0\) corresponds to the global monotone null over the full untreated-support set. Allowing \(\mathcal S_M\) to be a strict subset is useful when the substantive question concerns monotonicity only over a particular range of saturation levels.

We focus on the monotone increasing spillover null for untreated units,

equation[equation omitted — 161 chars of source]

The direction can be reversed if the substantive theory predicts that higher saturation weakly lowers untreated outcomes. Since the experiment randomizes over a finite set of saturation labels, (ref) is the operational design-based monotonicity restriction. The null is unitwise: it is stronger than monotonicity of the average spillover response.

One approach would be to test the adjacent inequalities \[ Y_i(0,q_\ell)\le Y_i(0,q_{\ell+1}), \qquad \ell=0,\ldots,M-1, \] and then combine the resulting \(p\)-values. This approach is valid if the multiple-testing step is handled appropriately, but it treats monotonicity as a collection of local statements. We instead construct a single unconditional randomization test for the joint null in (ref).

Pairwise-imputable monotone statistics

The test is based on pairwise comparisons between two assignments. For any assignments \(z=(a,d)\) and \(z'=(a',d')\) in \(\mathcal Z\), define

equation[equation omitted — 155 chars of source]

These are the units that are untreated under both assignments and whose two saturation exposures are both covered by the monotone null. If \(\mathcal S_M=\mathcal S_0\), this set contains all doubly untreated units whose two untreated exposures are feasible. If \(\mathcal S_M\subsetneq\mathcal S_0\), units whose pairwise exposures fall outside the tested saturation range are excluded because the null imposes no restriction on their pairwise outcome ordering.

definition[Pairwise-imputable spillover-monotone statistic] A statistic \[ T:\mathbb R^N\times\mathcal Z\times\mathcal Z\to\mathbb R\cup\{\infty\} \] is a pairwise-imputable spillover-monotone statistic for \(H_{0,\mathrm M}(\mathcal S_M)\) if, for every pair \(z=(a,d)\) and \(z'=(a',d')\), the following two conditions hold. \begin{enumerate}[(i)] • Pairwise imputability. If two outcome vectors \(y\) and \(y'\) agree on \(\mathbb I_M(z,z')\), then \[ T(y,z,z')=T(y',z,z'). \] • Monotonicity. If, for every \(i\in\mathbb I_M(z,z')\), \[ \begin{cases} y_i\ge y_i', & \text{when } a_{[i]}>a'_{[i]},\\ y_i\le y_i', & \text{when } a_{[i]}<a'_{[i]},\\ y_i= y_i', & \text{when } a_{[i]}=a'_{[i]}, \end{cases} \] then \[ T(y,z,z')\ge T(y',z,z'). \] \end{enumerate}

Condition (i) says that the statistic uses only units whose two untreated exposures are covered by the null. Condition (ii) says that the statistic respects the direction of the monotone restriction. When the saturation exposure is higher under \(z\) than under \(z'\), the corresponding outcome is weakly higher; when the saturation exposure is lower under \(z\), the corresponding outcome is weakly lower; and when the exposure is unchanged, the outcome is unchanged. The last case is natural because equal exposure labels correspond to the same untreated exposure. For a monotone decreasing null, the inequalities in Condition (ii) are reversed.

A simple example is a transformed difference in means. Let \(\psi_1,\psi_0:\mathbb R\to\mathbb R\) be nondecreasing functions.

equation[equation omitted — 362 chars of source]

with the convention \(0/0=0\). When \(\psi_1(u)=\psi_0(u)=u\), this statistic compares outcomes for units whose saturation is weakly higher under \(z\) than under \(z'\) with outcomes for units whose saturation is lower under \(z\) than under \(z'\). One may alternatively omit unchanged-exposure units by replacing \(H(z,z')\) with \(\{i\in\mathbb I_M(z,z'):a_{[i]}>a'_{[i]}\}\).

A rank-based statistic can be defined similarly. Let \(r_i(y_{\mathbb I_M(z,z')})\) be the rank of \(y_i\) among \(\{y_\ell:\ell\in\mathbb I_M(z,z')\}\), with ties handled by average ranks. For a nondecreasing score function \(\varphi\), define

equation[equation omitted — 178 chars of source]

The transformed difference-in-means statistic satisfies Definition (ref) directly by the monotonicity of \(\psi_1\) and \(\psi_0\). The rank statistic also satisfies the definition: increasing outcomes in the weakly higher group and decreasing outcomes in the lower group weakly increases the ranks, and hence the rank scores, assigned to the weakly higher group. Thus both \(T^{\mathrm{dm}}\) and \(T^{\mathrm{rk}}\) are pairwise-imputable spillover-monotone statistics.

The definition implies the following pairwise ordering.

proposition[Pairwise ordering under monotonicity] Suppose Assumption (ref) holds and the monotone null \(H_{0,\mathrm M}(\mathcal S_M)\) in (ref) is true. If \(T\) is a pairwise-imputable spillover-monotone statistic, then, for every \(z,z'\in\mathcal Z\), \[ T\{Y(z),z,z'\} \ge T\{Y(z'),z,z'\}. \]

Unconditional PIRT for the monotone null

Proposition (ref) leads to an unconditional randomization test following zhong2024unconditional. For the realized assignment \(Z^{\mathrm{obs}}\), define

equation[equation omitted — 216 chars of source]

Every term in (ref) is observable because the statistic is pairwise imputable: for each reference assignment \(z\), both statistics depend only on units in \(\mathbb I_M(Z^{\mathrm{obs}},z)\). When \(|\mathcal Z|\) is too large for enumeration, the sum in (ref) can be approximated by Monte Carlo draws from the known design distribution \(P\). Only exact enumeration, or a conservative finite-Monte-Carlo implementation, should be described as finite-sample valid without numerical qualification.

testprocedure[PIRT for the monotone spillover null] Fix an ordered saturation set \(\mathcal S_M\subseteq\mathcal S_0\) and a pairwise-imputable spillover-monotone statistic \(T\). \begin{enumerate}[(1)] • Draw or enumerate reference assignments \(z\) from the design distribution \(P\). • For each reference assignment, compute \[ \Gamma(z;Z^{\mathrm{obs}},Y^{\mathrm{obs}}) = \mathbf 1\!\bigl[ T\{Y^{\mathrm{obs}},Z^{\mathrm{obs}},z\} \ge T\{Y^{\mathrm{obs}},z,Z^{\mathrm{obs}}\} \bigr]. \] • Compute \[ p_M^{\mathrm{PIRT}}(Z^{\mathrm{obs}}) = \sum_{z\in\mathcal Z} \Gamma(z;Z^{\mathrm{obs}},Y^{\mathrm{obs}})P(z), \] or a Monte Carlo analogue. • Reject \(H_{0,\mathrm M}(\mathcal S_M)\) at level \(\alpha\) when \[ p_M^{\mathrm{PIRT}}(Z^{\mathrm{obs}})\le \alpha/2. \] Equivalently, \(\min\{2p_M^{\mathrm{PIRT}},1\}\) is the reported finite-sample valid \(p\)-value under exact enumeration. \end{enumerate}
theorem[Finite-sample validity of monotone PIRT] Suppose Assumption (ref) holds and the monotone null \(H_{0,\mathrm M}(\mathcal S_M)\) in (ref) is true. If \(T\) is a pairwise-imputable spillover-monotone statistic, then the rejection rule in Procedure (ref) satisfies \[ \operatorname*{E}_P\! \left( \mathbf 1\{p_M^{\mathrm{PIRT}}(Z^{\mathrm{obs}})\le \alpha/2\} \right) \le \alpha \qquad \text{for every } \alpha\in(0,1). \] Thus Procedure (ref) is a finite-sample valid unconditional randomization test of the monotone spillover null over \(\mathcal S_M\).

The factor \(\alpha/2\) is the price of converting the pairwise ordering in Proposition (ref) into an unconditional finite-sample test. The resulting test is conservative in general. This conservativeness should be weighed against the benefit of testing the monotone null directly over the chosen saturation set, without decomposing it into adjacent contrasts and combining multiple \(p\)-values.

remark[Full-support and subset monotonicity] When \(\mathcal S_M=\mathcal S_0\), the null in (ref) rules out nonmonotonicity over the full feasible untreated-support set used by the experiment. When \(\mathcal S_M\subsetneq\mathcal S_0\), the null is weaker and concerns only the specified saturation range. The PIRT remains valid because the pairwise imputable set in (ref) excludes units whose two exposures are not both covered by the null. The choice of \(\mathcal S_M\) is therefore part of the null hypothesis: a smaller set can sharpen the substantive target but may reduce the number of units contributing to each pairwise comparison.
remark[General ordered-exposure nulls] The PIRT construction does not rely on any special feature of the randomized saturation design beyond the exposure mapping and the known assignment distribution. More generally, suppose an experiment has assignment space \(\mathcal Z\), assignment law \(P\), and exposure mapping \(e_i:\mathcal Z\to\mathcal E\). Let \(\mathcal E_0\subseteq\mathcal E\) be the subset of exposures on which the null is stated, and let \(\preceq\) be a partial order on \(\mathcal E_0\). Consider the ordered-exposure null \[ Y_i(e)\le Y_i(e') \quad\text{for all }i\text{ and all }e,e'\in\mathcal E_0\text{ with }e\preceq e'. \] For a pair \(z,z'\), the pairwise imputable set is \[ \mathbb I_{\preceq}(z,z') = \left\{ i\in[N]: e_i(z),e_i(z')\in\mathcal E_0 \text{ and } \bigl(e_i(z)\preceq e_i(z')\text{ or }e_i(z')\preceq e_i(z)\bigr) \right\}. \] Units whose two exposures are not both covered by the null, or whose exposures are incomparable, are excluded from the pairwise comparison. Any statistic satisfying the analogues of Definition (ref) with respect to \(\mathbb I_{\preceq}(z,z')\) yields the same PIRT validity argument. The saturation-design null in (ref) is the special case \(e_i(z)=(D_i,A_{[i]})\), \( \mathcal E_0=\{(0,q):q\in\mathcal S_M\}, \) and \( (0,q)\preceq (0,q') \Longleftrightarrow q\le q'. \)

Empirical Application: Spillovers in the Zomba Cash Transfer Experiment

We conclude with an application to the Zomba Cash Transfer Program in Malawi. The experiment is useful for illustration because it combines cluster-level variation in treatment intensity with individual-level treatment assignment. baird2011cash and baird2018optimal use this design to study direct and spillover effects of cash transfers. Our purpose is narrower: we use the public randomized saturation structure to illustrate how the finite-sample randomization tests developed above can be implemented in a realistic multi-saturation experiment.

Design, exposure labels, and hypotheses

The original experiment assigned enumeration areas (EAs) to cash-transfer intervention status and then varied schoolgirl offers within treatment EAs. We focus on the unconditional cash-transfer (UCT) side of the schoolgirl intervention, together with treatment EAs in which no baseline schoolgirls were offered transfers. This UCT-plus-zero sample keeps the analysis within one schoolgirl intervention arm while retaining a zero schoolgirl-offer reference cell. Table (ref) summarizes the resulting support. The \(100\%\) cell remains part of the assignment support but contributes no untreated focal schoolgirls to the spillover contrasts.

table[table omitted — 732 chars of source]

The randomization tests condition on the observed set of 42 EAs in Table (ref) and on the fixed compound-cell margins \(15,9,9,9\). Under the published complete-randomization description, the schoolgirl-offer labels $\{\ell_0, \ell_{0.33}, \ell_{0.66}, \ell_1\}$ are completely randomized across these EAs subject to those margins. The local CRTs further condition on the selected untreated focal sets and relabel the relevant EA-level policy labels within the corresponding fixed margins. Thus the empirical relabelings should be read as relabelings of compound schoolgirl-offer policies, not as separate randomizations of the background dropout intervention or transfer amounts.

The exposure notation below makes this compound-label interpretation explicit. Let \(\ell_0\) denote the no-schoolgirl-offer treatment-EA label, and let \(\ell_s\) denote the UCT schoolgirl-offer label at saturation \(s\in\{0.33,0.66,1\}\). We write \(Y_i(d,\ell_s)\) for the potential outcome of baseline schoolgirl \(i\) under own schoolgirl offer status \(d\) and compound schoolgirl-offer label \(\ell_s\). This is a reduced exposure mapping: spillovers are assumed to depend on the randomized schoolgirl-offer policy label, with other components absorbed into or conditioned on by that label.

For each outcome, let \(\mathcal I_{\mathrm{app}}\) denote the baseline schoolgirls in the UCT-plus-zero EAs with that Round 3 outcome observed in the public analysis file.\footnote{The application is therefore a design-based analysis of the outcome-specific public-data finite population. The public files contain the assignment labels, EA identifiers, schoolgirl offer indicators, and Round 3 outcomes needed for this illustration; components used for other purposes in the published studies are not required for the randomization tests reported here.} For \(s>s'\), we test the unit-level bounded null \[ H_{0,B}^{s,s'}:\qquad Y_i(0,\ell_s)\le Y_i(0,\ell_{s'}) \quad\text{for every }i\in\mathcal I_{\mathrm{app}}. \] The three pairwise contrasts are \((s,s')=(0.33,0)\), \((0.66,0.33)\), and \((0.66,0)\). On the raw outcome scale, a small pairwise \(p\)-value is evidence that the higher compound label raises the untreated outcome for at least one schoolgirl in the finite population. We also test the increasing monotone null \[ Y_i(0,\ell_0)\le Y_i(0,\ell_{0.33})\le Y_i(0,\ell_{0.66}) \quad\text{for every }i\in\mathcal I_{\mathrm{app}}. \] For this null, a small monotone \(p\)-value is evidence against weakly increasing raw outcomes over the ordered labels. The four outcomes are current enrollment, English literacy, ever married, and ever pregnant.

The pairwise bounded-null tests use a focal-count weighted version of the least-favorable difference-in-means statistic from Section (ref). This statistic has the same least-favorable monotonicity property as \(T_{U,\delta}^{\mathrm B}\), so it is covered by Theorem (ref). For the global monotone PIRT, we use the strict changed-exposure variant of the transformed difference-in-means statistic in (ref), applied to \(-Y_i\) rather than \(Y_i\). This sign reversal makes large values correspond to violations in which a higher ordered label lowers the raw outcome. In the application code, both statistics are computed as focal-count weighted cluster-mean coefficients. This regression-style implementation is only a convenient way to compute the relevant mean contrasts; the validity of the tests comes from their respective randomization reference distributions.

Calibrated simulations

We use calibrated simulations to benchmark the candidate procedures in a design close to the application. The simulations are semi-synthetic: they fix the observed EA sizes, saturation-cell counts, and outcome-specific untreated baseline risks, and vary only the untreated spillover response surface. For outcome \(k\) in EA \(j\), the baseline risk is the smoothed untreated mean \[ \widehat p_{jk} = \frac{\sum_{i:j(i)=j,D_i=0}Y_i^k+\lambda\bar Y_0^k}{n_{j0}+\lambda}, \qquad \lambda=4, \] where \(\bar Y_0^k\) is the overall untreated mean and \(n_{j0}\) is the number of untreated schoolgirls in EA \(j\). We add an assignment-independent heavy-tailed EA component \[ \eta_j = 0.15\sqrt{n_j/\operatorname{median}(n_j)}\,t_{2j}/\sqrt{2}, \] where \(t_{2j}\) are independent Student-\(t\) draws with two degrees of freedom. The simulated untreated potential outcomes are generated as \[ Y_i^k(0,a;\tau) = \mathbf 1\!\left[ U_i^0 \le \operatorname{clip}_{[\epsilon,1-\epsilon]} \left\{ \widehat p_{j(i)k}+\eta_{j(i)} +\sigma\tau\frac{\min(a,0.66)}{0.66} \right\} \right], \qquad U_i^0\sim \operatorname{Unif}(0,1), \] with \(\epsilon=0.001\), \(\tau\in\{0,0.10,0.20,0.30,0.40,0.50\}\), and \(\sigma=1\) for the pairwise bounded-null simulations. For the monotone-null simulations, we set \(\sigma=-1\), so positive \(\tau\) generates violations of weakly increasing untreated outcomes. The smoothing, clipping, and latent EA components are used only to construct the semi-synthetic designs; the randomization tests applied to the observed data use the observed outcomes and the known assignment mechanism.

Figure (ref) compares three procedures for the pairwise bounded spillover nulls: a unit-level linear probability model with EA-clustered CR1 standard errors, the focal-set conditional CRT, and PIRT. The regression benchmark has the largest rejection rates under alternatives but is oversized under the zero-spillover design, with average size \(0.078\) and maximum size \(0.120\) across outcome-contrast cells. The CRT is much closer to nominal size, with average size \(0.051\) and maximum size \(0.077\). PIRT is more conservative, with average size \(0.009\) and maximum size \(0.020\). Over the positive effect grid, the average rejection rates are \(0.397\), \(0.315\), and \(0.180\) for the regression benchmark, CRT, and PIRT, respectively. We therefore emphasize the conditional CRT for the pairwise bounded-null application.

figure[figure omitted — 360 chars of source]

Figure (ref) compares the global monotone PIRT with adjacent CRTs combined by Bonferroni. In the empirical three-level support \(\{0,0.33,0.66\}\), adjacent CRT plus Bonferroni is stronger near the null: at \(\tau=0.10\) and \(\tau=0.20\), its rejection rates are \(0.163\) and \(0.400\), compared with \(0.090\) and \(0.367\) for PIRT. PIRT overtakes the adjacent CRT procedure for larger violations. In a four-level stress support \(\{0,0.22,0.44,0.66\}\), which keeps the same calibrated outcome structure but uses a richer ordered support, PIRT has higher average power over the positive effect grid: \(0.559\), compared with \(0.403\) for adjacent CRT plus Bonferroni.

These simulations describe operating characteristics; they do not determine the validity of the observed-data tests. For the observed pairwise bounded-null application, we emphasize the focal-set conditional CRT. For the observed monotone application, we report both the global monotone PIRT and the adjacent CRT-Bonferroni procedure. The latter uses the same local CRT building blocks as the pairwise analysis, but it tests a different directional null: small pairwise \(p\)-values support increases in raw outcomes at higher labels, whereas small monotone \(p\)-values reject weakly increasing raw outcomes.

figure[figure omitted — 355 chars of source]

Application to observed outcomes

Table (ref) reports the observed-data randomization \(p\)-values. The first three columns report one-sided focal-set CRT \(p\)-values for the pairwise bounded nulls. The last two columns report valid \(p\)-values for the raw-scale monotone null over \(\{\ell_0,\ell_{0.33},\ell_{0.66}\}\): the adjusted monotone PIRT value \(\min\{2p_M^{\mathrm{PIRT}},1\}\) and the adjacent-CRT Bonferroni value.

table[table omitted — 1,385 chars of source]

None of the reported randomization tests rejects at the \(5\%\) level. The smallest pairwise \(p\)-value is \(0.064\), for the \(\ell_{0.33}\) versus \(\ell_0\) contrast on ever married. The smallest monotone \(p\)-value is \(0.102\), from the adjacent-CRT procedure for current enrollment. Thus the pairwise CRTs do not reject the unit-level non-increase bounds in favor of higher untreated outcomes at higher compound labels. The monotone PIRT and adjacent-CRT Bonferroni procedures also do not reject the weakly increasing raw-outcome monotonicity null. These nonrejections should not be read as two-sided evidence of no spillover effects; they are evidence only with respect to the specified one-sided bounded and monotone nulls. The exercise illustrates the practical value of exact randomization procedures in a design with few clusters per saturation cell, heterogeneous cluster sizes, and partially sharp spillover nulls.

Conclusion

This paper develops a randomization-based toolkit for randomized saturation designs. The main message is that the conditioning step and the statistic should be separated. For a fixed pair of saturation levels, conditioning on untreated focal units turns the spillover comparison into a cluster-level relabeling problem. That relabeling distribution yields finite-sample validity for partially sharp nulls, asymptotic validity for weak nulls when paired with a studentized statistic, and finite-sample validity for bounded nulls when paired with a least favorable shifted statistic. For multiple saturation levels, a pairwise-imputation approach yields a finite-sample valid unconditional test of the global monotone null.

Several extensions are natural. First, the same ideas can be applied to total-effect nulls and to contrasts involving treated units, provided the focal set is chosen to preserve imputability. Second, the weak-null theory can be extended to alternative weighting schemes and regression-adjusted focal-cluster statistics. Third, the monotone PIRT framework can be adapted to more general exposure mappings, including network exposure mappings and continuous dose-response designs. These extensions reinforce the main point: randomization inference under interference is most transparent when the null hypothesis, the imputable units, and the statistic are designed together.

\spacingset{1}

\spacingset{1.5}