EconBase
← Back to paper

Randomization Tests in Randomized Saturation Designs

The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.

64,176 characters

Randomization Tests in Randomized Saturation Designs



\def\spacingset#1{\small\normalsize}
\spacingset{1}

\title{Randomization Tests in Randomized Saturation Designs}

\author{
Jizhou Liu \\
PHBS Business School\\
Peking University\\
\url{[email removed]}
\and
Azeem M.\ Shaikh \\
Department of Economics\\
University of Chicago\\
\url{[email removed]}
\and
Liang Zhong \\
Faculty of Business and Economics\\
The University of Hong Kong\\
\url{[email removed]}
}

\maketitle

\onehalfspacing

\begin{abstract}
Randomized saturation designs are widely used to study spillover effects in clustered populations.
In these designs, clusters are first assigned to treatment saturation levels, and units are then randomized within clusters according to the assigned saturation.
This paper develops randomization tests for such experiments under several null hypotheses that arise naturally in spillover analysis.
For a fixed pair of saturation levels, we first study two individual-level hypotheses: a partially sharp null of no spillover effect for every untreated unit and a bounded null that restricts individual spillover effects by a prespecified constant.
Both hypotheses can be tested using a common conditional randomization framework, with finite-sample validity obtained by combining the same focal-unit relabeling distribution with null-specific statistics.
We then study weak average-spillover nulls and show that, although these nulls do not yield finite-sample exact conditional tests, studentized relabeling statistics deliver asymptotically valid randomization-based inference.
Finally, for multiple ordered saturation levels, we develop a finite-sample valid unconditional pairwise-imputation test for global monotonicity of spillover effects.
Simulations and an application to the Zomba Cash Transfer experiment illustrate the finite-sample behavior and practical implementation of the methods.
\end{abstract}

\noindent \textbf{Keywords:} causal inference; conditional randomization test; interference; randomized saturation design; spillover effects; studentization.

\hypersetup{pageanchor=false}
\thispagestyle{empty}
\newpage
\hypersetup{pageanchor=true}
\setcounter{page}{1}

\spacingset{1.7}


\section{Introduction}

Randomized saturation designs are a central tool for studying spillovers in clustered populations.
In a typical design, clusters are first randomized to different treatment saturation levels, and then units within each cluster are randomized to treatment subject to the assigned saturation.
The first stage creates variation in the intensity of exposure to treated peers, while the second stage separates own treatment from peer treatment intensity.
Such designs have been used in a wide range of applications and have generated an active literature on estimation and inference under partial interference \citep{hudgens2008,toulis2013,Basse2018,Basse2019,Imai2021,liutwostage}.

Despite this flexibility, valid inference in randomized saturation designs remains challenging for two related reasons.
First, the number of clusters is often too small to justify large-sample approximations, and cluster sizes may be highly heterogeneous (see Table~\ref{tab:empirical}).
Fisherian randomization tests \citep{Fisher1935Design} are attractive in this setting because they rely only on the known assignment mechanism and can deliver exact finite-sample \(p\)-values under sharp null hypotheses \citep[see][]{imbens2015causal}.
Second, however, many hypotheses of substantive interest concern \emph{spillover effects} or \emph{total effects}, and these hypotheses are typically only partially sharp on the full assignment space under interference.
They restrict selected exposure conditions and therefore do not impute the full schedule of missing potential outcomes \citep{zhong2024unconditional}.
As a result, a naive randomization test that resamples the full assignment vector need not generate a valid reference distribution.
This motivates conditioning on \emph{focal units} and restricting the reference distribution to assignments for which the relevant potential outcomes are imputable \citep{Athey2018,Basse2019}.

This paper develops randomization tests for randomized saturation designs under several spillover null hypotheses.
For a fixed pair of saturation levels \(s\) and \(s'\), we first study two individual-level nulls.
The first is a partially sharp null requiring \(Y_i(0,s)=Y_i(0,s')\) for every unit \(i\), where \(Y_i(0,s)\) denotes unit \(i\)'s potential outcome under control treatment status \(0\) and cluster saturation level \(s\).
The second is a bounded null requiring \(Y_i(0,s)-Y_i(0,s')\le \delta\) for every unit \(i\), where \(\delta\) is a researcher-chosen constant.
For these two nulls, conditioning on untreated focal units in clusters whose realized saturation is either \(s\) or \(s'\) converts the problem into a cluster-level relabeling problem.
Under the partially sharp null, the focal outcomes are invariant across relabelings.
Under the bounded null, the equality boundary provides a least-favorable imputation for the one-sided alternative.
Both tests therefore have finite-sample conditional validity.

We then consider weak average-spillover nulls.
These nulls do not require every individual spillover effect to satisfy a unit-level restriction; instead, they set a prespecified weighted average of spillover effects equal to zero.
Leading examples include the unit-average spillover effect and the equally weighted cluster-average spillover effect.
Such weak nulls are scientifically less restrictive than the individual-level equality and bounded nulls, but they do not impute the missing focal potential outcomes.
Consequently, the conditional relabeling distribution is not an exact finite-sample null distribution.
We show that, when paired with an appropriately studentized cluster-level statistic, the same relabeling device yields an asymptotically valid randomization-based test under regularity conditions for the corresponding weighted cluster-level array.


We also study hypotheses involving more than two saturation levels.
In many applications the object of interest is not a single pairwise contrast, but the entire shape of the spillover response as the saturation level changes.
For example, researchers may want to test whether untreated outcomes are monotone in the saturation level, or whether the data reveal nonmonotonicity.
For this global monotone null, pairwise conditional tests would require a multiple-testing or intersection procedure.
We instead develop an unconditional randomization test based on pairwise imputation \citep{zhong2024unconditional}.
The test compares the observed assignment with each reassignment in the design support, using only units that are untreated under both assignments and whose exposures lie in the ordered set covered by the null.
This construction yields finite-sample size control for the global monotone null and extends naturally to ordered-exposure nulls beyond the randomized saturation design.

The paper makes three contributions.
First, it develops finite-sample conditional randomization tests for partially sharp and bounded spillover nulls in multi-saturation two-stage designs.
Existing work on randomized saturation designs has primarily studied estimation and large-sample inference for direct and spillover effects \citep{hudgens2008,Basse2018,Imai2021,Gonzalo-two-stage,liutwostage}, while the conditional randomization-test literature has emphasized nonsharp nulls under interference more generally \citep{Aronow2012,Athey2018,Basse2019,puelz2021,basse2024,liu2026randomizationinferencetwosidedmarket,liu2026randomizationtestsswitchbackexperiments}.
Closest to our setting, \citet{Basse2019} study a two-stage design with a binary first stage and exactly one treated unit in each treated cluster.
We extend this line of work to multi-saturation designs with multiple treated units per cluster, which are common in empirical applications but create a richer assignment space and a more complicated partial-imputation problem.

Second, the paper contributes to inference for weak null hypotheses on average spillover effects.
Weak nulls restrict only average treatment or spillover effects, so exact finite-sample Fisherian validity is generally unavailable without additional structure \citep{chung2013,diciccio2017robust,ZHAO2021278,Wu2021,toulis2025asymptotic}.
In randomized saturation designs, weak spillover nulls are especially natural because researchers often want to test whether average outcomes for untreated units differ across saturation levels.
We show that the conditional relabeling construction can be paired with studentized cluster-level statistics to obtain asymptotically valid tests for such average spillover nulls. This connects
the Fisherian randomization-test perspective with Neymanian large-sample inference
in saturation designs.

Third, the paper extends randomization inference in randomized saturation designs beyond equality nulls.
Recent work has emphasized that scientifically meaningful null hypotheses need not require exact equality of potential outcomes, but may instead impose bounded, ordered, or monotonic restrictions \citep{Caughey2023,huang2025randomization}.
For bounded spillover nulls, we show how conditional randomization tests can be adapted to allow nonzero but bounded spillover effects.
For monotone spillover nulls across ordered saturation levels, we develop a randomization test that targets the global monotonicity restriction directly.
Relative to approaches based on local or pairwise ordered-exposure comparisons, our procedure tests the global null in a single randomization-based framework and applies beyond randomized saturation designs to more general ordered-exposure settings.

The rest of the paper is organized as follows.
Section~\ref{sec:setup} introduces the randomized saturation design, feasible untreated exposures, and the null hypotheses.
Section~\ref{sec:two-level-crt} develops finite-sample conditional tests for the partially sharp and bounded two-saturation nulls.
Section~\ref{sec:weak-null} presents the weak average null as an asymptotic extension.
Section~\ref{sec:monotone} studies the multiple-saturation monotone null and presents the pairwise-imputation test.
Section~\ref{sec:application_malawi} gives the empirical illustration.
Section~\ref{sec:conclusion} concludes.
The appendices contain the formal conditioning construction, proofs, and implementation details.

\section{Setup and Notation}\label{sec:setup}

We work under a finite-population framework, following \citet{Basse2018,Basse2019}.
There are \(K\) clusters indexed by \(j=1,\ldots,K\).
Cluster \(j\) contains the unit set \(\mathcal I_j\), with \(|\mathcal I_j|=n_j\), and the total number of units is \(N=\sum_{j=1}^K n_j\).
Let \([i]\) denote the cluster containing unit \(i\).
Potential outcomes are treated as fixed throughout; randomness comes only from the known assignment mechanism.

\subsection{Randomized saturation design}
\label{subsec:design}

The assignment proceeds in two stages: clusters are first assigned to treatment saturation labels, and units are then randomized within clusters conditional on the assigned label.
Let
\[
\mathcal S=\{s_0,s_1,\ldots,s_L\}\subset[0,1]
\]
be the set of cluster-level saturation labels, where \(s_0=0\) denotes zero treatment saturation.
The label \(s_0\) is a pure control label only when no other intervention components are present.
Fix integers \(K_0,\ldots,K_L\) satisfying
\[
\sum_{\ell=0}^L K_\ell=K.
\]
These counts are fixed by design, with \(K_\ell\) denoting the number of clusters assigned to saturation label \(s_\ell\).

Let \(A_j\in\mathcal S\) denote the saturation assigned to cluster \(j\), and write \(A=(A_1,\ldots,A_K)\).
In the first stage, \(A\) is assigned by complete randomization over
\[
\mathcal A
=
\Bigl\{
 a\in\mathcal S^K:
 \#\{j:a_j=s_\ell\}=K_\ell,\ \ell=0,\ldots,L
\Bigr\},
\]
so that
\[
\Pr(A=a)
=
\binom{K}{K_0,\ldots,K_L}^{-1}\mathbf 1\{a\in\mathcal A\}.
\]

For each cluster \(j\) and saturation label \(s\in\mathcal S\), let
\[
m_j(s)\in\{0,1,\ldots,n_j\}
\]
denote the prespecified number of treated units if \(A_j=s\), with \(m_j(0)=0\).
When \(s n_j\) is an integer, a natural choice is \(m_j(s)=s n_j\); otherwise \(m_j(s)\) may be determined by any rule fixed ex ante, such as \(\lfloor s n_j\rfloor\) or \(\lceil s n_j\rceil\).
When the labels are interpreted as ordered treatment intensities, we additionally require
\[
q<q' \quad\Longrightarrow\quad m_j(q)\le m_j(q')
\quad\text{for every cluster }j.
\]
Without this restriction, monotonicity in the label remains mathematically meaningful, but it need not coincide with monotonicity in the number of treated peers.

Let \(D_i\in\{0,1\}\) denote the treatment indicator for unit \(i\), and write \(D=(D_i)_{i=1}^N\).
Conditional on \(A=a\), the second stage randomizes independently across clusters, selecting exactly \(m_j(a_j)\) treated units uniformly without replacement within cluster \(j\).
Define
\[
\mathcal D_j(s)
=
\left\{
 d_{\mathcal I_j}\in\{0,1\}^{\mathcal I_j}:
 \sum_{i\in\mathcal I_j} d_i=m_j(s)
\right\}.
\]
Then
\[
\Pr(D=d\mid A=a)
=
\prod_{j=1}^K
\binom{n_j}{m_j(a_j)}^{-1}
\mathbf 1\{d_{\mathcal I_j}\in\mathcal D_j(a_j)\}.
\]
The full assignment is \(Z=(A,D)\), with support
\[
\mathcal Z
=
\left\{
(a,d):
 a\in\mathcal A,
 d_{\mathcal I_j}\in\mathcal D_j(a_j)\ \text{for all }j
\right\}.
\]
The designs studied in \citet{Basse2018,Basse2019,liutwostage} arise as the special case \(\mathcal S=\{0,\pi_2\}\) for some \(\pi_2\in(0,1)\).

\subsection{Potential outcomes and feasible untreated exposures}\label{subsec:potential-outcomes}

For each unit \(i\) and feasible assignment \(z\in\mathcal Z\), let \(Y_i(z)\) denote the potential outcome of unit \(i\) under assignment \(z\).
The observed outcome is
\[
Y_i^{\mathrm{obs}}=Y_i(Z^{\mathrm{obs}}),
\qquad i=1,\ldots,N.
\]

We impose the standard homogeneous partial-interference restriction for randomized saturation designs \citep[see, e.g.,][]{hudgens2008,Basse2018,Basse2019,Forastiere2021,Imai2021,Gonzalo-two-stage,liutwostage}.

\begin{assumption}[Homogeneous partial interference]\label{ass:hpi}
For any unit \(i\) and any two feasible assignments \(z=(a,d),z'=(a',d')\in\mathcal Z\),
\[
Y_i(z)=Y_i(z')
\quad\text{whenever}\quad
d_i=d_i'
\quad\text{and}\quad
\sum_{k\in\mathcal I_{[i]}}d_k
=
\sum_{k\in\mathcal I_{[i]}}d_k'.
\]
\end{assumption}

Assumption~\ref{ass:hpi} has two components.
First, it rules out interference across clusters: assignments outside unit \(i\)'s cluster do not affect \(i\)'s outcome.
Second, within a cluster, the outcome depends on the assignment only through unit \(i\)'s own treatment status and the number of treated units in the cluster.
Conditional on these two quantities, the identities of the treated peers are irrelevant.

Because the design fixes the number of treated units in cluster \(j\) at \(m_j(s)\) whenever \(A_j=s\), Assumption~\ref{ass:hpi} allows us to index potential outcomes by own treatment status and cluster saturation whenever that exposure is feasible.
For cluster \(j\), define the feasible exposure set
\[
\mathcal E_j
=
\{(0,s):m_j(s)<n_j\}
\cup
\{(1,s):m_j(s)>0\}.
\]
For \(i\in\mathcal I_j\), the shorthand \(Y_i(r,s)\) is used only for \((r,s)\in\mathcal E_j\).
The observed outcome can therefore be written as
\[
Y_i^{\mathrm{obs}}
=
Y_i\!\left(D_i^{\mathrm{obs}},A_{[i]}^{\mathrm{obs}}\right),
\]
where the realized exposure is necessarily feasible.

The spillover tests in this paper compare untreated exposures.
To avoid off-support untreated potential outcomes in the main theory, define the common untreated-support set
\[
\mathcal S_0
=
\{s\in\mathcal S:m_j(s)<n_j\ \text{for every }j=1,\ldots,K\}.
\]
All pairwise untreated contrasts in Sections~\ref{subsec:two-level-nulls}--\ref{sec:weak-null} use saturation levels \(s,s'\in\mathcal S_0\).
The monotone null in Section~\ref{sec:monotone} is stated over ordered subsets \(\mathcal S_M\subseteq\mathcal S_0\).
Thus a full-saturation label with \(m_j(s)=n_j\) for some cluster is part of the assignment support, but it is not part of an untreated-spillover null unless the null is modified to use a cluster-specific feasible domain.

Assumption~\ref{ass:hpi} is a substantive exposure-mapping assumption, not a consequence of random assignment.
It combines a form of partial interference, which rules out cross-cluster effects, with a form of stratified interference, which reduces within-cluster interference to the cluster saturation.
Thus, the randomization justifies the assignment probabilities used below, but the interpretation of the resulting tests depends on whether the exposure mapping captures the relevant channels of interference.
If the exposure mapping is misspecified, a rejection may reflect either a violation of the stated exposure-defined null or a failure of the exposure mapping itself.
This is the standard interpretational issue that arises in randomization inference under exposure mappings \citep{basse2024}.

\subsection{Individual-level nulls for a pair of saturation levels}\label{subsec:two-level-nulls}

Fix two distinct saturation levels \(s,s'\in\mathcal S_0\).
Throughout this subsection, the target comparison is the untreated spillover contrast
\[
Y_i(0,s)-Y_i(0,s').
\]
The sign convention is arbitrary but fixed: a positive value means that the untreated potential outcome is larger at saturation \(s\) than at saturation \(s'\).

The first null is the individual-level no-spillover null
\begin{equation}\label{eq:null-ps}
H_{0,\mathrm{PS}}^{s,s'}:
\qquad
Y_i(0,s)=Y_i(0,s')
\quad\text{for all } i=1,\ldots,N.
\end{equation}
This null is partially sharp on the full assignment space.
It links two untreated exposures but does not determine treated outcomes or outcomes under other saturation levels.
It becomes sharp for a suitable set of untreated focal units after the conditioning step in Section~\ref{sec:two-level-crt}.

The second individual-level null is a bounded spillover null.
For a prespecified constant \(\delta\in\mathbb R\), define
\begin{equation}\label{eq:null-bounded}
H_{\delta,\mathrm{B}}^{s,s'}:
\qquad
Y_i(0,s)-Y_i(0,s')\le \delta
\quad\text{for all } i=1,\ldots,N.
\end{equation}
This null is useful when the researcher wants to test whether the spillover effect exceeds a substantively meaningful threshold.
The equality version
\begin{equation}\label{eq:null-bounded-equality}
H_{\delta,\mathrm{E}}^{s,s'}:
\qquad
Y_i(0,s)-Y_i(0,s')=\delta
\quad\text{for all } i=1,\ldots,N
\end{equation}
will serve as the least favorable configuration for the finite-sample bounded-null test.


\section{Conditional Tests for Two Saturation Levels}\label{sec:two-level-crt}

This section develops the finite-sample results for pairwise nulls.
For a fixed pair \(s,s'\in\mathcal S_0\), the construction has two steps.
First, we select untreated focal units from clusters whose observed saturation is either \(s\) or \(s'\).
Second, we generate the reference distribution by relabeling the two saturation labels across those clusters while preserving the observed numbers of \(s\)- and \(s'\)-clusters.
Under the partially sharp null, this conditioning makes the focal outcomes imputable.
Under the bounded null, the equality boundary gives a least-favorable imputation.
The weak average nulls do not share this finite-sample imputability property and are treated separately in Section~\ref{sec:weak-null}.

\subsection{Focal units and relabeling distribution}\label{subsec:conditioning-event}

Fix two distinct saturation levels \(s,s'\in\mathcal S_0\).
The comparison uses only clusters whose observed saturation is one of these two levels.
Let
\[
J^{\mathrm{obs}}
:=
\{j:A_j^{\mathrm{obs}}\in\{s,s'\}\}
\]
be this set of contrasted clusters.
For each cluster \(j\), choose an integer \(k_j\) before observing the assignment such that
\begin{equation}\label{eq:kj-condition}
1\le k_j\le \min\{n_j-m_j(s),\,n_j-m_j(s')\}.
\end{equation}
Because \(s,s'\in\mathcal S_0\), the right-hand side is positive for every cluster.
Condition~\eqref{eq:kj-condition} guarantees that cluster \(j\) has at least \(k_j\) untreated units available whether its saturation label is \(s\) or \(s'\).
After observing the assignment, for each \(j\in J^{\mathrm{obs}}\), sample
\[
U_j\subseteq \{i\in\mathcal I_j:D_i^{\mathrm{obs}}=0\},
\qquad |U_j|=k_j,
\]
uniformly without replacement, independently across contrasted clusters.
Let
\[
U=\bigcup_{j\in J^{\mathrm{obs}}}U_j
\]
be the focal unit set.
Thus the test uses exactly \(k_j\) observed untreated focal units from each contrasted cluster and uses no units from clusters with other observed saturation labels.

The reference distribution is obtained by permuting the two saturation labels \(s\) and \(s'\) across the contrasted clusters.
Let
\[
A_J^{\mathrm{obs}}=(A_j^{\mathrm{obs}}:j\in J^{\mathrm{obs}}),
\]
and define the observed margins
\[
K_s:=\sum_{j\in J^{\mathrm{obs}}}\mathbf 1\{A_j^{\mathrm{obs}}=s\},
\qquad
K_{s'}:=\sum_{j\in J^{\mathrm{obs}}}\mathbf 1\{A_j^{\mathrm{obs}}=s'\}.
\]
The relabeling space is
\begin{equation}\label{eq:relabel-space}
\mathcal A_J^{\mathrm{obs}}
=
\left\{
 a_J\in\{s,s'\}^{|J^{\mathrm{obs}}|}:
 \sum_{j\in J^{\mathrm{obs}}}\mathbf 1\{a_j=s\}=K_s
\right\}.
\end{equation}
The relabeling distribution is uniform over \(\mathcal A_J^{\mathrm{obs}}\).
Equivalently, a reference draw selects exactly \(K_s\) of the contrasted clusters to carry label \(s\), assigns label \(s'\) to the remaining \(K_{s'}\) contrasted clusters, and leaves all other clusters outside the comparison.
This is a cluster-level relabeling distribution: focal outcomes are not moved across units or clusters, and the unit-level treatment assignment is not re-randomized in the implementation.

For each contrasted cluster, define the observed focal-cluster mean
\begin{equation}\label{eq:focal-mean}
\widetilde Y_j^{\mathrm{obs}}
:=
\frac{1}{k_j}\sum_{i\in U_j}Y_i^{\mathrm{obs}},
\qquad j\in J^{\mathrm{obs}}.
\end{equation}
Under any relabeling \(a_J\in\mathcal A_J^{\mathrm{obs}}\), the selected focal units remain untreated, and only their cluster saturation label changes from the observed label to the relabeled value.
Thus the focal outcomes are the relevant observed quantities for comparing the untreated exposures \((0,s)\) and \((0,s')\).
Under the partially sharp null \eqref{eq:null-ps}, these focal outcomes are invariant to the relabeling.
For the bounded null, they are combined with the least-favorable imputation described below.

The appendix gives the formal conditional-randomization construction under the framework of \cite{Basse2019} that justifies the uniform cluster-level relabeling distribution.
The main text only uses the resulting focal-unit and relabeling objects because those are the objects needed to implement the test.

\subsection{Statistics for the finite-sample nulls}\label{subsec:test-statistics}

Throughout, write \(J=J^{\mathrm{obs}}\), and let
\[
K_q=\sum_{j\in J}\mathbf 1\{A_j^{\mathrm{obs}}=q\},
\qquad q\in\{s,s'\}.
\]
The relabeling space \(\mathcal A_J^{\mathrm{obs}}\) preserves these margins.
We use upper-tail statistics, so large values provide evidence that untreated outcomes are larger at saturation \(s\) than at saturation \(s'\).

For the partially sharp null \(H_{0,\mathrm{PS}}^{s,s'}\), the conditioning step makes the focal outcomes imputable over \(\mathcal A_J^{\mathrm{obs}}\).
Hence any focal statistic can be used.
Our default choice is the difference in focal-cluster means,
\begin{equation}\label{eq:stat-diffmean}
\hat\tau_U(a_J)
=
\frac{1}{K_s}
\sum_{j\in J}
\widetilde Y_j^{\mathrm{obs}}\mathbf 1\{a_j=s\}
-
\frac{1}{K_{s'}}
\sum_{j\in J}
\widetilde Y_j^{\mathrm{obs}}\mathbf 1\{a_j=s'\}.
\end{equation}

For the bounded null \(H_{\delta,\mathrm B}^{s,s'}\), we use the equality boundary \(H_{\delta,\mathrm E}^{s,s'}\) in \eqref{eq:null-bounded-equality} only to construct the least-favorable imputation.
Under this equality boundary, the focal-cluster mean under saturation \(s'\) is imputed as
\[
\widetilde Y_{j,\delta}^{s'}
:=
\widetilde Y_j^{\mathrm{obs}}
-
\delta\,\mathbf 1\{A_j^{\mathrm{obs}}=s\},
\]
and the corresponding imputed focal-cluster mean under saturation \(s\) is
\[
\widetilde Y_{j,\delta}^{s}
=
\widetilde Y_{j,\delta}^{s'}+\delta .
\]
Thus, for a relabeling \(a_J\), the imputed observed focal-cluster outcome is
\[
\widetilde Y_{j,\delta}(a_j)
=
\widetilde Y_{j,\delta}^{s'}
+
\delta\,\mathbf 1\{a_j=s\}.
\]
The literal imputed difference-in-means statistic is therefore
\[
T_{U,\delta}^{\mathrm{full}}(a_J)
=
\frac{1}{K_s}
\sum_{j\in J}
\left(\widetilde Y_{j,\delta}^{s'}+\delta\right)\mathbf 1\{a_j=s\}
-
\frac{1}{K_{s'}}
\sum_{j\in J}
\widetilde Y_{j,\delta}^{s'}\mathbf 1\{a_j=s'\}.
\]
Because every relabeling in \(\mathcal A_J^{\mathrm{obs}}\) has exactly \(K_s\) clusters labeled \(s\), this statistic can be written as
\[
T_{U,\delta}^{\mathrm{full}}(a_J)
=
\delta+T_{U,\delta}^{\mathrm B}(a_J),
\]
where
\begin{equation}\label{eq:stat-bounded}
T_{U,\delta}^{\mathrm B}(a_J)
=
\frac{1}{K_s}
\sum_{j\in J}
\widetilde Y_{j,\delta}^{s'}\mathbf 1\{a_j=s\}
-
\frac{1}{K_{s'}}
\sum_{j\in J}
\widetilde Y_{j,\delta}^{s'}\mathbf 1\{a_j=s'\}.
\end{equation}
The additive constant \(\delta\) is common to all relabelings and therefore does not affect the randomization \(p\)-value.
We can therefore compute the bounded-null test using the simpler \(s'\)-anchored statistic in \eqref{eq:stat-bounded}.
At the observed relabeling,
\[
T_{U,\delta}^{\mathrm B}(A_J^{\mathrm{obs}})
=
\hat\tau_U(A_J^{\mathrm{obs}})-\delta,
\]
whereas
\[
T_{U,\delta}^{\mathrm{full}}(A_J^{\mathrm{obs}})
=
\hat\tau_U(A_J^{\mathrm{obs}}).
\]
Thus the \(s'\)-anchored formulation does not ignore saturation \(s\); it simply removes the common shift \(\delta\) from the full imputed statistic. The same construction may be applied to focal-count weighted differences in
means, which are often convenient in applications with unequal focal-set sizes,
provided the statistic is monotone in the same least-favorable direction.

For any statistic \(T_U(a_J)\) defined on \(\mathcal A_J^{\mathrm{obs}}\), the corresponding one-sided conditional randomization \(p\)-value is
\[
p_T
=
\frac{1}{|\mathcal A_J^{\mathrm{obs}}|}
\sum_{a_J\in\mathcal A_J^{\mathrm{obs}}}
\mathbf 1\!\left\{
T_U(a_J)\ge T_U(A_J^{\mathrm{obs}})
\right\}.
\]
For the partially sharp and bounded nulls, we use \(T_U=\hat\tau_U\) and \(T_U=T_{U,\delta}^{\mathrm B}\), respectively.

\subsection{The conditional randomization test}\label{subsec:crt-procedure}

The following procedure summarizes the finite-sample conditional testing algorithm.
The permutation step is at the cluster level: we permute the saturation labels of the contrasted clusters and keep the focal units and their observed outcomes fixed.

\begin{testprocedure}[Two-saturation conditional randomization test]\label{proc:two-level-crt}
Fix two saturation levels \(s\neq s'\) in \(\mathcal S_0\), focal-set sizes \(\{k_j\}\) satisfying \eqref{eq:kj-condition}, and a statistic \(T_U(a_J)\) defined on \(\mathcal A_J^{\mathrm{obs}}\).
\begin{enumerate}[(1)]
\item \textit{Select focal controls.}
For each \(j\in J^{\mathrm{obs}}\), sample \(k_j\) units uniformly without replacement from
\[
\{i\in\mathcal I_j:D_i^{\mathrm{obs}}=0\}.
\]
Let \(U_j\) be the selected set in cluster \(j\), and let \(U=\cup_{j\in J^{\mathrm{obs}}}U_j\).

\item \textit{Construct the relabeling space.}
Form \(\mathcal A_J^{\mathrm{obs}}\) as in \eqref{eq:relabel-space}.

\item \textit{Evaluate the statistic.}
For each \(a_J\in\mathcal A_J^{\mathrm{obs}}\), compute \(T_U(a_J)\).
For the partially sharp null, one may use \eqref{eq:stat-diffmean} or any other focal statistic.
For the bounded null, use the shifted statistic \eqref{eq:stat-bounded}.

\item \textit{Compute the conditional \(p\)-value.}
For a one-sided alternative with large values of the statistic unfavorable to the null, set
\begin{equation}\label{eq:generic-crt-pvalue}
p_T
=
\frac{1}{|\mathcal A_J^{\mathrm{obs}}|}
\sum_{a_J\in\mathcal A_J^{\mathrm{obs}}}
\mathbf 1\left\{
T_U(a_J)\ge T_U(A_J^{\mathrm{obs}})
\right\}.
\end{equation}
\end{enumerate}
\end{testprocedure}

The test can be implemented by enumerating \(\mathcal A_J^{\mathrm{obs}}\) or by drawing Monte Carlo relabelings uniformly from \(\mathcal A_J^{\mathrm{obs}}\).
For a Monte Carlo implementation of the finite-sample tests, the usual plus-one correction can be used to preserve conservativeness of the simulated relabeling \(p\)-value.


\begin{theorem}[Finite-sample conditional validity for individual-level two-saturation nulls]\label{thm:two-level-exact}
Fix \(s\neq s'\) in \(\mathcal S_0\), and suppose the randomized saturation design in Section~\ref{subsec:design} is used with focal-set sizes satisfying \eqref{eq:kj-condition}.
Let \(U\) be the focal set generated in Procedure~\ref{proc:two-level-crt}.
Then the following statements hold.

\begin{enumerate}[(a)]
\item \textit{Partially sharp null.}
Under \(H_{0,\mathrm{PS}}^{s,s'}\) in \eqref{eq:null-ps}, the permutation test in Procedure~\ref{proc:two-level-crt} is finite-sample valid for any statistic based only on the focal outcomes and relabeling vector.
That is, for every \(\alpha\in[0,1]\),
\[
\mathbb P\{p_T\le \alpha\mid U\}\le \alpha.
\]

\item \textit{Bounded null.}
Under the bounded null \(H_{\delta,\mathrm B}^{s,s'}\) in \eqref{eq:null-bounded}, the permutation test in Procedure~\ref{proc:two-level-crt} with the shifted statistic \(T_{U,\delta}^{\mathrm B}\) in \eqref{eq:stat-bounded} is finite-sample valid:
\[
\mathbb P\{p_T\le \alpha\mid U\}\le \alpha
\qquad
\text{for every } \alpha\in[0,1].
\]
The same conclusion holds for any one-sided statistic satisfying the same least-favorable monotonicity property as \(T_{U,\delta}^{\mathrm B}\).
\end{enumerate}
\end{theorem}

Theorem~\ref{thm:two-level-exact} is the finite-sample core of the paper.
The focal-unit selection and the cluster-level relabeling distribution are common to the two tests, but the statistic is tailored to the null.
For the partially sharp null, the null itself makes the focal outcomes invariant to relabeling, so any focal statistic is valid.
For the bounded null, the equality null \(H_{\delta,\mathrm E}^{s,s'}\) is least favorable for the one-sided alternative, so imputing under this equality yields a conservative finite-sample test of the inequality null.

\begin{remark}[Choice of focal-set sizes]\label{rem:choice-kj}
The integers \(k_j\) must be chosen independently of the realized labels \(A_j\in\{s,s'\}\).
This ensures that the probability of selecting the observed focal set is the same under every admissible relabeling.
In practice, a simple choice is
\[
k_j=\min\{n_j-m_j(s),\,n_j-m_j(s')\},
\]
which uses as many untreated focal units as possible while preserving symmetry across the two saturation labels.
Smaller choices may be useful when computation is costly or when the analyst wants identical focal-set sizes across clusters.
\end{remark}

\section{Weak Average Nulls as an Asymptotic Extension}\label{sec:weak-null}

The preceding section relies on unit-level restrictions that either impute the focal potential outcomes or admit a least-favorable imputation, yielding finite-sample conditional tests.
We now replace these restrictions with weaker zero-average restrictions.
These nulls are scientifically less restrictive but do not determine the missing focal potential outcomes.
Consequently, the conditional relabeling distribution is no longer an exact finite-sample null distribution; it is instead used to calibrate a studentized statistic with unconditional asymptotic size control.

\subsection{Weak average nulls for a pair of saturation levels}\label{subsec:weak-nulls}

The weak nulls replace the unit-level restrictions above with average restrictions.
For cluster \(j\), define
\[
\mu_{a,j}
:=
\frac{1}{n_j}\sum_{i\in\mathcal I_j}Y_i(0,a),
\qquad a\in\{s,s'\},
\]
and
\[
\tau_j^{s,s'}
:=
\mu_{s,j}-\mu_{s',j}
=
\frac{1}{n_j}\sum_{i\in\mathcal I_j}\{Y_i(0,s)-Y_i(0,s')\}.
\]
Let \(\lambda_{1K},\ldots,\lambda_{KK}\) be nonnegative, nonrandom weights chosen by the researcher, with
\[
\sum_{j=1}^K\lambda_{jK}=1.
\]
The weighted weak null is
\begin{equation}\label{eq:null-weak-weighted}
H_{0,\mathrm W}^{s,s'}(\lambda):
\qquad
\bar\tau_{\lambda,K}^{s,s'}
:=
\sum_{j=1}^K\lambda_{jK}\tau_j^{s,s'}
=
0.
\end{equation}

Two special cases are useful.
If \(\lambda_{jK}=1/K\), then \eqref{eq:null-weak-weighted} is the equally weighted cluster-average weak null,
\begin{equation}\label{eq:null-weak-ew}
H_{0,\mathrm W,\mathrm{eq}}^{s,s'}:
\qquad
\bar\tau_{K}^{s,s'}
:=
\frac{1}{K}\sum_{j=1}^K\tau_j^{s,s'}
=
0.
\end{equation}
If \(\lambda_{jK}=n_j/N\), then \eqref{eq:null-weak-weighted} is the unit-average weak null,
\begin{equation}\label{eq:null-weak-cw}
H_{0,\mathrm W,N}^{s,s'}:
\qquad
\bar\tau_N^{s,s'}
:=
\frac{1}{N}\sum_{j=1}^K n_j\tau_j^{s,s'}
=
\frac{1}{N}\sum_{j=1}^K\sum_{i\in\mathcal I_j}\{Y_i(0,s)-Y_i(0,s')\}
=
0.
\end{equation}
The two targets coincide when cluster sizes are constant.

Unlike the partially sharp and bounded nulls, the weighted weak null does not determine the missing focal potential outcomes.
The relabeling distribution is therefore not an exact finite-sample null distribution in general.
The remainder of this section treats weak-null inference as an asymptotic extension based on studentized relabeling statistics.

\subsection{Studentized relabeling statistic and asymptotic validity}\label{subsec:weak-statistic}

Fix \(s\neq s'\) in \(\mathcal S_0\), and use the focal set and relabeling space from Section~\ref{subsec:conditioning-event}.
Write \(J=J^{\mathrm{obs}}\).
For the weighted weak null \(H_{0,\mathrm W}^{s,s'}(\lambda)\), define
\[
w_{jK}:=K\lambda_{jK},
\qquad
X_j^{\mathrm{obs}}:=w_{jK}\widetilde Y_j^{\mathrm{obs}}.
\]
The normalization \(w_{jK}=K\lambda_{jK}\) is convenient because \(K^{-1}\sum_{j=1}^K w_{jK}=1\).

For \(q\in\{s,s'\}\) and \(a_J\in\mathcal A_J^{\mathrm{obs}}\), define
\[
\bar X_U(q;a_J)
=
\frac{1}{K_q}
\sum_{j\in J}
X_j^{\mathrm{obs}}\mathbf 1\{a_j=q\},
\]
and, when \(K_q\ge 2\),
\[
\hat S_{X,U}^2(q;a_J)
=
\frac{1}{K_q-1}
\sum_{j\in J}
\left\{X_j^{\mathrm{obs}}-\bar X_U(q;a_J)\right\}^2
\mathbf 1\{a_j=q\}.
\]
The Neyman variance estimator for the transformed focal-cluster outcomes is
\[
\hat V_{X,U}^{\mathrm{Ney}}(a_J)
=
\frac{\hat S_{X,U}^2(s;a_J)}{K_s}
+
\frac{\hat S_{X,U}^2(s';a_J)}{K_{s'}}.
\]
The studentized statistic is
\begin{equation}\label{eq:stat-studentized}
T_{U,\lambda}^{\mathrm{Ney}}(a_J)
=
\frac{\bar X_U(s;a_J)-\bar X_U(s';a_J)}
{\sqrt{\hat V_{X,U}^{\mathrm{Ney}}(a_J)}}.
\end{equation}
When the denominator is zero, we use a fixed deterministic convention, such as setting the statistic to zero.
Assumption~\ref{ass:weak-regularity} in Appendix~\ref{app:weak-regularity}
rules out this degeneracy in the limit.
The corresponding upper-tail relabeling \(p\)-value is defined by \eqref{eq:generic-crt-pvalue} with \(T_U=T_{U,\lambda}^{\mathrm{Ney}}\); write this \(p\)-value as \(p_{T,\lambda}\).

The equally weighted cluster-average weak null corresponds to \(w_{jK}=1\), in which case \(T_{U,\lambda}^{\mathrm{Ney}}\) reduces to the unweighted studentized statistic.
The unit-average weak null corresponds to
\[
w_{jK}=\frac{n_j}{N/K}.
\]
Since multiplication of all transformed cluster outcomes by a common positive constant cancels after studentization, the unit-average version can equivalently be implemented by replacing each focal-cluster mean by \(n_j\widetilde Y_j^{\mathrm{obs}}\).



\begin{theorem}[Unconditional asymptotic validity for weighted weak average nulls]\label{thm:weak-null-asymptotic}
Fix \(s\neq s'\) in \(\mathcal S_0\), and suppose the randomized saturation design in Section~\ref{subsec:design} is used with focal-set sizes satisfying \eqref{eq:kj-condition}.
Let \(\lambda_{1K},\ldots,\lambda_{KK}\) be nonnegative, nonrandom weights summing to one, and let \(p_{T,\lambda}\) be the relabeling \(p\)-value computed from \(T_{U,\lambda}^{\mathrm{Ney}}\) in \eqref{eq:stat-studentized}.
If the weighted weak null \(H_{0,\mathrm W}^{s,s'}(\lambda)\) in \eqref{eq:null-weak-weighted} is true and certain regularity conditions defined in Assumption~\ref{ass:weak-regularity} hold, then
\[
\limsup_{K\to\infty}
\mathbb P\{p_{T,\lambda}\le \alpha\}
\le \alpha
\qquad
\text{for every } \alpha\in(0,1/2).
\]
\end{theorem}


Theorem~\ref{thm:weak-null-asymptotic} uses the same focal set and relabeling space as the finite-sample tests, but its validity statement is different.
The randomization distribution used for calibration is conditional on the realized contrasted clusters and focal outcomes, whereas the size guarantee is asymptotic and unconditional over the contrasted cluster set, focal sampling, and saturation assignment.
The distinction matters because the population weak null \(\sum_j\lambda_{jK}\tau_j^{s,s'}=0\) does not imply that the realized contrasted clusters have exactly zero weighted average spillover, nor does it determine the missing focal potential outcomes.


\section{Multiple Saturation Levels and Monotone Nulls}
\label{sec:monotone}

The preceding sections study a fixed pair of saturation levels.
This is the right object when the research question concerns a particular contrast.
In other applications, the substantive question concerns the shape of the spillover response across several saturation levels.
For example, a researcher may want to test whether untreated outcomes are monotone in treatment saturation over a specified range of saturation levels.
This section develops a finite-sample valid unconditional randomization test for such monotone spillover nulls following \citet{zhong2024unconditional}.

\subsection{Monotone nulls over an ordered set of saturation levels}
\label{subsec:monotone-null}

Let
\[
\mathcal S_M=\{q_0,q_1,\ldots,q_M\}\subseteq \mathcal S_0,
\qquad
q_0<q_1<\cdots<q_M,
\]
be the ordered set of feasible untreated saturation levels over which the researcher wants to test monotonicity.
The case \(\mathcal S_M=\mathcal S_0\) corresponds to the global monotone null over the full untreated-support set.
Allowing \(\mathcal S_M\) to be a strict subset is useful when the substantive question concerns monotonicity only over a particular range of saturation levels.

We focus on the monotone increasing spillover null for untreated units,
\begin{equation}\label{eq:null-monotone}
H_{0,\mathrm M}(\mathcal S_M):
\qquad
Y_i(0,q_0)\le Y_i(0,q_1)\le\cdots\le Y_i(0,q_M)
\quad
\text{for all } i=1,\ldots,N.
\end{equation}
The direction can be reversed if the substantive theory predicts that higher saturation weakly lowers untreated outcomes.
Since the experiment randomizes over a finite set of saturation labels, \eqref{eq:null-monotone} is the operational design-based monotonicity restriction.
The null is unitwise: it is stronger than monotonicity of the average spillover response.

One approach would be to test the adjacent inequalities
\[
Y_i(0,q_\ell)\le Y_i(0,q_{\ell+1}),
\qquad \ell=0,\ldots,M-1,
\]
and then combine the resulting \(p\)-values.
This approach is valid if the multiple-testing step is handled appropriately, but it treats monotonicity as a collection of local statements.
We instead construct a single unconditional randomization test for the joint null in \eqref{eq:null-monotone}.

\subsection{Pairwise-imputable monotone statistics}
\label{subsec:monotone-statistics}

The test is based on pairwise comparisons between two assignments.
For any assignments \(z=(a,d)\) and \(z'=(a',d')\) in \(\mathcal Z\), define
\begin{equation}\label{eq:pairwise-imputable-set}
\mathbb I_M(z,z')
=
\left\{
 i\in[N]:
 d_i=d_i'=0
 \text{ and }
 a_{[i]},a'_{[i]}\in\mathcal S_M
\right\}.
\end{equation}
These are the units that are untreated under both assignments and whose two saturation exposures are both covered by the monotone null.
If \(\mathcal S_M=\mathcal S_0\), this set contains all doubly untreated units whose two untreated exposures are feasible.
If \(\mathcal S_M\subsetneq\mathcal S_0\), units whose pairwise exposures fall outside the tested saturation range are excluded because the null imposes no restriction on their pairwise outcome ordering.

\begin{definition}[Pairwise-imputable spillover-monotone statistic]
\label{def:pairwise-monotone-stat}
A statistic
\[
T:\mathbb R^N\times\mathcal Z\times\mathcal Z\to\mathbb R\cup\{\infty\}
\]
is a pairwise-imputable spillover-monotone statistic for \(H_{0,\mathrm M}(\mathcal S_M)\) if, for every pair \(z=(a,d)\) and \(z'=(a',d')\), the following two conditions hold.
\begin{enumerate}[(i)]
\item \textit{Pairwise imputability.}
If two outcome vectors \(y\) and \(y'\) agree on \(\mathbb I_M(z,z')\), then
\[
T(y,z,z')=T(y',z,z').
\]

\item \textit{Monotonicity.}
If, for every \(i\in\mathbb I_M(z,z')\),
\[
\begin{cases}
y_i\ge y_i', & \text{when } a_{[i]}>a'_{[i]},\\
y_i\le y_i', & \text{when } a_{[i]}<a'_{[i]},\\
y_i= y_i', & \text{when } a_{[i]}=a'_{[i]},
\end{cases}
\]
then
\[
T(y,z,z')\ge T(y',z,z').
\]
\end{enumerate}
\end{definition}

Condition (i) says that the statistic uses only units whose two untreated exposures are covered by the null.
Condition (ii) says that the statistic respects the direction of the monotone restriction.
When the saturation exposure is higher under \(z\) than under \(z'\), the corresponding outcome is weakly higher; when the saturation exposure is lower under \(z\), the corresponding outcome is weakly lower; and when the exposure is unchanged, the outcome is unchanged.
The last case is natural because equal exposure labels correspond to the same untreated exposure.
For a monotone decreasing null, the inequalities in Condition (ii) are reversed.

A simple example is a transformed difference in means.
Let \(\psi_1,\psi_0:\mathbb R\to\mathbb R\) be nondecreasing functions.
\begin{equation}\label{eq:transformed-diffmean}
T^{\mathrm{dm}}(y,z,z')
=
\frac{
\sum_{i\in\mathbb I_M(z,z')}
\mathbf 1\{a_{[i]}\ge a'_{[i]}\}\psi_1(y_i)
}{
\sum_{i\in\mathbb I_M(z,z')}
\mathbf 1\{a_{[i]}\ge a'_{[i]}\}
}
-
\frac{
\sum_{i\in\mathbb I_M(z,z')}
\mathbf 1\{a_{[i]}<a'_{[i]}\}\psi_0(y_i)
}{
\sum_{i\in\mathbb I_M(z,z')}
\mathbf 1\{a_{[i]}<a'_{[i]}\}
}
\end{equation}
with the convention \(0/0=0\). When \(\psi_1(u)=\psi_0(u)=u\), this statistic compares outcomes for units whose saturation is weakly higher under \(z\) than under \(z'\) with outcomes for units whose saturation is lower under \(z\) than under \(z'\). One may alternatively omit unchanged-exposure units by replacing \(H(z,z')\) with \(\{i\in\mathbb I_M(z,z'):a_{[i]}>a'_{[i]}\}\).

A rank-based statistic can be defined similarly.
Let \(r_i(y_{\mathbb I_M(z,z')})\) be the rank of \(y_i\) among
\(\{y_\ell:\ell\in\mathbb I_M(z,z')\}\), with ties handled by average ranks.
For a nondecreasing score function \(\varphi\), define
\begin{equation}\label{eq:rank-statistic}
T^{\mathrm{rk}}(y,z,z')
=
\sum_{i\in\mathbb I_M(z,z')}
\mathbf 1\{a_{[i]}\ge a'_{[i]}\}
\varphi\!\left(r_i(y_{\mathbb I_M(z,z')})\right).
\end{equation}
The transformed difference-in-means statistic satisfies Definition~\ref{def:pairwise-monotone-stat} directly by the monotonicity of \(\psi_1\) and \(\psi_0\).
The rank statistic also satisfies the definition: increasing outcomes in the weakly higher group and decreasing outcomes in the lower group weakly increases the ranks, and hence the rank scores, assigned to the weakly higher group.
Thus both \(T^{\mathrm{dm}}\) and \(T^{\mathrm{rk}}\) are pairwise-imputable spillover-monotone statistics.

The definition implies the following pairwise ordering.

\begin{proposition}[Pairwise ordering under monotonicity]
\label{prop:pairwise-ordering}
Suppose Assumption~\ref{ass:hpi} holds and the monotone null \(H_{0,\mathrm M}(\mathcal S_M)\) in \eqref{eq:null-monotone} is true.
If \(T\) is a pairwise-imputable spillover-monotone statistic, then, for every \(z,z'\in\mathcal Z\),
\[
T\{Y(z),z,z'\}
\ge
T\{Y(z'),z,z'\}.
\]
\end{proposition}

\subsection{Unconditional PIRT for the monotone null}
\label{subsec:pirt}

Proposition~\ref{prop:pairwise-ordering} leads to an unconditional randomization test following \citet{zhong2024unconditional}.
For the realized assignment \(Z^{\mathrm{obs}}\), define
\begin{equation}\label{eq:pirt-pvalue}
p_M^{\mathrm{PIRT}}(Z^{\mathrm{obs}})
=
\sum_{z\in\mathcal Z}
\mathbf 1\!\bigl[
T\{Y^{\mathrm{obs}},Z^{\mathrm{obs}},z\}
\ge
T\{Y^{\mathrm{obs}},z,Z^{\mathrm{obs}}\}
\bigr]
P(z).
\end{equation}
Every term in \eqref{eq:pirt-pvalue} is observable because the statistic is pairwise imputable: for each reference assignment \(z\), both statistics depend only on units in \(\mathbb I_M(Z^{\mathrm{obs}},z)\).
When \(|\mathcal Z|\) is too large for enumeration, the sum in \eqref{eq:pirt-pvalue} can be approximated by Monte Carlo draws from the known design distribution \(P\).
Only exact enumeration, or a conservative finite-Monte-Carlo implementation, should be described as finite-sample valid without numerical qualification.

\begin{testprocedure}[PIRT for the monotone spillover null]
\label{proc:pirt}
Fix an ordered saturation set \(\mathcal S_M\subseteq\mathcal S_0\) and a pairwise-imputable spillover-monotone statistic \(T\).
\begin{enumerate}[(1)]
\item Draw or enumerate reference assignments \(z\) from the design distribution \(P\).
\item For each reference assignment, compute
\[
\Gamma(z;Z^{\mathrm{obs}},Y^{\mathrm{obs}})
=
\mathbf 1\!\bigl[
T\{Y^{\mathrm{obs}},Z^{\mathrm{obs}},z\}
\ge
T\{Y^{\mathrm{obs}},z,Z^{\mathrm{obs}}\}
\bigr].
\]
\item Compute
\[
p_M^{\mathrm{PIRT}}(Z^{\mathrm{obs}})
=
\sum_{z\in\mathcal Z}
\Gamma(z;Z^{\mathrm{obs}},Y^{\mathrm{obs}})P(z),
\]
or a Monte Carlo analogue.
\item Reject \(H_{0,\mathrm M}(\mathcal S_M)\) at level \(\alpha\) when
\[
p_M^{\mathrm{PIRT}}(Z^{\mathrm{obs}})\le \alpha/2.
\]
Equivalently, \(\min\{2p_M^{\mathrm{PIRT}},1\}\) is the reported finite-sample valid \(p\)-value under exact enumeration.
\end{enumerate}
\end{testprocedure}

\begin{theorem}[Finite-sample validity of monotone PIRT]
\label{thm:pirt-valid}
Suppose Assumption~\ref{ass:hpi} holds and the monotone null \(H_{0,\mathrm M}(\mathcal S_M)\) in \eqref{eq:null-monotone} is true.
If \(T\) is a pairwise-imputable spillover-monotone statistic, then the rejection rule in Procedure~\ref{proc:pirt} satisfies
\[
\operatorname*{E}_P\!
\left(
\mathbf 1\{p_M^{\mathrm{PIRT}}(Z^{\mathrm{obs}})\le \alpha/2\}
\right)
\le \alpha
\qquad
\text{for every } \alpha\in(0,1).
\]
Thus Procedure~\ref{proc:pirt} is a finite-sample valid unconditional randomization test of the monotone spillover null over \(\mathcal S_M\).
\end{theorem}

The factor \(\alpha/2\) is the price of converting the pairwise ordering in Proposition~\ref{prop:pairwise-ordering} into an unconditional finite-sample test.
The resulting test is conservative in general.
This conservativeness should be weighed against the benefit of testing the monotone null directly over the chosen saturation set, without decomposing it into adjacent contrasts and combining multiple \(p\)-values.

\begin{remark}[Full-support and subset monotonicity]\label{rem:subset-monotonicity}
When \(\mathcal S_M=\mathcal S_0\), the null in \eqref{eq:null-monotone} rules out nonmonotonicity over the full feasible untreated-support set used by the experiment.
When \(\mathcal S_M\subsetneq\mathcal S_0\), the null is weaker and concerns only the specified saturation range.
The PIRT remains valid because the pairwise imputable set in \eqref{eq:pairwise-imputable-set} excludes units whose two exposures are not both covered by the null.
The choice of \(\mathcal S_M\) is therefore part of the null hypothesis: a smaller set can sharpen the substantive target but may reduce the number of units contributing to each pairwise comparison.
\end{remark}

\begin{remark}[General ordered-exposure nulls]\label{rem:general-ordered-exposure}
The PIRT construction does not rely on any special feature of the randomized saturation design beyond the exposure mapping and the known assignment distribution.
More generally, suppose an experiment has assignment space \(\mathcal Z\), assignment law \(P\), and exposure mapping \(e_i:\mathcal Z\to\mathcal E\).
Let \(\mathcal E_0\subseteq\mathcal E\) be the subset of exposures on which the null is stated, and let \(\preceq\) be a partial order on \(\mathcal E_0\).
Consider the ordered-exposure null
\[
Y_i(e)\le Y_i(e')
\quad\text{for all }i\text{ and all }e,e'\in\mathcal E_0\text{ with }e\preceq e'.
\]
For a pair \(z,z'\), the pairwise imputable set is
\[
\mathbb I_{\preceq}(z,z')
=
\left\{
 i\in[N]:
 e_i(z),e_i(z')\in\mathcal E_0
 \text{ and }
 \bigl(e_i(z)\preceq e_i(z')\text{ or }e_i(z')\preceq e_i(z)\bigr)
\right\}.
\]
Units whose two exposures are not both covered by the null, or whose exposures are incomparable, are excluded from the pairwise comparison.
Any statistic satisfying the analogues of Definition~\ref{def:pairwise-monotone-stat} with respect to \(\mathbb I_{\preceq}(z,z')\) yields the same PIRT validity argument.
The saturation-design null in \eqref{eq:null-monotone} is the special case \(e_i(z)=(D_i,A_{[i]})\),
\(
\mathcal E_0=\{(0,q):q\in\mathcal S_M\},
\)
and
\(
(0,q)\preceq (0,q')
\Longleftrightarrow
q\le q'.
\)
\end{remark}

\section{Empirical Application: Spillovers in the Zomba Cash Transfer Experiment}
\label{sec:application_malawi}

We conclude with an application to the Zomba Cash Transfer Program in Malawi.
The experiment is useful for illustration because it combines cluster-level
variation in treatment intensity with individual-level treatment assignment.
\citet{baird2011cash} and \citet{baird2018optimal} use this design to study
direct and spillover effects of cash transfers.  Our purpose is narrower: we use
the public randomized saturation structure to illustrate how the finite-sample
randomization tests developed above can be implemented in a realistic
multi-saturation experiment.

\subsection{Design, exposure labels, and hypotheses}
\label{subsec:malawi_design_hypotheses}

The original experiment assigned enumeration areas (EAs) to cash-transfer
intervention status and then varied schoolgirl offers within treatment EAs.  We
focus on the unconditional cash-transfer (UCT) side of the schoolgirl
intervention, together with treatment EAs in which no baseline schoolgirls were
offered transfers.  This UCT-plus-zero sample keeps the analysis within one
schoolgirl intervention arm while retaining a zero schoolgirl-offer reference
cell.  Table~\ref{tab:malawi_design_cells} summarizes the resulting support.
The \(100\%\) cell remains part of the assignment support but contributes no
untreated focal schoolgirls to the spillover contrasts.

\begin{table}[h]
\centering
\begin{threeparttable}
\caption{UCT-plus-zero application sample}
\label{tab:malawi_design_cells}
\begin{tabular}{@{}llrrr@{}}
\toprule
Label & Schoolgirl policy & EAs & Offered & Untreated \\
\midrule
\(\ell_0\)      & No schoolgirl offer & 15 & 0   & 201 \\
\(\ell_{0.33}\)& UCT, \(33\%\)       & 9  & 68  & 135 \\
\(\ell_{0.66}\)& UCT, \(66\%\)       & 9  & 87  & 44  \\
\(\ell_1\)     & UCT, \(100\%\)      & 9  & 130 & 0   \\
\bottomrule
\end{tabular}
\begin{tablenotes}
\footnotesize
\item Notes: Counts are computed from the public Round 3 analysis file.  The
zero cell is not a pure control group; it is a treatment-EA cell with no
baseline schoolgirl offers.
\end{tablenotes}
\end{threeparttable}
\end{table}

The randomization tests condition on the observed set of 42 EAs in
Table~\ref{tab:malawi_design_cells} and on the fixed compound-cell margins
\(15,9,9,9\).  Under the published complete-randomization description, the
schoolgirl-offer labels $\{\ell_0, \ell_{0.33}, \ell_{0.66}, \ell_1\}$
are completely randomized across these EAs subject to those margins.  The local
CRTs further condition on the selected untreated focal sets and relabel the
relevant EA-level policy labels within the corresponding fixed margins.  Thus
the empirical relabelings should be read as relabelings of compound
schoolgirl-offer policies, not as separate randomizations of the background
dropout intervention or transfer amounts.

The exposure notation below makes this compound-label interpretation explicit.
Let \(\ell_0\) denote the no-schoolgirl-offer treatment-EA label, and let
\(\ell_s\) denote the UCT schoolgirl-offer label at saturation
\(s\in\{0.33,0.66,1\}\).  We write \(Y_i(d,\ell_s)\) for the potential outcome
of baseline schoolgirl \(i\) under own schoolgirl offer status \(d\) and
compound schoolgirl-offer label \(\ell_s\).  This is a reduced exposure mapping:
spillovers are assumed to depend on the randomized schoolgirl-offer policy
label, with other components absorbed into or conditioned on by that label.

For each outcome, let \(\mathcal I_{\mathrm{app}}\) denote the baseline schoolgirls
in the UCT-plus-zero EAs with that Round 3 outcome observed in the public
analysis file.\footnote{The application is therefore a design-based analysis of
the outcome-specific public-data finite population.  The public files contain
the assignment labels, EA identifiers, schoolgirl offer indicators, and Round 3
outcomes needed for this illustration; components used for other purposes in
the published studies are not required for the randomization tests reported
here.}  For \(s>s'\), we test the unit-level bounded null
\[
H_{0,B}^{s,s'}:\qquad
Y_i(0,\ell_s)\le Y_i(0,\ell_{s'})
\quad\text{for every }i\in\mathcal I_{\mathrm{app}}.
\]
The three pairwise contrasts are
\((s,s')=(0.33,0)\), \((0.66,0.33)\), and \((0.66,0)\).  On the raw outcome
scale, a small pairwise \(p\)-value is evidence that the higher compound label
raises the untreated outcome for at least one schoolgirl in the finite
population.  We also test the increasing monotone null
\[
Y_i(0,\ell_0)\le Y_i(0,\ell_{0.33})\le Y_i(0,\ell_{0.66})
\quad\text{for every }i\in\mathcal I_{\mathrm{app}}.
\]
For this null, a small monotone \(p\)-value is evidence against weakly increasing
raw outcomes over the ordered labels.  The four outcomes are current
enrollment, English literacy, ever married, and ever pregnant.

The pairwise bounded-null tests use a focal-count weighted version of the
least-favorable difference-in-means statistic from
Section~\ref{subsec:test-statistics}.  This statistic has the same
least-favorable monotonicity property as \(T_{U,\delta}^{\mathrm B}\), so it is
covered by Theorem~\ref{thm:two-level-exact}.  For the global monotone PIRT, we use the strict changed-exposure variant of the
transformed difference-in-means statistic in \eqref{eq:transformed-diffmean},
applied to \(-Y_i\) rather than \(Y_i\).  This sign reversal makes large values
correspond to violations in which a higher ordered label lowers the raw outcome.  In the
application code, both statistics are computed as focal-count weighted
cluster-mean coefficients.  This regression-style implementation is only a
convenient way to compute the relevant mean contrasts; the validity of the tests
comes from their respective randomization reference distributions.

\subsection{Calibrated simulations}
\label{subsec:malawi_simulation}

We use calibrated simulations to benchmark the candidate procedures in a design
close to the application.  The simulations are semi-synthetic: they fix the
observed EA sizes, saturation-cell counts, and outcome-specific untreated
baseline risks, and vary only the untreated spillover response surface.  For
outcome \(k\) in EA \(j\), the baseline risk is the smoothed untreated mean
\[
\widehat p_{jk}
=
\frac{\sum_{i:j(i)=j,D_i=0}Y_i^k+\lambda\bar Y_0^k}{n_{j0}+\lambda},
\qquad
\lambda=4,
\]
where \(\bar Y_0^k\) is the overall untreated mean and \(n_{j0}\) is the number
of untreated schoolgirls in EA \(j\).  We add an assignment-independent
heavy-tailed EA component
\[
\eta_j
=
0.15\sqrt{n_j/\operatorname{median}(n_j)}\,t_{2j}/\sqrt{2},
\]
where \(t_{2j}\) are independent Student-\(t\) draws with two degrees of
freedom.  The simulated untreated potential outcomes are generated as
\[
Y_i^k(0,a;\tau)
=
\mathbf 1\!\left[
U_i^0
\le
\operatorname{clip}_{[\epsilon,1-\epsilon]}
\left\{
\widehat p_{j(i)k}+\eta_{j(i)}
+\sigma\tau\frac{\min(a,0.66)}{0.66}
\right\}
\right],
\qquad
U_i^0\sim \operatorname{Unif}(0,1),
\]
with \(\epsilon=0.001\), \(\tau\in\{0,0.10,0.20,0.30,0.40,0.50\}\), and
\(\sigma=1\) for the pairwise bounded-null simulations.  For the monotone-null
simulations, we set \(\sigma=-1\), so positive \(\tau\) generates violations of
weakly increasing untreated outcomes.  The smoothing, clipping, and latent EA
components are used only to construct the semi-synthetic designs; the
randomization tests applied to the observed data use the observed outcomes and
the known assignment mechanism.

Figure~\ref{fig:pairwise-power} compares three procedures for the pairwise
bounded spillover nulls: a unit-level linear probability model with
EA-clustered CR1 standard errors, the focal-set conditional CRT, and PIRT.  The
regression benchmark has the largest rejection rates under alternatives but is
oversized under the zero-spillover design, with average size \(0.078\) and
maximum size \(0.120\) across outcome-contrast cells.  The CRT is much closer to
nominal size, with average size \(0.051\) and maximum size \(0.077\).  PIRT is
more conservative, with average size \(0.009\) and maximum size \(0.020\).  Over
the positive effect grid, the average rejection rates are \(0.397\), \(0.315\),
and \(0.180\) for the regression benchmark, CRT, and PIRT, respectively.  We
therefore emphasize the conditional CRT for the pairwise bounded-null
application.

\begin{figure}[t]
\centering
\includegraphics[width=\textwidth]{paper_figure1_pairwise_power_average}
\caption{Calibrated power for pairwise bounded spillover nulls.  The horizontal
line marks the nominal \(5\%\) level.  Each panel corresponds to one saturation
contrast.  Rejection rates are averaged over the four Round 3 outcomes.}
\label{fig:pairwise-power}
\end{figure}

Figure~\ref{fig:monotone-power} compares the global monotone PIRT with adjacent
CRTs combined by Bonferroni.  In the empirical three-level support
\(\{0,0.33,0.66\}\), adjacent CRT plus Bonferroni is stronger near the null:
at \(\tau=0.10\) and \(\tau=0.20\), its rejection rates are \(0.163\) and
\(0.400\), compared with \(0.090\) and \(0.367\) for PIRT.  PIRT overtakes the
adjacent CRT procedure for larger violations.  In a four-level stress support
\(\{0,0.22,0.44,0.66\}\), which keeps the same calibrated outcome structure but
uses a richer ordered support, PIRT has higher average power over the positive
effect grid: \(0.559\), compared with \(0.403\) for adjacent CRT plus
Bonferroni.

These simulations describe operating characteristics; they do not determine the
validity of the observed-data tests.  For the observed pairwise bounded-null
application, we emphasize the focal-set conditional CRT.  For the observed
monotone application, we report both the global monotone PIRT and the adjacent
CRT-Bonferroni procedure.  The latter uses the same local CRT building blocks as
the pairwise analysis, but it tests a different directional null: small pairwise
\(p\)-values support increases in raw outcomes at higher labels, whereas small
monotone \(p\)-values reject weakly increasing raw outcomes.

\begin{figure}[t]
\centering
\includegraphics[width=\textwidth]{paper_figure2_monotone_power}
\caption{Calibrated power for monotone spillover nulls.  The left panel uses
the three-level support in the empirical application; the right panel uses a
four-level stress support.  The horizontal line marks the nominal \(5\%\) level.}
\label{fig:monotone-power}
\end{figure}

\subsection{Application to observed outcomes}
\label{subsec:malawi_results}

Table~\ref{tab:malawi_application_pvalues} reports the observed-data
randomization \(p\)-values.  The first three columns report one-sided focal-set
CRT \(p\)-values for the pairwise bounded nulls.  The last two columns report
valid \(p\)-values for the raw-scale monotone null over
\(\{\ell_0,\ell_{0.33},\ell_{0.66}\}\): the adjusted monotone PIRT value
\(\min\{2p_M^{\mathrm{PIRT}},1\}\) and the adjacent-CRT Bonferroni value.

\begin{table}[t]
\centering
\small
\begin{threeparttable}
\caption{Randomization \(p\)-values in the UCT-plus-zero application}
\label{tab:malawi_application_pvalues}
\begin{tabular*}{\textwidth}{@{\extracolsep{\fill}}lccccc@{}}
\toprule
Outcome
& \(\ell_{0.33}\) vs. \(\ell_0\)
& \(\ell_{0.66}\) vs. \(\ell_{0.33}\)
& \(\ell_{0.66}\) vs. \(\ell_0\)
& PIRT
& Adj. CRT \\
\midrule
Currently enrolled & 0.855 & 0.116 & 0.545 & 0.782 & 0.102 \\
English literacy   & 0.554 & 0.470 & 0.659 & 0.820 & 1.000 \\
Ever married       & 0.064 & 0.749 & 0.943 & 0.788 & 0.669 \\
Ever pregnant      & 0.112 & 0.832 & 0.471 & 0.936 & 0.348 \\
\bottomrule
\end{tabular*}
\begin{tablenotes}
\footnotesize
\item Notes: The first three columns report focal-set CRT \(p\)-values for the
unit-level bounded null \(Y_i(0,\ell_s)\le Y_i(0,\ell_{s'})\) for every
schoolgirl in the application finite population.  Small pairwise \(p\)-values
are evidence that the higher compound label raises the untreated raw outcome.
The last two columns concern the monotone null
\(Y_i(0,\ell_0)\le Y_i(0,\ell_{0.33})\le Y_i(0,\ell_{0.66})\).  PIRT reports
\(\min\{2p_M^{\mathrm{PIRT}},1\}\); Adj. CRT reports the Bonferroni-adjusted
\(p\)-value from adjacent local CRTs.  Small monotone \(p\)-values are evidence
against weakly increasing raw outcomes over the ordered compound labels.
\end{tablenotes}
\end{threeparttable}
\end{table}

None of the reported randomization tests rejects at the \(5\%\) level.  The
smallest pairwise \(p\)-value is \(0.064\), for the
\(\ell_{0.33}\) versus \(\ell_0\) contrast on ever married.  The smallest
monotone \(p\)-value is \(0.102\), from the adjacent-CRT procedure for current
enrollment.  Thus the pairwise CRTs do not reject the unit-level non-increase
bounds in favor of higher untreated outcomes at higher compound labels.  The
monotone PIRT and adjacent-CRT Bonferroni procedures also do not reject the
weakly increasing raw-outcome monotonicity null.  These nonrejections should not
be read as two-sided evidence of no spillover effects; they are evidence only
with respect to the specified one-sided bounded and monotone nulls.  The
exercise illustrates the practical value of exact randomization procedures in a
design with few clusters per saturation cell, heterogeneous cluster sizes, and
partially sharp spillover nulls.

\section{Conclusion}\label{sec:conclusion}

This paper develops a randomization-based toolkit for randomized saturation designs.
The main message is that the conditioning step and the statistic should be separated.
For a fixed pair of saturation levels, conditioning on untreated focal units turns the spillover comparison into a cluster-level relabeling problem.
That relabeling distribution yields finite-sample validity for partially sharp nulls, asymptotic validity for weak nulls when paired with a studentized statistic, and finite-sample validity for bounded nulls when paired with a least favorable shifted statistic.
For multiple saturation levels, a pairwise-imputation approach yields a finite-sample valid unconditional test of the global monotone null.

Several extensions are natural.
First, the same ideas can be applied to total-effect nulls and to contrasts involving treated units, provided the focal set is chosen to preserve imputability.
Second, the weak-null theory can be extended to alternative weighting schemes and regression-adjusted focal-cluster statistics.
Third, the monotone PIRT framework can be adapted to more general exposure mappings, including network exposure mappings and continuous dose-response designs.
These extensions reinforce the main point: randomization inference under interference is most transparent when the null hypothesis, the imputable units, and the statistic are designed together.

\spacingset{1}
\bibliography{biblio}

\spacingset{1.5}
\small
\newpage