Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
76,946 characters · 16 sections · 38 citation commands
Randomization Inference in Two-Sided Market Experiments
\def\spacingset#1{ {#1}} \spacingset{1}
\onehalfspacing
KEYWORDS: Causal inference; Conditional Randomization Test; Two-sided Experiments; Multiple Randomization Design.
\hypersetup{pageanchor=false} \thispagestyle{empty} \hypersetup{pageanchor=true} \setcounter{page}{1}
\spacingset{1.7}
Randomized experiments are increasingly used to assess the impact of policies in various online domains, including online marketplaces, and streaming media platforms Kohavi2020,Athey2019,thomke2003experimentation. These environments generate complex datasets through two-sided interactions between multiple types of actors, such as viewers and content creators or buyers and sellers. In this context, policy evaluation is complicated due to market interference, where the policy treatment on certain units may affect the outcomes of different units in the market, including those never assigned to treatment whymarketplace, toulis2016long. Traditional randomized experiments, designed for single-population settings, often fall short in addressing the complexities arising from interference in two-sided markets imbens2021,johari2021,Wager2021.
To overcome these challenges, a novel class of experimental designs, termed Multiple Randomization Designs, was developed independently by imbens2021,bajari2023 and johari2021, specifically for marketplace experimentation. For example, in an online marketplace with buyers and sellers, the experimenter might randomly assign half of the buyers and half of the sellers to treatment groups, applying a policy intervention (e.g., free shipping) to transactions where both the buyer and the seller are treated. Transactions where either the buyer or the seller (or both) are untreated would not receive the intervention. This design creates variation in the exposure to treatment across different transactions, which facilitates the analysis of causal effects in the presence of interference.
In this paper, we propose Fisherian-style randomization tests Fisher1935Design tailored specifically for analyzing two-sided market experiments of this kind. We formulate sharp null hypotheses on spillover effects that are naturally aligned with the two-sided structure of these experiments. Our approach adopts the conditional randomization testing framework that restricts attention to a set of “focal units” associated with each tested hypothesis Athey2018, Basse2019, puelz2021. In this framework, the spillover hypotheses are not sharp in the traditional sense as they cannot be used to impute all potential outcomes. However, with proper conditioning, the hypotheses become sharp and are thus amenable to standard Fisherian-style randomization.
We propose two valid randomization testing procedures: a permutation-based procedure, which is straightforward to implement but relies on specific symmetry conditions on the design that may not universally hold; and an alternative procedure that is valid under arbitrary designs but requires sampling from a potentially complex space. For a third null hypothesis on total treatment effects, we identify a novel validity-power trade-off among multiple candidate tests and provide heuristic guidance on improving power.
Furthermore, we extend our framework to test weak null hypotheses on average spillover and total treatment effects by incorporating studentized statistics into our randomization tests. While studentization is a well-established technique for ensuring asymptotic validity under weak nulls chung2013, diciccio2017robust, ZHAO2021278, Wu2021, toulis2025asymptotic, we demonstrate that its effectiveness in two-sided experiments depends crucially on how the null hypothesis is formulated and what variance estimators are used for studentization. Specifically, we show that the commonly used Neyman-style studentization is asymptotically valid under a null hypothesis concerning focal average effects, but it may fail to control size under a global average null due to the additional variance introduced by the two-sided randomization. Then, we propose a modified, two-way variance estimator for studentization that restores asymptotic validity for the global weak null by absorbing the additional sampling variation. Our key technical contribution involves extending the proof techniques of ZHAO2021278 to account for these two-sided dependencies, a finding we validate through extensive simulation studies.
This paper contributes to the growing literature on randomized experiments in multi-sided online marketplaces. bajari2023 pioneered the introduction of multiple randomization designs to account for complex spillover effects occurring within marketplaces. imbens2021 delves deeper into such designs, adopting a Neymanian perspective that emphasizes randomization-based point estimation and proposes conservative variance estimators for statistical inference. By contrast, our work considers the Fisherian perspective that instead focuses on finite-sample exact $p$-values via randomization-based tests.
The rest of the paper is organized as follows. Section (ref) describes the setup and notation. Section (ref) presents the main results under sharp null hypotheses. Section (ref) discusses the randomization tests for the weak null hypotheses. Section (ref) examines the finite-sample behavior of our methods through simulations. Section (ref) illustrates the proposed inference methods in an empirical application based on the experiment conducted in Comola2021. Section (ref) concludes the paper.
We begin by introducing the two-sided randomized design following the two-population buyer-seller example in imbens2021. Consider a marketplace comprising \(I\) buyers and \(J\) sellers, indexed by \(i = 1, \ldots, I\) and \(j = 1, \ldots, J\), respectively. In the marketplace, buyers and sellers interact, producing an outcome of interest, such as the total number of transactions between the buyer-seller pair within a given time period. We denote this observed outcome between buyer $i$ and seller $j$ as \(Y_{i,j} \in \mathbb{R}\). We denote the set of buyer-seller pairs as $\mathbb{U}= \{(i,j):i=1,\dots,I, j=1,\dots,J\}$. The \(I \times J\) matrix \(\mathbf{Y} = [Y_{i,j}]\) captures the outcomes observed in the marketplace.
To examine the impact of various interventions, researchers conduct randomized experiments at the level of individual buyer-seller pairs, using a binary treatment assignment, \(W_{i,j} \in \{0, 1\}\). These interventions might include incentives such as free shipping or discounts on transaction fees. The matrix $\mathbf{W} = [W_{i,j}]\in\{0,1\}^{I\times J}$ denotes the full treatment assignment. We use $Y_{i,j}(\mathbf{w})$ to denote the potential outcome for the pair \((i, j)\) under treatment assignment $\mathbf{w}\in\{0,1\}^{I\times J}$, and $\mathbf{Y}(\mathbf{w})$ for the full $I\times J$ matrix of potential outcomes under assignment \(\mathbf{w}\). As usual, the observed outcomes and the potential outcomes are related to treatment assignment by the consistency relationship $\mathbf{Y}=\mathbf{Y}(\mathbf{W})$.
As highlighted by imbens2021, treatments in marketplace experiments need not be uniformly applied across all buyers associated with a particular seller, nor across all sellers interacting with a specific buyer. In this paper, we recast the Multiple Randomization Design introduced by bajari2023 as an independent two-sided randomized design, emphasizing that two independent randomizations occur separately on the two sides (buyers and sellers) of the marketplace. Following imbens2021, a buyer-seller pair—which constitutes the unit of analysis—is considered treated only if both the buyer and the seller are individually assigned treatment. However, we deviate slightly from Definition 3.4 in imbens2021 by relaxing the requirement that the proportion treated on each side must remain fixed. Consequently, in our setting, the full treatment matrix $\mathbf{W}$ can be represented as the outer product of the buyer treatment vector and the seller treatment vector, formalized in the following definition.
In Example (ref), we demonstrate the concept of an independent two-sided experiment under complete randomization. In this design, treatment assignments for each population are uniformly distributed, ensuring a fixed fraction of treated units. Another variant of the two-sided design assigns treatments based on independent and identically distributed (i.i.d.) Bernoulli trials for each buyer and seller. As discussed later, both designs enjoy a nice property known as “design symmetry” or “exchangeability”, which motivates the use of permutation tests as a computationally straightforward approach.
Next, we introduce the concept of local interference, which forms the basis for the two null hypotheses of interest. Under the classical Stable Unit Treatment Value Assumption (SUTVA) Rubin1972, Rubin1980, the potential outcome can be represented by only the treatment status of a given pair, i.e. \(Y_{i,j}(W_{i,j})\). However, recent literature on online marketplace experiments whymarketplace, toulis2016long, johari2021, imbens2021 has pointed out that this assumption is untenable as there is likely interference due to a single seller interacting with multiple buyers or a single buyer interacting with multiple sellers. On the other hand, without any restrictions on the potential outcome function, the analysis would be intractable as the treatment space is exponentially large. To make progress, we follow imbens2021 and assume that interference may only occur between the activities of a given buyer or seller in any given buyer-seller pair.
Following Lemma 3.5 in imbens2021 and under Assumption (ref), we can express the potential outcomes in terms of four distinct treatment exposure cases as follows:\footnote{ To be more specific, by Assumption (ref), the potential outcomes can be represented as \(Y_{i,j}(W_{i,j}, \bar w^B_i, \bar w^S_j)\), where \(\bar w^B_i = \frac{1}{J} \sum_{j=1}^J W_{i,j}\) and \(\bar w^S_j = \frac{1}{I} \sum_{i=1}^I W_{i,j}\) are observed treated fractions. Since $\bar w^B_i = \frac{1}{J} \sum_{j=1}^J w_i^B w_j^S = w_i^B r^S$ and $\bar w^S_j = w_j^S r^B$, where $r^S$ and $r^B$ are the treated fractions of sellers and buyers, respectively, we have \(Y_{i,j}(W_{i,j}, \bar w^B_i, \bar w^S_j) = Y_{i,j}(w_i^B w_j^S, w_i^B r^S, w_j^S r^B)\). Therefore, the potential outcomes are a function of only $w_i^B$ and $w_j^S$ whenever our testing procedures do not alter the treated fractions. This justifies the simplified notation of Equation (ref). }
Certain constrasts between the above potential outcomes express spillover effects from either the buyer side or the seller side. For example, the buyer spillover effect examines the impact on a buyer-seller pair \((i,j)\) when buyers are treated versus when no buyers are treated, holding the seller untreated. Similarly, the seller spillover effect compares the outcomes for a seller when sellers are treated versus untreated, assuming the buyer remains untreated. In Hypotheses (ref) and (ref), which we define below, we formalize the sharp null hypotheses to test for the existence of spillover effects from both the buyer and seller sides. Rejecting these two spillover null hypotheses provides evidence for the existence of interference, suggesting that spillover effects likely exist in the given direction. For instance, if the buyer spillover effect is non-zero, it could indicate that treated buyers' overall shopping experience is influenced by interactions with treated sellers, causing changes in behavior even when shopping from untreated sellers.
The total treatment effect (Hypothesis (ref)) aims to capture the difference between pairs where both the buyer and the seller are treated and pairs where neither is treated. We call this a “total effect” because it encompasses both spillover and direct treatment effects. For a clearer understanding, we discuss this concept through a linear outcome model in Example (ref) that follows.
Finally, we introduce Assumption (ref), which imposes symmetry on the experimental design. This symmetry allows us to construct permutation tests that are computationally efficient, but it is not necessary to construct valid randomization procedures in general.
The classical Fisher Randomization Test simulates the sampling distribution of the test statistic according to the actual treatment variation in the experiment imbens2015causal. While this procedure is valid for testing the sharp null hypothesis under interference, it is generally not valid for non-sharp hypotheses, which do not specify the complete schedule of potential outcomes. The null hypotheses defined in the previous section fall into this category. For instance, if we observe outcome $Y_{ij}(1, 0)$ for a treated buyer $i$ and control seller $j$, then we cannot impute the outcome $Y_{ij}(0,0)$ under $H_0^{\text{seller}}$---this null hypothesis is not sharp in the strict sense.
To address this issue, recent literature has proposed the use of conditional randomization tests that execute the resampling procedure on a subset of units, $\mathcal U \subseteq \mathbb{U}$ and a subset of assignments, $\mathcal{W} \subseteq \mathbb{W}$, such that the potential outcomes become fully specified Aronow2012, Athey2018, Basse2019. Note also that a unit is a buyer-seller pair in our case. In the terminology of Basse2019, $\mathcal{C} = (\mathcal U, \mathcal{W})$ is the conditioning event, which may depend on the observed treatment assignment in a random way. Let $p(\mathcal{C} \mid \mathbf{W})$ denote its distribution, which is under the analyst's control. The conditional Fisher Randomization Test then randomizes treatment according to its conditional distribution:
Such a test is finite-sample valid even for a non-sharp null hypothesis, provided that the conditioning event, $\mathcal{C}$, is constructed in a way such that the potential outcomes for all units and assignments in $\mathcal{C}$ can be imputed under the null hypothesis Basse2019. Sampling from the conditional distribution in (ref) may be challenging in general. However, under certain conditions on the design, $p(\mathbf{W})$, and the conditioning mechanism, $p(\mathcal{C}|\mathbf{W})$, the distribution in (ref) can be reduced to a permutation distribution that is easy to sample from. We will consider this approach in the following sections.
In this section, we develop conditional randomization tests for the spillover hypotheses $H_0^{\text{buyer}}$ and \(H_0^{\text{seller}}\). Under symmetric designs, our randomization procedure entails straightforward permutations of treatment assignments on one side (buyer or seller), conditioned on the assignment of the other side (seller or buyer, respectively). We begin with the buyer spillover hypothesis and describe the associated randomization procedure in detail. The analysis of the seller spillover hypothesis will be brief as it is completely symmetrical to the buyer case.
For the buyer spillover hypothesis (ref), we propose to condition on control sellers and fix those sellers to control across randomizations to ensure that the potential outcomes can be imputed. Moreover, we will condition on the marginal number of treated buyer-sellers observed in the sample. Under Assumption (ref), the conditional randomization distribution is exchangeable, resulting in a permutation test that can be efficiently implemented.
Formally, we define the following conditioning events:
where $\mathcal{M} = \{v u^\top: v\in \mathbb{W}^B, u \in \mathbb{W}^S \text{ s.t. } \sum_{i=1}^I v_i = \sum_{i=1}^I w^{obs, B}_i, \sum_{j=1}^J u_j = \sum_{j=1}^J w^{obs, S}_j\}$. Here, we use the superscript `obs' to emphasize that we are referring to the actual treatments realized in the observed data. Thus, $\mathcal U$ denotes the buyer-seller pairs for which the seller part of the pair is assigned control. Set $\mathcal{W}$ denotes the assignments that these sellers in control, while $\mathcal{M}$ is the set of assignments that match the marginal number of treated buyers and sellers observed in the sample. We note that all these sets are random as they all depend on the observed treatment assignment vector, $\mathbf{W}^{obs}$. We are now ready to define the main randomization test for the buyer spillover hypothesis.
Figure (ref) is a graphical depiction of Procedure (ref) for testing the buyer spillover effect. The shaded parts of the array in the figure illustrates that the test conditions on the set of sellers that were not treated. Procedure (ref) then randomizes the treatments of buyers (i.e., rows in the array). Under certain assumptions, such randomization is equivalent to permutations as explained in the following remark.
The validity of our proposed randomization tests in Procedure (ref) follows by adapting Theorem 2 of puelz2021 in the setting of two-sided experiments. We provide the statement of the validity result in the following theorem.
In this section, we proceed to discuss the testing procedure for total effects. We propose a randomization procedure that relies on a \(k\)-Block Conditioning Event. The idea is to split focal units into \(k\)-by-\(k\) blocks, with the distinctive arrangement ensuring that units within a specific block do not share rows or columns with units from other blocks. Furthermore, every unit within a block shares the same treatment status for both buyers and sellers.
Figure (ref) illustrates an example of a $k$-block conditioning event corresponding to the shaded diagonal blocks. The treatment assignments within each block are either $(w_i^B, w_j^S) = (0,0)$ or $(w_i^B, w_j^S) = (1,1)$, with no overlap of blocks across columns or rows. Such arrangement is important as it allows the use of efficient permutation procedures under symmetric two-sided randomized designs. In other words, under symmetric treatment assignment designs, such as complete randomization and Bernoulli trials, our proposed tests involve permuting the treatment assignments of focal units across blocks.
Towards testing the total effect, let $\mathcal{I}^{(k)}(\mathbf{W}^{obs})$ denote a random partition of the buyer set $[I]$ into $k$ equal-sized non-overlapping subsets with all buyers having the same treatment status within each subset.\footnote{We assume $I$ and $\sum_{i=1}^Iw_i^{obs, B}$ are both divisible by $k$ for simplicity. When this is not the case, our methods remain valid, and the framework can be extended accordingly, though at the cost of more cumbersome notation.} That is,
such that for all $1 \leq s, s^\prime \leq I/k$:
Similarly, we define $\mathcal{J}^{(k)}(\mathbf{W}^{obs}) $ as a random partition of the seller set $[J]$ into $k$-sized subsets with sellers having the same treatment status within each subset. Next, define
where $\mathcal{I}_s, \mathcal{J}_{s}$ are defined above. That is, $\mathcal{U}^{(k)}(\mathbf{W}^{obs})$ is the collection of buyer-seller pairs constructed from the cross-product of $\mathcal{I}^{(k)}(\mathbf{W}^{obs})$ and $\mathcal{J}^{(k)}(\mathbf{W}^{obs})$, making sure that all buyers and sellers have the same treatment status in the group. Figure (ref) below illustrates this construction. In the figure, the set $\mathcal{U}^{(k)}(\mathbf{W}^{obs})$ therefore corresponds to the diagonal blocks in the shaded area of the buyer-seller array. Note how these blocks are non-overlapping, they partition the sets $[I]$ and $[J]$, and, within each block, the individual treatments of every buyer and seller are identical.
We are now ready to define our main randomization procedure for testing the total null hypothesis.
Intuitively, to test $H_0^{\text{total}}$, Procedure (ref) conditions on a $k$-block conditioning event ---e.g., the $k$ diagonal blocks shown in Figure (ref)--- and permutes the block treatments $(0,0)$, $(1,1)$ at the block level. We note that the diagonal structure in the blocks is not essential and alternative constructions can be valid as long as the blocks are non-overlapping in neither the buyer nor the seller dimension.
The following result shows that Procedure (ref) is valid under randomized designs that satisfy local interference and design symmetry.
In the case of non-symmetric designs where Assumption (ref) does not hold, we can still construct valid randomization tests, albeit in a more complex form. Specifically, suppose without loss of generality that $I =J$ and the $k$-partitioning on two populations are identical, i.e. $\mathcal{I}^{(k)}(\mathbf{W}^{obs}) = \mathcal{J}^{(k)}(\mathbf{W}^{obs})$. Then, the permutation in step 2 of Procedure (ref) with sampling from the conditional
and setting $\mathbf{W}^{(l)} = w^{B, (l)} (w^{B, (l)})^\top$. Despite its simplicity, this approach may be complex to implement since sampling from the conditional (ref) could be computationally challenging under arbitrary non-symmetric designs.
Theorem (ref) establishes that Procedure (ref) is valid for any block size \(k\). However, the statistical power of the procedure likely depends on a complex trade-off relating to the value of $k$. A smaller value of $k$ leads to using more blocks and thus is beneficial for the test's power since it increases the support of the randomization distribution. However, it also leads to smaller block sizes, and thus it may decrease power due to the reduced sample size. Conversely, a higher value of $k$ leads to the reverse effect.
To illustrate further, consider a completely randomized two-sided design following Example (ref), with \(I_1\) and \(J_1\) treated rows and columns, respectively. For the sake of simplicity, let's assume the observed data matrix is square with half of the rows and columns assigned to the treatment group, such that \(I = J = 2n\) and \(I_1 = I_0 = J_1 = J_0 = n\), with \(n/k\) being an integer. Under these conditions, the total number of unique treatment assignments ---i.e., the support of the randomization distribution--- in a $k$-block conditioning event equals \(C(2n/k, n/k)=O(2^{2n/k}/\sqrt{2n/k})\), where \(C(\cdot, \cdot)\) represents the binomial coefficient. Moreover, the sample size of focal units is $|\mathcal{U}^{(k)}(\mathbf{W}^{obs})| = 2nk$. Thus, as $k$ increases, the sample size increases but the support of the randomization distribution decreases, which have opposing effects on the test power. Conversely, as $k$ decreases, the total sample size decreases but the support of the randomization distribution increases.
To navigate this trade-off, we can leverage the results of puelz2021, who established a lower bound for the power of general conditional randomization tests under some plausible assumptions, which we detail in Appendix (ref). Let \(\phi_{n,k}\) denote the power of a conditional randomization test in our setting. The following proposition provides the bound under complete randomization, as in Example (ref):
In Figure (ref), we illustrate the lower bound on power as a function of the block size \( k \), highlighting the trade-off between sample size and the number of unique treatment assignments as \( k \) increases. While the plot identifies the optimal \( k \) that maximizes power, in practice, the parameters of the lower bound function are unknown. Consequently, the optimal \( k \) cannot be directly determined using Proposition (ref). However, a heuristic approach can be used to select the largest feasible \( k \) based on a predetermined maximum power threshold. Based on Proposition (ref), $O\left(C(2n/k, n/k)^{-0.5 + \delta}\right)$ controls the maximum power of the test, while $2nk$ determines how rapidly the power function reaches its maximum (sensitivity). Therefore, to optimize the detection of treatment effects, we recommend selecting the largest feasible block size $k$ based on a predetermined maximum power. For example, setting this power to 0.95 ensures the maximal utilization of observations to detect the minimal size of the treatment effect. In practice, we can choose $k$ to satisfy $C(2n/k,n/k)\approx\frac{1}{(1-\beta)^2}$, where $\beta$ is the predetermined power level. For example, consider the completely randomized two-sided design discussed before, where choosing a block size $k = n/6$ results in $|\mathcal{W}| = C(12,6) = 924$, achieving a maximum power approximately equal to 0.967.
There has been a growing number of works on the study of randomization tests under weak null hypotheses; see, for example, chung2013,Ding2017,Canay2017,diciccio2017robust,DiCiccio&Romano2017,ZHAO2021278,Wu2021. This literature has pointed out that randomization tests based on sharp nulls are not necessarily valid under the weak null. Nevertheless, in classic one-sided experiments, validity can often be restored by studentizing the test statistic. The most commonly used approach relies on a Neyman-style variance estimator, which coincides with the usual two-sample $t$-test variance.
In this section, we study use of conditional randomization tests with suitable test statistics and studentization for testing weak null hypotheses in two-sided randomization design. We document that the validity of these tests hinges on both the specific form of the null as well as the way in which studentization is carried out.
Specifically, we consider two types of weak nulls. The first, which we call the population average spillover effect from buyers, imposes a single mean-zero restriction on the average treatment effect after pooling over all buyers and sellers.
The second, which we call the seller-specific average spillover effect from buyers, imposes a collection of mean-zero restrictions, one for each seller, requiring that the buyer-level average treatment effect vanish seller by seller.
We show that the natural test statistic, together with commonly used Neyman-style studentization, leads to a valid test for the seller-specific null, but not for the population null. Therefore, we propose a modified, two-way (buyer $\times$ seller) variance estimator that restores asymptotic validity for the population-average weak null by absorbing the additional sampling variation induced by seller randomization.
We begin by presenting the testing procedure under standard Neyman-style studentization:
The following theorem summarizes the asymptotic properties of Procedure (ref) under the two weak null hypotheses. Importantly, it establishes (i) asymptotic validity under the seller-specific weak null $H_0^{wb,2}$ in (ref), and (ii) the failure of asymptotic validity in general under the population-average weak null $H_0^{wb,1}$ in (ref).
The divergence in the asymptotic validity results between $H_0^{wb,1}$ and $H_0^{wb,2}$ can be understood through the standard comparison between the randomization and sampling distributions of the studentized test statistic. In the literature on randomization tests for weak null hypotheses, validity is typically established by showing that the randomization distribution---here, the conditional permutation distribution induced by $p(\mathbf{W} \mid \mathcal{C})$---asymptotically stochastically dominates the sampling distribution induced by the original design. A common route is to show that the randomization distribution converges to $N(0,1)$, while the sampling distribution converges to a centered normal distribution with variance no larger than 1.
In our setting, the key distinction between $H_0^{wb,1}$ and $H_0^{wb,2}$ is a mismatch between two average treatment effects. The conditional sampling distribution is centered at the focal average effect:
whereas the weak null hypothesis $H_0^{wb,1}$ in (ref) concerns the global average effect, $\tau$, defined as:
Under $H_0^{wb,2}$, both quantities vanish by construction, so the centering mismatch disappears: $\tau(w^S)=\tau=0$. Under $H_0^{wb,1}$, however, only the global restriction $\tau=0$ is imposed, and $\tau(w^S)$ remains random. In particular, it can be shown that $|\tau(w^S)-\tau|=O_p(1/\sqrt{J})$. Consequently, under $I\asymp J$, the studentized statistic admits the decomposition
where $T^B := T^{B}(\mathbf{W} \mid \mathbf{Y}, \mathcal{C})$ and $V^B := V^{B}(\mathbf{W} \mid \mathbf{Y}, \mathcal{C})$. As we show in Appendix (ref), both terms converge to normal limits: \[ A_N \xrightarrow{d} N(0, \sigma_A^2), \qquad B_N \xrightarrow{d} N(0, \sigma_B^2), \] and hence
The sampling distribution of $B_N$ is stochastically dominated by $N(0,1)$, whereas $A_N$ reflects additional variation induced by seller randomization. Under $H_0^{wb,2}$, the term $A_N$ vanishes, and the usual stochastic-dominance argument goes through. Under $H_0^{wb,1}$, by contrast, $A_N$ contributes non-negligibly to the sampling distribution, so the total variance $\sigma_A^2+\sigma_B^2$ need not be bounded by 1, the asymptotic variance of the randomization distribution. This creates the possibility of over-rejection under $H_0^{wb,1}$. Figure (ref) illustrates the limit distributions of the relevant terms in the Neyman-style test statistic under the population-average weak null.
To restore validity under $H_0^{wb,1}$, we augment the standard Neyman variance estimator by an additional term that estimates the sampling variance of the seller-specific mean effect. The resulting variance estimator resembles a design-based version of a two-way (buyer $\times$ seller) cluster-robust variance: the original $V^B$ captures buyer-randomization variation, while the new add-on captures seller-randomization variation.
For each control seller $j$ with $w_j^S=0$, define the within-seller difference-in-means across buyers,
and let $\bar{\hat{\mu}}^{\Delta} := J_0^{-1}\sum_{j:w_j^S=0}\hat{\mu}^{\Delta}_j$. Define the sample variance across control sellers, $s^2_{\hat{\mu}^{\Delta}}(\mathbf{W}\mid \mathbf{Y},\mathcal{C}):= \frac{1}{J_0-1}\sum_{j:w_j^S=0}\left(\hat{\mu}^{\Delta}_j-\bar{\hat{\mu}}^{\Delta}\right)^2$, and the seller-side add-on
Finally define the two-way variance estimator and the corresponding studentized statistic:
To test the population-average weak null $H_0^{wb,1}$, we implement the same permutation procedure as in Procedure (ref), but replacing $T^{WB}$ with $T^{WB,TW}$ in Steps (i)--(iii). The next theorem states that this two-way studentization restores asymptotic validity under the weak null $H_0^{wb,1}$.
In this subsection, we examine the finite-sample behavior of our randomization tests for both total and spillover effects under the sharp null hypotheses. Suppose that there are $3n$ units in each population, i.e. $I=J=3n$, and the design is a completely randomized two-sided design with $n$ units assigned to treatment group in both populations. We generate the potential outcomes as follows:
where $F_\ell$ are normal distributions such that $F_{\ell}(\cdot)=\mathcal{N}\left(\mu_{\ell}, \sigma_{\ell}^2\right)$ for $\ell \in \{0,B,S,1\}$. Under the sharp null hypothesis, the parameters are set as follows: $\mu_0 = \mu_S = \mu_B = \mu_1 = 0$, $\sigma_0= 0.2$ and $\sigma_B=\sigma_S = \sigma_1 = 0$. Under the alternative hypotheses, we set the parameters in a similar way but set the spillover effect to be 0.01 and total effects to be 0.02, i.e. $\mu_B = 0.01, \mu_1 = 0.02$.
Table (ref) displays the rejection probabilities under the sharp null and alternative hypotheses, computed from 5,000 Monte Carlo replications with the $p$-values approximated by 500 independent permutations of the treatment vector in each replication. “FRT” stands for the Fisher Randomization Test procedure based on sharp null hypotheses as in Procedure (ref) and (ref). “FRT adjusted” stands for the Neyman-style studentized randomization tests as in Procedure (ref).\footnote{We choose not to present the two-way studentized test statistics in this subsection, because it is developed for the weak null concerning global average effects and only available for buyer spillover effects.} “Neymanian” stands for $t$-tests using conservative estimators of variances from imbens2021. The block size for Procedure (ref) is set to be $k=\lfloor n/4 \rfloor$, resulting in $|\mathcal{W}|=C(12,4)=495$ and a maximum power approximately equal to 0.955.
The results show that the rejection probabilities of both “FRT” and “FRT adjusted” are universally around the nominal 5% level under the null hypothesis, which verifies the finite-sample exactness of our tests across all designs. Meanwhile, the Neymanian inference method is conservative as expected by imbens2021. Under the alternative hypotheses, the rejection probabilities of our tests are higher than the Neymanian method when testing spillover effects, but lower when testing total effects. The power gain by Procedure (ref) likely comes from the test's exactness, whereas the power loss from Procedure (ref) is due to the loss in sample size used for calculating the block test statistic.
We further investigate robustness by examining the influence of block size on both the size and power of the testing procedures. Table (ref) presents our analysis of rejection probabilities under the sharp null and alternative hypotheses, following the same Monte Carlo setup in Table (ref). We observe that up to a point the statistical power of our tests increases monotonically with increasing block size. Specifically, when block size is set at \(k=25\), where the randomization space encompasses \(|\mathcal{W}| = C(12,4) = 495\), near-maximum power is achieved. However, a further increase in block size to \(k=50\) results in a diminished randomization space of \(|\mathcal{W}| = C(6,2) = 15\), leading to a notable decline in test power. These results are consistent with the pattern observed in the graphical illustration of the theoretical lower bound in Figure (ref).
In this subsection, we consider two weak null hypotheses for buyer spillover effects introduced in Section (ref): $H_0^{wb,1}$, which concerns the population average, and $H_0^{wb,2}$, which concerns the seller-specific average. Appendix (ref) presents corresponding results for weak null hypotheses on total effects.
Notably, when treatment effects are i.i.d sampled as in Section (ref), the weak null hypotheses of seller-specific average $H_0^{wb,2}$ are satisfied by the data generating process (DGP). Because the DGP generates treatment effects in a dyadic i.i.d. fashion, the average effects over any sufficiently large subset of the population—including the focal set of untreated sellers—will converge to zero asymptotically. Therefore, we adopt the same DGP from the previous subsection and change the variance parameters as follows: $\sigma_0=0.2, \sigma_B=\sigma_1 = 0.4$ and $\sigma_S = 0$.
In contrast, to examine the population average spillover effect $H_0^{wb, 1}$, we consider a new DGP that satisfies $H_0^{wb,1}$ but violates $H_0^{wb,2}$. Specifically, let $\tilde{Y}_{ij}(0,0)$ and $\tilde{\Delta}_{ij}$ denote the baseline i.i.d.\ components generated as in the previous subsection: $\tilde{Y}_{ij}(0,0)\stackrel{\text{ind}}{\sim}\mathcal{N}(0,0.2^2)$, $\tilde{\Delta}_{ij}\stackrel{\text{ind}}{\sim}\mathcal{N}(0,0.4^2)$. Then, we define $Y_{ij}(1,0)=Y_{ij}(0,0)+\Delta_{ij}$ with \[ Y_{ij}(0,0)=\tilde{Y}_{ij}(0,0)+\alpha_i, \qquad \Delta_{ij}=\tilde{\Delta}_{ij}+\beta_j, \] where $\alpha_i$ and $\beta_j$ are independent fixed effects, i.e. $\alpha_i \stackrel{\text{ind}}{\sim} \mathcal{N}(0,0.1^2)$ and $\beta_j \stackrel{\text{ind}}{\sim} \mathcal{N}(0,0.4^2)$.\footnote{The remaining potential outcomes $Y_{ij}(0,1)$ and $Y_{ij}(1,1)$ are generated as in the previous subsection and play no role in the buyer-spillover test considered here.} This DGP violates $H_0^{wb,2}$ because seller-specific average treatment effects vary across $j$ through $\beta_j$, so $\tau(w^S)$ need not be close to zero for a realized seller focal set. However, the global average effect $\tau=0$ remains satisfied because both $\tilde{\Delta}_{ij}$ and $\beta_j$ have mean zero.
Table (ref) reports rejection probabilities at the nominal 5% level based on 5,000 Monte Carlo replications, each using 500 independent permutations. The results in the first half of the table indicate that, under $H_0^{wb,2}$, the adjusted FRT using the Neyman-style studentized statistic successfully maintains the nominal level. So does the two-way adjusted FRT. In contrast, the standard FRT (without studentization) fails to provide valid inference in this setting. This corroborates our theoretical analysis: when seller-specific average effect is zero, the Neyman-style studentization restores the asymptotic validity of the randomization test. In contrast, the results in the second half of the table indicate that, the two-way studentized procedure (“FRT two-way”) controls size well, whereas the standard Neyman-style studentization (“FRT adjusted”) fails due to focal-average randomness induced by seller heterogeneity, as predicted by Theorem (ref)-(ref).
In this section, we illustrate our methodology using a dataset from a randomized field experiment conducted by Comola2021, which offered access to formal savings accounts to a random sample of 915 households across 19 villages near Pokhara, Nepal. The treatment is defined at the household level as whether a household was offered a savings account. The primary outcome of interest is the change in a network link indicating whether one household lends to another. These links are measured using survey-based adjacency matrices.\footnote{To be more specific, the network link is constructed from survey questions such as “Who would you ask for help in case of need?” and “Who did you ask for help?”} The matrices are transformed into semi row-standardized versions, $\mathbf{G}^{(t)} = \{G_{i,j}^{(t)}\}$, where each row sums to one for non-isolated households and to zero otherwise and $t \in \{0,1\}$ denotes the baseline and endline periods. The final outcome is defined as $Y_{i,j} = G_{i,j}^{(1)} - G_{i,j}^{(0)}$, representing the change in standardized network links between households $i$ and $j$.
We follow the empirical question of interest in Comola2021 and test whether access to a savings account induces financial exchanges between households. In particular, we are interested in exploring heterogeneity of the treatment effect with respect to whether one of the households has experienced a negative shock (death or livestock loss). To formalize this, we construct a binary variable \( X_i \in \{0,1\} \), where \( X_i = 1 \) if household \( i \) experienced a death or livestock shock at baseline and \( X_i = 0 \) otherwise.\footnote{The original dataset includes only household‑pair covariates, but we recover household‑level covariates by aggregating information across all pairwise records involving each household.} This variable acts as a covariate because it is realized before treatment is assigned. The rationale behind this choice is that households experiencing shocks may be in greater financial need and are therefore interpreted as “buyers” of informal loans, while those without shocks are potential lenders (“sellers”).
This setting leads naturally to the same potential outcomes structure as in our two-sided experiment framework, with potential outcomes $Y_{i,j}(w_i^B, w_j^S)$, where \( w_i^B, w_j^S \in \{0,1\} \) denote the treatment assignments of households \( i \) and \( j \), respectively. The justification for this structure relies on a different version of the local interference assumption: the outcome for a given household pair depends only on the treatment assignments of the two households in that pair and not on the assignments of any other households in the network.
The null hypotheses also retain the same structural form as before but carry different empirical interpretations. Specifically, we test three hypotheses corresponding to different treatment comparisons:
Although the original experiment was not designed as a two-sided market intervention, the setting naturally fits our framework. The outcome of interest is dyadic—measured at the household-pair level—so the experimental design can be interpreted as a two-sided randomization problem where the “buyer” and “seller” sides originate from the same population. In this setting, the outcome $\mathbf{Y}$ is a $915 \times 915$ matrix and the assignment vector is 915‑dimensional, with the buyer‑side and seller‑side assignments coinciding, i.e., $w^B = w^S$. In our analysis, we take a further step to partition households into “buyers” and “sellers” using baseline information on which households experienced shocks. This not only allows us to more directly apply our two-sided randomization framework but also reflects a meaningful economic distinction that enables analysis of treatment effect heterogeneity.
To test $H_0^{(1,0)}$ and $H_0^{(0,1)}$, we apply the one-sided permutation procedure described in Procedure (ref). For $H_0^{\text{total}}$, we leverage the fact that the network is censored as survey data record only within-village links. This allows us to use the block-wise permutation procedure in Procedure (ref), where each village defines a permutation block. Table (ref) reports $p$-values from our randomization-based tests alongside those from a standard two-sample $t$-test.\footnote{We omit the results for “FRT two-way” because they are more conservative than those for FRT adjusted” and are therefore not statistically significant.} We include the $t$-test for comparison, noting that the Neymanian-style methods (imbens2021) are not directly applicable due to the censored nature of the data. In contrast, our approach remains valid by conditioning on the observed household pairs.
We fail to reject all three null hypotheses, indicating that providing savings accounts did not significantly affect households’ risk-sharing relationships. This finding complements Comola2021, who document significant effects on network behavior in the full population, which suggests that such effects are not driven by households experiencing shocks. Consistent with prior evidence, the impacts of savings accounts appear highly context-dependent: some studies find limited usage among poor households Dupas2016, while others report high take-up but no clear effects on aggregate expenditure, assets, or income PRINA2015. Overall, these results suggest that the effectiveness of savings accounts depends critically on how households use them, and that access alone may be insufficient to change financial behavior.
Motivated by recent advances in experimentation within online marketplaces, this paper develops randomization-based inference procedures for two-sided market experiments. Our proposed tests are finite-sample valid under sharp null hypotheses of no treatment effect, and asymptotically valid for weak null hypotheses on average treatment effects. Additionally, we offer practical guidance for test implementation based on power considerations. Promising directions for future research include extending our framework to accommodate more complex interference structures beyond the cross-product (buyer-seller) interactions examined here. Another extension would incorporate explicit market-clearing mechanisms, such as pricing rules.
\spacingset{1}