EconBase
← Back to paper

Inference for Group Interaction Experiments

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

85,011 characters · 18 sections · 54 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Inference for Group Interaction Experiments

abstractA common experimental research design is one in which individuals are randomly allocated into groups that then interact under different group-level treatment conditions. We develop design-based inference for such “group interaction” experiments, covering scenarios in which groups are either fixed or randomly formed and in which potential outcomes are either fixed relative to others' group assignments or subject to interference. For each scenario, we characterize the causal estimand that the design targets and the inferential strategy appropriate to it. Working in a sparse-sampling asymptotic regime, we show that cluster-robust inference remains consistent and accounts for dependencies from various sources when interference is present, delivering valid inference on marginalized exposure effects. When interference is absent and groups are formed randomly, the design reduces to an individually randomized experiment, and individual-level heteroskedasticity-robust inference suffices for the average treatment effect. Our results on the asymptotic distribution of commonly used estimators rely on a novel coupling strategy that may be useful for design-based inference in other complex experiments.

Keywords: Causal Inference, Interference, Marginalized Exposure Effects, Group Interaction Experiments, Design-Based Inference, Cluster-Robust Inference

\doublespacing

Introduction

A common experimental design in social science has individuals assigned into groups where they interact with their group mates under different group-level treatment assignments. We call these “group interaction” experiments.\footnote{lohr2014partially use “partially nested” design and pals2008individually use “individually randomized group treatment trials” for the same design, although these phrasings does not evoke group interaction as explicitly.} Prominent examples include mendelberg2014does, who assigned individuals to deliberation groups varying by gender composition and decision rule, and iacovone2022improving, who assigned business owners to individual or group-based consulting.

Despite similar designs, analytical approaches often differ. mendelberg2014does used cluster-robust standard errors, appealing to the intuition that group interaction creates dependence, as others have on similar grounds pals2008individually, candel2025efficient, sometimes via multilevel models lohr2014partially. iacovone2022improving, by contrast, relied on individual-level randomization to justify inference without clustering abadie2023should. This discrepancy reflects a broader confusion---spanning education, psychology, and economics---about whether individual assignment to groups warrants clustering when treatments involve interaction.\footnote{A recent Bluesky exchange with prominent applied microeconomists exemplifies issues that this paper tries to clarify: \url{https://bsky.app/profile/seema.bsky.social/post/3lbi3yzytfk2g}}

To our knowledge, this paper provides the first design-based analysis of inference for group interaction experiments. Group interaction implies the potential for {\it interference} such that an individual's potential outcomes may depend on {\it how others are assigned} to groups and treatments aronow2021spillover. This represents a departure from the standard “stable unit treatment value assumption” (SUTVA) upon which standard inferential results are based imbens2015causal. To build up insight on the factors that affect inference in such experiments, we consider scenarios in which interference is present or not, and when groups are randomly assigned or fixed. We define valid estimands in each of these scenarios, using the concept of average marginalized exposure effects when interference is present. Marginalized exposure effects are a now-common concept in the literature on causal inference with interference aronow2021spillover, hudgens2008toward, li2019randomization, hu2022average, savje2021average. They refer to effects defined in terms of “averages of averages”---that is, the population average of individual-level averages over different potential exposures.

Next, we derive convergence results for commonly used treatment effect estimators---namely, difference-in-means, difference in inverse probability weighted (IPW) means, and covariate-adjusted treatment effects regression---in relation to the estimands in each scenario. These results rest on a sparse-sampling asymptotic regime, in which the group sizes remain fixed and the individual-level sampling proportion goes to zero at a fast rate. Generalization to covariate-controlled regression estimators and generalized linear model predictions follows naturally. This regime operates similarly to the superpopulation regime of bai2022optimality and bai2024primer for analyzing experimental designs under SUTVA.

Building up to analyzing the implications of interference and random group formation, we first consider {\it fixed} groups (rather than randomly constructed ones) in the {\it absence} of interference. In this case, the results are essentially identical to what abadie2023should and su2021model have derived: the treatment effect estimators given above and cluster-robust standard errors are consistent for inference on the average treatment effect (ATE) in cluster-randomized trials. When we instead introduce {\it random group formation} while maintaining no interference, the distribution of the treatment effect estimators match what one obtains from an individual-level randomized trial targeting the ATE. This case corresponds to the intuition, invoked by some authors li2019randomization, iacovone2022improving, that the data can be analyzed on the basis of “individual random assignment.” The two cases differ in their variance: it is larger under fixed groups than under random group formation and, in the former case, increases with the group-level intraclass correlation induced by outcome homophily.

However, interference changes things. With {\it fixed} groups and interference, the treatment effect estimators are consistent for an {\it average marginalized exposure effect}, analogous to the “average total effect” of hudgens2008toward under full group-level treatment saturation. Interference adds a second source of intraclass correlation, on top of any already induced by outcome homophily in the pre-existing groups. Because the cluster-robust standard error captures within-group dependence whatever its source, it remains consistent for inference on this effect under our sparse-sampling regime. Finally, with random allocation into groups, the homophily-induced intraclass correlation is absent by design, but the one generated by interference still presents. Treatment effect estimators now target an average marginalized exposure effect, averaging over potential outcomes under different groupings of individuals. The cluster-robust standard error is consistent for inference on this estimand, while the heteroskedasticity-robust one is not. A key take-away from our analysis is that, when interference is present, the advice to “cluster at the level of randomization” can be misleading. Rather, what is required is to cluster at the level of exposure variation, and this depends on the full list of potential outcome inputs.

We further derive the asymptotic distribution of the treatment effect estimators under random group formation, both with and without interference, through a new coupling argument hajek1960limiting, polyanskiy2025information. We view random group formation as sampling from the collection of all possible groups that could be formed from the population, subject to the restriction that overlapping groups cannot be sampled together. We then show that this restricted sampling process can be approximated by one that draws groups independently, so that a standard central limit theorem applies at the group level ohlsson1989asymptotic. This approach may be of independent interest for design-based inference in other complex designs. Taken together, our analysis unifies the four design--outcome cases within a single framework and provides valid inference for each.

The sections that follow explicate these results formally. Section (ref) lays out the inferential setting, defining the design and outcome cases. Section (ref) introduces the estimands, along with the treatment effect estimators and their variance estimators. Section (ref) presents our main result, characterizing the limiting distribution of these estimators and the validity of each variance estimator across the design and outcome cases, as well as a test for interference based on the within-group intraclass correlation. Section (ref) presents evidence from a simulation study to validate the results. Finally, Section (ref) applies these methods to the mendelberg2014does and iacovone2022improving experiments, illustrating their practical importance, and Section (ref) concludes.

Inferential settings

When treatments are administered in group settings, units may interact within their groups. Such interaction means that a unit's potential outcomes may depend on who else is assigned to its group---a form of interference cox1958planning, sobel2006randomized. Group formation may lie outside the experimenter's control and be homophily-based, in which case potential outcomes typically exhibit within-group (intraclass) correlation even before any interaction; interaction can then add further correlation after treatment. Alternatively, the experiment may assign group membership at random. This latter case is our main focus, but we also treat (i) fixed groups and (ii) settings without interference, to clarify the distinct implications of random group formation, interference, and their combination. Throughout the paper, we use upper-case letters for random variables, lower-case letters for their realizations, and calligraphic letters for sets.

Experimental designs

Let $\mathcal{U}$ be a reference population of $n = |\mathcal{U}|$ units. We consider two research designs, distinguished by whether groups are fixed or randomly formed.

description• Units in the population belong to pre-existing groups. Using simple random sampling without replacement, we draw $G_N$ groups in sequence from $\mathcal{U}$; call this sample $\mathcal{N}$, with $|\mathcal{N}| = N$. We index the sampled groups $g=1,...,G_N$, assigning $g=1,...,G_{N1}$ to “treatment” and $g=G_{N1}+1,...,G_N$ to “control.” Treatment status is captured by a group treatment indicator $Z_g \in \{0,1\}$, equal to $1$ for treated groups and $0$ for control groups. Each group $g$ has size $M_g$, so the sample contains $N \equiv \sum_{g=1}^{G_N} M_g$ units, of which $N_1$ are under treatment and $N_0$ under control. • The design is intended to construct $G_N$ randomly formed groups, indexed by $g=1,...,G_N$. Each group $g$ has a pre-specified target size $M_g \in \mathcal{M} = \{m_1,...,m_K\}$, where $\mathcal{M}$ is the set of $K$ distinct admissible group sizes and $m_k$ is the $k$-th such size. The simplest case is $K=1$ with $M_g = M$ for all $g$ (homogeneous group sizes).\footnote{With a slight abuse of notation, we also use $M$ to denote the single admissible group size in the homogeneous case.} Using simple random sampling without replacement, we draw $N \equiv \sum_{g=1}^{G_N} M_g$ {\it individual} units from $\mathcal{U}$, call this sample $\mathcal{N}$, and randomly partition them into the $G_N$ groups. The partition is created through a random permutation of an index vector of length $N$ of the form $\underbrace{1,...,1}_{M_1},...,\underbrace{g,...,g}_{M_g},...,\underbrace{G_N,...,G_N}_{M_{G_N}}$, in which each index value $g$ appears $M_g$ times, so that each unit is randomly assigned to a group. Groups $g=1,...,G_{N1}$ are then assigned to “treatment” and $g=G_{N1}+1,...,G_N$ to “control,” with group treatment indicator $Z_g \in \{0,1\}$ equal to $1$ for treated groups and $0$ for control groups. We write $G_N(m)$ for the number of groups with $M_g = m \in \mathcal{M}$, $G_{N1}(m)$ for the number of these assigned to treatment, and $G_{N0}(m) = G_N(m) - G_{N1}(m)$ for the number assigned to control. Because group sizes can differ across arms, we let $\mathcal{M}_1$ and $\mathcal{M}_0$ denote the distinct group sizes appearing in the treatment and control arms, respectively; under homogeneous group sizes, both reduce to the same singleton. $N_1$ and $N_0$ denote the numbers of units under treatment and control, respectively, while $N_z(m)$ denotes the number of units in size-$m$ groups assigned to treatment status $z$.

Given random sampling, Design Case 0 is equivalent in distribution (by exchangeability of the sampling sequence) to a design in which we first randomly sample $G_N$ groups and then randomly assign them to treatment or control. The same would hold for Design Case 1 (by exchangeability of the sampling and permutation orders) if group sizes were homogeneous ($M_g = M$ for all $g$); but our description allows group sizes to vary, as in both of the motivating examples from mendelberg2014does and iacovone2022improving. The treatment and control arms may contain different group sizes, as in iacovone2022improving, where control units are left ungrouped, each forming a singleton ($\mathcal{M}_0 = \{1\}$).

In Design Case 1, the number of groups of each size, $G_N(m)$, is itself a design choice that researchers can set for different purposes---for instance, to equalize the number of groups across sizes ($G_N(m)$ constant), or to equalize the number of units across sizes ($G_N(m) \propto 1/m$). We have described complete randomization of groups to treatment and control for illustration; the framework also accommodates more sophisticated assignment mechanisms, such as randomizing treatment separately within each group size (i.e., blocking on group size), which the varying-size results below handle through size-specific treatment probabilities.

For both designs, let $\mathcal{A}_g$ denote the set of units in group $g$, with $|\mathcal{A}_g| = M_g$. For any unit $i$ in the sample, we let $A_i$ denote the index of the group to which unit $i$ belongs, so that $A_i \in \{1,...,G_N\}$ and $A_i = g$ if and only if $i \in \mathcal{A}_g$, or equivalently, $$ A_i = \sum_{g=1}^{G_N} g \cdot \mathbf{1}\{i \in \mathcal{A}_g\}. $$ The variable $A_i$ thus captures the random group assignment induced by the sampling and, in Design Case 1, the partitioning. When groups are fixed (Design Case 0), $\mathcal{A}_g$ is the pre-existing group that receives index $g$ based on the order in which groups were sampled; when groups are formed randomly (Design Case 1), $\mathcal{A}_g$ is the set of units assigned to position $g$ by the partitioning process. Unit $i$'s treatment status is that of its group, $Z_{A_i} \in \{0,1\}$.

Potential outcomes

For each $i \in \mathcal{U}$ define the potential outcome $Y_i\left(z, \mathcal{A}\right) \in \mathbb{R}$, where $z \in \{0,1\}$ is the treatment status of $i$'s group and $\mathcal{A} \subseteq \mathcal{U}$ with $i \in \mathcal{A}$ is the group to which $i$ belongs. This is a general definition of potential outcomes that allows for arbitrary heterogeneity on the basis of treatment and group membership. Below, we will impose restrictions on the potential outcomes to define the cases with and without interference. First, we assume the following high level assumptions:

assnFor all $i$ in $\mathcal{U}$ and all assignment conditions, we have \begin{enumerate} • (bounded potential outcomes) $|Y_i(z, \mathcal{A})| < C$ for all $z \in \{0,1\}$ and $\mathcal{A} \ni i$, where $C\in \mathbb{R_+}$, and • (lower bound on variance) The distribution of $Y_i(z, \mathcal{A})$ in $\mathcal{U}$ is non-degenerate such that sample variances of potential outcomes are bounded away from zero. \end{enumerate}

Bounded potential outcomes ensures fast convergence rates of sample statistics. The non-degeneracy condition ensures that true sampling and randomization variances are bounded away from zero.

We consider two types of assumptions on interference, which have implications for the potential outcomes and the observed outcomes:

description• For every $i \in \mathcal{U}$, potential outcomes depend only on the realized group-level treatment assignment for unit $i$; that is, $Y_i(z, \mathcal{A}) = Y_i(z)$ for all $\mathcal{A} \ni i$, and for each individual in the sample, we observe the outcome $Y_i = Y_i(Z_{A_i})$. • For every $i \in \mathcal{U}$, potential outcomes depend on both the group-level treatment assignment and which other individuals are assigned to the same group as unit $i$; that is, $Y_i(z, \mathcal{A})$ may vary across $\mathcal{A} \ni i$, and for each individual in the sample, we observe the outcome $Y_i = Y_i(Z_{A_i}, \mathcal{A}_{A_i})$.

The dependence of $Y_i(z, \mathcal{A})$ on $\mathcal{A}$ lets unit $i$'s potential outcomes vary with which other units share its group---a form of interference cox1958planning, sobel2006randomized restricted to within groups. This restriction is similar to partial interference of hudgens2008toward, but the source of interference differs: Hudgens and Halloran fix the group structure and vary the within-group treatment allocation, whereas we fix treatment within groups and let group composition vary randomly, allowing for group interaction, outcome-based contagion, or other composition-driven heterogeneity. Table (ref) summarizes the four cases that arise under the different design and outcome assumptions, which our analysis treats in turn.

table[table omitted — 469 chars of source]

Illustrations

Figure (ref) provides a toy illustration of Design Case 1 with $N=4$ sampled units, $G_N = 2$ groups, and $M_1 = M_0 = 2$ units per group. Panel A shows how the group index vector $(1,1,2,2)$ is randomly permuted across the four unit positions, yielding $\binom{4}{2}=6$ equally likely partitions. Panel B shows, for unit $i=1$, the potential outcome inputs under each permutation: without interference, the potential outcome depends only on the group-level treatment status $z_{a_1}$; with group interference, it also depends on the identity of unit 1's groupmate. The combination of group-level treatment status and groupmate assignments constitutes the “exposures” that the design generates aronow2017estimating.

figure[figure omitted — 2,533 chars of source]

Table (ref) shows the experimental design table from mendelberg2014does, which illustrates several features that our setup covers. Because the outcomes of interest concern women's deliberative behavior, we take the population of analysis to be all women and view the design as one that partitions these women into groups of varying size and then assigns each group a treatment---an instance of random group formation (Design Case 1). The group set $\mathcal{A}$ for a woman is the set of women in her deliberation group, and its size is the number of women that group contains. Since a five-person group may hold anywhere from zero to five women, there are six possible sizes; the all-male groups contain no member of the population of analysis and drop out, leaving five admissible group sizes, $\mathcal{M} = \{1, 2, 3, 4, 5\}$. Each row of Table (ref) collects the groups of one size, while the decision rule---the group-level treatment status $z$---is given by the first two columns. A cell of the table therefore corresponds to a size-by-treatment combination, each restricting the set of $(z, \mathcal{A})$ conditions to which a woman could be assigned.

Whether a woman's outcome is sensitive to the group set $\mathcal{A}$ is precisely the distinction between our two outcome cases. Under no interference (Outcome Case 0), her potential outcome depends only on the decision rule assigned to her group, so which other women share her group---and how many---is immaterial. Under group interference (Outcome Case 1), the potential outcome $Y_i(z, \mathcal{A})$ depends on the group set as well, so two women assigned the same decision rule but with different group partners (whether in terms of the identities of the group members or the number of group members) occupy different exposure conditions.

table[table omitted — 640 chars of source]

Estimands and estimators

We consider a variety of estimands motivated by the design and outcome cases, all defined with respect to the reference population $\mathcal{U}$. For any set of units $\mathcal{S}$, let $\mathcal{G}_{\mathcal{S}}(m) = \{\mathcal{A} \subseteq \mathcal{S} : |\mathcal{A}| = m\}$ denote the collection of all $m$-member groups that can be formed from units in $\mathcal{S}$, and let $\mathcal{G}_{\mathcal{S},i}(m) = \{\mathcal{A} \subseteq \mathcal{S} : i \in \mathcal{A},\ |\mathcal{A}| = m\}$ denote those that can be formed from $\mathcal{S}$ and contain unit $i$. We index a generic group from the population by $\omega$, whether fixed or randomly formed---the population-level counterpart of the index $g$ used for groups in the sample---and write $\mathcal{A}_\omega$ for its set of members.

Estimands under fixed groups

When SUTVA imbens2015causal holds (no interference), as in Outcome Case 0, we target the population average treatment effect, $$ \mathrm{ATE} := \frac{1}{n} \sum_{i \in \mathcal{U}}\left[Y_i(1)-Y_i(0)\right]. $$ Under group interference (Outcome Case 1), this expression is no longer well-defined: it is written in terms of $Y_i(1)$ and $Y_i(0)$, potential outcomes indexed by treatment alone, whereas interference makes a unit's outcome depend on its group composition as well. Recall that the general potential outcome is $Y_i(z, \mathcal{A})$, indexed by both the group's treatment status $z$ and its membership $\mathcal{A}$. Holding $i$'s group fixed at $\mathcal{A}_{\omega_i}$ while switching the group's treatment status defines the within-group contrast $$ Y_i(1, \mathcal{A}_{\omega_i}) - Y_i(0, \mathcal{A}_{\omega_i}). $$ This contrast is analogous to what hudgens2008toward call the individual-level “total effect” under full ($100\%$) versus zero ($0\%$) group-level treatment saturation, given the group $\mathcal{A}_{\omega_i}$. The difference accounts for arbitrary interference between $i$ and the other units in $\mathcal{A}_{\omega_i}$. Averaging over the population gives the population average total effect, $$ \mathrm{TOT} := \frac{1}{n}\sum_{i \in \mathcal{U}} \left[Y_i(1, \mathcal{A}_{\omega_i}) - Y_i(0, \mathcal{A}_{\omega_i})\right]. $$

Estimands under random group formation

Allowing for random group formation requires that we consider potential outcomes for any group to which unit $i$ might belong. We first define the marginalized potential outcome under homogeneous group sizes and then describe how it extends to varying sizes.

\oldparagraph{Homogeneous group sizes.} When all groups have a common size $M$, each group in $\mathcal{G}_{\mathcal{U},i}(M)$ is determined by its $M-1$ members besides $i$, chosen from the $n-1$ other units in $\mathcal{U}$, so $|\mathcal{G}_{\mathcal{U},i}(M)| = \binom{n-1}{M-1}$. The individualistic marginalized potential outcome for unit $i$ averages its potential outcome uniformly over these groups: $$ \mu_i(z) = \frac{1}{\binom{n-1}{M-1}} \sum_{\mathcal{A} \in \mathcal{G}_{\mathcal{U},i}(M)} Y_i(z, \mathcal{A}). $$ The individual marginalized effect is $\tau_i = \mu_i(1) - \mu_i(0)$, and the population average marginalized effect (PAME) is $$ \tau_{\mathrm{PAME}} = \frac{1}{n} \sum_{i \in \mathcal{U}} \tau_i. $$ The PAME aggregates these unit-level effects, each marginalized over the variation in group composition that random group formation induces. Because that design makes $\mathcal{A}_{A_i}$ uniform over $\mathcal{G}_{\mathcal{U},i}(M)$ and independent of $Z_i$, the marginalized potential outcome coincides with the design-conditional expectation, $\mu_i(z) = \text{E}\,[Y_i(z, \mathcal{A}_{A_i}) \mid Z_i = z]$.

\oldparagraph{Varying group sizes.} When group sizes vary within an arm, the two characterizations of $\mu_i(z)$---as a uniform average over candidate groupings and as the design-conditional expectation $\text{E}\,[Y_i \mid Z_i = z]$---do not always coincide. Generalizing the PAME therefore requires a choice about how to weight potential outcomes across group sizes, and the right choice depends on substantive considerations.

We first define the within-size-$m_k$ marginalized potential outcome under treatment status $z$: $$ \mu_i(z; m_k) = \frac{1}{\binom{n-1}{m_k-1}}\sum_{\mathcal{A} \in \mathcal{G}_{\mathcal{U},i}(m_k)} Y_i(z, \mathcal{A}), $$ which reduces to $\mu_i(z)$ when group sizes are homogeneous. Recall that $\mathcal{M}_z$ is the set of admissible group sizes in arm $z$; for instance, if groups of size 2, 3, or 4 are admissible under treatment while only singletons are admissible under control, then $\mathcal{M}_1 = \{2,3,4\}$ and $\mathcal{M}_0 = \{1\}$. We combine the within-size outcomes through a weighted average, $$ \bar\mu^{(\phi)}_i(z) = \sum_{m_k \in \mathcal{M}_z}\phi(z;m_k)\, \mu_i(z;m_k), $$ where the weights $\phi(z;m_k) \ge 0$ sum to one over $m_k \in \mathcal{M}_z$ and encode how much each admissible size counts toward the estimand. The unit-level marginalized exposure effect under these weights is $\tau^{(\phi)}_i = \bar\mu^{(\phi)}_i(1) - \bar\mu^{(\phi)}_i(0)$, and the population average marginalized effect under $\phi$ is $$

aligned\tau^{(\phi)}_{PAME} = & \frac{1}{n}\sum_{i \in \mathcal{U}}\tau^{(\phi)}_i = \frac{1}{n}\sum_{i \in \mathcal{U}}\left(\bar\mu^{(\phi)}_i(1) - \bar\mu^{(\phi)}_i(0)\right) \\ = & \sum_{m_k \in \mathcal{M}_1}\phi(1;m_k) \mu(1;m_k) - \sum_{m_k \in \mathcal{M}_0}\phi(0;m_k) \mu(0;m_k),

$$ where $\mu(z;m_k) = \frac{1}{n}\sum_{i \in \mathcal{U}} \mu_i(z;m_k)$. Different choices of $\phi$ give different forms of this single estimand; we highlight two of them.

The first weights each size by the design's own assignment probability, $\phi(z;m_k) = \Pr(|\mathcal{A}_{A_i}| = m_k \mid Z_i = z) = N_z(m_k) / N_z$, the share of arm-$z$ units the design places in size-$m_k$ groups. Under this choice $\bar\mu^{(\phi)}_i(z) = \text{E}\,[Y_i(z, \mathcal{A}_{A_i}) \mid Z_i = z]$ is exactly the design-conditional expectation from above, and $\tau^{(\phi)}_{PAME}$ is the probability limit of the unweighted difference in means. Its drawback is that this form is not a stable population quantity: it depends on the experimental assignment probabilities, so adjusting the design to change how often individuals land in groups of different sizes changes the target.

The second weights all admissible sizes evenly, $\phi(z;m_k) = 1/|\mathcal{M}_z|$. This form is stable: since the weights ignore the design's allocation across sizes, $\tau^{(\phi)}_{PAME}$ is unchanged by that allocation. It is natural when the researcher regards each group size as a substantively meaningful condition deserving equal weight. The two forms coincide under homogeneous group sizes, and more generally whenever the design assigns the same number of units to each admissible size within each arm, so that the design shares are themselves uniform. Other weightings are possible and may be preferable in particular applications; Appendix A.2 develops the results for a general $\phi$.

Estimators

We begin with the difference-in-means estimator, which we use across most of the design and outcome cases, and then generalize it to target the weighted PAME. Let $\bar Y_z = N_z^{-1}\sum_{i \in \mathcal{N}}\mathbf{1}\{Z_i = z\}Y_i$ be the mean outcome of arm-$z$ units. The difference in means is $\bar Y_1 - \bar Y_0$. When the treatment probability varies across groups, we replace it with the more general IPW or H\'ajek estimator. Details are deferred to the Appendix.

To target $\tau^{(\phi)}_{PAME}$ under weights $\phi$, we first estimate, for each arm $z$ and admissible size $m_k \in \mathcal{M}_z$, the within-size mean outcome $$ \hat\mu(z;m_k) = \frac{\sum_{i \in \mathcal{N}}\mathbf{1}\{|\mathcal{A}_{A_i}| = m_k\}\,\mathbf{1}\{Z_i = z\}\,Y_i}{G_{Nz}(m_k)\,m_k}. $$ The plug-in estimator for $\tau^{(\phi)}_{PAME}$ is then $$ \hat\tau_N = \sum_{m_k \in \mathcal{M}_1}\phi(1;m_k)\,\hat\mu(1;m_k) \;-\; \sum_{m_k \in \mathcal{M}_0}\phi(0;m_k)\,\hat\mu(0;m_k). $$ Setting $\phi$ to the realized design shares $\hat\phi(z;m_k) = N_z(m_k)/N_z$ makes the within-size means aggregate back to the arm means, so $\hat\tau_N = \bar Y_1 - \bar Y_0$ recovers the difference in means and targets the design-weighted form. Setting $\phi(z;m_k) = 1/|\mathcal{M}_z|$ instead targets the evenly-weighted form from the previous subsection.

The estimator $\hat\tau_N$ also has a regression representation: it can be computed as a (weighted) least squares regression of $Y_i$ on treatment, and covariate adjustment enters as in lin2013agnostic, with the regression estimator interpretable as a (weighted) difference in adjusted means.

\oldparagraph{Variance estimators.} We also need variance estimators for $\hat\tau_N$. We give them first for the difference in means $\bar Y_1 - \bar Y_0$, which is $\hat\tau_N$ under homogeneous group sizes, and then generalize. The heteroskedasticity-robust (“HC2” or Neyman) estimator splawa1990application, samii2012equivalencies, abadie2023should and the cluster-robust (“CR2”) estimator pustejovsky2018small are, respectively, $$

aligned\hat V^{HR}_N &= \frac{\sum_{i \in \mathcal{N}}Z_i(Y_i - \bar Y_1)^2}{N_1(N_1-1)} + \frac{\sum_{i \in \mathcal{N}}(1-Z_i)(Y_i - \bar Y_0)^2}{N_0(N_0-1)}, \\ \hat V^{CR}_N &= \sum_{g: Z_g=1}\frac{M_g^2(\bar Y_g - \bar Y_1)^2}{N_1(N_1 - M_g)} + \sum_{g: Z_g=0}\frac{M_g^2(\bar Y_g - \bar Y_0)^2}{N_0(N_0 - M_g)},

$$ with $\bar Y_g = M_g^{-1}\sum_{i \in \mathcal{A}_g}Y_i$ the sample mean of group $g$. When clusters are of equal size within each arm, $\hat V^{CR}_N$ reduces further to the cluster-level HC2 analogue \citep{su2021model}, $\hat V^{CR}_N = G_{N1}^{-1}(G_{N1}-1)^{-1}\sum_{g:Z_g=1}(\bar Y_g - \bar Y_1)^2 + G_{N0}^{-1}(G_{N0}-1)^{-1}\sum_{g:Z_g=0}(\bar Y_g - \bar Y_0)^2$. The two are also the sandwich variance estimators from regressing $Y_i$ on $Z_i$, at the unit and group level respectively, and differ only in whether the within-group cross-products of outcome residuals are retained: keeping them is exactly how the cluster-robust estimator accounts for group-level intra-cluster correlation, a point to which we return below.

With varying group sizes the same construction applies cell by cell. Random sampling and random assignment render the within-arm, within-size means $\hat\mu(z;m_k)$ independent across cells as the population goes to infinity. Thus, the variance of $\hat\tau_N$ obeys the following as the population size goes to infinity: $$ \frac{\sum_{m_k \in \mathcal{M}_1}\phi(1;m_k)^2\,\text{Var}\,[\hat\mu(1;m_k)] + \sum_{m_k \in \mathcal{M}_0}\phi(0;m_k)^2\,\text{Var}\,[\hat\mu(0;m_k)]}{\text{Var}\,[\hat\tau_N]} \to 1 \text{ as } n \to \infty, $$ with cells defined separately within each arm, so the two arms may differ in their size structure. We estimate each cell variance $\text{Var}\,[\hat\mu(z;m_k)]$ by the within-cell analogue of the heteroskedasticity-robust or cluster-robust estimator above, aggregating with the weights $\phi(z;m_k)^2$ that define $\hat\tau_N$. Which of the two is appropriate---in particular, when the heteroskedasticity-robust form remains valid---is the subject of the next section. We give the exact per-cell forms, together with the weighted-least-squares sandwich representation that yields both, in Appendix A.2.

Theory

Our inferential conclusions are based on a sparse-sampling asymptotic regime in which the admissible group sizes $\mathcal{M}$ are fixed, $n, N, G_N \to \infty$, and $\frac{N^2}{n} \to 0$. This regime yields results similar to those of the design-based superpopulation sampling analysis of bai2022optimality and bai2024primer. In either case, the goal is to characterize the relationship between what transpired in a particular experiment and a more general population.

We begin by stating our two main results. Theorem (ref) establishes the asymptotic normality of $\hat\tau_N$ in all four cases, and Theorem (ref) characterizes the two variance estimators. We then discuss how to construct confidence intervals based on these theorems and conduct inference in practice.

theorem[Asymptotic normality] Given Assumption (ref), the Design Cases indexed by $d \in \{0,1\}$ and Outcome Cases indexed by $o \in \{0,1\}$, and fixed admissible group sizes $\mathcal{M}$, with $n, N, G_N \to \infty$ and $N^2/n \to 0$, the following holds in all four Design--Outcome cases: \[ \frac{\hat\tau_N - \tau_{do}}{\sqrt{\mathrm{Var}[\hat\tau_N]}} \rightsquigarrow \mathcal{N}(0,1), \] where $\tau_{00} = \tau_{10} = \mathrm{ATE}$, $\tau_{01} = \mathrm{TOT}$, and $\tau_{11} = \tau_{\mathrm{PAME}}$ (homogeneous group sizes) or $\tau^{(\phi)}_{\mathrm{PAME}}$ (varying sizes).
theorem[Variance estimation] Under the conditions of Theorem (ref), the cluster-robust estimator is ratio-consistent for $\mathrm{Var}[\hat\tau_N]$ in all four cases, \[ \frac{\hat V^{CR}_N}{\mathrm{Var}[\hat\tau_N]} \xrightarrow{P} 1, \] while the heteroskedasticity-robust estimator $\hat V^{HR}_N$ is ratio-consistent in case 1-0 only.

In every Design--Outcome case except the random-groups, no-interference one (1-0), $\mathrm{Var}[\hat\tau_N]$ carries a within-group (intra-cluster) correlation component. The cluster-robust estimator retains the within-group cross-products and so captures this component whatever its source---homophily in fixed groups (case 0-0) or interference (cases 0-1 and 1-1)---and is therefore consistent throughout. The heteroskedasticity-robust estimator discards those cross-products and is inconsistent for $\mathrm{Var}[\hat\tau_N]$ whenever the correlation is present. Only in case 1-0 does it vanish: there the design reduces to unit-level randomization and the variance collapses to the corresponding unit-level Neyman variance, so the two estimators share a probability limit and both are consistent, with the heteroskedasticity-robust estimator the more efficient of the two. Explicit forms for all cases appear in Appendix A.1 and A.2.

Table (ref) collects the implications of the two theorems. The estimand depends on the cell: any commonly used estimator targets the $\mathrm{ATE}$ in both no-interference cells, the $\mathrm{TOT}$ under fixed groups with interference, and $\tau_{\mathrm{PAME}}$ under random groups with interference (its weighted form $\tau^{(\phi)}_{\mathrm{PAME}}$ when group sizes vary).

table[table omitted — 485 chars of source]

Combining the two theorems through Slutsky's theorem yields the studentized statements used in practice, which we present as a corollary:

corollary[Asymptotic normality of the t-statistic] Under the conditions of Theorem (ref), the following holds in all four Design--Outcome cases: \[ \frac{\hat\tau_N - \tau_{do}}{\sqrt{\hat V^{CR}_N}} \rightsquigarrow \mathcal{N}(0,1), \] where $\tau_{00} = \tau_{10} = \mathrm{ATE}$, $\tau_{01} = \mathrm{TOT}$, and $\tau_{11} = \tau_{\mathrm{PAME}}$ (homogeneous group sizes) or $\tau^{(\phi)}_{\mathrm{PAME}}$ (varying sizes). In case 1-0 only, the same limit holds with $\hat V^{HR}_N$ in place of $\hat V^{CR}_N$.

The recommended $(1-\alpha)$ confidence interval is \[ CI^{BM} = \left(\hat\tau_N + t^{K_{BM}}_{\alpha/2}\sqrt{\hat V_N},\; \hat\tau_N + t^{K_{BM}}_{1-\alpha/2}\sqrt{\hat V_N}\right), \] where $K_{BM}$ is the Bell--McCaffrey degrees-of-freedom adjustment of imbens2016robust, and $\hat V_N = \hat V^{CR}_N$ in all four cases, with $\hat V_N = \hat V^{HR}_N$ available in case 1-0.

Sketch of the argument with homogeneous group sizes

We provide a sketch of the proofs for Theorems (ref) and (ref) in the simplest case, in which all groups share a common size $M = N/G_N$. We focus on the novel combination of Design Case 1 and Outcome Case 1, for which the difference in means is the appropriate estimator, and indicate how the remaining design--outcome combinations follow as special cases. The full arguments appear in Appendix A.1, and the generalization to varying group sizes is taken up in the next subsection. The analysis proceeds in five steps: (i) a group-level representation of the design, (ii) unbiasedness of the difference in means, (iii) the group-level variance and its estimator, (iv) the collapse to unit-level randomization under no interference, and (v) asymptotic normality via a coupling argument. We write $a_N \simeq b_N$ to mean that $a_N$ and $b_N$ share the same limit.

Step 1: a group-level representation of the design. The key analytical move is to recast a random group formation experiment as a sampling design over a population of potential groups that could be formed by the individuals in the population. Recall that $\mathcal{G}_{\mathcal{U}}(M)$ denotes the collection of all size-$M$ subsets of $\mathcal{U}$, with $|\mathcal{G}_{\mathcal{U}}(M)| = \binom{n}{M}$. Sequential random sampling of units followed by random group formation (Design Case 1) is equivalent in distribution to a restricted group-level design in which $G_N$ non-overlapping members of $\mathcal{G}_{\mathcal{U}}(M)$ are drawn without replacement and then assigned to treatment and control as in a classical cluster-randomized experiment. Writing $W_\omega = 1$ if potential group $\omega$ is realized, each group enters with equal marginal probability $\text{E}\,[W_\omega] = \pi_N = G_N/|\mathcal{G}_{\mathcal{U}}(M)|$, and the only departure from independent (with-replacement) sampling is that two overlapping groups can never be drawn together. That restriction is immaterial as the population size goes to infinity: the share of groups overlapping any fixed $\mathcal{A}_\omega$ is $1 - \binom{n-M}{M}/\binom{n}{M} \simeq M^2/n \to 0$ as $n \to \infty$ with $M$ fixed, so realized groups are sampled as if uniformly and independently from $\mathcal{G}_{\mathcal{U}}(M)$. This large-population independence is the feature that drives every result below. Because all members of a realized group share a treatment status and composition, the estimand reduces to a simple average of potential-group effects, \[ \tau_{PAME} = \frac{1}{|\mathcal{G}_{\mathcal{U}}(M)|}\sum_{\omega=1}^{|\mathcal{G}_{\mathcal{U}}(M)|}\bigl[\bar{Y}_\omega(1) - \bar{Y}_\omega(0)\bigr], \qquad \bar{Y}_\omega(z) = \frac{1}{M}\sum_{i\in\mathcal{A}_\omega}Y_i(z,\mathcal{A}_\omega). \]

Step 2: unbiasedness. In this representation, unbiasedness of $\hat \tau_N$ is almost immediate. Conditional on the realized groups, $\hat{\tau}_N$ is the classical difference-in-means estimator applied to group-level outcomes, hence unbiased for the sample average treatment effect at the group level. Averaging over the sampling of potential groups then gives $\text{E}\,[\hat{\tau}_N] = \tau_{PAME}$. The same conclusion applies at the unit level, which explains why the difference in means targets a marginalized estimand: under Design Case 1 the group realized for unit $i$ is uniform over its $\binom{n-1}{M-1}$ candidate size-$M$ groups, so $\text{E}\,[Y_i \mid Z_i = z] = \mu_i(z)$, the individualistic marginalized potential outcome, and the difference in means therefore targets $n^{-1}\sum_i[\mu_i(1)-\mu_i(0)] = \tau_{PAME}$. The other cases follow by the similar sampling logic: under Design Case 0 with interference $\hat{\tau}_N$ is unbiased for the $TOT$ given random sampling of the fixed groups, and under either design with no interference it is unbiased for the $ATE$---for fixed groups because the group sizes are common\footnote{With varying sizes, the statement weakens to consistency as $n, N \to \infty$su2021model.}, and for random groups because random formation makes the group-level assignment independent of potential outcomes, mimicking an individual-level trial.

Step 3: the group-level variance and its estimator. Working at the group level yields a clean variance and a natural estimator for it. Decomposing $\text{Var}\,[\hat{\tau}]$ by the law of total variance into a treatment-assignment component and a group-sampling component, the cross term and the sampling-overlap contribution are negligible under the sparse-sampling regime, leaving \[ G_N \text{Var}\,[\hat{\tau}] \simeq \frac{1}{p}\,\frac{1}{|\mathcal{G}_{\mathcal{U}}(M)|}\sum_{\omega}\bigl(\bar{Y}_\omega(1)-\bar{Y}(1)\bigr)^2 + \frac{1}{1-p}\,\frac{1}{|\mathcal{G}_{\mathcal{U}}(M)|}\sum_{\omega}\bigl(\bar{Y}_\omega(0)-\bar{Y}(0)\bigr)^2. \] This is exactly the limit of the Neyman variance of a completely randomized experiment run on group-level outcomes imbens2015causal. The group-level HC2 estimator $\hat{V}^{CR}_N$ retains the within-group cross-products and is therefore consistent for the true variance. At the unit level, the cluster-robust estimator is equal in value to the group-level HC2 estimator, and is therefore consistent, whereas the heteroskedasticity-robust $\hat{V}^{HR}_N$ discards those cross-products and is not consistent.

Step 4: collapse to unit-level randomization under no interference. Under Outcome Case 0, groupmates do not affect potential outcomes, and Design Case 1 becomes a completely randomized experiment at the unit level, so the group-level variance collapses to the unit-level Neyman variance. To see this, write the group-level sampling variance through an intra-cluster decomposition, $G_{Nz}\text{Var}\,[\bar{Y}_z] \simeq M^{-1} S^2_{\mathcal{G}}\,(1 + \rho_{\mathcal{G}})$, where $S^2_{\mathcal{G}}$ is the potential-group analogue of the population variance of potential outcomes and $\rho_{\mathcal{G}}$ is the potential-group intra-cluster correlation. With no interference the deviations of $Y_i(z)$ about its population mean sum to zero across $\mathcal{U}$, and a short counting argument over the $\binom{n-1}{M-1}$ and $\binom{n-2}{M-2}$ groups containing a given unit or pair (Appendix A.1) gives \[ \rho_{\mathcal{G}} = -\frac{M-1}{n-1} \to 0 \text{ as } n \to \infty. \] There is thus no substantive within-group clustering of outcomes under no interference, only a negligible $O(1/n)$ negative correlation induced by the finite population. As a result, the heteroskedasticity-robust estimator becomes consistent and is the more efficient of the two; together with the cluster-robust consistency established in Step 3, this completes Theorem (ref). The vanishing of $\rho_{\mathcal{G}}$ is also what the ICC-based interference test developed below exploits.

Step 5: asymptotic normality via coupling. We now turn to the limiting distribution of $\hat \tau_N$. The difficulty is that the restricted design's inclusion indicators $W_\omega$ are dependent, so a central limit theorem for independent draws does not apply directly. We instead couple the restricted design to an auxiliary one that draws $G_N$ groups independently and with replacement from $\mathcal{G}_{\mathcal{U}}(M)$. Writing $Q$ and $Q^*$ for the laws of the ordered $G_N$-tuple of realized groups under the restricted and independent designs, we have $\mathrm{supp}(Q) \subseteq \mathrm{supp}(Q^*)$, and the total variation distance between them vanishes under $N^2/n \to 0$: \[ d_{TV}(Q, Q^*) = 1 - \prod_{r=0}^{G_N-1}\frac{\binom{n-rM}{M}}{\binom{n}{M}} \simeq \frac{M^2 G_N^2}{2n} \to 0, \] again because the overlapping share is negligible, so the two designs are asymptotically indistinguishable. By the data-processing inequality the same convergence holds for the law of any function of the tuple, in particular the standardized group-level average; since that average is asymptotically normal under the independent design (Lindeberg--Feller, the summands being i.i.d.\ and bounded), the restricted design inherits the same normal limit. Combining this with the ohlsson1989asymptotic two-stage martingale central limit theorem for the assignment stage delivers the conclusion of Theorem (ref).

Generalization to variable group sizes

The five-step argument extends to varying group sizes with two modifications, both developed at length in Appendix A.2.

First, the group-level sampling representation extends to a per-size-class version: the population of potential groups becomes $\bigcup_{k=1}^{K} \mathcal{G}_{\mathcal{U}}(m_k)$, and the design realizes $G_N(m_k)$ groups of size $m_k$ with within-class sampling probability $\pi_N(m_k) = G_N(m_k)/|\mathcal{G}_{\mathcal{U}}(m_k)|$. The variance, central limit theorem, and consistency arguments carry through within each size class, and the simple-average aggregation across classes inherits the within-class asymptotic distributions since sampling across classes is asymptotically independent. The cluster-robust estimator $\hat{V}^{CR}_N$ (in its varying-sizes form above) remains consistent for $\text{Var}\,[\hat\tau_N]$, while the heteroskedasticity-robust estimator $\hat V^{HR}_N$ is consistent only under no interference and random group formation.

Second, which $\tau^{(\phi)}_{PAME}$ an estimator is unbiased for depends on the weighting it places across size classes. The unweighted difference in means $\bar Y_1 - \bar Y_0$ is unbiased for $\tau^{(\phi)}_{PAME}$ at the design-share weights $\phi(z; m_k) = N_z(m_k)/N_z$, while the IPW estimator $\hat\tau_N$ reweights each size class evenly, targeting $\tau^{(\phi)}_{PAME}$ at $\phi(z; m_k) = 1/|\mathcal{M}_z|$ (the second weighting highlighted in Section (ref)). In general, within a size class $m_k$, the permutation-based group allocation in Design Case 1 assigns $\mathcal{A}_{A_i}$ uniformly over the $\binom{n-1}{m_k-1}$ size-$m_k$ groups in the population containing $i$, so $\text{E}\,[Y_i \mid Z_i = z, |\mathcal{A}_{A_i}| = m_k] = \mu_i(z; m_k)$. Different estimators weight these conditional expectations differently and so converge to different members of the $\tau^{(\phi)}_{PAME}$ family.

Testing for interference

When groups are randomly formed but interference is absent, both cluster-robust and heteroskedasticity-robust inference are asymptotically valid, but the heteroskedasticity-robust one is more efficient. It is therefore useful to determine which outcome regime we are in (Outcome Case 0 or 1), since this affects both the validity and the efficiency of inference.

One could simply test for a difference in the variance estimates, but a more direct test works with the intra-cluster correlations (ICCs) of the treated and control outcomes. Under Design Case 1 and Outcome Case 0, randomly formed groups do not generate persistent within-group dependence: within any treatment-by-size cell, the intraclass correlation is asymptotically zero, as in the homogeneous-size calculation yielding $\rho_{\mathcal{G}}=-(M-1)/(n-1)\to 0$. Under Outcome Case 1, by contrast, interaction among members of the realized group $\mathcal{A}_g$ may induce within-group dependence in observed outcomes. We therefore compute ICC statistics within treatment-by-size cells and combine the resulting evidence across cells.\footnote{We cannot pool the treatment and control groups, because the ICC is sensitive to the mean difference induced by the treatment effect. To see this, consider a balanced design with a constant effect and homogenous potential outcomes, such that $Y_i(1) = 10$ and $Y_i(0) = 5$ for all $i$. Within each treatment arm, for any group partition, the between group variation is zero, and so the arm-specific ICCs are zero. However, when we pool the two arms, the between group variation is no longer zero, and so the ICC is greater than zero.} This is not a direct test of interference, but it provides suggestive evidence.

For $z\in\{0,1\}$ and $m_k\in\mathcal{M}_z$, define the treatment-by-size cell \[ \mathcal{C}_{z,m_k} = \left\{ g\in\{1,\ldots,G_N\}: Z_g=z,\; M_g=m_k \right\}, \] so that $G_{Nz}(m_k)=|\mathcal{C}_{z,m_k}|$ and $N_z(m_k)=m_kG_{Nz}(m_k)$. We restrict attention to informative cells \[ \mathcal{C}_{\mathrm{ICC}} = \left\{ (z,m_k): z\in\{0,1\},\; m_k\in\mathcal{M}_z,\; m_k\ge 2,\; G_{Nz}(m_k)\ge 2 \right\}, \] since singleton groups do not contain within-group variation and cells with only one group do not identify between-group variation. For any $(z,m_k)\in\mathcal{C}_{\mathrm{ICC}}$, let $ \bar Y_g = \frac{1}{m_k}\sum_{i\in\mathcal{A}_g}Y_i $ denote the mean outcome in group $g\in\mathcal{C}_{z,m_k}$, and recall that the corresponding cell mean is $ \hat\mu(z;m_k) = \frac{1}{N_z(m_k)} \sum_{g\in\mathcal{C}_{z,m_k}} \sum_{i\in\mathcal{A}_g}Y_i = \frac{1}{G_{Nz}(m_k)} \sum_{g\in\mathcal{C}_{z,m_k}}\bar Y_g . $ Define the between-group and within-group mean squares in cell $(z,m_k)$ as \[ MSB_{z,m_k} = \frac{ m_k\sum_{g\in\mathcal{C}_{z,m_k}} \left(\bar Y_g-\hat\mu(z;m_k)\right)^2 }{ G_{Nz}(m_k)-1 }, \] and \[ MSW_{z,m_k} = \frac{ \sum_{g\in\mathcal{C}_{z,m_k}} \sum_{i\in\mathcal{A}_g} \left(Y_i-\bar Y_g\right)^2 }{ N_z(m_k)-G_{Nz}(m_k) }. \] When $MSW_{z,m_k}>0$, the cell-level ANOVA statistic and ICC estimate are $ F_{z,m_k} = \frac{MSB_{z,m_k}}{MSW_{z,m_k}}, $ and $ \hat\rho_{z,m_k} = \frac{MSB_{z,m_k}-MSW_{z,m_k}} {MSB_{z,m_k}+(m_k-1)MSW_{z,m_k}} = \frac{F_{z,m_k}-1}{F_{z,m_k}+m_k-1} . $

Under the traditional balanced one-way ANOVA model within cell $(z,m_k)$ and the null of zero within-group dependence, $ F_{z,m_k} \sim F_{G_{Nz}(m_k)-1,\;N_z(m_k)-G_{Nz}(m_k)}. $ Without the normality assumptions of the ANOVA model, and with fixed $m_k$ and $G_{Nz}(m_k)\to\infty$, we use the asymptotic approximation $$ T_{z,m_k} = \frac{ \sqrt{G_{Nz}(m_k)}\left(F_{z,m_k}-1\right) }{ \sqrt{2m_k/(m_k-1)} } \rightsquigarrow \mathcal{N}(0,1), $$ or equivalently, $$ \sqrt{\frac{G_{Nz}(m_k)m_k(m_k-1)}{2}}\; \hat\rho_{z,m_k} \rightsquigarrow \mathcal{N}(0,1). $$ Let $p_{z,m_k}$ denote the corresponding upper-tail $p$-value for positive within-group dependence, using either the finite-sample $F$ reference distribution or the normal approximation, $p_{z,m_k}=1-\Phi(T_{z,m_k})$.

To combine evidence across informative cells, let $J=|\mathcal{C}_{\mathrm{ICC}}|$ and define Fisher's combined statistic $$ S_{\mathrm{ICC}} = -2 \sum_{(z,m_k)\in\mathcal{C}_{\mathrm{ICC}}} \log p_{z,m_k}. $$ Asymptotically, $ S_{\mathrm{ICC}} \rightsquigarrow \chi^2_{2J} $ under the null of zero within-cell ICCs. We reject at level $\alpha$ when $ S_{\mathrm{ICC}} \ge \chi^2_{2J;\,1-\alpha}. $ In the homogeneous group-size case, $\mathcal{M}_0=\mathcal{M}_1=\{M\}$ and, provided $M\ge2$ with at least two groups in each arm, $J=2$, so the procedure reduces to combining the two treatment-stratum tests with a $\chi^2_4$ reference distribution.

Simulation study

We use a Monte Carlo simulation study to demonstrate that our convergence results provide reasonable approximations for inference even with a finite sample. The simulation setting is as follows (see the Appendix for a full description of the Monte Carlo specifications):

itemize$|\mathcal{U}|=400,000$, $N \in \{80, 200, 400, 800\}$, $M=4$; each unit $i$ carries a uniformly drawn scalar covariate $X_i$. • Individuals in the {\it population} are already partitioned into groups of size $M$ prior to sampling. The potential outcomes contain a group-level random effect, representing homophily in the population. • When groups are {\it fixed} (no random group formation), the pre-existing groups are sampled, and the within-group dependency carries into the experiment. • When groups are randomly formed, the new groups formed in the experiment do not necessarily correspond to the groups in the population and so the population grouping becomes irrelevant. • When interference is present, {\it individual level treatment effects} are a function of other group members' covariates. • We compute averages over 2,000 Monte Carlo simulation runs.

Table (ref) shows the true estimand, the Monte Carlo mean and variance of the difference in means estimator ($\hat \tau$), and then intraclass correlation of potential outcomes for units in the control groups and treated groups. We see that the estimand ($\tau$) differs across scenarios: $\tau$ corresponds to the PATE in the cases with no interference, the TOT with interference but fixed groups, and the PAME with interference and random groups. The TOT averages over potential groups drawn from the population, whereas PAME averages over potential groups that can be formed from individuals in the population. Our simulation is such that these two averages are nearly the same, although this does not have to be the case with regard to actual populations. The difference in means is unbiased for these estimands.

table[table omitted — 2,313 chars of source]

For untreated control groups, Table (ref) shows substantial intraclass outcome correlation when groups are fixed. This is due to the homophily in the groups that exist in the population. The control group intraclass correlation is zero when groups are randomly formed. Recall that interference in the simulation comes only through treatment effects, which do not pertain to control group observations. This shows how, in the absence of interference, random group formation yields what is essentially an individually-randomized experiment. For the treated groups, when groups are fixed and there is no interference, unit-level heterogeneity in treatment effects makes the intra-class correlation smaller than for the control groups. The introduction of interference with fixed groups brings the intraclass correlation back up. The difference between the intraclass correlation in the first two scenarios is due to the within-group interference that translates into treatment effects. When we have random group formation and no interference, again, there is no group-level intraclass correlation due to homophily, showing that the scenario is equivalent to unit-level randomization without interference. But when interference is reintroduced (the last scenario), this creates a novel source of group-level intraclass correlation.

figure[figure omitted — 437 chars of source]

Figures (ref) and Table (ref) present simulation results for the variance estimators. Figure (ref) compares the performance of the heteroskedasticity-robust (HR2) and cluster-robust (CR2) variance estimators. We observe convergence of the cluster-robust variance estimator to the true variance across all scenarios. By contrast, the heteroskedasticity-robust variance estimator is accurate only in the case with no interference and randomly formed groups. Table (ref) reports the empirical coverage rates of confidence intervals based on HR2 and CR2. We find that the coverage rates using CR2 are very close to the nominal 95% level across all four cases. However, when examining the standard deviations of the variance estimators, we observe that under no interference and randomly formed groups, HR2 is more precise than CR2, as expected.

table[table omitted — 1,944 chars of source]

In the group interaction experiment, the treatment effect is determined by treatment assignment, individual covariates, group members’ covariates, and group-level covariates. As in other experimental analyses, we can collect and control for these covariates in the regression. Figure (ref) presents the results from regression estimators with control variables. We consider several specifications. We begin with no covariates, then add only individual-level covariates, only group-level covariates, and finally both. We also estimate the treatment effect using Lin's regression adjustment lin2013agnostic. The vertical black line marks the estimand, $\tau_{PAME}$, while the colored density curves show the distributions of the estimates. All estimates are centered around the true PAME, and the biases are negligible across specifications. However, the distributions become visibly tighter as more relevant covariates are included, especially under Lin's regression adjustment, where the RMSE is reduced from 0.62 to 0.16.

figure[figure omitted — 287 chars of source]

Table (ref) shows the simulation results for our proposed interference test. The column $\mathbb{P}(p \le 0.05)$ reports the proportion of p-values from the test that are smaller than 0.05. When there is no interference, we expect this value to be close to the nominal level of 0.05, indicating correct Type I error control. When interference is present, we expect this value to be larger and close to 1, indicating high power to detect interference and reject the null hypothesis. Overall, even when the sample size is 200, the test exhibits correct Type I error (0.051) and acceptable power (0.952).

table[table omitted — 813 chars of source]

Application

Application I: gender composition and decision rule on deliberation

We revisit the experiment presented in mendelberg2014does. The experiment randomly assigned individual men and women to 5-person deliberation groups who were instructed to “decide collectively on the `most just' principle of redistribution and set a poverty line in dollars” (p. 4). The groups varied in the share of women versus men (from 0 to 5 women). The groups also varied in whether their decision would be made via majority vote or consensus. The experimental design was shown in Table (ref) above. The main outcome in the paper was the extent to which women raised “care issues” (e.g., concerns about raising children) in the context of the deliberations.

table[table omitted — 2,982 chars of source]

The main analysis in the paper studied the interaction effect between the decision rule (unanimity versus majority vote) and the number of women in the group. It was implemented, for women in mixed groups (those with at least one woman and one man), by regressing the frequency with which a woman raised care issues on an indicator for being in a majority-rule group, the number of women in the group, and their interaction. Recognizing that the group structure might be relevant, mendelberg2014does used cluster-robust standard errors.

The first column in Table (ref) presents the results from this analysis, with standard errors from both the CR2 and HR2 estimators. The difference is pronounced---for the interaction term, the cluster-robust standard error is 23% larger---and tracks the scale of the intraclass correlations shown at the bottom of the table. Because individuals were assigned to groups randomly, this intraclass correlation is attributable to within-group interference. We also conduct the interference test of Section (ref), and the null of no interference is rejected at the 0.001 level. In the second column, we further interact majority rule with indicators for each number of women in the group. The gap between the two is again much larger for the interaction effects: for the group with three women, we may even reach different conclusions about the treatment effect's significance depending on which estimator is used.

The third and fourth results columns illustrate the role of weighting in targeting different marginal PAMEs under varying group sizes. Here we look at the effect of the majority-vote decision rule on another outcome, and one likely to be sensitive to the decision rule: the average time each woman spent in the deliberation. (As is apparent from the coefficients in columns 1 and 2, the marginal PAME is close to zero for the “care issues” outcome.) We marginalize over the different gender compositions. From Table (ref), the number of groups is roughly uniform across the five gender compositions (sizes $m_k = k$ for $k = 1, \dots, 5$ women). The number of women is not, because a composition-$k$ group contributes $k$ women: the 19 single-woman groups supply 19 women, while the 16 four-woman groups supply 64. The assignment is also not perfectly balanced across the unanimity and majority-vote arms. We consider the two estimands discussed in Subsection (ref): the design-weighted PAME ($\phi(z; m_k) = \Pr(|\mathcal{A}_{A_i}| = m_k \mid Z_i = z) = N_z(m_k)/N_z$), which one obtains with an {\it unweighted} regression and the PAME that gives uniform weight to each size class ($\phi(z; m_k) = 1/|\mathcal{M}_z|$), which requires that one weight the regression by the inverse of the probability of appearing in any given size class. Columns 3 and 4 of Table (ref) show the consequences. The estimate for the PAME that gives equal weight to each size class is 17% larger in absolute value than that for the design-weighted PAME. This reflects that treatment effects are of greater magnitude for women in groups with fewer other women (apparent from the interaction effect estimates in column 5), and that groups with fewer women proportionally less weight in the design-weighted estimand.

Application II: improving management with individual and group-based consulting

iacovone2022improving conduct an experiment to test two approaches to improving poor management in the Colombian auto parts manufacturing sector. The first uses intensive, individualized consulting; the second is a group-based approach that delivers consulting at roughly a third of the cost, inspired by agricultural extension, and aims to leverage group-learning dynamics.

In the experiment, a sample of 159 firms was sorted into 53 matched triplets to improve balance on observables, and within each triplet one firm was randomly assigned to each of three arms: a control group, an individual-consulting treatment group, and a group-consulting treatment group, yielding 53 firms per arm. Firms assigned to the individual-consulting treatment received individual support over a period of six months. Treatment began with training in five areas (logistics, human resources, finance, marketing and sales, and production) delivered by five different consultants. This involved a theoretical part aimed at familiarizing the firm's management with modern management concepts and methods, complemented with practical exercises to apply these concepts to their firm, and was then followed by individual consulting to help the firm implement the improvement plan developed during the diagnostic phase.

In the group-consulting treatment, groups were formed with 2 to 8 firms. Leaders from the firms in a group signed an agreement to work together and help each other improve. As in the individual treatment, the group treatment began with training classes covering theoretical aspects of management, but delivered to the group in a classroom setting rather than one-on-one in the firm, and was then followed by group consulting sessions.

table[table omitted — 1,991 chars of source]

The main outcome is what the original paper refers to as the “Anexo K management score,” the average across 141 management practices (each rated on a five-point scale), measured at three points: baseline, during the intervention (at the end of implementation), and post-intervention (at the one-year follow-up). Following the specification in the original paper, we replicate their main results using a saturated regression that interacts each of the two treatment indicators---individual and group consulting---with the during- and post-intervention periods, with control as the reference arm; the regression controls for the baseline management score, randomization-triplet fixed effects, and a post-period indicator. The estimates are reported in Table (ref). The four interaction coefficients---each treatment arm in each period---are the effects relative to control, and they show immediate and persistent positive effects on firm management.

For standard errors, iacovone2022improving cluster at the firm level. Based on our theory, we instead compute CR2 standard errors at the group level: for firms in the group treatment we treat each consulting group as a single cluster, while for firms in the individual-treatment and control arms each firm forms its own cluster. These CR2 standard errors are shown in the second standard-error row of each cell. As expected, they are larger than the original ones, with a substantial increase in some cells.

We also calculate the ICC and test for interference in the group treatment, with results reported in the last row of Table (ref). The null of no interference is rejected at the 0.001 level, and the overall intraclass correlation exceeds 0.5. For nearly all submeasures (except Marketing Practices), the group treatment induces substantial intraclass correlation. This is intuitive, since firms within a group receive the same treatment and sign an agreement to work together and help each other.

Conclusion

We have developed design-based inference for group interaction experiments, in which individuals are allocated to groups that interact under group-level treatment conditions. Organizing the problem around two design cases (fixed versus randomly formed groups) and two outcome cases (no interference versus group interference), we characterized for each of the four cells the estimand that commonly used estimators target and the variance estimator that is valid for it. These estimators target the average treatment effect (ATE) when interference is absent, the total effect (TOT) under fixed groups with interference, and an average marginalized exposure effect (the PAME) under randomly formed groups with interference. The cluster-robust variance estimator (CR2) is consistent in all four cells; the heteroskedasticity-robust estimator (HR2) is consistent only when groups are randomly formed and interference is absent, where it is also the more efficient choice.

These results revise a familiar rule of thumb. Rather than clustering at the level of treatment assignment, as is commonly inferred from abadie2023should, one should cluster at the level of the potential outcome inputs\footnote{We thank Jonne Kamphorst for suggesting this phrasing.}---the arguments that determine which potential outcome an assignment pattern reveals. Under SUTVA, a unit's own treatment is the only such input. Once interference is present, the inputs include other units' treatment or group assignments, and valid inference requires clustering at the level of the resulting interference neighborhood even when treatment is assigned individually. In group interaction experiments that neighborhood is the group, so clustering at the group level is required whenever interaction generates interference; because such experiments typically assign individuals to groups with the expectation of interaction between individuals, cluster-robust inference should be the default. The intraclass-correlation test we develop helps researchers determine which regime they face, and hence which variance estimator to use. Our reanalyses of mendelberg2014does and iacovone2022improving bear these points out: in both, the null of no interference is rejected and group-level cluster-robust standard errors are materially larger than the standard errors that do not account for group interaction.

Methodologically, our asymptotic results rest on a sparse-sampling regime and a coupling argument that recasts random group formation as sampling from a population of candidate groups, shows the no-overlap restriction to be asymptotically immaterial when $N^2/n \to 0$, and thereby transports a standard central limit theorem to this design. The same device may prove useful for design-based inference in other experiments whose realized units are sampled subsets of a larger population.

Several extensions remain open for future research, and two seem particularly important. First, the analysis here suggests potential trade-offs that come from choosing a design with larger versus smaller groups. With a fixed sample size, one could have many small groups or a few larger groups. Having many smaller groups yields more independent clusters, however the group interaction may be more intense, and so the implications for estimator variance are ambiguous. A design with many small groups also targets a different estimand (marginalizing over small numbers of group-mates) than having fewer larger groups. Both statistical and substantive considerations are needed to determine which design is best. Second, the analysis here takes groupmate composition to be a source of unmeasured heterogeneity, and the estimands we consider here marginalize over that heterogeneity. Another target of inference, and one that is the focus of papers like li2019randomization, is the effect of measured variation in peer composition. The analysis here suggests ways that one might revisit the question of “peer effects”.

singlespace\begin{appendix}