Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
96,275 characters · 10 sections · 80 citation commands
Linear estimation of global average treatment effects
\pagenumbering{arabic} \onehalfspacing
Economists often study situations in which some units received a given treatment in order to estimate the consequences of treating more of them. For example, \citeapos{worms} eevaluation of deworming medication in a sample of 2,300 schoolchildren shaped subsequent debate over whether to administer similar medication to all schoolchildren in Kenya and beyond. A key quantity for policy-making in such settings is the global average treatment effect (GATE), or average causal effect of treating all (relevant) units.
Learning the GATE is challenging, even when treatment is randomly assigned, because it compares two counterfactual states which are never directly observed for any unit: one in which it, and all other units, are treated, and another in which none are. The econometrician must therefore extrapolate from, for example, the observed outcome for a unit with many treated neighbors to its counterfactual outcome had everyone been treated. Such extrapolation requires an assumption controlling interference or “spillovers" between units.
The predominant approach to dealing with spillovers has been to assume they are local via exposure mappings Manski13 which imply that treating one unit affects only a small number of neighbors, such as those within the same classroom or village. (The Stable Unit Treatment Value Assumption is an extreme example, implying that no other units are affected.) But such assumptions are sometimes in uncomfortable tension with physical or economic logic. Parasitic worms, as worms emphasized, need not stay neatly confined within a classroom. Economic models often imply that everyone's behavior affects everyone else to some degree through prices, strategic interactions, and other mechanisms. Recent, large-scale field experiments have illustrated these kinds of far-reaching general-equilibrium effects MuralidharanNiehaus2017. Reforming a public workfare program in India, for example, affected market wages and employment, land rents, and firm entry, and these effects did not stop at sub-district boundaries Muralidharanetal2021GE. Causal channels like these are difficult to reconcile with strictly local spillovers.
In this paper we study the estimation of the GATE via bounds on the magnitude of spillovers, rather than exposure mappings. Following Leung22, the specific assumption we work with is that interference decays with at least a power $\gamma$ of Euclidean distance.\footnote{An earlier version of the paper arxivprevious extends the results to “nearly Euclidean” spaces in which the relevant notion of distance can be asymmetric (e.g. if pollution travels from upwind to downwind sites) and need not exactly satisfy the triangle inequality (e.g. trade flows in a gravity model), or where the researcher is uncertain what notion of distance is most relevant and so uses the minimum of several proper metrics.} This yields a statistical framework better-suited to many economic settings: endogenous economic interactions decay with a power of distance in the gravity AllenArkolakisTakahashiUniversal and market access DH16 traditions, for example. We allow for $d$-dimensional spaces, arbitrary experimental designs, and for any estimator that can be written as a linear function of observed outcomes---a class which includes those commonly used in both applied and theoretical work.
In this setting we show that no linear estimator, regardless of the experimental design, can be guaranteed to converge to the GATE faster than $n^{-\frac{1}{2 + d/ \gamma}}$. The fact that this is slower than the parametric rate $n^{-\frac{1}{2}}$ illustrates the intrinsic difficulty of estimating the GATE, while the fact that it is faster for large $\gamma$ illustrates the intuitive idea that the problem becomes less difficult when the population is economically less interconnected. The optimal rate can be achieved using inverse probability weighting (IPW) when the experimental design is cluster-randomized with clusters that grow with the population at the right rate. The key idea is to prevent the probability of any one unit having all of its neighbors (within some growing radius) treated or untreated from decaying to zero.
These results build on and extend Leung22's (Leung22) important recent contribution, which characterizes the optimal rate when using a Horvitz-Thompson estimator in $\mathbb{R}^2$ and demonstrates that it is achievable using clusters that take the form of growing squares if $\gamma > 2$. We extend this result to $\mathbb{R}^d$, optimize over all linear estimators as well as all experimental designs, and allow $\gamma \in (0,\infty)$. The key to the last generalization in particular is showing that, while a phase change in the estimation error does occur at $\gamma = d$, the problematic terms this introduces shrink with the radius of the estimator and are never larger than the asymptotic bias. The phase change therefore does not affect the rate of convergence or disrupt inference.
Achieving the optimal rate does require assigning treatment in clusters that are “large” in the sense that they grow rapidly alongside the experiment. If instead they grow at a slower-than-optimal rate, then units with fully treated (“saturated”) or untreated (“dissaturated”) neighborhoods become too rare.\footnote{One can think of this as analogous to the “limited overlap” problem where individual propensity scores are close to zero or one imbens_limited_overlap.} The IPW estimator can then capture spillovers only in small neighborhoods around each unit, so that it cannot achieve the optimal rate. Moreover, its bias will tend to dominate its variance, posing challenges for inference. These issues are exacerbated under “random saturation” designs, where spatially grouped units can have different treatment assignments; here the probability that even one neighborhood is fully saturated or dissaturated vanishes, so that IPW cannot be guaranteed to consistently estimate the GATE at any polynomial rate. In short, when free to choose, researchers prioritizing the GATE should use large clusters.
In practice, however, researchers have often used small-cluster or even unit-level randomization MuralidharanNiehaus2017. There are several reasons for this. First, tradeoffs between estimands can arise; there may be scientific motives for separately identifying direct and indirect treatment effects, for example, as well as the policy-relevant GATE. Second, researchers may face practical constraints, such as a requirement to randomize treatment across pre-existing administrative units. For such scenarios we consider positing a linear causal model: potential outcomes are linear functions of treatment indicators. This assumption compensates for the scarcity of (dis)saturated neighborhoods in small-cluster designs by allowing learning about them from neighborhoods that are less than fully (dis)saturated. We view it as implicit in prominent applied papers which have reported results from linear regressions of outcomes on (among other things) the mean treatment status of nearby units worms, EggeretalGE, Muralidharanetal2021GE.
The linear causal model does not change the optimal rate, which can still be achieved using IPW if the rate of cluster growth is unconstrained. If, on the other hand, clusters must be “small” in the sense above, then linearity lets the researcher improve on the IPW rate using an OLS estimator. This estimator is formed by regressing outcomes on the share of treated clusters intersecting a neighborhood around each unit---a regression related to those in the applied studies above, but distinct in that (among other things) they regress on the share of nearby treated units. These regressions concentrate around a weighted average of the spillover effects, rather than the unweighted GATE.\footnote{The weights are non-negative, and in that sense less problematic than those which arise in two-way fixed effect designs TWFEChaise,GOODMANBACON2021254. However, since the weights depend on both the treated unit and the affected unit, they have no welfarist interpretation.} The OLS estimator proposed here is consistent for the GATE, and remains so even if the true data-generating process is non-linear under weak conditions on the experimental design. Overall these results imply the following tradeoff for practitioners: if feasible, a large-cluster experimental design paired with the IPW estimator is recommended for the GATE. If large clusters are not available, and if it is plausible that the underlying DGP is not too non-linear, then an OLS approach may be attractive.\footnote{The OLS-based approach is also amenable to the introduction of prior information the researcher may have about the pattern of economic interactions, which can be incorporated in a TSLS setup. We omit this approach here for brevity but study it in an earlier working paper version arxivprevious.}
For small-cluster designs there remains the practical problem of choosing the “radius" of an OLS estimator in a given, finite population. Should the researcher regress outcomes on the average treatment status of clusters within 1km, 10km, or some other distance? Choosing the radius that minimizes mean square error (MSE) is infeasible, since bias is always unknown. We propose a minimax procedure that is asymptotically guaranteed to select the radius that minimizes the worst-case MSE over a credible set of potential outcomes. This procedure does not require the use of outcome data, but only that the researcher specify a conservative lower bound on the rate of spatial decay of spillovers. This in turn facilitates inference. Inference faces two challenges: the bias is unknown, and the variance includes the usual unidentified component attributable to heterogeneous treatment effects in fixed populations. Undersmoothing is a natural consequence of this method because the researcher never knows the exact rate of spatial spillover decay and therefore will always use a conservative lower bound. This results in a wider radius than if they knew the exact rate for sure, addressing the first challenge. We provide a new asymptotically valid bound on the unidentified variance term to address the second.
We apply these methods to data from the large-scale cash transfer experiment in rural Kenya studied by EggeretalGE (henceforth, EHMNW). We first estimate the GATE on annualized household consumption. Using our OLS estimator with the minimax radius, we obtain an estimate of \$445 ($p < 0.01$), 22% larger and 65% more precise than when we adopt the radius used by EHMNW. We then simulate IPW and OLS performance holding geography and observed features of the outcome distribution fixed, but under a linear causal model and two different experimental designs: the one actually used by EHMNW, and a counterfactual one with clusters 4-5 times larger. Under the actual design the MSE of IPW is several multiples that of OLS, while with larger clusters the two estimators perform comparably. We also confirm that OLS confidence intervals provide appropriate coverage at their nominal level.
The primary theoretical contribution of the paper is to show that, broadly speaking, estimating the GATE is more realistic than had previously been understood. Theorists have cast the GATE as an important estimand because it is interpretable even when the researcher is uncertain about the exact channels over which spillovers propagate. For example, in response to \citeapos{SavjeMisspec21} \ characterization of estimands available when exposure mappings are somewhat misspecified, auerbachcomment point out that these estimands need not have causal interpretations, to which michaelcomment responds that the GATE is one of the few that does, but that it is difficult to estimate in practice. Our results show that the GATE can be rate-optimally estimated using standard IPW estimators under weaker conditions than those in Leung22, and that alternative OLS-based estimators converge faster in the presence of common constraints on design.
This approach complements recent theoretical work that has focused on consistency or on optimal experimental design, leaving open the question of how to jointly choose a design and estimator. auerbach2023localapproachcausalinference provide consistency results for a class of estimands including the GATE under very general assumptions, for example, but do not address how choices of design and estimator map into the rate of convergence and what rates are possible. leung2025crosscluster and viviano2024causalclusteringdesigncluster find optimal experimental designs to estimate the GATE assuming an IPW or difference-in-means estimator will be used. We emphasize the joint choice of estimator and design to achieve the best rate, the ways this is affected by the linearity assumptions on potential outcomes that are implicit in applied work, and the design constraints researchers often face.
An adjacent line of work has focused on estimands other than the GATE, such as the average causal effect of receiving some specific level of neighborhood exposure to treatment Aronow17,savjearonow, Leung21,auerbach2024regressiondiscontinuitydesignspillovers, which can be consistently estimated and interpreted even when the exposure mapping is mildly misspecified Zhang21, SavjeMisspec21.\footnote{Work in this genre grew out of an earlier literature on peer effects in which outcomes are linear in exogenous and endogenous variables and either the coefficients are known up to scale Manski93, Lee07, GPandI13, Brom14 or exposure mappings rule out interference from faraway units volf,Sussman. Our results under linearity show how one can consistently and rate-optimally estimate global average treatment effects even without such assumptions.} It targets only estimands that can be written in terms of local exposures (and thus not including the GATE). More generally, Airoldi18 show that there is no consistent estimator of the GATE within such a framework.\footnote{See Section 3.3 of savjearonow, for example, for discussion of this point within their framework.}
For applied work we provide a new theoretical foundation and practical guidance for program evaluation in the presence of spillovers. At least since worms it has been understood that capturing spillovers can be crucial to getting policy inferences right, and that the decay of spillovers with distance can be exploited to construct estimators that capture them worms_at_work_16,EggeretalGE,Muralidharanetal2021GE. But most studies have limited themselves to estimating the total effect of the experiment actually conducted, rather than the policy-relevant GATE, and done so using estimators and designs for which consistency and optimality results were not available. Our results allow for consistent, rate-optimal estimation of the parameter arguably most relevant to program evaluation, and provide guidance for tuning and inference. In doing so our broader aim is to provide a bridge between the experimental approach to economics---which has rested on the assumption that a “pure control group” exists---and economic theory, which typically implies that one does not.
Notational conventions are as follows. For sequences $A_n,B_n$ we say $A_n \lesssim B_n$ if $|A_n/B_n| =\mathcal{O}_p(1)$. $A_n\sim B_n$ means that both $A_n\lesssim B_n$ and $B_n\lesssim A_n$. A subscripted vector ${b}_i$ denotes the $i$th element while a subscripted matrix $G_i$ denotes the $i$th row.
A researcher observes a finite population $\mathcal{N}_n$ of $n$ units and randomly assigns each unit $i$ a binary treatment $D_i$. Let $\mathbf{D}\equiv (D_i)_{i\in\mathcal{N}_n}$ denote the full $n\times 1$ vector of treatments. We call the probability distribution of $\mathbf{D}$ the {\it experimental design}. The potential outcome for unit $i$ is the scalar-valued function $Y_i(\mathbf{d})$ where the argument $\mathbf{d}$ is a particular $n\times 1$ vector of counterfactual treatment assignments. The observed outcome is the random variable $Y_i(\mathbf{D})$. Since our estimand is the causal effect of treating an entire population and not a random sample, we model potential outcomes $Y_i(\cdot)$ as deterministic functions and the treatment assignment $\mathbf{D}$ as random (and correspondingly will study asymptotics with respect to a sequence of growing finite populations).\footnote{Fixing potential outcomes is standard in this literature, see for example Aronow17, Sussman, Leung21, Leung22.} We remain largely agnostic about the origins of the potential outcomes $Y_i(\cdot)$, which could be the realization of some unspecified spatial process, and in particular we impose no exposure mappings.
Our first assumption states that potential outcomes are uniformly bounded. While less restrictive assumptions could be sufficient, Assumption (ref) eases exposition and is standard Aronow17,MANTA2022109331,Leung22,leung2025crosscluster.
Our estimand of interest is the average causal effect of assigning {\it every} unit in the entire population to treatment vs assigning {\it no} units at all to treatment. This is called the {\it global average treatment effect} (GATE) because it sums both the own-effects and spillover effects of treating all units (within some relevant class).
The GATE is often the estimand most relevant for policy decisions; a policymaker using RCT data to decide whether to scale up a program to the control group, for example, should ideally base this decision on the GATE rather than the direct treatment-on-treated effect. Note that we assume the researcher observes outcomes for the entire population for whom she wishes to estimate the GATE.\footnote{Appendix (ref) extends to settings where only a subset of the entire population is eligible for treatment.} Results concerning the limiting properties of our estimators will require a notion of a growing population, and our estimand $\theta_n$ can change as the population grows. Later assumptions will characterize the conditions this sequence of populations must follow.
To solve the core identification challenge, we consider restrictions on the magnitude of spillovers, as opposed to less economically plausible restrictions on their sparsity. Specifically, we consider a restriction on the (maximum) impact that faraway treatments have on each unit. To that end, we define $\rho(i,j)$ as the distance metric between units $i$ and $j$, which in turn induces a natural notion of neighborhoods as $s$-balls:
Here $s > 0$ parameterizes the radius of the neighborhood. This will allow us to study estimators that make use of increasingly large neighborhoods as $n$ grows.
For expositional clarity our focus here will be on the purely spatial case where units are located in $\mathbb{R}^d$ and $\rho$ is the Euclidean distance. This captures many economic applications: economic interactions decay geometrically with spatial distance in the gravity AllenArkolakisTakahashiUniversal and market access DH16 traditions, for example. Space need not be the only factor governing interactions between units, provided that a spatial bound on the rate of decay holds. In an earlier version of the paper arxivprevious we showed more explicitly that the main results still hold if we relax the triangle inequality or symmetry property of distance,\footnote{Another case we discuss in that draft in which the more general results are useful is when the researcher is uncertain about which of several potential channels for spillovers matter. For example, training small business owners might lead to knowledge spillovers for which the social network defines the relevant notion of distance, businesses stealing for which physical distance and product substitutability between firms define distance, input demand effects for which the buyer-supplier network defines distance, and so on. We thank David McKenzie for suggesting this example. See SavjeMisspec21, auerbachcomment, and michaelcomment for further discussion of the desirability of robustness to misspecification.} while in ongoing work we are applying similar methods to a “small-world” network setting in which results change more fundamentally.\footnote{In this setting neighborhoods instead have the “small-world” feature that their size grows exponentially rather than polynomially with their radius. Our proof techniques extend naturally to this setting and yield an analogous (but slower) bound on the rate of convergence.}
In addition to locating units in Euclidean space, we require that the population is not too densely or sparsely distributed. Assumption (ref) says that (a) the population can fit inside a ball with volume proportional to $n$ and (b) every pair of units is at least distance $\rho_0>0$ apart.
Our next assumption captures the idea of spatial decay in the spillover effects. Specifically, it states that changing the treatment assignments of any subset of units more than distance $s$ away from unit $i$ changes the potential outcome of unit $i$ by an amount upper-bounded by a function that decays with $s^{-\gamma}$. The researcher need not know the exact fashion in which spillovers decay; they need only accept that spillovers are upper-bounded in this way. Assumption (ref) is identical to Assumption 3 of Leung22 and leung2025crosscluster except that, as we discuss further below, we require only $\gamma > 0$ rather than $\gamma > d$. This also implies that our inference results will relax fast spatial decay requirements common in the spatial central limit theorem (CLT) literature JENISH_NED.
Assumption (ref) bounds spillovers deterministically, as is standard in the literature on the GATE and related estimands Leung22,leung2025crosscluster.\footnote{Other papers bound long-range spillovers deterministically in a more subtle way via exposure mappings that apply directly to the potential outcome function itself rather than to expectations over it Manski13,Aronow17.} For readers familiar with the spatial econometrics literatures it may be helpful to contrast it with the Near Epoch Dependence used by (for example) JENISH_NED, which instead bound spillovers using expectations:
where $\mathcal{F}_{i,s}$ is the $\sigma$-algebra generated by all the treatments within distance $s$ of unit $i$ and the expectation is taken over the distribution of treatment assignments $\mathbf{D}$. While closely related to Assumption (ref), condition ((ref)) is not suitable for GATE estimation, for two reasons. First, ((ref)) is not primitive in this setting because it depends on the experimental design chosen by the researcher, over which we want to allow for optimization. Second, and more importantly, ((ref)) is not strong enough to identify the GATE, as Remark (ref) below illustrates. This is why we (and others, such as Leung22,leung2025crosscluster) bound spillovers deterministically using Assumption (ref).
We study estimators that are linear, i.e. that take the form of weighted averages of observed outcomes $Y_i$ where the weights $\omega_{in}(\mathbf{D})$ are arbitrary functions of the full vector of treatment assignments $\mathbf{D}$. Weights can be specific to units $i$ and may change with $n$, though we will sometimes suppress the latter subscript for brevity. Such estimators can thus be written as:
The class of linear estimators includes the IPW estimators common in theoretical analysis Aronow17,Leung22 as well as (after conditioning on covariates) the regression-based estimators common in applied work.
Theorem (ref) states that no estimator taking the form of a weighted mean of outcomes can be guaranteed to converge to the GATE at a rate strictly faster than $n^{-\frac{1}{2+ d /\gamma}}$ for all potential outcomes satisfying Assumptions (ref) and (ref). The following sketch illustrates the workings of the proof. Posit a sequence $a_n\to \infty$ such that $a_n(\widehat{\theta}_n-\theta_n)={o}_p(1)$ for any set of potential outcomes satisfying our assumptions. Let $\mathcal{J}_n\subset \mathcal{N}_n$ be a minimal subset of the population such that every unit $i\in \mathcal{N}_n$ is within distance $a_n^{1/\gamma}$ of some $j\in \mathcal{J}_n$ called $j(i)$; call these the “focal units.” Now let the potential outcomes be $Y_i(\mathbf{D})=a_n^{-1}D_{j(i)}$, so that they are entirely determined by the treatment statuses of the focal units. (Note that these satisfy Assumption (ref), as no focal unit $j(i)$ is too far from $i$.\footnote{Although the GATE here is decreasing, it is not $o\left(a_n^{-1}\right)$ and therefore must be meaningfully estimated by hypothesis.}) The effective sample size is consequently $|\mathcal{J}_n|$. And this quantity is bounded: by Assumption (ref) the population is contained by a ball of radius $n^{1/d}$ in $\mathbb{R}^d$, and it can be shown that the number of balls of radius $a_n^{1/\gamma}$ required to cover a larger ball of radius $n^{1/d}$ is just $\mathcal{O}(na_n^{-d/\gamma})$, i.e. the volume of the whole space divided by the volume of the balls used to cover it. We thus have: $a_n \lesssim |\mathcal{J}_n|^{1/2}\lesssim {n^{1/2}a_n^{-d/(2\gamma)}}$. This relation is solved for $a_n$ to yield the final result $a_n\lesssim n^{\frac{1}{2+d/\gamma}}$.\footnote{This sketch is partial; another set of potential outcomes is needed to resolve special cases, e.g. estimators where the weights $\omega_{i,n}(\mathbf{D})$ converge rapidly to zero.}
The rates allowed by Theorem (ref) are slower than the parametric rate $n^{-\frac{1}{2}}$ because $d/\gamma >0$. The result thus illustrates a sense in which GATE estimation in the presence of global spillovers is a more difficult problem than estimation with localized spillovers, where the parametric rate can be guaranteed Aronow17. Theorem (ref) also links the spatial rate of decay of the spillover effects with the rate of convergence of linear estimators. The economic principle this suggests is that estimators based on Assumption (ref) will tend to converge faster in less integrated economies, where distance is an important constraint on economic interactions, than in highly integrated ones where all units are “close” to each other.
Theorem (ref) extends the main results of Leung22 and leung2025crosscluster in two ways. First, it extends the rate limit result to the case $\gamma < d$. We will show below that this “speed limit" is still achievable in that case. Small values of $\gamma$ will also play a key role in our approach to bandwidth selection and inference because they can serve as conservative lower bounds which facilitate undersmoothing. Second, Theorem (ref) finds the optimal rate jointly over all experimental designs as well as all linear estimators. This is relevant given that applied work has used linear estimators other than IPW.
We next show that the optimal rate is achievable using an experimental design from a class of “Scaling Clusters” designs we define, and a suitable estimator.
Consider cluster-randomized experimental designs in which the researcher partitions the population $\mathcal{N}_n$ into $m_n$ mutually exclusive clusters $\{\mathcal{P}_c\}_{c=1}^{m_n}$ and assigns treatment cluster-by-cluster, treating each cluster independently of the rest with probability $p$, so that all units within a cluster share the same treatment status, denoted $W_c \in \{0,1\}$.
We focus on such designs throughout the main text for expositional clarity, and following the literature Leung22,leung2025crosscluster. In practice researchers often use variants of this design where exactly fraction $p$ of clusters are treated within predefined strata (“stratified randomization”). Our empirical application (EHMNW) is one prominent example, among many others.\footnote{For instance, worms (which randomized schools stratified by administrative zone), CaseyGlennersterMiguel2012reshaping (villages stratified by ward), or Alatasetal2016selftargeting (villages stratified by groups of one or more subdistricts).} We defer analysis of such cases to Appendix (ref) as the arguments are substantially more involved but the results are substantively the same. In particular we show that the same rates of convergence can be achieved when the design is spatially stratified provided clusters still grow, and provide consistency and CLT results for the OLS estimator applied to designs such as that in our empirical application.\footnote{Notice the absence of an overlap (or “positivity”) assumption bounding away from zero the probability that any given unit experiences full exposure to treatment or to control conditions in some relevant neighborhood. Such restrictions on the design are often used when the estimand compares the effects of “local" exposures. In these settings it (helpfully) ensures that the probability that the researcher observes both units whose relevant neighborhoods are fully treated and units whose relevant neighborhoods are fully untreated goes to one, and that the researcher can construct consistent estimators by comparing these units. In our setting, however, the relevant neighborhood is the entire population, so we can never observe both such units at once: if one unit is exposed to universal treatment, then there cannot be a comparison unit that was exposed to universal non-treatment. As a result a positivity assumption would hurt rather than help us.}
To consistently estimate the GATE we will need clusters to be of a shape and size such that the number of clusters within any $s$-neighborhood does not grow too rapidly with $s$. To achieve this, consider a class of {\it Scaling Clusters} designs where the clusters are shaped like neighborhoods and grow with $n$. The rate of cluster growth is governed by a non-decreasing sequence of positive numbers $\{g_n\}_{n=1}^\infty \subseteq \mathbb{R}^+$ as follows:
A Scaling Clusters design is thus one where treatment clusters both contain and are contained by neighborhoods that grow in proportion to $g_n$. Assumption (ref) guarantees that the number of clusters $m_n$ satisfies $m_n\sim\frac{n}{g_n^d}$, i.e. it scales with the ratio of volume of the whole space to the volumes of the individual clusters. This requirement is loose, ruling out only clusters that become progressively more “gerrymandered" as the population grows. One simple example of a Scaling Clusters design is a grid of expanding squares in $\mathbb{R}^2$. Another is provided by leung2025crosscluster, whose Algorithm SA.1.1 constructs clusters using unsupervised learning to optimize the rate at which MSE shrinks; these satisfy Definition (ref) by his Theorem 1.\footnote{While leung2025crosscluster optimizes the rate at which MSE shrinks and not its finite-sample level, we view his proposal as intuitively appealing as it aims to minimize the number of units near cluster boundaries which should lower estimator bias. Alternatively, we conjecture that it may be possible to extend the minimax risk procedure for radius selection we develop in Section (ref) to also select among a small finite set of candidate designs.} Given a specific experimental design in a fixed population, one can think of it as satisfying Scaling Clusters if the largest circle contained within each cluster is close to the smallest circle containing it. Figure (ref) illustrates this: in the left panel the solid (contained) and dashed (containing) circles are close together because the clusters are “convex enough," while in the right panel the solid and dashed circles are far apart because the clusters are very “gerrymandered."
The researcher can achieve the optimal rate without further assumptions using an inverse probability weighting (IPW) estimator. This estimator is conveniently tractable and standard in theoretical work on network effects Aronow17, Sussman, Leung21, Leung22,leung2025crosscluster. To define it, fix any sequence of neighborhood radii $\{\kappa_n\}_{n\in \mathbb{N}} \subset \mathbb{R}^+$ chosen by the researcher which grow to infinity with the sample size. Define the random variables $T_{i1},T_{i0}$ as indicators that all of $i$'s neighbors within distance $\kappa_n$ are treated or untreated:
The IPW estimator compares units with saturated $\kappa_n$-neighborhoods to units with dissaturated neighborhoods, following the recommendation in HiranoImbensIPW to take a difference between means rather than the simple Horvitz-Thompson procedure.
For the IPW estimator to converge at a polynomial rate it is necessary to control the probabilities in the denominators of the weights, which is equivalent to upper-bounding the number of clusters $\phi(i,\kappa_n)$ that each neighborhood $\mathcal{N}(i,\kappa_n)$ intersects. Lemma (ref) in Appendix (ref) shows that a Scaling Clusters design makes $\phi(i,\kappa_n)$ scale with $(\kappa_n^d+g_n^d)/g_n^d$. In order for IPW to be consistent, the radius of the estimator thus cannot outpace the cluster size. Theorem (ref) then shows that the IPW estimator paired with a Scaling Clusters design achieves the optimal rate from Theorem (ref) when the radius and cluster size grow at the appropriate rate.
In contrast to the IPW estimators for local effects studied in Aronow17, the IPW estimator of the GATE here converges at a subparametric rate. Extending results for the Horvitz-Thompson estimator and square cluster design in $\mathbb{R}^2$ studied by Leung22, Theorem (ref) applies to $\mathbb{R}^d$ for any $d$ and (in conjunction with Theorem (ref)) shows that pairing any Scaling Clusters design with an IPW estimator is rate-optimal over all pairs of experimental designs and linear estimators when the clusters and radius grow at the right rates.
Most importantly, Theorem (ref) allows for any $\gamma > 0$, relaxing the requirement that $\gamma > d$ from Leung21,leung2025crosscluster. A phase change does occur at $\gamma = d$, adding a possibly non-normal stochastic term to the estimation error arising from sums of long-range spillovers. However, we show that even though the “bad" stochastic term can converge more slowly than $n^{-1/2}$, it is controlled by the radius of the estimator and in fact shrinks with $\kappa_n^{-\gamma}$ (see Step 2 of the proof in Section (ref)). Since this is the same rate, for any $\gamma > 0$, as the bias, the overall rate of convergence is unaffected by the phase change. Inference will also be unaffected when the researcher undersmooths, i.e. selects a $\kappa_n$ large enough that variance dominates bias.\footnote{On the other hand, $\gamma>d$ would be required for many estimands other than the GATE. For example, JENISH_NED study inference for a population average estimated by a simple sample mean in a non-experimental setting. When $\gamma < d$ in this setup, the “bad" stochastic term converges slower than $n^{-1/2}$. Since there is neither a “radius" nor “clusters" to grow, there is no way to ensure that this term is dominated by the variance of the sample mean, posing a major problem for inference.}
This positive result regarding IPW is sensitive, however, to the experimental design. Theorem (ref) shows that when treatment clusters grow more slowly than the optimal rate, and barring implausible topologies,\footnote{Specifically, the theorem requires that units are not spatially arranged into dense and well-separated clumps that all fly apart as the population grows. A study that grows across a map covered by evenly-spaced villages meets this requirement, for example.} IPW will converge slowly. Growing $\kappa_n$ faster will not prevent this. Appendix (ref) shows additionally that in this case IPW's bias will usually dominate its standard error, posing challenges for inference.
We have seen that GATE estimation with IPW requires that randomization clusters be “large,” in the sense that their radii grow rapidly. But this contrasts with practice, where many field experiments use small clusters. In their review, for instance, MuralidharanNiehaus2017 documented a median cluster size of just 26 units, with 34% of designs not clustered at all. Where the primary purpose of the study is to inform a decision about scale-up, so that the GATE is the relevant estimand, such designs are unlikely to be optimal. But researchers may face constraints: equity considerations or the logistics of treatment delivery may require that the experimental design be based on pre-existing administrative units. This was true, for example, of the large-scale experiments referenced above: schools in worms, villages in EggeretalGE, and sub-districts in Muralidharanetal2021GE. Or they may choose to randomize in small clusters---or even at the individual level---due to a scientific interest in studying direct and indirect effects as well as the GATE.
How should researchers approach GATE estimation in such scenarios? Existing work---while targeting other, related estimands---rarely uses IPW: none of the studies covered by MuralidharanNiehaus2017, nor any of the other examples we have provided, do so. Instead, studies that take a spatial approach typically report results from OLS regressions on measures of the average treatment intensity in a surrounding neighborhood.
We next study an analogous approach to GATE estimation. The core idea is that, to learn about the GATE from an experiment with small clusters, we must learn about the outcomes a unit would experience if its neighborhood were fully (dis)saturated by extrapolating from the outcomes of units whose neighborhoods are less than fully (dis)saturated. Assumption (ref), which we introduce next, allows for such extrapolation.
Regression-based estimators are an intuitive way to take advantage of this posited linear structure, and are common practice in applied work. One could consider various regression-based estimators; we focus here on the simplest, which is to regress outcomes on a single measure of neighborhood treatment intensity.\footnote{Applied papers often regress outcomes on an indicator for own-cluster treatment status as well as a measure of treatment intensity in a neighborhood beyond one's own cluster; more generally, one could regress outcomes on measures of treatment intensity in multiple bands around each unit. In our empirical application we will see that separating out the own-cluster treatment status in the regression does not substantially change the results (see appendix Table (ref)).} To that end, define the scalar-valued regressor $X_i$, which is the centered sum of treated clusters within distance $\kappa_n$ of unit $i$ divided by the average number of clusters $\overline{\phi}$ within distance $\kappa_n$ over all units.
Here $\phi(i,\kappa_n)\equiv \sum_{c=1}^{m_n} \mathbbm{1}\left\{\mathcal{P}_{c} \cap \mathcal{N}(i,\kappa_n) \neq \emptyset\right\}$ denotes the number of clusters within $\kappa_n$ of unit $i$, and $\overline{\phi}\equiv \frac{1}{n}\sum_{i=1}^n\phi(i,\kappa_n)$ denotes its mean. Then define the OLS estimator as the following simple linear regression of $Y$ on $X$.
Our next pair of results demonstrates that imposing Assumption (ref) does not change the optimal rate, but does allow the OLS estimator to estimate the GATE consistently and to achieve that optimal rate. For the latter result we require that $\frac{\kappa_n^d+g_n^d}{n}\to 0$, which simply says that the number of units in any cluster or $\kappa_n$-neighborhood grows more slowly than $n$. We also impose the mild condition $\kappa_n \to \infty$ simply to shorten the proofs.
\newenvironment{thmbis}[1] { \addtocounter{theorem}{-1}
}
The last statement of Theorem (ref) is particularly significant in that, unlike Theorem (ref), it requires no lower bound on the rate $g_n$ at which clusters grow. It thus shows that OLS converges faster than IPW when clusters grow slower than the optimal rate, and in fact that OLS provides a consistent estimator even when the design is unclustered, i.e. randomization is at the individual level. (To see this, consider Theorem (ref) with $g_n = \rho_0/2$, or half the minimum distance between two units from Assumption (ref).b, which ensures that each cluster contains just one unit.) Intuitively, OLS is able to solve the small-clusters problem by using the linearity assumption to extrapolate from nearly-(dis)saturated neighborhoods to fully-(dis)saturated ones. (Our variance estimator, on the other hand, will require that clusters grow, but this can be at an arbitrarily slow rate.)
Taken together, these results suggest the following tradeoff for practitioners. If nonlinearities are expected to be economically unimportant, and the design of an experiment requires small clusters for reasons other than GATE estimation, then the linear approach developed here will be attractive. If, on the other hand economic interactions between treatment effects are expected to be important, and large clusters are feasible, then the IPW-based approach developed in the previous section is preferable from the point of view of the GATE.
We turn next to the problem of selecting a radius $\kappa_n$ for the OLS estimator given a specific, finite study population and experimental design. Radius selection for IPW is largely determined by the experimental design, since saturated neighborhoods will not occur with reasonable probability if the IPW radius much exceeds the cluster size. But this logic does not apply to OLS, since it does not depend on saturated neighborhoods; nor do the convergence results above select a radius, as they hold provided that $\kappa_n=C_{\kappa} n^{\frac{1}{2\gamma+d}}$ for an arbitrary constant $C_{\kappa}$. In fact, without further structure nearly any choice of radius and clustering can be justified by some DGP satisfying Assumptions (ref), (ref), and (ref). To see this, notice that for any $\kappa_n>0$ and any experimental design, if we set potential outcomes to be $Y_i = \frac{\overline{\phi}c\kappa_n^{-\gamma}}{\max_j \phi(j,\kappa_n)}X_i$, then $\widehat{\theta}_{n,\kappa}^{OLS}=\theta_n$ with probability one. Thus, any choice of radius (and design) is optimal under some circumstances.
We aim to provide a default selection method that requires minimal contextual knowledge and does not depend on units of measurement. Data from a pilot experiment could of course be used to guide radius selection if available, but typically they are not. Alternatively, a researcher willing to commit to specific values of $c,\gamma$ and to accept bias of up to $B$ could simply choose $\kappa_n=(B/c)^{-1/\gamma}$, but this is demanding; choosing $c$ in particular amounts to choosing the scale of the spillover effects overall---which is essentially what the experimenter wishes to learn in the first place. Instead we will develop a rule that requires only that the researcher place a conservative lower bound $\widetilde{\gamma}$ on the rate of spatial decay $\gamma$. This approach takes advantage of the fact that our convergence results hold for any $\gamma>0$, allowing the bound to be arbitrarily conservative provided only that it is positive.
To define and motivate the approach, we first isolate the leading bias and variance terms in OLS. Proposition (ref) decomposes OLS into an expectation-zero stochastic sum of regression residuals $[\mathcal{A}]$, a deterministic bias term $[\mathcal{B}]$, and a dominated term.
To help interpret these terms, recall that $\mathcal{F}_{i,\kappa_n}$ is the $\sigma$-algebra generated by all the treatments within distance $\kappa_n$ of unit $i$. The random variables $\mathbb{E}\left[Y_i|\mathcal{F}_{i,\kappa_n}\right]$ can thus be interpreted as the outcomes after ignoring “long-range” spillovers, and so the proposition shows (among other things) that these may affect bias but do not affect the leading variance term, which will be useful for radius selection as well as obtaining a CLT in the next section.
We would ideally like to directly minimize the sum of the variance of $[\mathcal{A}]$ and the square of the bias $[\mathcal{B}]$, to which we refer henceforth as the risk.\footnote{To be precise, it is equal to the expected MSE modulo the vanishingly rare event in which OLS's denominator is near zero. This approach is standard; see for example the term $\sigma_n^2$ defined prior to Theorem 4 in leung2025crosscluster.} For an OLS estimator $\widehat{\theta}_{n,\kappa_n}^{OLS}$ that uses radius $\kappa_n$, given potential outcomes determined by $A$ and $\epsilon$ as defined in Assumption (ref), we write this as
where the variance is taken with respect to random assignments induced by the design $\mathcal{D}$. A larger radius will tend to increase variance and decrease bias. Finding the exact minimizer that balances this tradeoff is not feasible, however, since $A,\epsilon$ are unknown. Instead we will consider the radius that minimizes the worst-case risk over a set $\mathcal{G}_n$ of plausible DGPs. To motivate that set, we first need some insight into how the potential outcomes map into risk asymptotically. Proposition (ref) isolates the leading terms of the risk:
Proposition (ref) reveals a helpful insight: provided that no single unit's treatment has an unbounded effect on the population ($||A||_1\lesssim 1$), the leading variance term does not depend on all features of the spillover matrix $A$. Instead, it depends asymptotically only on the global treatment effect heterogeneity, i.e. variation in the row sums $A_i\mathbf{1}$. Similarly, the bias depends only on the tail behavior of the rows of $A$. Exploiting these facts, we define below a large set of DGPs $\mathcal{G}_n$ over which it will be straightforward to maximize risk:
Condition 1 simply says that all DGPs must satisfy the previous assumptions. Condition 2 strengthens Assumption (ref) from a condition on sums of spillovers to a condition on individual spillovers. Specifically, it stipulates that individual spillovers decay with distance no slower than a classic spatial moving average. Conditions 3 and 4 impose uniform bounds on the magnitudes of the error term and treatment effect heterogeneity, respectively. This broad class of linear DGPs includes, e.g. the spatial moving average models from Leung22.
Define the minimax radius ${\kappa}^*_n(\mathcal{G}_n)$ as the minimizer of the worst-case risk over these DGPs.\footnote{Since increasing $\kappa_n$ by less than the minimum distance between units $\rho_0$ does not necessarily change the risk, minimizers cannot be unique. Let the $\operatorname*{arg\,min}$ here denote the largest minimizer.}
To solve this minimax problem, we will show that the maximized risk can be asymptotically approximated by the deterministic function $ R(\kappa_n,\mathcal{G}_n)$ defined by:
Theorem (ref) establishes that minimaxing $ R(\kappa_n,\mathcal{G}_n)$ is asymptotically equivalent to minimaxing the true risk $\mathcal{R}^*$. This result requires that clusters grow slower than their optimal rate (which is the relevant case since, if they grew optimally or faster, the researcher should use IPW instead).
To use this rule, the researcher must select the five scalar parameters that Definition (ref) uses to define $\mathcal{G}_n$. Since $\overline{Y}$ does not affect the minimax radius so long as it is sufficiently large, the researcher can simply set it to an arbitrary large number and ignore it. For the others we suggest a simple default: setting them to maximize the resulting minimax radius. This is conservative from the point of view of inference, since large $\kappa_n$ leads to less bias, more variance, and thus better coverage of the confidence intervals.
Doing so involves setting $\sigma_e$ and $\tau$ as small as possible, as this reduces the effect of increasing $\kappa_n$ on the worst-case variance and these parameters do not affect worst-case bias. For $\sigma_e$ this means setting $\sigma_e = 0$, as it affects only Condition 3 of Definition (ref). Note that this does not cause the suggested radius to diverge because treatment effect heterogeneity can still drive estimator variance. In fact, since the variance term of $R(\kappa_n,\mathcal{G}_n)$ is now proportional to $\tau^2$ and the bias term to $\widetilde{c}^2$, only the relative magnitude of these two parameters (and not their individual values) will determine the minimizing radius of $R(\kappa_n,\mathcal{G}_n)$. Structurally, $\tau$ ought to scale with $\widetilde{c}$ since the latter governs the overall scale of the spillover effects and the former controls their heterogeneity. We suggest setting $\tau= \max_{i\in \mathcal{N}_n}\left|\sum_{j\neq i} \widetilde{c}\rho(i,j)^{-d-\widetilde{\gamma}} - \frac{1}{n}\sum_{k=1}^n\sum_{j=1}^n\widetilde{c}\rho(k,j)^{-d-\widetilde{\gamma}}\right|$, which is the smallest value that does not rule out the spatial moving average model from Leung22 with spillovers as large as Condition 2 of Definition (ref) allows. Given these defaults, the value of $\widetilde{c}$ no longer matters, as $R(\kappa_n,\mathcal{G}_n)$ is now proportional to $\widetilde{c}^2$ for all $\kappa_n$.
This leaves the choice of $\widetilde{\gamma}$. Lower values will tend to yield wider minimax radii. Any value $\widetilde{\gamma}$ below the true value of $\gamma$ will cause the estimator to be undersmoothed, ensuring that confidence intervals cover at nominal rates asymptotically.\footnote{Undersmoothing is guaranteed because minimizing $ R(\kappa_n,\mathcal{G}_n)$ will set the standard deviation to scale with a bias term $\mathcal{O}\left(\kappa_n^{-\widetilde{\gamma}}\right)$ that dominates actual bias $\mathcal{O}\left(\kappa_n^{-\gamma}\right)$.} While in principle this might also yield undesirably large radii, the effect is moderated by the fact that when $\widetilde{\gamma}$ is small the worst-case bias is not only large but also insensitive to small changes in the radius.
Taken as a whole, this approach trades off between bias and variance with relatively little calibration; does not depend on any random quantities or on the units in which distance is measured; and does not interfere with inference. Provided the researcher can commit to a lower bound on $\gamma$, it will undersmooth by picking a radius that grows fast enough to make bias dominated, ensuring asymptotic coverage of confidence intervals in large experiments. And we will see shortly in our empirical application that this radius can be surprisingly small in practice: even when we set $\widetilde{\gamma}=0.01$, the resulting radius is substantially smaller than that originally used by EHMNW.
This section presents a central limit theorem and variance estimator in order to construct confidence intervals, focusing on the OLS case (which is more involved). Appendix (ref) provides a CLT for IPW when $\gamma < d$, and we recommend the standard errors from Leung22,leung2025crosscluster.
Proposition (ref) above decomposed OLS's estimation error into a variance term $[\mathcal{A}]$ with expectation zero, a nonstochastic bias term $[\mathcal{B}]$, and a dominated term. Assumption (ref) guarantees that the bias is of order: $[\mathcal{B}]\lesssim \kappa_n^{-\gamma}$. We control the asymptotic distribution of the variance term $[\mathcal{A}]$ using the following CLT.
Theorem (ref) shows that, provided the estimator does not exhibit degenerate variance, any choice of $\kappa_n$ that delivers consistency and grows relative to cluster size will ensure that $[\mathcal{A}]$ is asymptotically normal, no matter how small $\gamma$ or the clusters are. The proof is similar to the CLT in Leung22, but permits any $\gamma >0$.
If we knew $\mathbb{V}\left([\mathcal{A}]\right)$ then Theorem (ref) would allow us to conduct inference. The challenge is that $\mathbb{V}\left([\mathcal{A}]\right)$ is unidentified. Specifically, the variance of OLS is equal to the sum of an identified component and the unidentified term
where $\Lambda$ denotes the adjacency matrix that connects two units if and only if their $\kappa_n$-neighborhoods intersect a common cluster and $\widetilde{A}_{ij}\equiv A_{ij}\mathbf{1}\left\{\mathcal{P}_{c(j)}\cap \mathcal{N}(i,\kappa_n) \neq \emptyset\right\}$ and $c(j)$ denotes the cluster containing unit $j$. $H_n$ is unidentified because we cannot identify the treatment effect of any unit individually---a well-known consequence of unobserved heterogeneous treatment effects in a fixed-population setting (and not of spillovers, or of the choice of estimator, per se). Moreover, $H_n$'s sign is not known, as $\Lambda$ may have both positive and negative eigenvalues.
A few approaches have been proposed to solving the analogous problem in the IPW case. One can always bound the unidentified term using Young's Inequality Aronow17, but this can lead to very conservative confidence sets when clusters are large. Alternatively, if potential outcomes are allowed to be random and treatment effect heterogeneity is not too strongly correlated over large enough neighborhoods, the unidentified term vanishes entirely in large samples Leung22.\footnote{leung2025crosscluster shows that in fixed populations $H_n$ is asymptotically non-negative for IPW when $\kappa_n\sim g_n$. An analogous approach for OLS would not be useful, however, since OLS has a rate advantage only when $\kappa_n/g_n\to \infty$.} To provide a viable alternative for the OLS case, and holding potential outcomes fixed, we provide a sharper bound on $H_n$ than Young's Inequality affords by leveraging Assumption (ref) and (slow) treatment cluster growth. In the following variance estimator the first term has a form analogous to a traditional heteroskedasticity and autocorrelation consistent (HAC) variance estimator, while the second term bounds $H_n$.
Here $Q$ denotes the block-diagonal matrix that connects two units if and only if they are in the same cluster, and
Theorem (ref) shows that this is asymptotically conservative for the variance of $[\mathcal{A}]$, the key step being to show that the second term in ((ref)) is an asymptotically valid bound on $H_n$. We require several additional mild conditions for this to hold. First, there is the standard requirement that the variance of the estimator not converge too fast. Second, we require $\frac{\kappa_n^{2d}+g_n^{2d}}{n g_n^{d}} \to 0$, which is satisfied whenever Theorem (ref) guarantees that OLS is consistent. Third, $\kappa_n$ must grow faster than $g_n$. If this were not the case, then IPW would be preferred to OLS because the two estimators would converge at the same rate. Finally, in order for the upper bound on $H_n$ to hold asymptotically, we need the fraction of the population lying in the “outer crusts" of the clusters to shrink and the clusters to grow in size, but these rates can be arbitrarily slow.
When implementing this approach in small samples we recommend two conservative adjustments. First, we follow leung2025crosscluster in using the maximum of the “HAC" term and the usual cluster-robust standard error where $\Lambda$ is replaced with $Q$. This stabilizes the variance estimator, whose variance can otherwise become nontrivial, when the radius is large. Second, we set the second term to zero if it is numerically positive. This prevents rare cases in which the entire variance can become negative due to volatility in the second term when the radius is large. These adjustments preserve the guarantee from Theorem (ref) that standard errors are asymptotically conservative, while improving coverage in simulation.\footnote{Applying these adjustments, the variance estimator we use in our application is
}
A central theme in the theoretical results has been that large clusters are beneficial for GATE estimation, while an OLS-based approach can be appealing when the original design was not optimized for this. We turn now to an empirical examination of these issues, re-analyzing data from EHMNW. We first cover here the prerequisite task of selecting a radius for the OLS estimator given the design actually implemented. In Section (ref) we then compare the performance of IPW and OLS estimators under a counterfactual design with larger clusters.
EHMNW studied a cluster RCT of cash transfers worth USD 1,871 PPP allocated among 5,419 treatment-eligible households in 653 villages situated within 68 geographic strata formed from 84 sublocations in rural Kenya. The design was as follows: first, half of the strata were randomly assigned to “high saturation" and half to “low saturation." Second, in high (low) saturation groups, 2/3 (1/3) of villages were randomly assigned to treatment. All eligible households in treated villages were then treated. Given this design we interpret households as units, villages as clusters, and (groups of) sublocations as strata.\footnote{We account for this stratified structure using the modified regression described in Appendix (ref) although row 3 of Table (ref) shows that this does not change any point estimate by a meaningful amount.} EHMNW's design can be interpreted as satisfying our Scaling Clusters condition (Definition (ref)) if we imagine that, as the map expanded, they would have expanded strata and assigned treatment within each stratum in growing groups of 2 villages, then 3 villages, and so on. This design was not selected with GATE estimation in mind, however---EHMNW report estimates of the total effect of the intervention as implemented, but not of the GATE---and is in this sense a useful test case for the OLS-based approach.
We use the radius selection procedure from Section (ref) under the default values of the parameters there. We set $\widetilde{\gamma} = 0.01$, which is available (since Theorem (ref) allows $\gamma <d$) but which we view as very conservative. Assuming that some bound $\gamma > \widetilde{\gamma}$ on the rate of spatial decay exists, this ensures our estimator is undersmoothed and that bias adjustment is asymptotically unnecessary for inference. Given this, the minimax method selects a radius of 500m.\footnote{Increasing $\widetilde{\gamma}$ to 0.25, 0.50, 1.50, and 2.00 leads to suggested radii of 500m, 250m, 250m, and 250m respectively.}
Table (ref) presents results, focusing for simplicity on one outcome---annualized household per-capita consumption---from among those that EHMNW consider. It reports estimates from the OLS estimator using the selected 500m radius along with others for comparison, including the 2,000m radius used by EHMNW in constructing their (related but distinct) regressors. It also reports for comparison estimates from an OLS estimator $\widehat{\theta}^{OLS,UNITS}$ constructed along the lines discussed in Remark (ref), where the regressor is not the treated share of nearby clusters but the treated share of nearby units (the method that EHMNW used to construct their measures of neighborhood exposure). All estimates use the same inverse-probability sampling weights as in the original study.\footnote{Table (ref) provides additional comparisons to regression-based estimators that ignore stratification, e.g. one that includes in the regression a separate indicator for own-cluster treatment status. These estimates differ only modestly from our preferred estimates.}
We draw two main conclusions. First, radius selection matters. While all the estimates point to substantial effects, the largest point estimate within the range we consider (at $\kappa_n = 1000$m) is 46% larger than the smallest (at $\kappa_n = 0$m). Of particular interest, the estimates at our selected 500m radius is 22% larger and 65% more precise than that at the 2,000m radius which EHMNW used based on their priors.\footnote{To be precise, EHMNW prespecified that they would include a 2,000m-neighborhood regressor and possibly additional “rings” if this minimized an empirical information criterion.} Second, the consistent estimator $\widehat{\theta}^{OLS,STRAT}$ and the inconsistent estimator $\widehat{\theta}^{OLS,UNITS}$ differ, but only modestly, by at most 14% of the former.
We now turn to questions of cluster size and estimator performance. To that end, this section continues to use the EHMNW data but replaces the outcomes $Y_i$ and treatment assignments $D_i$ with simulated values. Knowing these, we know the true GATE and can compare the performance of IPW and OLS, both under the randomization design that was actually carried out and under a counterfactual one with larger clusters.
The data-generating process for the simulation is as follows. The population again consists of the 5,419 eligible households with their geographic locations, village membership, and sublocation membership unchanged. We impose Assumption (ref) by setting $Y_i(\mathbf{D}) = \beta_0+\sum_{j=1}^n A_{ij}(D_{j}-p)+\epsilon_i$. We impose Assumption (ref) by giving the $n\times n$ treatment effect matrix $A$ the spatial moving average structure $A_{ij} \propto \max\{\rho(i,j),\overline{m}\}^{-d-\widetilde{\gamma}}$ for $i \neq j$.\footnote{Here $\overline{m}$ is the median distance to the closest neighbor and the maximization is to prevent GPS measurement error from generating extremely large spillover values for close-together units.} We set $\gamma = 1.0$, consistent with the lower bound $\widetilde{\gamma} = 0.01$ used above to select a radius but still a very slow rate of decay, in order to challenge the performance of the estimators in a slower-decay environment and generate clearly contrasting performance across radii. We set the diagonal of $A$ such that $A_{ii} = \frac{1}{2}\mathbf{1}'A\mathbf{1}/n$ to generate a multiplier of 2, slightly larger than the fiscal multipliers found by Ramey_multiplier,Gabriel_multiplier. Together these parameters pin down the relative sizes of $A$'s entries. We then calibrate its scale so that the average effect of having one's own village treatment matches the value (\$300) we actually obtain from a regression of the outcomes on own (village) treatment status (i.e., a regression with $\kappa_n = 0$). Finally, we set the deterministic quantities $\beta_0 + \epsilon_i = Y_i^{Factual} - A(\mathbf{D}^{factual}-p)$ where $Y_i^{factual}$ is the outcome and $\mathbf{D}^{factual}$ the vector of treatment assignments observed in the real data. This aims to ensure that they mimic the spatial patterns and magnitudes in the real-world data. Together these parameter choices imply a true value of the GATE of \$410.
For each iteration of the simulation, we randomly redraw the treatment assignment vector $\mathbf{D}$. We randomize using two different designs for comparison. For the factual design, we use the original STATA code from EHMNW to re-draw treatment from the exact same distribution as the actual experiment. For a counterfactual design with larger clusters, we divide each of the 68 EHMNW strata in half (specifically, northern and southern halves) and treat all members of each of the resulting 136 clusters with probability $0.5$ i.i.d. We view this as a natural way to scale clusters while keeping to the spirit of the original design, which was based on pre-existing administrative units, though of course custom clusters are also feasible. It yields a substantial increase in cluster size compared to the factual design, as the 136 clusters consist of 4--5 villages each. For each iteration of the simulation we draw two independent treatment vectors, one for each design; compute outcomes using each of the two treatment vectors; and then estimate the GATE using IPW and OLS over radii $\kappa$ ranging from 1m radius to 2000m.
Figure (ref) displays the results. The left-hand panel shows root mean square error (RMSE), on a log scale, over different radii for the two designs and two estimators. To interpret magnitudes, recall that the true GATE is \$410 and note that the factual outcome variable has standard deviation \$1,802. We draw three main conclusions. First, the IPW estimator has RMSE far higher than that of OLS under the factual, small-cluster design. Second, this performance gap narrows substantially under the counterfactual, large-cluster design. Together these points illustrate the interdependence of design and estimator: IPW can perform well with a design suited for GATE estimation, while OLS performs far better with the small-cluster design EHMNW actually implemented. Third, we note that the RMSE-minimizing OLS radius is 250m, which is shorter than the worst-case-minimizing radius we obtained above (500m) which is itself smaller than that which EHMNW selected based on prior intuition (2,000m). This illustrates the importance of radius selection: using the worst-case-minimizing radius under the factual design, for example, reduces RMSE by 54% compared to the 2000m radius used by EHMNW.
The right-hand panel then shows the coverage of 95% confidence intervals computed using the OLS standard errors defined in Section (ref), again over different radii. Coverage improves as the radius increases, which is as expected since this reduces bias. Using the factual, small-cluster design and the recommended 500m radius we achieve a coverage level of 96.2%. The large-cluster counterfactual design yields coverage that is often closer to the nominal 95% level for all radii.
Appendix (ref) shows analogous figures for alternative values of the spatial decay rate $\gamma$. When $\gamma = 1.5$ (Figure (ref)), performance is better across the board: the estimators yield lower RMSE, and coverage of the OLS confidence intervals improves for small radii. This confirms empirically one of the broad and intuitive themes in the theory, that GATE estimation becomes less difficult in populations that are economically less interconnected. When $\gamma = 0.5$ (Figure (ref)), an extremely slow rate, coverage falters for the tightest radii but remains strong for 500m onward.
Considering these results together, a takeaway for practitioners is that the cluster sizes used in EHMNW---while “large,” by contemporary standards---were nevertheless too small for reliable IPW estimation. At a 1,000m radius, for instance, IPW's RMSE exceeds the true GATE. Assuming that linearity holds (as EHMNW implicitly did) and estimating the GATE via OLS might then be the best option. Alternatively, a design with larger clusters would have mitigated the need for this assumption.
Our analysis of the problem of GATE estimation has found that it is feasible---i.e., that consistent and rate-optimal estimator-design pairs exist---under more general conditions than was previously understood; that jointly optimizing over designs as well as linear estimators does not fundamentally change the rate bound; that there are reasonable justifications for using estimators based on OLS regressions like those used in practice to study large-scale field experiments, particularly when clusters must be “small,” albeit with the tradeoff that convergence may slow down when the underlying DGP is non-linear; and that disciplined radius selection for and valid inference on such procedures is feasible. Applying these methods to data from one such large-scale experiment, we obtain estimates 22% larger and 65% more precise than those obtained using the original authors' choice of radius. Overall, the results demonstrate the viability of GATE estimation without reliance on assumptions about strictly local spillovers, which are often in uncomfortable tension with general equilibrium economic reasoning.
\singlespacing \typeout