Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
101,668 characters · 62 sections · 73 citation commands
Where to Experiment? Site Selection Under Distribution Shift via Optimal Transport and Wasserstein DRO
\setcounter{tocdepth}{4} \setcounter{secnumdepth}{4}
Multi-site experimental studies have become central to cumulative learning and program evaluation across a number of disciplines, as they allow researchers to generate transportable, externally valid estimates of treatment effects that can inform policy-making, theory development, and testing Bloom2017, dunning2019information.
Across political science, economics, education, public health and medicine, multi-site experimental studies are supported by funding bodies with a view to providing insights that both generalize across multiple contexts, and support cumulative learning about a phenomenon of interest. In political science, EGAP's Metaketa initiatives have tested voter information campaigns, community policing programs, and natural resource governance interventions across multiple countries dunning2019voter, dunning2019information, blair2021community, hyde2022metaketa,blair2024crime, slough2021adoption. In economics, J-PAL, inter alia, coordinated the Graduation Program for the ultra-poor across six countries banerjee2015multifaceted and Teaching at the Right Level initiatives across multiple education systems banerjee2017proof, banerjee2016mainstreaming, banerjee2007remedying. Public health and medical researchers regularly conduct coordinated trials such as the WHO Solidarity trial for COVID-19 treatments across 35 countries who2021repurposed, who2022remdesivir, the Women's Health Inititative (WHI) rossouw2002risks and the Antihypertensive and lipid-lowering treatment to prevent heart attack trial (Allhat) Appel2002, rossouw2002risks. Multisite experiments are also common in education research, where interventions are naturally targeted at the level of schools multischool_gottfredson, Raudenbush2020, multischol_causal, Orr2019.
In each of these multi-site experimental designs, researchers face the following problem: given a finite budget and a universe of potential experimental sites, where should they actually conduct an experiment, given their downstream objective of calculating an unbiased, policy-relevant causal quantity with the smallest possible variance?
Further, what should researchers do when deployment populations differ from their observed populations? How should they take into account their limited information about target populations? And how should their decision change when they care about heterogeneity, and ensuring that diverse populations are included in the study sample?
External validity concerns whether findings from an experimental study can be generalized beyond the specific sample, setting, and time period in which the study was conducted shadish2002experimental, Findley2021,EGAMI_HARTMAN_2023. One central goal within the external validity literature is to develop methods for transporting causal estimates from experimental samples to target populations of policy interest, ensuring that conclusions drawn from studies remain valid when applied to new contexts pearl_bareinboim_2011, Pearl_2014, Rudolph2021, Egami2021, rudolph2024improvingefficiencytransportingaverage. Transporting results from one study context is difficult, because it requires us to make invariance assumptions about lack of change between source and target context. In practice, however, we are likely to encounter distribution shift: systematic differences in the distribution of observed covariates between the experimental sample and the target population rothenhausler2023distributionallyrobustgeneralizableinference, jin2024reweightingpredictiverolecovariate,wiles2021finegrainedanalysisdistributionshift, taori2020measuringrobustnessnaturaldistribution, koh2021wildsbenchmarkinthewilddistribution. When the covariate distributions differ substantially, estimates derived from the experimental sample may not accurately represent treatment effects in the target population, undermining the external validity (and practical utility) of research findings.
For the experimental planner, distribution shift can take on a number of concrete forms. Composition of available experimental sites may differ systematically from the policy-relevant population on which the researcher wishes to experiment Allcott_2015. Population characteristics may change in the time period between study planning and implementation Saville2022,bansak2023learningrandomdistributionalshifts. Observed covariates may be measured with error Bound_2001, and minority groups may be systematically underrepresented in selected experimental units Tan2022,hu2024minimaxregretsampleselectionrandomized. In political science, differences in institutional quality BOLD_2018, trust in institutions Cheeseman_2023, rural-urban mix dehejia2019localglobalexternalvalidity, and racial and ethnic context Anoll_Davenport_2024, Hassell_2021, can each be the source of substantive differences in the transportability of conclusions from one context to another. In education, school districts enrolled in an experiment may be unrepresentative of the distribution of schools, leading to overly optimistic assessments of policy impacts Olsen2016, Olsen2022. Suppose, for instance, we observe rural and urban sites, and use census data at a particular time to estimate population totals in both locations. Census data may suffer from selection bias: it may systematically undercount population totals in hard-to-reach, rural areas. There may also be temporal drift: in-migration from rural to urban areas may have shifted the composition of sites. In each case, the data we observe (historic census data) may not be representative of the deployment population: the actual rural-urban mix.
How should researchers account for the routine fact that the data they have collected may not accurately represent the population they are in fact interested in taori2020measuringrobustnessnaturaldistribution,rothenhausler2023distributionallyrobustgeneralizableinference, bansak2023learningrandomdistributionalshifts, cai2023diagnosingmodelperformancedistribution, jin2024reweightingpredictiverolecovariate?
A goal of recent research in site selection is to choose experimental locations that are, in relevant sense, robust to distribution shift, or designed with external validity in mind gechter2024selecting, Egami_Dainlee_2024, olea2024externallyvalidselectionexperimental. There a number of different ways we might want to formalize this idea in practice, using different statistical and theoretical tools.
Weighting-based methods aim to improve the external validity of an estimate by reweighting source data so that it more closely matches a prespecified target population Egami2021,Huang2023,zhang2024minimaxregretestimationgeneralizing. These methods require that the analyst has a specific transport target in mind and has collected covariate data from the target location. They also require assumptions about the stability of the mapping from source to target: that there is a unique map from source to target, which can be estimated in practice.
In contrast, Distributionally Robust Optimization (DRO) is a set of methods developed in operations research that find solution sets with guarantees against worst-case performance within the radius of a given solution Ben-Tal_2013,esfahani2017datadrivendistributionallyrobustoptimization, kuhn2024wassersteindistributionallyrobustoptimization, blanchet2021statisticalanalysiswassersteindistributionally, Blanchet_Murthy_2018, blanchet2024distributionallyrobustoptimizationrobust, duchi2020learningmodelsuniformperformance, levy2020, bertsimas2023distributionallyrobustcausalinference. DRO methods approach the problem of distributional uncertainty by providing statistical guarantees that a given solution is robust to a worst-case shift of the data. These methods provide insurance against poor performance within a specified neighborhood of the empirical solution Luo2020, duchi2020learningmodelsuniformperformance. Instead of asking, under what assumptions can we transport a valid conclusion from context A to context B, these approaches ask, what solution would we pick if we wanted it to still hold for any context that was “sufficiently close” to the context we actually saw?
These methods build on the optimal transport literature, which is an elegant body of applied mathematics that studies the abstract problem of moving (probability) mass from one location to another villani2003topics,villani2008optimal, peyre2019computational, christensen2023optimal, santambrogio2015optimal.
I contrast these methods with sampling-based site selection methods. Throughout this paper, I distinguish between three selection approaches: (1) simple random sampling without stratification, (2) stratified random sampling that randomizes within predefined strata, and (3) optimization-based selection using covariate information. The benefit of simple random sampling is that it does not require prior information about sites, and has optimality and robustness guarantees in general settings. Random sampling is minimax optimal when the analyst has no prior information about experimental units kallus2020optimal, and when the analyst knows the true underlying treatment response only with some error Wu1981.
Stratification compromises between random sampling and use of prior information and randomization Thompson_2022, Parsons2017. Stratification requires splitting the covariate space into strata, where we presume that the stratification occurs along dimensions of high treatment effect heterogeneity, so that the resulting strata capture meaningful variation in treatment response, and guarantee good coverage of the covariate space.
It turns out that Optimal Transport methods can be interpreted as a data-adaptive form of stratification: these methods simultaneously solve for optimal strata, and for optimal representatives within each stratum. I show this formally in (ref). In practice, stratification often involves analyst-driven choices about what to stratify on. And in high dimensions, it becomes less clear how the analyst should make high-dimensional stratification choices (though see Tipton_2013).
Optimization methods instead seek to exploit prior information about experimental sites in the form of covariate information. The goal of the methods outlined in this paper is to choose experiments robust to distribution shift by leveraging what we know about existing sites. This comes with a trade-off: when our prior information is good, that is, highly prognostic, optimization does better. When our prior information is bad, randomization methods are more robust, and have better worst-case performance. This is the so-called `price of robustness' Bertsimas_sim_2004_priceofrobustness, or a version of the no free lunch theorem Wolpert_nofreelunch_1997. I study these trade-offs by simulation in (ref).
A fundamental challenge in site selection is that we typically observe only a subset $P$ of the universe of potential experimental sites $\mathscr{P}$. The distance between $P$ and $\mathscr{P}$ is often unmeasurable, yet this is the population to which we ultimately wish to generalize. While no method can fully address this limitation, our approach provides robustness guarantees for deployment populations within a specified distance of the observed data.
In practice, researchers must play close attention to the informativeness of covariates collected in order to make decisions about site selection shpitser2012validitycovariateadjustmentestimating,VanderWeele2011,Stuart2013,bicalho2022conditional. Optimization can be a powerful tool to aid study design -- if researchers engage in significant efforts to collect data at the planning stage.
\paragraph*{I use optimal transport theory to formulate the problem of selecting sites optimal for the Population Average Treatment Effect and Conditional Average Treatment Effect.} Optimal transport is a rich body of applied mathematics with many possible applications in causal inference and machine learning Villani2003, santambrogio2015optimal, peyre2019computational. Optimal transport is concerned with the efficient shifting of mass between distributions, and gives rise to an intuitive notion of distance between distributions, the Wasserstein distance, which measures the shortest-cost transport distance between two distributions.
The Wasserstein distance quantifies how much “work” is required to transform one probability distribution into another, where work is measured as probability mass times the distance it must travel. Formally, for two distributions $P$ and $Q$, the $p$-Wasserstein distance is $$W_p(P,Q) = \inf_{\pi} \left(\int \|x-y\|^p d\pi(x,y)\right)^{1/p}$$ where the infimum is taken over all transport plans $\pi$ with marginals $P$ and $Q$. Intuitively, if we think of $P$ as describing the locations of piles of sand, $Q$ as describing where we want to move that sand, and $\pi$ as any given set of paths used to move sand from $P$ to $Q$, the Wasserstein distance gives the minimum total cost of the move under the best routing from $P$ to $Q$.
This metric is particularly well-suited for site selection because it directly captures the representativeness of selected sites. When we select experimental sites, we want them to “represent” the broader population in the sense that every population unit is adequately proxied by nearby selected sites. The Wasserstein distance formalizes this intuition: it measures how well a sparse set of selected sites can approximate a dense population by finding the optimal assignment of population units to selected sites while minimizing total “representation error.”
For the special case of $p=1$, this cost equals the population-weighted average distance that points must travel under the optimal assignment. For $p=2$, we minimize the sum of squared distances, so $W_2^2$ equals the population-weighted average squared distance, and $W_2$ itself is the square root of this quantity analogous to a population-weighted Euclidean distance. Just as standard deviation captures spread differently than mean absolute deviation, $W_2$ penalizes outliers more heavily than $W_1$: leaving any population point far from its nearest selected site contributes quadratically rather than linearly to the total cost.
In our site selection context, $W_1(P_X, S_X)$ measures the average distance from population units to their assigned experimental sites, while $W_2(P_X, S_X)$ is the root mean square Euclidean distance under optimal assignment, being more sensitive to ensuring no subpopulation is left too far from representation.
\paragraph*{I derive new upper bounds on the errors of the PATE and CATE estimator in terms of Wasserstein distances.} By using the tools of optimal transport to analyze the Mean Squared Error of the PATE estimate, and the Precision in Estimated Heterogeneous Effect Hill_2011, shalit2017estimatingindividualtreatmenteffect, I derive upper bounds for the PATE and CATE errors in terms of the Wasserstein distance (Theorem (ref) and (ref)).
\paragraph*{These bounds give us intuition about what our substantive goals are when choosing experimental sites for PATE and CATE estimation.}
When estimating the PATE, we seek a single number: the average treatment effect across the population. This means we want selected sites that, when averaged together, closely approximate the population average. Think of this as finding sites whose collective “center of gravity” matches the population's center.
When estimating the CATE, we want to estimate an entire function: how treatment effects vary across different covariate values. This requires accurate interpolation across the entire covariate space. We need sites spread throughout the population to avoid large gaps where we must extrapolate rather than interpolate.
These different goals lead to different selection strategies. For the PATE, the 1-Wasserstein distance naturally emerges because we care about average representation. Sites can compensate for each other, and modest coverage gaps in outlying areas are acceptable as long as the average is well-represented. For the CATE, the 2-Wasserstein distance emerges because outlying regions contribute quadratically to estimation error: leaving any subpopulation far from a selected site severely degrades our ability to estimate treatment effects in that region.
Both optimization problems seek to create balanced partitions of the covariate space that maximize representativeness, but they do so using different distance metrics that encode different notions of what “good representation" means.
\paragraph*{These upper bounds motivate a Mixed Integer Linear Program formulation of the PATE and CATE selection problems.} Because our bounds contain Wasserstein distance terms, our objective then becomes to choose experimental sites that minimize the Wasserstein distance between the observed population of experimental sites and the selected sample of experimental sites, subject to a budget constraint of sites. Wasserstein distance minimization can be tractably reformulated in terms of Mixed Integer Linear Programs. These are straightforward to solve using commercial solvers like Gurobi. I develop software to implement this approach.
\paragraph*{Empirical Performance versus Randomization and Optimization} These optimization-based methods outperform simple random sampling when covariates are sufficiently informative about treatment effects, as I show via simulation in Section (ref). A critical finding from our analysis is that optimization-based site selection requires observed covariates to explain more than approximately $50\%$ of treatment effect variation ($R > 0.5$). This threshold has important practical implications: researchers should validate covariate informativeness before investing in optimization-based selection, as uninformative covariates can lead to worse performance than randomization.
\paragraph*{I extend the site selection problem using Wasserstein distributionally robust optimization (DRO).} Rather than optimizing for the observed distribution, we can solve a more conservative problem that hedges against a set of plausible population distributions. Formally, our problem becomes: $$\min_{S:|S|\leq K} \sup_{P' \in \mathcal{B}(P,\rho)} W_p(P, S)$$ where $\mathcal{B}(P,\rho) = \{P' \: : \: W_p(P, P') \leq \rho \}$ is an ambiguity set: the collection of all population distributions within radius $\rho$, measured in terms of the Wasserstein distance, around the empirical distribution. This provides worst-case performance guarantees when the true population lies within $\rho$ of the observed data.
The ambiguity radius $\rho$ is a scalar that represents the total “transportation budget” available to an adversary that seeks to perturb the observed distribution. Specifically, $\rho$ bounds the total cost of moving probability mass in the covariate space, measured in the same units as the covariates themselves. For example, if covariates are standardized, then $\rho = 0.5$ allows Nature to move each population unit up to $0.5$ standard deviations on average, or to make larger moves for some units while keeping others fixed, so long as the total transportation cost remains within budget. This provides worst-case performance guarantees when the true population lies within $\rho$ of the observed data.
\paragraph*{I solve the Wasserstein DRO site selection problem using a novel cutting-plane algorithm\footnote{A cutting-plane algorithm solves optimization problems by iteratively adding constraints that eliminate infeasible regions. Rather than solving the full problem at once, the algorithm starts with a simplified version, finds a candidate solution, then checks if this solution satisfies all constraints of the original problem. If not, it adds a new constraint (a “cut”) that rules out this solution and similar infeasible ones, then resolves the simplified problem. This process continues until the candidate solution satisfies all original constraints bradley1977applied.} that exploits the minimax game structure of the optimization problem.} Formulating the DRO problem as a game theory problem directly suggests an algorithm for its implementation: the Researcher chooses a site selection; the adversary perturbs the observed data, subject to a budget on how far it can move points; the Researcher observes the adversary's new site selection and resolves the problem; and so on until neither the adversary nor the Researcher change their choices. (See (ref) for an explicit description of the equivalence.) Here, the Wasserstein DRO solution is interpretable as Nash Equilibrium in a game between Researcher and Nature; the algorithm proceeds by `playing' the game between Nature and the Researcher until there are no further moves left. This removes the need to enumerate all elements of the (infinite) Wasserstein ball; instead, we identify only the set of adversarial best responses to a given site selection.
This game-theoretic cutting-plane approach is novel in the Wasserstein DRO literature. Existing methods for solving Wasserstein DRO problems typically rely on dual reformulations that convert the minimax problem into a single optimization esfahani2017datadrivendistributionallyrobustoptimization, entropic regularization techniques that approximate the Wasserstein distance using Sinkhorn iterations to make the problem computationally tractable cuturi2013sinkhorn, or moment-based approaches that replace Wasserstein constraints with simpler moment constraints gao2020wassersteindistributionallyrobustoptimization.
The key insight of this approach is that we never need to characterize the full (infinite) ambiguity set $\mathcal{B}(P,\rho)$. Instead, we exploit the sequential structure: at each iteration, Nature reveals only the single adversarial distribution that is a best response to the current site selection, and we accumulate these best responses over iterations.
\paragraph*{I introduce a novel data-adaptive procedure for selecting the uncertainty radius in Wasserstein DRO problems.} A separate technical contribution is the introduction of a novel data-driven calibration method for selecting the robustness parameter $\rho$. A fundamental challenge in applying distributionally robust optimization is choosing an appropriate robustness radius: too small provides insufficient protection against distribution shift, while too large yields overly conservative selections that sacrifice performance. Theoretical results provide guidance on how to select a robustness radius in the presence of sampling variability, based on the rate of convergence of empirical measures fournier2013rateconvergencewassersteindistance,Blanchet_2019,blanchet2021statisticalanalysiswassersteindistributionally. However, it is difficult to formulate a theoretically principled way to choose a robustness radius in the face of unknown distribution shift beyond sampling variability: by design, we intend to guard against out-of-sample shifts, and so are limited in how we can use in-sample data to construct a plausible radius. This is because distribution shift in the wild induces Knightian Uncertainty Knight1921, Sunstein2023: we cannot really know, without making assumptions, how much shift to guard against.
An alternative approach is to provide the option to guard against shifts that are benchmarked by the observed variation in the data. My procedure, detailed in Section (ref), first constructs an empirical Wasserstein grid based on empirical distances in the covariate data. Intuitively, given any data set, there is a maximum radius beyond which an adversarial solution will not change. This motivates the heuristic procedure of 1) greedily searching for the maximum radius $\rho^{\text{max}}$ and 2) performing adaptive grid search over the line $[0, \rho^{\text{max}}]$. Site selection methods will produce different solution sets over this line: the goal is to identify when the output solutions exhibit small, moderate, and large differences from the baseline solution set. We can then define a series of $\rho$ thresholds in terms of these different solution sets. Rather than requiring the user to specify $\rho$ values, this procedure automatically generates $\rho$ values that answer the question, “What would small, medium, and large distributional shocks look like for my specific dataset?.” This makes DRO methods useful for practitioners without the need for arbitrary priors about size of the robustness radius.
\paragraph*{Empirical performance} I demonstrate the performance of these methods by reanalyzing Crepon_2015, who conduct a randomized microcredit experiment in Morocco, in which rural villages were randomized into receiving access to loans. I use as an outcome profits earned by individuals who did and did not take out the loan, and generate semi-synthetic treatment effects using observed covariates and a linear model. I first study the properties of site selections generated by my proposed methods, SPS, and random and stratified sampling on the full sample, evaluating the performance of these methods in terms of the $MSE_{PATE}$ and the PEHE. I then implement a simulation study, in which treatment effects vary with signal strength (the informativeness of observed covariates), and in which I induce distribution shift by moving observed covariates away from their actual values. I show that my nonrobust methods outperform SPS under distribution shift, and in high-signal environments.
This paper introduces four methods for different practical use cases in site selection. First, the researcher should decide whether they are interested in PATE estimation or CATE estimation. Second, the researcher should decide how concerned they are about distribution shift: are they willing to pay `the price of robustness' Bertsimas2004 to trade-off accuracy in minimizing observed error against potential unobserved distribution shifts?
Egami_Dainlee_2024 introduced explicit optimization methods for site selection in political methodology, and contributed significantly to defining the problem of site selection. Their approach, based on the synthetic control method, uses optimization to select included sites that closely approximate sites that are not included in the selection, by estimating balancing weights Abadie_2003, abadie_2010, abadie2025. The goal is to have a high-quality weighted average representation of non-selected sites; in practice, this can be thought of as ensuring that non-selected sites are within the convex hull of selected sites. The default implementation contains a penalty term that additionally penalizes using outlying sites in the final selection.
The goal of this paper is to use a set of different technical tools to address the site selection problem motivated by Egami_Dainlee_2024. Whereas they use an approach based on synthetic controls intended to select experiments for the PATE, I i) show that the PATE and CATE have different optimization problems ii) use the theoretical resources of optimal transport to state and implement the minimization problem iii) use Wasserstein Distributionally-Robust Optimization to induce robustness to distribution shift.
Tipton_2013, Tipton2013b propose a cluster-then-stratify approach to site selection, which we study via simulation, and is weakly dominated by 2-transport, as I show in (ref).
olea2024externallyvalidselectionexperimental solve the site selection problem, by defining it as the $k$-median problem. This is similar to the PATE transport solution, but the PATE solution implicitly imposes a balance constraint: that each site receive $\frac{1}{K}$ of the overall population mass. $k$-medians is not constrained in this way.
Optimal transport has a large number of possible applications for core causal inference tasks Galichon_2016. Studying the changes-in-changes model Athey2006_CiC, torous2024optimaltransportapproachestimating use optimal transport methods to estimate control group trends over time, and apply this same transformation to predict what the treatment group would have looked like without intervention. charpentier2023optimaltransportcounterfactualestimation propose using optimal transport methods to estimate counterfactual distributions, while dunipace2022optimaltransportweightscausal use optimal transport methods to solve IPW-type problems hajek:1971, Horvitz-Thompson, benmichael2021balancingactcausalinference.
The conceptual background of this paper is closely related to Response Surface Methodology Box_1951,Box_1975, box1987empirical. In RSM, the goal is to choose experiments based on their location on the surface that determines how covariates map onto outcomes. This yields applied optimization problems, where we want to learn, say, the maximum of a given output function given inputs: this may correspond to an efficient configuration of industrial inputs, for instance. In our context, we can think of the treatment effect surface $\tau(X)$ as our response surface, and note that we want to choose experiments that are informative about the treatment effect surface, in a sense we will explore below.
Section 2 motivates the problem of site selection, and studies the case where the population of sites is observed, describes the assumptions needed to use covariates to select sites, states theoretical upper bounds on the downstream errors in estimating the PATE and CATE due to site selection, formulates the optimization problems associated with each estimand, and states algorithms to implement each procedure. Section 3 describes the application of Wasserstein DRO to the problem, motivates robust upper bounds, and describes a cutting-plane algorithm to implement Wasserstein DRO that leverages a game theoretic interpretation of the DRO problem. Section 4 studies the behavior of the site selection procedures by simulation. I study the performance of the methods against randomization as a function of signal strength, and show that these methods have good performance relative to randomization methods even for relatively weak signal strengths. I also characterize the robustness behavior of Wasserstein DRO empirically, and show that increasing the robustness radius in practice increases the coverage of the selected set. Section 5 reanalyses Crepon_2015, an experiment in Morocco that randomized encouragement to access microcredit. I generate semi-synthetic treatment effects based on this data, and assess the behavior of the optimal transport and DRO methods compared to Synthetic Purposive Sampling and randomization methods as a function of problem size, signal strength, and distribution shift. Section 6 concludes.
Consider a researcher who is faced with a universe of sites $\mathscr{P}$, from which they must choose a subset $S$ of sites, subject to the constraint that they can choose at most $K$ sites.
The researcher's goal is to choose $K$ sites that `best represent' the population $\mathscr{P}$, in a sense that we will consider more specifically below.
We can formalize this by saying that the researcher must choose $K$ sites that minimize a specific objective problem. The researcher is interested in the results of a downstream analysis of an experiment: they will eventually conduct an experiment and get an estimate of their population estimand of interest. The goal is to minimize the error of this estimate of the population quantity by selecting the `best' sites at the planning stage of the experiment.
The choice of objective function depends on the research context. First, the researcher must choose an estimand: they may be interested in the Population Average Treatment Effect (PATE), or the Conditional Average Treatment Effect (CATE).
For notational simplicity, I will write $\tau \equiv Y(1) - Y(0)$ and $\tau(x) \equiv Y(1) - Y(0) | X = x$, which are related by $\tau = \int \tau(x) dx$.
These represent fundamentally different statistical objectives that lead to different site selection strategies. In selecting the sites for the PATE, the downstream task is to estimate a functional: we seek to estimate a single number $\tau = E[Y(1) - Y(0)]$ that summarizes the average treatment effect across the population.
Estimating the CATE is a function estimation problem: we seek to estimate the entire function $\tau(x) = \mathbb{E}[Y(1) - Y(0)|X = x]$ that describes how treatment effects vary across the covariate space.
This distinction has direct implications for site selection. For parameter estimation (PATE), we want sites that provide an efficient estimate of the population average. This requires representative sampling that balances coverage of different population subgroups. For function estimation (CATE), we want sites that enable accurate interpolation of $\tau(x)$ across the entire support of $X$. This requires broad coverage of the covariate space to minimize extrapolation error when predicting treatment effects at unobserved covariate values.
First, consider the case where the full population of sites is known to the researcher, the researcher has collected covariate information about all possible sites, and they can choose to run an experiment in any of those sites.\footnote{The formal analysis in (ref) does not make this assumption.} This describes the case where $P = \mathscr{P}$. In this case, the expectations described in Definitions (ref) and (ref) are taken over the observed subpopulation $P$, because the population and subpopulation exactly coincide.
The below errors are `downstream', because they are not realized until the analyst actually conducts the experiment. These quantities can be defined in advance of the experiment, however, and the infeasible problem that the analyst would like to solve can be stated.
For the PATE, we suppose that the researcher wants to minimize the Mean Squared Error of the downstream treatment effect estimate:
Where the expectation is taken over randomness in treatment assignment and downstream estimation.\footnote{Note that $S$ here is a set, not an index: we are optimizing over possible selections $S$, and calculating the PATE given that site selection.}
For the CATE, we suppose that the researcher wants to minimize the expected Precision in Estimation of Hetereogeneous Effect Hill_2011, shalit2017estimatingindividualtreatmenteffect.
This gives us the researcher's minimization problem:
Because these errors are downstream, they are unobserved, and this exact minimization problem is infeasible. We can, however, use covariates to study feasible versions of these problems, and provide guarantees about how close the solution to these feasible problems are to the infeasible problems.
In words, treatment effects vary across the covariate space, making site selection based on covariates meaningful.
This stipulates that covariates have the same effect on treatment effect values across sites.
This ensures that treatment effects vary smoothly with covariates. When covariate values change, treatment effects must vary within an envelope defined by the size of the change of covariate values. This assumption is important, because it allows us to move from claims about covariates to claims about treatment effects.
In the next section, we use the tools of optimal transport to derive bounds on the errors of the $MSE_{PATE}$ and $PEHE$. First, I introduce some terminology and notation, and a brief sketch of relevant concepts needed to state and solve our minimization problem. Optimal transport is a powerful methodological framework with broad application to problems in causal inference.
Optimal transport is concerned with moving mass between a source and a target in the most efficient way. An original motivating example, known as the Monge-Kantorovich Problem monge1781,ambrosio2003optimaltransportmapsmongekantorovich, Vershik2013, can be heuristically described as follows. Given a set of Parisian bakeries with specific production schedules and a set of cafes with specific consumption demands, located across Paris, what is the most efficient way to route bread from bakeries to cafes that minimizes the total transport distance? A transport map formalizes the idea of one possible solution to this problems: a collection of routes from bakeries to cafes, stored as a matrix. More formally, we have:
Where $\delta$ is the Dirac delta. Notice that $\pi_{ij}$ has row sums equal to $p_i$, the total mass of the empirical distribution $P_X$, and column sums equal to $r_i$, the total mass of the empirical distribution $Q_X$. In order to evaluate different transport plans, we need a way to assess the costs of a given proposed transport plan. A cost function describes the cost of travelling from $X$ to $Y$. We use $\ell^p$ distances as our cost function, so that $c(X,Y) = d_p(X,Y) = ||X - Y||^p$. For $p = 1$, this gives us the absolute distance, and for $p = 2$, this is the squared distance between $X$ and $Y$.
The optimal transport plan is the plan $\pi^*$ that in fact minimizes the distance between $P$ and $Q$, for a given cost function $c(X, Y)$. That is,
That is, if $\pi^*$ minimizes the cost of transporting mass from $P$ to $Q$ measured in the $p$-norm.
We can think of the solution to the optimal transport as being the shortest possible distance between $X$ and $Y$, given the distributions P and Q. The $p$-Wasserstein distance formalizes the notion of the shortest possible distance between $P$ and $Q$, and is specified in terms of an optimal transport plan:
In our bakery example, this is defined in terms of the best possible solution to the routing problem between bakeries and cafes. Note the duality between the Wasserstein distance and transport plans: the Wasserstein distance is the shortest distance from $P$ to $Q$.\footnote{A political science versions of the optimal transport problem. Suppose we have a set of precincts and a finite set of campaign workers with different home locations. What is the most efficient way to assign campaign workers to precincts to minimize total distance traveled? }
I use the tools of optimal transport to derive upper bounds on the site selection problem: the Wasserstein distance is central to the theory that follows. I use $P_X$ to denote the empirical distribution of covariates in the population, and $S_X$ to denote the empirical distribution of covariates in the sample.
In order to minimize the error on the $MSE_\text{PATE}$ and PEHE, we want to find a feasible upper bound on the problem that we can minimize via an optimization procedure. I derive two such bounds below. These bounds have the following properties:
\paragraph*{The bounds do not depend on a specific model of treatment effects.} That is, they are generically applicable to any site selection problem (as long as treatment effects vary smoothly with covariates). \paragraph*{The bounds make explicit the role of unmeasured heterogeneity.} This allows us to be explicit about what our site selection tools can and cannot achieve, and to assess their performance under unmeasured heterogeneity empirically.
We can upper bound the errors of the $MSE_\text{PATE}$ and the $PEHE$ by the 1-Wasserstein and 2-Wasserstein Distances between $P_X$ and $S_X$, respectively.
In each case we have a sensitivity parameter $\eta_p$, which measures how much the conditional distribution of unobserved covariates $U$ differs between population $P$ and selected sample $S$, given observed covariates $X$, which I call unmeasured heterogeneity.\footnote{This captures two factors: signal-to-noise ratio, or how much treatment effects depend on unobserved $U$; and unobserved covariate shift: how the distribution of $U$ between population and sample differs. The experimental planner observes only covariates and wants to know i) whether $X$ is sufficient for treatment effects and ii) whether their sample is similar to the population on unobserved dimensions. This is unmeasured heterogeneity in the sense of site-level selection bias, rather than the more usual individual-level treatment assignment bias.}
Where $\eta_1 = \mathbb{E}_{P_X}[W_1(P_{U|X}, S_{U|X})]$ represents the degree of unmeasured heterogeneity, and $\sigma^2_S$ represents irreducible estimation error.
Note that $\eta_1$ conditions on observed covariates $X$, capturing only the residual unobserved variation. When unobservables are independent of observables ($U \perp\!\!\!\perp X$), we have $\eta_1 = W_1(P_U, S_U)$, the full distance between unconditional distributions. When unobservables are perfectly predictable from observables ($U = f(X)$), we have $\eta_1 = 0$. Thus $\eta_1$ automatically adjusts for observable-unobservable correlation, representing only the unobserved heterogeneity that remains after accounting for what we can measure.
Where $\eta_2 = \mathbb{E}_{P_X}[W_2(P_{U|X}, S_{U|X})]$ represents the effect of unmeasured heterogeneity, and $\sigma^2_S$ represents irreducible estimation error.\footnote{ Why 1-Wasserstein for the PATE and 2-Wasserstein for the CATE? There is both a technical explanation and a substantive explanations. In the proofs of (ref) and (ref), we get two upper bounds. In the first case, we note that the difference in estimated ATEs is a difference of linear functionals, and apply Kantorovich-Rubinstein to this difference. This is upper bounded by ther 1-Wasserstin distance.
In the second case, the PEHE is the integral of the squared pointwise errors in estimating $\tau(x)$ over $X$. The intuition is that squared pointwise errors $|\tau(x) - \hat{\tau}(x)|$ are bounded by $L||x - y||$ by Lipschitz continuity, so squared pointwise errors are bounded by $L^2||x - y||^2$; integrating both sides yields the 2-Wasserstein distance.
Another way to compare this is that in the PATE case, we are interested in linear function approximation, which yields linear penalties. In the CATE case, we are interested in an integral of squared pointwise errors -- which has the same form as the 2-Wasserstein distance by construction. }
\paragraph*{These bounds allow us to specify site selection as an optimization problem.} The goal of these bounds is to find a feasible target for us to minimize via optimization. In both cases, our losses are upper-bounded by:
$$W_p(P_X, S_X) \: \text{for } p \in \{1,2\}$$
The $p$-Wasserstein distance between empirical distribution of covariates in the population and the sample. It is straightforward to minimize this quantity by choice of $S$ using linear programming, as I show below.
\paragraph*{Optimal site selections for the PATE and CATE differ.}
These bounds also help us to understand the difference in goals between selecting sites optimal for the PATE and selecting sites optimal for the CATE. The $1$-Wasserstein distance places more weight on location, rather than variance; whereas the 2-Wasserstein distance more heavily penalizes outliers.
\paragraph*{The bounds include sensitivity parameters $\eta_p$, which describe the effect of unobserved heterogeneity.}
Specifically, varying $\eta_p$ through simulation, we can empirically assess when site selection methods outperform sampling, which, because they are randomized, are broadly robust to unobserved heterogeneity.
This also allows us heuristically to think about the role of data collection in the site selection process. In the best case scenario, when we have perfect data collection, covariates are sufficient for treatment effects, so that $\eta_p = 0$, and site selection using observable covariates is a good idea. In the worst case, observed covariates are completely uninformative about unobserved covariates, so that $U \perp\!\!\perp X$, and $\mathbb{E}[W_p(P_{U|X}, S_{U|X})] = \mathbb{E}[W_p(P_{U}, S_{U})]$.
The bounds derived in the previous section give us clear objectives. If we want to select sites optimal for the PATE, we choose the sites $S$ that minimizes the 1-Wasserstein distance between the empirical distribution of covariates in the selected sites $S_X$ and the empirical distribution of the covariates in the population $P_X$. For the CATE, we select the sites that minimize the 2-Wasserstein distance.
From (ref) we now have the following optimization problem to minimize the upper bound on $MSE_{\text{PATE}}$: $$\underset{S}{\min} \quad W_1(P_X, S_X) \quad \text{subject to} \quad |S| \leq K$$
An important result from Kantorovich2006 is that optimal transport problems are linear programs: that is, we can find optimal transport plans by writing out and solving a corresponding linear program (minimize $\sum c_{ij}\pi_{ij}$ subject to marginal constraints) using the optimization toolkit.
To solve our Wasserstein distance minimization problem, where we want to select discrete numbers of sites, we can therefore formulate it as a Mixed Integer Linear Program (MILP). Define the site selection indicator $s_i = \mathbb{I}\{s \in S\}$. Then, our optimization problem is:
\paragraph*{Implementation details}
Because the optimal transport problem can be written as a linear program, it can be implemented and solved directly. I implement this using the R Optimization Infrastructure (ROI) framework with multiple solver backends. The primary fallback solver is GLPK, which is freely available and provides reliable solutions for moderately-sized problems. For larger instances, the implementation calls Gurobi, a commercial solver that typically provides faster solution times and better numerical stability. The solver selection is automatic: the code attempts to use Gurobi if available, falling back to GLPK otherwise.
I use LP relaxation and warm starting to improve computational performance for larger problem instances. LP relaxation replaces the binary site selection variables $z_j \in \{0,1\}$ with continuous variables $z_j \in [0,1]$, converting the converting the MILP to a linear program that can be solved in polynomial time. This relaxed solution then provides a warm start for the exact MILP solver by initializing binary variables to rounded values of the relaxed solution. Runtime experiments show this makes a significant difference in practice (see (ref)). For problems with $n > 100$ sites, LP relaxation is used as the default.
In the previous section, we studied the problem of selecting sites optimal for the PATE and the CATE given observed information about the covariates. We can think of this as the full-information case: we assume that we have good knowledge of the data-generating process that determines treatment effects, and can have enough information to actually minimize the MSE of the PATE and the PEHE.
Now, however, we explicitly take account of the fact that the empirical distribution $P_X$ is not guaranteed to be a perfect representation of the underlying distribution that generated the data.
To motivate Wasserstein DRO, we first define an ambiguity set, or Wasserstein ball:
We can incorporate our uncertainty about the underlying distribution into our optimization problem via the ambiguity set. In particular, we want to minimize the worst-case risk\footnote{Note that this differs from the sense of worst-case risk described in Egami_Dainlee_2024. They mean that they optimize an upper bound analogous to our results in the previous section; here I mean that we minimize the risk over an adversarially chosen distribution in the ambiguity set.}, in the following formal sense:
Where, by plugging in $p \in \{1,2\}$, we recover the site selection problems for the PATE and CATE respectively.
Wasserstein DRO has a useful game-theoretic interpretation. Writing out the DRO problem again, we can see:
$$\underbrace{\underset{S\: : |S| \leq K}{\min} \: \overbrace{\underset{P \in B(P, \rho)}{\sup} W_p(P, S)}^{\text{Inner problem: Nature selects worst-case distribution}}}_{\text{Outer problem: Researcher selects sites}}$$
The inner supremum is an action by adversarial Nature, to choose the worst-case distribution $P$, subject to the constraint that they can reallocate mass equal to at most $\rho$. In practice, this means that Nature can choose to relocate points adversarially (in practice, as outliers), selecting the worst-case distribution Q, and our result will still represent a valid upper bound on the chosen minimand. The outer minimization represents our best response to this adversarial perturbation. In short, $\rho$ represents the budget of covariate shift that the researcher wishes to insure against.
This game-theoretic interpretation is not just a point of theoretical interest: it in fact motivates the algorithm I use to implement the DRO version of site selection.
We have:
The ambiguity set $B(P, \rho)$ is built constructively out of Nature's best responses to the Researcher's site selections. We do not need to enumerate all elements of the Wasserstein ball, which is an infinite set; we need only enumerate the adversarial perturbations that increase the Researcher's observed loss.
Incorporating the robustness parameter $\rho$ allows to describe new bounds on our estimates. This gives us the robust upper-bounds:
Where these guarantees are given over a Wasserstein ball\footnote{Constraining shifts to be within a Wasserstein ball simply limits the total mass that can be moved around, and specifies a cost -- either an $\ell^1$ or $\ell^2$ penalty, depending on the estimand -- for doing so.} around the observed distribution. DRO ensures that the solution is robust to distribution shift -- that is, robust to changes in the distribution of observed covariates. We can also think of this as measurement error: our solution should be robust to a specified degree of mismeasurement $\rho$. This is in contrast to the parameter $\eta_p$, which represents outcome model error due to unmeasured heterogeneity. procedure.
The Wasserstein ambiguity set $B(P_X, \rho) = \{Q : W_p(Q, P_X) \leq \rho\}$ contains all distributions that can be reached by moving the observed covariate distribution's mass by at most $\rho$ units. Each distribution $Q$ in this set represents a different way our observed site characteristics could be wrong: measurement error, temporal drift, or systematic misrepresentation of the target population.
The core idea is that an adversary creates gaps in how representative our sample is by strategically relocating probability mass. $\rho$ represents uncertainty about where the population is located in covariate space. The larger the budget $\rho$, the more mass the adversary can relocate to regions poorly served by our specific site selection.
The “worst case” distribution $Q^*$ is the one that maximally exploits differences in site characteristics. For example: the worst-case distribution might concentrate all mass in rural extremes, given an initial selection of urban sites.
By optimizing against the worst case, we obtain a site selection that is robust to every distribution in the ambiguity set. This is because our selection must perform well against the adversary's best response: which means it performs at least as well against any other distribution the adversary could have chosen. In this sense, Wasserstein DRO provides insurance against all possible covariate shifts of magnitude $\rho$, representing all the ways we could have mismeasured the true site characteristics.
The robustness parameter $\rho$ controls shifts in observed covariates, while our bounds include an additive term $\eta_p = E_{P_X}[W_p(P_{U|X}, S_{U|X})]$ capturing unobserved heterogeneity. This formulation already accounts for observable-unobservable correlation: when $X$ and $U$ are independent, $\eta_p$ equals the Wasserstein distance between unconditional distributions; when they are perfectly correlated, $\eta_p$ approaches zero. The additive structure $(W_p(P,S) + \rho + \eta_p)$ separates robustness to observable shifts ($\rho$) from residual unobserved heterogeneity ($\eta_p$).
In practice, the adversary shifts the entire observed distribution, including components correlated with unobservables. This means that choosing $\rho$ based on empirical variation provides implicit protection against correlated unobserved factors. This is not explicitly stated in the bounds, which are conservative and describe the pessimistic case where $X$ and $U$ are orthogonal. When $X$ and $U$ are correlated, the effective robustness exceeds what the additive bound suggests, as the $\rho$-ball constrains both observable variation and its correlated unobservable components.
How should one choose the degree of robustness in practice? Previous work uses theory based on bootstrapping to estimate a radius based on observed variation in the data Blanchet_Murthy_2018. A problem with this approach is that it essentially assumes away distribution shift: bootstrapping relies on asymptotics based on resampling, where the bootstrap distribution converges to the underlying distribution of the data Bickel1981,EfroTibs93.
In (ref), I describe a procedure for selecting $\rho$ based on finding `stable' solution sets at different levels of empirical variation in the data, and allowing the user to specify what degree of variation they would like the solution to be robust to.
The idea is to run grid search over possible values of $\rho$, and evaluate the stability (in terms of Jaccard stablility) of the solution set as $\rho$ changes. Intuitively, there must be a maximum value, $\rho^{\text{max}}$, beyond which the solution set cannot get any more `extremal'; the first step of the algorithm is to do a greedy search for this value. Then, given a value $\rho^{\text{max}}$, the algorithm looks for stable solution sets on the interval $[0, \rho^{\text{max}}]$, and outputs them, corresponding to ordered degrees of robustness.
The simulations address three questions: (1) How do solution sets differ between 1-Wasserstein (PATE)and 2-Wasserstein (CATE) optimization? (2) How do site selection solutions change as the robustness parameter $\rho$ increases? and (3) How do optimization methods compare to selection methods based on random sampling and stratified sampling?
First, I provide visual characterizations of solution sets generated by different objectives, illustrating how the 1-Wasserstein objective (PATE) trades off between central location and coverage while the 2-Wasserstein objective (CATE) more heavily penalizes leaving any region uncovered. Second, I plot the evolution of solution sets, using synthetic data, as the parameter $\rho$ changes. These simulations show that increasing $\rho$ increases the size of the convex hull spanned by the solution set. This comes at a price, however, as the `naive' PATE estimate taken by aggregating sites is shifts as $\rho$ increases.
Finally, I use simulations to evaluate when optimization-based selection outperforms randomization approaches. For these performance comparisons, I generate candidate populations with covariates $X_s \sim \mathcal{N}(0, I_5)$ and site-level treatment effects $\tau_s = \sqrt{1-\eta^2}f(X_s) + \eta\varepsilon_s$, where $\eta \in [0,1]$ controls the signal-to-noise ratio: the fraction of treatment effect variation unexplained by observed covariates. I compare simple random sampling (uniform selection), stratified random sampling ($k$-means clustering followed by within-stratum sampling), and the optimization methods across varying covariate signal-to-noise ratios. The key finding is that optimization methods dominate when $\eta < 0.7$ (equivalently, when observable covariates explain more than 50% of treatment effect variation).
The above bounds show that there are different site selection objectives for the PATE and the CATE. In the PATE case, we care about the 1-Wasserstein Distance, and in the CATE case the 2-Wasserstein distance.
Recall that the 1-Wasserstein distance contains the absolute norm, and the 2-Wasserstein distance is a function of the $\ell^2$ norm. This entails that while the cost of increasing distance is linear in the 1-Wasserstein case, the cost of increasing distance from unselected points to selected points is quadratic in the difference of distances.
This should penalize selections that are far away from unselected points more in the 2-Wasserstein case, leading to a more compact set for the 1-Wasserstein solution and a larger set for the 2-Wasserstein solution.
This is intuitively appealing in the causal inference context, since the 1-Wasserstein distance is associated with the PATE, where our best guess of the PATE is the centroid of our observed sites. The CATE problem involves estimating a function over the support of X, and so, intuitively, we would want a solution set with improved coverage over the support of X.
To test these theoretical predictions, I generate synthetic datasets with known covariate distributions and compare the geometric properties of optimal site selections under both objectives. The simulation uses $|P| = 30$ candidate sites distributed across a two-dimensional covariate space, from which $K = 5$ sites are selected.
In practice, for small-sized problem instances, the solution sets are fairly similar. This is because, for sufficiently well-behaved data, site selections that minimize the 1-Wasserstein distance also minimize the 2-Wasserstein distance and vice versa. This behavior is analogous to that of Least Absolute Deviations versus Ordinary Least Squares -- while using the $\ell^1$ distance rather than the $\ell^2$ distance does in fact produce different solutions, these solutions may not be qualitatively different.
However, as the dimensionality and complexity of the covariate space increases, the differences become more pronounced. The CATE solutions exhibit systematically larger convex hull areas and greater dispersion, consistent with the goal of function estimation over the support of the space rather than centroid approximation.
In our causal inference context, the practical implication is that, for small sized problem sets, solution sets that are optimal for the PATE are likely also to be optimal for the CATE. The CATE objective, in principle, prioritizes coverage over the space, so that we can learn $\mathbb{E}[\tau | X = x]$ for a large support $X$. The PATE objective prioritizes coverage of the center, so that we learn the average location with high probability. In practice, however, good coverage of the space implies good coverage of the average, and a solution that minimizes absolute distance from selected sites to non-selected sites will also provide good coverage of the support of the covariates.
The robustness parameter $\rho$ controls the budget allocated to the adversary in the distributional robustness problem. As $\rho$ increases, the DRO framework hedges against increasingly severe distribution shifts by selecting more dispersed site configurations. This section demonstrates how robustness considerations systematically alter the geometry of optimal selections.
To illustrate this behavior, I solve the DRO problem across a range of $\rho$ values and track the evolution of site selection patterns. The simulation uses a two-dimensional covariate space with $30$ candidate sites, selecting $5$ sites at different robustness levels.
This robustness-coverage trade-off has implications for experimental design under uncertainty. Researchers facing potential distribution shift should choose $\rho$ values that balance the benefits of robustness against the costs of suboptimal site allocation. The Jaccard radius selection procedure, described in Section 3.5, provides an automated way to select this radius, with implications for the size of the hull selected.
Randomization is minimax optimal for experimental selection when the researcher has no prior information about experiments kallus2020optimal. We are essentially using prior information, in the form of covariates, to choose sites, and would expect that the quality of our site selection improves as covariates become more informative.
The key question is: at what threshold of covariate informativeness do optimization methods cease to provide benefits over simpler approaches? This threshold determines the practical applicability of the optimization procedures.
To evaluate this, I run a simulation in which the site selections are evaluated over a grid of $\eta$ values, where $\eta$ controls the degree of unmeasured confounding, as in the upper bounds derived above. There is a mild reparameterization, as $\eta$ is now defined on the support $[0,1]$, with the interpretation that $\eta = 0$ implies that covariates are sufficient, and there are no unobserved determinants of treatment effect, while $\eta = 1$ implies that covariates are completely uninformative about treatment effects, and the optimization methods are essentially fitting to noise.
The simulation generates treatment effects using the parameterization detailed in Appendix (ref), which allows systematic variation of signal strength while maintaining realistic correlation structures between covariates and outcomes.
The goal is to compare the optimization procedures to 1) complete randomization, in which sites are selected at random and 2) stratification, in which $k$-means is first used to separate the sites into strata, and sites are then sampled from the $k$ clusters. This is the procedure suggested in Tipton_2013.
These represent two different assumptions about our prior information. Complete randomization implies that we have no information about potential outcomes from covariates. Stratification implies that we have some information about covariates: we know that some covariates are important enough that we should condition our randomization on them. Stratification can be understood as a compromise between complete randomization and optimization approaches: it is a constrained randomization approach.
Results are displatyed in (ref) and (ref). Optimization methods outperform randomization when covariates are informative up to $\eta \approx .7$. We can translate $\eta \approx .7 \implies R^2 \approx .5$. The Cr\'epon study below has an $R^2$ of $.66$, which would mean we had good enough covariates to consider optimization-based selection methods.
This breakdown point has important practical implications. Researchers should validate covariate informativeness before relying heavily on optimization-based site selection. This suggests a straightforward moral: optimization methods outperform random assignment when covariates are sufficiently informative about potential outcomes.
I show this result formally in (ref), and it can be observed empirically in (ref). The intuition is that to select sites that provide optimal coverage of the support of the function, 2-Wasserstein transport simultaneously selects an optimal partition and optimal representatives of the space. This is in distinction to stratification, where optimal representatives are identified given a partition. Hence, 2-Wasserstein transport provides a weak lower bound on the error of the stratified sampling solution.
The connection to stratified sampling arises through the geometric structure of optimal transport solutions. The 2-Wasserstein optimal transport problem induces a Voronoi partition of the covariate space, where each selected site serves as “the local representative” for all sites in its Voronoi cell. Formally, for optimal sites $\{s_1^*, \ldots, s_K^*\}$, the induced partition is $\mathcal{V}_j = \{x \in \mathcal{X} : ||x - s_j^*|| \leq ||x - s_k^*|| \text{ for all } k \neq j\}$.
We can think of these partitions as optimal strata that are learned from the data and adapt to the dimensionality of the covariates.
These simulations demonstrate that optimization methods outperform random selection when observable covariates explain at least $50\%$ of treatment effect variation ($R^2 > 0.5$ or $\eta < 0.7$). This threshold is achieved in many policy-relevant settings for instance, the Cr pon et al. microcredit study analyzed below has $R^2 = 0.66$. Researchers should use optimization when they have strong priors or evidence about treatment effect predictors; otherwise, stratified random sampling is a sensible choice as a compromise between `informed' and `ex ante impartial' site selection methods.
This also highlights the important of collecting prognostic covariates prior to experimental deployment bicalho2022conditional.
Cr\'epon et al. (2015) studied the effects of a randomized microcredit intervention in Morocco. They considered a population of $162$ villages, which were randomized into $81$ matched pairs. Treatment consisted of an encouragement campaign to take out credit from Al Amana banks: “door-to-door campaigns, meetings with current and potential clients, contact with village associations, cooperatives, and women's centers, etc." (129).
These villages that were randomized into treatment were a population of sites that were on the periphery of catchment areas of existing branches: the goal was to assess whether taking up microcredit had an impact on a number of economic variables.
In this simulation, we take household self-employment activity profits as the outcome. We estimate the effect of treatment, site-level and individual covariates on profits, and estimate synthetic treatment effects for every individual in the sample using observed information. Sites are selected on the basis of aggregate-level site data, and we then estimate the error in terms of $MSE_{\text{PATE}}$ and $PEHE$ for each site selection. A more detailed description of the simulation procedure can be found in (ref).
For each of 500 simulation runs, we:
Distribution shift is induced by modifying the covariate distributions of candidate sites relative to the deployment population.
Site level covariates are shifted in the following way:
$$\mathbf{X}^{\text{shifted}}_s = \mathbf{X}_s + \varsigma \cdot \frac{d_{\text{med}}}{2} \cdot \frac{\mathbf{X}_s - \bar{\mathbf{X}}}{\|\mathbf{X}_s - \bar{\mathbf{X}}\|}$$
where: $$\varsigma \in \{0.0, 0.4, 0.6, 0.9, 1.7, 3.4\}, \quad d_{\text{med}} = \text{median}_{s,s'}\|\mathbf{X}_s - \mathbf{X}_{s'}\|$$
We calculate the actual variation in the data -- the observed median shift and perturbation of each site s, and use this as a benchmark level of variation. This then allows us to `stretch' sites away from their current location, as parameterized by our choice of the (true, underling degree of shift) $\varsigma$.
The simulation is run for two signal-to-noise ratio levels: $.3$, $.9$. These correspond to a low signal and high signal case respectively.
We implement five site selection methods:
Each method selects $K$ sites from a pool of $N$ candidate sites, with $(N,K) \in {(20,4), (25,5)}$. These are small site selection sizes, but are nonetheless sufficient to demonstrate the scale advantages of the optimal transport methods.
The simulation results demonstrate three main patterns. First, site selection method choice produces larger performance differences for PATE estimation than for CATE estimation. Second, the relative performance of methods depends on signal strength and problem size. Third, distributionally robust methods become preferred under realistic degrees of distribution shift.
For PATE estimation, performance advantages range from 3.4% to 71.9% . Under low signal strength (0.3), SPS dominates when distribution shift is minimal, but Wasserstein DRO becomes optimal when shift exceeds 1.7 times empirical variation. Under high signal strength (0.9), Optimal Transport methods generally outperform alternatives, except under large distribution shift where DRO maintains advantages.
For CATE estimation, performance differences between methods are substantially smaller, with most advantages below 1%. This pattern holds across signal strength and shift conditions, indicating that CATE performance depends more on fundamental signal-to-noise constraints than on site selection method choice. Our results show that site selection for the PATE is qualitatively different to site selection for the CATE. In (ref), I show that there are theoretical equivalences between optimal transport methods and familiar survey sampling approaches.
SPS methods have an advantage in the $\binom{20}{4}$ case under low signal strength, but are dominated by Optimal Transport methods for the larger problem size of $\binom{25}{5}$. This is practically important, as convex hull methods suffer from the curse of dimensionality as the sample size increases.
Optimal transport methods strictly dominate in the signal $= .9$ case. This was true for both the original and shifted problems, with performance advantages over SPS ranging from 10.1% to 43.6%.
The crossover point where DRO methods become preferred occurs at shift levels of $1.7$ times observed empirical variation. This is in part because DRO is specifically designed for the distribution shift context; the synthetic control method does not come with specific robustness guarantees against adversarial distribution shift.
For CATE estimation, both methods perform equivalently well, with Optimal Transport methods weakly dominant.
This is largely because of the Nature of the CATE estimation task, in which the goal is to smoothly interpolate a function over a large covariate space. In this setting, the optimal site selection is a regularly spaced grid over the support of the covariates.
Estimating the CATE is a fundamentally difficult problem, because it requires that we are able to well-estimate $\tau(x)$ at every `cell' $X = x$. In the low-signal regime, our estimates will be inherently noisy.
The limited difference between CATE and PATE methods may be an artifact of the simulation structure. dehejia2019localglobalexternalvalidity argue that macro-level variables are, in the case they study, more significant moderators of treatment effects. By aggregating up individual level treatment effects, it is likely that we are constructing macro level variables with little realistic variation between sites, instead of supposing that treatment effects vary significantly as a function of macro variables.
When within-site variance of treatment effects is large relative to between-variance, selecting sites based on aggregate-level data is not very informative. This will naturally be the case when selecting sites based on aggregated data: we lose the individual-level information that ultimately determines how precise our estimate of the PEHE is.
In essence, even though we are in a high-signal regime, our site selection covariates are not especially predictive of individual treatment effects. We essentially need to study the behavior of the CATE method when treatment effects contain large, site-moderated effects.
How can researchers assess whether their covariates are sufficiently informative (i.e., $\eta < 0.7$ or $R^2 > 0.5$) without running the full experiment? Several approaches are available: (1) Prior experiments: Use treatment effect estimates from similar interventions to assess how much variation is explained by observable characteristics. For instance, in educational interventions, prior multi-site trials can reveal whether school-level characteristics predict treatment effects. (2) Pre-treatment outcomes: When available, the relationship between covariates and pre-treatment outcomes is suggestive about the relationship between covariates and treatment, under the assumption that these covariates are also effect mediators (3) Domain expertise and theory: In many fields, accumulated knowledge suggests which characteristics drive heterogeneity: for example, baseline health status in medical trials or institutional capacity in policy interventions. (4) Pilot studies: Small-scale pilots across diverse sites can help researchers estimate treatment effect heterogeneity before committing to the full experiment.
To apply these methods in practice:
The Cr\'epon reanalysis demonstrates that distributionally robust optimization provides insurance against population misspecification at realistic uncertainty levels. In our simulation, DRO methods become preferred when deployment populations differ from candidate sites by margins exceeding 1.7 times observed empirical variation in candidate sites. This is a useful heuristic.
A practitioner objection to these methods might be that collecting data before engaging in an RCT is expensive or difficult, and that large-scale, policy-relevant RCTs are already difficult enough. I argue however that pre-emption is better than cure: given the expense and scale of many modern RCTs, improving pre-execution data collection may significantly increase the efficiency of the actual experimental estimate, making it much less likely that the experiment will fail due to random features of the selected experimental population, rather than the absence of a treatment effect.
Optimal transport methods are likely to scale better than convex hull methods to large experimental design problems. This is because, in high dimensions, the volume of a convex hull is concentrated at its surface: this is one version of the curse of dimensionality. Optimal transport methods estimate pairwise distances, are not computationally expensive to calculate, and can be solved by linear programming. The computational advantages become more pronounced as numbers of sites and numbers of covariates increase, making these approaches particularly suitable for large experimental site selection problems in the 100s.
We found that optimization methods perform well compared to randomization when covariates were moderately informative ($R^2 > .5$).
A fundamental challenge in site selection is that we typically observe only a subset $P \subset \mathcal{P}$ of the universe of potential experimental sites, and the gap between $P$ and $\mathcal{P}$ represents a form of Knightian uncertainty Knight1921, Sunstein2023. While our distributionally robust optimization methods provide insurance against distribution shifts within a Wasserstein ball of radius $\rho$ around the observed data, choosing $\rho$ itself requires confronting irreducible uncertainty about the nature of unobserved sites. This uncertainty differs qualitatively from the statistical risk we can quantify within $P$: we cannot assign probabilities to different ways $\mathcal{P}$ might differ from $P$ without making untestable assumptions. For instance, if infrastructure constraints systematically exclude remote rural sites from $P$, we face true uncertainty about how treatment effects might differ in these unobserved contexts. This limitation is not specific to our methods but reflects an inherent constraint in experimental site selection: optimization can only operate within the bounds of what we observe. In practice, expanding the set of feasible experimental sites $P$ -- the extensive margin -- also matters significantly for the quality of downstream inferences.
If we have information about individual covariates in a given site, it would be possible to incorporate this information into the site selection problem Neyman_1934,rosenman2022designingexperimentsshrinkageestimation. Intuitively, for the PATE, we would want to minimize the within-variance of selected sites: this simply increases the error of our downstream estimate. But for the CATE, within-variance of selected sites is heterogeneity to be exploited downstream. In both cases, we could incorporate prior information about the informativeness of sites into the objective function of the minimization problem Bertsimas2015.
We can adapt this method to select individuals to enroll in an experiment, not just sites. This is a topic of particular interest in experimental planning in industry settings, where user bases may be large, and understanding the behavior of specific market segments is of core interest arbour_2021. Sinkhorn regularization can be used to make optimal transport methods scalable to problem instances with $N$ in the 1000s cuturi2013sinkhorn.
A wide variety of core tasks in causal inference can be described as optimal transport problems: achieving balance between treatment and control distributions, matching, and synthetic control-type approaches dunipace2022optimaltransportweightscausal, bruns-smith2022. Distributionally Robust Optimization methods could also be practically useful for applied researchers in political science, where there is uncertainty about the quality of data collection, or, broadly, about differences between trial and deployment populations. Here, the connection with sensitivity analysis is germane: researchers can find treatment effect estimates with guarantees on their stability under worst-case distribution shift.