Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
116,294 characters · 19 sections · 90 citation commands
Learning What to Learn: Experimental Design when Combining Experimental with Observational Evidence
\onehalfspacing
Randomized controlled trials (RCTs) have improved empirical economics by providing internally valid estimates. However, in practice, feasibility constraints often confine trials to localized effects such as the effect in a specific site or subpopulations, or of a particular mechanism, as muralidharan2017experimentation document.\footnote{Experiments in top economics journals are in median “representative of a population of 10,885 units [...] and randomized in clusters of 26 units per cluster” muralidharan2017experimentation.}These effects, although useful, are often not sufficient to answer relevant policy questions about external validity and general equilibrium (GE) effects of at-scale interventions.
In response, a large literature in development meghir2022migration, ToddWolpin2006AER, attanasio2012education, de2025decoupling, education and labor economics allende2019approximating, larroucau2024college, chetty2016effects, and more recently industrial organization katz2025digital, allcott2025sources combines experiments with external evidence—observational and/or results from experiments in other settings—to extrapolate policy-relevant counterfactuals that no single experiment can identify. In our review of AEA journals over the past decade, about one-third of experimental papers complement experimental with observational inputs, either reduced-form or structural (Figure (ref)). Since budget constraints often bind, such common practice raises the question of how to choose the most informative experiment(s) and design for policy-relevant estimands from a (typical large) menu of possible experiments ravallion2012fighting, niederle2025experiments.
Our main contribution is a framework for choosing which experiment to run and how to design it in this common setting of combining experimental with external evidence. By external evidence we mean reduced-form or structural estimates based on observational data, and/or results from experiments in other contexts. External evidence may be biased in ways not known ex ante (e.g., due to confounding or lack of external validity). We provide a procedure that (1) helps researchers disentangle the role of each experiment in this setting, (2) selects which experiments to run subject to a budget (e.g., which treatment arms and/or sub-populations), (3) optimally allocates sample size, and (4) prescribes how to combine experimental and external evidence for estimation. We allow for arbitrary (user-specific) constraints in each step, for both the choice of the design and of the estimator.\footnote{For (4), researchers may constrain the estimator to either only use experimental evidence when available or optimally combine experimental with (biased) observational evidence for estimating certain parameters.}
For an illustrative example, consider a government piloting a cash-transfer program in a small set of districts to increase children's school attendance. The trial delivers a local average effect, but policy decisions often require counterfactuals at scale, allowing prices and wages to adjust ToddWolpin2006AER, egger2022general. Researchers therefore combine experimental evidence with a supply-demand model that leverages observational data to map local impacts into economy-wide outcomes. The design question is (i) which mechanisms are the most valuable to learn experimentally given constraints on the size of the experiment and (ii) how to allocate sample size across treatment arms and sites.
The first step is to specify the estimand and the underlying parameters. Define $\tau(\theta)$ the (policy-relevant) object of interest, where $\tau$ is a known smooth function and $\theta$ is a vector of unknown parameters; $\tau$ may be scalar or vector-valued. Researchers also observe external baseline estimates of $\theta$--henceforth observational estimates--which may be biased. Each feasible experiment can identify a subset of parameters without bias. This parameterization makes explicit which components of $\theta$ are learned experimentally and which may rely on (biased) observational evidence; not all components need or can be learned experimentally.
Returning to the cash-transfer example, $\tau(\theta)$ is the effect on schooling when all poor households in rural Kenya receive the subsidy. The entries of $\theta$ consist of (i) the direct schooling response to a conditional cash transfer (CCT), (ii) the income effect of transfers in partial equilibrium, and (iii) wage adjustments that shift the returns to schooling. Due to cost constraints, researchers may only run small, partial-equilibrium experiments and then combine these with prior evidence to recover $\tau(\theta)$. They may choose between a CCT arm to estimate the direct effect, or an unconditional-cash arm (UCT) to measure the income elasticity, and extrapolate the GE effects using evidence from a previous study in Mexico ToddWolpin2006AER. (When site selection is also part of the design, $\theta$ additionally collects site-specific effects and $\tau(\theta)$ targets the average effect across sites.)
The second step is to select the experiment(s), sample size and estimator. A natural starting point would be to design the experiment to minimize the mean-squared error (MSE) for $\tau(\theta)$. However, in practice, the bias is unknown prior to the experiment. A worst-case approach would be overly-conservative and disregard information about the variance. We instead adapt the definition of adaptation regret, previously studied for combining two estimators under a fixed design armstrong2024adaptingmisspecification, tsybakov1998pointwise, to our experimental design problem (with multiple parameters). Here, the regret is the ratio of the researcher's worst-case mean-squared error relative to an oracle that knows the worst-case bound on the observational bias and chooses both the design and the estimator. Taking its supremum over the worst-case bias yields a robust procedure without requiring a prior about the bias or on its bound. However, our task of also choosing the experimental design and the lack of an unbiased estimate for $\tau(\theta)$ when not all parameters can be learned experimentally make existing solutions not applicable.
Our main result provides an explicit and novel characterization of adaptation regret for (asymptotically) linear estimators in experimental and observational moments. These include generalized methods of moments where we may also optimize over the weighting matrix. For any design and weighting matrix that combine the experimental and observational point estimates, the regret is the maximum of two normalized components. The first is a variance regret: the estimator’s sampling variance under the chosen design and weights, divided by the smallest attainable variance. The second is a bias regret: the worst-case squared bias induced by using external evidence, divided by the smallest feasible worst-case bias over the class of designs. The variance regret depends on the (expected) variance-covariance matrix of the observational and experimental estimates implied by the design, the weighting matrix, and the sensitivity of the policy parameter, i.e., the gradient of $\tau$ with respect to the parameters evaluated, under mild conditions, at the observational estimates.\footnote{For this result to hold, we assume that $\tau(\theta)$ is twice differentiable with bounded derivatives. For non-linear $\tau(\theta)$ in $\theta$, we also require that the observational estimates bias is local to zero in the spirit of andrews2020transparency, armstrong2021sensitivity, BonhommeWeidner2021arXiv, where its rate of convergence can be the same, faster, or even slower than the estimators' standard error. Section (ref) provides details.} The bias regret depends only on the sensitivity vector and the weighting matrix, known ex-ante.
Our characterization makes the bias-variance trade-off intuitive. At the optimum, the two normalized components tend to be equalized: when the bias dominates, the design prioritizes high-sensitivity parameters; when variance dominates, the procedure invests sample where it most reduces variance. This yields a nested procedure that jointly selects the estimator, the sample allocation and the experiment(s) to run, trading off precision with bias sensitivity. The resulting design and estimator only require knowledge of the expected variance-covariance matrix and observational estimates. It can be reported in a pre-analysis plan. That is, for given observational estimates, the inputs are the same as those required by standard power-analysis plans for $\tau(\theta)$ GerberGreen2012.
We then generalize our framework to general moment selection problems that combine experimentally identified moments with observational (misspecified) moments. With parameters defined by non-linear but smooth moments our results directly apply under standard local asymptotics, where the bias is local to zero at same or possibly different rate than the standard error andrews2017, andrews2020transparency. We also generalize to partially identified $\tau$ and settings where researchers have priors (confidence sets) about the bias.
From a theoretical (and computational) perspective, our key insight is that, as we restrict the class of estimators to be (asymptotically) linear in the moments, the worst-case adaptation regret is quasi-convex in the radius of the bias, for any norm characterizing the bias' uncertainty set. This result yields a transparent variance-bias trade-off and differs from existing solutions under the fixed-design optimal estimators as in armstrong2024adaptingmisspecification, tsybakov1998pointwise; there, quasi-convexity fails due to non-linearity of the estimator, and optimization corresponds to a minimax program over worst-case priors. In our joint design and estimation problem, quasi-convexity substantially simplifies optimization which is quadratic for the choice of the estimator, with a closed form expression for sample allocation and a mixed-integer formulation for the experiment choice only (typically from a finite set). This program is solved via off-the-shelf mixed-integer quadratic programming.
We illustrate the framework in a cash‐transfer application in Kenya, where the researcher combines new experimental evidence with external evidence from Mexico’s PROGRESA program ToddWolpin2006AER, attanasio2012education, where these preliminary estimates may lack external validity. We compare three designs and their combinations: (i) a CCT arm to estimate direct schooling effects; (ii) a UCT arm to estimate income effects; and (iii) an employment program that identifies individual wage elasticities. Experimentation costs are binding, as cash transfers and employment programs require distinct infrastructures and implementation teams. We show that the wage experiment is the most informative for sufficiently large sample sizes, while for smaller samples a combination of CCT and UCT—allocating most of the sample to the CCT arm—is optimal. With one thousand participants, relative to the variance-minimizing (Neyman) design, our optimal two-arm design delivers more than a 400% reduction in bias (and adaptation regret) with variance no larger than $\sim$30% of the Neyman variance; for a single-arm design, our procedure yields a 140% reduction in bias at the cost of only about a 12% increase in variance. These comparisons illustrate the benefit of our method in balancing robustness and precision.
As a second application, we study where to run a microfinance experiment in India, integrating evidence from earlier nonrandomized microfinance introductions. We target the average effect in the entire region, but constraint the researcher to choose one or two smaller areas for randomization due to coordination (fixed) costs constraints. We use observational estimates in banerjee2024changes to select one or more areas for randomization. Because banerjee2024changes also implement a separate randomized expansion, we can calibrate performance of our design (agnostic of the bias) by comparing each candidate design’s MSE to that of an oracle that knows the bias, using the experimental estimates to calibrate the bias. Compared to the variance-optimal benchmark, we reduce MSE by more than 250%.
\paragraph{Related literature} This paper links the experimental-design literature—which mostly focuses on settings where all parameters are identified within the experiment and leaves aside questions of data combination—to recent work that integrates experimental and (reduced-form or structural) observational evidence to extrapolate effects in complex scenarios.
Recent advances for experimental design include balancing and variance-minimizing allocations tabord2018stratification, bai2019optimality, cytrynbaum2021optimal, kallus2018optimal, bertsimas2015power, adaptive designs for policy choice kasy2019adaptive, russo2018tutorial, cesa2025adaptive, and experimental design under correct model specification silvey2013optimal, chaudhuri1993nonlinear, chaloner1995bayesian, higbee2024experimental, kiefer1959optimum, reeves2024model, viviano2020experimental, kasy2016experimenters. Traditional work on experimental design for robust-model estimation either focused on testing competing models atkinson1975design, lopez2007optimal, or on using a-priori knowledge of (worst-case) bias for the design of an experiment box1959basis, wiens1998minimax, tsirpitzi2023robust, sacks1984some. All these references leave aside questions about observational data combination. In the context of minimizing the variance of an experiment, rosenman2021designing proposes to use observational studies to construct high-confidence bounds on the experimental variance, and then focus on variance-optimal stratification. This problem differs from our design problem and analysis, which instead studies the use of observational studies in combination with experiments for extrapolation.
In summary, we complement this literature by studying which experiment to run (and how) when not all parameters are learned experimentally, a common problem in economics.
Our minimax regret criterion connects to a long-standing decision-theoretic tradition for experimental design. References include manski2016sufficient, banerjee2020theory, manski2004, dominitz2017more, olea2024externally, hu2024minimax, breza2025generalizability among others. These references focus on settings where researchers only have access to experimental variation, instead of combining it with external (observational) evidence, motivating different designs and objective functions. We also complement literature that leverages correctly-specified models for site-selection gechter2024selecting, abadie2021synthetic by allowing for misspecification in observational estimates.
A related line of work analyzes estimation under misspecification armstrong2024adaptingmisspecification, armstrong2018optimal, andrews2025structural, andrews2020transparency, BonhommeWeidner2021arXiv, donoho1994statistical, kwon2025estimating, and combining existing experiments with observational studies gechter2022combining, athey2025experimental, athey2020combining, kallus2018removing, dutz2021selection, rosenman2022propensity, de2020empirical, rambachan2024program. We provide a novel characterization of adaptation regret as a quasi-convex function in the bias bound (and therefore as the maximum of the variance and bias regret) by leveraging linearity of the estimator. Such characterization shows how, in our problem, misspecification is captured through the first-order sensitivity of the estimand to each parameter, an important measure for sensitivity analysis in andrews2020transparency; it provides a decision-theoretic foundation for simultaneously protecting against a correctly specified and misspecified model bickel1984parametric, kempthorne1988controlling, here through the maximum of a normalized variance and bias components. Unlike the references above, where the design is fixed and researchers optimize over the choice of the estimator, our main contribution is to optimize the design itself. This changes the optimization and regret, computed against the best design-estimator.
We will focus first on a simple shrinkage estimator and a univariate estimand. Our results generalize directly to general methods of moments and multivariate targets in Section (ref).
Consider a researcher interested in an arbitrary target estimand $\tau(\theta)\in\mathbb{R}$, indexed by a low-dimensional unknown parameter vector $\theta\in\mathbb{R}^p$ and a known mapping $\tau$. The goal is to construct an estimator $\hat{\tau}$ that accurately approximates $\tau(\theta)$ by combining existing evidence (e.g., observational studies) with experimental variation designed by the researcher. Our question is how to design such experiments under feasibility constraints.
\paragraph{Basic setup} We assume that researchers have access to estimators (and their variance) of $\theta$ denoted as $\tilde{\theta}^{\mathrm{obs}}\in\mathbb{R}^p$. We impose no restrictions on how $\tilde{\theta}^{\mathrm{obs}}$ is formed: it can based on arbitrary exclusion restrictions implied by an economic or statistical model. However, $\tilde{\theta}^{\mathrm{obs}}$, defined as observational estimates, have biases collected in a vector $b\in\mathbb{R}^p$ unknown at the experimental design stage (see Remark (ref)). Examples include an estimate from an instrumental variable regression that may fail the exclusion restriction, or an experimental estimate from a country different from the one of interest that may lack external validity. Researchers can then collect experimental evidence for a subset of parameters, producing estimates which are unbiased by design.
Here, $(\mathcal{E},\Sigma(\mathcal{E}))$ characterizes the experimental design, encoding which parameters are learned and with what precision. When clear, we omit the subscript $\Sigma$ on $\tilde{\theta}^{\mathrm{exp}}$. We assume $\Sigma(\mathcal{E})$ is known, which is standard in experimental planning GerberGreen2012; when $\Sigma$ is unknown, we can interpret $\Sigma$ as the expected variance-covariance matrix.\footnote{In this case, we interpret the MSE in Equation (ref) as the expected MSE under a common expectation over the experimental variance-covariance matrix. It is also possible to extend our framework by allowing $\Sigma$ to be a function of $\tilde{\theta}^{\mathrm{obs}}$ provided that $\Sigma(\tilde{\theta}^{\mathrm{obs}}) \rightarrow_p \Sigma(\mathrm{plim}(\tilde{\theta}^{\mathrm{obs}}))$, with $\Sigma(\mathrm{plim}(\tilde{\theta}^{\mathrm{obs}}))$ denoting the common expectation of the variance-covariance experimental matrix. In this case our analysis should be interpreted in an asymptotic regime similar to what discussed in Section (ref). }
The class of experiments and estimators is assumed to be in a user-specific set $\mathcal{D}$.
We let $(\mathcal{E}, \Sigma(\mathcal{E}))$ be within arbitrary constraint sets that may encode feasibility or budget constraints. In our framework, restrictions on the experiments researchers can run $\mathcal{E}$ and on their precision $\Sigma(\mathcal{E})$ can take any desired form. In most applications, we think of such constraints as arising from fixed costs of running an experiment and/or constraints on their power duflo2007using, athey2017econometrics, list2011so. We assume that $\Sigma(\mathcal{E})$ is uniformly bounded, which implies that once we commit to learn a set of experimental parameters $\tilde{\theta}_{\mathcal{E}}^{\mathrm{exp}}$, their variance is finite, and strictly positive definite, therefore assuming that the variance is bounded away from zero.\footnote{For this latter condition to hold in an asymptotic framework with growing sample size, $\theta$ and $\tilde{\theta}$ can be defined without loss as parameters of interest after appropriately rescaling by the square-root of the sample size; in this case $\Sigma$ denotes the asymptotic variance. See Section (ref) for details.} The shrinkage estimator can be unrestricted (i.e., $\gamma$ can take any value on the real line) or restricted (when for example researchers aim to only use experimental evidence when available, setting $\gamma_j = 1$ for $j \in \mathcal{E}$).
\paragraph{Main design problem} Our design problem can be described in three steps:
Choosing $\gamma$ (step 1) for a fixed design follows a long-standing tradition in econometrics and statistics armstrong2024adaptingmisspecification,andrews2021model,donoho1994statistical,athey2025experimental. We study this question here in combination with the experimental design problem, step 2 and 3, which, as we show will change the decision problem. The challenge is that performance depends on the bias $b$, unknown when designing the experiment.
\paragraph{Properties of $\tau$} A key feature of our framework is that researchers explicitly parameterize $\tau(\cdot)$, thereby pre-specifying and making transparent which biases drive the design problem; this feature clarifies the role of the experiment in complementing existing evidence. The sensitivity of $\tau$ to $\theta$, captured by its gradient $\omega(\theta)=\partial\tau(\theta)/\partial\theta$, determines how (to first order) the bias propagates onto the final estimand. For simplicity, we impose exact linearity, $\tau(\theta)=\omega^\top\theta$; all results extend to smooth, nonlinear $\tau$ via a first-order expansion. When $\tau$ is unknown, Section (ref) studies settings where researchers only partially identify $\tau$.
Assumption (ref) holds exactly when $\tau(\cdot)$ is linear and serves as a first-order approximation for smooth non-linear $\tau(\cdot)$. Section (ref) (and Example (ref)) extend the framework to non-linear estimands encompassing as special cases the local-misspecification setups in andrews2020transparency,armstrong2021sensitivity. For non-linear $\tau$, we replace $\omega$ with the gradient of $\tau$ evaluated at a preliminary observational estimate, i.e., $ \omega \;=\; \frac{\partial \tau(\theta)}{\partial \theta}\Big|_{\theta = \tilde{\theta}^{\mathrm{obs}}} \;+\; o_p(1), $ which is typically available before the experiment is conducted. That is, building on references above, $\omega$ captures the first-order effect of the observational bias $b$ on the bias of $\hat{\tau}$.\footnote{Intuitively, second order effects are negligible under local asymptotics where, for given observational sample size $n$, $||b||^2 \sqrt{n} \rightarrow 0$ (or equivalently $||b||^2/\sqrt{n} \rightarrow 0$ when we define $\theta$ the parameter appropriately multiplied by its rate of convergence $\sqrt{n}$). This implies that Assumption (ref) holds asymptotically when $||b||^2$ is proportional to, smaller or even much larger than the variance of $\tau(\hat{\theta})$. Section (ref) presents a formalization. }
We conclude with three examples in simple two-parameter models.
Next, we introduce an experimental design focusing on the mean-squared error (MSE) of the estimator $\hat{\tau}$. We consider confidence interval length in Section (ref). The MSE is defined, for given bias level $b$, design $(\mathcal{E},\Sigma(\mathcal{E}))$, and choice of weights $\gamma$ as
where $\mathbb{E}_{\mathcal{E},\Sigma,b}$ denotes expectation under the design $(\mathcal{E},\Sigma(\mathcal{E}))$ and observational bias $b$.
Ideally, one would minimize $\mathrm{MSE}_b$. However, because $b$ is unknown, we consider an uncertainty set given by an $\ell_\infty$-ball,
with $B\ge 0$ an upper bound on the largest coordinate-wise bias. In our framework, a key challenge is that biases may arise from multiple parameters. Here, for exposition we first focus on the $\ell_\infty$-ball given its interpretability, while Section (ref) generalizes the framework to arbitrary norms. The $\ell_\infty$ norm imposes symmetric restrictions on the biases and it does not force a trade-off across coordinates (a large bias in one component need not be “offset” by a small bias elsewhere), which is attractive when biases may be positively correlated across parameters. The drawback is that worst-case solutions depend on the radius $B$ which may be unknown in practice.
If an oracle knew $B$, a natural choice would be to pick the minimax design \[ \mathrm{MSE}^*(B)\;\equiv\; \inf_{(\mathcal{E},\Sigma,\gamma)\in\mathcal{D}} \ \sup_{b\in\mathcal{B}(B)}\ \mathrm{MSE}_b(\mathcal{E},\Sigma,\gamma). \] However, having to specify $B$ can pose a large burden on the researchers and make the choice of the experiment sensitive to $B$. Therefore, we seek designs that perform as close as possible to the oracle that knows $B$, uniformly over the values of $B$, following a long-standing tradition in decision theory manski2004, manski2007admissible, montiel2023decision, KitagawaTetenov_EMCA2018. Specifically, we minimize the worst-case proportional regret
defined adaptation regret by armstrong2024adaptingmisspecification and tsybakov1998pointwise for estimation under fixed designs (see Remark (ref) for a comparison with our experimental design problem).
Our main results show that the adaptation regret provides natural trade-offs between the variance and the bias. To do so, denote $\mathbb{V}_{\mathcal{E},\Sigma}(\cdot)$ the variance operator for given $(\mathcal{E},\Sigma)$.
The variance regret denotes the ratio between the variance of a given design and estimator relative to the smallest achievable variance. We provide a similar definition for the bias.
To gain insight, note that $B^2 \beta = B^2 \Big(\|\omega\|_1 - \|\omega_{\mathcal{E}}\|_1\Big)^2$ denotes the worst-case bias for a given design, value $B^2$ and shrinkage estimator. $B^2 \beta^\star$ denotes its minimizer.
An ideal design would set the variance equal to $\alpha^\star$ and bias equal to $\beta^\star$. Unfortunately, this may be infeasible as it might require different choices of experiments and estimators to achieve one or the other. This raises the question of how to trade-off these two components.
This trade-off is formally characterized in our main theorem below. We will use the convention throughout that 0/0 = 1.
Theorem (ref) offers a simple and novel expression of the regret which does not require researchers to specify $B$. As we show below, Theorem (ref) will significantly simplify the optimization program. The key insight is that the adaptation regret for linear estimators and arbitrary designs leads to a (strictly) quasi-convex objective function, whose worst-case solution is the maximum between the variance and the bias regret. This provides a simple and intuitive trade-off between the two. Specifically, provided that the class of designs $\mathcal{D}$ is sufficiently flexible (outside boundary solutions), it follows that at the optimum we equalize $$ \frac{\alpha}{\alpha^\star} \;=\; \frac{\beta}{\beta^\star}. $$ That is, we allocate sample size to noisier treatment arms when the variance dominates, and select treatment arms with largest sensitivity when the bias dominates.
Next, we provide an explicit optimization routine that can be solved using off-the-shelf software. We consider observational estimates independent of the experimental ones (but possibly dependent with each other), and experimental estimates independent of each other. This occurs when conducting experiments on independent samples different from the observational sample as in our applications. More general dependences are possible but omitted.
\paragraph{Notation and decision variables} For given $(\mathcal{E},\Sigma,\gamma)$, denote by $\Sigma_{\mathrm{obs}}$ the submatrix of $\Sigma$ corresponding to the variance–covariance matrix of the observational estimates (possibly non-diagonal), and let $\mathbb{V}(\tilde{\theta}_j^{\mathrm{exp}}) = v_j^2/n_j$ be the variance of the experimental estimator for component $j$, where $n_j$ denotes the sample size allocated to that experiment. Let $x_j = 1\{j \in \mathcal{E}\} \in \{0,1\}$ indicate whether the experiment identifies component $j$, and let $x \in \mathcal{X}$, where $\mathcal{X}$ denotes the constraint set of feasible experiments. We optimize over $(x,\gamma,n_j)$ under a total budget constraint $ \sum_{j\in\mathcal{E}} c_j n_j = n, $ where $c_j > 0$ denotes a per-unit cost for experiment.
\paragraph{Variance, bias and optimal sample size} Under independence of experimental and observational estimators, the variance of $\tau(\theta)$ can be written as
where $\odot$ denotes the elementwise product. The sample sizes $n_j$ enter only through the first (experimental) component. By the standard Neyman allocation with variable costs GerberGreen2012, the minimizer over $(n_j)$ for fixed $s$ is given by
Substituting (ref) into the experimental variance and with a slight abuse of notation, let
the variance and bias contributions.
\paragraph{Oracle solutions} The next step is to compute $\alpha^\star$ and $\beta^\star$. These solve \[
\] The constraints for $s$ are standard inequalities enforcing $s_j = \gamma_j x_j$ when $x_j\in\{0,1\}$ and $0 \le \gamma_j \le 1$. Because $\alpha(s)$ is convex quadratic and $\mathcal{X}$ typically admits a linear (or quadratic) description, the first problem is a mixed-integer convex quadratic program, while the second is a mixed-integer linear program. Both can be solved using off-the-shelf solvers.
\paragraph{Regret-optimal solution} We are now ready to solve for the regret-optimal design. By Theorem (ref), we want to minimize $\max\{\alpha/\alpha^\star, \beta/\beta^\star\}$. The corresponding optimization program can be written as
The constraints (ref) again enforce $s_j = x_j \gamma_j$ for $x_j\in\{0,1\}$ and $0 \le \gamma_j \le 1$. The variable $t$ satisfies $t \ge \alpha/\alpha^\star$ and $t \ge \beta/\beta^\star$, so at the optimum we have $t = \max\{\alpha/\alpha^\star, \beta/\beta^\star\}$ as desired. Because both $\alpha(s)$ and $\beta(s)$ are convex quadratic functions, the optimization problem (ref)--(ref) is a mixed-integer quadratic program (MIQP) with linear objective and convex quadratic constraints, and can be passed to standard MIQP solvers.
In many scenarios, we may expect that experiments only identify relevant moments in the data, and observational estimates are often based on model-implied moment conditions. We now extend the framework to generalized method of moments (GMM) estimators. In addition, we allow for a more general class of ambiguity sets and multivalued estimands.
Consider a vector of moment conditions $g(\theta)\in\mathbb{R}^p$ that stacks the “menu of moments” (observational and experimental), i.e. all moments potentially available for estimation. Let $\Lambda$ denote the Jacobian of the moment function with respect to $\theta$, evaluated at $\theta$ (i.e., $\Lambda:=\partial g(\theta)/\partial\theta^\top$). For simplicity, we will assume that $\Lambda$ is known (i.e., moments are linear in $\theta$). This simplifies exposition but it is not necessary in the local asymptotic framework we present in Section (ref), where $\Lambda$ can be replaced by the Jacobian of $g(\theta)$ evaluated at the observational estimates. Let $\bar g_{\Sigma}$ denote the sample analogue with \[ \mathbb{E}[\bar g_{\Sigma}]=b,\qquad \mathbb{V}(\bar g_{\Sigma})=\Sigma, \] where the vector $b$ captures the biases in the moments. The researcher may access only some of the entries of $\bar{g}$ by selecting a subset of experiments due to budget constraints.
\paragraph{General misspecification set.} Let $\mathcal I\subseteq\{1,\dots,p\}$ index all potential experimental moments researchers may select (unbiased by design) and $\mathcal{I}^c$ the observational ones (potentially misspecified). We posit bounded misspecification under arbitrary norms
where the ambiguity set is defined relative to an arbitrary norm $||\cdot||_l$. Examples include the $\ell_1, \ell_2, \ell_\infty$ norms as special cases, as well as norms that may weight observational biases differently, or may even assume no bias for some observational moments (see Example (ref)).
\paragraph{Estimator and design.} Given the empirical moments $\bar g_{\Sigma}$, we consider (asymptotically) linear estimators of the form
where $W$ and $\Sigma$ are chosen by the researcher from a feasible set. The structure of the estimator follows a standard GMM asymptotic expansion andrews2020transparency, andrews2017.\footnote{The weighting matrix $W$ may be singular. The expansion $\hat\theta-\theta \;\dot{=}\; -(\Lambda^\top W \Lambda)^{-1}\Lambda^\top W\,\bar g_{\Sigma}$ holds whenever $\Lambda^\top W \Lambda$ is nonsingular, which can be forced through the set of constraints in Assumption (ref).}
Here, $W$ may have columns of zeros, which excludes the corresponding moments from the estimator. This occurs when researchers face a trade-off in which experiment to run, leading to different moments $\bar{g}$ observed in the experiment.
The matrix $\Sigma$ depends on the experimental design (e.g., sample-size allocations across experimental moments). The constraints on $W$ and $\Sigma$ effectively limit the class of experiments, sample size allocations and estimators the researcher may consider.
Assumption (ref) requires that the variance of the selected (reweighted) moments is uniformly bounded and strictly positive definite (even when $W$ does not select some moments). This ensures that any subset of moments the researcher chooses to use is nondegenerate. Note that for now we are abstracting from settings with incomplete models studied in Section (ref).
\paragraph{Target and MSE.} Let the estimand be potentially multivalued $\tau(\theta) \in \mathbb{R}^q$, where $\tau(\theta)= \Omega\theta$ with $\Omega\in\mathbb{R}^{q \times d}$ with $\Omega$ different from the zero matrix. Our objective is the sum of MSE for each entry of $\tau$, that is equal to
The key distinction now is that the variance and bias regret depend on $\Omega$ and $\Lambda$.
The bias now depends on the interaction between the sensitivity $\Omega$, the moment Jacobian $\Lambda$ and the identity of which moments are observational and therefore potentially biased. For unidimensional $\tau(\theta)$ the bias corresponds to the squared dual norm of $\Omega \Gamma_\Lambda(W)$ with respect to $||\cdot||_l$. The weighting matrix $W$ determines how these moments influence $\hat{\theta}$.
As before, we consider the adaptation regret with respect to an oracle that knows $B$:
The next theorem generalizes Theorem (ref) to this more general framework.
Theorem (ref) shows how our framework generalizes to (i) choosing the moments and the estimator (via $W$); and (ii) arbitrary norms characterizing the ambiguity set. The key insight is that the regret is quasi-convex under (approximate) linearity in the moments relative to the dual norm of the ambiguity set.
This generalization encompasses, in addition to choosing which experiment to run, problems that involve whether to collect additional data (e.g., through surveys or data acquisition) since these can be directly incorporated as part of the moment conditions.
Next, we generalize our framework to minimizing confidence interval length and partially identified $\tau(\theta)$. For simplicity, we let $\tau(\theta) \in \mathbb{R}$. Researchers may not know $\tau(\cdot)$ exactly and only know functions $\bar{\tau}, \underline{\tau}$ with \[ \bar{\tau}(\theta) = \bar{\omega}^\top \theta, \qquad \underline{\tau}(\theta) = \underline{\omega}^\top \theta, \quad \tau(\theta) \in [\underline{\tau}(\theta), \bar{\tau}(\theta)] \] where linearity can be justified as a first-order linear approximation for $\bar{\tau}(\theta)$ and $\underline{\tau}(\theta)$ under the local asymptotics in Section (ref) for smooth functions $\bar{\tau},\underline{\tau}$.\footnote{When the envelope for $\tau(\theta)$ is non-smooth, we recommend replacing it with smooth upper and lower envelopes. In this case, following Section (ref), $\bar{\omega}$ and $\underline{\omega}$ can be interpreted as the gradients of these smooth bounds evaluated at the observational estimates.} We implicitly assume that at least one of $\bar{\omega}$ and $\underline{\omega}$ differ from the zero vector to avoid trivial solutions.
For a generic weight vector $\omega$ let $ \alpha_{\omega}(W,\Sigma), \beta_{l,\omega}(W) $ as in Definition (ref).
\paragraph{Confidence intervals under bias.} For a candidate design $(W,\Sigma)$ and a bias vector $b\in\mathbb{R}^p$, define two-sided $(1-\eta)$ lower and upper bounds for $\tau(\theta)$ by \[
\] where $z_{1-\eta/2}$ is the standard normal $(1-\eta/2)$ quantile and $\hat\theta(W)$ is the GMM estimator in (ref). These bounds adjust for the bias $\omega^\top\Gamma_\Lambda(W)b$ in $\omega^\top\hat\theta(W)$ for each envelope separately.
As before, the Jacobian $\Lambda$ is known; when estimated see Section (ref) and Remark (ref).
An audience, indexed by a worst-case bias radius $B\ge 0$, forms a worst-case confidence interval (i.e., a bias-aware confidence interval a la armstrong2018optimal) by choosing the most conservative endpoints over $b$ in the ambiguity set $\mathcal{B}_l(B)$ in (ref): \[ L_{l,B}(W,\Sigma) ~\equiv~ \left[ \inf_{b\in\mathcal{B}_l(B)}\ell_b(W,\Sigma),\; \sup_{b\in\mathcal{B}_l(B)}u_b(W,\Sigma) \right], \] and we denote by $|L_{l,B}(W,\Sigma)|$ its total length.
The researcher does not need to know the audience's specific choice of $B$ and can simply report the point estimates and associated standard errors; the audience (indexed by $B$) can then form their own bias-aware confidence set from these statistics. The researcher would like to minimize the expected length of the confidence interval formed by the audience.
An oracle that knows $B$ minimizes the worst-case expected length of the audience's confidence length relative to the size of the sharp identified set (worst-case under the audience's prior $b \in \mathcal{B}_l(B)$). Its (ex-ante) expected loss is defined as
Here $\mathcal{L}_{l,B}(W,\Sigma)$ measures how informative the expected confidence interval is, relative to the length of the identified set $\bar{\tau}(\theta) - \underline{\tau}(\theta)$, worst-case over $b_0$ in $\mathcal{B}_l(B)$. Here, $b_0$ determines the expectation of $\hat{\theta}(W)$, and ultimately the length of the bias-aware identified set.
As in previous sections, we define the corresponding regret as \[ \tilde{\mathcal{R}}_l(W,\Sigma) ~\equiv~ \sup_{B\ge 0}\; \frac{\mathcal{L}_{l,B}(W,\Sigma)}{ \inf_{(W',\Sigma')\in \mathcal{D}'} \mathcal{L}_{l,B}(W',\Sigma') }\!. \]
\paragraph{Regret optimal design (and estimator).} Define \[ A(W,\Sigma) \equiv z_{1-\eta/2}\big(\sqrt{\alpha_{\bar{\omega}}(W,\Sigma)} + \sqrt{\alpha_{\underline{\omega}}(W,\Sigma)}\big), \quad C_l(W) \equiv \sqrt{\beta_{l,\bar{\omega}}(W)} + \sqrt{\beta_{l,\underline{\omega}}(W)} + \sqrt{\beta_{l,\bar{\omega} - \underline{\omega}}(W)}, \] and the corresponding oracle solutions \[ A^\star \equiv \inf_{(W,\Sigma)\in\mathcal{D}'} A(W,\Sigma), \qquad C_l^\star \equiv \inf_{(W,\Sigma)\in\mathcal{D}'} C_l(W). \] The first component $A$ depends on the variance of the estimated parameters (and envelopes) and the second component $C_l$ depends on the bias. We now show how our results directly extend to this scenario.
Theorem (ref) shows that the confidence-interval-length regret can be decomposed into a variance component $A(W,\Sigma)$ and a bias component $C_l(W)$, each divided by its smallest feasible value. This result mimics Theorem (ref) allowing for partially identified models. It provides a ready-to-use expression for experimental design (and moment selection).
The following corollary shows that whenever the model is point identified the solutions that minimize the regret relative to confidence interval length coincides with the MSE optimal solution we derived in Section (ref).
In summary, when $\tau(\theta)$ is point identified, minimizing the regret of the confidence interval length leads to the same optimal design as minimizing MSE-based regret. In partially identified models, Theorem (ref) shows that the confidence interval length maintains the same bias-variance decomposition, a desirable property of the MSE solution in Section (ref).
In this section we extend our results to settings where researchers have prior information about the bias upper bound. Following the framework in Section (ref), suppose researchers know that $B \le \bar{B}$ for a known upper bound $\bar{B}$, and consider an arbitrary (possibly weighted) norm $\|\cdot\|_l$ on the bias. The theorem below establishes the resulting characterization of the regret.
Theorem (ref) shows that when an upper bound $\bar{B}$ on the bias is known, the worst-case proportional regret over all $B\in[0,\bar{B}]$ is again the maximum of two terms. The first term, $\alpha_\Omega(W,\Sigma)/\alpha_\Omega^\star$, is the familiar variance regret. The second term compares the MSE evaluated at the largest feasible bias level, $\bar{B}$, to the smallest achievable MSE at the same bias level across all feasible designs and weight matrices. The optimization problem is therefore analogous to our baseline case, but with the “bias” component of regret evaluated at $B=\bar{B}$ rather than in the limit as $B\to\infty$.
From an interpretation standpoint, this extension interpolates between two polar cases. When $\bar{B}$ is small, the second term is close to the variance ratio, and designs that minimize variance dominate. As $\bar{B}$ increases, the second term places more weight on robustness to misspecification, and the objective approaches the fully robust case studied in our baseline results. In practice, the theorem implies that any prior information that narrows $B$ to $[0,\bar{B}]$ can be incorporated simply by replacing the worst-case bias component in our original criterion with the MSE at $B=\bar{B}$, while keeping the same quasi-convex structure.
Next, we extend our framework for (i) nonlinear moment functions $g(\theta)$ used in Section (ref) and (ii) a smooth, potentially nonlinear estimand $\tau(\theta)$. Although we focus on known $\tau(\cdot)$, similar reasoning applies to partially identified sets in Section (ref), omitted for brevity.
\paragraph{Linearization of the moments} Assume the sample moments obey root-$n$ scaling, i.e., $\mathbb{V}(\bar g_{\Sigma})=O(n^{-1})$. Let $b=b_n$ be a sequence where we retain the same misspecification structure as in Section (ref): the moment vector has mean $b=b_n$ with $(b_n)_j=0$ for $j\in\mathcal I$ (experiment-based moments) and $(b_n)_{\mathcal{I}^c}$ potentially nonzero. Our key assumption is a local condition on the moments bias of the form
This condition builds on the local asymptotic frameworks in andrews2020transparency and armstrong2021sensitivity, where $||b_n||^2 = 1/n$ matching the rate of convergence of the variance. In our case, $b_n$ can grow faster, slower or at the same rate of the standard error (i.e., $b_n \propto n^{-\alpha}, \alpha > 1/4$), encompassing $||b_n||^2 = 1/n$ as a special case.
Under standard smoothness conditions, a GMM estimator with weighting matrix $W$ satisfies \[ \sqrt n\big(\hat\theta(W,\Sigma)-\theta\big) = -(\Lambda^\top W \Lambda)^{-1}\Lambda^\top W\,\sqrt n\big(\bar g_{\Sigma}-b_n\big) + o_p(1). \] The condition $\|b_n\|^2\sqrt n\to 0$ ensures that higher order terms are negligible under sufficient smoothness conditions, justifying the use of the linear expansion (ref) for our design analysis. Because $b_n \to 0$, the Jacobian $\Lambda$ can be consistently estimated by evaluating the derivative of $g(\theta)$ at a preliminary observational estimate $\tilde{\theta}^{\mathrm{obs}} \rightarrow_p \theta$ andrews2020transparency.
\paragraph{Linearization of the estimand.} Let $\tau:\mathbb{R}^p\to\mathbb{R}$ be twice continuously differentiable near $\theta$, with gradient $\omega(\theta):=\partial\tau(\theta)/\partial\theta$ and bounded Hessian. A Taylor expansion around $\theta$ yields \[ \sqrt n\big(\tau(\hat\theta)-\tau(\theta)\big) = \omega(\theta)^\top\,\Gamma_\Lambda(W)\,\sqrt n\big(\bar g_{\Sigma}-b_n\big) \;+\; o_p(1), \] because $\|\hat\theta-\theta\|=O_p(n^{-1/2}+\|b_n\|)$ and the local condition $\sqrt n\,\|b_n\|^2\to 0$ renders the quadratic remainder $o_p(1)$ after $\sqrt n$ scaling.
Consequently, the regret expressions derived for the linear case continue to apply up to $o_p(1)$, with $\omega$ interpreted as the gradient $\omega(\theta)$ (or any consistent plug-in $\omega(\tilde\theta^{\mathrm{obs}})$ from a preliminary estimator $\tilde\theta^{\mathrm{obs}}\to_p\theta$).
Because large-scale experimentation is often infeasible due to cost constraints muralidharan2017experimentation, researchers often combine experimental and observational data to extrapolate effects from small-scale experiments to learn general-equilibrium effects ToddWolpin2006AER,attanasio2012education, meghir2022migration. The goal of this subsection is to present a step-by-step practitioner's guide with an example, and compare our regret-optimal design to a standard variance-minimizing (Neyman) benchmark.
Consider a researcher evaluating conditional cash transfers for sending children to school in rural Kenya. Due to budget constraints, the researcher can only randomize at a small scale (partial equilibrium), but ultimately wishes to predict GE effects. Researchers have access to preliminary (external) estimates from the Mexican PROGRESA experiment ToddWolpin2006AER, although these preliminary estimates may lack external validity in Kenya.
The first step for a researcher is to describe how observational/external and experimental estimates in Kenya are jointly used for estimation. Here, we consider a stylized model of school choice following BonhommeWeidner2021arXiv, ToddWolpin2006AER.
\paragraph{Individual choices} Let $S\in\{0,1\}$ denote school attendance, $C$ consumption, $Y$ (pre-transfer) household income, $W$ the child's potential wage, and $t$ the stipend when enrolled. Abstracting from covariates (we will introduce covariates in the estimation), utility is
with budget $C=Y+W(1-S)+tS$. The parametrization $(\xi_2-\xi_3-\xi_1)$ is without loss and simplifies expressions below. Enrollment satisfies $ S=1\big\{U(Y+t,1,t,\varepsilon)>U(Y+W,0,0,0)\big\}. $ Letting $ Z(Y,W,t)\equiv \xi_3 W-\xi_1 Y-\xi_2 t-\xi_4, $ we have $P(S=1\mid Y,W,t)=\Phi\!\big(-Z(Y,W,t)\big)$, with $\Phi(\cdot)$ denoting the standard Gaussian CDF.
\paragraph{Model for GE effects} We are interested in the effect of a small stipend $t$ to all eligible (poor) households in rural Kenya. GE feedback is allowed through income and wages which we write as functions of a common transfer $t$ to all eligible households in the population. Specifically for functions $y(t),w(t)$ and mean-zero idiosyncratic income and wage shocks $\varepsilon_{Yi}, \varepsilon_{Wi}$, for a common transfer $t$, denote the potential individual income and wage\footnote{Because we focus on the effect of a common transfer, we keep a scalar notation for $t$ for expositional convenience. Alternatively, here $y(t), w(t)$ can also be interpreted as exposure mappings aronow2017estimating and $t$ can be defined as a $n$ dimensional vector of transfers. } $$
$$ Let $y_0\equiv \partial_t y(t)\vert_{t=0}$ and $w_0\equiv \partial_t w(t)\vert_{t=0}$. Define $\phi_0\equiv \mathbb{E}\!\left[\phi\!\big(-Z\big(Y(0),W(0),0\big)\big)\right]$. Our estimand of interest is the marginal effect of an increase in a small transfer $t$ to all eligible households
which decomposes into the direct stipend effect $\phi_0\xi_3$ and the indirect GE effects via income and wages, $\phi_0\xi_1 y_0$ and $-\phi_0\xi_1 w_0$. Under standard market clearing for labor supply and demand, we write\footnote{This follows from the following: set $L_s(W,t)=L_d(W)$, the labor supply equals the labor demand, and differentiate at $t=0$: $\frac{\partial L_s}{\partial W}w_0+\frac{\partial L_s}{\partial t}=\frac{\partial L_d}{\partial W}w_0$. Suppose we have a constant number of hours worked for working child, denoted as $H$, and any child not at school is working. From the probit index, $\partial_t \mathbb{E}[S\,|\,\cdot] = \phi(-Z)\,\xi_2$ and $\partial_W \mathbb{E}[S\,|\,\cdot]=-\phi(-Z)\,\xi_3$ at $t=0$. Hence $\partial L_s/\partial t=-\phi_0\xi_2 H$ and $\partial L_s/\partial W=\phi_0\xi_3 H$. Dividing the equation by $H$ we obtain the desired result.}
where $d$ denotes the slope of the demand curve divided by total hours of work. Wages in equilibrium depend both on the demand's slope and the supply effect, captured through $\phi_0 \xi_2$ (assuming that all non-students are employed for a constant number of hours).
The following step is to define the parameters of interest, here defined as
with corresponding estimand
where the explicit form of $w_0(\theta)$ follows from (ref), a smooth function of the parameters.
We treat $d$ as known for simplicity, using meta-analyses on demand elasticities, and $y_0 = 1.5$ based on experimental estimates in Kenya from egger2022general.\footnote{egger2022general report a total income/consumption multiplier $M_{\text{tot}}=\partial_t \mathbb{E}\big[Y_{\text{pre}}(t)+t\big]\vert_{t = 0} = 2.5$; therefore the target derivative of pre-transfer income is $y_0=\partial_t \mathbb{E}\big[Y_{\text{pre}}(t)\big]\vert_{t = 0}=M_{\text{tot}}-1$.} Specifically, $d = 0.5 \frac{1 - S_0}{W_0}$, where $S_0$ and $W_0$ are baseline probability to go to school (assuming all children not in school go to work) and baseline wages calibrated to the pre-experimental data from PROGRESA and $0.5$ obtained from meta-studies in espey2000farm.\footnote{One could also calibrate $S_0, W_0$ to pre-experimental data in Kenya directly when this data is available.}
Here, the primary concern is bias in $\tilde{\theta}^{\mathrm{obs}}$, abstracting from potential biases in $y_0, d$. When biases of $(y_0, d)$ are also considered first-order, researchers should let $y_0,d$ be additional parameters in $\theta$, which is allowed in our framework. This difference illustrates a feature of our framework: the chosen parameterization ensures transparency about which bias concerns drive the design; these considerations can be collected in a pre-analysis plan.
Using experimental data from the large-scale Mexico’s PROGRESA program, we estimate $(\xi_1,\xi_2,\xi_3,\xi_4)$ through probit with standard controls (age, distance to school, eligibility, year, highest grade achieved) and use these estimates to obtain $\tilde{\theta}^{\mathrm{obs}}$. Estimation is conducted separately by gender; we will focus here on the effects on female students.
We pool treated and control observations to estimate the schooling effect $\xi_4$, controlling for the individual-level subsidy. In turn, we obtain observational estimates $\big(\tilde{\theta}_1^{\mathrm{obs}},\tilde{\theta}_2^{\mathrm{obs}},\tilde{\theta}_3^{\mathrm{obs}}\big)$ that map to the model parametrization. Standard errors are constructed using the Delta method with clustering at the village level. Following Section (ref), we then construct $\omega = \frac{\partial \tau}{\partial \theta}\Big|_{\theta = \tilde{\theta}^{\mathrm{obs}}}$. Point estimates, estimated sensitivity $\omega$, and standard errors are reported below.
In this example, we consider a researcher who can afford only small, partial-equilibrium experiments in Kenya, and examine three stylized experiments or combinations thereof:
The per-unit variance in the experiment is calibrated to the corresponding entry of $\sigma^2$ in observational study, appropriately multiplied by the observational study sample size.
We consider two scenarios. First, the researcher may run only one of the three programs. Second, the researcher may run combinations of two of the three programs allocating a fixed total sample size to each, but not all three simultaneously, reflecting higher fixed costs and institutional constraints. For instance, the job program requires a different implementation infrastructure than a cash-transfer program; at the same time, conducting conditional and unconditional cash transfers in the same experimental location is often infeasible due to fairness concerns wingfield2023experiences, which raises costs for each cash-transfer arm.
Finally, we design the experiment, with total number of experimental observations $n_{\mathrm{tot}}$; let $n_{\mathrm{obs}}$ denote instead the number of units in the original Mexican PROGRESA experiment.
\paragraph{Choice of the experiment and sample size} In Figure (ref) we show the optimal sample size allocation as a function of the total experimental sample size, under the constraints that either one (left panel) or two (right panel) treatment arms can be selected. When the total sample size is small and only one arm can be chosen, the optimal design selects the UCT arm. Intuitively, the UCT strikes a good balance between sensitivity and precision of the experimental estimate when $n_{\text{tot}}$ is small.
When two treatment arms can be selected, the optimal design initially chooses both UCT and CCT for $n_{\mathrm{tot}} \le 500$. In this region, almost all of the experimental sample is allocated to the CCT arm (which is more sensitive but noisier), with a smaller fraction allocated to the more precise UCT arm. For total sample sizes above five-hundred, the optimal design instead selects the job program as the main experiment, reflecting the high sensitivity of the estimand to the wage effect (Appendix Figure (ref) reports the bias and variance regret of the design). With two experiments and large sample sizes, roughly 90% of the total sample is allocated to the job program and the remaining fraction to the UCT arm, using the precise UCT effect to help balance variance against bias.
In addition to sample allocation, our method optimizes the shrinkage weights of the parameters of interest. Appendix Figure (ref) collects the estimated shrinkage weight across designs. We find that the shrinkage weight is above 0.7, and closer to one for most cases. This shows how the experimental variation gets significantly higher weight.
\paragraph{Main bias and variance regret comparisons} In Figure (ref) we compare the proposed optimal design to the variance-optimal (Neyman) design in terms of both variance and bias. This comparison clarifies what is gained by our procedure relative to a benchmark that chooses the experiment, shrinkage weights, and sample size solely to minimize the variance of the estimated effect $\hat{\tau}$ (and which therefore achieves the smallest possible variance). We find that the variance ratio between the Neyman and optimal designs (the inverse of the variance regret) is close to one for all sample sizes and converges to one as the sample size grows. For experimental sample $n_{\mathrm{tot}} = 1000$, the variance of the Neyman and our proposed procedure is only about $14\%$ smaller when choosing one treatment arm and $35\%$ when choosing two treatment arms. In contrast, the bias of the Neyman procedure is about 1.4 times larger for one treatment arm and about 4.5 times larger with two. As a result, the adaptation regret of the Neyman allocation is similarly 1.4 and 4.5 times larger than our proposed design. These comparisons illustrate the benefit of our method: its variance is only slightly larger than the Neyman variance, while the bias is more than four times smaller for two treatment arms.
Next, we show how observational evidence can inform whom to recruit into an experiment. We consider the setting in banerjee2024changes, who study how the expansion of microfinance affects village social networks in Karnataka, India. We focus on the first outcome reported by banerjee2024changes corresponding to the density of the network, i.e., the percentage of connections a random household in a village has relative to the village size.
\paragraph{Background and research question} In their observational analysis, the authors assemble panel data from 75 villages in Karnataka, 43 of which were exposed to microfinance. Because program rollout was not randomized, they estimate effects using Difference-in-Differences (DiD). Although DiD is informative, the lack of randomized assignment can bias estimates in the presence of selection ghanem2022selection. Therefore, the authors subsequently conducted an experimental evaluation in one metropolitan area, randomizing microfinance access across 104 urban neighborhoods. We only use observational variation for choosing the experimental design and estimator. We then use experimental variation from banerjee2024changes to validate our procedure by calibrating its worst-case bias.
In practice, implementation costs often depend on how geographically dispersed the study sites are due to coordination and survey costs. Using preliminary observational estimates from banerjee2024changes in Karnataka, we ask: Which area(s) in Karnataka should be prioritized for an experimental evaluation, and how many villages should be enrolled in each area? We focus on the experiment randomizing treatments in one or two areas.
\paragraph{Observational study} We partition the observational sample into four geographically contiguous areas that group nearby villages; each area contains 11--12 villages that were exposed to microfinance during the study period. Areas are heterogeneous in terms of their overall population size. Figure (ref) maps these four areas and reports the corresponding DiD point estimates. Three of the four areas display negative estimated effects, with meaningful variation in magnitudes across areas. Table (ref) summarizes, for each area, samples sizes, variances clustered at the village level, and population sizes ($\omega$). These area-specific DiD estimates and their variances serve as our primary observational inputs for experimental design.\footnote{We use DiD estimates to mimic the estimator used by the authors. An analogous analysis can be conducted after empirical-Bayes shrinkage of area-level estimates, omitted for brevity.} Using the 2011 population census, matched with the geographic boundaries of the areas obtained from merging village positions with the SHRUG platform asher2019socioeconomic, we compute $\omega$ as the population share of a given area.
\paragraph{Experimental design} We consider a family of designs that (i) select $E\in\{1,2\}$ geographic areas in Karnataka from which to recruit experimental sites and (ii) assign $n_1$ villages to treatment and $n_0=n_1$ to control (so the total sample is $n=2n_1$); (iii) optimizes over shrinkage weights as in Algorithm (ref). We examine a grid of values for $n_1$ ranging from 40 to 80 (with 52 corresponding to the size of the experiment in banerjee2024changes). For variance calibration, we take $v_{\mathrm{pre},a}^2, a \in \{1, \cdots, 4\}$ to be the pre-intervention, area-level variance of network density in Table (ref). We assume that the variance of a single treated-control difference in the experiment is $2 v_{\mathrm{pre},a}^2$.
\paragraph{Choice of the areas} Figure (ref) reports the optimal sample size allocation across areas when either one or two areas can be selected. When only one area can be chosen, the method selects the area with the largest weight $\omega$, which in this application corresponds to the area containing Bengaluru (about 76% of the population). When two areas can be selected, the procedure chooses the two areas with the largest $\omega$ (Area 2 and Area 3, the latter containing Tumakuru), where the second area has weight $\omega \approx 12\%$. The method also optimally allocates the sample size across these areas, assigning roughly 10 to 20 villages in total (half of which treated) to Tumakuru and the remaining villages to the larger area containing Bengaluru. Figure (ref) shows the optimal shrinkage weights: the shrinkage weight for the Bengaluru area is about 93%, reflecting an interior solution, while it is essentially one for Tumakuru. The resulting pattern is simple and intuitive: assign most of the sample size to the area with the largest weight, but, when possible, include a second area with substantial weight, letting the allocation of villages across areas balance robustness and variance.
\paragraph{MSE comparisons} We compare two strategies: (i) our robust design, which jointly chooses the area(s), the sample allocation, and the shrinkage weights $\gamma^\star$ as described in Section (ref); and (ii) the variance–optimal (Neyman) allocation. We evaluate each strategy using the worst-case MSE over bias vectors satisfying $||b||_\infty \le B$. Rather than treating $B$ as a free sensitivity parameter, we calibrate it using experimental evidence from banerjee2024changes. Specifically, we set $B$ equal to the largest absolute discrepancy between the Hyderabad RCT ATE reported in the follow-up experiment of banerjee2024changes and the corresponding Karnataka DiD estimate in each area. Importantly, our design procedure does not use this calibrated bound $B$ as an input; it is used only for the ex-post MSE comparison.
Figure (ref) plots regret (the ratio $\text{MSE}/\text{MSE}^\star$, $\text{MSE}^\star$ is the smallest achievable worst-case MSE for a bias upper bound at most $B$, and MSE is the worst case MSE for $||b||_{\infty} \le B$). As the number of treated villages $n_1$ increases, all designs improve; because the denominator also falls with $n_1$, the ratio need not be monotone. Across $E\in\{1,2\}$, the Neyman benchmark remains well above the oracle— with regret around 2.5 or more. On the other hand, our design tracks the oracle more closely, with the ratio closer to one and no larger than 1.2. Appendix Figure (ref) collects similar comparisons for the worst-case regret over all possible values of $B$.
This paper studies experimental design in the presence of observational evidence. A key challenge is that the bias of the observational estimators is unknown in practice. We adopt a minimax adaptation regret criterion that compares the mean-squared error (MSE) of a candidate design to that of an oracle that knows the worst-case bias. This reveals a fundamental trade-off between precision and robustness. The optimal design balances the design’s variance normalized by the smallest achievable variance (variance gap) and its worst-case bias normalized by the smallest attainable bias (bias gap). We propose a procedure that jointly determines: (i) how to combine observational and experimental evidence; (ii) how to allocate precision across experiments given budget constraints; and (iii) which treatment arm and/or sub-population to include in the experiment with fixed experimental costs.
In practice, the workflow is:
The applicability of our framework spans a wide range of settings in economics and beyond. Examples include estimating general equilibrium effects or structural models ToddWolpin2006AER, attanasio2012education, meghir2022migration, kreindler2023optimal, de2025decoupling; choosing among alternative treatment arms in factorial designs muralidharan2019factorial, bandiera2025illusion; and deciding where to run the next experiment for external validity gechter2024selecting, olea2024externally. In industrial organization, applications include choices between demand-side and supply-side interventions bergquist2020competition, the effects of information acquisition in markets allende2019approximating, larroucau2024college, and decisions about which additional data source to acquire to improve statistical analysis allcott2025sources. Beyond economics, medical applications may include allocating sample size across subgroups and allocating doses across treatment arms porter2024phase, morita2017simulation, manski2025using.
Several open questions remain for future research. These include settings where researchers have well-specified priors about bias such as from Bayesian models gechter2024selecting, or other modeling choices olea2024externally. This may be potentially incorporated through additional moment restrictions as in Section (ref). Additional future extensions include sequential or adaptive experimental choices cesa2025adaptive.
\singlespacing
\onehalfspacing