EconBase
← Back to paper

When and How to Pilot: Design Rules for Two-Wave Experiments

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

107,847 characters · 23 sections · 45 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

-2.5cm When and How to Pilot: Design Rules for Two-Wave Experiments

bibunit\\ { Brown University }} \bgroup \let\footnoterule\relax \begin{singlespace} \begin{abstract} Experimenters often run pilots, but how much a small pilot should shape the main-wave design has no settled answer. This paper shows how noisy pilot evidence should guide treatment assignment probabilities in two-wave experiments. Two canonical rules mark the extremes. Balanced assignment guards against worst cases but ignores evidence that one arm is noisier. Feasible Neyman allocation adapts, but with a finite pilot it can overreact to noise, producing arbitrarily large precision losses. We propose a Conditional Minimax Regret (CMR) rule that minimizes worst-case regret over a finite-sample confidence set for the treatment and control variances. CMR retains balance's worst-case protection with high probability, converges to the Neyman allocation as the pilot grows, and attains the minimax-regret rate up to constants. It extends to multi-arm and stratified designs, and simulations calibrated to four field experiments show it avoids feasible Neyman's severe small-pilot losses while capturing most of its large-pilot gains. \end{abstract} \end{singlespace} Keywords: Experimental design; Statistical decision theory; Minimax regret.\\ JEL Codes: C44, C90, C93, D81. \thispagestyle{empty} \egroup \setcounter{page}{1}

Introduction

Nearly one in ten of the 10{,}905 trials registered in the AEA RCT Registry over the past decade reports running a pilot.\footnote{Between 2016 and 2025, 991 of the 10{,}905 first-registered trials, or 9.1%, report running or using a pilot. Only 24, 0.22% of all trials in the window, report using the pilot to set the treatment-assignment probability. Online Appendix (ref) provides full details.} Experimenters use them to test logistics, refine survey instruments, and learn about the study population before launching the main wave. Pilot data can also inform a consequential design choice, the treatment assignment probability. By determining how many observations each arm receives, this probability directly affects the variance of the average treatment effect estimator and hence the experiment's power. The promise of pilot-based assignment is that it can use early evidence to allocate more observations where they are most valuable, and hence deliver more precise estimates from the same sample size.

If the potential-outcome variances under treatment and control were known, the variance-minimizing assignment rule would be the Neyman allocation neyman1992two. It assigns a larger share of the experimental sample to the noisier arm, reflecting the principle that precision requires more observations where outcomes are more variable. In practice, these variances are unknown when the main-wave design is chosen. A natural rule estimates them from the pilot and substitutes the resulting sample variances into the Neyman formula, yielding the feasible Neyman allocation (FNA). When the pilot is large enough for the estimates to be reliable, this plug-in logic is well founded. For example, hahn2011adaptive show that pilot-based estimates of conditional variances can deliver asymptotically efficient assignment probabilities when both pilot and main-wave samples grow large.

When the pilot is small, however, the premise of this plug-in logic fails, and the experimenter must decide how much the design should respond to variance estimates that are not necessarily trustworthy. Keeping the main wave evenly split ensures that neither arm ends up with too few observations, which caps the possible precision loss whatever the true variances. The cost is that it forgoes the gains the pilot could deliver. Following the pilot estimates as if they were true removes this protection, and can tilt the main wave too far, or toward the wrong arm, precisely when the pilot is least informative.

These small-sample concerns are empirically relevant, since in the AEA RCT Registry the median reported pilot has about 300 observations and roughly a tenth have fewer than 30, sizes at which variance estimates can remain very noisy. In fact, cai2024performance show that a fixed pilot can leave the feasible Neyman allocation less precise than balance even when the main wave is large. This makes pilot-based assignment a finite-sample design problem, not a plug-in calculation.

We formalize the experimenter's choice as a finite-sample statistical decision problem. In a two-wave experiment with binary treatment, bounded potential outcomes, and no parametric restrictions, the experimenter observes a pilot and then chooses the main-wave assignment probability. The goal is to minimize the variance of the average treatment effect estimator. We evaluate decision rules by their regret, the excess variance relative to the infeasible Neyman allocation that would be chosen if the variance pair were known. Although the setting is deliberately simple, the question it isolates is general. Design rules should be optimized for the information actually available at the moment of design, not for an asymptotic world in which pilot estimates are treated as known. Pilot-based assignment is the tractable model case in which that principle can be developed exactly.

In this setup, the two canonical assignment rules mark the endpoints of the finite-pilot design problem. Balanced assignment, the no-data rule, is optimal under minimax risk and is uniquely minimax-regret optimal without a pilot. It therefore has a formal foundation for caution, but it gives the realized pilot no authority over the assignment. Feasible Neyman gives the pilot's point estimates full authority. This plug-in logic attains the large-pilot efficiency benchmark armstrong2022asymptotic, but with a finite pilot nothing limits how far noisy estimates can tilt the assignment, so the precision loss can be arbitrarily large---in the extreme, an arm whose pilot observations all coincide is estimated to have zero variance and receives no main-wave observations at all. In short, balance refuses to move, while feasible Neyman can move too far. The finite-pilot question is how much authority the realized evidence has earned, or equivalently how far the assignment should move from balance toward the feasible Neyman allocation.

Exact minimax regret is the natural ex ante criterion for finite-pilot adaptation savage1951theory, manski2021econometrics. It asks, before the pilot is drawn, which complete mapping from pilot realizations to assignments has the smallest worst-case expected regret. By this criterion, with as few as two observations per pilot arm, the minimax-regret value falls strictly below its no-pilot level, so any exact minimax-regret rule must move the assignment away from balance for some pilot realizations. The criterion does not, however, yield an operational design rule for this problem. Regret depends on the population distribution only through the treatment and control variances, yet two populations with the same variances can generate different pilot data, so the minimax problem ranges over full outcome distributions and is not directly computable. More importantly, the criterion evaluates complete pilot-to-assignment mappings, averaging regret over every pilot a population could generate. The authority question starts at the other end, from the single pilot actually drawn, and asks which variance pairs the evidence has ruled out and how much movement from balance the surviving configurations justify.

The Conditional Minimax Regret (CMR) rule answers the authority question with a principle from the literature on inference for decision making, acting on what the data have not ruled out rather than on a point estimate manski2021econometrics, chernozhukov2025policy. Once the pilot is observed, the only uncertainty that matters for the loss is which variance pair is true. CMR therefore builds a finite-sample confidence set for the treatment and control variances, the configurations still consistent with the realized pilot, and chooses the main-wave assignment that minimizes worst-case regret over that set. In the binary-treatment case the rule has a closed form, the Neyman formula applied to the midpoint of each arm's standard-deviation confidence interval. When the set is wide, CMR stays close to balance. As the set contracts around the true variances, CMR moves toward the Neyman allocation. Alongside the assignment, CMR reports a certificate, a bound on the regret the chosen assignment can incur andrews2025certified. Movement from balance is thus earned by what the pilot rules out, not imposed by a point estimate.

The CMR rule comes with finite-sample and large-pilot guarantees. In finite samples, the certificate bounds the realized regret of the chosen assignment with probability at least \(1-\alpha\), turning a loss the experimenter cannot observe into a number computed from the pilot alone. This certificate is never larger than the guarantee available without a pilot, and it tightens as soon as the pilot rules out one of the adversarial configurations that make a large tilt dangerous. As the pilot grows, the caution that holds CMR near balance relaxes, and at interior variance pairs the rule converges to the infeasible Neyman allocation at the same rate as feasible Neyman. Thus asymptotic efficiency and finite-sample safety hold together rather than trading off. Specifically, CMR's worst-case expected regret matches the minimax-regret rate up to constants, the best rate any pilot-based rule can achieve.

The CMR construction is flexible, and it remains simple to compute across the designs experimenters actually run. In multi-arm experiments with a shared control and in stratified designs, the assignment probability becomes an allocation vector, and CMR again moves from the no-information allocation toward the Neyman allocation as the pilot narrows the confidence set. The same recipe adapts to how outcomes are measured. Binary outcomes admit exact confidence sets that sharpen the rule precisely at the small pilots where it matters most, a kurtosis condition can replace the assumption of a known outcome bound, and the construction carries over when the design targets several primary outcomes or when the pilot observes only a short-run proxy of the outcome that defines the loss.

We also ask when a pilot is worth running for design in the first place, and how large it should be. Observations spent on the pilot could have gone to the main wave instead, so adaptation has to repay that diversion. A worthwhile pilot must be large enough for CMR to move away from balance at all, yet small enough that even perfect adaptation recovers the observations it cost. Folding pilot observations into the final estimator softens this trade-off. The worst-case optimal design cannot be computed, but a simple rule approximates it. A balanced pilot at the two-thirds power of the budget, followed by CMR, comes within a constant factor of the best worst-case performance any two-wave design can achieve, and any design coming close must size its pilot the same way, so larger experiments pilot more but devote a smaller share of their sample to piloting.

Calibrated simulations document that these finite-sample concerns are quantitatively important at realistic pilot sizes. In data-generating processes built from the public microdata of four well-known field experiments, feasible Neyman allocation incurs infinite or very large losses at small pilots, while CMR stays at balance until the pilot makes the confidence rectangle informative, then captures most of the attainable gain. In the two-arm designs that attainable gain is below one percent, so the realistic case for a pilot-based rule there is insurance, protection against the plug-in's unbounded downside at essentially no cost in upside. The pattern sharpens in the multi-arm and stratified designs, where the same pilot is split across more cells. There balance leaves more precision on the table, plug-in Neyman becomes more fragile, and CMR's discipline is worth correspondingly more.

The broader contribution is a template for building statistical uncertainty into experimental design itself. Everything in the paper runs on three features of the assignment problem. First, the unknown parameters enter the design loss only through a low-dimensional vector of variances. Second, the loss is linear in each variance, which makes regret convex, so worst-case regret over any rectangle of plausible values is attained at one of its corners. Third, a small pilot delivers a finite-sample confidence bound for each variance at any prespecified level. Any design problem with these features admits the same treatment, a decision paired with a finite-sample certificate for the loss it can incur, and every extension in this paper is an instance of that recipe.

Two literatures frame these contributions. The first is adaptive experimental design and pilot-based assignment. hahn2011adaptive set assignment probabilities from pilot data, tabord2023stratification builds adaptive stratification rules, and bai2022optimality and cytrynbaum2021optimal design matched and stratified experiments from pre-experimental information, with armstrong2022asymptotic characterizing the efficiency frontier these designs target. These guarantees are asymptotic. They identify the efficient design once the pilot estimates the variances reliably, but they are silent when the pilot is too small to be trusted. Building on the finite-pilot warning of cai2024performance, we formulate pilot-based assignment as a finite-sample decision problem and derive a computable post-pilot rule that pairs the treatment share with a regret certificate.\footnote{A rapidly growing statistics and machine-learning literature studies adaptive Neyman allocation in fully sequential or many-batch experiments, including finite-sample Neyman-regret guarantees dai2023clip, zhao2024adaptive. Those rules adapt the assignment continuously as outcomes arrive. Ours solves a one-shot design problem, choosing the main-wave assignment after a single finite pilot.} The contribution is a characterization of how far noisy pilot evidence should move a design while the variances remain uncertain.

The second is inference for decision making. manski2021econometrics examines as-if decisions with set estimates, which act on the parameter values a set estimate has not ruled out. ishihara2021evidence aggregate estimates from prior studies into a minimax-regret treatment choice, chernozhukov2025policy select policies by balancing estimated welfare against estimation risk, and andrews2025certified pair recommended decisions with high-probability bounds on their loss. In these settings, the data are already in hand and the decision is which policy to implement. Our decision comes one stage earlier. The action is the main-wave assignment probability, and the loss is the precision of an estimator whose data do not yet exist.\footnote{The regret criterion follows the decision-theory tradition of savage1951theory, manski2004statistical, stoye2009minimax,stoye2012minimax, tetenov2012statistical, and manski2016sufficient. That literature usually focuses on treatment choice; we focus instead on experimental design. The closest is hu2024minimax, who use minimax regret for sample selection, but derive their rules in a local asymptotic framework rather than our finite-sample one.} This structure yields a rule that is closed form in the binary case, immune to the failures of plug-in rules, and accompanied by a certificate for the realized pilot.

The rest of the paper is organized as follows. Sections (ref) and (ref) set up the decision problem and the canonical rules, Section (ref) develops the CMR rule and its guarantees, Section (ref) treats multi-arm and stratified designs, Section (ref) reports the calibrated simulations, and Section (ref) concludes. Online Appendix (ref) extends the rule to binary, unbounded, multiple, and delayed outcomes, and Online Appendix (ref) analyzes when a pilot is worth running and how large it should be.

The Two-Wave Design Problem

Experimental Setup

\paragraph{Population, Outcomes, and Estimand.}

An experimenter aims to estimate the average treatment effect (ATE) of a binary treatment in a target population. The population is characterized by the joint distribution \(F\) of the potential outcomes \((Y(1),Y(0))\), where \(Y(1)\) denotes the outcome under treatment and \(Y(0)\) denotes the outcome under control. Throughout, we assume that potential outcomes take values in the unit interval, \(Y(d)\in[0,1]\) for each treatment status \(d\in\{0,1\}\), so \(F\in\mathcal F=\mathcal P([0,1]^2)\), the collection of all Borel probability measures on the compact set \([0,1]^2\).\footnote{Since outcomes lie in \([0,1]\), all moments of \(F\) exist and are finite. In particular, the marginal variances are bounded by \(1/4\).} Normalizing outcomes to \([0,1]\) is without loss of generality whenever outcomes are supported on a finite interval, and the particular choice of zero and one is adopted only to simplify notation and exposition. Online Appendix (ref) relaxes boundedness, replacing the known outcome bound with a bound on kurtosis.

The estimand of interest is the ATE, \(\operatorname{ATE}(F)=\mathbb E_F[Y(1)-Y(0)]\), where the expectation is taken with respect to \(F\in\mathcal F\). To increase the precision of the ATE estimator, the experiment proceeds in two waves. A smaller pilot precedes a larger main wave and serves only to inform its design. The single design choice is the main-wave assignment probability. Lower estimator variance translates directly into higher power and a smaller minimum detectable effect at any target power, so variance-minimizing assignment makes the experiment as informative as possible about the ATE.

\paragraph{Pilot Wave.}

The pilot is a sample of \(M\) units whose potential outcomes \(\{(Y_i(1),Y_i(0))\}_{i=1}^M\) are i.i.d. draws from \(F\). For each unit \(i\), the experimenter observes the treatment indicator \(D_i\in\{0,1\}\) and the realized outcome \(Y_i=D_iY_i(1)+(1-D_i)Y_i(0)\). Treatment in the pilot is assigned by a completely randomized design (CRD).\footnote{Fixed pilot arm sizes are imposed only to keep notation simple. The same arguments extend to Bernoulli assignment conditional on both realized arm sizes being at least two.} In the pilot CRD, exactly \(M_1\) units are assigned to treatment and the remaining \(M_0\) to control, with \(M_1+M_0=M\). The assignment vector \((D_1,\ldots,D_M)\) is sampled uniformly at random from all binary vectors with exactly \(M_1\) treated units.

The pilot realization is denoted by \(\omega=\{(Y_i,D_i)\}_{i=1}^M\), and the pilot sample space is \(\Omega=([0,1]\times\{0,1\})^M\), with the fixed-arm-size restriction imposed by the pilot design. Each pilot arm holds at least two units, \(M_1,M_0\geq 2\), so the within-arm sample variances are well defined. This minimum is maintained wherever the pilot is fixed in advance, and is relaxed only in Appendix (ref), where the pilot size is itself the object of choice. The extensions in Section (ref) impose the same minimum on each relevant cell.

\paragraph{Main Wave.}

The main wave is a new sample of \(N\) units, independent of the pilot, whose potential outcomes \(\{(Y_j(1),Y_j(0))\}_{j=1}^N\) are i.i.d. draws from \(F\). The main wave also uses a completely randomized design. For a treatment assignment probability \(\pi\in(0,1)\), the design assigns \(N_1=N\pi\) units to treatment and \(N_0=N(1-\pi)\) units to control.\footnote{We suppress integer constraints on \(N\pi\) throughout, with \(\pi\) understood as the target treatment assignment probability implemented by the closest feasible fixed-arm design. The same logic applies under Bernoulli assignment with the estimator of hajek1971.} The experimenter estimates the ATE using the difference in means, \(\widehat{\operatorname{ATE}}=\bar Y_1-\bar Y_0\), where \(\bar Y_d\) is the main-wave average outcome among units assigned to arm \(d\).

Main-Wave Variance and the Neyman Allocation

The marginal potential-outcome variance in arm \(d\) is \(\sigma_d^2(F)=\operatorname{Var}_F(Y(d))\), abbreviated \(\sigma_d^2\) when \(F\) is clear from context. Under the main-wave CRD, \(\operatorname{Var}_F(\widehat{\operatorname{ATE}})=\sigma_1^2/N_1+\sigma_0^2/N_0\), with no cross-arm covariance term because treatment and control are measured on independent units. With \(N_1=N\pi\) and \(N_0=N(1-\pi)\), this equals \(V(\pi,\sigma_1^2,\sigma_0^2)/N\), where \(V(\pi,\sigma_1^2,\sigma_0^2)=\sigma_1^2/\pi+\sigma_0^2/(1-\pi)\) for \(\pi\in(0,1)\). The main-wave sample size \(N\) does not depend on the assignment, so the sampling variance is proportional to \(V\) with fixed constant \(1/N\), and minimizing the variance over \(\pi\) is the same as minimizing \(V\), which we take as the variance criterion for the assignment problem. Boundary recommendations, \(\pi\in\{0,1\}\), are assigned infinite variance by convention, since assigning all main-wave units to one arm leaves the other arm mean unobserved.

Minimizing \(V(\cdot,\sigma_1^2,\sigma_0^2)\) over \(\pi\in(0,1)\) yields the Neyman allocation neyman1992two, which for strictly positive variances is \(\pi^*(\sigma_1^2,\sigma_0^2)=\sigma_1/(\sigma_1+\sigma_0)\). The Neyman allocation gives the larger share of the main wave to the arm with the more variable outcomes, because that arm's sample mean is noisier and benefits more from a larger allocation.

The benchmark is the smallest variance an interior assignment can achieve,

equation[equation omitted — 145 chars of source]

attained at the Neyman allocation when both variances are strictly positive. Writing \(V^*\) as an infimum keeps it well defined when one variance is zero, since the minimizer then lies on the boundary and is only approached from the interior.\footnote{With \(\sigma_0=0\), for instance, the Neyman formula returns the boundary value \(\pi^*=1\). When both variances are zero, every interior assignment attains the value zero, and we normalize \(\pi^*(0,0)=1/2\).} Because \(V^*\) depends on the unknown variance pair, it is the infeasible benchmark against which feasible assignment rules are evaluated.

remark[How much adaptation can gain] The benchmark caps what any pilot-based rule can gain over balance. For every variance pair, \(V(1/2,\sigma_1^2,\sigma_0^2)/V^*(\sigma_1^2,\sigma_0^2)=2(\sigma_1^2+\sigma_0^2)/(\sigma_1+\sigma_0)^2\le 2\), with equality only when one variance is zero. Even perfect adaptation can therefore at most halve the estimator's variance, equivalent to doubling the effective sample size, and realistic configurations deliver far less. Balance's excess variance relative to the benchmark is \((\sigma_1-\sigma_0)^2/(\sigma_1+\sigma_0)^2\), about eleven percent when one standard deviation is twice the other and four percent when it is fifty percent larger. The ceiling rises in the designs of Section (ref), where the no-information allocation must divide the sample more finely. With \(K\) treatment arms sharing a control, perfect adaptation can cut variance by up to a factor of \(K+\sqrt K\) rather than two. In a stratified design the factor is two divided by the smallest stratum share, so with \(S\) equally sized strata it is \(2S\).

States, Actions, and Pilot-Based Rules

The two-wave design problem is a statistical decision problem. The unknown state of the world is the population distribution \(F\). It both fixes the main-wave variance of any assignment and generates the pilot data through which the experimenter learns about that variance. Because \(F\) is left unrestricted beyond bounded support, the parameter is the full distribution and the parameter space is the nonparametric class \(\mathcal F\). However, the main-wave variance depends on \(F\) only through the pair of marginal potential-outcome variances \(\theta(F)=(\sigma_1^2(F),\sigma_0^2(F))\), which ranges over \(\Theta=[0,1/4]^2\). This pair is the payoff-relevant parameter.

Crucially, the variance pair governs the payoff but not the data. At a fixed assignment, the loss is the same under any two distributions sharing a variance pair, since it depends on \(F\) only through \(\theta(F)\). The expected loss of a rule can still differ between them, because the rule chooses its assignment from the pilot, whose distribution depends on the arm-specific marginal distributions of \(Y(1)\) and \(Y(0)\) rather than on the variance pair alone. Under the fixed-arm pilot design each unit reveals only one potential outcome, so the pilot distribution depends on \(F\) only through these two marginals, and never on the within-unit dependence between \(Y(1)\) and \(Y(0)\). The statistical experiment is thus indexed by the pair of marginal outcome distributions and not by the variance pair.

The experimenter's action is the main-wave assignment probability \(\pi\in\mathcal A=[0,1]\). The boundary actions \(\pi\in\{0,1\}\) are kept in \(\mathcal A\) so that the framework can evaluate rules that would place the entire main wave in one arm, and the infinite-variance convention of Subsection (ref) applies to them.

The experimenter does not know the variance pair and estimates it from the pilot. For each arm \(d\), the natural estimate of \(\sigma_d^2\) is the within-arm sample variance

equation*[equation* omitted — 151 chars of source]

Under the fixed-arm pilot design, $\hat\sigma_d^2$ is unbiased for $\sigma_d^2$. The main analysis focuses on variance-based decision rules, meaning rules that use the pilot only through the two within-arm sample variances. Formally, the class of decision rules is

equation*[equation* omitted — 228 chars of source]

The class includes balance, feasible Neyman, trimmed feasible Neyman, and the Conditional Minimax Regret rules developed below. This focus is deliberate, reflecting how pilot-based assignment is typically implemented in practice and keeping the finite-sample design problem low-dimensional enough to yield tractable rules. However, the restriction is substantive since, relative to the unrestricted class $\mathcal{D}_0 = \{p \in \mathcal{A}^\Omega \mid p \text{ is measurable}\}$, features of the full pilot beyond $(\hat\sigma_1^2,\hat\sigma_0^2)$ may contain additional information about $(\sigma_1^2,\sigma_0^2)$ in the nonparametric model.

Once a decision rule is fixed, the realized pilot outcome $\omega$ remains random under $F$. For each $F\in\mathcal F$, let $P_F$ denote the distribution of $\omega$ induced by the fixed-arm pilot CRD with arm sizes $(M_1,M_0)$, suppressing this dependence in the notation. The statistical experiment is the family $\{P_F:F\in\mathcal F\}$ on $\Omega$.

Loss and Regret

The variance criterion \(V(\pi,\sigma_1^2,\sigma_0^2)\) is the loss from using treatment assignment probability \(\pi\) when the variance pair is \(\theta=(\sigma_1^2,\sigma_0^2)\). Regret compares this loss with the infeasible benchmark in (ref). For an interior assignment probability, $r(\pi,\theta) = V(\pi,\sigma_1^2,\sigma_0^2) - V^*(\sigma_1^2,\sigma_0^2)$. For boundary assignments, regret is infinite by the same convention used for \(V\). For a pilot-based rule \(p\in\mathcal D\), the realized regret after observing pilot sample \(\omega\) is \(r(p(\omega),\theta(F))\). Regret measures excess main-wave sampling variance relative to the infeasible Neyman benchmark, so it is always nonnegative. For nondegenerate variance pairs with strictly positive variances, regret is zero exactly when the chosen assignment equals the Neyman allocation. At degenerate variance pairs, regret remains well defined because the infeasible benchmark is an infimum over interior assignments.

For every interior assignment \(\pi\in(0,1)\), regret has the equivalent representations

equation[equation omitted — 243 chars of source]

with the normalization \(\pi^*(0,0)=1/2\) in the first. Expression (ref) separates the size of the assignment mistake from the penalty attached to it. The squared distance \((\pi-\pi^*(\theta))^2\) to the Neyman target is weighted by \(1/[\pi(1-\pi)]\), so a given mistake is most costly near the boundary, where one arm is barely sampled, and the factor \((\sigma_1+\sigma_0)^2\) scales the loss with the outcome noise. The second form, the square of an expression affine in the two standard deviations, is the version the analysis of Section (ref) exploits.

Canonical Assignment Rules

This section studies the two assignment rules that current practice and the existing literature bring to the finite-pilot design problem. The first canonical rule is complete balance, which ignores the pilot and assigns treatment with probability one half. Although simple, balance has a decision-theoretic foundation as the solution to the minimax-risk problem. The second is the feasible Neyman allocation, which uses the pilot in the most direct way, replacing the unknown potential-outcome variances with their pilot estimates.

These two rules expose the central tension of the design problem. Balance is safe but does not adapt, while feasible Neyman adapts but is not safe. We then turn to exact minimax regret, the standard decision-theoretic criterion for resolving this tension. The criterion identifies the best worst-case performance any pilot-based rule can achieve, but it is not itself an operational assignment rule, because the game between the experimenter and Nature ranges over full outcome distributions rather than variance pairs.

Minimax Risk and Balanced Assignment

Intuitively, minimax risk asks for the safest assignment rule when the true variance configuration is unknown, even if the pilot is misleading. Formally, the experimenter may choose any pilot-based rule \(p\in\mathcal D_0\), and each rule is evaluated by its frequentist risk, $\mathcal{R}(p,F) = \mathbb E_{P_F} \!\left[ V\!\bigl(p(\omega),\sigma_1^2,\sigma_0^2\bigr) \right]$, the expected variance of the main-wave ATE estimator when the pilot is generated under \(F \in \mathcal F\) and assignment follows rule \(p\). The minimax-risk criterion ranks rules by their worst-case risk over the model class \(\mathcal F\). The following proposition shows that complete balance is a minimax-risk rule.

proposition[{\normalfont\hyperlink{proof:minimax_risk}{Minimax risk}}] The minimax-risk problem $\inf_{p \in \mathcal D_0} \sup_{F \in \mathcal F} \mathcal R(p,F)$ has value \(1\), and the constant balanced rule \(p_{\mathrm{mm}}(\omega) = 1/2\) for all \(\omega \in \Omega\) attains this value.

Proofs of the main results are collected in Appendix (ref). Proofs of the remaining main-text results are collected in Online Appendix (ref). The intuition behind the proof is that the least favorable distribution makes both arms as noisy as possible. With outcomes bounded in $[0,1]$, this happens when both potential outcomes are Bernoulli with success probability one half, so $\sigma_1^2=\sigma_0^2=1/4$. Under this distribution, the two arms are equally variable, so the best assignment is the balanced split $\pi=1/2$, and even it yields a main-wave variance of $1$. A rule that lets the pilot tilt the main wave is reacting to noise, protecting one arm only by taking observations from an equally noisy other. This single distribution therefore prevents every rule from having worst-case risk below $1$. Balance, for its part, never exceeds this value, because its risk is $2\sigma_1^2+2\sigma_0^2$ and each variance is capped at $1/4$. The two bounds meet at $1$, so balance is minimax.

This optimality gives balance a decision-theoretic foundation. It is not a default adopted by convention but the symmetric hedge against the worst-case distribution, which can make either arm maximally noisy. The same result, however, reveals the limits of the criterion. Its least favorable distribution is the same whatever the pilot shows, so minimax risk gives the pilot no value and keeps the main wave evenly split even after data strongly suggesting that one arm is noisier. The reason is that minimax risk evaluates total main-wave variance, most of which no assignment can avoid. The worst case is then driven by this unavoidable component rather than by the part the pilot can improve, a well-understood conservativeness of minimax criteria berger1985statistical.

Asymptotic Efficiency and the Feasible Neyman Allocation

The second canonical rule takes the opposite view from minimax risk. Instead of asking for the safest rule when the pilot may mislead, it treats the pilot as a source of estimates for the variance pair that determines the Neyman allocation. If those estimates are accurate, the natural choice is to plug them into the Neyman formula. The feasible Neyman allocation does so regardless, treating the sample variances as if they were known even when the pilot is small and noisy. With \(\hat\sigma_d(\omega)=\sqrt{\hat\sigma_d^2(\omega)}\) the pilot standard deviation in arm \(d\), the rule is $\hat p(\omega) = \frac{\hat\sigma_1(\omega)}{\hat\sigma_1(\omega)+\hat\sigma_0(\omega)}$ whenever the two are not both zero, with \(\hat p(\omega)=1/2\) when they are. It assigns the larger share of the main wave to the arm whose pilot outcomes are more variable, the tilt the Neyman allocation prescribes when the true standard deviations are known. The appeal of the rule is asymptotic. Outcomes are bounded, so the pilot sample variances are consistent, and whenever \(\sigma_1+\sigma_0>0\) the assignment \(\hat p(\omega)\) converges to the Neyman allocation \(\pi^*(\theta)\). Feasible Neyman thus inherits the precision of the infeasible Neyman allocation, which no design can improve on in large samples hahn1998role,hahn2011adaptive,armstrong2022asymptotic.

This large-sample appeal, however, masks a finite-sample fragility that comes from the ratio form of the rule. On the event that both pilot standard deviations are positive, substituting \(\hat p(\omega)\) into \(V\) gives \[ V\!\left(\hat p(\omega),\sigma_1^2,\sigma_0^2\right) = \sigma_1^2 \left( 1+\frac{\hat\sigma_0(\omega)}{\hat\sigma_1(\omega)} \right) + \sigma_0^2 \left( 1+\frac{\hat\sigma_1(\omega)}{\hat\sigma_0(\omega)} \right). \] Each term pairs a true variance with a ratio of pilot standard deviations that scales it. When the pilot estimates are close to the truth and both true standard deviations are positive, the two ratios approach \(\sigma_0/\sigma_1\) and \(\sigma_1/\sigma_0\), and the expression collapses to \(V^*=(\sigma_1+\sigma_0)^2\). When one pilot standard deviation is far too small, the ratio that divides by it grows without bound and turns a true variance of at most \(1/4\) into an arbitrarily large contribution to \(V\). The loss comes from trusting a small estimate, not from any large variance in the population. When a pilot standard deviation is exactly zero, \(\hat p(\omega)\) lands on the boundary, the main wave assigns no observations to one arm, and that arm's outcome mean cannot be estimated.

The finite-sample consequence, established in the next proposition, is an unbounded worst-case risk. The problem is not only that feasible Neyman sometimes makes a noisy adjustment. The problem is that, with a finite pilot, nothing limits how far a noisy estimate can move the assignment, up to and including assigning no main-wave observations to an arm whose true variance is positive.

proposition[{\normalfont\hyperlink{proof:fna_boundary}{Boundary vulnerability of feasible Neyman}}] For every finite pilot size, the feasible Neyman rule has infinite worst-case risk under the boundary convention for \(V\), $\sup_{F\in\mathcal F} \mathcal R(\hat p,F) = \infty$.

The same inverse-probability structure that gives the Neyman allocation its efficiency makes the plug-in rule vulnerable at the boundary.

This failure is not driven by an exotic distribution or by a knife-edge pilot sample. It can arise in the simplest binary-outcome experiment. The next remark makes this point explicit.

remark[Discrete outcomes and zero pilot variance] For a Bernoulli outcome \(Y(d)\sim\mathrm{Bernoulli}(q_d)\) with \(q_d\in(0,1)\), the population variance \(\sigma_d^2=q_d(1-q_d)\) is strictly positive, while the pilot sample variance is zero whenever every observed outcome in that arm is equal, an event with probability \(q_d^{M_d}+(1-q_d)^{M_d}\). Under independent sampling across pilot arms, the probability that exactly one arm shows zero pilot variance is therefore strictly positive for every finite \(M_0,M_1\ge 2\) and every \(q_0,q_1\in(0,1)\). On this event, FNA places the entire main wave in one arm despite both population variances being strictly positive. The difficulty is not confined to exact zeros. The smallest positive sample variance for a Bernoulli arm occurs when a single observation differs from the rest, giving \(\hat\sigma_d^2=1/M_d\) under the unbiased estimator, with probability \(M_d q_d(1-q_d)^{M_d-1}+M_d(1-q_d)q_d^{M_d-1}\). A draw of this kind makes an estimated standard deviation as small as \(M_d^{-1/2}\), which keeps the realized variance finite but, by the mechanism above, can still make it very large.

Remark (ref) suggests an immediate repair. If the problem is that feasible Neyman can assign an arm zero probability, one can force the rule to remain away from zero and one. This is the trimmed feasible Neyman rule, which is discussed in the next remark.

remark[Trimmed feasible Neyman] A natural fix trims feasible Neyman away from the boundary, $\hat p_\tau(\omega)=\min\{\max\{\hat p(\omega),\tau\},1-\tau\}$ with \(\tau\in(0,1/2)\), which caps the realized variance at \(1/(2\tau)\) and removes the infinite-risk pathology. The protection, however, is fixed before any data arrive rather than set by the strength of the pilot evidence, and no single \(\tau\) works well. A large \(\tau\) guards against extreme assignments but stays far from Neyman even after a pilot that has all but resolved the variances, while a small \(\tau\) tracks feasible Neyman but leaves only the weak guarantee \(1/(2\tau)\). The conflict persists even asymptotically. At \((\sigma_1^2,\sigma_0^2)=(1/4,0)\) the Neyman target \(\pi^*(\theta)=1\) exceeds the cap \(1-\tau\), so regret converges to \(\tfrac{\tau}{4(1-\tau)}>0\) rather than to zero, and shrinking \(\tau\) with the pilot removes this gap only by surrendering the finite-sample protection. Trimming is therefore a blunt safeguard against extreme assignments, not a criterion for how far the realized pilot justifies moving away from balance.

The feasible Neyman allocation therefore captures the large-pilot ideal but not the finite-pilot problem. It uses the pilot through point estimates and gives no way to distinguish a reliable variance imbalance from a noisy one.

Minimax Regret and Finite-Pilot Adaptation

Minimax risk is too conservative for pilot adaptation because it ranks rules by total main-wave variance, most of which no design can avoid. The regret of Subsection (ref) instead charges a rule only for the excess variance its assignment creates relative to the infeasible Neyman benchmark, isolating the component a pilot can reduce. Ranking rules by worst-case regret rather than worst-case risk is the standard, less conservative alternative and the more appropriate criterion here savage1951theory.

A statistical decision rule \(p \in \mathcal{D}_0\) is fixed before the pilot, prescribing the main-wave assignment probability \(p(\omega)\) at every realization \(\omega\). Its expected regret at \(F\), $R(p,F) := \mathbb E_{P_F}\!\left[ r\!\left(p(\omega),\theta(F)\right) \right]$ is the frequentist risk of \(p\) under regret loss. The exact minimax-regret criterion selects the rule with the smallest worst-case expected regret. For fixed pilot arm sizes \(M_1,M_0\), with their dependence suppressed elsewhere in the notation, the minimax-regret value is \[ R^{\mathrm{mmr}}_{M_1,M_0} := \inf_{\delta\in\mathcal D_0} \sup_{F\in\mathcal F} \mathbb E_{P_{F,M_1,M_0}}\!\left[ r\!\left(\delta(\omega),\theta(F)\right) \right] = \inf_{\delta\in\mathcal D_0} \sup_{F\in\mathcal F} R(\delta,F). \] The three operations read in order. The experimenter chooses a rule \(\delta\), Nature responds with the least favorable distribution \(F\), and the expectation averages the resulting regret over the pilots that \(F\) generates. The infimum ranges over all measurable pilot-to-assignment rules in \(\mathcal D_0\), not only those that use the pilot through its two sample variances, which makes this the most demanding ex ante minimax-regret criterion available. When the infimum is attained, an exact minimax-regret rule is any \(\delta^{\mathrm{mmr}}_{M_1,M_0}\in \mathcal D_0\) achieving the value.

Proposition (ref) records the basic properties of the minimax-regret value and shows when finite-pilot information must be used by an exact minimax-regret rule.

proposition[{\normalfont\hyperlink{proof:minimax_regret}{Exact minimax regret}}] The minimax-regret value satisfies the following properties. \begin{enumerate} • For every pilot size \((M_1,M_0)\), \(R^{\mathrm{mmr}}_{M_1,M_0} \leq 1/4\). In the no-pilot problem, the minimax-regret value is \(1/4\), and the unique minimax-regret action is balanced assignment. • If \(M_1,M_0 \geq 2\), then \(R^{\mathrm{mmr}}_{M_1,M_0} < 1/4\). Consequently, every exact minimax-regret rule \(\delta^{\mathrm{mmr}}_{M_1,M_0}\) satisfies $\Pr_{P_{F,M_1,M_0}}\!\left(\delta^{\mathrm{mmr}}_{M_1,M_0}(\omega)\neq \tfrac12\right)>0$ for some \(F\in\mathcal F\). • If \(M_d\to\infty\) for each \(d\in\{0,1\}\), then \(R^{\mathrm{mmr}}_{M_1,M_0}\to 0\). \end{enumerate}

Part (i) combines a universal upper bound with a no-pilot characterization. Balanced assignment is feasible at every pilot size and has worst-case regret exactly \(1/4\), attained when one arm has variance \(1/4\) and the other zero. With no pilot, the feasible rules are the constant assignments, and any \(\pi\neq 1/2\) undersamples one arm, which Nature punishes by placing all variance there. Only \(\pi=1/2\) equalizes the two opposing worst cases, all variance in treatment versus all in control.

Part (ii) shows that the slightest pilot overturns the no-pilot optimality of balance. With two observations per arm, the smallest size at which within-arm variation can appear, the minimax-regret value already drops below \(1/4\), so balance is no longer optimal. The reason is that Nature now faces two competing forces. Driving the arm variances far apart raises the regret balance suffers, because balance is optimal only when the two are equal. But that same asymmetry is what the pilot detects, so it also tells the experimenter which arm to favor. A slight tilt toward the arm with the larger pilot variance exploits this signal. Where the variances are far apart, the pilot usually identifies the noisier arm and the tilt reduces regret. Where they are close and balance is nearly optimal, the tilt costs almost nothing. Such a rule therefore has worst-case regret strictly below \(1/4\), so every exact minimax-regret rule must leave \(1/2\) with positive probability under some distribution.

Part (iii) describes the large-pilot limit. As both pilot arms grow, the variance estimates become accurate enough that a stabilized plug-in rule, constructed in the proof, approaches the Neyman allocation and its worst-case expected regret vanishes. Since the minimax-regret value is no larger than the worst-case regret of this rule, \(R^{\mathrm{mmr}}_{M_1,M_0}\to0\).

Despite the desirable properties recorded in Proposition (ref), Remark (ref) shows that the exact minimax-regret problem is not directly operational, because it does not reduce to a finite-dimensional optimization over the variance pair.

remark[Why exact minimax regret is not directly operational] The exact minimax-regret problem is a well-defined decision-theoretic object. The computational difficulty is not merely that the game is infinite-dimensional, but that the inner supremum admits no variance-pair reduction. Nature's choice of $F$ enters the problem twice, through the loss via the variance pair $\theta(F)$ and through the data via $P_{F,M_1,M_0}$. These two channels are not linked by the variance pair, as discussed in Subsection (ref). The inner supremum therefore cannot be taken over $\Theta$, and the problem does not collapse to the static problem $\inf_p\sup_{\theta\in\Theta}r(p,\theta)$ that settles the no-pilot case. The usual shortcuts do not restore finite-dimensional structure. The guess-and-verify approach of stoye2009minimax, which identifies a minimax-regret rule as Bayes against a least favorable prior, does not help here, because the least favorable prior would, for the same reason, have to range over $\mathcal F$ rather than over variance pairs, leaving no finite-dimensional family to guess within. The Bernoulli reduction for bounded mean-payoff problems schlag2006eleven,stoye2009minimax, which replaces $Y\in[0,1]$ by $B\mid Y\sim\mathrm{Bernoulli}(Y)$, preserves the mean but strictly increases the variance unless $Y$ is already binary, so it maps the state $\theta(F)$ to a different variance pair and defines a different variance-based game.\footnote{The marginal variance is $\operatorname{Var}(B)=\mathbb E[Y]\,(1-\mathbb E[Y]) =\operatorname{Var}(Y)+\mathbb E[Y(1-Y)]\ge\operatorname{Var}(Y)$.} Exact minimax regret is best read as a theoretical standard of comparison rather than an operational rule. It becomes a finite optimization only after one restricts the outcome distribution, discretizes the pilot, or restricts the class of rules, and any such computation solves the restricted game rather than the exact nonparametric problem.\footnote{With \(Y(d)\sim\mathrm{Bernoulli}(q_d)\), the state \((q_1,q_0)\in[0,1]^2\) pins down both the loss and the pilot distribution, yet the game remains a semi-infinite minimax problem with no apparent closed form, numerically solvable only for small pilots and case by case, since all assignments, one per pilot realization, must be optimized jointly against a worst case verified globally over a state space on which expected regret is nonconcave.}

Even so, the value \(R^{\mathrm{mmr}}_{M_1,M_0}\) remains informative, since it is the smallest worst-case expected regret attainable by any pilot-based assignment rule. The next proposition shows that even this ideal value cannot vanish faster than the inverse-square-root rate in the pilot arm sizes.

proposition[{\normalfont\hyperlink{proof:mmr-lower-bound}{Minimax regret lower bound}}] Fix pilot arm sizes \(M_1,M_0\ge 2\). There exists a universal constant \(c>0\) such that $R^{\mathrm{mmr}}_{M_1,M_0} \ge c\left(M_1^{-1/2}+M_0^{-1/2}\right)$.

The bound reflects a limit on what a finite pilot can reveal about the variance pair. Even the best possible rule faces variance configurations that generate nearly indistinguishable pilot data but call for different Neyman allocations. Because the pilot cannot reliably tell such configurations apart, any rule must make similar recommendations in states where different assignments would be optimal, and must pay regret in at least one of them. Since the minimax-regret value already optimizes over all measurable pilot-to-assignment rules, this cost applies to every pilot-based design. No rule can therefore improve the worst-case expected regret beyond order \(M_1^{-1/2}+M_0^{-1/2}\). Any tractable rule that attains this order matches the minimax-regret rate. In sum, exact minimax regret may be out of reach computationally, but matching its rate is enough for first-order worst-case optimality.

Conditional Minimax Regret Rule

Section (ref) leaves unanswered how far the main-wave assignment should move from balance toward the feasible Neyman allocation given the evidence in the realized pilot. A rule that answers this question must be computable from the realized pilot and must move from balance only as far as the evidence justifies.

The regret of an assignment depends on the population distribution only through the marginal potential-outcome variances, so once the pilot is observed, the uncertainty that matters for the main-wave loss is which variance pair is true. The CMR procedure summarizes the pilot's design-relevant information by a finite-sample confidence set, the variance pairs the evidence has not ruled out, and selects the assignment with the smallest worst-case regret over that realized set. The worst case thus runs over the configurations consistent with the realized pilot, not over the full model class and not averaged over pilot realizations. In this sense, CMR is the minimax-regret principle applied to the problem the experimenter faces after the pilot, with the surviving variance pairs in the role of the state space.

Alongside the assignment, CMR reports a certificate in the spirit of andrews2025certified, the largest regret the chosen assignment can incur over the configurations the set still allows. Because the set contains the true variance pair with probability at least \(1-\alpha\), the realized regret exceeds the certificate with probability at most \(\alpha\). The certificate is large when the pilot leaves the variances uncertain and small when the pilot has nearly pinned them down.

The Conditional Minimax Regret Procedure

The CMR rule depends on the pilot only through a confidence set for the variance pair. Because each pilot unit is observed in only one arm, treatment observations inform $\sigma_1^2$ and control observations inform $\sigma_0^2$. The set therefore combines two separate inferences, an interval for $\sigma_1^2$ and an interval for $\sigma_0^2$, pairing every treatment variance in the first with every control variance in the second. This product structure makes the confidence set a rectangle.

For each arm \(d\in\{0,1\}\) and one-sided error level \(b\in(0,1)\), let \(\underline{\sigma}_d^2(b;\omega)\) and \(\overline{\sigma}_d^2(b;\omega)\) denote a lower and an upper confidence bound for \(\sigma_d^2(F)\), computed from the pilot and taking values in \([0,1/4]\). They must satisfy $\Pr_{P_F}\!\left( \underline{\sigma}_d^2(b;\omega)\le \sigma_d^2(F) \right) \ge 1-b$ and $\Pr_{P_F}\!\left( \sigma_d^2(F)\le \overline{\sigma}_d^2(b;\omega) \right) \ge 1-b$ for every \(F\in\mathcal F\). Each bound is read as one reads a standard confidence bound, with the pilot as the sample and \(\sigma_d^2(F)\) the unknown parameter. Subsection (ref) constructs such bounds from finite-sample concentration inequalities. The general CMR procedure described below needs only the coverage property.

The rectangle contains the true variance pair \(\theta(F)\) exactly when all four one-sided bounds hold at once. To keep the overall miss probability at most \(\alpha\), the construction divides that allowance equally among the four sides and forms each at one-sided error level \(\alpha/4\). The realized confidence rectangle is

equation[equation omitted — 285 chars of source]

By Bonferroni's inequality bonferroni1936teoria, the probability of a miss is at most the sum of the four failure probabilities.\footnote{Independence of the two pilot arm samples would also permit a multiplicative, Šidák-type split of the error budget across arms vsidak1967rectangular.} Each is at most \(\alpha/4\), so the miss probability is at most \(\alpha\), and $\Pr_{P_F}\!\left( \theta(F)\in\widehat\Theta_\alpha(\omega) \right) \ge 1-\alpha$ for every $F\in\mathcal F$. The guarantee is frequentist in the usual sense, with the random rectangle covering the fixed pair \(\theta(F)\) in at least a fraction \(1-\alpha\) of repeated pilot samples.

Given the realized rectangle, CMR chooses the main-wave treatment assignment probability that minimizes worst-case regret over the variance pairs the pilot has not ruled out:

equation[equation omitted — 150 chars of source]

The inner supremum is the largest regret \(p\) could incur over the surviving variance pairs, and the outer minimization selects the treatment assignment probability that makes this worst case smallest. The key simplification is that, after the pilot is translated into the rectangle, the decision problem no longer ranges over full outcome distributions.

The certificate reported with the assignment is the worst-case regret at the chosen CMR assignment,

equation[equation omitted — 164 chars of source]

Whenever the rectangle contains the true variance pair, the realized regret \(r(p_{\mathrm{CMR}}(\omega),\theta(F))\) is therefore at most \(U_{\mathrm{CMR}}(\omega)\).

Geometry and Closed-Form Solution

The CMR assignment and certificate can be computed in closed form once the realized rectangle is in hand. The dependence on the pilot realization \(\omega\) is left implicit in what follows. By the second form of the regret identity (ref), regret is the square of \((1-\pi)\sigma_1-\pi\sigma_0\), which is affine in the two standard deviations, so for fixed \(\pi\) the worst case over the rectangle is attained at a corner. The relevant corners are the two off-diagonal ones, the corner where treatment is as variable as the rectangle allows and control as stable as it allows, and the corner where the roles are reversed. The proposition below uses this corner structure to solve the assignment problem in closed form.

proposition[{\normalfont\hyperlink{proof:cmr_assignment_rectangle}{Closed-form CMR rule}}] Let \(\underline\sigma_d=\sqrt{\underline\sigma_d^2}\) and \(\overline\sigma_d=\sqrt{\overline\sigma_d^2}\), and suppose \(\overline{\sigma}_1>0\) and \(\overline{\sigma}_0>0\). The CMR assignment problem (ref) has a unique solution, given by \begin{equation} p_{\mathrm{CMR}} = \frac{\overline{\sigma}_1+\sigma_1} {\overline{\sigma}_1+\sigma_1+\overline{\sigma}_0+\sigma_0} \in(0,1). \end{equation} The corresponding certificate is \begin{equation} U_{\mathrm{CMR}} = \frac{\left(\overline{\sigma}_1\overline{\sigma}_0-\sigma_1\sigma_0\right)^2} {\left(\overline{\sigma}_1+\sigma_1\right)\left(\overline{\sigma}_0+\underline{\sigma}_0\right)}. \end{equation}

To interpret the assignment formula (ref), multiply its numerator and denominator by one half, which shows that CMR applies the Neyman formula to the midpoint of each standard-deviation confidence interval. Treatment receives more than half the main wave exactly when its confidence-interval midpoint exceeds control's. For example, if \(\sigma_1\in[0.20,0.40]\) and \(\sigma_0\in[0.10,0.30]\), the midpoints are \(0.30\) and \(0.20\), and the treatment assignment probability prescribed by CMR is \(0.30/(0.30+0.20)=0.60\).

A further feature of the closed form is that its assignment is always interior, whereas feasible Neyman can collapse to the boundary. A zero pilot sample variance is not proof that the population variance is zero, and the finite-sample constructions of the next subsection keep the upper endpoint \(\overline{\sigma}_d\) strictly positive in that case. That alone keeps both arms randomized, without the ad hoc trim feasible Neyman requires.

The certificate has the same corner structure. The two off-diagonal corners pull the assignment in opposite directions, CMR equalizes their regrets, and \(U_{\mathrm{CMR}}\) is their common value. The certificate is therefore large when the rectangle still contains variance pairs that point to substantially different Neyman allocations, and it vanishes exactly when the two corners imply the same allocation, which occurs when \(\overline{\sigma}_1\overline{\sigma}_0=\underline{\sigma}_1\underline{\sigma}_0\). It measures the spread in implied Neyman allocations, not the raw width of the rectangle.

The rule and its certificate reduce to familiar values at the two extremes. When the rectangle contains no information beyond the maintained bounds and each standard-deviation interval is $[0,1/2]$, the two midpoints coincide at $1/4$, the rule returns balance, and the certificate equals $1/4$, the worst-case regret of the no-pilot problem. When the rectangle collapses to a single variance pair with both standard deviations positive, each confidence-interval midpoint equals the corresponding true standard deviation, the rule returns the Neyman allocation $\sigma_1/(\sigma_1+\sigma_0)$, and the certificate is zero.

Constructing the Confidence Rectangle

The geometry of the problem also implies that the CMR assignment and certificate depend on the pilot only through the four rectangle endpoints. The baseline construction of \(\widehat\Theta_\alpha(\omega)\) is distribution-free and finite-sample, using empirical Bernstein bounds for the arm-specific standard deviations and converting them into one-sided variance bounds valid uniformly over the bounded-outcome model. When outcomes are binary, the pilot distribution is discrete and exactly tractable, so the rectangle can be tightened by inverting the folded-binomial distribution of the pilot statistic, as developed in Online Appendix (ref). Either way, the assignment and certificate are computed from the resulting rectangle exactly as in Proposition (ref), and any endpoints satisfying the one-sided coverage requirements of Subsection (ref) deliver the same guarantees.

The relevant concentration inequality is the empirical Bernstein bound of maurer2009empirical for the sample standard deviation of bounded random variables.\footnote{Sharper, first-order optimal empirical Bernstein confidence intervals for the variance of bounded random variables are available martinez2025sharp. The theoretical guarantees extend to these tighter bounds, but the resulting expressions are considerably more involved.} For each arm \(d\) and one-sided error level \(b\in(0,1)\), write \(\eta_d(b)=\sqrt{2\log(1/b)/(M_d-1)}\) for the finite-sample concentration radius on the standard-deviation scale. Then \[ \Pr_{P_F}\!\left(\sigma_d(F)\le\hat\sigma_d(\omega)+\eta_d(b)\right)\ge 1-b \quad\text{and}\quad \Pr_{P_F}\!\left(\sigma_d(F)\ge\hat\sigma_d(\omega)-\eta_d(b)\right)\ge 1-b \] uniformly over \(F\in\mathcal F\). The radius \(\eta_d(b)\) depends only on the pilot size and the error level, shrinking at rate \(1/\sqrt{M_d}\) and growing only as \(\sqrt{\log(1/b)}\) as the error level falls.

The variance bounds follow by squaring these standard-deviation statements and projecting onto the maintained variance interval \([0, 1/4]\). For each arm \(d \in \{0, 1\}\) and error level \(b \in (0, 1)\), the lower and upper variance bounds are

equation[equation omitted — 268 chars of source]

Both endpoints are projected onto \([0,1/4]\). Because the true variance lies in that interval, the projection preserves the one-sided coverage inequalities.\footnote{At a degenerate pilot \(\hat\sigma_d^2(\omega)=0\), \(\overline{\sigma}_d^2(b;\omega)=\min\{1/4,\eta_d(b)^2\}>0\), so the rectangle never treats it as proof of zero variance.}

The four endpoints, each at error level \(\alpha/4\), assemble into the rectangle of (ref), and the union bound delivers its \(1-\alpha\) coverage, as the following lemma shows.

lemma[{\normalfont\hyperlink{proof:mp_rectangle_coverage}{Coverage of the Maurer--Pontil rectangle}}] Fix \(\alpha \in (0, 1)\). For each arm \(d \in \{0, 1\}\) and each one-sided error level \(b \in (0, 1)\), $\Pr_{P_F}\!\left( \underline{\sigma}_d^2(b; \omega) \le \sigma_d^2(F) \right) \ge 1-b$ and $\Pr_{P_F}\!\left( \sigma_d^2(F) \le \overline{\sigma}_d^2(b; \omega) \right) \ge 1-b$ for every \(F \in \mathcal F\). Consequently, the rectangle $\widehat\Theta_\alpha(\omega)$ satisfies $\Pr_{P_F}\!\left( \theta(F) \in \widehat\Theta_\alpha(\omega) \right) \ge 1-\alpha$ for every \(F \in \mathcal F\).

Statistical Properties

This subsection establishes three guarantees for the CMR assignment. First, in finite samples, the reported certificate bounds realized regret with probability at least \(1-\alpha\). Second, the caution built into the rule disappears as the pilot grows. At interior variance pairs, CMR converges to the Neyman allocation, with assignment error shrinking at the inverse-square-root rate and realized regret and the certificate at the inverse-pilot rate. Third, we derive an upper bound on CMR's worst-case expected regret that matches the minimax-regret lower bound up to constants.

\paragraph{Finite-sample validity.}

The link from \(U_{\mathrm{CMR}}\) to a finite-sample guarantee is coverage, the event that the realized rectangle contains the true variance pair. On that event, the true pair is among the configurations over which the certificate takes the worst case, so the realized regret of the CMR assignment cannot exceed \(U_{\mathrm{CMR}}\). Lemma (ref) shows that coverage holds with probability at least \(1-\alpha\). The next theorem combines these two observations and adds that the certificate never exceeds \(1/4\).

theorem[{\normalfont\hyperlink{proof:cmr-certified-optimality}{Finite-sample regret certificate}}] Fix \(\alpha\in(0,1)\). Then the following statements hold. \begin{enumerate} • For every \(F\in\mathcal F\), $\Pr_{P_F}\! \Big( r\!\left(p_{\mathrm{CMR}}(\omega),\theta(F)\right) \le U_{\mathrm{CMR}}(\omega) \Big) \ge 1-\alpha$. • For every pilot realization \(\omega\), $U_{\mathrm{CMR}}(\omega)\le \frac14$. Moreover, equality holds if and only if $(1/4,0)\in\widehat\Theta_\alpha(\omega)$ and $(0,1/4)\in\widehat\Theta_\alpha(\omega)$. \end{enumerate}

Part (i) shows that, for every distribution, the certificate \(U_{\mathrm{CMR}}(\omega)\) bounds the regret of \(p_{\mathrm{CMR}}(\omega)\) with probability at least \(1-\alpha\). The guarantee is finite-sample, holding at the realized pilot sizes rather than only in a large-pilot approximation.\footnote{Part (i) also implies quantile domination: for every \(F\) and every \(t\le 1-\alpha\), the \(t\)-quantile of realized regret is bounded by the \((t+\alpha)\)-quantile of the certificate.} The bound is not only valid but optimal in the sense that \(U_{\mathrm{CMR}}(\omega)\) is the smallest worst-case regret bound that any assignment can attain over the variance pairs still plausible after the pilot. In particular, for every alternative assignment \(\widetilde p(\omega)\in(0,1)\), $U_{\mathrm{CMR}}(\omega) \le \sup_{\theta\in\widehat\Theta_\alpha(\omega)} r\!\left(\widetilde p(\omega),\theta\right)$. This makes CMR a certified decision in the framework of andrews2025certified. In their terms, \(U_{\mathrm{CMR}}(\omega)\) is a \(P\)-certificate for the loss \(r(p_{\mathrm{CMR}}(\omega),\theta(F))\).

Part (ii) bounds the certificate itself. Balance is always among the candidates and its regret is at most \(1/4\) at every configuration, so the certificate never exceeds \(1/4\), the guarantee available with no pilot.\footnote{The same rectangle can also certify a move against balance directly, since an assignment whose variance is smaller than balance's at every pair in the rectangle is no worse than balance with probability at least \(1-\alpha\). Restricting CMR to such assignments yields a conservative variant with this stronger guarantee.} Equality holds exactly when the rectangle still contains the two adversarial corners \((1/4,0)\) and \((0,1/4)\), one placing all variability in treatment and the other all in control. Because these are opposite vertices of \([0,1/4]^2\), only the full variance space contains both, so any pilot that rules out anything at all earns a certificate strictly below \(1/4\).

\paragraph{CMR convergence rates.}

Now we ask how that guarantee evolves as the pilot becomes informative. In the interior of the variance space, where both arms have positive variance, the confidence rectangle collapses around the truth. The next theorem shows that, as both pilot arms grow large, the CMR rule converges to the Neyman allocation, and both realized regret and the certificate shrink to zero.

theorem[{\normalfont\hyperlink{proof:cmr-neyman-recovery}{Neyman convergence and interior rates}}] Fix \(F\in\mathcal F\), and suppose \(\sigma_1(F)>0\) and \(\sigma_0(F)>0\). Suppose also that \(\alpha\in(0,1)\) is fixed and that \(M_0,M_1\to\infty\), with \(M_0/M_1\) bounded away from zero and infinity. Then the CMR rule satisfies the following statements. \begin{enumerate} • $p_{\mathrm{CMR}} \overset{p}{\longrightarrow} \pi^*(\theta(F))$, $|p_{\mathrm{CMR}}-\pi^*(\theta(F))| = O_p\!\left(M_1^{-1/2}+M_0^{-1/2}\right)$. • $r\!\left(p_{\mathrm{CMR}},\theta(F)\right) = O_p\!\left(M_1^{-1}+M_0^{-1}\right)$. • $U_{\mathrm{CMR}} = O_p\!\left(M_1^{-1}+M_0^{-1}\right)$. \end{enumerate}

All three rates come from a single source, the contraction of the confidence rectangle around the truth. The Maurer--Pontil endpoints learn each arm's standard deviation at the inverse-square-root rate in the pilot arm sizes, shrinking the rectangle to the variance pair at rate \(M_1^{-1/2}+M_0^{-1/2}\). At an interior truth the rule varies smoothly with the rectangle and inherits this rate, which is statement (i). The other two rates are faster, and the reason is that, by the regret identity (ref), regret is quadratic in the gap between the assignment and the Neyman allocation, with a coefficient that is finite whenever both arms have positive variance. An inverse-square-root error in the assignment then enters regret squared, at rate \(M_1^{-1}+M_0^{-1}\), which is statement (ii). The certificate obeys the same bound through its closed form, which squares a product gap of inverse-square-root order over a denominator bounded away from zero in the interior, giving statement (iii).

The key implication of the theorem is that, at an interior truth, the finite-sample safety of the CMR rule does not slow the rate at which it converges to the Neyman allocation. In the corresponding large-main-wave limit, this convergence implies that the resulting assignment approaches the optimized Hahn-bound allocation, which armstrong2022asymptotic identifies as the first-order efficiency frontier across all designs that may adapt to covariates and past outcomes. The CMR rule approaches this frontier while certifying its own regret at every pilot size with probability at least \(1-\alpha\).\footnote{The feasible Neyman allocation approaches this frontier at the same inverse-square-root rate as the CMR rule, because the pilot estimates the two standard deviations at that rate and the Neyman formula responds smoothly to small errors in those inputs at an interior truth.} First-order asymptotic efficiency and finite-sample safety hold together, with no trade-off between them.

The interior rates assume both arms have positive variance. The following remark shows that when one arm is degenerate, realized regret and the certificate converge at the slower rate \(M_d^{-1/2}\) rather than \(M_d^{-1}\).

remark[Boundary rates] The interior rates rely on the Neyman allocation lying strictly between zero and one. There the denominator \(\pi(1-\pi)\) in the regret identity (ref) is bounded away from zero, so regret scales with the squared assignment error, and an error of order \(M_d^{-1/2}\) produces regret of order \(M_d^{-1}\). A degenerate arm moves the Neyman allocation to a corner and breaks this scaling, slowing realized regret and the certificate to the assignment-error rate itself. With \(\sigma_1(F)>0\) and \(\sigma_0(F)=0\), the Neyman allocation is \(\pi^*=1\). The upper confidence endpoint for \(\sigma_0\) does not collapse to zero, so CMR keeps a small control share \(1-p_{\mathrm{CMR}}\), which contracts at rate \(M_0^{-1/2}\) under the Maurer--Pontil construction. At this truth the regret identity reduces to $r(\pi,\sigma_1^2(F),0)=\tfrac{\sigma_1^2(F)\,(1-\pi)}{\pi}$, linear in the leftover control share rather than quadratic. The hedge that keeps the rule off the corner therefore costs order \(M_0^{-1/2}\) in realized regret, and the certificate is of the same order. The case \(\sigma_1(F)=0\) and \(\sigma_0(F)>0\) is symmetric.

\paragraph{Matching the Minimax-Regret Rate.}

The convergence rates just established are pointwise in $F$, describing how the rule and its realized regret behave at a fixed population as the pilot grows. The minimax-regret criterion of Subsection (ref) instead judges a rule by its worst-case expected regret across all populations. As the next theorem shows, the CMR rule's worst-case expected regret shrinks at the inverse-square-root rate in the pilot arm sizes, and at the inverse-pilot rate once both arms are bounded away from degeneracy.

theorem[{\normalfont\hyperlink{proof:cmr-competitive-risk}{Uniform expected regret}}] Fix \(\alpha\in(0,1)\). There exists a constant \(C_\alpha<\infty\), depending only on \(\alpha\), such that, for all pilot arm sizes \(M_1,M_0\ge2\), \[ \sup_{F\in\mathcal F} \mathbb E_{P_F}\!\left[ r\!\left(p_{\mathrm{CMR}}(\omega),\theta(F)\right) \right] \le C_\alpha \left(M_1^{-1/2}+M_0^{-1/2}\right). \] Moreover, for every \(\kappa>0\), there exists a constant \(C_{\alpha,\kappa}<\infty\), depending only on \(\alpha\) and \(\kappa\), such that \[ \sup_{\substack{F\in\mathcal F:\\ \sigma_1(F)\wedge\sigma_0(F)\ge \kappa}} \mathbb E_{P_F}\!\left[ r\!\left(p_{\mathrm{CMR}}(\omega),\theta(F)\right) \right] \le C_{\alpha,\kappa} \left(M_1^{-1}+M_0^{-1}\right). \]

Theorem (ref) evaluates CMR by the ex ante criterion of Section (ref), worst-case expected regret. Even though the assignment adapts to the realized pilot, its worst-case expected regret vanishes at the inverse-square-root rate in the pilot arm sizes. On any class of populations with both standard deviations bounded away from zero, the rate improves to the inverse-pilot rate. The slower uniform rate has the same source as the boundary rate of Remark (ref).

The key reason why the upper bound vanishes is that the assignment and its regret do not blow up when the rectangle \(\widehat\Theta_\alpha(\omega)\) misses the truth \(\theta(F)\). Both depend continuously on how far the pilot's variance estimates fall from the truth, not on whether the rectangle contains it, so a small estimation error keeps regret small, covered or not. The bound is therefore governed by the size of this estimation error, which shrinks at the inverse-square-root rate.

Combining Theorem (ref) with Proposition (ref) yields $\sup_{F\in\mathcal F} \mathbb E_{P_F}\!\left[ r\!\left(p_{\mathrm{CMR}}(\omega),\theta(F)\right) \right] \le \frac{C_\alpha}{c} R^{\mathrm{mmr}}_{M_1,M_0}$. The theorem gives the CMR upper bound, of order \(M_1^{-1/2}+M_0^{-1/2}\), and the proposition the matching lower bound for the exact minimax-regret value. CMR's worst-case expected regret is therefore at most a constant, depending only on \(\alpha\), times \(R^{\mathrm{mmr}}_{M_1,M_0}\), the smallest any pilot-based rule can attain. Strikingly, CMR achieves this although it optimizes only against the realized rectangle and never solves the ex ante minimax problem. Unlike an exact minimax-regret rule, CMR is closed form and reports a finite-sample certificate for the assignment it chooses.

Extensions: Multiple Treatments and Stratification

Multiple Treatments with a Shared Control

\paragraph{Setup.} Many experiments compare several treatments against a common control, testing versions of a program, prices, information treatments, or implementation modes. The design question is how to divide the main wave across all the treatments and the shared control at once. The control arm is indexed by \(k=0\) and the treatments by \(k=1,\ldots,K\). A unit's potential outcomes \((Y(0),\ldots,Y(K))\), with \(Y(k)\in[0,1]\), are drawn from a population distribution \(F\). The mean and variance of \(Y(k)\) are \(\mu_k\) and \(\sigma_k^2\), and \(\theta(F)=(\sigma_0^2, \sigma_1^2,\ldots,\sigma_K^2)\in[0,1/4]^{K+1}\) collects the arm variances. The main wave is set by the allocation vector \(\pi=(\pi_0,\pi_1,\ldots,\pi_K)\), where \(\pi_k>0\) is the fraction allocated to arm \(k\) and \(\sum_{k=0}^K\pi_k=1\). The estimands of interest are the \(K\) treatment-control contrasts \(\text{ATE}_k = \mu_k-\mu_0\).

For a main wave of size \(N\) with \(N_k\) units on arm \(k\), the difference-in-means estimator of contrast \(k\) has variance \(\sigma_k^2/N_k+\sigma_0^2/N_0\). With \(N_k=N\pi_k\), the sum of the marginal variances of the \(K\) contrasts equals \(V(\pi,\theta)/N\), where $V(\pi,\theta) = \frac{K\sigma_0^2}{\pi_0} + \sum_{k=1}^K \frac{\sigma_k^2}{\pi_k}$. Every term is an arm variance divided by the share on that arm, so the allocation problem retains the inverse-share structure of the binary case, with the control variance now counted \(K\) times. A pilot of fixed size must now be spread over \(K+1\) arms rather than two, so each arm variance is estimated with fewer observations.

\paragraph{Neyman Allocation.}

For known positive variances, minimizing \(V(\pi,\theta)\) over the simplex \(\Delta_K=\{\pi\in\mathbb R_+^{K+1}:\sum_{k=0}^K\pi_k=1\}\) yields the multi-arm Neyman allocation:

equation[equation omitted — 272 chars of source]

The benchmark value is \(V^*(\theta)=\inf_{\pi}V(\pi,\theta)=\left(\sqrt K\,\sigma_0+\sum_{k=1}^K\sigma_k\right)^2\), attained at the Neyman allocation (ref) when all arm variances are strictly positive, and regret at a feasible \(\pi\) is \(r(\pi,\theta)=V(\pi,\theta)-V^*(\theta)\).

\paragraph{CMR.}

The CMR construction extends by replacing the scalar assignment probability with the allocation vector. Each arm's variance receives a one-sided empirical Bernstein interval as in Section (ref), with the error budget now split equally across the \(2(K+1)\) one-sided endpoints, and stacking the intervals gives the hyperrectangle \(\widehat\Theta_\alpha(\omega)=\prod_{k=0}^{K}\bigl[\underline{\sigma}_k^2(\omega),\overline{\sigma}_k^2(\omega)\bigr]\subseteq[0,1/4]^{K+1}\), which covers the true variance vector with probability at least \(1-\alpha\) for every \(F\). The CMR allocation \(p_{\mathrm{CMR}}(\omega)\in\Delta_K\) minimizes worst-case regret over this hyperrectangle, and the reported certificate \(U_{\mathrm{CMR}}(\omega)\) is again the worst-case regret at the chosen allocation, carrying the same \(1-\alpha\) finite-sample guarantee as in the binary case.\footnote{Online Appendix (ref) collects the properties used across the extensions of Section (ref). Regret is convex in the variance vector, so worst-case regret over the hyperrectangle is attained at a vertex (Lemmas (ref) and (ref)), the CMR problem is a finite convex program (Lemma (ref)), and the reported certificate bounds realized regret with probability at least \(1-\alpha\) (Lemma (ref)) and is monotone under set inclusion (Lemma (ref)).} The next proposition traces the CMR allocation from the safe default, when the hyperrectangle is the full \([0,1/4]^{K+1}\), to the Neyman allocation, when it collapses to a point as the pilot grows.

proposition[{\normalfont\hyperlink{proof:multi_arm_cmr}{Multi-arm CMR between shared-control balance and Neyman}}] Consider the multi-arm shared-control design. \begin{enumerate} • If the pilot is uninformative, with \(\widehat\Theta_\alpha(\omega)=[0,1/4]^{K+1}\), the CMR allocation is \[ p_{0,\mathrm{CMR}} = \frac{1}{1+\sqrt K}, \qquad p_{k,\mathrm{CMR}} = \frac{1}{\sqrt K\,(1+\sqrt K)}, \quad k=1,\ldots,K . \] • Fix \(\sigma_k(F)>0\) for all \(k\). With \(\alpha\) fixed and \(M_{\min}=\min_{0\le k\le K} M_k\to\infty\), every measurable selection $p_{\mathrm{CMR}}(\omega)$ from the CMR solution set satisfies \( p_{\mathrm{CMR}} \overset{p}{\to}\pi^*(\theta(F))\), the multi-arm Neyman allocation (ref). \end{enumerate}

Part (i) is the safe default. Unable to treat any arm as noisier than another, CMR plays the equal-variance Neyman allocation, placing \(1/(1+\sqrt K)\) of the main wave on the control and \(1/(\sqrt K(1+\sqrt K))\) on each treatment. For example, with \(K=4\), the control receives a third of the main-wave sample and each treatment a sixth. Part (ii) is the multi-arm analogue of the interior convergence in Theorem (ref), driven by the same contraction of the confidence set around the truth. The worst case also changes shape relative to the binary case. The binding configurations are vertices of the hyperrectangle, each making some subset of arms as noisy as the pilot allows and the rest as stable as it allows, and CMR hedges against the worst of these patterns rather than against one arm's variance at a time. Away from the uninformative rectangle, neither this allocation nor its stratified counterpart has a closed form. Online Appendix (ref) shows how both are computed as finite convex programs over the vertices of the realized hyperrectangle.

Stratified Experiments

\paragraph{Setup.}

Many target populations divide into strata, such as schools, regions, or gender, known before the main wave is designed. A stratified design involves two decisions, how many observations each stratum receives and how each stratum's observations are split between treatment and control. The pilot informs both decisions.

Strata are indexed by \(x=1,\ldots,S\), a unit's pre-specified stratum is \(X\in\{1,\ldots,S\}\), and the target-population shares \(s_x=\Pr_F(X=x)>0\) sum to one. These shares are known and fixed before the main-wave design is chosen. Within stratum \(x\), the potential outcome under arm \(d\in\{0,1\}\) has conditional mean \(\mu_{dx}=\mathbb E_F[Y(d)\mid X=x]\) and variance \(\sigma_{dx}^2=\operatorname{Var}_F(Y(d)\mid X=x)\). As before, the estimand is the population average treatment effect, \(\operatorname{ATE}(F)=\sum_{x=1}^S s_x(\mu_{1x}-\mu_{0x})\), the average of the within-stratum effects \(\mu_{1x}-\mu_{0x}\) weighted by the population shares.

The two decisions combine into a single design object, the treatment-by-stratum cell shares \(\pi=\{\pi_{dx}:d\in\{0,1\},\ x=1,\ldots,S\}\), where \(\pi_{dx}\) is the fraction of the main wave drawn from stratum \(x\) and assigned to arm \(d\). These shares are nonnegative fractions of one fixed main wave, so they sum to one and \(\pi\) ranges over the simplex \(\Delta_{2S-1}=\{\pi:\pi_{dx}\ge 0,\ \sum_{x=1}^S(\pi_{1x}+\pi_{0x})=1\}\). Two summaries of $\pi$ recover the two decisions. The total share collected from stratum $x$, $\pi_{\cdot x} = \pi_{1x} + \pi_{0x}$, is the sampling margin, and the within-stratum treatment probability, $\pi_{1\mid x} = \pi_{1x}/(\pi_{1x} + \pi_{0x})$, is the assignment margin.\footnote{The assignment margin is defined where $\pi_{\cdot x} > 0$. As in the baseline model, boundary allocations are allowed as formal actions, but any allocation that leaves a treatment-by-stratum mean entering the estimand unobserved carries infinite loss by convention.}

The stratified difference-in-means estimator \(\widehat{\operatorname{ATE}}=\sum_{x=1}^S s_x(\bar Y_{1x}-\bar Y_{0x})\) estimates each within-stratum effect by the cell contrast \(\bar Y_{1x}-\bar Y_{0x}\) and aggregates these contrasts using the target-population shares \(s_x\), where \(\bar Y_{dx}\) is the sample mean in cell \((d,x)\) imbens2015causal. For positive cell shares, independent sampling across cells gives an estimator variance of \(V(\pi,\theta)/N\), where

equation[equation omitted — 283 chars of source]

The square \(s_x^2\) appears because the estimator multiplies each stratum contrast by \(s_x\).

\paragraph{Neyman Allocation.}

If the cell variances were known, the optimal stratified design would minimize \(V(\pi,\theta)\) over the cell-share simplex $\Delta_{2S-1}$. For positive variance configurations, the solution is

equation[equation omitted — 196 chars of source]

The optimal design assigns observations to cells in proportion to \(s_x\sigma_{dx}\), how much the cell matters for the target ATE times how noisy its mean is to estimate. A large stratum with a nearly deterministic outcome needs few observations, and a noisy cell receives little weight if it represents a small part of the population.

Summing (ref) over arms gives the Neyman sampling margin, \(\pi_{\cdot x}^*(\theta)=\tfrac{s_x(\sigma_{1x}+\sigma_{0x})}{\sum_{z=1}^S s_z(\sigma_{1z}+\sigma_{0z})}\), which samples a stratum more than proportionally to its population share exactly when its total standard deviation \(\sigma_{1x}+\sigma_{0x}\) exceeds the population-weighted average of total standard deviations across strata. The Neyman assignment margin is the standard two-arm Neyman allocation, \(\pi^*_{1\mid x}=\sigma_{1x}/(\sigma_{1x}+\sigma_{0x})\). Substituting (ref) into (ref) gives $V^*(\theta) =\left[ \sum_{x=1}^S s_x(\sigma_{1x}+\sigma_{0x})\right]^2$, the smallest variance any feasible design attains. Regret takes the same form as in the baseline model, $r(\pi,\theta) = V(\pi,\theta)-V^*(\theta)$.

\paragraph{CMR.}

From the \(M_{dx}\ge 2\) pilot observations in cell \((d,x)\), a one-sided empirical Bernstein interval of maurer2009empirical is built for each cell variance \(\sigma_{dx}^2\), with the error split equally across the \(4S\) one-sided endpoints at level \(\alpha/(4S)\). Stacking the \(2S\) cell intervals gives the hyperrectangle $ \widehat\Theta_\alpha(\omega) = \prod_{d\in\{0,1\}} \prod_{x=1}^S [ \underline{\sigma}_{dx}^2(\omega), \overline{\sigma}_{dx}^2(\omega) ] \subseteq [0,1/4]^{2S}$, which covers the true variance vector with probability at least \(1-\alpha\) by the union bound. The CMR rule and its certificate take the same form as before, now over the cell-share simplex \(\Delta_{2S-1}\), with \(p_{\mathrm{CMR}}(\omega)\) collecting the cell shares \(p_{dx,\mathrm{CMR}}(\omega)\).

proposition[{\normalfont\hyperlink{proof:stratified_cmr}{Stratified CMR between representative balance and Neyman}}] Consider the stratified design. \begin{enumerate} • If the pilot is uninformative, with \(\widehat\Theta_\alpha(\omega)=[0,1/4]^{2S}\), the CMR allocation is $ p_{1x,\mathrm{CMR}}(\omega) = p_{0x,\mathrm{CMR}}(\omega) = \frac{s_x}{2}$, $x=1,\ldots,S$. Equivalently, \(p_{\cdot x,\mathrm{CMR}}(\omega)=s_x\) and \(p_{1|x,\mathrm{CMR}}(\omega)=1/2\) for every stratum \(x\). • Suppose \(\sigma_{dx}(F)>0\) for all \(d\in\{0,1\}\) and \(x=1,\ldots,S\). With \(\alpha\) fixed and \(M_{\min}=\min_{d,x}M_{dx}\to\infty\), every measurable selection \(p_{\mathrm{CMR}}(\omega)\) from the CMR solution set satisfies $p_{\mathrm{CMR}}(\omega) \overset{p}{\longrightarrow} \pi^*(\theta(F))$, the stratified Neyman allocation in (ref). \end{enumerate}

With an uninformative pilot the rule has no reason to treat any stratum or arm as noisier than another, so the allocation is representative sampling with an even split within each stratum. As the pilot shrinks the confidence set, a stratum revealed to be noisier receives more than its population share, the split within each stratum moves toward its noisier arm, and in the limit the allocation becomes the stratified Neyman allocation (ref).

Calibrated Simulations

This section stress-tests the allocation rules at practice-relevant pilot sizes, between 30 and 500 total observations. We calibrate the data-generating processes to public microdata from four published field experiments: the deworming program of miguel2004worms (MK), the resume audit of bertrand2004emily (BM), the experiment on incentives to learn HIV results of thornton2008demand, and the reference-letter experiment of abel2020value.\footnote{We select published experiments in leading economics journals with public microdata, randomized assignment arms that map directly into our design problems, and headline outcomes from the original papers. Across studies, the retained outcomes span binary, count, and bounded continuous measurements.} The two-arm designs are built from MK, BM, and Thornton, the shared-control multi-arm design from Thornton's randomized incentive amounts, and the stratified design from the gender stratification in Abel et al.

Simulation Design

Each data-generating process fixes the outcome distributions that the original experiment induced in its randomized arms. For binary outcomes, treatment and control outcomes are Bernoulli draws with the arm means estimated from the microdata. For continuous and count outcomes, outcomes are drawn from the empirical distribution of the corresponding arm, rescaled to the unit interval using the observed range. The multi-arm design keeps Thornton's incentive levels and their shared control as five separate arms, and the stratified design applies the same construction within gender strata, holding the stratum shares fixed at their sample values.

Table (ref) reports the calibrated arm means, variances, and implied Neyman allocations for each design. In the two-arm designs, the infeasible Neyman allocation assigns between \(46\) and \(55\) percent of the sample to treatment, so balance is never far from optimal. Balance's excess variance over the infeasible allocation is at most \(0.84\) percent, which also caps the improvement any pilot-based rule can deliver, the ceiling of Remark (ref) materializing in calibrated data.

Each simulation replication draws a pilot of total size \(M\) from the calibrated distribution. The pilot itself is assigned without any information, evenly across arms in the two-arm and multi-arm designs, and representatively across strata with an even treatment-control split within each stratum in the stratified design. To make the designs comparable, our performance metric is the percentage efficiency loss, \(100\times[V(p,\theta)-V^*(\theta)]/V^*(\theta)\), where \(p\) is the main-wave assignment the rule chooses from the realized pilot. An entry of \(1\) means the assignment raises the estimator's variance by one percent, or equivalently that the experimenter would need one percent more main-wave observations to reach the same precision. Each entry averages \(500\) independent pilot draws, and the tables report \(M\in\{30,100,250,500\}\).

Two-Arm Results

Table (ref) compares the mean efficiency losses of balance, FNA, and CMR across the seven two-arm designs. Balance has a single entry per design because it does not use the pilot. Figure (ref) plots the comparison over the full grid of pilot sizes.

table[table omitted — 2,154 chars of source]

The FNA rows quantify the downside of trusting a small pilot. At \(M=30\), the efficiency loss is infinite in five of the seven designs and equals \(3.33\) and \(2.19\) percent in the other two, several times the \(0.84\) percent ceiling on what adaptation can gain in these designs. The infinite entries are the boundary failure of Proposition (ref) materializing in calibrated data. In the callback design, where both callback rates are below ten percent, the failure persists through \(M=100\), and even where every draw is finite convergence is slow, with the condom-count design still losing \(5.93\) percent at \(M=100\).\footnote{Common repairs of FNA do not resolve this. Table (ref) evaluates FNA with its assignment trimmed away from the boundary, together with the pre-test and regularized rules studied by cai2024performance. Trimming removes the infinite entries, but all four alternatives still lose substantially at \(M=30\).}

At \(M=30\), CMR's efficiency loss coincides with balance's in every design because the confidence rectangle still contains every variance pair and the rule returns balance. The same discipline appears at \(M=100\) in the callback design. There the confidence rectangle excludes at least one corner of the variance space in only \(3.2\) percent of pilot draws (Table (ref)), so the set remains nearly uninformative and the mean efficiency loss stays at the balance value of \(0.84\) percent.

Once the rectangle becomes informative, CMR's gains arrive in the designs with unequal variances. The efficiency loss falls from \(0.45\) to \(0.07\) percent in the MK infection design between \(M=30\) and \(M=100\), from \(0.59\) to \(0.18\) percent in the Thornton HIV-result design, and from \(0.84\) to \(0.47\) percent in the callback design by \(M=250\). In the designs with nearly equal variances there is nothing to find, so any movement the rectangle licenses chases sampling noise and CMR can lose to balance. In the condom-count design it loses between \(0.22\) and \(0.27\) percent against \(0.18\) for balance, a premium of at most \(0.09\) percentage points, in the same design in which FNA loses \(5.93\) percent at \(M=100\). By \(M=500\), the pilot variance estimates are accurate enough that FNA is modestly ahead of CMR in several designs.\footnote{Table (ref) reports the distribution of the efficiency loss across pilot draws, including medians, standard deviations, maxima, and FNA boundary frequencies. The worst CMR efficiency loss across every design and pilot size is \(5.95\) percent, against FNA boundary failures and finite losses as large as \(57.56\) percent. Once adaptation begins, CMR's loss distribution is right-skewed in some designs, because some draws still move the assignment in the wrong direction, though the rule's conservatism keeps these movements small.}

Table (ref) shows that, in the reported Monte Carlo draws, the confidence rectangle covers the true variance pair for every design and pilot size and the certificate is at least as large as the realized regret in every draw. Coverage exceeds its nominal \(1-\alpha\) level by a wide margin, reflecting the conservativeness of the distribution-free Maurer--Pontil bounds. The median certificate falls from its uninformative value of one quarter at \(M=30\) to between \(0.05\) and \(0.14\) at \(M=500\). Relative to the no-pilot guarantee of one quarter, a pilot of \(500\) observations thus cuts the certified worst case by between roughly one half and four fifths.

Table (ref) translates the losses into the design quantities applied experimenters plan with, additional subjects and statistical power. In the worst cases across the seven designs at \(M=30\), FNA requires \(64\) additional main-wave subjects per \(1{,}000\) to match the precision of the infeasible allocation, and its power against an effect that the infeasible allocation detects with \(80\) percent power falls to \(50\) percent. Under CMR, the extra subjects never exceed \(8.4\) per \(1{,}000\), and power stays within \(0.3\) percentage points of the target in every design and at every pilot size. The cost of overreacting to a small pilot is measured in dozens of subjects or many percentage points of power. The cost of CMR's caution is measured in a handful of subjects.

Multi-Arm and Stratified Results

The extension designs raise both the value and the risk of adaptation, since the allocation is now a vector rather than a single treatment share. Table (ref) reports the results for the two extensions of Section (ref). In the table, the balance row is the no-information allocation of each design rather than an equal split across cells. In Panel A it is the shared-control allocation of Proposition (ref), one third of the sample to control and one sixth to each treatment. In Panel B it is representative balance from Proposition (ref), each gender stratum at its population share, split evenly between treatment and control. Figures (ref) and (ref) plot both panels over the full grid of pilot sizes.

table[table omitted — 1,809 chars of source]

The first difference from the two-arm designs is that the infeasible benchmark adapts along more margins. In Panel A, it exploits variance differences across five arms. In Panel B, it both samples noisier strata more heavily and tilts the treatment split within each stratum. Balance accordingly loses \(1.76\) percent in Panel A and \(5.59\) percent in Panel B, well above the \(0.84\) percent maximum gain available in the two-arm designs.

The second difference is that the same pilot now feeds more variance estimates. At \(M=30\), each of the five Thornton arms contributes six pilot observations, and each treatment-by-stratum cell in the Abel design contributes fewer than ten. Consequently, FNA's mean efficiency loss is infinite in Panel A through \(M=250\), since a single zero-variance binary arm is enough to push it to the boundary. Panel B shows that the boundary event is not the whole problem. There FNA's losses are finite but severe, \(72.13\) percent at \(M=30\) and \(26.01\) percent at \(M=100\), because interior allocations built on noisy cell variances are also far from optimal.

The CMR efficiency loss equals the balance loss at \(M=30\) and \(M=100\) in both panels. The reason is that each cell's variance bound now rests on fewer observations, and the rectangle splits its error probability across more cell-level bounds, so the pilot size at which it first becomes informative grows with the number of cells. Once it does, the rule adapts, to \(1.48\) percent at \(M=250\) and \(0.34\) at \(M=500\) in Panel A, and to \(5.38\) and then \(2.94\) in Panel B.

Panel A includes a second CMR row because the outcome is binary, and it shows that the rule's caution resides in the confidence set rather than in the minimax-regret logic. CMR with the exact-inversion Bernoulli variance sets of Online Appendix (ref) becomes informative at much smaller pilots. It pays a small early premium, \(1.82\) percent against \(1.76\) at \(M=30\), and then adapts earlier than the bounded-outcome rectangle, reaching \(1.21\) percent at \(M=100\) and \(0.75\) at \(M=250\), where the Maurer--Pontil version still sits at \(1.76\) and \(1.48\).

The extensions sharpen the paper's main message.\footnote{Table (ref) translates the extension efficiency losses into subjects and power. At \(M=30\), FNA's joint-test power in Panel A is \(20.7\) percent against \(79.1\) for balance and CMR, and in Panel B FNA requires \(590\) additional main-wave subjects per \(1{,}000\) to match the infeasible benchmark, against \(56\) for balance and CMR. Balance itself is expensive here, and by \(M=500\) CMR cuts the extra subjects to \(27\) and raises power from \(77.8\) to \(78.9\) percent, while FNA's efficiency loss in Panel A falls to \(0.41\) percent.} Exactly where adaptation has the most to buy and plug-in rules are most dangerous, CMR captures a growing share of the attainable gain as the pilot informs, never pays more than a tenth of a percentage point for its caution, and is the only adaptive rule in these tables that arrives with a finite-sample guarantee for the assignment it selects.

Conclusions

The central question in pilot-based design is how much authority noisy preliminary evidence should have over the main experiment. Balanced assignment gives the pilot none, and feasible Neyman gives its point estimates full authority. Rather than acting on point estimates or forgoing adaptation altogether, the Conditional Minimax Regret rule acts on what the pilot has ruled out, staying at balance when the evidence is weak and approaching the Neyman allocation as the evidence accumulates. Before committing the main wave, the experimenter holds a certificate that, with high probability, bounds the precision lost to the realized design and never exceeds balance's no-pilot guarantee.

The same CMR logic applies to any design choice whose loss depends on parameters a small pilot estimates imprecisely. Two directions for future research seem most valuable. The first is to develop such rules for richer designs, including cluster-randomized, sequential, and imperfect-compliance experiments. The second is to treat the confidence set itself as part of the decision. CMR splits its error budget symmetrically and takes the coverage level as given, and neither choice is necessarily optimal. Sharper sets would let the rule earn the right to adapt sooner.

Much of the theory of adaptive experimentation justifies pilot-based designs by letting the pilot grow large. Real pilots are small, and at the sizes experimenters actually run, adaptation has modest upside and unbounded downside. The practical lesson of this paper is that the choice between ignoring the pilot and trusting it is a false one. A design can move exactly as far as the pilot's evidence warrants and no further, with a guarantee in hand before the main wave begins.

{ {3pt}

singlespace\putbib

}