The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.
107,784 characters
-2.5cm When and How to Pilot: Design Rules for Two-Wave Experiments
\begin{bibunit}
\title{\vspace*{-2.5cm} When and How to Pilot: \\ Design Rules for Two-Wave Experiments}
\author{\large Juan C. Yamin\thanks{Brown University, Department of Economics, Robinson Hall, 64 Waterman Street, Providence, RI 02912. Email: [email removed]}. I am grateful for the generous advice and support of Toru Kitagawa, Soonwoo Kwon, and Jonathan Roth. I thank Andrew Chesher, Yuchen Hu, Bobby Pakzad-Hurson, Chen Qiu, Panos Toulis, Stefan Wager, Kohei Yata, and participants at the 2025 Advances with Field Experiments (AFE) Conference and Brown's Econometric Coffee for their helpful comments.}\\ { Brown University }}
\date{\vspace*{0.5cm} \small \today}
\bgroup
\let\footnoterule\relax
\begin{singlespace}
\maketitle
\begin{abstract}
\noindent Experimenters often run pilots, but how much a small pilot should shape the main-wave design has no settled answer. This paper shows how noisy pilot evidence should guide treatment assignment probabilities in two-wave experiments. Two canonical rules mark the extremes. Balanced assignment guards against worst cases but ignores evidence that one arm is noisier. Feasible Neyman allocation adapts, but with a finite pilot it can overreact to noise, producing arbitrarily large precision losses. We propose a Conditional Minimax Regret (CMR) rule that minimizes worst-case regret over a finite-sample confidence set for the treatment and control variances. CMR retains balance's worst-case protection with high probability, converges to the Neyman allocation as the pilot grows, and attains the minimax-regret rate up to constants. It extends to multi-arm and stratified designs, and simulations calibrated to four field experiments show it avoids feasible Neyman's severe small-pilot losses while capturing most of its large-pilot gains.
\end{abstract}
\end{singlespace}
\vspace{0.15cm}
\noindent \textbf{Keywords:} Experimental design; Statistical decision theory; Minimax regret.\\
\noindent \textbf{JEL Codes:} C44, C90, C93, D81.
\thispagestyle{empty}
\clearpage
\egroup
\setcounter{page}{1}
\section{Introduction}
Nearly one in ten of the 10{,}905 trials registered in the AEA RCT Registry over
the past decade reports running a pilot.\footnote{Between 2016 and 2025, 991 of
the 10{,}905 first-registered trials, or 9.1\%, report running or using a pilot.
Only 24, 0.22\% of all trials in the window, report using the pilot to set the
treatment-assignment probability. Online Appendix~\ref{app:empirics} provides
full details.} Experimenters use them to test logistics, refine survey instruments, and learn about the study population before launching the main wave. Pilot data can also inform a consequential design choice, the treatment assignment probability. By determining how many observations each arm receives, this probability directly affects the variance of the average treatment effect estimator and hence the experiment's power. The promise of pilot-based assignment is that it can use early evidence to allocate more observations where they are most valuable, and hence deliver more precise estimates from the same sample size.
If the potential-outcome variances under treatment and control were known, the variance-minimizing assignment rule would be the Neyman allocation \citep{neyman1992two}. It assigns a larger share of the experimental sample to the noisier arm, reflecting the principle that precision requires more observations where outcomes are more variable. In practice, these variances are unknown when the main-wave design is chosen. A natural rule estimates them from the pilot and substitutes the resulting sample variances into the Neyman formula, yielding the feasible Neyman allocation (FNA). When the pilot is large enough for the estimates to be reliable, this plug-in logic is well founded. For example, \citet{hahn2011adaptive} show that pilot-based estimates of conditional variances can deliver asymptotically efficient assignment probabilities when both pilot and main-wave samples grow large.
When the pilot is small, however, the premise of this plug-in logic fails, and the experimenter must decide how much the design should respond to variance estimates that are not necessarily trustworthy. Keeping the main wave evenly split ensures that neither arm ends up with too few observations, which caps the possible precision loss whatever the true variances. The cost is that it forgoes the gains the pilot could deliver. Following the pilot estimates as if they were true removes this protection, and can tilt the main wave too far, or toward the wrong arm, precisely when the pilot is least informative.
These small-sample concerns are empirically relevant, since in the AEA RCT Registry the median reported pilot has about 300 observations and roughly a tenth have fewer than 30, sizes at which variance estimates can remain very noisy. In fact, \citet{cai2024performance} show that a fixed pilot can leave the feasible Neyman allocation less precise than balance even when the main wave is large. This makes pilot-based assignment a finite-sample design problem, not a plug-in calculation.
We formalize the experimenter's choice as a finite-sample statistical decision problem. In a two-wave experiment with binary treatment, bounded potential outcomes, and no parametric restrictions, the experimenter observes a pilot and then chooses the main-wave assignment probability. The goal is to minimize the variance of the average treatment effect estimator. We evaluate decision rules by their regret, the excess variance relative to the infeasible Neyman allocation that would be chosen if the variance pair were known. Although the setting is deliberately simple, the question it isolates is general. Design rules should be optimized for the information actually available at the moment of design, not for an asymptotic world in which pilot estimates are treated as known. Pilot-based assignment is the tractable model case in which that principle can be developed exactly.
In this setup, the two canonical assignment rules mark the endpoints of the finite-pilot design problem. Balanced assignment, the no-data rule, is optimal under minimax risk and is uniquely minimax-regret optimal without a pilot. It therefore has a formal foundation for caution, but it gives the realized pilot no authority over the assignment. Feasible Neyman gives the pilot's point estimates full authority. This plug-in logic attains the large-pilot efficiency benchmark \citep{armstrong2022asymptotic}, but with a finite pilot nothing limits how far noisy estimates can tilt the assignment, so the precision loss can be arbitrarily large---in the extreme, an arm whose pilot observations all coincide is estimated to have zero variance and receives no main-wave observations at all. In short, balance refuses to move, while feasible Neyman can move too far. The finite-pilot question is how much authority the realized evidence has earned, or equivalently how far the assignment should move from balance toward the feasible Neyman allocation.
Exact minimax regret is the natural ex ante criterion for finite-pilot adaptation \citep{savage1951theory, manski2021econometrics}. It asks, before the pilot is drawn, which complete mapping from pilot realizations to assignments has the smallest worst-case expected regret. By this criterion, with as few as two observations per pilot arm, the minimax-regret value falls strictly below its no-pilot level, so any exact minimax-regret rule must move the assignment away from balance for some pilot realizations. The criterion does not, however, yield an operational design rule for this problem. Regret depends on the population distribution only through the treatment and control variances, yet two populations with the same variances can generate different pilot data, so the minimax problem ranges over full outcome distributions and is not directly computable. More importantly, the criterion evaluates complete pilot-to-assignment mappings, averaging regret over every pilot a population could generate. The authority question starts at the other end, from the single pilot actually drawn, and asks which variance pairs the evidence has ruled out and how much movement from balance the surviving configurations justify.
The Conditional Minimax Regret (CMR) rule answers the authority question with a principle from the literature on inference for decision making, acting on what the data have not ruled out rather than on a point estimate \citep{manski2021econometrics, chernozhukov2025policy}. Once the pilot is observed, the only uncertainty that matters for the loss is which variance pair is true. CMR therefore builds a finite-sample confidence set for the treatment and control variances, the configurations still consistent with the realized pilot, and chooses the main-wave assignment that minimizes worst-case regret over that set. In the binary-treatment case the rule has a closed form, the Neyman formula applied to the midpoint of each arm's standard-deviation confidence interval. When the set is wide, CMR stays close to balance. As the set contracts around the true variances, CMR moves toward the Neyman allocation. Alongside the assignment, CMR reports a certificate, a bound on the regret the chosen assignment can incur \citep{andrews2025certified}. Movement from balance is thus earned by what the pilot rules out, not imposed by a point estimate.
The CMR rule comes with finite-sample and large-pilot guarantees. In finite samples, the certificate bounds the realized regret of the chosen assignment with probability at least \(1-\alpha\), turning a loss the experimenter cannot observe into a number computed from the pilot alone. This certificate is never larger than the guarantee available without a pilot, and it tightens as soon as the pilot rules out one of the adversarial configurations that make a large tilt dangerous. As the pilot grows, the caution that holds CMR near balance relaxes, and at interior variance pairs the rule converges to the infeasible Neyman allocation at the same rate as feasible Neyman. Thus asymptotic efficiency and finite-sample safety hold together rather than trading off. Specifically, CMR's worst-case expected regret matches the minimax-regret rate up to constants, the best rate any pilot-based rule can achieve.
The CMR construction is flexible, and it remains simple to compute across the designs experimenters actually run. In multi-arm experiments with a shared control and in stratified designs, the assignment probability becomes an allocation vector, and CMR again moves from the no-information allocation toward the Neyman allocation as the pilot narrows the confidence set. The same recipe adapts to how outcomes are measured. Binary outcomes admit exact confidence sets that sharpen the rule precisely at the small pilots where it matters most, a kurtosis condition can replace the assumption of a known outcome bound, and the construction carries over when the design targets several primary outcomes or when the pilot observes only a short-run proxy of the outcome that defines the loss.
We also ask when a pilot is worth running for design in the first place, and how large it should be. Observations spent on the pilot could have gone to the main wave instead, so adaptation has to repay that diversion. A worthwhile pilot must be large enough for CMR to move away from balance at all, yet small enough that even perfect adaptation recovers the observations it cost. Folding pilot observations into the final estimator softens this trade-off. The worst-case optimal design cannot be computed, but a simple rule approximates it. A balanced pilot at the two-thirds power of the budget, followed by CMR, comes within a constant factor of the best worst-case performance any two-wave design can achieve, and any design coming close must size its pilot the same way, so larger experiments pilot more but devote a smaller share of their sample to piloting.
Calibrated simulations document that these finite-sample concerns are quantitatively important at realistic pilot sizes. In data-generating processes built
from the public microdata of four well-known field experiments, feasible
Neyman allocation incurs infinite or very large losses at small pilots, while
CMR stays at balance until the pilot makes the confidence rectangle
informative, then captures most of the attainable gain. In the two-arm designs that attainable gain is below one
percent, so the realistic case for a pilot-based rule there is insurance, protection
against the plug-in's unbounded downside at essentially no cost in upside. The pattern sharpens in the multi-arm and
stratified designs, where the same pilot is split across more cells. There
balance leaves more precision on the table, plug-in Neyman
becomes more fragile, and CMR's discipline is worth correspondingly more.
The broader contribution is a template for building statistical uncertainty into experimental design itself. Everything in the paper runs on three features of the assignment problem. First, the unknown parameters enter the design loss only through a low-dimensional vector of variances. Second, the loss is linear in each variance, which makes regret convex, so worst-case regret over any rectangle of plausible values is attained at one of its corners. Third, a small pilot delivers a finite-sample confidence bound for each variance at any prespecified level. Any design problem with these features admits the same treatment, a decision paired with a finite-sample certificate for the loss it can incur, and every extension in this paper is an instance of that recipe.
Two literatures frame these contributions. The first is adaptive experimental design and pilot-based assignment. \citet{hahn2011adaptive} set assignment probabilities from pilot data, \citet{tabord2023stratification} builds adaptive stratification rules, and \citet{bai2022optimality} and \citet{cytrynbaum2021optimal} design matched and stratified experiments from pre-experimental information, with \citet{armstrong2022asymptotic} characterizing the efficiency frontier these designs target. These guarantees are asymptotic. They identify the efficient design once the pilot estimates the variances reliably, but they are silent when the pilot is too small to be trusted. Building on the finite-pilot warning of \citet{cai2024performance}, we formulate pilot-based assignment as a finite-sample decision problem and derive a computable post-pilot rule that pairs the treatment share with a regret certificate.\footnote{A rapidly growing statistics and machine-learning literature studies adaptive Neyman allocation in fully sequential or many-batch experiments, including finite-sample Neyman-regret guarantees \citep{dai2023clip, zhao2024adaptive}. Those rules adapt the assignment continuously as outcomes arrive. Ours solves a one-shot design problem, choosing the main-wave assignment after a single finite pilot.} The contribution is a characterization of how far noisy pilot evidence should move a design while the variances remain uncertain.
The second is inference for decision making. \citet{manski2021econometrics} examines as-if decisions with set estimates, which act on the parameter values a set estimate has not ruled out. \citet{ishihara2021evidence} aggregate estimates from prior studies into a minimax-regret treatment choice, \citet{chernozhukov2025policy} select policies by balancing estimated welfare against estimation risk, and \citet{andrews2025certified} pair recommended decisions with high-probability bounds on their loss. In these settings, the data are already in hand and the decision is which policy to implement. Our decision comes one stage earlier. The action is the main-wave assignment probability, and the loss is the precision of an estimator whose data do not yet exist.\footnote{The regret criterion follows the decision-theory tradition of \citet{savage1951theory}, \citet{manski2004statistical}, \citet{stoye2009minimax,stoye2012minimax}, \citet{tetenov2012statistical}, and \citet{manski2016sufficient}. That literature usually focuses on treatment choice; we focus instead on experimental design. The closest is \citet{hu2024minimax}, who use minimax regret for sample selection, but derive their rules in a local asymptotic framework rather than our finite-sample one.} This structure yields a rule that is closed form in the binary case, immune to the failures of plug-in rules, and accompanied by a certificate for the realized pilot.
The rest of the paper is organized as follows. Sections~\ref{sec:two_wave_design_problem} and~\ref{sec:benchmark_assignment_rules} set up the decision problem and the canonical rules, Section~\ref{sec:conditional_minimax_regret_rule} develops the CMR rule and its guarantees, Section~\ref{sec:extensions} treats multi-arm and stratified designs, Section~\ref{sec:sims} reports the calibrated simulations, and Section~\ref{sec:conclusion} concludes. Online Appendix~\ref{sec:appendix_extensions} extends the rule to binary, unbounded, multiple, and delayed outcomes, and Online Appendix~\ref{sec:appendix_when_to_pilot} analyzes when a pilot is worth running and how large it should be.
\section{The Two-Wave Design Problem}
\label{sec:two_wave_design_problem}
\subsection{Experimental Setup}
\label{sub:experimental_setup}
\paragraph{Population, Outcomes, and Estimand.}
An experimenter aims to estimate the average treatment effect (ATE) of a binary treatment in a target population. The population is characterized by the joint distribution \(F\) of the potential outcomes \((Y(1),Y(0))\), where \(Y(1)\) denotes the outcome under treatment and \(Y(0)\) denotes the outcome under control. Throughout, we assume that potential outcomes take values in the unit interval, \(Y(d)\in[0,1]\) for each treatment status \(d\in\{0,1\}\), so \(F\in\mathcal F=\mathcal P([0,1]^2)\), the collection of all Borel
probability measures on the compact set \([0,1]^2\).\footnote{Since outcomes lie in \([0,1]\), all moments of \(F\) exist and are finite. In particular, the marginal variances are bounded by \(1/4\).} Normalizing outcomes to \([0,1]\) is without loss of generality whenever outcomes are supported on a finite interval, and the particular choice of zero and one is adopted only to simplify notation and exposition. Online Appendix~\ref{sub:unbounded_continuous_outcomes} relaxes boundedness, replacing the known outcome bound with a bound on kurtosis.
The estimand of interest is the ATE, \(\operatorname{ATE}(F)=\mathbb E_F[Y(1)-Y(0)]\), where the expectation is taken with respect to \(F\in\mathcal F\). To increase the precision of the ATE estimator, the experiment proceeds in two waves. A smaller pilot precedes a larger main wave and serves only to inform its design. The single design choice is the main-wave assignment probability. Lower estimator variance translates directly into higher power and a smaller minimum detectable effect at any target power, so variance-minimizing assignment makes the experiment as informative as possible about the ATE.
\paragraph{Pilot Wave.}
The pilot is a sample of \(M\) units whose potential outcomes
\(\{(Y_i(1),Y_i(0))\}_{i=1}^M\) are i.i.d. draws from \(F\). For each unit \(i\), the
experimenter observes the treatment indicator \(D_i\in\{0,1\}\) and the realized
outcome \(Y_i=D_iY_i(1)+(1-D_i)Y_i(0)\). Treatment in the pilot is assigned by a completely randomized design (CRD).\footnote{Fixed pilot arm sizes are imposed only to keep notation simple. The same arguments extend to Bernoulli assignment conditional on both realized arm sizes being at least two.} In the pilot CRD, exactly \(M_1\) units are assigned to treatment and the remaining
\(M_0\) to control, with \(M_1+M_0=M\). The assignment vector \((D_1,\ldots,D_M)\) is sampled uniformly at random from all binary vectors with exactly \(M_1\) treated units.
The pilot realization is denoted by \(\omega=\{(Y_i,D_i)\}_{i=1}^M\), and the pilot
sample space is \(\Omega=([0,1]\times\{0,1\})^M\), with the fixed-arm-size restriction
imposed by the pilot design. Each pilot arm holds at least two units, \(M_1,M_0\geq 2\), so the within-arm sample
variances are well defined. This minimum is maintained wherever the pilot is fixed in
advance, and is relaxed only in Appendix~\ref{sec:appendix_when_to_pilot}, where the pilot size is
itself the object of choice. The extensions in Section~\ref{sec:extensions} impose the
same minimum on each relevant cell.
\paragraph{Main Wave.}
The main wave is a new sample of \(N\) units, independent of the pilot, whose potential
outcomes \(\{(Y_j(1),Y_j(0))\}_{j=1}^N\) are i.i.d. draws from \(F\). The main wave also
uses a completely randomized design. For a treatment assignment probability \(\pi\in(0,1)\), the
design assigns \(N_1=N\pi\) units to treatment and \(N_0=N(1-\pi)\) units to
control.\footnote{We suppress integer constraints on \(N\pi\) throughout, with \(\pi\)
understood as the target treatment assignment probability implemented by the closest feasible
fixed-arm design. The same logic applies under Bernoulli assignment with the estimator of \citet{hajek1971}.} The experimenter estimates the ATE using the difference in means,
\(\widehat{\operatorname{ATE}}=\bar Y_1-\bar Y_0\), where \(\bar Y_d\) is the main-wave
average outcome among units assigned to arm \(d\).
\subsection{Main-Wave Variance and the Neyman Allocation}
\label{sub:main_wave_variance_neyman_allocation}
The marginal potential-outcome variance in arm \(d\) is
\(\sigma_d^2(F)=\operatorname{Var}_F(Y(d))\), abbreviated \(\sigma_d^2\) when
\(F\) is clear from context. Under the main-wave CRD,
\(\operatorname{Var}_F(\widehat{\operatorname{ATE}})=\sigma_1^2/N_1+\sigma_0^2/N_0\), with no cross-arm
covariance term because treatment and control are measured on independent units. With
\(N_1=N\pi\) and \(N_0=N(1-\pi)\), this equals \(V(\pi,\sigma_1^2,\sigma_0^2)/N\), where \(V(\pi,\sigma_1^2,\sigma_0^2)=\sigma_1^2/\pi+\sigma_0^2/(1-\pi)\) for \(\pi\in(0,1)\).
The main-wave sample size \(N\) does not depend on the assignment, so the sampling
variance is proportional to \(V\) with fixed constant \(1/N\), and minimizing the
variance over \(\pi\) is the same as minimizing \(V\), which we take as the variance
criterion for the assignment problem. Boundary recommendations, \(\pi\in\{0,1\}\), are
assigned infinite variance by convention, since assigning all main-wave units to one arm
leaves the other arm mean unobserved.
Minimizing \(V(\cdot,\sigma_1^2,\sigma_0^2)\) over \(\pi\in(0,1)\) yields the Neyman
allocation \citep{neyman1992two}, which for strictly positive variances is \(\pi^*(\sigma_1^2,\sigma_0^2)=\sigma_1/(\sigma_1+\sigma_0)\).
The Neyman allocation gives the larger share of the main wave to the arm with the more
variable outcomes, because that arm's sample mean is noisier and benefits more from a
larger allocation.
The benchmark is the smallest variance an interior assignment can achieve,
\begin{equation}
\label{eq:neyman_value}
V^*(\sigma_1^2,\sigma_0^2)=\inf_{\pi\in(0,1)}V(\pi,\sigma_1^2,\sigma_0^2)
=(\sigma_1+\sigma_0)^2,
\end{equation}
attained at the Neyman allocation when both variances are strictly positive. Writing
\(V^*\) as an infimum keeps it well defined when one variance is zero, since the
minimizer then lies on the boundary and is only approached from the interior.\footnote{With
\(\sigma_0=0\), for instance, the Neyman formula returns the boundary value
\(\pi^*=1\). When both variances are zero, every interior assignment attains the value
zero, and we normalize \(\pi^*(0,0)=1/2\).} Because \(V^*\) depends on the unknown
variance pair, it is the infeasible benchmark against which feasible assignment rules are
evaluated.
\begin{remark}[How much adaptation can gain]
\label{rem:adaptation_ceiling}
The benchmark caps what any pilot-based rule can gain over balance. For every variance pair, \(V(1/2,\sigma_1^2,\sigma_0^2)/V^*(\sigma_1^2,\sigma_0^2)=2(\sigma_1^2+\sigma_0^2)/(\sigma_1+\sigma_0)^2\le 2\), with equality only when one variance is zero. Even perfect adaptation can therefore at most halve the estimator's variance, equivalent to doubling the effective sample size, and realistic configurations deliver far less. Balance's excess variance relative to the benchmark is \((\sigma_1-\sigma_0)^2/(\sigma_1+\sigma_0)^2\), about eleven percent when one standard deviation is twice the other and four percent when it is fifty percent larger. The ceiling rises in the designs of Section~\ref{sec:extensions}, where the no-information allocation must divide the sample more finely. With \(K\) treatment arms sharing a control, perfect adaptation can cut variance by up to a factor of \(K+\sqrt K\) rather than two. In a stratified design the factor is two divided by the smallest stratum share, so with \(S\) equally sized strata it is \(2S\).
\end{remark}
\subsection{States, Actions, and Pilot-Based Rules}
\label{sub:states_actions_pilot_based_rules}
The two-wave design problem is a statistical decision problem. The unknown state of the
world is the population distribution \(F\). It both
fixes the main-wave variance of any assignment and generates the pilot data through which
the experimenter learns about that variance. Because \(F\) is left unrestricted beyond
bounded support, the parameter is the full distribution and the parameter space is the
nonparametric class \(\mathcal F\). However, the main-wave variance depends on \(F\)
only through the pair of marginal potential-outcome variances
\(\theta(F)=(\sigma_1^2(F),\sigma_0^2(F))\), which ranges over \(\Theta=[0,1/4]^2\). This
pair is the payoff-relevant parameter.
Crucially, the variance pair governs the payoff but not the data. At a fixed assignment, the loss is
the same under any two distributions sharing a variance pair, since it depends on \(F\)
only through \(\theta(F)\). The expected loss of a rule can still differ between them,
because the rule chooses its assignment from the pilot, whose distribution depends on the
arm-specific marginal distributions of \(Y(1)\) and \(Y(0)\) rather than on the variance
pair alone. Under the fixed-arm pilot design each unit reveals only one potential outcome,
so the pilot distribution depends on \(F\) only through these two marginals, and never on the
within-unit dependence between \(Y(1)\) and \(Y(0)\). The statistical experiment is thus
indexed by the pair of marginal outcome distributions and not by the variance pair.
The experimenter's action is the main-wave assignment probability \(\pi\in\mathcal A=[0,1]\).
The boundary actions \(\pi\in\{0,1\}\) are kept in \(\mathcal A\) so that the framework can
evaluate rules that would place the entire main wave in one arm, and the infinite-variance
convention of Subsection~\ref{sub:main_wave_variance_neyman_allocation} applies to them.
The experimenter does not know the variance pair and estimates it from the pilot. For each
arm \(d\), the natural estimate of \(\sigma_d^2\) is the within-arm sample variance
\begin{equation*}
\hat\sigma_d^2(\omega)
=\frac{1}{M_d-1}\sum_{i:D_i=d}(Y_i-\bar Y_d)^2,
\qquad
\bar Y_d=\frac{1}{M_d}\sum_{i:D_i=d}Y_i.
\end{equation*}
Under the fixed-arm pilot design, $\hat\sigma_d^2$ is unbiased for $\sigma_d^2$. The main analysis focuses on variance-based decision rules, meaning rules that use the pilot only through the two within-arm sample variances. Formally, the class of
decision rules is
\begin{equation*}
\mathcal D
=\bigl\{\,p:\Omega\to\mathcal A
\;\big|\;
p(\omega)=g\bigl(\hat\sigma_1^2(\omega),\hat\sigma_0^2(\omega)\bigr)
\text{ for some measurable } g:\mathbb R_{+}^2\to\mathcal A\,\bigr\}.
\end{equation*}
The class includes balance, feasible Neyman, trimmed feasible Neyman, and the Conditional Minimax Regret rules developed below. This focus is deliberate, reflecting how pilot-based assignment is typically implemented in practice and keeping the finite-sample design problem low-dimensional enough to yield tractable rules. However, the restriction is substantive since, relative to the unrestricted class $\mathcal{D}_0 = \{p \in \mathcal{A}^\Omega \mid p \text{ is measurable}\}$, features of the full pilot beyond $(\hat\sigma_1^2,\hat\sigma_0^2)$ may contain additional information about $(\sigma_1^2,\sigma_0^2)$ in the nonparametric model.
Once a decision rule is fixed, the realized pilot outcome $\omega$ remains random under $F$. For each $F\in\mathcal F$, let $P_F$ denote the distribution of $\omega$ induced by the fixed-arm pilot CRD with arm sizes $(M_1,M_0)$, suppressing this dependence in the notation. The statistical experiment is the family $\{P_F:F\in\mathcal F\}$ on $\Omega$.
\subsection{Loss and Regret}
\label{sub:loss_and_regret}
The variance criterion \(V(\pi,\sigma_1^2,\sigma_0^2)\) is the loss from using treatment assignment probability \(\pi\) when the variance pair is \(\theta=(\sigma_1^2,\sigma_0^2)\). Regret compares this loss with the infeasible benchmark in \eqref{eq:neyman_value}. For an interior assignment probability, $r(\pi,\theta) = V(\pi,\sigma_1^2,\sigma_0^2) - V^*(\sigma_1^2,\sigma_0^2)$.
For boundary assignments, regret is infinite by the same convention used for \(V\). For a pilot-based rule \(p\in\mathcal D\), the realized regret after observing pilot sample \(\omega\) is \(r(p(\omega),\theta(F))\). Regret measures excess main-wave sampling variance relative to the infeasible Neyman benchmark, so it is always nonnegative. For nondegenerate variance pairs with strictly positive variances, regret is zero exactly when the chosen assignment equals the Neyman allocation. At degenerate variance pairs, regret remains well defined because the infeasible benchmark is an infimum over interior assignments.
For every interior assignment \(\pi\in(0,1)\), regret has the equivalent representations
\begin{equation}
\label{eq:regret_allocation_mistake_baseline}
r(\pi,\theta)
=
(\sigma_1+\sigma_0)^2
\frac{\bigl(\pi-\pi^*(\theta)\bigr)^2}
{\pi(1-\pi)}
=
\frac{\bigl((1-\pi)\sigma_1-\pi\sigma_0\bigr)^2}{\pi(1-\pi)},
\end{equation}
with the normalization \(\pi^*(0,0)=1/2\) in the first. Expression~\eqref{eq:regret_allocation_mistake_baseline} separates the size of the assignment mistake from the penalty attached to it. The squared distance \((\pi-\pi^*(\theta))^2\) to the Neyman target is weighted by \(1/[\pi(1-\pi)]\), so a given mistake is most costly near the boundary, where one arm is barely sampled, and the factor \((\sigma_1+\sigma_0)^2\) scales the loss with the outcome noise. The second form, the square of an expression affine in the two standard deviations, is the version the analysis of Section~\ref{sec:conditional_minimax_regret_rule} exploits.
\section{Canonical Assignment Rules}
\label{sec:benchmark_assignment_rules}
This section studies the two assignment rules that current practice and the existing literature bring to the finite-pilot design problem. The first canonical rule is complete balance, which ignores the pilot and assigns treatment with probability one half. Although simple, balance has a decision-theoretic foundation as the solution to the minimax-risk problem. The second is the feasible Neyman allocation, which uses the pilot in the most direct way, replacing the unknown potential-outcome variances with their pilot estimates.
These two rules expose the central tension of the design problem. Balance is safe but does not adapt, while feasible Neyman adapts but is not safe. We then turn to exact minimax regret, the standard decision-theoretic criterion for resolving this tension. The criterion identifies the best worst-case performance any pilot-based rule can achieve, but it is not itself an operational assignment rule, because the game between the experimenter and Nature ranges over full outcome distributions rather than variance pairs.
\subsection{Minimax Risk and Balanced Assignment \label{sub:minimax_risk_balanced_assignment}}
Intuitively, minimax risk asks for the safest assignment rule when the true variance configuration is unknown, even if the pilot is misleading. Formally, the experimenter may choose any pilot-based rule \(p\in\mathcal D_0\), and each rule is evaluated by its frequentist risk, $\mathcal{R}(p,F) = \mathbb E_{P_F} \!\left[ V\!\bigl(p(\omega),\sigma_1^2,\sigma_0^2\bigr) \right]$, the expected variance of the main-wave ATE estimator when the pilot is generated under \(F \in \mathcal F\) and assignment follows rule \(p\). The minimax-risk criterion ranks rules by their worst-case risk over the model class \(\mathcal F\). The following proposition shows that complete balance is a minimax-risk rule.
\begin{proposition}[{\normalfont\hyperlink{proof:minimax_risk}{Minimax risk}}]
\label{prop:minimax_risk}
The minimax-risk problem $\inf_{p \in \mathcal D_0} \sup_{F \in \mathcal F} \mathcal R(p,F)$ has value \(1\), and the constant balanced rule \(p_{\mathrm{mm}}(\omega) = 1/2\) for all \(\omega \in \Omega\) attains this value.
\end{proposition}
Proofs of the main results are collected in Appendix~\ref{sec:proofs_main_results}. Proofs of the remaining main-text results are collected in Online Appendix~\ref{sec:online_appendix_main_text_proofs}. The intuition behind the proof is that the least favorable distribution makes both arms as noisy as possible. With outcomes bounded in $[0,1]$, this happens when both potential outcomes are Bernoulli with success probability one half, so $\sigma_1^2=\sigma_0^2=1/4$. Under this distribution, the two arms are equally variable, so the best assignment is the balanced split $\pi=1/2$, and even it yields a main-wave variance of $1$. A rule that lets the pilot tilt the main wave is reacting to noise, protecting one arm only by taking observations from an equally noisy other. This single distribution therefore prevents every rule from having worst-case risk below $1$. Balance, for its part, never exceeds this value, because its risk is $2\sigma_1^2+2\sigma_0^2$ and each variance is capped at $1/4$. The two bounds meet at $1$, so balance is minimax.
This optimality gives balance a decision-theoretic foundation. It is not a default adopted by convention but the symmetric hedge against the worst-case distribution, which can make either arm maximally noisy. The same result, however, reveals the limits of the criterion. Its least favorable distribution is the same whatever the pilot shows, so minimax risk gives the pilot no value and keeps the main wave evenly split even after data strongly suggesting that one arm is noisier. The reason is that minimax risk evaluates total main-wave variance, most of which no assignment can avoid. The worst case is then driven by this unavoidable component rather than by the part the pilot can improve, a well-understood conservativeness of minimax criteria \citep{berger1985statistical}.
\subsection{Asymptotic Efficiency and the Feasible Neyman Allocation}
\label{sub:asymptotic_efficiency_feasible_neyman_allocation}
The second canonical rule takes the opposite view from minimax risk. Instead of
asking for the safest rule when the pilot may mislead, it treats the pilot as
a source of estimates for the variance pair that determines the Neyman
allocation. If those estimates are accurate, the natural choice is to plug
them into the Neyman formula. The feasible Neyman allocation does so regardless, treating the sample
variances as if they were known even when the pilot is small and noisy. With
\(\hat\sigma_d(\omega)=\sqrt{\hat\sigma_d^2(\omega)}\) the pilot standard deviation
in arm \(d\), the rule is $\hat p(\omega) = \frac{\hat\sigma_1(\omega)}{\hat\sigma_1(\omega)+\hat\sigma_0(\omega)}$
whenever the two are not both zero, with \(\hat p(\omega)=1/2\) when they are. It assigns the larger share of the main wave to the arm whose pilot outcomes are more
variable, the tilt the Neyman allocation prescribes when the true standard
deviations are known. The appeal of the rule is asymptotic. Outcomes are bounded, so the pilot sample variances are consistent, and whenever \(\sigma_1+\sigma_0>0\) the assignment \(\hat p(\omega)\) converges to the Neyman allocation \(\pi^*(\theta)\). Feasible Neyman thus inherits the precision of the infeasible Neyman allocation, which no design can improve on in large samples \citep{hahn1998role,hahn2011adaptive,armstrong2022asymptotic}.
This large-sample appeal, however, masks a finite-sample fragility that comes from
the ratio form of the rule. On the event that both pilot
standard deviations are positive, substituting \(\hat p(\omega)\) into \(V\) gives
\[
V\!\left(\hat p(\omega),\sigma_1^2,\sigma_0^2\right)
=
\sigma_1^2
\left(
1+\frac{\hat\sigma_0(\omega)}{\hat\sigma_1(\omega)}
\right)
+
\sigma_0^2
\left(
1+\frac{\hat\sigma_1(\omega)}{\hat\sigma_0(\omega)}
\right).
\]
Each term pairs a true variance with a ratio of pilot standard deviations that
scales it. When the pilot estimates are close to the truth and both true standard deviations are positive, the two ratios approach
\(\sigma_0/\sigma_1\) and \(\sigma_1/\sigma_0\), and the expression collapses to
\(V^*=(\sigma_1+\sigma_0)^2\). When one pilot standard deviation is far too small, the ratio that divides by it
grows without bound and turns a true variance of at most \(1/4\) into an
arbitrarily large contribution to \(V\). The loss comes from trusting a small estimate, not from any large variance in the population. When a pilot standard deviation is exactly zero, \(\hat p(\omega)\) lands on the boundary, the main wave assigns no observations to one arm, and that arm's outcome mean cannot be estimated.
The finite-sample consequence, established in the next proposition, is an unbounded worst-case risk. The problem is not only that feasible Neyman sometimes makes a noisy adjustment. The problem is that, with a finite pilot, nothing limits how far a noisy estimate can move the assignment, up to and including assigning no main-wave observations to an arm whose true variance is positive.
\begin{proposition}[{\normalfont\hyperlink{proof:fna_boundary}{Boundary vulnerability of feasible Neyman}}]
\label{prop:fna_boundary}
For every finite pilot size, the feasible Neyman rule has infinite worst-case risk under the boundary convention for \(V\), $\sup_{F\in\mathcal F} \mathcal R(\hat p,F) = \infty$.
\end{proposition}
The same inverse-probability structure that gives the Neyman allocation its efficiency makes the plug-in rule vulnerable at the boundary.
This failure is not driven by an exotic distribution or by a knife-edge pilot sample. It can arise in the simplest binary-outcome experiment. The next remark makes this point explicit.
\begin{remark}[Discrete outcomes and zero pilot variance]
\label{rem:bernoulli}
For a Bernoulli outcome \(Y(d)\sim\mathrm{Bernoulli}(q_d)\) with \(q_d\in(0,1)\), the population
variance \(\sigma_d^2=q_d(1-q_d)\) is strictly positive, while the pilot sample
variance is zero whenever every observed outcome in that arm is equal, an event
with probability \(q_d^{M_d}+(1-q_d)^{M_d}\). Under independent sampling across pilot arms, the probability that exactly one arm shows zero pilot variance is therefore strictly positive for every finite \(M_0,M_1\ge 2\) and every \(q_0,q_1\in(0,1)\). On this event, FNA places the entire main wave in one arm despite both
population variances being strictly positive.
The difficulty is not confined to exact zeros. The smallest positive sample
variance for a Bernoulli arm occurs when a single observation differs from the
rest, giving \(\hat\sigma_d^2=1/M_d\) under the unbiased estimator, with
probability \(M_d q_d(1-q_d)^{M_d-1}+M_d(1-q_d)q_d^{M_d-1}\). A draw of this kind
makes an estimated standard deviation as small as \(M_d^{-1/2}\), which keeps the
realized variance finite but, by the mechanism above, can still make it very large.
\end{remark}
Remark~\ref{rem:bernoulli} suggests an immediate repair. If the problem is that feasible Neyman can assign an arm zero probability, one can force the rule to remain away from zero and one. This is the trimmed feasible Neyman rule, which is discussed in the next remark.
\begin{remark}[Trimmed feasible Neyman]
\label{rem:trimmed_fna}
A natural fix trims feasible Neyman away from the boundary, $\hat p_\tau(\omega)=\min\{\max\{\hat p(\omega),\tau\},1-\tau\}$ with \(\tau\in(0,1/2)\), which caps the realized variance at \(1/(2\tau)\) and removes the infinite-risk pathology. The protection, however, is fixed before any data arrive rather than set by the strength of the pilot evidence, and no single \(\tau\) works well. A large \(\tau\) guards against extreme assignments but stays far from Neyman even after a pilot that has all but resolved the variances, while a small \(\tau\) tracks feasible Neyman but leaves only the weak guarantee \(1/(2\tau)\).
The conflict persists even asymptotically. At \((\sigma_1^2,\sigma_0^2)=(1/4,0)\) the Neyman target \(\pi^*(\theta)=1\) exceeds the cap \(1-\tau\), so regret converges to \(\tfrac{\tau}{4(1-\tau)}>0\) rather than to zero, and shrinking \(\tau\) with the pilot removes this gap only by surrendering the finite-sample protection. Trimming is therefore a blunt safeguard against extreme assignments, not a criterion for how far the realized pilot justifies moving away from balance.
\end{remark}
The feasible Neyman allocation therefore captures the large-pilot ideal but not the finite-pilot problem. It uses the pilot through point estimates and gives no way to distinguish a reliable variance imbalance from a noisy one.
\subsection{Minimax Regret and Finite-Pilot Adaptation}
\label{sub:minimax_regret_finite_pilot_adaptation}
Minimax risk is too conservative for pilot adaptation because it ranks rules by
total main-wave variance, most of which no design can avoid. The regret of
Subsection~\ref{sub:loss_and_regret} instead charges a rule only for the excess
variance its assignment creates relative to the infeasible Neyman benchmark,
isolating the component a pilot can reduce. Ranking rules by worst-case regret
rather than worst-case risk is the standard, less conservative alternative and the
more appropriate criterion here \citep{savage1951theory}.
A statistical decision rule \(p \in \mathcal{D}_0\) is fixed before the pilot, prescribing the main-wave
assignment probability \(p(\omega)\) at every realization \(\omega\). Its expected
regret at \(F\), $R(p,F) := \mathbb E_{P_F}\!\left[ r\!\left(p(\omega),\theta(F)\right) \right]$
is the frequentist risk of \(p\) under regret loss. The exact minimax-regret criterion selects the rule with the smallest worst-case
expected regret. For fixed pilot arm sizes \(M_1,M_0\), with their dependence
suppressed elsewhere in the notation, the minimax-regret value is
\[
R^{\mathrm{mmr}}_{M_1,M_0}
:= \inf_{\delta\in\mathcal D_0} \sup_{F\in\mathcal F}
\mathbb E_{P_{F,M_1,M_0}}\!\left[ r\!\left(\delta(\omega),\theta(F)\right) \right]
= \inf_{\delta\in\mathcal D_0} \sup_{F\in\mathcal F} R(\delta,F).
\]
The three operations read in order. The experimenter chooses a rule \(\delta\),
Nature responds with the least favorable distribution \(F\), and the expectation
averages the resulting regret over the pilots that \(F\) generates. The infimum
ranges over all measurable pilot-to-assignment rules in \(\mathcal D_0\), not only
those that use the pilot through its two sample variances, which makes this the
most demanding ex ante minimax-regret criterion available. When the infimum is
attained, an exact minimax-regret rule is any \(\delta^{\mathrm{mmr}}_{M_1,M_0}\in
\mathcal D_0\) achieving the value.
Proposition~\ref{prop:minimax_regret} records the basic properties of the minimax-regret value and shows when finite-pilot information must be used by an exact minimax-regret rule.
\begin{proposition}[{\normalfont\hyperlink{proof:minimax_regret}{Exact minimax regret}}]
\label{prop:minimax_regret}
The minimax-regret value satisfies the following properties.
\begin{enumerate}
\item[\rm (i)] For every pilot size \((M_1,M_0)\), \(R^{\mathrm{mmr}}_{M_1,M_0} \leq 1/4\). In the no-pilot problem, the minimax-regret value is \(1/4\), and the unique minimax-regret action is balanced assignment.
\item[\rm (ii)] If \(M_1,M_0 \geq 2\), then \(R^{\mathrm{mmr}}_{M_1,M_0} < 1/4\). Consequently, every exact minimax-regret rule \(\delta^{\mathrm{mmr}}_{M_1,M_0}\) satisfies $\Pr_{P_{F,M_1,M_0}}\!\left(\delta^{\mathrm{mmr}}_{M_1,M_0}(\omega)\neq \tfrac12\right)>0$ for some \(F\in\mathcal F\).
\item[\rm (iii)] If \(M_d\to\infty\) for each \(d\in\{0,1\}\), then \(R^{\mathrm{mmr}}_{M_1,M_0}\to 0\).
\end{enumerate}
\end{proposition}
Part~(i) combines a universal upper bound with a no-pilot characterization. Balanced
assignment is feasible at every pilot size and has worst-case regret exactly \(1/4\),
attained when one arm has variance \(1/4\) and the other zero. With no pilot, the feasible rules are the constant
assignments, and any \(\pi\neq 1/2\) undersamples one arm, which Nature punishes by
placing all variance there. Only \(\pi=1/2\) equalizes the two opposing worst cases,
all variance in treatment versus all in control.
Part~(ii) shows that the slightest pilot overturns the no-pilot optimality of balance. With two observations per arm,
the smallest size at which within-arm variation can appear, the minimax-regret value
already drops below \(1/4\), so balance is no longer optimal. The reason is that Nature
now faces two competing forces. Driving the arm variances far apart raises the regret
balance suffers, because balance is optimal only when the two are equal. But that same
asymmetry is what the pilot detects, so it also tells the experimenter
which arm to favor. A slight tilt toward the arm with the larger pilot variance exploits this signal. Where the variances are far apart, the pilot usually identifies the noisier arm and the tilt reduces regret. Where they are close and balance is nearly optimal, the tilt costs almost nothing. Such a rule therefore has worst-case regret strictly below \(1/4\), so every exact minimax-regret rule must leave \(1/2\) with positive probability under some distribution.
Part~(iii) describes the large-pilot limit. As both pilot arms grow, the variance
estimates become accurate enough that a stabilized plug-in rule, constructed in the
proof, approaches the Neyman allocation and its worst-case expected regret vanishes.
Since the minimax-regret value is no larger than the worst-case regret of this rule,
\(R^{\mathrm{mmr}}_{M_1,M_0}\to0\).
Despite the desirable properties recorded in Proposition~\ref{prop:minimax_regret},
Remark~\ref{rem:exact_mmr_computation} shows that the exact minimax-regret problem is
not directly operational, because it does not reduce to a finite-dimensional optimization over the variance pair.
\begin{remark}[Why exact minimax regret is not directly operational]
\label{rem:exact_mmr_computation}
The exact minimax-regret problem is a well-defined decision-theoretic object.
The computational difficulty is not merely that the game is
infinite-dimensional, but that the inner supremum admits no variance-pair
reduction. Nature's choice of $F$ enters the problem twice, through the loss via
the variance pair $\theta(F)$ and through the data via $P_{F,M_1,M_0}$. These two
channels are not linked by the variance pair, as discussed in
Subsection~\ref{sub:states_actions_pilot_based_rules}. The inner supremum
therefore cannot be taken over $\Theta$, and the problem does not collapse to
the static problem $\inf_p\sup_{\theta\in\Theta}r(p,\theta)$ that settles the
no-pilot case.
The usual shortcuts do not restore finite-dimensional structure. The
guess-and-verify approach of \citet{stoye2009minimax}, which identifies a
minimax-regret rule as Bayes against a least favorable prior, does not help here,
because the least favorable prior would, for the same reason, have to range over
$\mathcal F$ rather than over variance pairs, leaving no finite-dimensional
family to guess within. The Bernoulli reduction for bounded mean-payoff problems
\citep{schlag2006eleven,stoye2009minimax}, which replaces $Y\in[0,1]$ by
$B\mid Y\sim\mathrm{Bernoulli}(Y)$, preserves the mean but strictly increases the
variance unless $Y$ is already binary, so it maps the state $\theta(F)$ to a
different variance pair and defines a different variance-based
game.\footnote{The marginal variance is
$\operatorname{Var}(B)=\mathbb E[Y]\,(1-\mathbb E[Y])
=\operatorname{Var}(Y)+\mathbb E[Y(1-Y)]\ge\operatorname{Var}(Y)$.}
Exact minimax regret is best read as a theoretical standard of comparison rather than an
operational rule. It becomes a finite optimization only after one restricts the
outcome distribution, discretizes the pilot, or restricts the class of rules, and
any such computation solves the restricted game rather than the exact
nonparametric problem.\footnote{With \(Y(d)\sim\mathrm{Bernoulli}(q_d)\), the state
\((q_1,q_0)\in[0,1]^2\) pins down both the loss and the pilot distribution,
yet the game remains a
semi-infinite minimax problem with no apparent closed form, numerically
solvable only for small pilots and case by case, since all
assignments, one per pilot realization, must be optimized jointly against a
worst case verified globally over a state space on
which expected regret is nonconcave.}
\end{remark}
Even so, the value \(R^{\mathrm{mmr}}_{M_1,M_0}\) remains informative, since it is the smallest worst-case expected regret attainable by any pilot-based assignment rule. The next proposition shows that even this ideal value cannot vanish faster than the inverse-square-root rate in the pilot arm sizes.
\begin{proposition}[{\normalfont\hyperlink{proof:mmr-lower-bound}{Minimax regret lower bound}}]
\label{prop:mmr-lower-bound}
Fix pilot arm sizes \(M_1,M_0\ge 2\). There exists a universal constant \(c>0\) such that $R^{\mathrm{mmr}}_{M_1,M_0} \ge c\left(M_1^{-1/2}+M_0^{-1/2}\right)$.
\end{proposition}
The bound reflects a limit on what a finite pilot can reveal about the variance
pair. Even the best possible rule faces variance configurations that generate
nearly indistinguishable pilot data but call for different Neyman allocations.
Because the pilot cannot reliably tell such configurations apart, any rule must
make similar recommendations in states where different assignments would be
optimal, and must pay regret in at least one of them. Since the minimax-regret
value already optimizes over all measurable pilot-to-assignment rules, this cost
applies to every pilot-based design. No rule can therefore improve the worst-case
expected regret beyond order \(M_1^{-1/2}+M_0^{-1/2}\). Any tractable rule that
attains this order matches the minimax-regret rate. In sum, exact
minimax regret may be out of reach computationally, but matching its
rate is enough for first-order worst-case optimality.
\section{Conditional Minimax Regret Rule \label{sec:conditional_minimax_regret_rule}}
Section~\ref{sec:benchmark_assignment_rules} leaves unanswered how far the main-wave assignment should move from balance toward the feasible Neyman allocation given the evidence in the realized pilot. A rule that answers this question must be computable from the realized pilot and must move from balance only as far as the evidence justifies.
The regret of an assignment depends on the population distribution only through the marginal potential-outcome variances, so once the pilot is observed, the uncertainty that matters for the main-wave loss is which variance pair is true. The CMR procedure summarizes the pilot's design-relevant information by a finite-sample confidence set, the variance pairs the evidence has not ruled out, and selects the assignment with the smallest worst-case regret over that realized set. The worst case thus runs over the configurations consistent with the realized pilot, not over the full model class and not averaged over pilot realizations. In this sense, CMR is the minimax-regret principle applied to the problem the experimenter faces after the pilot, with the surviving variance pairs in the role of the state space.
Alongside the assignment, CMR reports a certificate in the spirit of \citet{andrews2025certified}, the largest regret the chosen assignment can incur over the configurations the set still allows. Because the set contains the true variance pair with probability at least \(1-\alpha\), the realized regret exceeds the certificate with probability at most \(\alpha\). The certificate is large when the pilot leaves the variances uncertain and small when the pilot has nearly pinned them down.
\subsection{The Conditional Minimax Regret Procedure \label{sub:conditional_minimax_regret_procedure}}
The CMR rule depends on the pilot only through a confidence set for the variance
pair. Because each pilot unit is observed in only one arm, treatment observations
inform $\sigma_1^2$ and control observations inform $\sigma_0^2$. The set therefore combines two separate inferences, an
interval for $\sigma_1^2$ and an interval for $\sigma_0^2$, pairing every treatment
variance in the first with every control variance in the second. This product
structure makes the confidence set a rectangle.
For each arm \(d\in\{0,1\}\) and one-sided error level \(b\in(0,1)\), let
\(\underline{\sigma}_d^2(b;\omega)\) and \(\overline{\sigma}_d^2(b;\omega)\)
denote a lower and an upper confidence bound for \(\sigma_d^2(F)\), computed from
the pilot and taking values in \([0,1/4]\). They must satisfy $\Pr_{P_F}\!\left( \underline{\sigma}_d^2(b;\omega)\le \sigma_d^2(F) \right) \ge 1-b$ and
$\Pr_{P_F}\!\left( \sigma_d^2(F)\le \overline{\sigma}_d^2(b;\omega) \right) \ge 1-b$
for every \(F\in\mathcal F\). Each bound is read as one reads a standard confidence bound, with the pilot as the sample and \(\sigma_d^2(F)\) the unknown parameter. Subsection~\ref{sub:constructing_confidence_rectangle}
constructs such bounds from finite-sample concentration inequalities. The general CMR procedure described below needs only the coverage property.
The rectangle contains the true variance pair \(\theta(F)\) exactly when all four
one-sided bounds hold at once. To keep the overall miss probability at
most \(\alpha\), the construction divides that allowance equally among the four sides
and forms each at one-sided error level \(\alpha/4\). The realized
confidence rectangle is
\begin{equation}
\label{eq:rectangular_confidence_set}
\widehat\Theta_\alpha(\omega)
=
\bigl[
\underline{\sigma}_1^2(\alpha/4;\omega),\,
\overline{\sigma}_1^2(\alpha/4;\omega)
\bigr]
\times
\bigl[
\underline{\sigma}_0^2(\alpha/4;\omega),\,
\overline{\sigma}_0^2(\alpha/4;\omega)
\bigr].
\end{equation}
By Bonferroni's inequality \citep{bonferroni1936teoria}, the probability of a
miss is at most the sum of the four failure probabilities.\footnote{Independence
of the two pilot arm samples would also permit a multiplicative, Šidák-type split of
the error budget across arms \citep{vsidak1967rectangular}.} Each is at most
\(\alpha/4\), so the miss probability is at most \(\alpha\), and $\Pr_{P_F}\!\left( \theta(F)\in\widehat\Theta_\alpha(\omega) \right) \ge 1-\alpha$ for every $F\in\mathcal F$.
The guarantee is frequentist in the usual sense, with the random rectangle covering
the fixed pair \(\theta(F)\) in at least a fraction \(1-\alpha\) of repeated pilot samples.
Given the realized rectangle, CMR chooses the main-wave treatment assignment
probability that minimizes worst-case regret over the variance pairs the pilot has
not ruled out:
\begin{equation}
\label{eq:cmr_assignment}
p_{\mathrm{CMR}}(\omega)
\in
\arg\min_{p\in(0,1)}
\sup_{\theta\in\widehat\Theta_\alpha(\omega)}
r(p,\theta).
\end{equation}
The inner supremum is the largest regret \(p\) could incur over the surviving variance pairs, and the outer minimization selects the treatment assignment probability that makes this worst case smallest. The key simplification is that, after the pilot is translated into the rectangle, the decision problem no longer ranges over full outcome distributions.
The certificate reported with the assignment is the worst-case regret at the chosen
CMR assignment,
\begin{equation}
\label{eq:cmr_certificate}
U_{\mathrm{CMR}}(\omega)
=
\sup_{\theta\in\widehat\Theta_\alpha(\omega)}
r\!\left(p_{\mathrm{CMR}}(\omega),\theta\right).
\end{equation}
Whenever the rectangle contains the true variance pair, the realized regret
\(r(p_{\mathrm{CMR}}(\omega),\theta(F))\) is therefore at most \(U_{\mathrm{CMR}}(\omega)\).
\subsection{Geometry and Closed-Form Solution}
\label{sub:geometry_closed_form_solution}
The CMR assignment and certificate can be computed in closed form once the
realized rectangle is in hand. The dependence on the pilot realization \(\omega\)
is left implicit in what follows. By the second form of the regret identity
\eqref{eq:regret_allocation_mistake_baseline}, regret is the square of
\((1-\pi)\sigma_1-\pi\sigma_0\), which is affine in the two standard deviations,
so for fixed \(\pi\) the worst case over the rectangle is attained at a corner.
The relevant corners are the two off-diagonal ones, the corner where treatment is
as variable as the rectangle allows and control as stable as it allows, and the
corner where the roles are reversed. The proposition below uses this corner
structure to solve the assignment problem in closed form.
\begin{proposition}[{\normalfont\hyperlink{proof:cmr_assignment_rectangle}{Closed-form CMR rule}}]
\label{prop:cmr_assignment_rectangle}
Let \(\underline\sigma_d=\sqrt{\underline\sigma_d^2}\) and \(\overline\sigma_d=\sqrt{\overline\sigma_d^2}\), and suppose \(\overline{\sigma}_1>0\) and \(\overline{\sigma}_0>0\). The CMR assignment problem \eqref{eq:cmr_assignment} has a unique solution, given by
\begin{equation}
\label{eq:cmr_closed_form_assignment}
p_{\mathrm{CMR}}
=
\frac{\overline{\sigma}_1+\underline{\sigma}_1}
{\overline{\sigma}_1+\underline{\sigma}_1+\overline{\sigma}_0+\underline{\sigma}_0}
\in(0,1).
\end{equation}
The corresponding certificate is
\begin{equation}
\label{eq:cmr_closed_form_certificate}
U_{\mathrm{CMR}}
=
\frac{\left(\overline{\sigma}_1\overline{\sigma}_0-\underline{\sigma}_1\underline{\sigma}_0\right)^2}
{\left(\overline{\sigma}_1+\underline{\sigma}_1\right)\left(\overline{\sigma}_0+\underline{\sigma}_0\right)}.
\end{equation}
\end{proposition}
To interpret the assignment formula \eqref{eq:cmr_closed_form_assignment}, multiply its
numerator and denominator by one half, which shows that CMR applies the Neyman formula
to the midpoint of each standard-deviation confidence interval. Treatment receives more
than half the main wave exactly when its confidence-interval midpoint exceeds
control's. For example, if \(\sigma_1\in[0.20,0.40]\) and \(\sigma_0\in[0.10,0.30]\),
the midpoints are \(0.30\) and \(0.20\), and the treatment assignment probability
prescribed by CMR is \(0.30/(0.30+0.20)=0.60\).
A further feature of the closed form is that its assignment is always interior, whereas
feasible Neyman can collapse to the boundary. A zero pilot sample variance is not proof
that the population variance is zero, and the finite-sample constructions of the next subsection keep the
upper endpoint \(\overline{\sigma}_d\) strictly positive in that case. That alone keeps
both arms randomized, without the ad hoc trim feasible Neyman requires.
The certificate has the same corner structure. The two off-diagonal corners pull
the assignment in opposite directions, CMR equalizes their regrets, and
\(U_{\mathrm{CMR}}\) is their common value. The certificate is therefore large
when the rectangle still contains variance pairs that point to substantially
different Neyman allocations, and it vanishes exactly when the two corners imply
the same allocation, which occurs when
\(\overline{\sigma}_1\overline{\sigma}_0=\underline{\sigma}_1\underline{\sigma}_0\).
It measures the spread in implied Neyman allocations, not the raw width of the
rectangle.
The rule and its certificate reduce to familiar values at the two
extremes. When the rectangle contains no information beyond the maintained bounds
and each standard-deviation interval is $[0,1/2]$, the two midpoints coincide at
$1/4$, the rule returns balance, and the certificate equals $1/4$, the worst-case
regret of the no-pilot problem. When the rectangle collapses to a single variance
pair with both standard deviations positive, each confidence-interval midpoint
equals the corresponding true standard deviation, the rule returns the Neyman
allocation $\sigma_1/(\sigma_1+\sigma_0)$, and the certificate is zero.
\subsection{Constructing the Confidence Rectangle}
\label{sub:constructing_confidence_rectangle}
The geometry of the problem also implies that the CMR assignment and certificate depend on the pilot only through the four
rectangle endpoints. The baseline construction of \(\widehat\Theta_\alpha(\omega)\) is
distribution-free and finite-sample, using empirical Bernstein bounds for the
arm-specific standard deviations and converting them into one-sided variance bounds
valid uniformly over the bounded-outcome model. When outcomes are binary, the pilot
distribution is discrete and exactly tractable, so the rectangle can be tightened by
inverting the folded-binomial distribution of the pilot statistic, as developed in Online Appendix~\ref{sub:binary_exact_rectangle}. Either way, the assignment and certificate are computed from the resulting rectangle exactly as in Proposition~\ref{prop:cmr_assignment_rectangle}, and any endpoints satisfying the one-sided coverage requirements of Subsection~\ref{sub:conditional_minimax_regret_procedure} deliver the same guarantees.
The relevant concentration inequality is the empirical Bernstein bound of
\citet{maurer2009empirical} for the sample standard deviation of bounded
random variables.\footnote{Sharper, first-order optimal empirical Bernstein
confidence intervals for the variance of bounded random variables are available
\citep{martinez2025sharp}. The theoretical guarantees extend to these tighter
bounds, but the resulting expressions are considerably more involved.} For each arm \(d\) and one-sided error level \(b\in(0,1)\), write
\(\eta_d(b)=\sqrt{2\log(1/b)/(M_d-1)}\) for the finite-sample concentration radius on
the standard-deviation scale. Then
\[
\Pr_{P_F}\!\left(\sigma_d(F)\le\hat\sigma_d(\omega)+\eta_d(b)\right)\ge 1-b
\quad\text{and}\quad
\Pr_{P_F}\!\left(\sigma_d(F)\ge\hat\sigma_d(\omega)-\eta_d(b)\right)\ge 1-b
\]
uniformly over \(F\in\mathcal F\). The radius \(\eta_d(b)\) depends only on the pilot
size and the error level, shrinking at rate \(1/\sqrt{M_d}\) and growing only as
\(\sqrt{\log(1/b)}\) as the error level falls.
The variance bounds follow by squaring these standard-deviation statements
and projecting onto the maintained variance interval \([0, 1/4]\). For each
arm \(d \in \{0, 1\}\) and error level \(b \in (0, 1)\), the lower and upper
variance bounds are
\begin{equation}
\label{eq:eb_bounds}
\underline{\sigma}_d^2(b;\omega)=\min\!\Big\{\tfrac14,\,\big(\hat\sigma_d(\omega)-\eta_d(b)\big)_{\!+}^{\!2}\Big\},
\qquad
\overline{\sigma}_d^2(b;\omega)=\min\!\Big\{\tfrac14,\,\big(\hat\sigma_d(\omega)+\eta_d(b)\big)^{\!2}\Big\}.
\end{equation}
Both endpoints are projected onto \([0,1/4]\). Because the true variance lies in that interval, the projection preserves the one-sided coverage inequalities.\footnote{At a degenerate pilot \(\hat\sigma_d^2(\omega)=0\), \(\overline{\sigma}_d^2(b;\omega)=\min\{1/4,\eta_d(b)^2\}>0\), so the rectangle never treats it as proof of zero variance.}
The four endpoints, each at error level \(\alpha/4\), assemble into the rectangle of
\eqref{eq:rectangular_confidence_set}, and the union bound delivers its \(1-\alpha\)
coverage, as the following lemma shows.
\begin{lemma}[{\normalfont\hyperlink{proof:mp_rectangle_coverage}{Coverage of the Maurer--Pontil rectangle}}]
\label{lem:mp_rectangle_coverage}
Fix \(\alpha \in (0, 1)\). For each arm \(d \in \{0, 1\}\) and each one-sided
error level \(b \in (0, 1)\), $\Pr_{P_F}\!\left( \underline{\sigma}_d^2(b; \omega) \le \sigma_d^2(F) \right) \ge 1-b$ and $\Pr_{P_F}\!\left( \sigma_d^2(F) \le \overline{\sigma}_d^2(b; \omega) \right) \ge 1-b$
for every \(F \in \mathcal F\). Consequently, the rectangle $\widehat\Theta_\alpha(\omega)$ satisfies $\Pr_{P_F}\!\left( \theta(F) \in \widehat\Theta_\alpha(\omega) \right) \ge 1-\alpha$ for every \(F \in \mathcal F\).
\end{lemma}
\subsection{Statistical Properties}
This subsection establishes three guarantees for the CMR assignment. First, in
finite samples, the reported certificate bounds realized regret with probability
at least \(1-\alpha\). Second, the caution built into the rule disappears as the
pilot grows. At interior variance pairs, CMR converges to the Neyman allocation, with assignment error shrinking at the inverse-square-root rate and realized regret and the certificate at the inverse-pilot rate. Third, we derive an upper bound on CMR's worst-case expected regret that matches the minimax-regret lower bound up to constants.
\paragraph{Finite-sample validity.}
\label{subsub:conditional_optimality}
The link from \(U_{\mathrm{CMR}}\) to a finite-sample guarantee is coverage, the
event that the realized rectangle contains the true variance pair. On that
event, the true pair is among the configurations over which the certificate takes
the worst case, so the realized regret of the CMR assignment cannot exceed
\(U_{\mathrm{CMR}}\). Lemma~\ref{lem:mp_rectangle_coverage} shows that coverage
holds with probability at least \(1-\alpha\). The next theorem combines these two
observations and adds that the certificate never exceeds \(1/4\).
\begin{theorem}[{\normalfont\hyperlink{proof:cmr-certified-optimality}{Finite-sample regret certificate}}]
\label{thm:cmr-certified-optimality}
Fix \(\alpha\in(0,1)\). Then the following statements hold.
\begin{enumerate}
\item[(i)] For every \(F\in\mathcal F\), $\Pr_{P_F}\! \Big( r\!\left(p_{\mathrm{CMR}}(\omega),\theta(F)\right) \le U_{\mathrm{CMR}}(\omega) \Big) \ge 1-\alpha$.
\item[(ii)] For every pilot realization \(\omega\), $U_{\mathrm{CMR}}(\omega)\le \frac14$. Moreover, equality holds if and only if
$(1/4,0)\in\widehat\Theta_\alpha(\omega)$ and $(0,1/4)\in\widehat\Theta_\alpha(\omega)$.
\end{enumerate}
\end{theorem}
Part~(i) shows that, for every distribution, the certificate
\(U_{\mathrm{CMR}}(\omega)\) bounds the regret of \(p_{\mathrm{CMR}}(\omega)\) with
probability at least \(1-\alpha\). The guarantee is finite-sample, holding
at the realized pilot sizes rather than only in a large-pilot approximation.\footnote{Part~(i) also implies quantile domination: for every \(F\) and every \(t\le 1-\alpha\), the \(t\)-quantile of realized regret is bounded by the \((t+\alpha)\)-quantile of the certificate.} The
bound is not only valid but optimal in the sense that
\(U_{\mathrm{CMR}}(\omega)\) is the smallest worst-case
regret bound that any assignment can attain over the variance pairs still
plausible after the pilot. In particular, for every alternative assignment
\(\widetilde p(\omega)\in(0,1)\), $U_{\mathrm{CMR}}(\omega) \le \sup_{\theta\in\widehat\Theta_\alpha(\omega)} r\!\left(\widetilde p(\omega),\theta\right)$.
This makes CMR a certified decision in the framework of
\citet{andrews2025certified}. In their terms, \(U_{\mathrm{CMR}}(\omega)\) is a
\(P\)-certificate for the loss \(r(p_{\mathrm{CMR}}(\omega),\theta(F))\).
Part~(ii) bounds the certificate itself. Balance is always among the candidates and its regret is at most \(1/4\) at every configuration, so the certificate never exceeds \(1/4\), the guarantee available with no pilot.\footnote{The same rectangle can also certify a move against balance directly, since an assignment whose variance is smaller than balance's at every pair in the rectangle is no worse than balance with probability at least \(1-\alpha\). Restricting CMR to such assignments yields a conservative variant with this stronger guarantee.} Equality holds exactly when the rectangle still contains the two adversarial corners \((1/4,0)\) and \((0,1/4)\), one placing all variability in treatment and the other all in control. Because these are opposite vertices of \([0,1/4]^2\), only the full variance space contains both, so any pilot that rules out anything at all earns a certificate strictly below \(1/4\).
\paragraph{CMR convergence rates.}
\label{subsub:cmr_convergence_rates}
Now we ask how that guarantee evolves as the pilot
becomes informative. In the interior of the variance space, where both arms have positive variance, the
confidence rectangle collapses around the truth. The next theorem shows
that, as both pilot arms grow large, the CMR rule converges
to the Neyman allocation, and both realized regret and the certificate shrink to zero.
\begin{theorem}[{\normalfont\hyperlink{proof:cmr-neyman-recovery}{Neyman convergence and interior rates}}]
\label{thm:cmr-neyman-recovery}
Fix \(F\in\mathcal F\), and suppose \(\sigma_1(F)>0\) and \(\sigma_0(F)>0\). Suppose also that \(\alpha\in(0,1)\) is fixed and that \(M_0,M_1\to\infty\), with \(M_0/M_1\) bounded away from zero and infinity. Then the CMR rule satisfies the following statements.
\begin{enumerate}
\item[(i)] $p_{\mathrm{CMR}} \overset{p}{\longrightarrow} \pi^*(\theta(F))$, $|p_{\mathrm{CMR}}-\pi^*(\theta(F))| = O_p\!\left(M_1^{-1/2}+M_0^{-1/2}\right)$.
\item[(ii)] $r\!\left(p_{\mathrm{CMR}},\theta(F)\right) = O_p\!\left(M_1^{-1}+M_0^{-1}\right)$.
\item[(iii)] $U_{\mathrm{CMR}} = O_p\!\left(M_1^{-1}+M_0^{-1}\right)$.
\end{enumerate}
\end{theorem}
All three rates come from a single source, the contraction of the confidence
rectangle around the truth. The Maurer--Pontil endpoints learn each arm's
standard deviation at the inverse-square-root rate in the pilot arm sizes, shrinking the rectangle to the
variance pair at rate \(M_1^{-1/2}+M_0^{-1/2}\). At an interior truth the rule
varies smoothly with the rectangle and inherits this rate, which is
statement~(i). The other two rates are faster, and the reason is that, by the regret identity~\eqref{eq:regret_allocation_mistake_baseline}, regret is quadratic in the gap between the assignment and the Neyman allocation,
with a coefficient that is finite whenever both arms have positive variance. An inverse-square-root error in the assignment then enters regret squared, at rate
\(M_1^{-1}+M_0^{-1}\), which is statement~(ii). The certificate obeys the same bound through its closed form, which squares a
product gap of inverse-square-root order over a denominator bounded away from zero in the
interior, giving statement~(iii).
The key implication of the theorem is that, at an interior truth, the
finite-sample safety of the CMR rule does not slow the rate at which it
converges to the Neyman allocation. In the corresponding large-main-wave limit,
this convergence implies that the resulting assignment approaches the optimized
Hahn-bound allocation, which \citet{armstrong2022asymptotic} identifies as the
first-order efficiency frontier across all designs that may adapt to covariates
and past outcomes. The CMR rule approaches this frontier while certifying its own
regret at every pilot size with probability at least \(1-\alpha\).\footnote{The
feasible Neyman allocation approaches this frontier at the same inverse-square-root rate
as the CMR rule, because the pilot estimates the two standard deviations at that
rate and the Neyman formula responds smoothly to small errors in those inputs at
an interior truth.} First-order asymptotic efficiency and finite-sample safety
hold together, with no trade-off between them.
The interior rates assume both arms have positive variance. The following remark shows that when one arm is degenerate, realized regret and the certificate converge at the slower rate \(M_d^{-1/2}\) rather than \(M_d^{-1}\).
\begin{remark}[Boundary rates]
\label{rem:cmr-boundary-rates}
The interior rates rely on the Neyman allocation lying strictly between zero and
one. There the denominator \(\pi(1-\pi)\) in the regret identity
\eqref{eq:regret_allocation_mistake_baseline} is bounded away from zero, so regret scales with the
squared assignment error, and an error of order \(M_d^{-1/2}\)
produces regret of order \(M_d^{-1}\). A degenerate arm moves the Neyman allocation to a corner and breaks this scaling, slowing realized regret and the certificate to the assignment-error rate itself.
With \(\sigma_1(F)>0\) and \(\sigma_0(F)=0\), the Neyman allocation is \(\pi^*=1\). The upper
confidence endpoint for \(\sigma_0\) does not collapse to zero, so CMR keeps a
small control share \(1-p_{\mathrm{CMR}}\), which contracts at rate \(M_0^{-1/2}\) under the Maurer--Pontil construction. At this truth the regret identity reduces to
$r(\pi,\sigma_1^2(F),0)=\tfrac{\sigma_1^2(F)\,(1-\pi)}{\pi}$,
linear in the leftover control share rather than quadratic. The hedge that keeps the rule off the corner therefore costs
order \(M_0^{-1/2}\) in realized regret, and the certificate is
of the same order. The case \(\sigma_1(F)=0\) and \(\sigma_0(F)>0\) is symmetric.
\end{remark}
\paragraph{Matching the Minimax-Regret Rate.}
\label{subsub:matching_minimax_regret_rate}
The convergence rates just established are pointwise in $F$,
describing how the rule and its realized regret behave at a fixed population as the
pilot grows. The minimax-regret criterion of
Subsection~\ref{sub:minimax_regret_finite_pilot_adaptation} instead judges a
rule by its worst-case expected regret across all populations. As the next theorem shows, the CMR rule's worst-case expected regret shrinks at
the inverse-square-root rate in the pilot arm sizes, and at the inverse-pilot
rate once both arms are bounded away from degeneracy.
\begin{theorem}[{\normalfont\hyperlink{proof:cmr-competitive-risk}{Uniform expected regret}}]
\label{thm:cmr-competitive-risk}
Fix \(\alpha\in(0,1)\). There exists a constant \(C_\alpha<\infty\), depending only on \(\alpha\), such that, for all pilot arm sizes \(M_1,M_0\ge2\),
\[
\sup_{F\in\mathcal F}
\mathbb E_{P_F}\!\left[
r\!\left(p_{\mathrm{CMR}}(\omega),\theta(F)\right)
\right]
\le
C_\alpha
\left(M_1^{-1/2}+M_0^{-1/2}\right).
\]
Moreover, for every \(\kappa>0\), there exists a constant
\(C_{\alpha,\kappa}<\infty\), depending only on \(\alpha\) and \(\kappa\), such that
\[
\sup_{\substack{F\in\mathcal F:\\ \sigma_1(F)\wedge\sigma_0(F)\ge \kappa}} \mathbb E_{P_F}\!\left[ r\!\left(p_{\mathrm{CMR}}(\omega),\theta(F)\right) \right] \le C_{\alpha,\kappa} \left(M_1^{-1}+M_0^{-1}\right).
\]
\end{theorem}
Theorem~\ref{thm:cmr-competitive-risk} evaluates CMR by the ex ante criterion of Section~\ref{sec:benchmark_assignment_rules}, worst-case expected regret. Even though the assignment adapts to the realized pilot, its worst-case expected regret vanishes at the inverse-square-root rate in the pilot arm sizes. On any class of populations with both standard deviations bounded away from zero, the rate improves to the inverse-pilot rate. The slower uniform rate has the same source as the boundary rate of Remark~\ref{rem:cmr-boundary-rates}.
The key reason why the upper bound vanishes is that the assignment and its regret do not blow up when the rectangle \(\widehat\Theta_\alpha(\omega)\) misses the truth \(\theta(F)\). Both depend continuously on how far the pilot's variance estimates fall from the truth, not on whether the rectangle contains it, so a small estimation error keeps regret small, covered or not. The bound is therefore governed by the size of this estimation error, which shrinks at the inverse-square-root rate.
Combining Theorem~\ref{thm:cmr-competitive-risk} with
Proposition~\ref{prop:mmr-lower-bound} yields $\sup_{F\in\mathcal F} \mathbb E_{P_F}\!\left[ r\!\left(p_{\mathrm{CMR}}(\omega),\theta(F)\right) \right] \le \frac{C_\alpha}{c} R^{\mathrm{mmr}}_{M_1,M_0}$.
The theorem gives the CMR upper bound, of order \(M_1^{-1/2}+M_0^{-1/2}\), and the
proposition the matching lower bound for the exact minimax-regret value. CMR's
worst-case expected regret is therefore at most a constant, depending only on
\(\alpha\), times \(R^{\mathrm{mmr}}_{M_1,M_0}\), the smallest any pilot-based rule
can attain. Strikingly, CMR achieves this although it optimizes only against the realized rectangle and never solves the ex ante minimax problem. Unlike an exact minimax-regret rule, CMR is closed form and reports a finite-sample
certificate for the assignment it chooses.
\section{Extensions: Multiple Treatments and Stratification \label{sec:extensions}}
\subsection{Multiple Treatments with a Shared Control}
\label{sub:multiple_treatments_shared_control}
\paragraph{Setup.}
Many experiments compare several treatments against a common control, testing
versions of a program, prices, information treatments, or implementation modes. The
design question is how to divide the main wave across all the treatments and the
shared control at once. The control arm is indexed by \(k=0\) and the treatments by \(k=1,\ldots,K\).
A unit's potential outcomes \((Y(0),\ldots,Y(K))\), with \(Y(k)\in[0,1]\), are
drawn from a population distribution \(F\). The mean and variance of \(Y(k)\)
are \(\mu_k\) and \(\sigma_k^2\), and \(\theta(F)=(\sigma_0^2,
\sigma_1^2,\ldots,\sigma_K^2)\in[0,1/4]^{K+1}\) collects the arm variances. The
main wave is set by the allocation vector \(\pi=(\pi_0,\pi_1,\ldots,\pi_K)\),
where \(\pi_k>0\) is the fraction allocated to arm \(k\) and
\(\sum_{k=0}^K\pi_k=1\). The estimands of interest are the \(K\) treatment-control contrasts
\(\text{ATE}_k = \mu_k-\mu_0\).
For a main wave of size \(N\) with \(N_k\) units on arm \(k\), the
difference-in-means estimator of contrast \(k\) has variance
\(\sigma_k^2/N_k+\sigma_0^2/N_0\). With \(N_k=N\pi_k\), the sum of the marginal variances of the \(K\) contrasts equals \(V(\pi,\theta)/N\), where
$V(\pi,\theta) = \frac{K\sigma_0^2}{\pi_0} + \sum_{k=1}^K \frac{\sigma_k^2}{\pi_k}$.
Every term is an arm variance divided by the share on that arm, so the allocation problem retains the inverse-share structure of the binary case, with the control variance now counted \(K\) times. A pilot of fixed size must now be spread over \(K+1\) arms rather than two, so each arm variance is estimated with fewer observations.
\paragraph{Neyman Allocation.}
For known positive variances, minimizing \(V(\pi,\theta)\) over the simplex \(\Delta_K=\{\pi\in\mathbb R_+^{K+1}:\sum_{k=0}^K\pi_k=1\}\) yields the multi-arm Neyman allocation:
\begin{equation}
\label{eq:multi_arm_neyman}
\pi_0^*(\theta)
=
\frac{\sqrt K\,\sigma_0}
{\sqrt K\,\sigma_0+\sum_{j=1}^K\sigma_j},
\qquad
\pi_k^*(\theta)
=
\frac{\sigma_k}
{\sqrt K\,\sigma_0+\sum_{j=1}^K\sigma_j},
\quad k=1,\ldots,K .
\end{equation}
The benchmark value is \(V^*(\theta)=\inf_{\pi}V(\pi,\theta)=\left(\sqrt K\,\sigma_0+\sum_{k=1}^K\sigma_k\right)^2\), attained at the Neyman allocation \eqref{eq:multi_arm_neyman} when all arm variances are strictly positive, and regret at a feasible \(\pi\) is \(r(\pi,\theta)=V(\pi,\theta)-V^*(\theta)\).
\paragraph{CMR.}
The CMR construction extends by replacing the scalar assignment probability with the allocation vector. Each arm's variance receives a one-sided empirical Bernstein interval as in Section~\ref{sub:constructing_confidence_rectangle}, with the error budget now split equally across the \(2(K+1)\) one-sided endpoints, and stacking the intervals gives the hyperrectangle \(\widehat\Theta_\alpha(\omega)=\prod_{k=0}^{K}\bigl[\underline{\sigma}_k^2(\omega),\overline{\sigma}_k^2(\omega)\bigr]\subseteq[0,1/4]^{K+1}\), which covers the true variance vector with probability at least \(1-\alpha\) for every \(F\). The CMR allocation \(p_{\mathrm{CMR}}(\omega)\in\Delta_K\) minimizes worst-case regret over this hyperrectangle, and the reported certificate \(U_{\mathrm{CMR}}(\omega)\) is again the worst-case regret at the chosen allocation, carrying the same \(1-\alpha\) finite-sample guarantee as in the binary case.\footnote{Online Appendix~\ref{sub:extension_auxiliary_lemmas} collects the properties used across the extensions of Section~\ref{sec:extensions}. Regret is convex in the variance vector, so worst-case regret over the hyperrectangle is attained at a vertex (Lemmas~\ref{lem:extension_regret_convexity} and~\ref{lem:extension_extreme_points}), the CMR problem is a finite convex program (Lemma~\ref{lem:extension_epigraph}), and the reported certificate bounds realized regret with probability at least \(1-\alpha\) (Lemma~\ref{lem:extension_certificate_validity}) and is monotone under set inclusion (Lemma~\ref{lem:extension_set_monotonicity}).} The next proposition traces the
CMR allocation from the safe default, when the hyperrectangle is the full
\([0,1/4]^{K+1}\), to the Neyman allocation, when it collapses to a point as the
pilot grows.
\begin{proposition}[{\normalfont\hyperlink{proof:multi_arm_cmr}{Multi-arm CMR between shared-control balance and Neyman}}]
\label{prop:multi_arm_cmr}
Consider the multi-arm shared-control design.
\begin{enumerate}
\item[(i)] If the pilot is uninformative, with
\(\widehat\Theta_\alpha(\omega)=[0,1/4]^{K+1}\), the CMR allocation is
\[
p_{0,\mathrm{CMR}}
=
\frac{1}{1+\sqrt K},
\qquad
p_{k,\mathrm{CMR}}
=
\frac{1}{\sqrt K\,(1+\sqrt K)},
\quad k=1,\ldots,K .
\]
\item[(ii)] Fix \(\sigma_k(F)>0\) for all \(k\). With \(\alpha\) fixed and
\(M_{\min}=\min_{0\le k\le K} M_k\to\infty\), every measurable selection
$p_{\mathrm{CMR}}(\omega)$ from the CMR solution set satisfies
\( p_{\mathrm{CMR}} \overset{p}{\to}\pi^*(\theta(F))\), the multi-arm Neyman allocation
\eqref{eq:multi_arm_neyman}.
\end{enumerate}
\end{proposition}
Part (i) is the safe default. Unable to treat any arm as noisier than another, CMR plays the equal-variance Neyman allocation, placing \(1/(1+\sqrt K)\) of the main wave on the control and \(1/(\sqrt K(1+\sqrt K))\) on each treatment. For example, with \(K=4\), the control receives a third of the main-wave sample and each treatment a sixth. Part (ii) is the multi-arm analogue of the interior convergence in Theorem~\ref{thm:cmr-neyman-recovery}, driven by the same contraction of the confidence set around the truth. The worst case also changes shape relative to the binary case. The binding configurations are vertices of the hyperrectangle, each making some subset of arms as noisy as the pilot allows and the rest as stable as it allows, and CMR hedges against the worst of these patterns rather than against one arm's variance at a time. Away from the uninformative rectangle, neither this allocation nor its stratified counterpart has a closed form. Online Appendix~\ref{sub:extension_computation} shows how both are computed as finite convex programs over the vertices of the realized hyperrectangle.
\subsection{Stratified Experiments}
\label{sub:stratified_experiments}
\paragraph{Setup.}
Many target populations divide into strata, such as schools, regions, or gender, known before the main wave is designed. A stratified design involves two decisions, how many observations each stratum receives and how each stratum's observations are split between treatment and control. The pilot informs both decisions.
Strata are indexed by \(x=1,\ldots,S\), a unit's pre-specified stratum is \(X\in\{1,\ldots,S\}\), and the target-population shares \(s_x=\Pr_F(X=x)>0\) sum to one. These shares are known and fixed before the main-wave design is
chosen. Within stratum \(x\), the potential outcome
under arm \(d\in\{0,1\}\) has conditional mean
\(\mu_{dx}=\mathbb E_F[Y(d)\mid X=x]\) and variance
\(\sigma_{dx}^2=\operatorname{Var}_F(Y(d)\mid X=x)\). As before, the estimand is the
population average treatment effect, \(\operatorname{ATE}(F)=\sum_{x=1}^S
s_x(\mu_{1x}-\mu_{0x})\), the average of the within-stratum effects
\(\mu_{1x}-\mu_{0x}\) weighted by the population shares.
The two decisions combine into a single design object, the treatment-by-stratum
cell shares \(\pi=\{\pi_{dx}:d\in\{0,1\},\ x=1,\ldots,S\}\), where \(\pi_{dx}\) is
the fraction of the main wave drawn from stratum \(x\) and assigned to arm \(d\).
These shares are nonnegative fractions of one fixed main wave, so they sum to one
and \(\pi\) ranges over the simplex
\(\Delta_{2S-1}=\{\pi:\pi_{dx}\ge 0,\ \sum_{x=1}^S(\pi_{1x}+\pi_{0x})=1\}\). Two summaries of $\pi$ recover the two decisions. The total share collected
from stratum $x$, $\pi_{\cdot x} = \pi_{1x} + \pi_{0x}$, is the sampling
margin, and the within-stratum treatment probability,
$\pi_{1\mid x} = \pi_{1x}/(\pi_{1x} + \pi_{0x})$, is the assignment
margin.\footnote{The assignment margin is defined where $\pi_{\cdot x} > 0$. As in the
baseline model, boundary allocations are allowed as formal actions, but any
allocation that leaves a treatment-by-stratum mean entering the estimand
unobserved carries infinite loss by convention.}
The stratified difference-in-means estimator
\(\widehat{\operatorname{ATE}}=\sum_{x=1}^S s_x(\bar Y_{1x}-\bar Y_{0x})\) estimates
each within-stratum effect by the cell contrast \(\bar Y_{1x}-\bar Y_{0x}\)
and aggregates these contrasts using the target-population shares \(s_x\), where
\(\bar Y_{dx}\) is the sample mean in cell \((d,x)\) \citep{imbens2015causal}. For positive
cell shares, independent sampling across cells gives an estimator variance of \(V(\pi,\theta)/N\), where
\begin{equation}
\label{eq:stratified_variance_criterion}
V(\pi,\theta)
=
\sum_{x=1}^S s_x^2
\left(
\frac{\sigma_{1x}^2}{\pi_{1x}}
+
\frac{\sigma_{0x}^2}{\pi_{0x}}
\right),
\qquad
\theta=\{\sigma_{dx}^2:d\in\{0,1\},\ x=1,\ldots,S\}.
\end{equation}
The square \(s_x^2\) appears because the estimator multiplies each stratum contrast by \(s_x\).
\paragraph{Neyman Allocation.}
If the cell variances were known, the optimal stratified design would minimize
\(V(\pi,\theta)\) over the cell-share simplex $\Delta_{2S-1}$. For positive variance
configurations, the solution is
\begin{equation}
\label{eq:stratified_neyman}
\pi_{dx}^*(\theta)
=
\frac{s_x\sigma_{dx}}
{\sum_{z=1}^S s_z(\sigma_{1z}+\sigma_{0z})},
\qquad
d\in\{0,1\},\quad x=1,\ldots,S .
\end{equation}
The optimal design assigns observations to cells in proportion to \(s_x\sigma_{dx}\), how much the cell matters for the target ATE times how noisy its mean is to estimate. A large stratum with a nearly deterministic outcome needs few observations, and a noisy cell receives little weight if it represents a small part of the population.
Summing \eqref{eq:stratified_neyman} over arms gives the Neyman sampling margin,
\(\pi_{\cdot x}^*(\theta)=\tfrac{s_x(\sigma_{1x}+\sigma_{0x})}{\sum_{z=1}^S s_z(\sigma_{1z}+\sigma_{0z})}\),
which samples a stratum more than proportionally
to its population share exactly when its total standard deviation
\(\sigma_{1x}+\sigma_{0x}\) exceeds the population-weighted average of total
standard deviations across strata. The Neyman assignment margin is the standard
two-arm Neyman allocation, \(\pi^*_{1\mid x}=\sigma_{1x}/(\sigma_{1x}+\sigma_{0x})\). Substituting \eqref{eq:stratified_neyman} into
\eqref{eq:stratified_variance_criterion} gives $V^*(\theta) =\left[ \sum_{x=1}^S s_x(\sigma_{1x}+\sigma_{0x})\right]^2$, the smallest variance any feasible design attains. Regret takes the same form as in the baseline model, $r(\pi,\theta) = V(\pi,\theta)-V^*(\theta)$.
\paragraph{CMR.}
From the \(M_{dx}\ge 2\) pilot observations in cell \((d,x)\), a
one-sided empirical Bernstein interval of \citet{maurer2009empirical} is built for
each cell variance \(\sigma_{dx}^2\), with the error split equally
across the \(4S\) one-sided endpoints at level \(\alpha/(4S)\). Stacking the
\(2S\) cell intervals gives the hyperrectangle $ \widehat\Theta_\alpha(\omega) = \prod_{d\in\{0,1\}} \prod_{x=1}^S [ \underline{\sigma}_{dx}^2(\omega), \overline{\sigma}_{dx}^2(\omega) ] \subseteq [0,1/4]^{2S}$,
which covers the true variance vector with probability at least \(1-\alpha\) by the
union bound. The CMR rule and its certificate take the same form as before, now over the cell-share simplex \(\Delta_{2S-1}\), with \(p_{\mathrm{CMR}}(\omega)\)
collecting the cell shares \(p_{dx,\mathrm{CMR}}(\omega)\).
\begin{proposition}[{\normalfont\hyperlink{proof:stratified_cmr}{Stratified CMR between representative balance and Neyman}}]
\label{prop:stratified_cmr}
Consider the stratified design.
\begin{enumerate}
\item[(i)] If the pilot is uninformative, with \(\widehat\Theta_\alpha(\omega)=[0,1/4]^{2S}\), the CMR allocation is $ p_{1x,\mathrm{CMR}}(\omega) = p_{0x,\mathrm{CMR}}(\omega) = \frac{s_x}{2}$, $x=1,\ldots,S$. Equivalently, \(p_{\cdot x,\mathrm{CMR}}(\omega)=s_x\) and \(p_{1|x,\mathrm{CMR}}(\omega)=1/2\) for every stratum \(x\).
\item[(ii)] Suppose \(\sigma_{dx}(F)>0\) for all \(d\in\{0,1\}\) and \(x=1,\ldots,S\). With \(\alpha\) fixed and \(M_{\min}=\min_{d,x}M_{dx}\to\infty\), every measurable selection \(p_{\mathrm{CMR}}(\omega)\) from the CMR solution set satisfies $p_{\mathrm{CMR}}(\omega) \overset{p}{\longrightarrow} \pi^*(\theta(F))$, the stratified Neyman allocation in \eqref{eq:stratified_neyman}.
\end{enumerate}
\end{proposition}
With an uninformative pilot the rule has no reason to treat any stratum or arm as noisier than another, so the allocation is representative sampling with an even split within each stratum. As the pilot shrinks the confidence set, a stratum revealed to be noisier receives more than its population share, the split within each stratum moves toward its noisier arm, and in the limit the allocation becomes the stratified Neyman allocation \eqref{eq:stratified_neyman}.
\section{Calibrated Simulations \label{sec:sims}}
This section stress-tests the allocation rules at practice-relevant pilot sizes,
between 30 and 500 total observations. We calibrate the data-generating
processes to public microdata from four published field experiments: the deworming program of \citet{miguel2004worms} (MK), the resume audit of \citet{bertrand2004emily} (BM), the experiment on incentives to learn HIV results of
\citet{thornton2008demand}, and the reference-letter experiment of
\citet{abel2020value}.\footnote{We select published experiments in leading
economics journals with public microdata, randomized assignment arms that map
directly into our design problems, and headline outcomes from the original
papers. Across studies, the retained outcomes span binary, count, and bounded
continuous measurements.} The two-arm designs are built from MK, BM, and Thornton, the shared-control multi-arm design from Thornton's randomized incentive amounts, and the stratified design from the gender stratification in Abel et al.
\subsection{Simulation Design \label{sub:sims_design}}
Each data-generating process fixes the outcome distributions that the original
experiment induced in its randomized arms. For binary outcomes, treatment and
control outcomes are Bernoulli draws with the arm means estimated from the
microdata. For continuous and count outcomes, outcomes are drawn from the
empirical distribution of the corresponding arm, rescaled to the unit interval
using the observed range. The multi-arm design keeps Thornton's incentive
levels and their shared control as five separate arms, and the stratified
design applies the same construction within gender strata, holding the stratum
shares fixed at their sample values.
Table~\ref{tab:section7-dgp-calibration} reports the calibrated arm means,
variances, and implied Neyman allocations for each design. In the two-arm designs, the infeasible Neyman allocation assigns between \(46\) and \(55\) percent of the sample to treatment, so balance is never far from optimal. Balance's excess variance over the infeasible allocation is at most \(0.84\) percent, which also caps the improvement any pilot-based rule can deliver, the ceiling of Remark~\ref{rem:adaptation_ceiling} materializing in calibrated data.
Each simulation replication draws a pilot of total size \(M\) from the calibrated distribution. The pilot itself is assigned without any information, evenly across arms in the two-arm and multi-arm designs, and representatively across strata with an even treatment-control split within each stratum in the stratified design. To make the designs comparable, our performance metric is the percentage efficiency loss, \(100\times[V(p,\theta)-V^*(\theta)]/V^*(\theta)\), where \(p\) is the main-wave assignment the rule chooses from the realized pilot. An entry of \(1\) means the assignment raises the estimator's variance by one percent, or equivalently that the experimenter would need one percent more main-wave observations to reach the same precision. Each entry averages \(500\)
independent pilot draws, and the tables report \(M\in\{30,100,250,500\}\).
\subsection{Two-Arm Results \label{sub:sims_two_arm}}
Table~\ref{tab:section7-main-efficiency-loss} compares the mean efficiency
losses of balance, FNA, and CMR across the seven two-arm designs. Balance has a
single entry per design because it does not use the pilot.
Figure~\ref{fig:section7-main-regret} plots the comparison over the full grid of
pilot sizes.
\begin{table}[tbp]
\centering
\caption{Mean efficiency loss in the calibrated two-arm simulations}
\label{tab:section7-main-efficiency-loss}
\begingroup
\small
\setlength{\tabcolsep}{5.0pt}
\begin{tabular}{@{}llrrrr@{}}
\toprule
Study/outcome & Rule & $M=30$ & $M=100$ & $M=250$ & $M=500$ \\
\midrule
\multirow{3}{*}{MK: attendance} & Balance & \multicolumn{4}{c}{0.02} \\
& CMR & 0.02 & 0.02 & 0.08 & 0.07 \\
& FNA & 3.33 & 0.71 & 0.27 & 0.12 \\
\addlinespace[0.38em]
\multirow{3}{*}{MK: test score} & Balance & \multicolumn{4}{c}{0.07} \\
& CMR & 0.07 & 0.07 & 0.05 & 0.05 \\
& FNA & 2.19 & 0.58 & 0.25 & 0.12 \\
\addlinespace[0.38em]
\multirow{3}{*}{MK: mod.-heavy infection} & Balance & \multicolumn{4}{c}{0.45} \\
& CMR & 0.45 & 0.07 & 0.08 & 0.08 \\
& FNA & $\infty$ & 0.20 & 0.06 & 0.03 \\
\midrule
\multirow{3}{*}{Thornton: HIV result} & Balance & \multicolumn{4}{c}{0.59} \\
& CMR & 0.59 & 0.18 & 0.13 & 0.12 \\
& FNA & $\infty$ & 0.36 & 0.14 & 0.06 \\
\addlinespace[0.38em]
\multirow{3}{*}{Thornton: condom purchase} & Balance & \multicolumn{4}{c}{0.28} \\
& CMR & 0.28 & 0.18 & 0.10 & 0.07 \\
& FNA & $\infty$ & 0.49 & 0.18 & 0.08 \\
\addlinespace[0.38em]
\multirow{3}{*}{Thornton: condom count} & Balance & \multicolumn{4}{c}{0.18} \\
& CMR & 0.18 & 0.22 & 0.27 & 0.22 \\
& FNA & $\infty$ & 5.93 & 2.14 & 1.09 \\
\midrule
\multirow{3}{*}{BM: callback} & Balance & \multicolumn{4}{c}{0.84} \\
& CMR & 0.84 & 0.84 & 0.47 & 0.44 \\
& FNA & $\infty$ & $\infty$ & 1.14 & 0.48 \\
\bottomrule
\end{tabular}
\vspace{0.25em}
\begin{minipage}{0.88\textwidth}
\footnotesize
\emph{Notes:} Entries are percent mean efficiency losses relative to the infeasible Neyman allocation:
$100\,\mathbb E_\omega[V(\hat p_a(\omega),F)-V(\pi^*(F),F)]/V(\pi^*(F),F)$.
Columns $M=30,100,250,500$ index the pilot size; the Balance row reports a single value because it does not use the pilot.
Each entry uses 500 pilot replications per DGP and pilot size.
$\infty$ indicates infinite unconditional mean efficiency loss, which occurs when FNA assigns zero probability to an arm with positive true variance in at least one pilot draw.
\end{minipage}
\endgroup
\end{table}
The FNA rows quantify the downside of trusting a small pilot. At \(M=30\), the efficiency loss is infinite in five of the seven designs and equals \(3.33\) and \(2.19\) percent in the other two, several times the \(0.84\) percent ceiling on what adaptation can gain in these designs. The infinite entries are the boundary failure of Proposition~\ref{prop:fna_boundary} materializing in calibrated data. In the callback design, where both callback rates are below ten percent, the failure persists through \(M=100\), and even where every draw is finite convergence is slow, with the condom-count design still losing \(5.93\) percent at \(M=100\).\footnote{Common repairs of FNA do not resolve this.
Table~\ref{tab:section7-appendix-rule-comparison} evaluates FNA with its
assignment trimmed away from the boundary, together with the pre-test and
regularized rules studied by \citet{cai2024performance}. Trimming removes the
infinite entries, but all four alternatives still lose substantially at
\(M=30\).}
At \(M=30\), CMR's efficiency loss coincides with balance's in every design because the confidence rectangle still contains every variance pair and the rule returns balance. The same discipline
appears at \(M=100\) in the callback design. There the confidence rectangle
excludes at least one corner of the variance space in only \(3.2\) percent of
pilot draws (Table~\ref{tab:section7-cmr-diagnostics}), so the set remains nearly uninformative and the mean efficiency loss stays at the balance value of \(0.84\) percent.
Once the rectangle becomes informative, CMR's gains arrive in the designs with
unequal variances. The efficiency loss falls from \(0.45\) to \(0.07\) percent in the MK infection design between \(M=30\) and \(M=100\), from \(0.59\) to \(0.18\) percent in the Thornton HIV-result design, and from \(0.84\) to \(0.47\) percent in the callback design by \(M=250\). In the designs with nearly equal variances there is nothing to find, so any movement the rectangle licenses chases sampling noise and CMR can lose to balance. In the condom-count design it
loses between \(0.22\) and \(0.27\) percent against \(0.18\) for balance, a
premium of at most \(0.09\) percentage points, in the same design in which
FNA loses \(5.93\) percent at \(M=100\). By \(M=500\), the pilot variance estimates are accurate enough that FNA is modestly ahead of CMR in several designs.\footnote{Table~\ref{tab:section7-regret-distribution} reports the distribution of the efficiency loss across pilot draws, including medians, standard deviations, maxima, and FNA boundary frequencies. The worst CMR efficiency loss across every design and pilot size is \(5.95\) percent, against FNA boundary failures and finite losses as large as \(57.56\) percent. Once adaptation begins, CMR's loss distribution is right-skewed in some designs, because some draws still move the assignment in the wrong direction, though the rule's conservatism keeps these movements small.}
Table~\ref{tab:section7-cmr-diagnostics} shows that, in the reported Monte Carlo draws, the confidence rectangle covers the true variance pair for every design and pilot size and the certificate is at least as large as the realized regret in every draw. Coverage exceeds its nominal \(1-\alpha\) level by a wide margin,
reflecting the conservativeness of the distribution-free Maurer--Pontil bounds. The median certificate falls from its uninformative value of one quarter at \(M=30\) to between \(0.05\) and \(0.14\) at \(M=500\). Relative to the no-pilot guarantee of one quarter, a pilot of \(500\) observations thus cuts the certified worst case by between roughly one half and four fifths.
Table~\ref{tab:section7-applied-implications} translates the losses into the
design quantities applied experimenters plan with, additional subjects and
statistical power. In the worst cases across the seven designs at \(M=30\),
FNA requires \(64\) additional main-wave subjects per \(1{,}000\) to match the
precision of the infeasible allocation, and its power against an effect that
the infeasible allocation detects with \(80\) percent power falls to \(50\)
percent. Under CMR, the extra subjects never exceed \(8.4\) per \(1{,}000\),
and power stays within \(0.3\) percentage points of the target in every design
and at every pilot size. The cost of overreacting to a small pilot is measured
in dozens of subjects or many percentage points of power. The cost of CMR's caution is
measured in a handful of subjects.
\subsection{Multi-Arm and Stratified Results \label{sub:sims_extensions}}
The extension designs raise both the value and the risk of adaptation, since the allocation is now a vector rather than a single treatment share.
Table~\ref{tab:section7-extension-efficiency-loss} reports the results for the
two extensions of Section~\ref{sec:extensions}. In the table, the balance row is
the no-information allocation of each design rather than an equal split across
cells. In Panel~A it is the shared-control allocation of
Proposition~\ref{prop:multi_arm_cmr}, one third of the sample to control and one
sixth to each treatment. In Panel~B it is representative balance from
Proposition~\ref{prop:stratified_cmr}, each gender stratum at its population
share, split evenly between treatment and control. Figures~\ref{fig:section7-multiarm-extension-regret}
and~\ref{fig:section7-stratified-extension-regret} plot both panels over the
full grid of pilot sizes.
\begin{table}[tbp]
\centering
\caption{Mean efficiency loss in the calibrated extension simulations}
\label{tab:section7-extension-efficiency-loss}
\begingroup
\small
\setlength{\tabcolsep}{5.0pt}
\begin{tabular}{@{}p{0.28\textwidth}@{}*{4}{>{\centering\arraybackslash}p{0.15\textwidth}@{}}}
\toprule
Rule & $M=30$ & $M=100$ & $M=250$ & $M=500$ \\
\midrule
\multicolumn{5}{@{}p{0.88\textwidth}@{}}{\textit{Panel A: Thornton (2008): learned HIV result, multi-arm shared-control}} \\
\addlinespace[0.10em]
Balance & \multicolumn{4}{c}{1.76} \\
CMR & 1.76 & 1.76 & 1.48 & 0.34 \\
CMR (Bernoulli) & 1.82 & 1.21 & 0.75 & 0.36 \\
FNA & $\infty$ & $\infty$ & $\infty$ & 0.41 \\
\midrule
\multicolumn{5}{@{}p{0.88\textwidth}@{}}{\textit{Panel B: Abel et al. (2020): job applications, gender-stratified}} \\
\addlinespace[0.10em]
Balance & \multicolumn{4}{c}{5.59} \\
CMR & 5.59 & 5.59 & 5.38 & 2.94 \\
FNA & 72.13 & 26.01 & 9.27 & 3.86 \\
\bottomrule
\end{tabular}
\vspace{0.25em}
\begin{minipage}{0.88\textwidth}
\footnotesize
\emph{Notes:} Entries are percent mean efficiency losses relative to the infeasible Neyman allocation:
$100\,\mathbb E_\omega[V(\hat p_a(\omega),F)-V(\pi^*(F),F)]/V(\pi^*(F),F)$.
Columns $M=30,100,250,500$ index total pilot size; each entry uses 500 pilot replications.
Balance is the appropriate non-adaptive benchmark for each design: shared-control balance-equivalent in Panel A and half treatment within each stratum in Panel B.
CMR is the Maurer--Pontil bounded-outcome rule; CMR (Bernoulli) uses Bernoulli variance confidence sets for the binary multi-arm outcome.
$\infty$ indicates infinite unconditional mean efficiency loss, which occurs when FNA assigns zero probability to a positive-variance arm or treatment-by-stratum cell in at least one pilot draw.
\end{minipage}
\endgroup
\end{table}
The first difference from the two-arm designs is that the infeasible benchmark adapts along more margins. In Panel~A, it exploits variance differences across five arms. In Panel~B, it both samples noisier strata more heavily and tilts the treatment split within each stratum. Balance
accordingly loses \(1.76\) percent in Panel~A and \(5.59\) percent in
Panel~B, well above the \(0.84\) percent maximum gain available in the two-arm designs.
The second difference is that the same pilot now feeds more variance
estimates. At \(M=30\), each of the five Thornton arms contributes six pilot
observations, and each treatment-by-stratum cell in the Abel design
contributes fewer than ten. Consequently, FNA's mean efficiency loss is infinite in Panel~A through \(M=250\), since a
single zero-variance binary arm is enough to push it to the boundary. Panel~B
shows that the boundary event is not the whole problem. There FNA's losses are
finite but severe, \(72.13\) percent at \(M=30\) and \(26.01\) percent at
\(M=100\), because interior allocations built on noisy cell variances are also
far from optimal.
The CMR efficiency loss equals the balance loss at \(M=30\) and \(M=100\) in both panels. The reason is that each cell's variance bound now rests on fewer observations, and the rectangle splits its error probability across more cell-level bounds, so the pilot size at which it first becomes informative grows with the number of cells. Once it does, the rule adapts, to \(1.48\) percent at \(M=250\) and \(0.34\) at
\(M=500\) in Panel~A, and to \(5.38\) and then \(2.94\) in Panel~B.
Panel~A includes a second CMR row because the outcome is binary, and it shows
that the rule's caution resides in the confidence set rather than in the
minimax-regret logic. CMR with the exact-inversion Bernoulli variance sets of Online Appendix~\ref{sub:binary_exact_rectangle} becomes informative at much smaller pilots. It pays a small early premium, \(1.82\)
percent against \(1.76\) at \(M=30\), and then adapts earlier than the
bounded-outcome rectangle, reaching \(1.21\) percent at \(M=100\) and
\(0.75\) at \(M=250\), where the Maurer--Pontil version still sits at
\(1.76\) and \(1.48\).
The extensions sharpen the paper's main message.\footnote{Table~\ref{tab:section7-extension-applied-implications} translates the extension efficiency losses into subjects and power. At \(M=30\), FNA's joint-test power in Panel~A is \(20.7\) percent against \(79.1\) for balance and CMR, and in Panel~B FNA requires \(590\) additional main-wave subjects per \(1{,}000\) to match the infeasible benchmark, against \(56\) for balance and CMR. Balance itself is expensive here, and by \(M=500\) CMR cuts the extra subjects to \(27\) and raises power from \(77.8\) to \(78.9\) percent, while FNA's efficiency loss in Panel~A falls to \(0.41\) percent.} Exactly where
adaptation has the most to buy and plug-in rules are most dangerous, CMR
captures a growing share of the attainable gain as the pilot informs, never
pays more than a tenth of a percentage point for its caution, and is the only
adaptive rule in these tables that arrives with a finite-sample guarantee for the
assignment it selects.
\section{Conclusions \label{sec:conclusion}}
The central question in pilot-based design is how much authority noisy preliminary evidence should have over the main experiment. Balanced assignment gives the pilot none, and feasible Neyman gives its point estimates full authority. Rather than acting on point estimates or forgoing adaptation altogether, the Conditional Minimax Regret rule acts on what the pilot has ruled out, staying at balance when the evidence is weak and approaching the Neyman allocation as the evidence accumulates. Before committing the main wave, the experimenter holds a certificate that, with high probability, bounds the precision lost to the realized design and never exceeds balance's no-pilot guarantee.
The same CMR logic applies to any design choice whose loss depends on parameters a small pilot estimates imprecisely. Two directions for future research seem most valuable. The first is to develop such rules for richer designs, including cluster-randomized, sequential, and imperfect-compliance experiments. The second is to treat the confidence set itself as part of the decision. CMR splits its error budget symmetrically and takes the coverage level as given, and neither choice is necessarily optimal. Sharper sets would let the rule earn the right to adapt sooner.
Much of the theory of adaptive experimentation justifies pilot-based designs by letting the pilot grow large. Real pilots are small, and at the sizes experimenters actually run, adaptation has modest upside and unbounded downside. The practical lesson of this paper is that the choice between ignoring the pilot and trusting it is a false one. A design can move exactly as far as the pilot's evidence warrants and no further, with a guarantee in hand before the main wave begins.
\clearpage
{\small
\setlength{\bibsep}{3pt}
\begin{singlespace}
\putbib
\end{singlespace}
}
\clearpage