Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
74,917 characters · 18 sections · 31 citation commands
Selecting the Best Arm in One-Shot Multi-Arm RCTs: The Asymptotic Minimax-Regret Decision Framework for the Best-Population Selection Problem
Randomized controlled trials (RCTs) are now indispensable in data-driven decision-making across public policy, business, and clinical trials. Most one-shot RCTs still use two-arm “A/B test” designs, for which simple threshold rules are known to be decision-theoretically optimal under usual conditions Karlin1956. When comparisons span more than two arms, decision-makers (DMs) typically rely on ad hoc rules, such as pairwise tests designed for two-arm settings, multiple-hypothesis testing procedures, or the “pick-the-winner” empirical success rule that ignores arm-specific variance. None of these is decision-theoretically optimal when the goal is to select the one best policy among multiple candidates in a one-shot, multi-arm RCT. This lack of an optimal rule for multi-arm comparison has left the two-arm “A/B test” designs pervasive for one-shot RCTs, leading to inefficiencies in RCT designs and suboptimality in the resulting decisions.
We address this inefficiency by characterizing the decision-theoretically optimal rule for the best-population selection problem in one-shot, multi-arm RCTs. Our analysis is cast in the frequentist minimax-regret (MMR) framework, where regret is the opportunity cost from not choosing the best arm ex post, and the MMR criterion minimizes the worst-case expected regret over the parameter space. We first characterize the MMR-optimal rule for any multivariate location family reward distribution with fully-supported density that shifts its location-parameter vector but keeps the same shape and is positive everywhere. We then specialize to multivariate normal (MVN) rewards with an arbitrary covariance matrix. Lastly, when the mean-reward estimator is only (locally uniformly) asymptotically normal and covariances are estimated, we show that a plug-in version of the MVN-MMR rule is locally asymptotically minimax. The resulting Asymptotic MMR (AMMR) procedure maps a consistent covariance-matrix estimate directly into decision boundaries for the mean-reward estimates, making implementation straightforward in practice.
Our results reveal a sharp and practically important distinction between two-arm and multi-arm designs. With two arms, the empirical success rule remains MMR-optimal, and heterogeneity in variances across arms does not alter the decision threshold. With three or more arms and heterogeneous variances, however, the empirical success rule is no longer optimal. The MMR decision boundaries become nonlinear and penalize high-variance arms, requiring stronger evidence to select them. Intuitively, for the two-arm case, arm-specific variance is irrelevant because the problem reduces to comparing only the realized reward difference between the two arms. However, arm-specific variance becomes critical once multiple options are compared simultaneously, as the winning arm must be superior over all other arms with different arm-specific variances.
Formally, we study the problem of selecting the best arm after observing one noisy signal vector whose distribution belongs to a location family with an unknown location parameter vector. We adopt the regret loss and the frequentist minimax criterion following the MMR analyses in two-arm setups Manski2004,Manski2016,Manski2019,Manski2021,Hirano2009,Tetenov2012,Stoye2009,Stoye2012,HuZhuBrunskillWager2024,JooForthcoming. Our frequentist approach to the decision problem yields a plug-in rule that does not require specifying or estimating a subjective prior, facilitating immediate practical applications.
Directly solving for the MMR problem beyond the two-arm setup is typically intractable. To make progress, we recast the problem as a zero-sum game against nature, the dual formulation of the problem following Wald1945. In this dual formulation, nature chooses a least-favorable prior to which the DM responds with the Bayes rule. We prove that nature’s least-favorable prior places exactly one support point in each region where a particular arm is optimal. We further prove that the corresponding Bayes rule is deterministic, unique, and therefore unique minimax for the original MMR problem. These results drastically simplify the numerical characterization and computation of the MMR-optimal policy.
In applications, the exact reward distribution is rarely known. We therefore provide a proof that the MVN-MMR rule and its regret risk uniformly approximate their exact finite-sample counterparts when the mean-reward estimator is locally uniformly asymptotically normal, a mild assumption satisfied by many smooth M-estimators including GMM and MLE, and by several shrinkage/robust estimators. Consequently, one can plug $\sqrt{n}$-scaled mean-reward estimates into the MVN decision boundaries of a consistent covariance estimate. This plug-in rule leads to the AMMR rule that is optimal in the local asymptotic sense, without requiring asymptotic efficiency. This complements the approach of Hirano2009 and facilitates operationalization.
We solve the MVN-MMR problem numerically and plot the resulting MVN-MMR decision boundaries. With two arms, the boundaries are linear and coincide with the empirical success rule, regardless of the arm-specific variances. With three or more arms and non-homoskedastic covariance matrix, the optimal boundaries curve and shrink the region for high-variance arms, formalizing the intuitive idea that higher arm-specific uncertainty requires stronger evidence.
Taken together, our results demonstrate that the MMR decision theory can directly inform regret-optimal decision-making using one-shot, multi-arm RCTs. Therefore, our multi-arm AMMR allows researchers and practitioners to compare multiple competing policies by designing and implementing multi-arm RCTs within a unified decision-theoretic framework, moving beyond the pervasive two-arm “A/B test” designs.
We are, to our knowledge, the first to characterize the general minimax-risk-optimal best-population selection rule under regret loss in one-shot, multi-arm experiments. Classic decision-theory results bahadur1950problem,bahadur1952impartial,lehmann1966theorem,eaton1967some show that the empirical success rule is minimax, admissible, and Bayes under permutation invariance, a symmetry that effectively requires homoskedastic rewards in the MVN case and certain discrete, permutation-invariant priors. In most applications, covariance matrices are heteroskedastic, breaking permutation invariance and eliminating the optimality properties of the empirical success rule. Indeed, later work shows that empirical success is not minimax under heteroskedasticity (Dhariyal1994 under 0-1 loss; Masten2023 under regret loss). The minimax rule in the general heteroskedastic case had remained unknown, with only partial results available for Bayes rules under special priors abughalous1995selecting. A common workaround is balanced design bechhofer1954single,somerville1954some, which equalizes realized arm-specific variances via sample allocation. However, balanced design cannot be applied to one-shot multi-arm RCTs because it typically requires pilot experiments or sequential experimentation.
Our problem is conceptually related to best-arm identification in pure-exploration bandits, where regret-based objectives and explore-then-commit designs are standard Audibert2010,Jamieson2014,Garivier2016,Kaufmann2016,Grover2018,Agrawal2019,Russo2020,Komiyama2021,Howard2021,Kasy2021,AdusumilliRESforthcoming. However, sequential bandit experimentation is often infeasible in RCT settings with delayed outcomes or when evaluating multiple policies retrospectively using completed trials. Even when feasible, bandit algorithms can be operationally demanding. Our AMMR rule for one-shot RCTs is therefore complementary and broadly applicable.
Our multi-arm characterization nests the two-arm MMR problem of Tetenov2012: the optimal rule in two-arm experiments is the empirical success threshold of zero, which our dual characterization reproduces. While directly maximizing the regret function as in Tetenov does not scale to multiple arms, our dual characterization of the multi-arm MMR-optimal policy that is a deterministic Bayes rule under the least-favorable prior restores computational tractability. Recent numerical approximation methods such as Fernandez2025,Guggenberger2025 could further improve computational efficiency.
Our local asymptotics and the associated plug-in approach are similar in spirit to Hirano2009. Hirano2009 derive the local asymptotic minimax bounds using Le Cam’s limits-of-experiments framework and show that rules based on asymptotically efficient (“best regular”) estimators attain those bounds, thereby justifying the plug-in of the efficient estimators into the optimal decision rules derived for normally distributed rewards. By contrast, our asymptotic approach is to begin with a mean-reward estimator that is locally uniformly asymptotically normal, and directly establish the local uniform convergence of both the risk function and the associated MMR decision rule. As a result, our plug-in approach can be broadly applied to smooth GMM/MLE and many shrinkage estimators that may not be asymptotically efficient.
Other related works include Kitagawa2022, who study least-favorable priors when regret is nonlinearly transformed, and broader ranking/subset-selection problems GuKoenker2023. We also do not address covariate-based treatment assignment or its multi-arm extensions Stoye2009,Stoye2012,Kitagawa2018,Athey2021.
$ $
The remainder of the article is organized as follows. Section (ref) characterizes the general DM's MMR-optimal decision rule when rewards follow a fully-supported multivariate location family. Section (ref) applies the general result to the MVN-distributed rewards and characterizes the properties of MVN-MMR decision rule. Section (ref) provides the asymptotic approximation results when the rewards are only asymptotically MVN distributed. Section (ref) provides numerical examples. Section (ref) concludes.
In this section, we characterize the DM's MMR-risk-optimal rule for the best-population selection problem. We define the best-population selection problem, regret loss function, and the minimax-risk solution concept. Because directly characterizing the DM’s MMR-risk-optimal strategy is intractable, we formulate the dual problem as a two-player zero-sum game against nature. We show that the game has a value and that a minimax theorem applies. Then we characterize nature's least-favorable prior in the dual maximin problem, and solve for the DM's optimal strategy under the least-favorable prior. The resulting DM's optimal strategy in the dual problem coincides with the DM's minimax decision rule in the primal problem.
\paragraph*{Setup, Notation, and Assumptions}
Let $\mathcal{J}=\left\{ 1,2,...,J\right\} $. We use boldfaces to denote vectors throughout. Denote the parameter set by $\Theta=\left[-B,B\right]^{J}\subset\mathbb{R}^{J}$, a compact hypercube. $B>0$ can be arbitrarily large but finite. We assume the true parameter vector $\bm{\theta}$ lies in the interior of $\Theta$.
For any $\bm{\theta}\in\Theta$ let $i^{*}\left(\bm{\theta}\right)=\min\left\{ l:\theta_{l}=\max_{j\in\mathcal{J}}\theta_{j}\right\} $, the smallest index such that the component $\theta_{l}$ attains $\max_{j\in\mathcal{J}}\left\{ \theta_{j}\right\} $. Define $\Theta^{i}=\left\{ \bm{\theta}\in\Theta:i^{*}\left(\bm{\theta}\right)=i\right\} $. $\Theta^{i}$ is the union of the set $\left\{ \bm{\theta}\in\Theta:-B\leq\theta_{j}<\theta_{i}\leq B\quad\forall j\neq i\right\} $ and certain boundary surfaces of ties, where $\Theta^{1}$ contains all the tie points of the form $\max_{k}\left\{ \theta_{k}\right\} =\theta_{1}=\theta_{j}$ for $j>1$ whereas $\Theta^{J}$ does not include any ties in the largest element. By construction, $\left\{ \Theta^{i}\right\} _{i\in\mathcal{J}}$ partitions $\Theta$, i.e., $\Theta=\bigcup_{i\in\mathcal{J}}\Theta^{i}$ and $\Theta^{i}\cap\Theta^{j}=\emptyset$ for any $i\neq j$.
We use superscript $i$ to designate that the true parameter belongs to $\Theta^{i}$, and subscript $i$ to denote $i$-th element of a vector. For instance, $\theta_{J}^{J}>\theta_{j}^{J}$ for all $j\neq J$ by construction because $\bm{\theta}^{J}\in\Theta^{J}$ and $\Theta^{J}$ does not include ties in the largest element. For the dual problem introduced later, we denote by $\pi^{i}$ the prior probability of $\Theta^{i}$, and by $\bm{\pi}=\left(\pi^{i}\right)_{i\in\mathcal{J}}$ the collection of these prior probabilities. We denote by $\pi\equiv\left\{ \pi^{i},\bm{\theta}^{i}\right\} _{i\in\mathcal{J}}$ a discrete prior distribution $\pi$ supported on $\left\{ \bm{\theta}^{j}\right\} _{j\in\mathcal{J}}$.
Let $\left\{ P_{\bm{\theta}}\right\} _{\bm{\theta}\in\Theta}$ be a multivariate location family. We assume each $P_{\bm{\theta}}$ is a probability measure on $\mathbb{R}^{J}$ dominated by the Lebesgue measure $\mu$, the densities $p_{\bm{\theta}}$ are continuous and fully supported on $\mathbb{R}^{J}$, and $p_{\bm{\theta}}\left(\mathbf{x}\right)\neq p_{\bm{\theta}'}\left(\mathbf{x}\right)$ $\mu$-a.e. whenever $\bm{\theta}\neq\bm{\theta}'$. When $\bm{\theta}^{i}\in\Theta^{i}$ for each $i$, we use the shorthand $p^{i}=p_{\bm{\theta}^{i}}$.
The best-population selection problem (the selection problem) is finding the largest element (arm) of the location-parameter vector $\bm{\theta}$ after observing $\mathbf{x}\sim P_{\bm{\theta}}$. A decision rule for the selection problem is a (measurable) mapping $\bm{\phi}:\mathbb{R}^{J}\rightarrow\bm{\Delta}$, where the codomain $\bm{\Delta}=\left\{ \bm{\phi}\in\mathbb{R}^{J}:\sum_{j=1}^{J}\phi_{j}=1\right\} $ is the probability simplex of $\mathbb{R}^{J}$. $\phi_{j}\left(\mathbf{x}\right)$ is the probability of selecting arm $j$ when an observation $\mathbf{x}$ is realized. Note that the decision rule $\bm{\phi}$ permits randomization as $\bm{\phi}\in\text{int}\left(\bm{\Delta}\right)$ is possible. Note, a nonrandomized decision rule can take values only among the vertices $\left\{ \mathbf{e}_{i}\right\} _{i\in\mathcal{J}}$ of the simplex $\bm{\Delta}$, where $\mathbf{e}_{i}$ is the $i$-th unit vector.
\paragraph*{Loss, Risk, and the Solution Concept}
The regret loss with the realization of $\mathbf{x}$ under $\bm{\phi}$ is defined as:
The risk of a decision rule $\bm{\phi}$ is the expected loss taking $\bm{\theta}$ as given, where the expectation is taken with respect to the measure $dP_{\bm{\theta}}\left(\mathbf{x}\right)=p_{\bm{\theta}}\left(\mathbf{x}\right)d\mu\left(\mathbf{x}\right)$:
The regret risk is continuous in $\bm{\theta}$ for any decision rule $\bm{\phi}$ under our setup (see Lemma (ref)).
The DM's objective is to find the MMR-risk-optimal decision rule $\bm{\delta}$ that minimizes the worst-case risk over the parameter space $\Theta$:
where $R\left(\bm{\theta},\bm{\phi}\right)$ is defined by ((ref)) and ((ref)).
Directly characterizing the minimax-risk decision rule is intractable except for a few special cases Manski2004,Stoye2009,Stoye2012,Kitagawa2022. Therefore, in what follows, we turn to the “dual” maximin problem of a two-player zero-sum game between the DM and nature to characterize the DM’s optimal decision rule.
In this subsection, we first establish the existence of nature’s least-favorable prior $\pi$ for the selection problem under regret loss, which is supported by at most $J$ distinct points in $\mathbb{R}^{J}$ (Theorem (ref)). Next, we characterize the form of the Bayes rule under this least-favorable prior, which is deterministic up to tie-breaking (Lemma (ref)). Then we establish that the support points of the least-favorable prior are exactly $J$ distinct points, with $\bm{\theta}^{i}\in\Theta^{i}$ for each $i$ and the corresponding prior probability vector $\left(\pi^{i}\right)_{i\in\mathcal{J}}\in\text{int}\left(\bm{\Delta}\right)$ (Theorem (ref)). Taken together, the results established in this subsection simplify the search for nature's least-favorable prior in the dual problem drastically.
Parts (i) and (iii) follow from the standard minimax theorem for statistical decision problems with finite action sets under bounded loss and dominated experiments (e.g., Blackwell1954;Berger1985;Liese2008).
Part (ii) asserts the existence of a least-favorable prior supported on at most $J$ distinct points. The geometric intuition behind this is as follows. Define the $J$-dimensional risk vector $\mathbf{R}\left(\bm{\theta}\right)\equiv\left(\max_{k\in\mathcal{J}}\left\{ \theta_{k}\right\} -\theta_{j}\right)_{j\in\mathcal{J}}=\left(R\left(\bm{\theta},\mathbf{e}_{j}\right)\right)_{j\in\mathcal{J}}$ and the risk set $S\equiv\left\{ \left(\max_{k\in\mathcal{J}}\left\{ \theta_{k}\right\} -\theta_{j}\right)_{j\in\mathcal{J}}:\bm{\theta}\in\Theta\right\} $. The convex hull $\text{co}\left(S\right)=\left\{ \int_{\Theta}\mathbf{R}\left(\bm{\theta}\right)d\pi\left(\bm{\theta}\right):\pi\text{ }\text{is Borel prob. measure over }\Theta\right\} $ collects the risk vectors induced by priors. For any mixed action $\mathbf{a}\in\bm{\Delta}$, nature's worst-case risk is $\sup_{\mathbf{u}\in\text{co}\left(S\right)}\mathbf{a}^{\top}\mathbf{u}$, so the value is $V_{S}\equiv\min_{\mathbf{a}\in\bm{\Delta}}\sup_{\mathbf{u}\in\text{co}\left(S\right)}\mathbf{a}^{\top}\mathbf{u}=\sup_{\mathbf{u}\in\text{co}\left(S\right)}\min_{\mathbf{a}\in\bm{\Delta}}\mathbf{a}^{\top}\mathbf{u}$. The maximizer $\mathbf{u}^{*}$ lies on a supporting hyperplane $\left\{ \mathbf{u}\in\text{co}\left(S\right):\mathbf{a}^{*\top}\mathbf{u}=V_{S}\right\} $, a $J-1$ dimensional affine subspace. Hence, Caratheodory's theorem yields a representation with at most $J$ support points. While the minimax/Bayes rule is a data-dependent Markov kernel $\bm{\phi}:\mathbb{R}^{J}\rightarrow\bm{\Delta}$, the forgoing $S$-game geometry is used to control nature's side of the problem, justifying the finite-support property; it does not restrict $\bm{\phi}$ to ignore the data.
Part (iv) can either be imposed as a constraint in finding the least-favorable prior numerically, or can be checked after a candidate least-favorable prior is found. We come back to this in Section (ref). We also note that Theorem (ref) (iv) under a slightly different set of assumptions can be found in Corollary to Theorem 2.2 of Kempthorne1987.
We next characterize the form of the Bayes rule under a finite prior, which is deterministic up to tie-breaking that can be chosen arbitrarily. For the following lemma and the proof, we use the superscript $\left(k\right)$ to denote the $k$-th element of the support point set $\left\{ \bm{\theta}^{\left(k\right)}\right\} _{k=1}^{m}$ and the prior probability set $\left\{ \pi^{\left(k\right)}\right\} _{k=1}^{m}$, where $\bm{\theta}^{\left(k\right)}$ may not necessarily belong to $\Theta^{k}$.
The following Theorem (ref) establishes that nature’s least-favorable prior is supported by exactly $J$ distinct points $\left\{ \bm{\theta}^{i}\right\} _{i\in\mathcal{J}}$ with $\bm{\theta}^{i}\in\Theta^{i}$, and with strictly positive prior probabilities for each arm. In turn, this implies the prior probabilities $\bm{\pi}\in\text{int}\left(\bm{\Delta}\right)$.
In Theorem (ref), we characterize the DM's Bayes decision rule $\bm{\delta}$ under the form of nature’s least-favorable prior established in Theorems (ref)-(ref), supported on $J$ distinct points with $\bm{\theta}^{j}\in\Theta^{j}$. Under a mild regularity condition on the multivariate location-family density $p_{\bm{\theta}^{k}}=p^{k}\left(\mathbf{x}\right)$, $\bm{\delta}$ is the unique Bayes rule and the unique minimax rule, implying that $\bm{\delta}$ is admissible. The unique MMR-risk rule $\bm{\delta}$ is deterministic up to a Lebesgue-null set, characterized by simple pointwise comparisons of weighted densities.
The additional condition $\mu\left(\left\{ \mathbf{x}:h_{ij}\left(\mathbf{x}\right)=0\right\} \right)=0$ imposed for (ii)-(iv) is very mild. This is imposed to rule out pathological behavior of the density $p^{k}\left(\cdot\right)$ having a strictly positive flat region. Sufficient conditions include $p^{k}\left(\cdot\right)$ being real analytic or being $C^{1}$ with $\nabla h_{ij}\left(\mathbf{x}\right)\neq\mathbf{0}$ whenever $h_{ij}\left(\mathbf{x}\right)=0$. Most importantly, the condition is satisfied for the MVN density that we analyze in detail below in Sections (ref)-(ref), because each $\pi^{k}\left(\theta_{i}^{k}-\theta_{j}^{k}\right)p^{k}\left(\mathbf{x}\right)$ is real analytic, $\left\{ p^{k}\left(\mathbf{x}\right)\right\} _{k\in\mathcal{J}}$ are linearly independent, and $\pi^{k}\left(\theta_{i}^{k}-\theta_{j}^{k}\right)$ is not identically zero for all $k$. We formalize the MVN case in the following Corollary.
$\left\{ p_{\bm{\theta}}\right\} _{\bm{\theta}\in\Theta}$ is a multivariate location family and the resulting decision problem is location invariant. Specifically, $\forall\mathbf{x},\bm{\theta}\in\mathbb{R}^{J}$ and $\forall t\in\mathbb{R}$, let $\mathbf{x}'=\mathbf{x}-t\mathbf{1}_{J}$ and $\bm{\theta}'=\bm{\theta}-t\mathbf{1}_{J}$. Then, $p_{\bm{\theta}}\left(\mathbf{x}\right)=p_{\mathbf{0}}\left(\mathbf{x}-\bm{\theta}\right)=p_{\bm{\theta}'}\left(\mathbf{x}'\right)$, $\bm{\delta}\left(\mathbf{x}'\right)=\bm{\delta}\left(\mathbf{x}\right)$, $L\left(\bm{\delta}\left(\mathbf{x}'\right),\bm{\theta}'\right)=L\left(\bm{\delta}\left(\mathbf{x}\right),\bm{\theta}\right)$, and $R\left(\bm{\theta}',\bm{\delta}\right)=R\left(\bm{\theta},\bm{\delta}\right)$.\cprotect\footnote{The parameter space $\Theta=\left[-B,B\right]^{J}$ is a compact set in $\mathbb{R}^{J}$. Therefore, technically, the transformation $\bm{\theta}'=\bm{\theta}-t\mathbf{1}$ is not a location group defined on $\Theta$. At the cost of complicating the notations, we may define the extended parameter space $\tilde{\Theta}=\mathbb{R}^{J}$ and define the value of the loss function on $\tilde{\Theta}\backslash\Theta$ as $2B$
. However, this upper bound is never attained under the DM's MMR-risk rule as long as $B$ is large enough.}
This implies that if $\left(\pi^{k},\bm{\theta}^{k}\right)_{k\in\mathcal{J}}$ is a least-favorable prior, then $\left(\pi^{k},\bm{\theta}^{k}-t\mathbf{1}_{J}\right)_{k\in\mathcal{J}}$ is also a least-favorable prior, i.e., the support points can slide along the subspace $\left\{ t\mathbf{1}_{J}:t\in\mathbb{R}\right\} $. To pin down the support points during implementation, either $\sum_{i=1}^{J}\theta_{i}^{j}=0$ or $\theta_{J}^{j}=0$ can be imposed for some fixed $j$.
In this section, we apply the results of Section (ref) to the important special case of MVN-distributed rewards with an arbitrary positive definite (p.d.) covariance matrix. We establish two main results for the MVN selection problem in this section. First result, Proposition (ref), is the $\sqrt{n}$-scaling of the value and the minimax decision rule. Second result, Proposition (ref), is the continuity of the MVN-MMR decision rule and value in the covariance matrix, respectively. These results will be used to establish the large-sample approximation results of the exact MMR problem in the subsequent section.
The definition of the MVN density with parameters $\left(\bm{\theta},\bm{\Sigma}\right)$ is:
In the MVN selection problem, the p.d. covariance matrix $\bm{\Sigma}$ is known to the DM whereas the location-parameter vector $\bm{\theta}$ is unknown. The goal of the MVN selection problem is to select the largest element of $\bm{\theta}$ after observing $\mathbf{x}\sim\mathcal{N}\left(\bm{\theta},\bm{\Sigma}\right)$.
In the remainder of the paper, we index the objects of interest by $\bm{\Sigma}$ when necessary: we denote by $\pi_{\bm{\Sigma}}=\left(\pi_{\bm{\Sigma}}^{k},\bm{\theta}_{\bm{\Sigma}}^{k}\right)_{k\in\mathcal{J}}$ the least-favorable prior under $\bm{\Sigma}$, by $R_{\bm{\Sigma}}\left(\cdot,\cdot\right)$ the Gaussian risk function, by $\bm{\delta}_{\bm{\Sigma}}$ the associated a.e.-unique MMR rule, and by $V_{\bm{\Sigma}}$ the associated value.
The following Proposition (ref) establishes the $\sqrt{n}$-scaling of the MVN selection problem. For Proposition (ref) and its proof, we use subscript $\left[\cdot\right]$ to denote the dependence of the covariance matrix on the sample size. Specifically, let $\bm{\Sigma}_{\left[1\right]}$ be a positive-definite covariance matrix for $n=1$, and for any $n\geq1$, denote by $\bm{\Sigma}_{\left[n\right]}=\frac{1}{n}\bm{\Sigma}_{\left[1\right]}$. Let $\Theta_{\bm{\Sigma}_{\left[1\right]}}=\Theta=\left[-B,B\right]^{J}$ be the parameter space for $n=1$.
Let $\mathcal{S}=\left\{ \bm{\Sigma}\in\mathbb{S}_{++}^{J}:\text{\underbar{\ensuremath{\lambda}}}\mathbf{I}_{J}\preceq\bm{\Sigma}\preceq\bar{\lambda}\mathbf{I}_{J}\right\} $ be a compact set of p.d. covariance matrix for some $0<\underbar{\ensuremath{\lambda}}<\bar{\lambda}<\infty$, where $\mathbf{A}\preceq\mathbf{B}$ means $\mathbf{B}-\mathbf{A}$ is p.s.d. and $\mathbf{I}_{J}$ is the $J$-dimensional identity matrix. $\underbar{\ensuremath{\lambda}}$ ($\bar{\lambda}$) can be arbitrarily small (large) but has to be finite. For the remainder of the paper, we assume the covariance matrix $\bm{\Sigma}\in\mathcal{S}$. The following Proposition (ref) establishes continuity of the values $V_{\bm{\Sigma}}$ and the MMR rule $\bm{\delta}_{\bm{\Sigma}}\left(\mathbf{x}\right)$ in $\bm{\Sigma}$.
In this section, we establish that the finite-sample regret function and the associated MMR-risk-optimal decision rule can be asymptotically approximated by the MVN counterpart studied in Section (ref). We assume that the estimator $\hat{\bm{\theta}}_{n}$ for $\bm{\theta}$ is $\sqrt{n}$-consistent and locally uniformly asymptotically Normal, and a consistent estimator $\hat{\bm{\Sigma}}$ for the covariance matrix $\bm{\Sigma}$ is available. Then, Theorems (ref)-(ref) establish that $\hat{\bm{\theta}}_{n}$ can be plugged into the MMR decision rule $\bm{\delta}\left(\cdot\right)$ to select the best arm as if $\hat{\bm{\theta}}_{n}\sim\mathcal{N}\left(\bm{\theta},\frac{1}{n}\hat{\bm{\Sigma}}\right)$ when $n$ is large.
We begin this subsection by imposing a regularity assumption on the decision rule $\bm{\phi}$ being considered.
\begingroup
\endgroup
Assumption (ref) is a restriction imposed on the decision rule $\bm{\phi}$ to rule out pathological behavior, such as $\bm{\phi}$ being discontinuous on sets of positive Lebesgue measure. This assumption does not rule out randomized decision rules a priori. Importantly, we note that the DM's minimax decision rule characterized in Theorem (ref) satisfies this assumption.
Next, we introduce the local uniform CLT assumption and scaling of the problem. We assume the true parameter $\bm{\theta}$ lies in the interior of $\Theta$. For every $\bm{\theta}\in\text{int}\left(\Theta\right)$ and for any local parameter vector $\mathbf{h}$ such that $\left\Vert \mathbf{h}\right\Vert \leq H<\infty$, we consider a locally shifted parameter sequence $\bm{\theta}+\frac{\mathbf{h}}{\sqrt{n}}$ throughout.
\begingroup
\endgroup
Assumption (ref) says the shifted and scaled reward estimator $\sqrt{n}\left(\hat{\bm{\theta}}_{n}-\bm{\theta}\right)$ is asymptotically normal uniformly in $\bm{\theta}$ over local neighborhoods of radii shrinking at the $\frac{1}{\sqrt{n}}$ rate. Heuristically, the assumption can be understood as $\sqrt{n}\left(\hat{\bm{\theta}}_{n}-\bm{\theta}\right)\stackrel{\mathbf{h}}{\rightsquigarrow}\mathcal{N}\left(\mathbf{h},\bm{\Sigma}\right)$
uniformly in $\mathbf{h}$ such that $\left\Vert \mathbf{h}\right\Vert \leq H<\infty$ for some $H>0$. When $\mathbf{h}=\mathbf{0}$, it reduces to the usual pointwise weak convergence in $\bm{\theta}$: $\mathbf{u}_{\bm{\theta},n}\equiv\sqrt{n}\left(\hat{\bm{\theta}}_{n}-\bm{\theta}\right)\stackrel{}{\rightsquigarrow}\mathcal{N}\left(\mathbf{0},\bm{\Sigma}\right)\equiv\mathbf{u}_{\bm{\theta}}$. This assumption is weaker than imposing the uniform CLT over the original parameter space $\Theta$ and it is the minimal condition required for the usual asymptotic power comparison to be valid vanderVaart1998. Regular (i.e., locally asymptotically shift-equivariant) estimators satisfy this assumption, including but not limited to the general class of MLE and GMM estimators under suitable regularity conditions Newey1994,vanderVaart1998.
In the following, we localize at $\bm{\theta}+\frac{\mathbf{h}}{\sqrt{n}}$ centered at $\bm{\theta}=t\mathbf{1}_{J}$, and set $t=0$ WLOG by invoking the location invariance discussed in Section (ref). $\bm{\theta}=t\mathbf{1}_{J}$ is where the competing arms are hardest to distinguish; if we localize at any other point $\bm{\theta}\neq t\mathbf{1}_{j}$, the arm that is best at $\bm{\theta}$ will be best for all the local alternatives $\bm{\theta}+\frac{\mathbf{h}}{\sqrt{n}}$ as the sample size grows, so the selection problem becomes trivial. The same localization is standard for the two-arm problem (e.g., Hirano2009; vanderVaart1998; Liese2008), which we are adapting to the multi-arm best-population selection problem.
Next, we scale the estimator $\hat{\bm{\theta}}_{n}$ under the local parameter sequence $\bm{\theta}_{n}\left(\bm{\vartheta}\right)\equiv\frac{\bm{\vartheta}}{\sqrt{n}}$. Let \[ \mathbf{y}_{n}=\sqrt{n}\hat{\bm{\theta}}_{n}\quad\text{and}\quad\mathbf{u}_{n}\left(\bm{\vartheta}\right)=\mathbf{y}_{n}-\bm{\vartheta}=\sqrt{n}\left(\hat{\bm{\theta}}_{n}-\bm{\theta}_{n}\left(\bm{\vartheta}\right)\right), \] and we omit the argument $\left(\bm{\vartheta}\right)$ when it is obvious from context. Denote by $P_{\bm{\vartheta},n}$ the law of $\mathbf{y}_{n}$, by $Q_{\bm{\vartheta},n}$ the law of $\mathbf{u}_{n}$. The decision rule $\bm{\phi}:\mathbb{R}^{J}\rightarrow\bm{\Delta}$ now takes $\mathbf{y}_{n}$ as its argument, matching the $\sqrt{n}$-scaling in Proposition (ref).
The following Proposition (ref) establishes that the CLT holds uniformly over the original parameter space $\Theta$ when the estimator $\hat{\bm{\theta}}_{n}$ is properly centered and “inflated” by $\sqrt{n}$, a direct consequence of Assumption (ref).
We establish our main Gaussian approximation results in this subsection. As before, we index the MVN-MMR rule $\bm{\delta}$, the Gaussian regret risk function $R\left(\bm{\vartheta},\bm{\delta}\right)$, and the minimax value $V$ by the covariance matrix $\bm{\Sigma}$ or its consistent estimate $\hat{\bm{\Sigma}}$, respectively. Denote by $P_{\bm{\vartheta},\bm{\Sigma}}$ the Gaussian measure $\mathcal{N}\left(\bm{\vartheta},\bm{\Sigma}\right)$, and for any decision rule $\bm{\phi}$ and ($\sqrt{n}$-inflated local) parameter $\bm{\vartheta}$, define:
Theorem (ref) below establishes, for any fixed rule $\bm{\phi}$, $R_{\left[n\right]}\left(\bm{\vartheta},\bm{\phi}\right)$ converges to $R_{\bm{\Sigma}}\left(\bm{\vartheta},\bm{\phi}\right)$ uniformly over $\bm{\vartheta}\in\Theta$. It immediately follows the MVN-MMR rule $\bm{\delta}_{\bm{\Sigma}}$ plugged into the Gaussian risk function achieves the minimax value $V_{\bm{\Sigma}}$. However, in practice, the covariance matrix $\bm{\Sigma}$ is unknown so the rule $\bm{\delta}_{\bm{\Sigma}}$ is infeasible. A feasible version of the decision rule is $\bm{\delta}_{\hat{\bm{\Sigma}}}$, which is the MVN-MMR decision rule obtained by plugging-in the consistent estimate $\hat{\bm{\Sigma}}$. Therefore, the subsequent Theorem (ref) establishes that the feasible version $\bm{\delta}_{\hat{\bm{\Sigma}}}$
also achieves the MVN minimax-risk value asymptotically.
Taken together, these two theorems justify plugging $\sqrt{n}\hat{\bm{\theta}}_{n}$ into the MVN decision boundaries derived using the consistent estimate $\hat{\bm{\Sigma}}$ (equivalently, plugging $\hat{\bm{\theta}}_{n}$ into the decision boundaries for $\frac{1}{n}\hat{\bm{\Sigma}}$) when $\hat{\bm{\theta}}_{n}$ is $\sqrt{n}$-consistent and (locally uniformly) asymptotically MVN.
We next prove a lemma that upgrades the pointwise (in $\bm{\phi}$) convergence of the risk function to the uniform convergence when we restrict the class of decision rules to MVN-MMR rules indexed by the covariance matrices $\bm{\Gamma}\in\mathcal{S}$. Note, this class of rules satisfy Assumption (ref) by Theorem (ref) and its Corollary.
Finally, the following Theorem (ref) establishes the plug-in MVN rule $\bm{\delta}_{\hat{\bm{\Sigma}}}$ also achieves $V_{\bm{\Sigma}}$ asymptotically.
In this section, we describe our numerical optimization approach for finding the MVN-MMR decision rule and report decision boundaries for several examples. In Section (ref), we briefly describe how we carry out the numerical optimization to find the least-favorable prior, which hinges on the theoretical results from Sections (ref)-(ref). In Section (ref), we illustrate the decision boundaries for the two-arm and three-arm MVN-MMR problems under homoskedasticity and an arbitrary covariance matrix, respectively.
Through the numerical analysis, we reconfirm that the decision boundaries are linear and that the “pick-the-winner” empirical success rule is optimal either (i) in two-arm experiments or (ii) when the covariance matrix is homoskedastic in multi-arm experiments. However, when the covariance matrix is not homoskedastic in multi-arm experiments, the problem is no longer permutation invariant and the decision boundaries become nonlinear; in this case the empirical success rule is no longer optimal. The optimal decision rule penalizes higher-variance arms, requiring stronger evidence to select a high-variance arm.
We characterized the least-favorable prior $\left(\pi^{k},\bm{\theta}^{k}\right)_{k\in\mathcal{J}}$ in the dual maximin problem, which is supported by $J$ distinct points with exactly one point in $\Theta^{j}$. The Bayes rule $\bm{\delta}$ under this least-favorable prior takes the deterministic form (Theorem (ref)). The form of $\bm{\delta}$ is tractable enough that it can be directly plugged into the Bayes risk. Therefore, we numerically solve nature's problem of maximizing the Bayes risk under $\bm{\delta}$:
In the operationalization, we use $2^{19}$ Sobol quasi Monte Carlo draws to approximate the Gaussian integral in ((ref)). We then use two algorithms to solve the maximization problem: (i) direct maximization using the derivative-free subplex algorithm; and (ii) a softmax smoothing of the indicator function inside the Gaussian integral combined with a derivative-based optimizer, supplying closed-form derivatives. The first approach works better with a small problem dimension $\left(J\leq4\right)$, whereas the second approach works better with modest parameter dimensions $\left(J\geq5\right)$. We confirm that both algorithms converge to the same least-favorable prior when they converge.
Because the objective function has many near-optimal points, we use many randomized starting points to increase the chance that the algorithm reaches the global maximum. The theoretical prediction from Theorem (ref) (iv) that $R\left(\bm{\theta}^{k},\bm{\delta}\right)=V$ for all $k$ serves as a diagnostic. We relegate the details of the numerical optimization to Appendix (ref).
It is known that the problem with $J=2$ can be reduced to a one-dimensional “testing” problem with $x_{2}-x_{1}$, and the MMR-optimal threshold for the difference $x_{2}-x_{1}$ is $0$ Tetenov2012,JooForthcoming. In Tetenov2012,JooForthcoming, the optimal decision threshold in the two-arm case is derived by directly maximizing the Gaussian risk over the parameter-coordinate difference $\theta_{2}-\theta_{1}$. In the following, we find that our decision boundaries, the minimax value, and the locations of the least-favorable prior’s support points calculated from the dual maximin problem exactly coincide with those reported in Tetenov2012,JooForthcoming.
For the two-arm experiments, we let \[ \bm{\Sigma}_{sym}=
,\quad\bm{\Sigma}_{asym}=
. \] For $\bm{\Sigma}_{\text{sym}}$, we have $\Sigma_{11}+\Sigma_{22}=1$ and the minimax value is known to be $0.170$ (Tetenov2012,JooForthcoming), which we confirm in the Bayes risk column of Table (ref). In both Tables (ref) and (ref), the risk values attained at the support points of the least-favorable prior are identical up to rounding error, reconfirming the theoretical prediction of Theorem (ref) (iv).
In Figure (ref), the decision boundaries align exactly with the $45^{\circ}$ line, corresponding to the empirical success rule. Therefore, we also reconfirm that our dual formulation of the problem yields the same decision boundary for both the symmetric and asymmetric cases. The asymmetric arm-specific variance does not affect the decision boundaries nor the prior probabilities in this two-arm case---the location of the least-favorable prior’s support points adjusts to equalize the risk values attained at those points.
A more interesting case is $J\geq3$, where homoskedasticity no longer holds. We illustrate this case for two three-arm experiments. Let the covariance matrices be as follows: \[ \bm{\Sigma}_{sym}=
,\quad\bm{\Sigma}_{asym}=
. \] Arm 3 of $\bm{\Sigma}_{\text{asym}}$ exhibits higher variance than the other two arms.
In Table (ref), we report the least-favorable prior’s support points with their prior probabilities and the associated risk values evaluated at those points. We reconfirm the conclusion of Theorem (ref) (iv) that the risk values attained are the same at all three support points of the least-favorable prior up to a rounding error. Unlike the two-arm experiment, however, both the prior probabilities and the support points of the least-favorable prior are no longer symmetric across the arms.
The decision boundaries are illustrated in Figure (ref) for the symmetric covariance matrix $\bm{\Sigma}_{\text{sym}}$ and in Figure (ref) for the asymmetric covariance matrix $\bm{\Sigma}_{\text{asym}}$. Each subfigure illustrates the decision boundaries for two coordinates, fixing the remaining coordinate at $\left(-2,0,2\right)$. We mark the $45^{\circ}$ line in the first quadrants as a benchmark representing the empirical success rule's boundary.
For the symmetric covariance matrix $\bm{\Sigma}_{\text{sym}}$, we find that the boundaries are linear and the empirical success rule is optimal, as seen in the middle panels of Figures (ref)-(ref). As reported in Table (ref), the least-favorable prior’s support points are symmetric with respect to the origin, with prior probabilities of 1/3, up to rounding error.
For the asymmetric covariance matrix $\bm{\Sigma}_{\text{asym}}$, we find that the empirical success rule is no longer optimal, as seen in the first quadrants of the middle panels of Figures (ref)-(ref). The decision boundaries deviate from the $45^{\circ}$ line. The areas of $\bm{\delta}\left(\mathbf{x}\right)=\mathbf{e}_{3}$ shrink, implying stronger evidence is required to select the higher-variance arm. Interestingly, both the support points and the prior probabilities of the least-favorable prior are no longer symmetric, as reported in Table (ref).
With $J=2$, the joint shift invariance discussed in Section (ref) eliminates the nuisance location $t\mathbf{1}_{2}$ and reduces the experiment to a single contrast $x_{2}-x_{1}$ with parameter $\theta_{2}-\theta_{1}$. Under any location family, $x_{2}-x_{1}|\theta_{2}-\theta_{1}$ is a one-dimensional location model with mean $\theta_{2}-\theta_{1}$ and variance $Var\left(x_{2}-x_{1}\right)$ that depends on the original covariance matrix $\bm{\Sigma}$ but not on the common shift. Regret for any rule depends on the location-parameter vector $\bm{\theta}$ only through $\theta_{2}-\theta_{1}$, and on the data only through $x_{2}-x_{1}$. Therefore, the MMR-risk problem collapses to a one-dimensional sign decision with symmetric noise, for which the unique MMR rule is a threshold in $x_{2}-x_{1}$. Equalization of the risk at the support points (Theorem (ref) (iv)) forces the threshold for $x_{2}-x_{1}$ to be 0. In the original $\left(x_{1},x_{2}\right)$ plane, this is precisely the $45^{\circ}$ bisector $x_{1}=x_{2}$. This is exactly what we see in Table (ref) and Figure (ref).
For $J\geq3$, the same invariance leaves a $\left(J-1\right)$-dimensional vector of contrasts $\left(x_{2}-x_{1},...,x_{J}-x_{1}\right)$ with a joint Gaussian law whose covariance inherits all heteroskedasticity and correlations from $\bm{\Sigma}$. Deciding that arm $i$ is the best requires simultaneously beating $J-1$ correlated comparisons (e.g., arm\,3 must beat both arms\,1 and\,2), so there is no single one-dimensional sufficient direction and no symmetry that pins each pairwise threshold to zero. In the dual problem, the $i$-$j$ boundary solves \[ h_{ij}\left(\mathbf{x}\right)\equiv\sum_{k=1}^{J}\pi^{k}p_{\bm{\theta}^{k}}\left(\mathbf{x}\right)\left(\theta_{i}^{k}-\theta_{j}^{k}\right)=0 \] (see Theorem (ref)). With two arms, $h_{12}\left(\mathbf{x}\right)$ contains two exponential terms and rearranges to a single linear condition, hence a straight line. With three or more arms, $h_{ij}\left(\mathbf{x}\right)$ is a sum of at least three exponentials with mixed signs, so the zero-level set is generically curved. When $\bm{\Sigma}$ is homoskedastic/permutation-invariant, symmetry still forces linear boundaries with the empirical success rule (Figure (ref) and Table\,(ref)). However, when $\bm{\Sigma}$ breaks permutation invariance, nothing compels symmetry across arms; the minimax rule bends and tilts the boundaries to penalize high-variance arms, shrinking their decision regions and requiring stronger evidence to select them. This is the inward shift of the $\mathbf{e}_{3}$ region in Figure\,(ref) and the asymmetric prior weights in Table\,(ref).
The prior probabilities $\pi^{k}$ are $1/J$ in Tables (ref), (ref), and (ref), whereas they deviate from $1/J$ in Table (ref). Specifically, even when the covariance matrix is asymmetric in two arms, $\pi^{k}=1/2$ for both arms. The intuition for this follows from Theorem (ref) (iv) combined with the Bayes risk expression $V=\sum_{k\in\mathcal{J}}\pi^{k}R\left(\bm{\theta}^{k},\bm{\delta}\right)$. Theorem (ref) (iv) states that the pointwise risks at the support points of the least-favorable prior, $R\left(\bm{\theta}^{k},\bm{\delta}\right)$, equal $V$. Hence, the outer weighted sum in the Bayes risk expression does not impose any restriction on the values of $\pi^{k}$. What determines $\pi^{k}$ are the decision boundaries of the MMR decision rule $\bm{\delta}$, which takes $\left(\pi^{k},\bm{\theta}^{k}\right)_{k\in\mathcal{J}}$ as its argument (see equation ((ref))). When the optimal MMR decision boundaries are symmetric, as in two-arm experiments or in multi-arm experiments with $\bm{\Sigma}_{\text{sym}}$, the prior probabilities satisfy $\pi^{k}=1/J$.
However, when the optimal MMR decision boundaries are asymmetric, the implied prior probabilities deviate from $1/J$. Intuitively, because the MMR rule shrinks the selection region for a high-variance arm, regret for the shrunk region would fall unless nature increases both the distance of that arm’s support from the line $t\mathbf{1}_{J}$ and its prior mass. The optimized solution therefore loads more prior weight on the high-variance arm and pushes its support deeper into $\Theta^{i}$. This is exactly what we see in Table\,(ref): the high-variance arm\,3 receives the largest prior weight ($\pi^{3}=0.433$) and a more distant support point from the line $t\mathbf{1}_{J}$ ($d\left(\bm{\theta}^{k},t\mathbf{1}_{J}\right)=1.679$), jointly restoring risk equalization stated in Theorem (ref) (iv).
We characterized the MMR-risk-optimal best-population selection rule, established local asymptotics when the mean rewards are only asymptotically normal, and illustrated the resulting decision boundaries. Our local asymptotic approximation allows the plug-in of mean-reward estimates directly into the MVN-MMR decision boundaries, where a consistent covariance estimator translates directly into decision boundaries for mean rewards. These multi-arm plug-in decision boundaries generalize the two-arm threshold MMR decision rules substantially. Our results allow researchers and practitioners to compare multiple competing policies by designing and implementing multi-arm RCTs within a unified decision-theoretic framework.