EconBase
← Back to paper

Minimax-Regret Sample Selection in Randomized Experiments

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

54,048 characters · 13 sections · 46 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Minimax-Regret Sample Selection in Randomized Experiments

\ifec

\else

abstractRandomized controlled trials are often run in settings with many subpopulations that may have differential benefits from the treatment being evaluated. We consider the problem of sample selection, i.e., whom to enroll in a randomized trial, such as to optimize welfare in a heterogeneous population. We formalize this problem within the minimax-regret framework, and derive optimal sample-selection schemes under a variety of conditions. Using data from a COVID-19 vaccine trial, we also highlight how different objectives and decision rules can lead to meaningfully different guidance regarding optimal sample allocation.

\fi

Introduction

Results from randomized controlled trials (RCTs) are often used to guide decision rules, especially in medical settings. It is thus natural to ask: How should a researcher design an RCT if their primary objective is to optimize for the expected welfare achieved via the induced decision rule? How should they reason about whom to enroll in the study?

The goal of this paper is to flesh out and investigate properties and tradeoffs for welfare-optimizing RCT designs in heterogeneous populations, i.e., in populations where different decision rules (also often referred to as treatment strategies) may be appropriate for different subgroups. There is a large existing literature that gives guidelines on {\it how many} study participants to enroll in a RCT. Such sample size calculations are usually driven by power considerations, and aim to guarantee statistically significant detection of the average treatment effect whenever this effect is larger than a practically relevant threshold altman1980statistics,moher1994statistical. More recently, manski2016sufficient,manski2019trial revisited sample size calculations with a focus on welfare considerations, and showed that relatively small sample sizes can still yield strong welfare guarantees under a minimax-regret criterion. azevedo2020b,azevedo2023b discuss properties of Bayesian rules for choosing the sample size. Perhaps surprisingly, however, the question of {\it whom} to enroll, or how to allocate a given number of spots in a study across different eligible groups, has not received similar attention.

In this paper, we consider sample selection for RCTs under the following simple model. The population of eligible study participants can be divided into $g = 1, \, \ldots, \, G$ groups, each with its own treatment effect $\tau_g$. The researcher has budget to run an RCT with $N$ total study participants, and needs to choose how many people $n_g$ to enroll from each group $g$. The researcher will use the data from the RCT to propose group-specific treatment rules. The researcher's ultimate goal is to achieve robust welfare guarantees and they seek to choose $n_g$ in line with this goal. For example, consider a company who wishes to make product design decisions and is not interested in precise estimates of treatment effects but instead in using a limited experimental budget to allocate samples to customer subgroups in order to optimize for downstream outcomes under an induced decision policy. Formally, as in manski2016sufficient,manski2019trial we specify the objective through the lens of minimax regret savage1951theory,manski2004statistical. We provide further details on our model in Section (ref).

Perhaps our most surprising finding is as follows. First let $U_g$ denote the $g$-th group's expected utility. Suppose we can assign different treatments to different groups, and seek to optimize the weighted average utility \smash{$U = \sum_{g = 1}^G \alpha_g U_g$} where the weights $\alpha_g$ are proportional to the size of each group in the population.\footnote{This aggregated utility objective is a scaled version of utilitarian social welfare that weights everyone in the population equally moulin2004fair.} Then, the minimax-regret study design allocates samples to the $g$-th group proportionally to \smash{$\alpha_g^{2/3}N$}. In particular, this recommendation deviates from the conventional practice of selecting samples in a stratified experiment in proportion to each group's share in the population singh2013fundamentals,sharma2017pros, and instead over-samples minority groups. In Section (ref), we discuss reasons for this finding. We also prove that more familiar sample-allocation schemes, like proportional sampling, are minimax optimal under alternate decision theoretic framings. We also consider a hypothetical application of our method guided by historical COVID-19 trial data, which illustrates the notably different sample allocations that arise under different framings (Section (ref)). Overall, our results highlight how different welfare-focused objectives have different implications as to optimal sample selection.

Related work

There has been considerable work on methods for learning treatment rules in heterogeneous populations, using data from either randomized or observational studies athey2021policy,kitagawa2018should,manski2004statistical,stoye2009minimax,swaminathan2015batch,zhao2012estimating. These methods assume that the researcher has access to a dataset of a given size, and then seeks to use the available data to discover heterogeneity patterns and learn treatment assignment rules. For example, ross2023estimated use electronic-health-record data from the US Department of Veterans Affairs to study heterogeneity in how different groups of patients with suicide ideation or recent suicide attempts respond to hospitalization, and argue that decision rules that leverage such heterogeneity could substantially improve outcomes. Here, in contrast, we do not take the dataset as given, and instead seek to provide guidance on data collection with heterogeneous study populations.

Our research question falls within literature on using decision theory to guide the design of RCTs. As noted above, manski2016sufficient,manski2019trial and azevedo2020b,azevedo2023b use decision theoretic criteria to choose the sample size in an RCT. Another line of work compares different randomization rules over the treatment and control sample allocation in terms of the resulting variance for the average treatment effect, and seeks intervention randomization rules with minimax variance bai2021randomize,kallus2021optimality,li1983minimaxity,wu1981robustness. banerjee2020theory study the value of randomization itself under a number of different decision theoretic frameworks. However, we are not aware of prior work in this literature that considers the question of sample allocation across groups.

Finally, our work also relates to the broader literature on fairness in data analysis. RCTs have often been run in blatantly unfair ways. For example, until recently, a substantial fraction of medical research in the United States was conducted on white men, while excluding women and racial minorities dresser1992wanted; and FDA-approved trials still under-sample black participants relative to their share of the population alsan2022representation. However, while it is easy to agree that some ways of running randomized trials are unfair, it can be much harder to say what makes a trial fair. For example, if there is a minority subgroup that has the potential for both disproportionate benefits and disproportionate risks from participating in a study, what considerations should go into deciding how to fairly include them in the study? In such settings results on welfare-driven sample selection may help provide useful insights into potential desiderata for fair RCT design. Similar considerations are also relevant to fairness-aware data collection when training machine learning and AI systems holstein2019improving.

Minimax-Regret Sample Selection

We start with a standard analysis in a stratified completely randomized controlled trial where participants within a stratum receive treatment or control in a 1:1 ratio. The experimenter has the freedom to enroll a maximum of $N$ participants from $G$ groups, with $n_g$ participants allocated to group $g$, $g=1,\dots, G$. In other words, the class of feasible sample selection $\mathcal{N}$ is composed of all $\bfn=(n_1,\dots,n_G)^\top$ such that $\sum_{g=1}^G n_g\le N$, $n_g\in 2\ZZ_{\ge 0}, \forall g$.

Given the sample selection $\bfn$, the experimenter conducts the experiment, and collects data $D=\cb{Y_g,W_g}_{g=1,\dots,G}$ on all of the groups from the experiment, in which $D$ is generated from some data-generating process $\mathcal{D}(\bfn,\tau)$ parameterized by $\tau$, where $Y_g=(Y_{g,1},\dots,Y_{g,n_g})^\top$ is the vector of observed outcomes and $W_g=(W_{g,1},\dots,W_{g,n_g})^\top$ is the vector of treatment assignments. Following the classic potential outcome framework, under the SUTVA condition imbens2015causal, the observed outcome $Y_{g,i}$ can be written as $Y_{g,i}=W_{g,i}Y_{g,i}(1)+(1-W_{g,i})Y_{g,i}(0)$, where $W_{g,i}$ is the treatment assigned to unit $i$ in group $g$, and $Y_{g,i}(w)$, $w\in\cb{0,1}$, is the potential outcome that we would have observed from unit $i$ in group $g$ if it was assigned to treatment group $w$.

Based on the collected dataset, the experimenter proceeds to estimate the treatment effect in each group with an estimator $\htau(D)$ that maps from the collected data $D$ to $G$ real-valued estimates of treatment effect $\htau_g$. They then make a decision $\delta(D)=(\delta_1,\dots,\delta_G)^\top\in \cb{0,1}^G$ regarding the administration of treatment within each group. Based on the quality of the decisions, the experimenter attains an aggregated total utility $U(\alpha,\tau,\delta(D))$ defined as

equation[equation omitted — 103 chars of source]

The corresponding regret of making such a decision is

equation[equation omitted — 86 chars of source]

where

equation[equation omitted — 81 chars of source]
figure[figure omitted — 1,532 chars of source]

Ideally, one would like to identify the optimal sample selection $\bfn\in\mathcal{N}$ by minimizing the expected total regret as given by $\EE[D]{R(\bfn,\delta(D))}$, with expectation taken over the data generation. However, the expected total regret is contingent upon the unknown quantity $\tau$, rendering it challenging to quantify during the design phase before observing any data.

Here we adopt a minimax-regret design approach where the sample selection $\bfn$ is chosen to minimize the maximum expected regret that could be incurred if $\tau$ is chosen by an adversary with knowledge of both the choice of sample selection $\bfn$ and the decision rule $\delta(D)$. An illustration of this strategy can be found in Figure (ref). Under this framework, we say that a sample selection rule $\bfn\in\mathcal{N}$ is minimax-regret if this sample selection, together with the choice of estimator, minimizes $\EE[D]{R(\bfn,\delta(D))}$ when the data generation is adversarial. For simplicity, we only consider sample allocation rules where each group is given an even number of samples; this lets us focus on completely randomized experiments with treatment (and control) perfectly balanced in each stratum.

definitionLet \smash{$\mathcal{N}=\{\bfn:\sum_{g=1}^G n_g\le N, n_g\in 2\ZZ_{\ge 0}, \forall g \}$}. The worst-case regret of any sample selection rule $\bfn \in\mathcal{N}$ is \begin{equation} H(\bfn) = \inf_{\delta(D)}\max_{\tau\in\RR^{G}} \EE[D]{R(\bfn,\delta(D))}. \end{equation} Such a sample selection is minimax-regret if $\bfn \in \argmin_{\bfn\in\mathcal{N} }H(\bfn)$.

Below, we give a near-optimal solution to this minimax-regret problem when the data-generating distribution $D$ is from a Gaussian\footnote{When the data-generating distribution is generalized to be from a sub-Gaussian class, an asymptotic limits-of-experiments analogue to Theorem (ref) holds in the large sample limit; see hirano2009asymptotics for a general discussion of results of this type.} class. The sub-optimality in our result is only due to rounding to achieve integer sample allocations, and our solution is optimal whenever no rounding is needed in (ref). We present a brief proof of the theorem, with proof of the technical lemmas deferred to the supplementary material. Compared with a conventional proportional selection, the minimax selection rule dilutes the influence of the weights $\alpha_g$, and oversamples the minority groups.

theoremSuppose $Y_{g,i}(0)$ and $Y_{g,i}(1)$ are generated i.i.d. from Gaussian distributions $N(b_g-\tau_g/2,s_{0,g}^2)$ and $N(b_g+\tau_g/2,s_{1,g}^2)$, respectively.\footnote{For simplicity, we assume the experimenter knows the standard errors. However, $s_{0,g}$ and $s_{1,g}$ can also be regarded as upper bounds on unknown standard errors, as the adversary would maximize the noise levels.} Suppose that $s_{0,g}^2,s_{1,g}^2,\alpha_g$ are bounded away from $0$ and $\infty$ for all $g$. Then among 1:1 completely stratified designs with at most $N$ total units, the sample selection $\bfn^*=\p{n^*_1,n^*_2,\cdots,n^*_G}^\top$ with \begin{equation} n^*_g=2\left\lfloor\frac{(s_{0,g}^2+s_{1,g}^2)^{1/3}\alpha_g^{2/3}N}{2\sum_{g'=1}^G(s_{0,g'}^2+s_{1,g'}^2)^{1/3}\alpha_{g'}^{2/3}}\right\rfloor, \qquad\qquad g=1,\dots,G \end{equation} is nearly minimax-regret in the sense that \begin{equation} \frac{H\p{\bfn^*}-\min_{\bfn\in\mathcal{N} }H\p{\bfn}}{\min_{\bfn\in\mathcal{N} }H\p{\bfn}} = \oo\p{N^{-1}}. \end{equation} Furthermore, the decision rule induced by a threshold of the difference-in-means estimator \begin{equation} \begin{split} &\delta^{DM}(D) = \p{I(\htau^DM_1> 0),\cdots,I(\htau^DM_G> 0)}^\top\\ &\htau^DM_g = \frac{2}{n_g}\sum_{\cb{i:W_{g,i}=1}} Y_{g,i} - \frac{2}{n_g}\sum_{\cb{i:W_{g,i}=0}} Y_{g,i},\qquad\qquad g=1,\dots,G \end{split} \end{equation} is the minimax-regret decision rule given data from the experiment.
proof[Proof of Theorem (ref)] We start by verifying the last part of the statement, i.e., that given the available data we should choose a decision based on the decision rule induced by a threshold of the difference-in-means estimator, and the induced regret is linear in standard errors. This result is common in statistical treatment choice with Gaussian errors; see stoye2012minimax and tetenov2012statistical for a discussion. \begin{lemma} Under the assumptions in Theorem (ref) and for any $\bfn \in \mathcal{N}$ with $n_g>0$ for all $g$, there exists a universal constant $C_0\approx 0.17$ such that \begin{equation} H(\bfn)=C_0\sum_{g=1}^G \alpha_g\sqrt{\frac{2(s_{0,g}^2+s_{1,g}^2)}{n_g}}. \end{equation} Furthermore, the thresholded difference-in-means estimator attains this bound, \begin{equation} \begin{split} &\EE[D]{R(\bfn,\delta^DM(D))} \leq C_0\sum_{g=1}^G \alpha_g\sqrt{\frac{2(s_{0,g}^2+s_{1,g}^2)}{n_g}} \end{split} \end{equation} for all $\tau \in \RR^G$. \end{lemma} Now, to lower-bound minimax regret for sample allocation rules $\bfn \in \mathcal{N}$, we start by considering the minimum of the regret bound (ref) over the the following convex relaxation of $\mathcal{N}$ $$\widetilde{\mathcal{N}}=\cb{\bfn:\sum_{g=1}^G n_g\le N, n_g\in \RR_{\ge 0}, \forall g }. $$ Doing so yields the following: \begin{lemma} Under the assumptions in Theorem (ref), \begin{equation} \begin{split} &\Tilde{n}_g = \argmin_{\bfn\in\widetilde{\mathcal{N}}} C_0\sum_{g=1}^G \alpha_g\sqrt{\frac{2(s_{0,g}^2+s_{1,g}^2)}{n_g}} \\ &\Tilde{n}_g=\frac{(s_{0,g}^2+s_{1,g}^2)^{1/3}\alpha_g^{\frac{2}{3}}N}{\sum_{g'=1}^G(s_{0,g'}^2+s_{1,g'}^2)^{1/3}\alpha_{g'}^{2/3}},\qquad\qquad g=1,\dots,G. \end{split} \end{equation} \end{lemma} Our proposed sample allocation rule $n^*_g$ is obtained by rounding $\Tilde{n}_g$. \begin{lemma} Under the assumptions in Theorem (ref), \begin{equation} \begin{split} \inf_{\delta(D)}\max_{\tau\in\RR^{G}}\EE[D]{R(\bfn^*,\delta(D))}-\min_{\bfn\in\mathcal{N} }\inf_{\delta(D)}\max_{\tau\in\RR^{G}} \EE[D]{R(\bfn,\delta(D))} = \oo\p{N^{-3/2}}. \end{split} \end{equation} \end{lemma} From Lemmas (ref) and (ref), $\min_{\bfn\in\mathcal{N} }H\p{\bfn}=\Theta\p{N^{-1/2}}$. Combining this with (ref) gives rise to the claimed result.

Intuitively, the reason that the optimal sampling scheme in this setting samples the smaller proportion group more, is that, in the absence of additional structure, adding a small number of additional samples to a small number of samples, is more helpful in reducing uncertainty compared to adding those same samples to a larger pool of data for the group with a higher proportion.

Conventional Sample Selections as Minimax-Regret Solutions

Above, we found that the minimax-regret sample selection suggests oversampling minority groups and employing a sample allocation proportional to \smash{$\alpha_g^{2/3}$}, which is not a common choice in the literature---at least currently. In this section, we discuss two widely adopted sample selection methods, proportional selection and equal selection. We demonstrate that these common selections can also be regarded as regret-minimax selections when the experimenter makes a single treatment decision for all groups, or, respectively, uses an alternate utility function from the one that we used in Section (ref).

Proportional Selection

Proportional sampling (or sometimes referred to as “stratified sampling”), is a longstanding and widely embraced approach in the literature of survey sampling singh2013fundamentals,sharma2017pros, and is often the default choice when designing stratified experiments. It involves sampling in each group proportionally to the group’s size when compared to the population. While the minimax solution to our defined problem in Section (ref) is different from the conventional proportional selection, we demonstrate below that proportional sampling is in fact minimax-regret in a different model where the researcher is constrained to make a single decision for the entire population based on the sign of a difference-in-means estimator calculated over the complete sample.

definitionConsider the decision class $\{\Tilde{\delta}(D)\}$ that maps from the collected data $D$ to a single decision in $\cb{0,1}$, i.e., only a joint decision can be made for all groups. Define $\Tilde{\delta}^{DM}(D)$ to be such a decision that is made according to the sign of a pooled difference-in-means estimator, i.e., \begin{equation} \begin{split} &\Tilde{\delta}^{DM}(D) = I(\htau^DM,p> 0),\\ &\htau^DM,p = \frac{2}{N}\sum_{g=1}^G\sum_{\cb{i:W_{g,i}=1}} Y_{g,i} - \frac{2}{N}\sum_{g=1}^G\sum_{\cb{i:W_{g,i}=0}} Y_{g,i}. \end{split} \end{equation} The worst-case regret of any sample selection rule $\bfn \in\mathcal{N}$ subject to this class of decision rules is \begin{equation} H^\dag(\bfn) = \max_{\tau\in\RR^{G}} \EE[D]{R(\bfn,\Tilde{\delta}^{DM}(D))}. \end{equation} Such a sample selection is joint-minimax-regret if $\bfn \in \argmin_{\bfn\in\mathcal{N} }H^\dag(\bfn)$.
theoremUnder the conditions of Theorem (ref), suppose in addition that $\alpha_gN$ is an even integer for all $g$. Then the sample selection $\bfn^{\dag}=\p{n^\dag_1,n^{\dag}_2,\cdots,n^{\dag}_G}^\top$ with \begin{equation} {n}^{\dag}_g=\alpha_gN, \qquad\qquad g=1,\dots,G \end{equation} is joint-minimax-regret in the sense that \begin{equation} {n}^{\dag} \in \argmin_{\bfn\in\mathcal{N} }H^\dag(\bfn). \end{equation}

The intuition behind this sample selection is that, when the decision-maker has to make a joint decision, any mismatch between the weight $\alpha_g$ in the population and the weight $n_g/N$ used in the estimator can lead to arbitrarily bad performance. Thus, under the decision process outlined in Definition (ref), any sample selection other than proportional selection may result in infinite regret.

An Egalitarian Solution

Although much work in learning and data-driven decision making is implicitly underpinned by a utilitarian ethical model, some applications may call for different ways of comparing benefits across people or groups. In particular, some researchers have argued that maximization of utilitarian social welfare may result in decisions that prioritize the overall population's well-being at the expense of utility for minority groups rawls1971atheory. In consideration of this, we also investigate minimax-regret sample selection guided by the egalitarian rule (also known as the max-min rule or the Rawlsian rule), which seeks to maximize the welfare of the worst-off individual in society rawls1971atheory,kolm1997justice,moulin2004fair,sen2018collective. Under this setup, the regret objective is no longer framed in terms of the average utility across the entire population, but rather at the group level---and we then focus on minimizing regret uniformly across all groups.

definitionDefine $U_g(\tau_g,\delta_g) = \tau_g\delta_g$, and $R_g(\bfn,\delta(D))=U_g(\tau_g,\delta_g^*)-U_g(\tau_g,\delta_g(D))$. The worst-case regret of any sample selection rule $\bfn \in\mathcal{N}$ for the egalitarian social welfare\footnote{Here, we focus on the single group with the smallest forecasted regret while averaging over the randomness in the data generation. Alternatively, one may consider defining $g^* = \argmax_g R_g(\bfn,\htau(D))$, where the worst-off group might vary with different realizations. While this is a valid formulation to consider, we do not pursue this path here.} is \begin{equation} H^\ddag(\bfn) = \inf_{\delta(D)}\max_{\tau\in\RR^{G}}\max_g \EE[D]{R_g(\bfn,\delta(D))}. \end{equation} Such a sample selection is egalitarian-minimax-regret if $\bfn \in \argmin_{\bfn\in\mathcal{N} }H^\ddag(\bfn)$.

Below, we show that, when one considers egalitarian social welfare, the minimax-regret problem indeed leads to an equal sample allocation across all groups if the signal-to-noise ratio in each group is the same.

theoremUnder the conditions of Theorem (ref), the sample selection $\bfn^{\ddag}=\p{n^\ddag_1,n^{\ddag}_2,\cdots,n^{\ddag}_G}^\top$ with \begin{equation} {n}^{\ddag}_g=2\left\lfloor\frac{s_{0,g}^2+s_{1,g}^2}{2\sum_{g=1}^G\p{s_{0,g}^2+s_{1,g}^2}}N\right\rfloor, \qquad\qquad g=1,\dots,G \end{equation} is nearly egalitarian-minimax-regret in the sense that \begin{equation} \frac{H^\ddag\p{\bfn^\ddag}-\min_{\bfn\in\mathcal{N} }H^\ddag\p{\bfn}}{\min_{\bfn\in\mathcal{N} }H^\ddag\p{\bfn}} = \oo\p{N^{-1}}. \end{equation}

One may notice that the resulting sample selection shares a common ground with the Neyman allocation (also known as the optimal allocation), which advocates sampling proportionally to the standard deviation neyman1934two. However, Neyman's goal was to minimize the variance of a sample mean estimator, whereas our focus is on the quality of the decision and we target directly the regret with egalitarian social welfare.

Case Study: A COVID-19 Vaccine Trial

To further explore the implications of different group sample size allocation practices, we conduct a semi-synthetic analysis using historical data from COVID-19 vaccine development. Different demographic subpopulations may experience different benefits and side effects from vaccines, and so it is of interest to consider experimental design to inform vaccination recommendations for subpopulations. In real-world clinical trials, there are many important considerations beyond the ones discussed in this paper; our goal here is simply to help illustrate how different decision-theoretic criteria may lead to notably different group allocations.

We use data from a large Phase 3 randomized, placebo-controlled COVID-19 vaccine trial conducted at 99 centers across the United States baden2021efficacy. To account for anticipated heterogeneity in the treatment effect across different demographic groups, the randomization is stratified into three groups: 1. $\ge18$ and $<65$ years and not at risk (for severe Covid-19); 2. $\ge18$ and $<65$ years and at risk; 3. $\ge 65$ years. In an effort to demonstrate how the proposed sample selection schemes can be used in practice, we explore the sample selections chosen with the three minimax sample selections given in Theorems (ref), (ref) and (ref), and compare their performances under both the worst-case incidence rates and the incidence rates reported in baden2021efficacy.

Motivated by the two-fold primary objective of the trial, we study a composite outcome $\Tilde{Y}_{g,i}=Y_{g,i,1}+\beta Y_{g,i,2}$, where $Y_{g,i,1}$ represents the incidence of severe COVID-19 of unit $i$ in group $g$, $Y_{g,i,2}$ represents the incidence of severe solicited adverse reaction of unit $i$ in group $g$ during the injections, and $\beta$ governs the tradeoff between the efficacy and the safety of the vaccine.\footnote{We define severe solicited adverse reactions as solicited adverse reactions with a toxicity grade for Erythema greater than or equal to 3 baden2021efficacy.} Since the incidence rates are only reported stratified by whether age is above or below 65 years, we will consider the case where there is only treatment effect heterogeneity between people with age $\ge 18$ to $<65$ yr (group 1), and with age $\ge 65$ yr (group 2). According to the 2020 U.S. Census Bureau's data on age census2020age, the estimated proportion of people aged 65 and over in the United States is nearly 17%, which motivates our choice of weights $\alpha_1=0.83$ and $\alpha_2=0.17$. Throughout, we consider two different values of $\beta$, with case 1 being $\beta=0.005$ (in which case it is beneficial to assign treatment to both groups under the reported incidence rates) and case 2 being $\beta=0.025$ (in which case it is only beneficial to assign treatment to people $>=65$ years old under the reported incidence rates).

Variance Approximation

To determine the minimax-regret sample selections, one needs to specify additionally the noise levels $s_{w,g}$, $g=1,2$, $w=0,1$. We obtain a conservative approximation of those noise levels based on empirical observations of baseline risks for severe COVID-19 reported by the Centers for Disease Control and Prevention (CDC) cdc2020covid and risk levels of severe adverse reactions reported in an earlier phase trial jackson2020mrna.

table[table omitted — 356 chars of source]

According to CDC's data recorded as of September 2020 (during the period when the protocol was being prepared), the incidence rate for newly admitted patients with confirmed COVID-19 was around $0.7\%$ for people $>=18$ to $<65$ yr, and $2.5\%$ for people $>=65$ yr cdc2020covid. Furthermore, according to a phase 1 trial of an mRNA vaccine, the incidence rate of severe solicited adverse reactions was around $6.7\%$ in a dose group of 100mcg jackson2020mrna. Given this, we consider a conservative approximation of the noise levels, where we take the incidence rates of severe COVID-19 in the treated group to be the same as the ones in the control group and the incidence rates of severe solicited adverse reactions in the control group to be the same as the ones in the treated group. Those approximated noise levels are summarized in Table (ref).

Budget Constraint

While there is no explicit requirement that sample sizes must be based on power calculations, it is generally considered good practice to conduct power calculations to determine the appropriate sample size for clinical trials food1988guideline. Thus, we determine the total sample size constraint by following baden2021efficacy that sets $N$ to be the sample size required to achieve a power of 0.9 under a specified one-sided hypothesis test with $\alpha=0.05$.

In baden2021efficacy, the sample size is chosen to be the minimum number of samples such that there is 90% power to detect a 60% reduction in hazard rate of COVID-19. Given that hazard ratios are commonly interpreted as the incidence rate ratios in practice hernan2010hazards, we translate this into a sample size calculation where $N$ is chosen to be the minimum number of samples required to achieve 90% power in detecting that $\tau_{\text{COVID}} = -0.6\,\EE{Y_{g,i,0}(0)}$. Utilizing the approximated incidence rates of severe COVID-19 in the control group as detailed in Section (ref), we have the following hypothesis test:

equation[equation omitted — 101 chars of source]

As in baden2021efficacy, the sample size is calculated by assuming a homogeneous treatment effect across the entire population, which is a standard assumption in such calculations.

To calculate the total sample size, we adopt the approach outlined in charan2013calculate and assume that $\Tilde{Y}_{g,i,1}(1)$ and $\Tilde{Y}_{g,i,1}(0)$ are generated i.i.d. from $N(b_{\text{COVID}}-\tau_{\text{COVID}}/2, s_{0,\text{COVID}}^2)$ and $N(b_{\text{COVID}}+\tau_{\text{COVID}}/2, s_{1,\text{COVID}}^2)$, respectively, with $s_{w,\text{COVID}}^2=\sum_{g=1,2}\alpha_g s_{w,g,\text{COVID}}^2$, $s_{w,g,\text{COVID}}$ denoting the standard deviations of $Y_{g,i,1}(w)$, and $b_{\text{COVID}}$ being a nuisance baseline parameter. We find the required total sample size $N$ that, with a difference-in-means estimator,

equation[equation omitted — 138 chars of source]

Note that under $H_1$,

equation[equation omitted — 111 chars of source]

and $Z_{p}$ is the $p$th quantile of a standard Gaussian random variable. This gives rise to a total sample size of

equation[equation omitted — 180 chars of source]

Thus, we set the constraint to be $n_1+n_2\le 9320$, $n_1, n_2>0$.

Worst-Case Performance

We start by comparing the worst-case performances of the minimax sample selections (Equations (ref), (ref) and (ref)) calculated based on the approximated noise levels in Table (ref). For demonstration purposes, we consider a data-generating distribution as stated in Theorem (ref), with noise levels fixed to be the ones given in Table (ref) and treatment effects chosen adversarially. To evaluate the worst-case performances, we calculate

equation[equation omitted — 62 chars of source]

the maximum expected regret using the DM estimator with group-level decisions,

equation[equation omitted — 70 chars of source]

the maximum expected regret using the DM estimator with a joint decision, and

equation[equation omitted — 70 chars of source]

the maximum expected regret using the DM estimator with group-level decisions under an egalitarian utility function, respectively.

table[table omitted — 1,729 chars of source]

Table (ref) presented the minimax sample allocations as well as their worst-case performances under different decision paradigms. We first highlight that different decision paradigms result in significantly distinct sample selections. The egalitarian and the Neyman approaches allocate a substantial portion of the sample to the older group (i.e., the group with a larger anticipated variance and a smaller weight), with the egalitarian selection being more extreme. In contrast, both the minimax and the proportional selections prioritize the younger group that possesses a large weight, with the latter assigning it a higher number of samples. Moreover, it is evident that the minimax-regret sample selections are indeed regret minimax, and they consistently outperform the other sample allocations with respect to the specific worst-case scenarios defined by their targeting decision paradigms. Meanwhile, extreme selections that allocate all samples to one group result in unbounded worst case regret. Furthermore, when the policy-maker expects to make a single decision for the entire population, any sample selection other than proportional selection may result in arbitrarily bad decisions. Nevertheless, proportional selection's performance deteriorates when group-level decisions are permitted, particularly when the policy-maker prioritizes the utility in the worst-off group.

Performance under Reported Incidence Rates

In addition to the worst-case performances, we also evaluate the performance of the minimax sample selections under the incidence rate reported in baden2021efficacy, reproduced in Table (ref). We then use Gaussian approximation to obtain the distribution of the difference-in-means treatment effect estimator, resulting in parameters reported in Table (ref).

table[table omitted — 647 chars of source]
table[table omitted — 483 chars of source]
table[table omitted — 1,504 chars of source]

Table (ref) summarizes the sample allocations and their performances. Despite the approximated noise levels used for guiding the selections being largely misspecified, the selections we investigated still achieve decent performances. In particular, although the minimax selection may appear conservative, it tends to outperform the proportional selection in most cases in this non-adversarial real-data example. In Case 1 where the minority group exhibits a higher signal-to-noise ratio, the minimax selection that allocates more equally outperforms under separate decisions; when only a joint decision is allowed, the egalitarian and the Neyman selections that allocate a large portion of the sample to the easier-to-learn group achieve a smaller regret. Nevertheless, in Case 2 where the optimal decisions for the groups differ, the egalitarian selection that aims to minimize regret in the worst-off group performs considerably worse, and it appears to be less robust to model misspecification than the Neyman selection in the setting we consider.

Bayes-Optimal Sample Selection

In Section (ref), we identified the minimax-regret sample selection by minimizing the expected regret with $\tau$ chosen adversarially. The minimax rule can also be interpreted as a Bayes-optimal rules with a least favorable prior on $\tau$ wasserman2013all. While the minimax rule has attractive worst-case guarantees, a decision-maker with meaningful information about $\tau$ may prefer to conduct Bayes-optimal sample selection using their chosen prior rather than with the least favorable one underlying the minimax rule. In this section, we briefly consider Bayesian sample selection under an informative prior, and examine alignment of the minimax-regret sample selections with Bayes-optimal ones in a simple example.

Within a Bayesian framework, the parameters are regarded as random variables governed by a prior distribution. Unlike in non-Bayesian approaches where the treatment effect $\tau$ is treated as an unknown fixed value, the experimenter holds a belief captured by a prior distribution $p$ that reflects their understanding of $\tau$ based on various sources such as historical data, literature, or common beliefs.\footnote{In the Bayesian framework, we assume knowledge of all other components of the data-generating distribution. However, it is also possible to extend our framework to incorporate priors on these components.} Below, we give a formal definition of the Bayes-optimal sample selection.

definitionGiven a prior $p$ on the unknown parameter $\tau$, let $\delta^p(D)$ be a decision that is made according to the sign of the posterior mean of $\tau$, i.e., \begin{equation} \begin{split} &\delta^{p}(D) = I(\htau^PM\ge 0),\\ &\htau^PM = \EE{\tau\cond D}. \end{split} \end{equation} The Bayesian expected regret of a sample selection rule $\bfn \in\mathcal{N}$ subject to this prior is \begin{equation} H^p(\bfn) = \EE[\tau\sim p]{ \EE[D]{R(\bfn,\delta^p)\cond \tau}}. \end{equation} Such a sample selection is Bayes-optimal with respect to prior $p$ if $\bfn \in \argmin_{\bfn\in\mathcal{N} }H^p(\bfn)$.

An advantage of the Bayesian decision-theoretic framework is its flexibility in modeling relationships between different groups. For example, one can account for information sharing across groups by adjusting the prior correlation between $\tau_g$'s. A high correlation implies that the treatment effects across groups are likely similar, and thus insights gained from one group can inform the others.

To obtain the Bayesian expected regret, both a prior and a data-generating distribution must be specified. When the prior is conjugate to the data-generating process, $\hat{\tau}^\text{PM}$ can be derived in closed form. Otherwise, computational methods such as Markov Chain Monte Carlo gelman2013bayesian are needed to approximate $\hat{\tau}^\text{PM}$. In the supplementary material, we provide derivations of the posterior distribution to facilitate computation when the prior $p$ and the data-generating distribution $\mathcal{D}$ form a normal-normal conjugate.

To get some insight about the effect of prior choice on sample selection, we consider the following synthetic two-group setting. For unit $i$ from group $g$, $g=1,2$, we consider the following data-generating distribution:

equation[equation omitted — 339 chars of source]

Throughout, we set $s_{0,g}=s_{1,g}=1$, and $\sigma_g=0.1$ for all $g$, and investigate the change in the resulting optimal sample selection as we vary $\alpha_1$. As in Section (ref), we determine the sample size using (ref) based on a power calculation for the hypothesis test $H_0:\tau=0$ versus $H_1:\tau=0.1$, resulting in a total sample size of $N=1713$.

figure[figure omitted — 524 chars of source]

Figure (ref) displays (1) the Bayes-optimal sample selection with $\rho=0$, (2) the Bayes-optimal sample selection with $\rho=0.9$, (3) the minimax-regret sample selection, and their corresponding performances under three different priors. Given the specific prior and budget constraints used here, the Bayes-optimal selection under Gaussian prior with $\rho=0$ is close to the minimax-regret selection that oversamples the minority group, while the Bayes-optimal selection under Gaussian prior with $\rho=0.9$ is close to the proportional selection.\footnote{As one would expect, using $\rho\in(0,0.9)$ here leads to selections in between proportional selection and minimax selection.} When evaluated under the least favorable prior and the Gaussian prior with $\rho=0$, the minimax-regret selection and the Bayes-optimal selection with $\rho=0$ perform better than the other two selections, especially when the minority group has a small weight. Under the Gaussian prior with $\rho=0.9$, all selections show similar performances. As anticipated, the expected regret evaluated under the prior associated with the minimax-regret selection is consistently the highest, since it is the least favorable one.

These findings suggest that, in a simple two-group setting where the sample size is chosen such that the experiment has power for typical effect sizes under the prior, our minimax rules form a reasonable starting point for a discussion on sample selection: With uncorrelated priors ($\rho = 0$) our recommendations closely match the Bayes-optimal ones, while with highly correlated priors ($\rho = 0.9$), all methods do comparably well (because it's possible to share information across groups). We do note, however, that it is also possible to construct examples where the Bayes-optimal rule diverges from our recommendations; for example, in the appendix we show a highly under-powered example in which the Bayes-optimal rule chooses to over-sample the majority group. Getting a deeper understanding of instance-level properties of the minimax-regret rules---and of when this approach is aligned with Bayesian sample selection with an informative prior---presents an interesting direction for further work.

Discussion

To guide sample selection subject to a fixed budget, we employ a minimax-regret framework that aims to minimize the maximum expected gap between the achieved utility and the maximum potential utility attainable under any decision rule had the true data-generating distribution been determined by an adversary. Within this minimax-regret framework, we explore the cases when the decision-maker can and cannot make separate decisions across groups, and the case when the target utility is utilitarian and egalitarian. We give a summary of all minimax-regret sample selections we investigated in Table (ref). Our results highlight that optimal sample selection can vary significantly based on different practices and objectives.

table[table omitted — 451 chars of source]

When considering utilitarian social welfare, the weights $\alpha_g$ stand as an important factor for determining the optimal sample selection. A prevalent approach is to sample proportionally to $\alpha_g$. Our results affirm that such proportional sampling is minimax-regret when a joint decision is guided by a sample-wise difference-in-means estimator, under which it is crucial to ensure that the sample is representative of the overall population.

One common critique of proportional sampling is that it fails to take into account the differentiation in subgroup variances sharma2017pros. Indeed, we show that when the decision-maker is not restricted to a joint decision, subgroup variances start to play an important role in optimal sample selections. Moreover, the minimax-regret solution with separate decisions suggests sampling proportionally to $\alpha_g^{-2/3}$ and thus oversamples the minority groups, highlighting the importance of making qualitative decisions for each of the subgroups in the population.

While utilitarian social welfare is a natural and common choice, it does favor groups with large proportion weights, which may sometimes be regarded as unfair. An alternative is to consider the egalitarian rule that suggests making decisions that maximize the minimum utility of all individuals in society under the least undesirable condition. Our findings demonstrate that this indeed results in equal sample allocation across all groups when the signal-to-noise ratio is the same across all groups. In scenarios where signal-to-noise ratios vary among groups, the resulting allocation suggests sampling proportionally to the variance of the treatment-control difference, ensuring an equal level of confidence in estimating the treatment effect across all groups.

One natural concern is whether the decision-theoretic mechanisms we propose for allocating samples to subgroups---given a fixed total sample budget---may interact with or change the sample size calculations commonly done for power analyses which typically assume population-proportional group allocation in the study design. Our resulting study designs will still satisfy the power calculations, but may be with respect to a reweighted subgroup population. We leave interesting considerations of experimental design that can jointly vary the total sample size, the subgroup allocation and the randomization over treatment and control, to future work.

Finally, in this paper, we considered welfare-optimizing sample selection through the lens of minimax-regret decision rules. This approach is attractive in terms of its robustness and generality. However, in settings where prior evidence about the effectiveness of a treatment is available, it may be desirable to instead use Bayesian methods that incorporate this prior knowledge. Another interesting direction would be to investigate optimal sample selection under non-linear objectives. Models for cooperative bargaining can be used to motivate targeting expected log-utility nash1950bargaining,kaneko1979nash; and baek2021fair recently studied implications of targeting such objectives in bandit experiments. It may also be of interest to consider a richer welfare model that takes into account potentially different risks and benefits for study participants (who may risk direct harms from trying out a potentially dangerous treatment) and the general population (whose risks and benefits are mediated by the overall study findings).

\ifec

acksThis research was supported by the Stanford Graduate School of Business initiative on Business, Government & Society.