EconBase
← Back to paper

Designing Persuasive Experiments

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

97,675 characters · 38 sections · 74 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Designing persuasive experiments

abstractIncentives in experimental design are often misaligned: experimenters design and finance experiments to seek regulatory approval, while regulators seek to maximize social-welfare. We propose a framework to resolve this conflict, wherein regulators set a minimum welfare threshold, and experimenters optimize designs subject to this constraint. It requires no knowledge of experimenters' private preferences or costs and mitigates strategic Bayesian persuasion. Under normal priors, Neyman-allocation is always the optimal-sampling strategy, regardless of specific objectives. We also characterize the optimal stopping-rule. A numerical study calibrated to clinical-trial data shows sample-size reductions of over 48% relative to classical designs attaining the same social-welfare.

Introduction

In many scientific domains, researchers design experiments with the goal of persuading a regulator to take a specific action. For instance, in clinical trials, pharmaceutical companies conduct studies to convince regulators---such as the FDA---to approve a new drug. Similarly, in development economics, researchers run experiments to influence policymakers to adopt particular interventions. In these settings, the preferences of the experimenter can often diverge significantly from those of the regulators. A pharmaceutical company, for example, may have a strong incentive to secure drug approval regardless of the drug's actual efficacy.

At the same time, in such strategic environments, the researcher typically bears the full cost of experimentation. The experimental design is proposed to the regulator before the experiment begins, and the researcher is able to fully commit to this design. The regulator, in turn, can either approve the proposed design or request a new one. This article studies which experiments a researcher should suggest and a regulator should approve when the interests of researchers and regulators may not be aligned, yet the cost and control over the design lie entirely with the researcher.

In practice, experiments in such settings are typically framed within the hypothesis testing paradigm. For example, in clinical trials, pharmaceutical companies commit to a specific test statistic and agree to report its value upon completion of the study. Regulatory approval is then granted if the test statistic exceeds a predetermined threshold. In designing such experiments, firms must navigate a trade-off: minimizing the cost of experimentation while maximizing the power of the test---which indirectly affects the probability of regulatory approval---subject to maintaining control over the test's size.

However, the hypothesis testing paradigm is often criticized for not explicitly aligning with social welfare objectives. For instance, it is standard practice to apply a 5% significance level across clinical trials, regardless of the specific context or stakes involved. Such a uniform threshold can be far from socially optimal. Over the past 15 years, the US Food and Drug Administration (FDA) has approved several drugs, such as Makena (Hydroxyprogesterone Caproate), Exondys 51 (Eteplirsen), and, perhaps more controversially, Aduhelm (Aducanumab), even though the supporting evidence of clinical benefit under conventional (i.e., 5% significance level) standards was limited. In each of these cases, the justification emphasized the urgent demand for new therapies, highlighting how welfare considerations are frequently used to justify the departure from strict adherence to traditional statistical thresholds. Moreover, even if there were consensus on the appropriate size of a test, designing experiments solely to maximize power may still fall short of achieving optimal outcomes from a societal perspective.

For these and related reasons, the FDA has recently promoted a transition toward Bayesian clinical trials. This stance builds on the commitments to Bayesian designs established by the Prescription Drug User Fee Act and the 21st Century Cures Act, both enacted by the US Congress in 2016, and has been codified in new draft guidance (us2026guidance) intended “to encourage the use of Bayesian statistics in clinical trial design and the readout of results” (doshi2026fda).\footnote{In an accompanying 1.09-minute video (Makary2026) shared on X on 12 January 2026, FDA commissioner Marty Makary stated: “I'm pleased to announce the FDA is now open to Bayesian statistics. We are putting out new guidance to encourage the use of Bayesian statistics in clinical trial design and readout of results... it is a leap forward beyond the frequentist model of analyzing data".} In parallel, the European Medicines Agency has also committed to releasing a draft `reflection paper' outlining its position on Bayesian clinical trials by 2027.

Within this move toward Bayesian inference, the “statistical threshold” for regulatory approval is usually defined as the posterior probability that the treatment effect exceeds a pre-specified value. For example, the Pfizer-BioNTech COVID-19 Vaccine (Comirnaty) was authorized because the posterior probability that its vaccine efficacy was greater than 30% surpassed a 99.5% threshold. Nevertheless, much like the conventional 5% significance level, it remains unclear how such thresholds ought to be selected, and it is also unlikely that a universal threshold of this kind would be socially optimal.

The choice of decision thresholds and associated stopping rules is particularly crucial in the context of adaptive clinical trials. In fact, encouraging the use of adaptive trials has been a key objective of the FDA since 2004, when it launched the `Critical Path Initiative' to increase the efficiency of drug development (see, e.g., the draft guidance on adaptive trials in us2019guidance). Since pharmaceutical companies propose the trial design, the combination of Bayesian and adaptive elements expands the scope for Bayesian persuasion---an inherent advantage experimenters enjoy in strategic settings due to their control over the information structure. If there were no restrictions on the kinds of designs that could be used, experimenters would always choose to end the trial exactly when regulators are just indifferent between approving and rejecting it, i.e., as soon as the estimated treatment effect becomes (barely) positive. As a result, in the absence of such limits, the overall social value the FDA derives from these trials could actually be lower than what is already achieved under traditional, non-adaptive Randomized Controlled Trial (RCT) designs.

In this article, we advocate using welfare maximization as the primary criterion for approving experimental designs. Concretely, we propose that the regulator impose a minimum requirement on the expected social welfare that a proposed design must generate in order to be approved. The experimenter, in turn, selects a design that maximizes their own objective, subject to this welfare constraint. We characterize the form of optimal experimental designs within this framework, explicitly allowing for misaligned preferences between experimenters and regulators.

Our proposal has several key features that make it particularly suitable for real-world application. First, imposing a welfare threshold curtails the scope for Bayesian persuasion by bounding the experimenter's ability to bias outcomes through strategic evidence design. We further recommend a natural benchmark for calibrating this threshold: the expected welfare that a regulator would obtain under the Bayesian prior if the experimenter were limited to running a conventional RCT.

Moreover, our proposal does not require the regulator to know the experimenter's preferences or the costs of running the experiment. This is crucial, because such information is typically unverifiable and experimenters have strong incentives to misrepresent it. Indeed, we suspect that the widespread adoption of the 5% significance level in hypothesis testing may arise precisely from the difficulty of eliciting true preferences and cost structures. While some recent decision-theoretic rationales (e.g., tetenov2016; viviano2021model) for hypothesis testing tie the choice of test size directly to experimental costs, such links are not usually observed in real-world applications.

In our framework, we allow the experimenter to jointly optimize the sampling allocation and the stopping rule; the latter endogenously determines the total sample size. Utilizing a continuous-time framework under Gaussian priors, we characterize an optimal experimentation policy that exhibits several notable properties.

First, the optimal sampling rule is always the Neyman allocation, just as in classical designs, regardless of the specific utility functions of the experimenter or the regulator. This result is based on an extension of liang-mu-syrgkanis-ecta-2022 to a multi-agent setting. The optimality of Neyman allocation arises from a fundamental alignment of interests regarding informational efficiency: despite their divergent ultimate objectives, both the experimenter and the regulator seek to minimize the posterior variance of the treatment effect relative to the cost of experimentation. Consequently, the allocation of subjects across treatment arms is governed solely by variance-minimization.

Second, the optimal stopping strategy diverges from the single-agent benchmark if and only if the experimenter derives an additive benefit from approval that is independent of the treatment effect's magnitude. In the presence of this benefit, the optimal stopping boundaries become asymmetric: the threshold for approval is shifted downward by a constant factor relative to the rejection threshold, reflecting the experimenter's systematic bias toward treatment adoption.

Third, the optimal stopping rule---and by extension, the entire experimental strategy---is invariant to the specific weighting of Type I (approving the treatment when the true treatment effect is negative) and Type II (rejecting the treatment when the true treatment effect is positive) errors. The structural form of the optimal experimentation strategy thus remains unchanged whether the regulator and experimenter are expected utility maximizers or regret-minimizers. This robustness stems from the fact that while error-weighting shifts the relative costs of misclassification, it does not alter the fundamental trade-off between the marginal cost of sampling and the marginal informational gain.

In a numerical analysis with priors calibrated using clinical trial data, we show that our proposed strategies reduce expected sample sizes by more than 48% compared with traditional RCT designs that achieve the same welfare level. For pharmaceutical firms, this implies nearly \$6 million in savings on experimentation costs.

While the preceding results are derived within a continuous-time framework, we demonstrate that the finite-sample counterparts of these strategies remain asymptotically optimal under `small-cost asymptotics' (adusumilli2025).

Related literature

If we remove the welfare constraint, the framework presented here collapses to the classic model of costly Bayesian persuasion (kamenica-gentzkow). It is precisely to emphasize this link that we refer to it as persuasive experimental design.

Our results add to the broad literature on information acquisition and optimal stopping with costly sampling. The seminal contributions of wald and arrow derive optimal stopping rules for welfare maximization in binary state spaces. Bather_Walker_1962 study the more general problem of selecting an optimal stopping time for a Brownian motion, where the cost of any path depends solely on its terminal time and position. More recent work has shifted toward richer distributional frameworks and more flexible experimentation. fudenberg-strack-strzalecki2018 characterize the welfare-maximizing stopping rule for choosing between two treatments under Gaussian priors. liang-mu-syrgkanis-ecta-2022 extend these insights by endogenizing the sampling strategy and demonstrate that Neyman allocation is optimal under symmetric Gaussian priors, regardless of the experimenter's utility function. A common element across these contributions, however, is the focus on a single decision-maker. Our work adds to this literature by determining the optimal experimentation policy in a multi-agent environment in which the experimenter and the regulator have conflicting objectives.

This article thus also contributes to a growing line of economic research that examines how incentives and the presence of multiple agents shape empirical work. andrews2021model study how statistical findings are communicated to an audience with diverse priors. tetenov2016, viviano2021model, and bates2023incentive investigate hypothesis testing when incentives are misaligned and emphasize how controlling size can help `screen out' experimenters who aim to put forward low-quality treatments. In our setting, the imposition of a prior and a welfare constraint serves a similar purpose, indirectly functioning as a screening mechanism for undesirable participants. In addition, we study the full experimental design problem---including sampling and stopping rules---rather than focusing exclusively on approval thresholds.

In more closely related work, yoder2022designing studies settings in which a principal hires an experimenter to perform research but is unable to contract on the experimenter's private type (such as their utility function, research costs, and so forth). This differs from our setup, where the experimenter undertakes research with the goal of persuading a regulator; however, we do share the same contracting constraints. henry2019research examine a related environment in which the experimentation strategy is determined as the Nash equilibrium of a game between the experimenter and the regulator. This differs from our framework, where the regulator is not permitted to directly optimize against any particular experimenter: the same constraints must apply to all experimenters, and the regulator cannot write contracts contingent on the experimenter's type. Consequently, in our setting, the regulator is only able to impose a minimum welfare threshold. Moreover, our state space is substantially richer than in henry2019research, featuring two continuous treatments rather than binary states, and the decision space encompasses sampling strategies in addition to optimal stopping.

Our analysis of the continuous-time experimental design problem makes use of the formalism introduced in adusumilli-continuous-arts. We then rely on continuous-time asymptotic representations, also developed in that work, to translate the results to discrete-time environments. The particular asymptotic regime we employ is based on `small-cost asymptotics', introduced in adusumilli2025. In this setting, one assumes that the marginal cost of each observation is negligible relative to the benefit of approval, with the cost-benefit ratio converging to zero at a specified rate. We demonstrate that this assumption is supported by empirical estimates from the clinical trial literature. More broadly, small-cost asymptotics fall under the umbrella of local asymptotics, which analyze local perturbations around reference parameter values. Other recent studies that use local asymptotic techniques to analyze adaptive experiments include wagerxu, fanglynn, and hirano2025asymptotic.

General Setup

We start by presenting a general framework for studying experimental design when the experimenter and the regulator have divergent objectives. Since the canonical application of this framework is the design of Phase 3 clinical trials, we use this as a running example throughout this article.

The model involves two agents: a regulator (Alice) and an experimenter (Bob). Bob proposes a project (e.g., a new drug) whose utility depends on an underlying state of the world $s \in \mathcal{S}$. A simple example could be the binary state space $\mathcal{S} = \{0, 1\}$, where $s=1$ indicates an effective drug and $s=0$ indicates an ineffective one; but ultimately, we are interested in continuous $s$.

Both agents are expected utility maximizers. Let $\Delta(\mathcal{S})$ denote the set of all beliefs, i.e., probability measures on $\mathcal{S}$, endowed with the metric of weak convergence. The space of distributions over beliefs (i.e., distributions over probability measures) is represented by $\Delta(\Delta(\mathcal{S}))$.

Experimental design

We assume that both Alice and Bob agree on a prior $p_0$. To convince Alice to approve his proposal, Bob submits an experimental protocol. If Alice approves the protocol, the experiment is conducted, yielding a signal that allows both agents to update their beliefs, ex-post, to a posterior $p \in \Delta(\mathcal{S})$.

Employing the terminology of Blackwell experiments, we characterize an experiment by the ex-ante distribution over posterior beliefs $Q \in \Delta(\Delta(\mathcal{S}))$ that it induces. Posterior consistency demands that $Q$ satisfy

equation[equation omitted — 73 chars of source]

This is the minimal and only constraint on $Q$ that experimentation imposes (at this point, we leave the class of possible experiments unrestricted and allow for garbling of signals).

Utilities and welfare

After the experiment ends, Alice selects an action: accept ($\delta = 1$) or reject ($\delta = 0$). Alice's utility, $u_A(s, \delta)$, is determined by the state $s$ and her chosen action. For a given posterior $p$, Alice selects $\delta^*(p) \in \arg\max_{\delta \in \{0,1\}} \int u_A(s, \delta) \, dp(s)$, which yields the following ex-post expected welfare: $$ V_A(p) \coloneq \max_{\delta \in \{0, 1\}} \int_{\mathcal{S}} u_A(s, \delta) \, dp(s). $$

Whenever Alice is indifferent between $\delta \in \{0,1\}$, we adopt the tie-breaking convention that she chooses Bob's preferred action, that is, $\delta = 1$. In fact, as shown in Section (ref), when information arrives gradually, we can, without mathematical loss of generality, suppose that Alice selects Bob's preferred action whenever she is indifferent, no matter what her actual tie-breaking rule may be.

Bob's overall utility, which is of the form $$ u_B(s,\delta) - C(Q), $$ consists of two components. The first component, $u_{B}(s,\delta)$, depends on the state of the world and Alice's action. For example, Bob may receive a constant benefit $B$ if the drug is approved: $u_B(s, \delta) = B \cdot \mathds{1}\{\delta = 1\}$. The second component captures the cost of the experiment, which is borne by Bob. Following denti2022experimental, we model this ex-ante, as a function, $C:\Delta(\Delta(\mathcal{S})) \to [0,\infty)$, of the distribution over posteriors $Q$. Intuitively, $C(Q)$ represents the least-expensive way for Bob to run a specific experiment---which corresponds to some unique distribution over posteriors $Q$---with the convention that $C(Q) = \infty$ if the posterior distribution is unattainable, e.g., due to limitations on what experiments are allowed or are feasible.

Often, the cost $C(Q)$ is posterior separable, i.e., it can be written as \[ C(Q) = \int\phi(p)dQ(p) - \phi(p_0), \] where $\phi(\cdot)$ is convex. Even though we measure information cost in an ex-ante sense, it is known that posterior separable costs can be induced by specific running cost functions (mor). Under a posterior separable cost, Bob's ex-ante expected welfare, for a given distribution over posteriors $Q$, is given by: $$ \int [V_B(p) - \phi(p)] \, dQ(p), \quad \text{where} \\ \quad V_B(p) \coloneq \int u_B(s, \delta^*(p)) \, dp(s). $$

The regulator constraint and Bob's optimization problem

Given the misalignment of incentives between Alice and Bob, it is suboptimal for Alice to unconditionally approve Bob's experimental protocols, as this allows Bob to leverage Bayesian persuasion to his advantage. At the same time, as in yoder2022designing, contracting directly on Bob's utility or cost functions is infeasible for two reasons. First, these parameters constitute private information and are subject to potential misrepresentation. For instance, Bob may claim that an experiment is more expensive than it really is in order to induce Alice to accept smaller sample sizes. Second, legal regulations often obligate Alice to impose identical conditions on all experimenters, which prevents her from tailoring her requirements specifically to Bob.

Given these constraints, we propose that Alice impose a minimum threshold $V_0$ for the ex-ante expected welfare. Specifically, we introduce the regulator constraint:

equation[equation omitted — 81 chars of source]

and suggest that Alice commit to clearing any experimental proposal submitted by Bob that satisfies this condition. The precise value of $V_0$ depends on how Alice wishes to split the welfare surplus from experimentation with Bob (and other potential experimenters). Among other things, this would depend on the weights Alice places on Bob's profits, incentives for research, and overall consumer welfare. Here, we abstract from these issues and treat $V_0$ as given. In fact, many of our policy implications turn out to be independent of $V_0$, but we also show how it can be calibrated using past data on clinical trials.

Given Alice's requirements, Bob's actions reduce to choosing an experiment $Q$ that maximizes his ex-ante expected welfare, subject to the martingale and regulator constraints ((ref)) and ((ref)):

align[align omitted — 184 chars of source]

Note that when $V_0 = -\infty$ (no regulatory oversight), this model reduces to costly Bayesian persuasion (kamenica-gentzkow).

A dual formulation of Bob's problem

When experimentation costs are posterior separable, Bob's experimental design problem admits the dual representation

align*[align* omitted — 200 chars of source]

where $L_{1}(\mathcal{S})$ is the set of all Lipschitz continuous functions on $\mathcal{S}$. The above is a slight extension of kolotilin2025persuasion.

The function $f(s)$ can be interpreted as the shadow price associated with state $s$. We may view the dual problem as describing a decentralized economy in which Alice and Bob outsource the creation of welfare to an entrepreneur, Carol, who charges a price $f(s)$ for each latent state $s$. Alice and Bob pay Carol an ex-ante fee of $\int f(s)\,dp_0(s)$, and in return, Carol commits to delivering at least $[V_B(p) - \phi(p)]+ \gamma V_A(p)$ in every possible contingency where the posterior is $p$. The solution to the dual problem then characterizes the minimum ex-ante payment Carol can charge Alice and Bob without suffering an expected loss.

The dual problem greatly simplifies the determination of the optimal strategy when $\mathcal{S}$ is discrete since, in that setting, it reduces to a linear programming problem. In this article, however, our main focus is on the case where $s$ is continuous, and the functions $V_A(\cdot)$ and $V_B(\cdot)$ are also difficult to evaluate explicitly. As a result, the dual problem is of limited practical use for our purposes, but we present it here because it is of conceptual interest.

Persuasive Experimental Design under Incremental Learning

We now focus on a specific class of problems that compare a candidate treatment to an active control or placebo. Unfortunately, solving the experimental design problem ((ref)) in general contexts is still infeasible, especially when $s$ is continuous. We therefore make three further restrictions: (1) information is constrained to arrive incrementally in the form of Gaussian processes, (2) the cost of sampling is proportional to the number of observations, and (3) the priors are Gaussian.

The first two restrictions are relatively mild. As we demonstrate in Section (ref), incremental signal processes can be formally motivated using local asymptotics. The assumption of linear sampling costs follows a long tradition in sequential analysis dating back to wald; see also arrow, fudenberg-strack-strzalecki2018, and adusumilli2025. In practice, clinical trials are often conducted in batches, and we can think of our framework as assuming that the marginal cost of each batch remains approximately constant. Furthermore, while firms do incur significant fixed costs to initiate trials, these are essentially sunk costs that do not influence optimal experimental design (although they may shift how Alice and Bob split the welfare surplus from experimentation).

The restriction to Gaussian priors is motivated by both technical and practical considerations. Formally, this assumption allows us to make use of the results of liang-mu-syrgkanis-ecta-2022 to characterize the optimal treatment assignment rule under very general conditions. Because the resulting assignment rule is history-independent and static, the complex dynamic optimization problem reduces to determining the optimal stopping rule. Empirically, meta-analyses of large-scale clinical trial databases suggest that the distribution of realized treatment effects is well-approximated by a Gaussian prior. Also, as suggested by the FDA draft guidance on Bayesian clinical trials (us2026guidance), the prior could be estimated from Phase 2 trial results; by the central limit theorem, this is typically well-approximated by a Gaussian distribution.

Setup under incremental learning

As discussed above, we consider a special case of the general setup from Section (ref), where the aim is to determine the best option out of two treatments $a \in \{0,1\}$. These treatments are associated with unknown mean rewards $\mu_0, \mu_1$ and known variances $\sigma^2_0, \sigma^2_1$. The state of the world is therefore $s = (\mu_1, \mu_0)$. We assume that both Bob and Alice agree on a Gaussian prior $(\mu_1, \mu_0) \sim \mathcal{N}(\boldsymbol\mu^0, \Sigma)$.

We employ the formalism of adusumilli-continuous-arts to describe adaptive experiments under incremental learning. The underlying information environment consists of two independent Gaussian signal processes corresponding to each treatment $a$, given by

align[align omitted — 65 chars of source]

where $W_0(\cdot), W_1(\cdot)$ are independent standard Brownian motions. Intuitively, $z_a(\gamma)$ corresponds the cumulative sum of outcomes generated by treatment arm $a$ when it is sampled for a duration $\gamma$. The signal processes $\{z_a(\cdot)\}_a$ are supplemented by an exogenous random variable \(U \sim \text{Uniform}[0,1]\), which accounts for the randomness inherent in any potentially randomized experimentation strategy. Let $\mathbb{P}$ denote the joint probability distribution over the state space and the sample paths, composed of the prior $p_0$ over $s$, together with the conditional law of $\{z_1(\cdot), z_0(\cdot), U\}$ given $s$. Also, take $\mathcal{H}^{(a)}_\gamma \coloneq \sigma\{z_a(r) : r \le \gamma \}$ to be the natural filtration generated by the sample paths of $z_a(\cdot)$ between $0$ and $\gamma$, and set:

equation[equation omitted — 135 chars of source]

The experimental design consists of specifying both a sampling strategy and a stopping time. Following adusumilli-continuous-arts, we represent the sampling strategy by an allocation process $\bm{q} \equiv \{q_a(t)\}_a$, which records the cumulative amount of attention assigned to treatment $a$ up to time $t$. For example, if $q_a(t) = \gamma_a$, then at time $t$ we have observed the path of $z_a(t)$ over the interval $[0,\gamma_a]$. We work with allocation processes rather than standard sampling rules or policies because the latter are generally not measurable in continuous time. The requirements of an allocation process are that $\{q_a(\cdot)\}_a$ are monotone, satisfy $q_1(t) + q_0(t) = t$ for all $t$, and obey the informational constraint that the event $\{q_1(t) \le \gamma_1,\; q_0(t) \le \gamma_0\}$ is $\mathcal{G}_{\gamma_1,\gamma_0}$-measurable for every choice of $\gamma_1,\gamma_0,t$. The observed signal at time $t$ is given by $$ x_a(t) \coloneq z_a(q_a(t)), $$ and the information available at time $t$ under a given experiment is summarized by $\mathcal{F}_t^{\bm{q}} = \mathcal{G}_{q_1(t), q_0(t)}$. The filtration $\mathcal{F}_t^{\bm{q}}$ is indexed by $\bm{q}$ because it depends on the particular sampling strategy in use. A stopping time $\tau$ that is adapted to $\mathcal{F}_t^{\bm{q}}$ then specifies when the experiment terminates. Since $\mathcal{F}_t^{\bm{q}}$ incorporates the exogenous randomization $U$, the framework accommodates randomized stopping times. We denote the overall experimentation strategy, combining both the allocation process and the stopping time, by $\bm{d} = \{\{q_a(\cdot)\}_a, \tau \}$.

Following the experiment, Alice chooses an implementation rule $\delta \in \{0, 1\}$. Given an experimentation strategy $\bm{d}$, we represent Alice's (Bayes) optimal response by $$ \delta^*_{\bm{d}} := \mathds{1}\left\{ \mathbb{E}\left[u_A(s,1) \vert \mathcal{F}_\tau^{\bm{q}} \right] \ge \mathbb{E}\left[u_A(s,0) \vert \mathcal{F}_\tau^{\bm{q}} \right] \right\}. $$

Since the priors are Gaussian, standard results in stochastic filtering imply that the posterior distribution $p^{\bm{q}}(t)$ of $s$ induced by $\bm{q}$ is also Gaussian and is given by $$ p^{\bm{q}}(t) \propto \left[ \prod_{a\in \{0,1\}} \phi \left(x_a(t) \vert \mu_a q_a(t), \sigma_a^2q_a(t) \right) \right] \times \phi\left(\mu_1, \mu_0 \vert \bm{\mu}^0, \Sigma \right), $$ where $\phi(\cdot|m, A)$ represents the normal pdf with mean $m$ and covariance matrix $A$.

Denote the set of all distributions over posteriors inducible through incremental learning by $ \mathcal{Q} \equiv \{\textrm{ex-ante law of }p^{\bm{q}}(\tau)\}_{\bm{d}}. $ As noted earlier, we simplify the structure of experimentation costs by assuming it is given by $c\tau$ for any stopping time $\tau$, where $c > 0$ is some constant denoting the marginal cost of each observation. We can relate this to the ex-ante cost of information $C(\cdot)$ from Section (ref) as $$ C(Q) =

cases\inf_{\{(\bm{q},\tau): p^{\bm{q}}(\tau) \sim Q\}} c\tau &if $Q\in \mathcal{Q}$\\ \infty &otherwise.

$$ Apart from this structure on $C(\cdot)$, our setup is otherwise the same as Section \ref{Sec:General_setup}; indeed, all the constraints on experimentation are subsumed into the form of $C(\cdot)$.

Characterizing the optimal sampling strategy

We now characterize the optimal sampling strategy. As we demonstrate below, this can be achieved under quite broad conditions.

In what follows, let $\mathbb{P}_{\bm{d}}$ represent the restriction of the measure $\mathbb{P}$ to the $\sigma$-algebra $\mathcal{F}_\tau^{\bm{q}}$, and $\mathbb{E}_{\bm{d}}[\cdot]$ its associated expectation (this is the probability induced by the sampling strategy).

asm(i) The utility functions $u_A(\cdot), u_B(\cdot)$ depend on $s$ only through $\mu_1 - \mu_0$.\\ (ii) The functions $\bar{u}_A(s)\coloneq \vert\max_a u_A(s,a)\vert$ and $\bar{u}_B(s)\coloneq\vert\max_a u_B(s,a)\vert$ are integrable with respect to the prior $p_0$, and are also bounded whenever $\vert s \vert$ is bounded.\\ (iii) There exists some experimentation strategy $\bm{d}$ such that $\mathbb{E}_{\bm{d}}[u_A(s, \delta^*_{\bm{d}})] \geq V_0$ and $\mathbb{E}_{\bm{d}}[u_B(s, \delta^*_{\bm{d}}) -c\tau] \ge -L$ for some $L < \infty$.\\ (iv) The priors over $\mu_1, \mu_0$ are independent, i.e., $\mathcal{N}(\boldsymbol\mu^0, \Sigma) = \mathcal{N}(\mu_1^0, \Sigma_{11}) \times \mathcal{N}(\mu_0^0, \Sigma_{00})$. Furthermore, $\Sigma_{11}/\sigma_1^2 = \Sigma_{00}/\sigma_0^2$.

Assumption (ref)(i) posits that Alice and Bob care exclusively about the relative difference between $\mu_1$ and $\mu_0$ rather than their absolute levels. In effect, this serves as a standard normalization for welfare that simplifies the parameter space. Assumption (ref)(ii) is a mild regularity condition on the growth rates of $u_A(\cdot), u_B(\cdot)$; given that a Gaussian prior possesses exponential tails, this condition is satisfied by most standard utility functions, which generally exhibit sub-exponential growth relative to $|s|$. Assumption (ref)(iii) represents Bob's participation constraint by ensuring that his loss is bounded from below. Assumption (ref)(iv) states that $\mu_1/\sigma_1$ and $\mu_0/\sigma_0$ have the same (independent) prior. It is introduced primarily to simplify the algebra because, as we demonstrate below, it implies that the Neyman allocation is always optimal. The case of more general covariance structures is analyzed in Appendix (ref). Even in this broader setting, the optimal sampling rule is qualitatively similar: it prescribes sampling from a single treatment for some initial period of time before switching to the Neyman allocation.

Under Assumption (ref), we can rewrite Bob's experimental design problem ((ref)) as

equation[equation omitted — 248 chars of source]

The (Lagrangian) dual of the problem is

align[align omitted — 228 chars of source]

The dual problem is very similar to the class of information acquisition problems analyzed by liang-mu-syrgkanis-ecta-2022. It is, therefore, a lot easier to analyze than the primal.

thmSuppose Assumption (ref) holds. Then: \begin{enumerate}[(i)] • The primal and dual problems, ((ref)) and ((ref)), attain the same bounded value. • The optimal sampling strategy is the Neyman allocation: $$ q_a^*(t) = \frac{\sigma_a}{\sigma_1 + \sigma_0}t \ \forall\ t. $$ • Under the Neyman allocation, the posterior variance $\varrho_t^*$ of $\mu_1 - \mu_0$ is deterministic and given by $$ \varrho_t^* = \frac{\sigma^2}{\sigma^2 \varrho_0^{-1} + t}, $$ where $\sigma^2 \coloneq (\sigma_1 + \sigma_0)^2$ and $\varrho_0$ is the prior variance of $\mu_1 - \mu_0$. Furthermore, the posterior mean, $m_t$, of $\mu_1 - \mu_0$ evolves as \begin{equation} m_t = \mu_1^0 - \mu_0^0 + \int_0^t \frac{\sigma^{-1}}{\sigma^{-2}t + \varrho_0^{-1}} dW(t), \end{equation} where $W(\cdot)$ is standard Brownian motion. \end{enumerate}

The first part of Theorem (ref) is new to this paper. We demonstrate the equivalence between the primal and dual problems by establishing a minimax theorem. The proof is subtle because the space of decisions $\bm{d}$ is quite rich and infinite dimensional.

The second and third parts are extensions of results obtained by liang-mu-syrgkanis-ecta-2022. They are based on analyzing the dual problem. Remarkably, the second part states that the optimal sampling strategy is independent of both sampling costs $c$ and the form of the utility functions $u_A(\cdot), u_B(\cdot)$. The strategy is also deterministic and extremely simple: one should always employ the Neyman allocation. The Neyman allocation is optimal because it minimizes the posterior variance of $\mu_1 - \mu_0$ at every possible time point. Intuitively, it is optimal to reach precise beliefs, even unfavorable ones, as quickly as possible to minimize experimentation costs. This is achieved by choosing the strategy that minimizes uncertainty.

To understand the third part of Theorem (ref), observe that under the Neyman allocation and Assumption (ref)(i), the process

equation[equation omitted — 187 chars of source]

is a sufficient statistic for $\mu_1 - \mu_0$. Indeed, representing the allocation process under the Neyman allocation by $\bm{q}^*$, the posterior distribution of $\mu_1 - \mu_0$ is given by

equation[equation omitted — 242 chars of source]

Equation ((ref)) then follows from standard results in stochastic filtering.

Characterizing the optimal stopping time

Utility specifications

In contrast to the optimal sampling strategy, the optimal stopping rule inherently depends on the specific utility functions of Alice and Bob. For tractability, we therefore adopt specific functional forms for $u_A(\cdot)$ and $u_B(\cdot)$.

asmAlice's utility function is of the form: \[u_A(s,\delta) \coloneq u(\mu_1 - \mu_0,\delta;\alpha) = \begin{cases} \alpha(\mu_1 - \mu_0),\ \text{if }\delta = 1,\\ \ (1- \alpha)(\mu_0-\mu_1),\ \text{if }\delta=0, \end{cases} \] where $\alpha \in [0,1]$ is a known parameter.

Assumption (ref) specifies that Alice's utility is linear in $|\mu_1 - \mu_0|$, conditional on the decision $\delta$. Within this framework, Alice faces two distinct errors: a Type I error occurs when approving a treatment ($\delta = 1$) despite $\mu_1 < \mu_0$, and a Type II error occurs when rejecting a treatment ($\delta = 0$) when $\mu_1 > \mu_0$. We allow for asymmetric utility across these error types, with the degree of asymmetry governed by the parameter $\alpha$. Two cases are particularly significant: $\alpha = 1$ recovers the standard utilitarian welfare, whereas $\alpha = 1/2$ is equivalent to the negative of welfare-regret.

Under Assumption (ref), Alice's optimal action is to choose the treatment with the higher posterior mean, i.e. $\delta^* = \mathds{1}[m_\tau \geq 0]$. This results in an ex-post Bayes welfare of $\max\{\alpha m_\tau, -(1-\alpha)m_\tau\}$.

Regarding Bob's preferences, we assume he derives some utility solely from the approval of the treatment.

asmBob's utility is of the form: $$ u_B(s,\delta) = B\mathds{1}[\delta = 1] + \gamma u(\mu_1 - \mu_0,\delta;\alpha^\prime), $$ where $\gamma, B$ are positive quantities, and $u(s,\delta;\alpha^\prime)$ follows Alice's functional form with a potentially different asymmetry parameter $\alpha^\prime$.

Beyond the fixed benefit $B$, Assumption (ref) allows Bob to derive additional utility (or disutility) from the magnitude of the treatment effect conditional on approval. For instance, setting $\alpha' = 1$ implies $u(\mu_1 - \mu_0, \delta; \alpha') = (\mu_1 - \mu_0)\delta$. However, as we demonstrate below, the fundamental incentive misalignment between Alice and Bob stems exclusively from the additive benefit $B$. While the players may weigh Type I and Type II errors differently (when $\alpha \neq \alpha'$), this preference heterogeneity does not alter the structural form of the optimal stopping rule.

The intuition is as follows: even with divergent $\alpha, \alpha'$, both players share an underlying incentive to learn $s$ as precisely as possible, subject to the costs of experimentation. While they may disagree on how to distribute the welfare surplus from experimentation, the shared incentive for learning $s$ implies that the basic form of the optimal stopping rule remains constant when $B$ is fixed. Indeed, setting $B=0$ reduces the model to the single-player experimentation setup in fudenberg-strack-strzalecki2018.

To formally understand how these properties come about, consider the Lagrangian (dual) objective, which represents a weighted average of the players' utilities: $$ u_B(\cdot, \delta) + \lambda u_A(\cdot, \delta) = B \cdot \mathds{1}[\delta = 1] + \gamma u(\cdot, \delta; \alpha') + \lambda \left( u(\cdot, \delta; \alpha) - V_0 \right). $$ It is easy to verify that there exist “effective” parameters $\tilde{V}_0 \in (0, \infty)$, $\tilde{\lambda} \in [0, \infty)$, and $\tilde{\alpha} \in [-1, 1]$ such that the combination of the players' error-weighted utilities maps to a single representative utility function: $$ \gamma u(\cdot, \delta; \alpha') + \lambda u(\cdot, \delta; \alpha) = \tilde{\lambda} (u(\cdot, \delta; \tilde{\alpha}) - \tilde{V}_0). $$ Consequently, the solution to the dual problem is observationally equivalent to a model where $\gamma = 0$ and $\alpha = \alpha' = \tilde{\alpha}$. Moreover, Lemma (ref) below establishes that the optimal solution is invariant to the specific choice of $\alpha$ for a given $\lambda$.

Variations in $\gamma, \alpha,$ and $\alpha'$ thus influence the experimental design only in an indirect manner, via their possible influence on the calibration of $V_0$, which in turn modifies the optimal choice of $\lambda$. Because our theoretical results already describe the optimal solution for the full range of $\lambda$ values, tracking these parameters separately does not yield any further explanatory insight. Consequently, for the rest of the analysis, we can, without loss of generality, impose $\gamma = 0$.

Properties of the optimal stopping time

Despite the functional form restrictions in Assumptions (ref) and (ref), it is well known that a closed-form characterization of the optimal stopping time is intractable under a Gaussian prior (Bather_Walker_1962; fudenberg-strack-strzalecki2018). Instead, we follow the usual approach of identifying the optimal stopping rule with the solution to a boundary problem and use this to characterize its properties.

Consider the dual problem ((ref)) with a fixed value of the multiplier $\lambda$. The optimal stopping time $\tau$ then solves:

equation[equation omitted — 173 chars of source]

Because $\lambda$ is fixed, we omit the $V_0$ term in what follows. In Section (ref), we show how $\lambda$ can be numerically calibrated given a level of $V_0$. By standard results in convex optimization, each strictly positive $\lambda$ corresponds to a unique $V_0$ value when the duality gap is 0 and the primal problem has a bounded value (both these conditions are verified in Theorem (ref)).

Since the posterior distribution of $\mu_1 - \mu_0$ is Gaussian, it is uniquely characterized by its posterior mean and variance. Under the Neyman allocation, the evolution of the posterior mean $m_t$ is described in Theorem (ref), while the posterior variance is a deterministic function of time $t$. Consequently, the sufficient statistics for optimal stopping are $(t, m_t)$. Define $$ S_\alpha(x) = \max\{x,0\} - (1-\alpha)x, $$ and let $V(t,m)$ denote the continuation value in state $(t,m)$ when $m$ is the current value of the posterior mean:

equation[equation omitted — 158 chars of source]

As in oksendal2013stochastic, the optimal stopping rule for Markov problems of this kind has the form

equation[equation omitted — 132 chars of source]

Equivalently, we may define two time-dependent boundaries and take the optimal stopping rule, $\tau^* = \min\{\tau^+,\tau^-\}$, to be the first exit time from these boundaries:

align[align omitted — 294 chars of source]

Recall that we chose to set $\gamma = 0$ without loss of generality. As referenced earlier, a remarkable feature of the boundaries $b^+(t), b^-(t)$ is that they are the same for any $\alpha \in [0,1]$. Thus, we need only express their properties under $\alpha = 1$ or $\alpha = \frac{1}{2}$, whichever is convenient.

lemLet $b^+(t;\alpha), b^-(t;\alpha)$ be the stopping boundaries given $\alpha \in [0,1]$. Then, \[b^+(t;\alpha) = b^+(t;1),\quad b^-(t;\alpha) = b^-(t; 1).\]

Lemma (ref) implies that allowing for asymmetric welfare under approval and rejection leaves the optimal experimental design unchanged for a given $\lambda$. It further allows us to extend the results of fudenberg-strack-strzalecki2018 beyond the welfare-regret objective. Intuitively, while $\alpha$ controls the `tilt' of the welfare function, this tilt does not modify the relative difference between the options' means, and the optimal boundaries thus remain unaffected.

Since the optimal $\lambda$ is uniquely determined by the value of $V_0$, it follows that, conditional on knowing $V_0$, one does not need to know $\alpha$. In the numerical analysis (Section (ref)), we calibrate $\lambda$ through $V_0$ so that it matches the expected welfare of a Randomized Controlled Trial (RCT) with a fixed sample size ($t = 1$). In fact, when $m_0 = 0$, we obtain the striking result that the welfare generated by the RCT---and therefore the welfare under the calibrated stopping rule---remains the same for every $\alpha \in [0,1]$.

We now characterize several additional properties of the stopping boundaries.

thmUnder Assumptions 1-3, $(b^+(t),b^-(t))$ have the following properties: \begin{enumerate}[(i)] • $b^+(t), b^-(t)$ are well defined, with $ b^+(t), \vert b^-(t)\vert < \infty$. • $b^+(t)$ and $|b^-(t)|$ are weakly decreasing in $t$, with $\lim_{t\to\infty}b^+(t) = \lim_{t\to\infty}b^-(t)=0$. • $b^+(t) \leq |b^-(t)|$. Also, whenever $B > 0$, $b^+(t) = 0\ \forall\ t\ge t^*$, where $t^*$ is such that $b^-(t^*) = -B/\lambda$. • $b^+(t), b^-(t)$ are continuous in $t$, with the latter being Lipschitz continuous. • $b^+(t)$ and $|b^-(t)|$ are weakly increasing in $\lambda$. • $b^+(t), b^-(t)$ are weakly decreasing in $B$ for a fixed $\lambda$. • $b^-(t)$ is strictly less than 0. \end{enumerate}

Theorem (ref) is based on the idea that the stopping rule identifies the smallest treatment-effect magnitude at which stopping weakly dominates continuation.

Parts (ref) and (ref) are technical results stating that the lower and upper stopping boundaries, $b^-(t)$ and $b^+(t)$, are well defined and continuous (in fact, $b^-(t)$ is Lipschitz continuous). Part (ref) further establishes that these boundaries converge monotonically to zero over time. The intuition is that the incremental value of new information shrinks as beliefs about the true mean become more precise. Since the sampling cost remains constant, there comes a point at which it is optimal to stop for any given belief. This pattern is driven by the Gaussian prior, which assigns little prior mass to large values of $s$. By contrast, adusumilli2025 shows that with a two-point prior, beliefs can remain as uninformative in the future as they are at present.

Part (ref) shows that the stopping thresholds are asymmetric: in absolute value, the acceptance threshold is always lower than the rejection threshold. This asymmetry is driven by Bob's private benefit $B$, which gives him an incentive to continue the experiment even when early signals are weakly favorable, as he hopes for a possible reversion to a positive mean. Part (ref) further implies that $b^+(t)$ eventually hits zero. Under standard Bayesian persuasion, which corresponds to a nonbinding welfare constraint ($\lambda = 0$), the approval threshold would be $b^+(t) = 0$ for all $t$, i.e., the experiment would stop whenever the posterior mean becomes nonnegative. By contrast, for sufficiently large $V_0$, Bob chooses $b^+(t) > 0$ for a finite time interval, though this boundary always lies strictly below the welfare-optimal one, $b^+(t) = -b^-(t)$, corresponding to $B = 0$. Ultimately, however, the Bayesian persuasion motive prevails from $t \ge t^*$ onward.

Since $\lambda$ is in one-to-one correspondence with the welfare constraint $V_0$, part (ref) states that tighter welfare constraints necessarily push the boundaries outward.

Part (ref) is intuitive: for a fixed $\lambda$, a larger benefit $B$ makes waiting less appealing when beliefs are favorable $(m \ge 0)$ and more appealing when beliefs are unfavorable $(m < 0)$. It is important to note, however, that if we hold the welfare constraint $V_0$ constant, an increase in $B$ actually raises the corresponding value of $\lambda$. In combination with part (ref), this suggests that the overall effect of raising $B$ on $b^-(t)$ should be strictly negative: increasing the benefit always makes waiting more attractive. By contrast, the impact on $b^+(t)$ is ambiguous: a lower $\lambda$ tends to reduce $b^+(t)$, so in total $b^+(t)$ may either expand or contract, depending on which mechanism dominates.

On tie-breaking

Up to this point, our analysis has implicitly assumed that Alice breaks ties in favor of Bob's preferred action, $\delta = 1$, whenever she is indifferent, i.e., when $m_\tau = 0$. It turns out, however, that the specific tie-breaking rule Alice adopts is irrelevant. Because $m_t$ can be expressed as a time-changed Brownian motion, it satisfies the immediate-crossing property: whenever the process reaches $0$, it will cross it within an arbitrarily short time interval. Consequently, Bob can always ensure that Alice accepts the proposal as soon as $m_\tau = 0$ (since it implies that $m_t > 0$ almost immediately).

In Section (ref), we develop finite-sample analogs of our optimal policies. In that context, to prevent any ambiguity, we impose that $m_\tau$ be strictly bounded away from 0 by a factor $\xi$. We then demonstrate that $\xi$ can be chosen sufficiently small so that Bob's overall welfare is only minimally affected.

Lessons for experimental design

In the previous sections, we have characterized how experimental design changes when there is incentive misalignment between the experimenter (Bob) and the regulator (Alice). We summarize some key takeaways below.

Neyman allocation is always optimal

Incentive misalignment (or even alignment) does not alter the optimal sampling strategy. Both Alice and Bob share a common objective to learn the state of the world, $s$, as efficiently as possible relative to the cost of experimentation. Consequently, the allocation of subjects across treatment arms is determined by variance-minimization rather than divergence in preferences.

Incentive misalignment only arises from additive benefits

The form of the optimal stopping rule is altered if and only if there exists a private benefit from approval, $B$, that Bob receives regardless of the size of the treatment effect ($\mu_1 - \mu_0$). A careful analysis of Section (ref) and the proof of Lemma (ref) shows that this feature is, in fact, a structural feature of the decision problem, and not an artifact of Gaussian priors: indeed, it is a consequence of the players agreeing on a common prior. In the absence of such an additive benefit ($B=0$), the players' interests are aligned regarding the form of the optimal stopping time, regardless of their individual risk tolerances ($\alpha, \alpha'$) or benefits that are linear in $\vert \mu_1 - \mu_0\vert$.

Asymmetry in decision thresholds

A significant departure from standard sequential designs for clinical trials---such as those discussed in wassmer2016group---is the emergence of asymmetric thresholds for acceptance and rejection. The rejection threshold is shifted downward in absolute terms relative to the approval threshold. This asymmetry reflects Bob's `pro-approval' bias: because he captures a fixed gain from approval, he is incentivized to persist with the experiment even when early evidence is underwhelming, as he hopes for a reversion toward a positive mean.

Terminal convergence to zero of approval thresholds

In classical fixed-sample experiments, where the number of observations is characterized by time $t$, a treatment is approved if its mean effect exceeds a threshold of $1.96\sigma/\sqrt{t}$. In the commonly used group sequential designs of o1979multiple, this threshold declines at a rate proportional to $1/t$. In our framework, the optimal threshold for approval declines even more aggressively, eventually reaching zero at some finite time $t^*$. This rapid convergence is driven by two factors. First, under Gaussian priors, a low posterior mean at a large $t$ suggests the true effect is likely near zero; thus, the potential `downside' from a Type I error diminishes. Second, lowering the threshold for late-stage stopping increases Bob's `option value' of approval, and incentivizes him to continue costly experimentation for a longer duration.

Potential limitations

Our analysis rests on a couple of important structural assumptions, which have economic content.

Common prior

The first assumption is that Alice and Bob share a common prior. Our preferred view of this prior is that it represents an objective quantity that can be estimated from historical data. For instance, in its current guidance on Bayesian clinical trials, the FDA notes that one option is to estimate the prior for Phase 3 designs using data from Phase 2 studies of the same drug (us2026guidance).

In the same document, the FDA also indicates that the prior may alternatively be derived from data obtained in earlier clinical trials. This aligns with the Empirical Bayes approach. The central assumption here is exchangeability: the treatment effect in the current trial is presumed to arise from the same distribution as the effects observed in earlier, comparable studies. In clinical research, the development of such an objective prior is supported by the existence of extensive meta-analytic repositories, most notably the Cochrane Database of Systematic Reviews (CDSR). The Empirical Bayes method is likely most appropriate for evaluating `follow-up' treatments or therapies within established drug classes, rather than for genuinely novel or `first-in-class' interventions.

More generally, pharmaceutical firms are legally obligated to disclose all pertinent information about a drug to the FDA before initiating a Phase 3 trial. Once this prior, together with the information used to construct it, becomes common knowledge, Alice and Bob can engage in a process of `communication' that will ultimately lead them to share a common prior, as implied by the classic Aumann agreement theorem (aumann1976agreeing). Indeed, FDA guidance (us2026guidance) explicitly notes that companies must account for the time required for this prior alignment when planning and developing the trial.

Risk neutrality and linear utility

For the analysis of optimal stopping times, we employed a second fundamental assumption that the Bernoulli utility functions of Alice and Bob are linear in $|\mu_1 - \mu_0|$, conditional on $\delta$.

For Alice, this implies risk neutrality with respect to the magnitude of the treatment effect. This is a stance consistent with that of a utilitarian social planner, with the important caveat that we do not restrict Alice's risk preferences regarding the binary decision itself; she may still exhibit significant asymmetry in her weighting of Type I versus Type II errors.

For Bob, linear utility implies that the private returns from approval scale proportionally with the magnitude of the treatment effect. In practice, one might expect profits to be non-linear and convex in $\mu_1 - \mu_0$; indeed, a breakthrough treatment effect might yield exceptionally high returns due to market dominance or reduced competition. However, such non-linearities would exacerbate, rather than mitigate, the incentive misalignments identified in Section (ref). For example, if Bob derives increasing marginal utility from the treatment effect, his `pro-approval' bias would strengthen. This would lead to even more pronounced asymmetry between approval and rejection thresholds and an even faster convergence of the approval threshold toward zero.

Numerical Calibration for Clinical Trials

Determining the optimal stopping time requires information on the marginal cost of experimentation $c$, the private fixed benefit from approval $B$, the treatment-effect dependent benefit $\gamma$, the prior variance $\varrho_0$, and the welfare lower bound $V_0$. In this section, we calibrate these parameters using evidence from the clinical trial literature and then use them to derive optimal stopping rules. This allows us to measure the percentage reduction in the expected sample size needed to reach the same expected welfare as that delivered by a standard Randomized Controlled Trial (RCT).

Up to this point, our analysis has been framed in continuous time, whereas in practical settings data are collected at discrete intervals. As outlined in Section (ref), the continuous-time framework serves as an approximation---under so-called `small-cost asymptotics' (adusumilli2025)---to a discrete sampling scheme in which time $t$ corresponds to the number of observations divided by $n$, with $n \to \infty$. In this asymptotic setup, the continuous-time parameters $c, B, \gamma, \varrho_0$ emerge as renormalized versions of the underlying structural parameters:

equation[equation omitted — 185 chars of source]

where $C$ denotes the true per-observation cost, $B_n$ the private fixed benefit from regulatory approval, $\gamma_n$ the scaling factor on treatment-effect–adjusted benefits, and $\varrho_{0,n}$ the prior variance.

The scaling can be rationalized as follows: following adusumilli2025, we interpret $n^{3/2}$ as the relevant population size for the drug. As argued in adusumilli2025, welfare for such a population size is maximized when the design is tuned to detect treatment effects of order $1/\sqrt{n}$. The true sampling cost $C$ is evidently invariant in $n$. By contrast, treatment-adjusted benefits must be multiplied by the population size $n^{3/2}$ to obtain aggregate profits, implying $\gamma_n = O(n^{3/2})$. Simultaneously, to keep the private fixed benefit and the treatment-adjusted benefits on the same scale, we impose $B_n/\gamma_n = O(1/\sqrt{n})$. This is also intuitive: $B$ represents the pure profit from approval even if the drug has no therapeutic effect, and in realistic scenarios, this should constitute only a small share of the overall profits from approval.

For ease of exposition, we assume throughout this section that $\sigma^2 := (\sigma_1^2 + \sigma_0^2) = 1$. Under this convention, $\varrho_{0}$ should be interpreted as the prior variance of the standardized treatment effect $\mu/\sigma$. While in practice $\sigma$ would be estimated from the data, all of our later results remain valid for arbitrary values of $\sigma$ if the $m_t$ terms appearing in this section are understood as $m_t/\sigma$.

Calibration of $n$, $c$ and $B$

chen2025investigating use data from all published and unpublished Phase 3 clinical trials registered on ClinicalTrials.gov over the past 20 years to estimate a median sample size of $300$ participants for Phase 3 studies. We therefore take this value as our scaling factor $n$.

moore report that the median cost per participant is approximately $C = \$41{,}000$.

tetenov2016 estimates that the value of an approved drug to pharmaceutical firms is approximately \$802 million. This figure, however, incorporates both the pure profit from regulatory approval and the benefits adjusted for treatment effects. As a result, the precise value of $B_n$ is not directly observable. Using the scaling argument outlined above, we regard it as reasonable to divide this total by $\sqrt{n}$, yielding an estimate of $B_n \approx \$46.3$ million. Under this approach, the pure benefit of approval amounts to roughly $5.5\%$ of total profits, a share we consider to be of a realistic order of magnitude.\footnote{This figure also aligns with the estimated value of approximately \$50 million for drugs in the preclinical phase, as reported by aryal2022valuing.}

Combining these values yields a cost-to-benefit ratio of $c/B = 0.265$. The appendix outlines two alternative calibration approaches: one assuming $B_n = 0$, and another assuming $B_n =$ \$802 million, meaning that the entire profit is represented as a purely approval-based benefit (although we consider this second calibration to be implausible).

Estimating the prior distribution

The Cochrane Database of Systematic Reviews (CDSR) is the principal repository for systematic reviews and meta-analyses in the clinical trials literature. vanzwet2021 use the Cochrane Database to estimate a prior distribution for scaled treatment effects ($\mu/\sigma$) in clinical trials via Empirical Bayes methods.\footnote{vanzwet2021 fit a four-component mixture of normal distributions, although one component receives only a negligible weight. We estimate $\varrho_0$ by taking $\varrho_0/n$ to be the variance of this distribution.} On the basis of their estimates, we set $\varrho_0 = 3.12^2 = 9.7344$.

vanzwet2021 contend that choosing a prior mean of $0$ for $\mu_1 - \mu_0$ is reasonable, as it treats the active treatment and control symmetrically. In fact, their empirical estimates also yield a prior mean that is close to zero.

Calibration of the welfare lower bound

A natural baseline choice for $V_0$ is the expected social welfare, $V_0^*$, that regulators could expect to obtain under a standard RCT. Recall that we normalize $t$ so that $t = 1$ matches the median sample size of typical clinical trials. With a $\mathcal{N}(0,\varrho_0)$ prior over $\mu_1 - \mu_0$, the posterior mean at $t = 1$ is distributed as $m_1 \sim \mathcal{N}(0, \varrho_0/(1+\varrho_0^{-1}))$.

Observe that $V_0^* = E[S_\alpha(m_1)]$. Now, $E[S_\alpha(m_1)] = E[\max\{m_1, 0\}]$ does not depend on $\alpha$ since $E[m_1] = 0$. Substituting the estimated value of $\varrho_0$ thus yields \[ V_0^* = E[S_\alpha(m_1)] = \sqrt{\frac{\varrho_0}{2\pi(1+\varrho_0^{-1})}} \approx 1.1853. \] In our numerical work below, we also vary $V_0$ by trying out different multiplicative factors of $V_0^*$.

Results

table[table omitted — 606 chars of source]

Table (ref) summarizes the values of the key parameters used in our analysis. A simple binary search pinpoints the value of $\lambda$ under which the optimal experimentation strategy achieves an ex-ante welfare of $V_0^*$. The resulting stopping rule is plotted in Figure (ref). The stopping boundaries clearly satisfy the properties highlighted in Theorem (ref). For comparison, we also plot the classical approval rule (accept if the sample mean, scaled by $\sigma$, is above $1.96/\sqrt{t}$; see the discussion in Section (ref)) as $b_{RCT}(t)$. As expected, $b^+(t)$ truncates to 0. In our calibration, the truncation occurs at around $t \approx 1.75$ (530 units); for the same sample size, $b_{RCT}(t)$ requires an effect size for approval at which the adaptive experiment would have ended much earlier.

figure[figure omitted — 697 chars of source]

In Figures (ref) and (ref), we plot the distributions of $\tau$ and $m_\tau$, obtained by running 1 million simulations of the hypothetical adaptive experiment implied by the optimal experimentation strategy. The results show that the proposed strategy can reach the social welfare level $V_0^*$---the same level delivered by a conventional RCT---while using, on average, about $48\%$ fewer observations (i.e., $\mathbb{E}[\tau^*] \approx 0.52$). Under our parameter estimates $C_n = 41{,}000\$$ and $n = 300$, this implies expected cost savings of nearly 6 million dollars. In addition, approximately $86\%$ of the simulated paths stop before $t = 1$; recall that this corresponds to the median sample size in clinical trials.

figure[figure omitted — 643 chars of source]

The preceding analysis presumes that Bob appropriates the entire welfare surplus generated by more efficient experimentation. Alternatively, Alice might instead choose to increase social welfare beyond $V_0^*$. Figure (ref) plots $V_0$ as a function of the multiplier $\lambda$, Figure (ref) shows how variations in $V_0$ affect the optimal boundaries, and Figures (ref) and (ref) present the expected and median stopping times over the same range of welfare values. The last two figures reveal that the mapping from $V_0$ to the expected sample size $\mathbb{E}[\tau]$ is highly non-linear. As discussed above, Alice can replicate the welfare level of a traditional RCT while using $48\%$ fewer observations. However, if she instead raises $V_0$ up to the point where Bob's average sample size matches that of a classical RCT (i.e., $\mathbb{E}[\tau^*] = 1$), this only yields a 3.2% welfare gain relative to $V_0$. This welfare level is denoted by $V^*_1$ in Figure (ref). These findings appear indicate that, under our calibration, pharmaceutical firms bear most of the inefficiency associated with relying on standard RCTs in clinical trial design: switching to more efficient approaches would significantly reduce firms' experimentation costs, while delivering only relatively small gains in overall social welfare.

Figure (ref) illustrates how the stopping boundaries vary with the benefit when $V_0$ is held constant. Reducing $B$ by half or increasing it two-fold has surprisingly little effect on the magnitudes of the stopping boundaries.

We describe additional comparative statics exercises in Appendix (ref).

figure[figure omitted — 1,070 chars of source]
figure[figure omitted — 199 chars of source]

Local Asymptotics and Parametric Models

Up to this point, we have studied the optimal experimental design problem under incremental learning (in continuous time). In this section, we demonstrate that the continuous-time formulation can be viewed as an approximation to the corresponding experimental design problem in discrete time, under so-called `small-cost asymptotics'.

The central idea, following adusumilli2025, is to consider a setting in which Bob and Alice care about detecting treatment effect differences on the order of $1/\sqrt{n}$ because the population size is of the order $O(n^{3/2})$. This motivates scaling time $t$ in units of $n$ observations, and employing the scaling for $B_n, \gamma_n$ described in ((ref)). For simplicity, in what follows, we rescale units so that $B = 1$.

Experimental design in discrete time

As before, Bob runs an experiment aimed at persuading Alice to take a particular action. The new feature is that the experiment now unfolds in discrete time.

Following lattimore, we adopt the “stack-of-rewards” representation to formalize experimental design. Concretely, we imagine that there exists an infinite stack of observations \(\bm{y}_{a} := \{Y_{i}^{(a)}\}_{i=1}^{\infty}\) for each treatment $a$, generated at the outset as i.i.d draws from a parametric family \(\{P_{\theta^{(a)}}^{(a)}\}\), where \(\theta^{(a)} \in \mathbb{R}^d\) is unknown. Initially, neither Alice nor Bob observes any of these realizations. Each time a treatment is selected, we can imagine that Bob sees the current top element of the corresponding treatment stack; this element is then taken out of consideration.

Bob's sampling strategy is described by a policy \(\{\pi_{n,j}\}_{j} \equiv \{\pi_{n,\lfloor nt\rfloor}\}_{t}\), which specifies, at each period \(j\), the probability of assigning that observation to treatment 1 as a function of past data. The treatment assignment is then given by \(A_{j} \sim \text{Bernoulli}(\pi_{n,j})\). Define \[ q_{n,a}(t) := \frac{1}{n}\sum_{j=1}^{\lfloor nt\rfloor} \mathbb{I}\{A_{j} = a\} \] as the number of observations that have been allocated to treatment \(a\) by time \(t\), normalized by $n$. We refer to \(\{q_{n,a}(\cdot)\}_{a}\) as the empirical allocation process. The policy rule \(\{\pi_{n,j}\}_{j}\) can be viewed as a function that takes the stacks \(\ensuremath{(\bm{y}_{1}, \bm{y}_{0})}\) and an exogenous random variable \(U \sim \text{Uniform}[0,1]\) as inputs and produces the realized trajectory of the empirical allocation process \(\{q_{n,a}(\cdot)\}_{a}\). The sampling strategy can thus be characterized by \(\{q_{n,a}(\cdot)\}_{a}\) instead of \(\{\pi_{n,j}\}_{j}\).

Let $\mathcal{F}_t^{\bm{q}_n}$ denote the $\sigma$-algebra generated by \[ \xi_t \coloneq \left\{U, \{A_j\}_{j=1}^{\lfloor nt\rfloor},\{Y_i^{(1)}\}_{i=1}^{\lfloor nq_{n,1}(t)\rfloor}, \{Y_i^{(0)}\}_{i=1}^{\lfloor nq_{n,0}(t)\rfloor}\right\}, \] that is, the sequence of actions and realized rewards up to time $\lfloor nt\rfloor$. In addition to the sampling strategy, Bob chooses a stopping time $\tau_n$ that is $\mathcal{F}_{t}^{\bm{q}_n}$-adapted.

Overall, Bob's decision rule $\bm{d}_n$ consists of the combination of the empirical allocation process $\bm{q}_n(\cdot)$ and the stopping time $\tau_n$. After the experiment concludes, Alice makes a binary decision---either approval $(\delta_n = 1)$ or rejection $(\delta_n = 0)$---with the requirement that $\delta_n$ be $\mathcal{F}_{\tau_n}^{\bm{q}_n}$-measurable.

Local asymptotics

Following hirano2025asymptotic, we assess decisions under local perturbations of the form $\{\theta_{0}^{(a)} + h_a/\sqrt{n} : h_a \in \mathbb{R}^{d}\}$, where $\theta_{0}^{(a)}$ is a reference parameter. Let the mean outcomes be defined by $\mu_a(\theta) \coloneq \mathbb{E}_{P_\theta^{(a)}}[Y_i^{(a)}]$. The reference parameters are selected so that $\mu_1\bigl(\theta_0^{(1)}\bigr) - \mu_0\bigl(\theta_0^{(0)}\bigr) = 0$. This setup is motivated by the consideration that Alice and Bob generally must be able to discriminate between treatment effects of order $1/\sqrt{n}$ in order to justify running experiments with sample size on the order of $n$.

To simplify notation, we normalize the means so that $\mu_1\bigl(\theta_0^{(1)}\bigr) = \mu_0\bigl(\theta_0^{(0)}\bigr) = 0$. Consequently, under local perturbations, we have

equation[equation omitted — 143 chars of source]

where $\dot\mu_a \coloneq \nabla_{\theta}\mu_a(\theta_0^{(a)})$.

Let $\nu$ denote a dominating measure for $\{P_{\theta}^{(a)}:\theta\in\mathbb{R}^{d},a\in\{0,1\}\}$, and set $p_{\theta}^{(a)}\coloneq dP_{\theta}^{(a)}/d\nu$ . We require $\{P_{\theta}^{(a)}\}_{\theta}$ to be quadratic mean differentiable (qmd):

asmThe class $\{P_\theta^{(a)}:\theta \in \mathbb{R}^d\}$ is differentiable in quadratic mean around $\theta^{(a)}_0$ for each $a \in \{0,1\}$, i.e., there exists a score function $\psi_a(\cdot)$ such that for each $h_a \in \mathbb{R}^d$, \[\int\left[\sqrt{p^{(a)}_{\theta^{(a)}_0 + h_a}} - \sqrt{p^{(a)}_{\theta_0^{(a)}}} - \frac{1}{2}h_a^{\intercal}\psi_a\sqrt{p^{(a)}_{\theta_0^{(a)}}}\right]^2d\nu = o\left(|h_a|^2\right).\]Furthermore, the information matrix $I_a \coloneq \mathbb{E}_0[\psi_a\psi_a^\intercal]$ is invertible for $a \in \{0,1\}$.

This assumption is fairly weak and holds for nearly all standard distributions, such as the Normal, Cauchy, Exponential, and Poisson distributions.

In what follows, define $P^{(a)}_h \coloneq P^{(a)}_{\theta_0^{(a)}+h/\sqrt{n}}$ and let $\mathbb{E}_{h_a}[\cdot]$ denote the corresponding expectation. Furthermore, let $P^{(a)}_{n,h}$ be the joint distribution of $\bm{y}_{n}^{(a)} = \left\{Y_1^{(a)},Y_2^{(a)}, \ldots\right\}$, where $Y_i^{(a)} \sim \text{i.i.d. } P^{(a)}_h$. We also set $\textbf{\textit{h}} \coloneq (h_1,h_0)$ and define $$ P_{n,\textbf{\textit{h}}} := P^{(1)}_{n,h_1} \times P^{(0)}_{n,h_0}, $$ with $\mathbb{E}_{n,\textbf{\textit{h}}}[\cdot]$ as its associated expectation.

For each treatment \(a\), the score process is defined as \[ X_{n,a}(t):=\frac{I_{a}^{-1/2}}{\sqrt{n}}\sum_{i=1}^{\left\lfloor nq_{n,a}(t)\right\rfloor }\psi_{a}(Y_{i}^{(a)}). \] adusumilli-continuous-arts demonstrates that the sample paths of \(\{X_{n,a}(\cdot),q_{n,a}(\cdot)\}\) up to time \(t\) form asymptotically sufficient statistics for the adaptive experiment.

Payoffs, priors and welfare

Let $\mu_n(\bm{h}) = \mu_{n,1}(h_1) - \mu_{n,0}(h_0)$. We assume that Alice and Bob's Bernoulli utility functions (exclusive of sampling costs) take the form $$ \sqrt{n} u\left(\mu_n(\bm{h}),\delta_n; \alpha \right),\quad B_n \delta_n + \gamma_n u\left(\mu_n(\bm{h}),\delta_n; \alpha^\prime \right), $$ where the function $u(\cdot)$ is specified in Section (ref). To understand the scaling of Alice's welfare, recall from ((ref)) that $\mu_{n,0}(h_0) = O(1/\sqrt{n})$, so we effectively scale the treatment effects by $\sqrt{n}$ to prevent them from vanishing asymptotically.

Observe that $\sqrt{n} u\left(\mu_n(\bm{h}),\delta_n; \alpha \right) = u\left(\sqrt{n}\mu_n(\bm{h}),\delta_n; \alpha \right)$, so for a given $\bm{h}$, Alice's expected utility under the strategy pair $(\bm{d}_n, \delta_n)$ is given by

align[align omitted — 136 chars of source]

For Bob's expected utility under a given $\bm{h}$, we scale by the factor $1/B_n$ and define

equation[equation omitted — 463 chars of source]

where the second line uses the asymptotic scaling regime from ((ref)).

Prior choice

We assume that Alice and Bob share a common Gaussian prior $\Gamma_0(\bm{h})$ on the local parameter $\bm{h}$. This prior induces a corresponding prior $p_0$ on the scaled treatment effects $(\dot{\mu}_1^\intercal h_1, \dot{\mu}_0^\intercal h_0)$. We further assume that $\Gamma_0(\bm{h})$ factorizes into the prior $p_0$ on $(\dot{\mu}_1^\intercal h_1, \dot{\mu}_0^\intercal h_0)$ and an independent prior $\tilde{\Gamma}_0$ on the remaining components of $\bm{h}$. The multiplicative separability of $p_0$ and $\tilde{\Gamma}_0$ can be justified by an invariance requirement: inference about $(\dot{\mu}_1^\intercal h_1, \dot{\mu}_0^\intercal h_0)$ should not depend on the values of `nuisance parameters' associated with the other components of $\bm{h}$.

For our asymptotic framework, we additionally assume that $\Gamma_0$ remains fixed as $n$ increases. As highlighted in adusumilli2025bandits, such local priors provide a more accurate description of the asymptotic behavior of decisions because their influence does not vanish with growing sample size.

Finite sample welfare

Let $\delta_n^*$ denote Alice's optimal strategy under the given prior, given Bob's choice of $\bm{d}_n$. Alice requires that social welfare exceed a predetermined level $V_0$, up to a slackness term $\epsilon_n \to 0$. Formally, her requirement is $$ \int W_n^A((\bm{d}_n, \delta_n^*), \bm{h}) \, d\Gamma_0(\bm{h}) \ge V_0 - \epsilon_n. $$ Introducing an arbitrarily small relaxation of the welfare constraint makes it less stringent, allowing it to be met in the limit rather than exactly. Given Alice's choice, Bob's experimental design problem is to select $\bm{d}_n$ to maximize his own expected welfare, i.e., Bob solves

equation[equation omitted — 278 chars of source]

Limit approximations and upper bounds on welfare

Consider a limit experiment in which the underlying informational environment consists of signal processes of the form

equation[equation omitted — 105 chars of source]

where $\bar{W}_1(\cdot), \bar{W}_0(\cdot)$ are independent $d$-dimensional Brownian motions. As before, in this limit experiment, Bob chooses an experimental strategy $\bm{d}$---consisting of an allocation strategy $\{q_a(\cdot)\}_a$ and a stopping time $\tau$---while Alice makes a binary decision $\delta$.

Suppose further that Alice and Bob's expected utilities in the limit experiment, under a strategy pair $(\bm{d}, \delta)$, take the form

equation[equation omitted — 342 chars of source]

where $\mu(\bm{h}) := \dot{\mu}_1^\intercal h_1 - \dot{\mu}_0^\intercal h_0$. Based on the above, we can write Bob's experimental design problem in this limit-experiment as:

equation[equation omitted — 266 chars of source]

Observe that $W^A((\bm{d}, \delta), \bm{h})$ and $W^B((\bm{d}, \delta), \bm{h})$ depend on $\bm{h}$ solely through the terms $\dot{\mu}_1^\intercal h_1$ and $\dot{\mu}_0^\intercal h_0$. Together with the assumption that we restrict attention to multiplicatively separable priors $\Gamma_0$ (as specified in Section (ref)), this implies that the signal processes $\dot{\mu}_1^\intercal I_1^{-1/2}Z_1(\cdot)$ and $\dot{\mu}_0^\intercal I_0^{-1/2} Z_0(\cdot)$ constitute sufficient statistics for the limit experiment. Consequently, we can directly relate this experiment to the one in Section (ref) by equating $\mu_a$ with $\dot{\mu}_a^\intercal h_a$, $z_a(\cdot)$ with $\dot{\mu}_a^\intercal I_a^{-1/2} Z_a(\cdot)$, and $\sigma_a^2$ with $\dot{\mu}_a^\intercal I_a^{-1} \dot{\mu}_a$. See Lemma (ref) for a formal proof of the equivalence between these two experiments.

The asymptotic representation theorem of adusumilli-continuous-arts enables us to match the expected welfare of Alice and Bob under any sequence of strategy pairs, $\{(\bm{d}_n, \delta_n)\}_n$, with that from a strategy pair, $(\bm{d},\delta)$, in the limit experiment described above. Here, we do not necessarily require $\delta_n, \delta$ to be Bayes-optimal relative to $\bm{d}_n, \bm{d}$; they just represent permissible strategies for Alice. The result on the matching of expected welfare relies on the following assumptions:

asmThere exists $T < \infty$ such that $\tau_n \le T$ for all $n$.
asmThere exists $\dot\mu_a \in \mathbb{R}^d$ and $\varepsilon_n \to 0$ independent of $(a, \bm{h})$ such that $\sqrt{n}\mu_{n,a}(\bm{h}) = \dot\mu_a^\intercal h_a + \varepsilon_n|h_a|^2$.

Assumption (ref) requires that the stopping times are bounded. While our analysis of the incremental learning regime allows for unbounded stopping times, it poses technical challenges for asymptotic approximations; therefore, we restrict our attention to stopping times bounded by some arbitrarily large, but finite, $T$. Assumption (ref) is a mild requirement on the smoothness properties of $\mu_{n,a}(\textbf{\textit{h}})$.

thmSuppose Assumptions (ref), (ref), and (ref) hold. Then, for any sequence of strategy pairs $(\bm{d}_n, \delta_n)$ and corresponding welfare functions $W^A_n(\cdot, \bm{h}), W_n^B(\cdot, \bm{h})$, there exists a subsequence $(\bm{d}_{n_k}, \delta_{n_k})$ and a strategy pair $(\bm{d},\delta)$ in the limit experiment, with welfare functions $W^A(\cdot, \textbf{\textit{h}}), W^B(\cdot, \bm{h})$, such that $W^A_{n_k}(\cdot, \textbf{\textit{h}}) \to W^A(\cdot, \textbf{\textit{h}})$ and $W^B_{n_k}(\cdot, \textbf{\textit{h}}) \to W^B(\cdot, \textbf{\textit{h}})$ for each $\textbf{\textit{h}}$.

The quantities $\bar{W}_n^*$ and $\bar{W}^*$, defined in ((ref)) and ((ref)), denote Bob's optimal welfare in the finite-$n$ and limit settings, respectively. We next establish $\limsup_{n \to \infty} \bar{W}_n^* \le \bar{W}^*$, which shows that Bob's welfare in the limit experiment serves as an asymptotic upper bound for his welfare in the finite-$n$ case.

thmSuppose Assumptions (ref), (ref), and (ref) hold, and the prior $\Gamma_0$ is Gaussian.\footnote{An inspection of the proof of this theorem reveals that the extra conditions imposed on the Gaussian prior $\Gamma_0$ in Section (ref) are in fact not needed here.} Then, $\lim_{T \to \infty} \limsup_{n \to \infty} \bar{W}_n^* \le \bar{W}^*$.

Attaining the upper bound

We now demonstrate that the upper bound stated in Theorem (ref) can be attained using finite-sample counterparts of the optimal strategies derived in Section (ref).

In what follows, we assume that the normal prior $\Gamma_0(\bm{h})$ induces a distribution $p_0$ over $(\dot{\mu}_1^\intercal h_1, \dot{\mu}_0^\intercal h_0)$ for which Assumption (ref)(iv) is satisfied. Recall that under this condition, the optimal sampling strategy in the limit experiment reduces to the Neyman allocation in every period.

Let $\Sigma_{aa}$ denote the prior variance of $\dot{\mu}_a^\intercal h_a$ and define $$

aligned\mu_{n,a}(t) &:= \frac{\dot{\mu}_a^\intercal I_a^{-1/2} X_{n,a}(t) + \Sigma_{aa}^{-2}\mu_0^{(a)}}{\sigma_a^{-2} q_{n,a}(t) + \Sigma_{aa}^{-2}}, \\ m_{n,t} &:= \mu_{n,1}(t) - \mu_{n,0}(t).

$$ In this notation, $m_{n,t}$ is the sample analogue of $m_t$, the posterior average treatment effect in the limit experiment. Furthermore, let $b^+(t; \lambda)$ and $b^-(t; \lambda)$ be the optimal stopping boundaries defined in Theorem \ref{thm:optimal-stopping-rule}, where we now make their dependence on $\lambda$ explicit.

Take $\lambda^*$ to be the Lagrange multiplier corresponding to $V_0$ in the limit experiment. For some $T < \infty$ and $\xi > 0$, we construct finite sample analogs of $\bm{d}^*$ as $\bm{d}^*_{n,T,\xi} = (\pi^*_n, \tau^*_{n,T,\xi})$, where

equation[equation omitted — 329 chars of source]

and it may be recalled from Section (ref) that $\sigma_a^2 := \dot{\mu}_a^\intercal I_a^{-1} \dot{\mu}_a$. The policy rule $\pi^*_{n,a}(t)$ realizes the Neyman allocation $q_{n,a}^*(t) = \sigma_a t /(\sigma_1 + \sigma_0)$, up to an approximation error of order $O(1/(nt))$. The stopping time $\tau^*_{n,T,\xi}$ is a finite-sample counterpart of $\tau^*$ in which $m_{n,t}$ substitutes for $m_t$, but with two further modifications: we cap the stopping time at an arbitrarily large horizon $T$, and we enlarge the upper boundary by an arbitrarily small $\xi > 0$. The latter adjustment ensures that Alice never enters her indifference region, which corresponds to $b^+(t) = 0$. From a formal standpoint, this is required to eliminate discontinuity issues when proving convergence of welfare. Practically, it also means that the exact tie-breaking rule Alice uses within her indifference region is no longer relevant.

The theorem below shows that, for sufficiently large $T$ and sufficiently small $\xi$, the policy $\bm{d}_{n,T,\xi}^*$ asymptotically satisfies Alice's welfare constraint while also delivering Bob's asymptotically optimal welfare level $\bar{W}^*$. Therefore, this experimentation strategy is asymptotically optimal.

thmSuppose Assumptions (ref), (ref), and (ref) hold. Further assume that the prior $\Gamma_0$ is Gaussian and can be factored into a prior $p_0$ over $(\dot{\mu}_1^\intercal h_1, \dot{\mu}_0^\intercal h_0)$, which fulfills Assumption (ref)(iv), and an independent prior $\tilde{\Gamma}_0$ over the remaining components of $\bm{h}$. Then $$ \lim_{T \to \infty} \lim_{\xi \to 0}\lim_{n \to \infty} \int W_n^A\left((\bm{d}_{n,T,\xi}^*, \delta_{n,T,\xi}^*), \bm{h} \right) d\Gamma_0(\bm{h}) \ge V_0, $$ where $\delta_{n,T,\xi}^*$ denotes Alice's Bayes-optimal strategy given $\bm{d}_{n,T,\xi}^*$, and $$ \lim_{T \to \infty} \lim_{\xi \to 0}\lim_{n\to\infty} \int W^B_n\left((\bm{d}^*_{n,T,\xi}, \delta^*_{n,T,\xi}), \bm{h} \right)d\Gamma_0(\bm{h}) = \bar W^*. $$

In applications, we recommend choosing $\xi$ to be a small multiple of $\sigma$, e.g., $0.05\sigma$. Regarding $T$, we advise taking $T = \infty$ in practice, as the restriction to a finite $T$ in Theorems (ref) and (ref) is primarily to simplify the theory.

Unknown variances

Up to this point, our analysis has taken $\sigma_1, \sigma_0$ as known. These objects are the information matrices evaluated at the reference parameters $\theta_{0}^{(1)}, \theta_{0}^{(0)}$. Conceptually, within the local asymptotic framework, these reference parameters are treated as if they are known beforehand. As emphasized in adusumilli2025, although in applications one would ideally design procedures that either adapt to or are invariant with respect to these quantities, the local asymptotic framework itself cannot capture the effect of estimating them.

In principle, one could use a ‘forced exploration' phase (see, for example, lattimore): for the first $\bar{n} = n^{c}$ observations, with $c \in (0,1)$, we set $\pi_{n,1}^{*}(t) = 1/2$. This is equivalent to applying the equal allocation strategy until time $\bar{t} = n^{c-1}$. The data collected in this preliminary stage are then used to obtain consistent estimators $\hat{\sigma}_{1}^{2}, \hat{\sigma}_{0}^{2}$ of $\sigma_{1}^{2}, \sigma_{0}^{2}$. From time $\bar{t}$ onward, we apply the asymptotically optimal experimentation rule $\bm{d}_{n,T,\xi}^*$, replacing $\sigma_{1}, \sigma_{0}$ with their estimates $\hat{\sigma}_{1}, \hat{\sigma}_{0}$. In practical applications, since clinical trials are typically conducted in groups, we suggest using $\pi_{n,1}^{*}(t) = 1/2$ for the first group of participants.

In the special case of Bernoulli outcomes, the local asymptotic structure forces $\sigma_1 = \sigma_0$, which in turn implies that $\pi_{n,1}^{*}(t) = 1/2$ throughout.

Simulation with Bernoulli outcomes

To study the finite-sample behavior of our proposed procedures, we simulate the stopping times under Bernoulli outcomes when using the asymptotically optimal strategy from ((ref)).

Let $Y^{(a)} \sim \text{Bernoulli}(\theta^{(a)})$ represent the outcome under treatment $a$. As noted earlier, conducting a local asymptotic analysis with Bernoulli outcomes requires specifying reference parameters $\theta^{(1)}_0$, $\theta^{(0)}_0$ such that $\theta^{(1)}_0 = \theta^{(0)}_0 := \theta_0$. We then posit that each $\theta^{(a)}$ is drawn independently from a Gaussian prior, $\theta^{(a)} \sim \mathcal{N}(\theta_0, \sigma^2 \nu^2/n)$, where $\sigma^2 = 4\theta_0(1-\theta_0)$.\footnote{Recall that $\sigma^2 := (\sigma_1 + \sigma_0)^2$, and in the Bernoulli case $\sigma_1^2 = \sigma_0^2 = \theta_0(1-\theta_0)$.} In our simulations, we set $\theta_0 = 0.5$ and $\nu^2 = \varrho_0/2$, where $\varrho_0 \approx 9.734$ is the value estimated in Section (ref). With this choice, we have $\sqrt{n}(\theta^{(1)} - \theta^{(0)})/\sigma \sim \mathcal{N}(0,\varrho_0)$, in agreement with the prior specification obtained in Section (ref).

For Bernoulli outcomes, the score function takes the form $\psi_a(Y) = (\theta_0(1-\theta_0))^{-1}(Y^{(a)} - \theta_0)$, so the (normalized) score process can be written as $$ X_{n,a}(t):=\frac{1}{\sqrt{n\theta_0(1-\theta_0)}}\sum_{i=1}^{\left\lfloor nq_{n,a}(t)\right\rfloor } (Y_i^{(a)} - \theta_0). $$ We then use this specification of $X_{n,a}(\cdot)$ in ((ref)). Figure (ref) illustrates how Alice and Bob's finite-sample welfare under the asymptotically optimal strategy converges as $n$ grows.\footnote{Due to numerical approximations used to compute the asymptotic welfare, Bob's finite-sample welfare appears slightly higher than the approximate asymptotic benchmark.}

figure[figure omitted — 914 chars of source]

Conclusion

In this article, we propose that regulators directly target social welfare when regulating experimental designs. We characterize the optimal design in a continuous-time setting with two treatments and Gaussian priors over the mean effects of these treatments. Finally, we show that the optimal design in continuous time is asymptotically optimal under parametric outcome distributions and construct a finite sample analog of the optimal design that can be used in practical settings.