Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
86,839 characters · 14 sections · 12 citation commands
Learning about Treatment Effects with Prior Studies: A Bayesian Model Averaging Approach
\thispagestyle{empty}
\setcounter{page}{1}
\doublespacing
Governments, firms, and researchers often conduct experiments to estimate the causal effects of a given policy or intervention. In practice, they rarely run a single, isolated experiment; instead, policies are piloted, refined, and expanded in sequential waves. In these settings, new experiments begin with substantial prior information---results from earlier pilots, studies conducted in related populations, or expert assessments about likely effect sizes. Incorporating this information can dramatically reduce the cost of experimentation and accelerate inference. However, prior evidence typically varies in relevance and external validity, and na\"{i}vely pooling it with new data can lead to biased estimates and misguided policy decisions. A central question is therefore how to systematically incorporate prior sources for learning the expected effect of treatment, while allowing for the possibility that some may be biased, misspecified, or only partially informative for the current environment.
This paper provides a principled solution by merging standard estimation of treatment effects with treating each prior experiment (or expert assessment) as a distinct “model” in the Bayesian model averaging (BMA) tradition.\footnote{For excellent reviewss of BMA, see KassRaftery1995,Hoeting1999,Wasserman2000,Steel2020 among others.} The experimenter uses new data to update the posterior probability that each source is externally valid for the current environment, so the resulting estimator --- a weighted average of source-specific posterior means for each treatment-covariate pair, where the weights are BMA posteriors --- automatically assigns more weight to sources supported by the data and downweights those that are inconsistent or biased. This approach offers desirable properties: it yields faster learning when externally valid sources exist, and it remains robust when they do not. These properties arise under an asymptotic framework that fundamentally differs from the standard asymptotics used in Bayesian model averaging.
While our analysis is formally expressed using BMA and Bayes model posteriors, their role here differs fundamentally from that in the standard BMA literature. Classical theoretical results rely on an asymptotic regime in which the current experiment grows large while the information content of each prior source remains fixed Schwarz1978,Walker2004,Wasserman2000,Steel2020. Although this regime is useful in many model-selection settings, it can be a coarse or even misleading approximation in environments such as multi-site RCTs, sequential pilots, and phased rollouts, where prior studies may be comparable in size to—or larger than—the ongoing experiment. In these cases, standard BMA asymptotics offer limited guidance on how posterior weights evolve across sources and on the resulting speed of learning of the estimator.
To address this gap, the paper introduces a new asymptotic framework in which the precision of each prior source is allowed to grow with the sample size of the ongoing experiment. This device captures settings where all available evidence---past and present---contains substantial information, and it allows us to characterize how uncertainty from both the new data and the priors is resolved. This delivers a novel form of discrimination that reflects learning about external validity of sources rather than differences in likelihood fit. This discrimination yields oracle-type and robustness results that emerge as implications of modeling empirically relevant environments with large and heterogeneous sources of prior information, rather than as objectives built into the asymptotic design.
The oracle-type result implies that when at least one prior source is unbiased---its prior mean coincides with the true treatment effect---the posterior asymptotically identifies these sources and assigns weight only to them. As a result, our estimator converges strictly faster than the standard estimator based solely on the new data. The magnitude of this improvement depends on the relative size of the unbiased prior sources: when such sources are of comparable scale to the current experiment, the estimator achieves a convergence rate that is twice as fast, representing a substantial reduction in estimation error.
Equally important, the framework delivers a form of robustness. If all prior sources are biased but at least one diffuse (low-precision) source is included, the procedure automatically downweights the biased sources and the estimator converges to the truth at the standard rate. In this way, the procedure never performs worse than using only the new experiment, and performs strictly better whenever externally valid sources exist. This robustness property is, to our knowledge, absent from the classical BMA literature, which --- under global misspecification --- lacks a mechanism that guarantees performance no worse than using the new experiment alone. Here, robustness follows from a simple safeguard: the inclusion of a deliberately diffuse source, which can always be incorporated by the researcher.
The mechanism behind these results differs fundamentally from existing BMA theory. In classical applications---such as Bayesian variable selection---models differ in their likelihood or parameter dimension, and asymptotic selection is driven by differences in goodness-of-fit penalized by model complexity KassRaftery1995,RafteryMadiganHoeting1997,RafteryMadiganHoeting2002. In our setting, by contrast, all sources correspond to the same scalar parameter (the expected outcome for each treatment-covariate pair) and share the same likelihood; they differ only in their prior means and precisions. Under standard asymptotics, these differences vanish asymptotically and offer no discriminatory power. Under our asymptotic regime, however, posterior discrimination operates through a continuous external-validity index that depends jointly on a source's bias and effective precision. Biased sources are exponentially downweighted, while unbiased sources dominate with weights proportional to their asymptotic information content. These results demonstrate how, under our nonstandard asymptotic framework, Bayesian model averaging provides a coherent and transparent method for incorporating prior experimental evidence when external validity is uncertain.
Why focus on learning/convergence rates? From a theoretical perspective, convergence rates are the basic building blocks for downstream frequentist results, including asymptotic normality, coverage guarantees, and the behavior of plug-in decision rules. Establishing sharp convergence rates is therefore an important step to determine the statistical precision available at each sample size that underpins all subsequent inferential and decision-theoretic guarantees. In sequential experimentation, however, rates are also interesting in their own right. They determine how quickly the true treatment effects are learned, and therefore how soon an experimenter can credibly stop, scale up, or revise a policy. At a given sample size, faster rates translate directly into tighter policy recommendations, and they quantify the value of borrowing from prior evidence relative to collecting additional observations.
The paper also extends these results in two directions that are especially relevant in practice. First, we allow outcomes to be binary and develop a Bernoulli version of the model. In that case, the same logic continues to govern posterior source weights, but external validity is no longer summarized by a quadratic loss. Instead, it is measured by a Kullback-Leibler projection index that captures the likelihood cost of reconciling a source's prior belief with the target environment. This delivers the same qualitative conclusions as in the Gaussian benchmark: biased sources are exponentially downweighted, while unbiased or sufficiently diffuse sources determine the asymptotic behavior of the estimator.
Second, we allow different sources to bring different sets of control variables in a Gaussian linear regression framework. This changes the geometry of learning because posterior updating is no longer arm-by-arm: treatment effects and nuisance coefficients are learned jointly, and the relevant external-validity object becomes matrix-valued after partialling out controls. Even so, the same oracle and robustness logic survives. Together, these extensions show that our framework is not tied to the simplest Gaussian setup but applies more broadly to environments researchers actually face.
\paragraph{Related literature.} Perhaps closest in spirit to our asymptotic framework are the ideas behind Zellner's $g$-prior Zellner1986. The $g$-prior is used in a different problem than ours---variable selection in a Gaussian linear regression model---but it has the notable feature of a parameter (the “$g$” factor) that regulates the prior precision relative to sample size. The literature has explored many choices and hyperpriors for $g$, often motivated by model-selection consistency and predictive performance; see, for example, FernandezLeySteel2001 and LiangEtAl2008 for overviews and prominent proposals. For fixed values of $g$, the influence of the prior does not vanish, even asymptotically; in that sense, this feature resembles ours. However, the motivations and implications of these choices are different from ours: in our setting, scaling prior precision with the experiment is a device to formalize learning about external validity across heterogeneous prior sources.
One might ask whether our framework is equivalent to BMA with a Zellner $g$-prior, possibly under a different scaling of $g$. This is not the case. In classical variable-selection settings, Zellner priors are centered at zero for all models, so prior means coincide across specifications. As a result, all models are equally biased whenever the true parameter differs from zero, and Bayes factors discriminate models only through differences in model dimension. In contrast, our framework allows prior means to differ across sources and treats these means as potentially misspecified objects. Combined with an asymptotic regime in which prior precision grows at the same order as the experiment, this makes bias itself an object of learning: biased sources are exponentially downweighted, while unbiased sources dominate. This selection mechanism has no analogue under Zellner-type priors.
Regarding model selection results in BMA, there is a large body of results developed in the context of covariate selection in regression and moment‐condition models.\footnote{See LI2016132,Johnson01062012 and references therein, as well as Wasserman2000 for a review.} However, as noted above, the mechanism driving those selection results is fundamentally different from ours. In standard BMA, model‐selection consistency arises from differences in likelihood fit and model dimension under fixed‐prior asymptotics. In contrast, selection in our setting emerges from a nonstandard asymptotic regime in which prior precision grows with sample size, allowing the posterior to learn about external validity and to exponentially downweight biased sources while concentrating on unbiased ones. Moreover, to the best of our knowledge, existing literature does not deliver an analogue of our finding that such selection directly translates into faster concentration rates for the resulting estimator.
\paragraph{Roadmap.} Section (ref) describes the setup. Section (ref) presents the main results. Section (ref) presents simulation evidence. Section (ref) develops the binary-outcome and control-variable extensions. Section (ref) concludes. All proofs are related to Appendix (ref).
In this section, we describe the experiment and how prior information is incorporated into the estimation of our parameter of interest. The design is intentionally general, applying both to sequential experiments and to standard randomized controlled trials.
\paragraph{The Experiment.}
Consider an experiment in which individuals are assigned to a set of treatments, $\mathbb{D} := \{0,\dots,M\}$ based on observable characteristics, $x \in \mathbb{X}$, where both sets are finite. For each treatment–covariate pair $(d,x)$, let $Y(d,x) \in \mathbb{R}$ denote the potential outcome, and define the parameter of interest as the mean potential outcome
At each instance $t \in \mathbb{N}$, the observed outcome covariate profile, $x$ is $Y_t(x) := Y_t(D_t(x),x)$ where $D_{t}(x) \in \mathbb{D}$ is the assigned treatment, which is assigned according to a policy rule
which specifies a probability distribution over treatments as a function of the past history of outcomes and assignments.\footnote{The index $t$ should be interpreted as indexing experimental stages rather than calendar time. In a standard RCT, $t$ can be viewed as labeling independent experimental units or cohorts drawn under a fixed randomization scheme, so that $\delta_t$ does not vary with the realized history. In sequential or adaptive experiments, by contrast, $t$ indexes decision rounds, and the assignment rule $\delta_t$ is allowed to depend on past outcomes and treatment assignments. Our analysis accommodates both interpretations.} Thus, $\delta_t(d \mid x)$ gives the probability that an individual with covariates, $x$, receives treatment, $d$, at instance, $t$. When no confusion arises, we omit the conditioning on past history.
The policy rule encompasses standard randomized controlled trials (RCTs), in which the assignment rule $\delta_t(y^{t-1}, d^{t-1})(\cdot \mid x)$ is independent of past outcomes and treatments, as well as more sophisticated sequential experimentation designs in which treatment assignment adapts to accumulated information. We deliberately refrain from taking a stand on the desirability of any particular policy rule. Instead, our objective is to maintain sufficient generality so that the learning rates derived below apply uniformly across a wide class of commonly used assignment mechanisms, including both static and adaptive designs.
\paragraph{Prior sources of information.}
For each $(d,x) \in \mathbb{D} \times \mathbb{X}$, the experimenter has access to a collection of prior sources $\mathcal{S} := \{0,\dots,L\}$. Each source $s \in \mathcal{S}$ is represented by a Gaussian prior
where $\phi(\cdot;a,b)$ denotes the normal density with mean $a$ and variance $b$. The quantity $\zeta^{s}_{0}(d,x)$ represents the prior value of the expected outcome under treatment $d$ for individuals with covariates $x$, while $\nu^{s}(d,x)$ reflects the precision of source $s$, with larger values indicating greater confidence.
The prior sources may derive either from previous experiments or from expert judgments. When $s$ corresponds to a past experiment, the experimenter simply collects the estimated average outcome for each treatment–covariate pair, which becomes $\zeta^{s}_{0}(d,x)$, along with the number of units with covariates $x$ assigned to treatment $d$, which becomes $\nu^{s}(d,x)$. In contrast, when $s$ represents an expert opinion or recommendation, the experimenter elicits the expert’s assessment of the expected outcome -- serving as $\zeta^{s}_{0}(d,x)$ -- and the expert’s confidence in that assessment, which naturally maps into the precision parameter $\nu^{s}(d,x)$. In all cases, the pair $(\zeta^{s}_{0}, \nu^{s})_{s \in \mathcal{S}}$ is treated as non-random.
\paragraph{Posterior updating.}
Assume that the experimenter models the outcome distribution as belonging to the Gaussian family
After observing treatments and outcomes up to instance ($t$), the experimenter computes, for each source $s$, the posterior distribution for $\theta(d,x)$ using Bayes’ rule. Under conjugacy, this posterior is Gaussian with mean
and the posterior precision is $N_t(d,x) + \nu^{s}(d,x)$.
\paragraph{Model posterior probabilities.} Faced with $L+1$ sources for each $(d,x)$, the experimenter assigns posterior model weights
These quantities are Bayesian model posterior probabilities in the sense of Bayesian model averaging (BMA). In our context, they can be interpreted as the probabilities that each source is externally valid for the pair $(d,x)$. A formal and detailed discussion is provided in Section (ref).
\paragraph{Estimator of average effects.}
For each treatment–covariate pair $(d,x)$, the estimator for $\theta(d,x)$ is given by the BMA-weighted average
The Gaussian–Gaussian framework yields a simple and intuitive estimator: a weighted average of the source-specific posterior means, where the weights adaptively reflect each source’s fit to the observed data. Importantly, we do not assume that the true distribution of outcomes is Gaussian. The Gaussian likelihood is a working model used only for tractable inference. Because the parameter of interest is the mean outcome, and because sample averages are unbiased under very general conditions (e.g., finite second moments), the estimator remains consistent even if the Gaussian working model is misspecified.
In this section, we present our asymptotic framework and discuss how the posterior weights, $\alpha^{s}_{t}$, can be interpreted as a measure of the external validity of source $s$ for the current experiment. We then derive the learning rates of our estimator and show that its rate of convergence is proportionally faster than the standard one whenever at least one unbiased source exists.
To establish these results, we impose the following assumptions. The first assumption describes the data-generating process for potential outcomes.
The second assumption concerns the assignment mechanism. Beyond the structure discussed in the setup, we impose the following minimal restriction.
Assumption (ref) guarantees that the number of times treatment $d$ is assigned to covariate profile $x$ diverges almost surely as $t$ grows. Intuitively, the assumption allows the probability of assignment to decay but not too quickly.\footnote{See Lemma (ref) in Appendix (ref) for a formal statement and a discussion of its role.} It generalizes the standard overlap assumptions in randomize control trials and it is satisfied by standard heuristic policies widely used in sequential experimentation. For generalized $\epsilon$-greedy algorithms, we have $\delta_{i}(d \mid x) \ge \epsilon_{i}$, so Assumption (ref) requires only that $(\epsilon_{i})_{i}$ decays more slowly than $1/i$. For Thompson Sampling and UCB algorithms, it is well known that \[ \sum_{i=1}^{t} \delta_{i}(d \mid x) \asymp t \quad \text{for the optimal arm}, \qquad \sum_{i=1}^{t} \delta_{i}(d \mid x) \asymp \log t \quad \text{for suboptimal arms}, \] under mild regularity conditions; see, for example, auer2002finite for UCB and pmlr-v23-agrawal12 and Kaufman2012 for Thompson Sampling. Hence, all three heuristics satisfy Assumption (ref).
In general, Assumption (ref) allows for decreasing assignment probabilities but rules out excessively fast decay. The experimenter must continue to explore each treatment--covariate pair infinitely often, though possibly at a slowly vanishing rate.
To simplify the analysis, we rely on asymptotic techniques. However, in order to better approximate the empirical settings we consider, we deviate from the standard asymptotic framework.
In many applications, the size of the treatment groups in the target experiment and in the prior source $s$ are of comparable magnitude. In such cases, the usual asymptotics --- where $\nu^{s}(d,x)$ is fixed while $N_{t}(d,x)$ diverges --- may not provide an accurate approximation to finite-sample behavior. To address this issue, we adopt a nonstandard asymptotic framework in which the prior precision $\nu^{s}(d,x)$ is allowed to depend on $t$. We write $\nu^{s}_{t}(d,x)$ to make this dependence explicit.
We assume that $(\nu^{s}_{t}(d,x))_{t}$ diverges and satisfies
Allowing the prior precision to depend on $t$ should be understood as a mathematical device used to approximate empirically relevant settings in which both the prior sample size and the sample size of the current experiment are large. This asymptotic framework nests the standard one as a special case: setting $c^{s}(d,x)=0$ corresponds precisely to the usual assumption that the prior precision remains fixed while $N_{t}(d,x)$ diverges.
As presented in the setup, the estimator uses BMA to aggregate across the $L+1$ sources. The resulting weights $\alpha^{s}_{t}(d,x)$ --- the posterior probability that model $s$ best fits the observed data --- can be interpreted as the experimenter's subjective probability that source $s$ is externally valid for the pair $(d,x)$ in the current experiment. To formalize this, we introduce a quantitative measure of external validity and relate it to the asymptotic behavior of the weights $(\alpha^{s}_{t}(d,x))_{s=0}^{L}$.
Let $E : \mathbb{R}_{+} \times \mathbb{R}_{+} \to \overline{\mathbb{R}}$ be defined by
The first argument represents the “misspecification cost" (mc) associated with a source, and the second represents the precision associated with its marginal likelihood for $(d,x)$. The value $E(mc,p)$ provides a continuous measure of external validity: unbiased sources (with $p\geq 1$) satisfy $E(0,p)>0$, and this value increases with precision, with the extreme case corresponding to a degenerate source with arbitrarily large $p$. Biased sources instead satisfy $E(mc,p)<0$, and their external validity worsens as precision increases, since their probability mass becomes more tightly concentrated around an incorrect value. Thus, unlike frequentist treatments where external validity is typically binary, this Bayesian measure varies continuously with both the bias and the precision of the source.
For the Gaussian model, for any source $s$ and any treatment-covariate pair $(d,x)$, the misspecification cost is captured by $mc^{s}(d,x) : = (bias^{s}(d,x))^{2} : = (\theta(d,x) - \zeta^{s}_{0}(d,x))^{2}$ --- the square of the difference between the prior value and the true expected outcome.\footnote{As it turns out, what matters for external validity is not the bias in the sense of a difference in parameters, but bias in the sense of systematic likelihood loss. This loss is quantified by the Kullback–Leibler divergence. In the Gaussian case with known variance, this divergence is exactly proportional to the squared difference in means, so squared bias emerges as a convenient and exact summary of misspecification. We refer the reader to Section (ref) for a more thorough explanation and additional instances of the misspecification cost for other settings.}
The next result provides an asymptotic equivalence between the Bayesian posterior weights and the mapping $E$.
The intuition behind Proposition (ref) is as follows. The weight $\alpha^{s}_{t}(d,x)$ represents the posterior probability assigned to source $s$, and its behavior is governed by the marginal likelihood of that source relative to the others. Under standard asymptotics—where $N_{t}(d,x)$ diverges while $\nu^{s}_{t}(d,x)$ remains fixed—it is well known that $\alpha^{s}_{t}(d,x)$ becomes asymptotically proportional to the marginal likelihood evaluated at the true parameter. In this regime, one source of uncertainty, the sampling uncertainty, is resolved at rate $N_{t}(d,x)$, but the other source, the prior uncertainty, never disappears because the prior precision is fixed. The latter prevents perfect separation across sources, and asymptotically the weights remain proportional to $e^{\frac{1}{2} E((bias^{s}(d,x))^{2},\nu^{s}_{t}(d,x)) }$.
Under our asymptotics, however, $\nu^{s}_{t}(d,x)$ is permitted to diverge, which changes this behavior. Let us first analyze the case where $\nu^{s}_{t}(d,x)$ diverges but at a slower rate than $N_{t}(d,x)$, so that $c^{s}(d,x)=0$. In this case, perfect discrimination among sources becomes possible asymptotically, and the determining factor in the asymptotic behavior of $\alpha^{s}_{t}(d,x)$ becomes the bias --- distance between $\theta(d,x)$ and the prior mean $\zeta^{s}_{0}(d,x)$ ---, with smaller bias yielding exponentially larger posterior weight. The rate of decay is given by $\nu^{s}_{t}(d,x)$, the rate at which precision grows. Thus, although the weights remain asymptotically proportional to $e^{\frac{1}{2} E(bias^{s}(d,x),\nu^{s}_{t}(d,x)) }$, the interpretation differs because the prior precision is no longer fixed.
A third case arises when $\nu^{s}_{t}(d,x)$ diverges proportionally with $N_{t}(d,x)$, so that $\nu^{s}_{t}(d,x)/N_{t}(d,x) \to c^{s}(d,x) > 0$. In this setting, the precision stemming from the prior remains relevant asymptotically, but now the effective precision of the marginal likelihood becomes $N_{t}(d,x)\,\frac{c^{s}(d,x)}{1+c^{s}(d,x)}$. This quantity reflects the combined contribution of both the experimental data and the prior source, with the factor $\frac{c^{s}(d,x)}{1+c^{s}(d,x)}$ determining how much of the total precision is attributable to the prior. When $c^{s}(d,x)$ is large, the overall precision approaches that of the current experiment; when $c^{s}(d,x)$ is small, residual prior uncertainty persists. The term $E(bias^{s}(d,x), \nu^{s}_{t}(d,x)/(1+c^{s}(d,x)))$ captures this relationship exactly: the bias penalizes the source, and the effective precision amplifies or mitigates this penalty depending on the informativeness of source $s$. For this reason, $\nu^{s}_{t}(d,x)/(1+c^{s}(d,x))$ is an appropriate asymptotic measure of the precision of source $s$ when its precision grows proportionally to that of the target experiment.
Based on this discussion, we define the external validity of source $s$, for pair $(d,x)$ at instance $t$ as
where the inputs are the bias and the precision.
Proposition (ref) provides, asymptotically and almost surely, an isomorphism between the posterior odds ratio of the sources and their external validity, i.e.,
This relationship allow us to obtain a source selection in terms of their external validity. To see this, let, for any pair $(d,x)$, $\mathcal{U}(d,x) : = \{ s \in \mathcal{S} \colon bias^{s}(d,x) =0 \}$ be the set of unbiased sources (which could be empty). For any biased source, $b \notin \mathcal{U}(d,x)$, $\mathbb{EV}^{b}_{t}(d,x)$ diverges to minus infinity with its size, $\nu^{b}_{t}(d,x)$, while an unbiased source is diametrically opposite, diverging to plus infinity with its size. Therefore, if unbiased sources exist, equation (ref) implies that bias sources receive exponentially vanishing weight.
This result has an oracle-type feature: asymptotically, our estimator assigns weight only to unbiased sources. However, this result is silent about what happens if there are no unbiased sources --- all sources are mis-specified. In this case, we are able to obtain a robustness-type result provided that diffuse sources are included. A source is considered to be diffuse if $\sup_{t} \nu_{t}(d,x) \leq K < \infty$ for some small constant $K$.\footnote{Asymptotically, the constant $K$ need not be small, simply finite. However, for finite sample and to capture the idea of a “diffuse" source, $K$ should be chosen to be small.} For such sources the $\mathbb{EV}^{b}_{t}(d,x)$ remains bounded (recall that for biased sources this quantity diverges to minus infinity). Therefore, when all sources are biased, equation (ref) implies the diffuse source will accumulate all the weight exponentially fast and thus our estimator will be essentially equal to the standard one. In this sense, we view this result as a robustness property: Our estimator performs no worse than the standard one even if all sources are mis-specified.
Even though the result relies on the presence of a diffuse source, the experimenter can always include one. The condition should therefore be interpreted as a practical recommendation rather than a formal restriction.
The next proposition formalizes this discussion.
We conclude by pointing out that this result does not rank sources within the class of unbiased sources. For two unbiased sources, $u$ and $u'$, equation (ref) implies that
So the procedure assigns larger weight to more informative sources, but the weights do not collapse onto a single source.
As mentioned above, the object of interest is the average effect of each treatment. At each instance $t$ and for each $(d,x) \in \mathbb{D} \times \mathbb{X}$, the experimenter estimates this effect using
The next result is the main result in the paper and establishes the rate at which this estimator concentrates around the true expected outcome $\theta(d,x)$. Henceforth, let $\ell : [1,\infty) \to \mathbb{R}_{+}$ be any increasing function such that $\int_{1}^{\infty} 1/(x \ell(x)^{2}) dx < \infty $.\footnote{This function serves as a scaling factor for our almost sure concentration rates, and it stems from classical results; its role is explained in Lemmas (ref) and (ref) in the Appendix (ref).}
\paragraph{Remarks.}
The term $\frac{1}{\sqrt{N_{t}(d,x)}}$ is the standard almost-sure rate for estimating $\sum_{i=1}^{t} \mathbf{1}\{D_{i}(x)=d\} Y_{i}(d,x)/N_{t}(d,x)$, and the scaling factor $\ell(N_{t}(d,x))$ is a standard "loss" in almost sure results, e.g., $\ell(x) = \log x$; see Lemma (ref) in Appendix (ref).
When $\mathcal{U}(d,x)$ is non-empty, the convergence rate of $\widehat{\theta}_{t}(d,x)$ improves proportionally relative to the standard rate by the factor $\mathcal{A}_{t}(\mathcal{U}(d,x))$. This factor is an average of $(1+c^{s}(d,x))^{-1}$ across unbiased sources, weighted by their posterior probabilities. It always lies in $[0,1]$ and is strictly less than one whenever at least one unbiased source has size proportional to the target experiment. Only unbiased sources contribute to this improvement because, as shown in Propsotion (ref), the posterior eventually assigns positive weight exclusively to unbiased sources whenever they exist. For example, if $c^{s}(d,x)$ takes the common value $c(d,x)$ across all unbiased sources, then the rate is accelerated by the multiplicative factor $(1+c(d,x))^{-1}$. Even moderate relative precision of the prior source can therefore generate a meaningful proportional gain. In this sense, our estimator enjoys an oracle-type property by (asymptotically) putting all the weight on unbiased sources.
If $\mathcal{U}(d,x)$ is empty the concentration rate is asymptotically equal to the standard one, provided a diffuse source is included --- a diffuse source can always be included by adding a source with arbitrarily small precision. In this case, as shown in Proposition (ref), the diffuse source dominates all biased sources and $\alpha^{diffuse}_{t}(d,x) \to 1$. The estimator converges at the standard rate, thereby yielding a natural robustness property: even when all sources are biased, the aggregation procedure performs no worse (up to constants) than the usual estimator based solely on the target experiment.
Thus, with our procedure the experimenter does not need to identify in advance which sources are unbiased or correctly specified. When unbiased sources of comparable size exist, the procedure automatically assigns them weight and converges at a strictly faster rate—behaving as if it were an “oracle” with knowledge of which sources are unbiased. When no unbiased sources exist, the procedure remains robust and achieves the standard convergence rate.
\paragraph{PAC interpretation.} A PAC interpretation of our results further clarifies the gains delivered by incorporating unbiased sources. The convergence rate in Theorem (ref) implies that achieving a target precision $\varepsilon$ requires on the order of $(1/\varepsilon^2)\log(1/\varepsilon)$ observations under the standard regime, which matches the canonical PAC rate for estimating a mean. By contrast, when an unbiased source of relative size $c$ exists, the effective rate is scaled by the factor $A= 1/(1+c) <1$, so the required sample size decreases to approximately $(A/\varepsilon^2)\log(1/\varepsilon)$. Thus the presence of unbiased sources reduces the number of observations needed to attain a given accuracy by a proportional factor $A$, reflecting the larger effective sample size generated by external information. In this sense, the procedure behaves as an adaptive PAC learner: in the presence of unbiased sources it attains a strictly smaller required sample size for a given precision, while in misspecified settings (all sources biased) it reverts to the standard PAC rate.
This section reports Monte Carlo evidence on the finite-sample performance of our estimator with multiple sources that vary in bias and effective sample size --- for simplicity, we focus on the case of no covariates. We focus on the behavior of the scaled absolute error in treatment arm $d$, \[ \sqrt{N_T(d)}\,|\hat{\theta}(d)-\theta(d)|, \] where $N_T(0) = N_T(1) = : N_T$ is the number of observations per arm (in our balanced design, $N_T=T/2$). Throughout, potential outcomes satisfy $Y(0)\sim \mathcal{N}(1,1)$ and $Y(1)\sim \mathcal{N}(1.3,1)$, and each experiment is replicated $1{,}000$ times for each design point. For ease of exposition and without loss of generality, we restrict attention to the control arm, $d=0$.
In each experiment, we observe $N_T=T/2$ draws per arm and compute two estimators of $\theta(0)$: (i) the standard sample mean $\bar{Y}(0)$ and (ii) our estimator that averages across sources using posterior model weights as defined in Equation (ref). Each source $s$ is characterized by an initial mean $\zeta_s(0)$ and a precision (effective sample size) parameter $\nu_s(0)$. We parameterize the strength of the non-diffuse sources as scaling with the experiment sample size via $e\in\{0.5,1,2\}$, so that (holding fixed baseline multipliers) $e$ can be interpreted as the source's effective sample size relative to the per-arm sample size in the experiment.\footnote{In the simulations, the diffuse source is kept weak with a fixed precision $\nu=1$ to represent a low-information baseline.}
We consider three configurations:
For each model, Figures (ref)--(ref) report mean scaled absolute errors in arm $0$ for ours and the standard estimator across $T\in\{50,100,250,500,750\}$ (i.e., $N_{T}(0)\in\{25,50,125,250,375\}$). Each figure contains three panels corresponding to $e\in\{0.5,1,2\}$. Appendix Figures (ref)--(ref) report the full distribution of scaled errors using box plots, to verify that mean effects are not driven by a small set of outliers.
The scaled absolute error captures the speed of learning. According to Theorem (ref), for models 1 and 3, our estimator will present a faster learning rate than the standard one --- faster by a factor of $1/(1+e)$. Whereas for model 2, our estimator will present a learning rate comparable to the standard one, despite not having unbiased sources.
\paragraph{A model with an unbiased source.} Figure (ref) shows that our estimator delivers sizable gains when an informative source is correctly centered. Across sample sizes, the mean scaled error of the standard estimator is approximately stable (around $0.8$), consistent with Gaussian sampling and $\sqrt{N_T}$ scaling. In contrast, our estimator's errors are uniformly lower and decline modestly with $T$ within each $e$-panel. For example, mean scaled error for our estimator falls from $0.622$ at $N_{T}(0)=25$ to $0.551$ at $N_{T}(0)=375$ when $e=0.5$, from $0.536$ to $0.442$ when $e=1$, and from $0.425$ to $0.332$ when $e=2$.
The alpha-weight table reinforces this interpretation. Table (ref) shows that the posterior weight on the unbiased source rises with the experiment sample size: for instance, in Model 1 the average $\alpha$-weight on the unbiased source increases from about $0.71$--$0.75$ at $N_{T}(0)=25$ to about $0.90$--$0.91$ by $N_{T}(0)=375$ (with slightly higher weights for larger $e$). Thus, the improvements in Figure (ref) reflect systematic reweighting toward the externally valid informative source as data accumulate.
\paragraph{A model with a biased source.} Figure (ref) shows that when the only informative alternative is biased, our estimator can perform worse in small samples---particularly when the biased source is not very informative in effective sample size (low $e$). At $N_{T}(0)=25$, mean scaled error of our estimator exceeds that of the sample mean for all $e$ (e.g., $0.879$ vs.\ $0.818$ for $e=0.5$; $0.864$ vs.\ $0.830$ for $e=1$; $0.804$ vs.\ $0.796$ for $e=2$). Intuitively, with limited experimental data, the posterior model weights can still place nontrivial mass on the biased source, and the resulting posterior mean inherits some of that bias, worsening finite-sample accuracy.
As $N_{T}(0)$ increases, our estimator approaches the performance of the standard estimator. By $N_{T}(0)=50$, our estimator and the sample mean are already close, and for larger sample sizes the two estimators are nearly indistinguishable in mean scaled error. This convergence is mirrored in the alpha weights: Table (ref) shows that the mean $\alpha$-weight on the biased source in Model 2 is small even at $N_{T}(0)=25$ (about $0.09$ for $e=0.5$, $0.04$ for $e=1$, and $0.018$ for $e=2$) and becomes essentially zero by $N_{T}(0)=50$ and beyond. Thus, the small-sample underperformance arises precisely in the range where the biased source still receives some posterior weight and is most pronounced when the source is relatively weak (low $e$).
\paragraph{A model with an unbiased source and a biased source.} Model 3 allows our estimator to choose among an unbiased informative source and a biased informative source, in addition to the diffuse baseline. Figure (ref) shows that adding an unbiased competitor restores robustness: Our estimator performance closely tracks Model 1 for moderate and large samples, with only modest degradation in the smallest sample. For instance, at $N_{T}(0)=25$ mean scaled error is $0.667$ for $e=0.5$, $0.570$ for $e=1$, and $0.446$ for $e=2$, compared to $0.622$, $0.536$, and $0.425$ in Model 1; by $N_{T}(0)\ge 50$, mean scaled errors in Model 3 are essentially identical to those in Model 1.
The alpha table clarifies why. Table (ref) shows that the biased source receives only a small initial weight in Model 3 (e.g., $0.034$ at $N_{T}(0)=25$ for $e=0.5$, and smaller for larger $e$) and this weight collapses quickly with $N_{T}(0)$. Meanwhile, the $\alpha$-weight on the unbiased source rises with $N_{T}(0)$ and converges to the same levels as in Model 1. As a result, Model 3 inherits the precision gains from the valid informative source while rapidly discarding the biased alternative.
\paragraph{Distributional evidence.} Appendix Figures (ref)--(ref) show that these patterns hold across the distribution of simulation outcomes, not only in means. In Model 1, our estimator’s error distribution is uniformly shifted downward relative to the standard estimator, with tighter interquartile ranges as $e$ increases. In Model 2, the main discrepancy occurs in the smallest sample size, where our estimator's distribution exhibits a modest rightward shift relative to the sample mean (especially for low $e$); for larger samples the distributions largely coincide. In Model 3, the small-sample penalty is small and disappears quickly, with the BMA distribution converging to the Model 1 benchmark as the biased source weight vanishes.
\paragraph{Summary.} Across models, the figures and alpha weights jointly show that our theoretical results accurately captures the behavior of our estimator, even in small samples, and that our estimator’s finite-sample performance is governed by the interaction between (i) the informativeness of external sources (controlled by $e$), (ii) their external validity (bias vs.\ unbiasedness), and (iii) posterior weight concentration. When an unbiased informative source is available (Models 1 and 3), posterior weights tilt toward it and our estimator yields substantial efficiency gains. When only a biased informative source is available (Model 2), our estimator can underperform in small samples---particularly when the source effective sample size is small---but it rapidly learns to downweight the biased source, leading to performance that becomes nearly indistinguishable from the sample mean as $N(0)$ increases.
In this section we consider two extensions of the baseline model. The first replaces Gaussian outcomes with binary outcomes and leads to a Bernoulli specification. The second keeps the Gaussian environment but allows each source to use its own set of control variables.
The main message is that Proposition (ref) and Theorem (ref) continue to hold in both settings. As in the baseline model, the BMA weights $\alpha_t$ are governed by external validity, understood as the cost of reconciling the source prior value with the target parameter. The form of the cost, however, changes. This cost is naturally expressed in terms of Kullback-Leibler divergence, so its functional form changes outside the homoskedastic Gaussian benchmark: in the Bernoulli model it is no longer quadratic in the mean, and with source-specific controls it becomes matrix-valued because learning is no longer treatment-by-treatment. Even so, the core logic is unchanged: sources that fit the target environment better receive more weight.
\paragraph{Data and parameter.} Outcomes satisfy $Y_t(d,x) \in \{0,1\}$, and the object of interest is the mean success probability \[ \theta(d,x) := \mathbb E[Y(d,x)] \in (0,1). \] Let $N_t(d,x) := \sum_{i=1}^t \mathbf 1\{D_i(x)=d\}$ and $K_t(d,x) := \sum_{i=1}^t Y_i(d,x)\mathbf 1\{D_i(x)=d\}$, and denote the sample mean by $\bar Y_t(d,x) := \frac{K_t(d,x)}{N_t(d,x)}$.
\paragraph{Prior sources and Updating.} Each source $s \in \mathcal{S} = \{0,\dots,L\}$ is represented by a Beta prior over $\theta(d,x)$, $\mathrm{Beta}\big(a^s_{0}(d,x),\, b^s_{0}(d,x)\big)$. As in the Gaussian case, we reparametrize priors by their mean and precision: $a^s_{0} = \nu^s \zeta^s_0$ and $b^s_{0} = \nu^s (1-\zeta^s_0)$. That is, these parameters are the successes and failures associated to the prior source $s$, so $\zeta^{s}_{0}(d,x)$ represents the prior value of the probability of outcome being equal to one, while $\nu^{s}(d,x)$ reflects the precision.
For each $\theta$, the outcome distribution is given by $p_\theta(y) = \theta^y (1-\theta)^{1-y},~ y \in \{0,1\}$. Hence, conjugacy implies the posterior is $\mathrm{Beta}\big(a^s_{0}(d,x)+K_t(d,x),\; b^s_{0}(d,x)+N_t(d,x)-K_t(d,x)\big)$ with posterior mean
This equation is analogous to expression (ref) for the Gaussian model.
\paragraph{The BMA weights.} As above, the $\alpha$ weights depend on the integrated likelihood. For each source $s \in \mathcal{S}$, \[ \alpha^s_t(d,x) = \frac{\mathcal M^s_t(d,x)}{\sum_{s' \in \mathcal{S}} \mathcal M^{s'}_t(d,x)}, \] where \[ \mathcal M^s_t(d,x) = \frac{ B\!\left(a^s_{0}(d,x)+K_t(d,x),\; b^s_{0}(d,x)+N_t(d,x)-K_t(d,x)\right) }{ B\!\left(a^s_{0}(d,x),\; b^s_{0}(d,x)\right) }, \] where $B(\cdot,\cdot)$ denotes the Beta function. The BMA estimator is defined as before, \[ \widehat\theta_t(d,x) := \sum_{s \in \mathcal{S}} \alpha^s_t(d,x)\,\zeta^s_t(d,x). \]
Finally, we impose the same nonstandard asymptotics as in the Gaussian case: Allowing $\nu^{s}(d,x)$ to be indexed by $t$ and $\frac{\nu^s_t(d,x)}{N_t(d,x)} \to c_s(d,x) \in \mathbb R_+,~\text{a.s.}$. Sources with $c_s(d,x)=0$ are asymptotically diffuse, while $c_s(d,x)>0$ corresponds to sources whose information content is comparable to the experiment.
\paragraph{External Validity Measure.} The external validity of source $s$ for $(d,x)$ at instance $t$ is given by
where
Expression (ref) is a generalization of expression (ref) to non-Gaussian frameworks wherein the KL divergence is not quadratic. The quantity $\Psi^{s}(\theta(d,x),\zeta_0^{s}(d,x))$ therefore represents the (asymptotic) penalty incurred when the source posterior is required to reconcile its prior belief $\zeta_0^{s}(d,x)$ with the target truth $\theta(d,x)$. Larger values of $c^{s}(d,x)$ (more dogmatic sources) magnify the cost of discrepancy.
However as expression (ref) in the Gaussian model, expression (ref) also acts as a notion of distance between $\theta(d,x)$ and $\zeta^{s}_{0}(d,x)$. Indeed, it is not hard to show that for $\Psi^{s}(\theta(d,x),\zeta_0^{s}(d,x)) =0$ if $\theta(d,x)=\zeta_0^{s}(d,x)$ and positive otherwise.
\paragraph{Theoretical Guarantees for the Bernoulli Model.} We conclude by extending all our results extend to the binary outcomes case. The next result is analogous to Proposition (ref).
The concept of unbiased source remains unchanged in this new setup, because, as pointed out above, $\Psi^{s}\!\big(\theta(d,x), \zeta_{0}^{s}(d,x)\big)$ is naught only if the source is unbiased (or diffuse). Therefore, an analogous result to Proposition (ref) holds: Bias sources will be discarded exponentially fast --- at rate given by $N_t(d,x) \Psi^{s}(\theta(d,x), \zeta^{s}_0(d,x))$ --- in favor of unbiased ones (or diffused should all sources be biased). Consequently, Theorem (ref) also holds for the Bernoulli model.
In addition to extending Theorem (ref) --- and our theory --- to binary outcomes, the main takeaway of this section is that the posterior weights continue to be governed by the same external validity logic as in the Gaussian case, even though the squared-Euclidean metric is replaced by the information-theoretic divergence $\Psi^{s}(\theta(d,x),\zeta_0^{s}(d,x))$, which emerges from the joint KL projection of the target and source beliefs. Thus, the geometry of comparison changes but the economic content of the weights remains identical: sources are rewarded or penalized according to how costly it is, in information terms, to reconcile their prior belief with the target environment, with the penalty scaled by their effective precision $c^{s}(d,x)$.
\paragraph{Model.} Let $\boldsymbol{\theta}:=(\theta(0),\ldots,\theta(M))^{\top}$ denote the vector of treatment-specific mean parameters, let $Z_{t} = (1\{ D_{t} = 0 \},\ldots , 1\{D_{t} = M\})^{\top}$ be the $(M+1)\times 1$ treatment-indicator vector, and let $W_t^{s}\in\mathbb R^{p_s}$ be the predetermined vector of controls used by source $s$, with associated nuisance coefficient $\gamma^{s}\in\mathbb R^{p_s}$. We work with the linear Gaussian regression
where $\varepsilon_t^s \sim \mathcal N(0,1)$. The parameter of interest is $\boldsymbol{\theta}$, so the average treatment effect of arm $d$ relative to the baseline arm is $\theta(d)-\theta(0)$. Allowing $W_t^s$ and $\gamma^s$ to vary across sources captures the fact that different prior studies may use different controls, while keeping the treatment means comparable across sources. As in the benchmark Gaussian model, normality is imposed for analytical convenience; for identification of the conditional mean, the essential requirement is a mean-zero error with finite variance.
\paragraph{Likelihoods, Sources, and Posteriors.} The posterior in this model is analogous to the one for no-controls but with an important caveat. Due to the fact the coefficients $\gamma^{s}$ can be common across treatment, learning does not longer take place “treatment-by-treatment" but jointly. We now formalize this.
For each source $s \in \mathcal{S}$, let \[ X_{t}^{s}:=\big[\,Z_{t}\;\; W_{t}^{s}\,\big]\in\mathbb R^{1\times((M+1)+p_s)}, \qquad \beta^{s}:=
\in\mathbb R^{((M+1)+p_s) \times 1}. \] Then expression (ref) implies the Gaussian working likelihood
Each source $s$ is represented by Gaussian prior over its model-specific parameter $\beta^{s}$:
where $\boldsymbol{\zeta}_0^{s}\in\mathbb R^{M+1}$ is the source mean for $\theta$ and $\eta_0^{s}\in\mathbb R^{p_s}$ is the source mean for $\gamma^{s}$. We allow the prior covariance $\Sigma_{0,t}^{s}$ to depend on $t$ so as to reproduce the same “growing information in the sources” regime as in the benchmark model.
Specifically,\footnote{Here and throughout, for any vector $X$, $\operatorname{Diag}[X]$ denotes a diagonal matrix with diagonal components given by the vector $X$.}
where $\lambda^{s}_{t}$ is a $p_{s} \times 1$ vector uniformly bounded away from zero.
The term $\operatorname{Diag}[1/\nu^{s}_{t}]$ is completely analogous to $\nu^{s}_{t}(d,x)$ in the no-controls case. The vector $\lambda^{s}_{t}$ represents the prior precision associated to the control coefficients. We assume that $\lambda_{t}^{s}/t = o(1/\sqrt{t})$, so the effect of this prior precision vanishes asymptotically. This assumption is only used to simplify the asymptotic expressions.
By conjugacy, for each source $s\in\mathcal S$, the posterior over $\beta^{s}$ is Gaussian, \[ \beta^s \mid Y_{1:t},D_{1:t},W_{1:t}^s \sim \mathcal N(\beta_t^s,\Sigma_t^s), \] where $\Sigma_t^s$ is the posterior covariance matrix and the posterior mean is $\beta_t^s=((\boldsymbol\zeta_t^s)^\top,(\eta_t^s)^\top)^\top$.\footnote{See Appendix (ref) for a proof of this result and the expressions for the posterior mean and covariance matrix.}
To isolate the treatment-effect component, partition the posterior precision matrix as
where $A_t^s$ is the treatment-treatment block, $B_t^s$ is the treatment-control block, and $C_t^s$ is the control-control block. Then Appendix (ref) provides expressions for these terms, in the main text we only need the induced expression for the treatment posterior mean.
\paragraph{Estimator.} Our goal is to estimate $\boldsymbol{\theta}$. Without prior sources, the standard estimator is the OLS coefficient on the treatment indicators in a regression of $Y_t$ on $(Z_t,W_t)$, or equivalently the Frisch-Waugh-Lovell residualized estimator that partials out the controls before averaging within treatment arms. Based on this intuition, a natural source-specific estimator uses the same logic while incorporating the prior information. This is formalized in the next lemma. Define $N_t := (N_t(0),\ldots,N_t(M))^\top$ and $\boldsymbol m_t := (m_t(0),\ldots,m_t(M))^\top$ with $m_t(d) := \frac{1}{N_t(d)}\sum_{i=1}^t 1\{D_i=d\}Y_i$.
Lemma (ref) is the special case of the more general block posterior system that is relevant for our treatment-effect analysis; the full statement is reported in Appendix (ref). The expression combines a vector-valued version of (ref), where each component averages the prior mean $\zeta_0^s$ and the empirical treatment mean $m_t$, with the Frisch-Waugh-Lovell adjustment that partials out the controls. If the precision matrix $(\Sigma_t^s)^{-1}$ were diagonal, the treatment effects would update independently across arms. With controls, however, the off-diagonal block $B_t^s$ captures the empirical link between treatment assignments and controls, so updating must be done jointly.
In particular, the term \[ B_t^s (C_t^s)^{-1} \left( \operatorname{Diag}[\lambda_t^s]\eta_0^s + \sum_{i=1}^{t} (W_i^s)^\top Y_i \right) \] subtracts the part of $ \operatorname{Diag}[\nu^{s}_{t}] \boldsymbol{\zeta}_0^s + \operatorname{Diag}[N_{t}] \boldsymbol{m}_{t} $ that is explained by the controls. The matrix $T_t^s$ is the Schur complement of the control block and represents the effective information about the treatment effects after accounting for the controls. In this sense, the posterior treatment effects correspond to a precision-weighted update based on treatment information that has been residualized with respect to the controls, mirroring the role of the Frisch--Waugh--Lovell theorem in classical regression.
The estimator of $\boldsymbol{\theta}$ takes the ensemble of estimators $(\boldsymbol{\zeta}^{s}_{t})_{s \in \mathcal{S}}$ and combines them using BMA weights,
where $\alpha_t^{s} \;:=\; \frac{\pi_s\, m_t^{s}(Y_{1:t}\mid D_{1:t},w_{1:t}^{s})} {\sum_{r\in\mathcal S}\pi_r\, m_t^{r}(Y_{1:t}\mid D_{1:t},w_{1:t}^{r})}$, with $m_t^{s}(Y_{1:t}\mid D_{1:t},w_{1:t}^{s}) :=\int \prod_{i=1}^{t} \phi(Y_{i} ; X^{s}_{i} \beta , 1 ) \phi (\beta;\beta_{0}^{s},\Sigma_{0,t}^{s})\,d\beta$ being the source marginal likelihood. Unlike the no-controls case, these source weights are not treatment-specific because the controls make posterior updating joint across treatment arms.
\paragraph{Theoretical guarantees.} In this section we present the theoretical guarantees for this model. The next proposition links the weights of the source to the External Validity measure.
The matrix $\bar H^{s}$ represents the asymptotic curvature of the marginal likelihood with respect to deviations of the treatment parameters from the benchmark value $\beta_0^{s}$. It shows that this curvature arises from the interaction between the prior precision contributed by source $s$ and the information about the treatment parameters available in the experimental data after accounting for the nuisance parameters. The matrix $\operatorname{Diag}[c^{s}\cdot\delta]$ captures the prior precision of the source, while $K$ represents the information matrix for the treatment parameters after partialling out the controls via the Schur complement --- $K$ measures how informative the experiment is about treatment after accounting for the controls. Consequently, $\bar H^{s}$ can be interpreted as the effective prior precision of source $s$ once it is filtered through the experimental information about the treatment parameters.
The term $ \frac{1}{t} \left( 2d_{s} \log t + \log | \bar{Q}^{s} (\Sigma^{s}_{0,t}/t) | \right) $ is the standard “log $t$" penalization term obtained in BMA regression models (e.g. BIC criteria, Fernandez2001) but extended to incorporate the fact that prior precision may not be vanishing.
Thus, the leading term of the source weights for the Gaussian regression with controls is given by
To shed more light on this expression, consider the absence of controls. In this case the matrix $K$ simplifies to $K = \operatorname{Diag}[(1+c^s)\cdot\delta]$. Substituting this expression into the definition of $\bar H^s$ yields (up to negligible $o_{as}(1)$ terms) \[ (t\bar H^s) = \operatorname{Diag}\left[ t \delta \cdot \frac{c^s}{1+c^s} \right] = \operatorname{Diag}\left[ N_{t} \cdot \frac{c^s}{1+c^s} \right] = \operatorname{Diag}\left[ \frac{\nu^{s}_{t} }{1+c^s} \right] \] Thereby yielding a leading term in the EV measure of $ \left( \boldsymbol{\theta} - \boldsymbol{\zeta}^{s}_{0} \right)^{\top} \operatorname{Diag}\left[ \frac{\nu^{s}_{t} }{1+c^s} \right] \left( \boldsymbol{\theta} - \boldsymbol{\zeta}^{s}_{0} \right) $ which is a vector-valued version of the EV measure for the no-controls case given in expression (ref).
We now show that an analogous result to Theorem (ref) holds in this framework. To see this, we provide a characterization of $ \boldsymbol{\zeta}_t^s$ based on Lemma (ref).
Suppose $\mathcal{U}$ is non-empty. Then, by proposition (ref), biased sources will be given a weight that converge to zero at least at rate $1/t$. Hence,
Any source in $\mathcal{U}$ satisfies, by definition, $ \boldsymbol{\zeta}^{s}_{0} = \boldsymbol{\theta} $. This and Lemma (ref) imply that
By definition of $(V_t^s+M_t^s)$ and strong LLN it is easy to see that
where $P_{\boldsymbol{W}^{s}}$ is the projection matrix onto $\boldsymbol{W}^{s}$.
The fact that $\lambda^{s}_{t}/t = o(t^{-1/2})$ and the strong LLN imply that $r_{t} = o_{as}(\ell(t)/\sqrt{t})$ where $\ell$ is as in Theorem (ref). Moreover, $E[(Z)(Z)^{\top}] = t \operatorname{Diag}[\delta]$. Putting all this results together, we obtain that
This expression is completely analogous to the conclusion obtained in Theorem (ref) for the no controls case. The standard rate of convergence is given $ \frac{\ell(t)}{\sqrt{t} } \left \Vert (Var(Z - P_{\boldsymbol{W}^{s}} Z) )^{-1} \right \Vert$, and this is precisely the rate we obtain if $\mathcal{U}$ were empty but we have at least one diffuse source. However, if there are unbiased non-diffuse source, our rate is faster than the standard one due to the $\operatorname{Diag}[c^{s}] E[(Z)(Z)^{\top}] $ factor.
This paper studies learning about treatment effects in experiments (both RCTs and sequential ones) when each study starts multiple prior sources --- past pilots, related studies, or expert assessments --- whose external validity is uncertain. We formalize this environment by treating each source as a distinct model within Bayesian model averaging (BMA), and we introduce a nonstandard asymptotic framework in which the precision of each source can grow with the sample size of the ongoing experiment. This scaling captures empirically relevant settings where prior studies may be as informative as the current trial, and where standard asymptotics provide little guidance about how posterior weights should evolve.
Within this framework, we characterize posterior weights and the estimator's convergence rate through an external-validity index that depends jointly on bias and effective precision. The resulting rate exhibits an oracle property: as formalized in Theorem (ref), if at least one source is unbiased, posterior weight concentrates asymptotically on the unbiased set and the estimator converges strictly faster than the benchmark rate based only on the new experiment, with the speedup determined by the effective precisions of the valid sources. At the same time, the same theorem delivers robustness at the level of rates: when all informative sources are biased, the presence of a deliberately conservative (diffuse) prior forces posterior weight onto that source, and the estimator reverts to the standard convergence rate, avoiding contamination by precise but misspecified information. Simulations illustrate how these oracle and robustness properties manifest in finite samples and quantify the magnitude of the resulting efficiency gains and protections.
\setcounter{section}{0}