Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
105,629 characters · 15 sections · 112 citation commands
Evidence Aggregation for Treatment Choice
Keywords: Meta-analysis, program evaluation, statistical decision theory, minimax regret.
An increasing number of policy-making authorities are interested in making their policy decisions evidence-based. In evidence-based decision-making, it is crucial for a planner to acquire credible evidence of a policy's causal impact on the affected population. Obtaining credible evidence for better policy decision-making is, however, challenging in many contexts. For instance, although randomized control trials (RCTs) are considered to be ideal for obtaining evidence of the causal impact of a policy, conducting an RCT can be costly in terms of budget, time, or administrative resources. Moreover, ethical or legal constraints can prevent the use of an RCT in certain institutional environments. In contrast, observational data can be both more accessible and easier to collect, but the credibility of any resulting causal estimates is limited if the validity of these estimates relies on restrictive identifying assumptions. In scenarios where the planner faces difficulties in collecting direct evidence, a practical alternative is to analyze the publicized results of intervention studies performed for similar policies on different populations. With this approach in mind, how should the planner make use of and aggregate existing evidence to reach her policy decision?
Statistical methodologies to aggregate evidence from multiple studies have been considered in the literature of meta-analysis and research synthesis. See, for instance, HO85 and a recent handbook volume, CooperHandbook2019. Since the seminal works of Rubin81 and DL86, a common approach to aggregation of evidence is the hierarchical Bayesian approach, in which the typical objective of analysis is to infer hyper-parameters indexing the population of studies. This framework of meta-analysis is useful for “summarizing what has been learned and quantifying how results differ across the studies beyond the sampling error” (DL15). Its use is, however, limited when it comes to the planner's policy choice because the output of meta-analysis mainly concerns the population of studies rather than the particular population that is of interest to the planner. This point is made in manski2020towards:
We pursue this paradigm of `patient-centered meta-analysis' to develop a method to aggregate existing studies for the purpose of making an optimal treatment decision on the local population that is of interest to the planner (hereafter, the target population). Building on the framework of statistical treatment choice proposed by Manski2000, manski2004statistical, we formulate the planner's problem as a statistical decision problem \`{a} la Wald50. The basic formulation of the decision problem analyzed in this paper is as follows. Let $\tau_0$ be the average welfare effect of introducing a new policy to the target population. There is no data from which the planner can directly infer $\tau_0$, but she does have access to the results of existing intervention or observational studies that are indexed by $k=1, 2, \dots, K$, $K \geq 1$. Each study $k$ reports a point estimate $\hat{\tau}_k$ for the average welfare effect $\tau_k$ in the study population, and an associated estimate $\hat{\sigma}_k$ of the standard error. We allow the study population to be different from the target population, so that the average welfare effects can differ, i.e., $\tau_k \neq \tau_{k'}$ for $k \neq k'$, $0 \leq k, k' \leq K$. The planner's decision problem, which we solve in this paper, is whether or not to adopt the new policy for the target population upon observing a meta-sample, $(\hat{\tau}_k, \hat{\sigma}_k)$, $k=1, \dots, K$. That is, the statistical treatment choice rule we consider in this paper is a function $\hat{\delta}$ that maps the meta-sample to the binary choice of whether to adopt the policy or not.
Following manski2004statistical, Manski2007, stoye2009minimax, stoye2012minimax, and tetenov2012statistical, we apply the minimax regret criterion of Savage51 to obtain a minimax-regret treatment choice rule for the planner. We assume that the planner's objective function (social welfare function) is linear in $\tau_0$ and consider the class of non-randomized statistical treatment choice rules that select the treatment based on the sign of linear aggregation of $(\hat{\tau}_k : k = 1, \dots, K)$:
where $\bm{w}=(w_1, \dots, w_K)'$ is a vector of weights assigned to each estimate in the pool of studies, which does not depend on the data. Restricting the feasible rules to non-randomized (non-fractional) ones can be attractive in the following contexts. First, in the real-world practice of treatment choice or drug approval decisions, non-randomized allocation of treatments are easier than randomized ones for policy authorities to administer. Second, non-fractional allocations are guaranteed to attain parity within the target population since either everyone or no one is treated. On the other hand, fractional rules do not attain the ex-post parity. The restriction to non-randomized rules, however, sacrifice the value of minimax regret since unconstrained minimax regret rules are known to be randomized for some realization of data as shown by stoye2012minimax, yata2021optimal, and montiel2023decision.
Assuming a Gaussian sampling distribution for $(\hat{\tau}_1, \dots, \hat{\tau}_K)$ with known variances and imposing certain symmetry and invariance conditions on the parameter space for $(\tau_0, \tau_1, \dots, \tau_K)$, we derive the aggregation weights $\bm{w}_{\text{minimax}}$ leading to a minimax-regret treatment choice rule among the non-randomized rules. Analytical characterization and computation of the exact minimax regret rule often become challenging in the context of statistical treatment choice. Our approach to the planner's minimax regret aggregation rule, in contrast, overcomes these challenges by showing that some mild restrictions on the parameter space and the class of decision rules deliver analytically and computationally tractable minimax regret rules.
We assert that the perspective and tools of statistical decision theory are particularly appealing in the meta-analysis setting for the following reasons. First, if each study in the pool reports a consistent estimate using a sample of moderate to large size (e.g., the difference-in-means estimator for the average treatment effect) then, by its asymptotic normality, it is plausible to assume that $\hat{\tau}_k$ follows a Gaussian distribution centered at $\tau_k$. Hence, the standard and well-studied framework of Gaussian experiments fits well to the current meta-analysis setting. Second, it is common for the meta-sample to consist of only a small number of studies. In such instances, asymptotic analysis with $K \to \infty$ can be misleading, and deriving finite-$K$ optimal procedures, which statistical decision theory is particularly suitable for, is desirable.
As an alternative to the minimax regret treatment choice rule, one could consider using a plug-in rule that chooses the treatment according to the sign of an estimate of $\tau_0$. The plug-in rule that uses a minimax mean squared error (MSE) optimal estimate of $\tau_0$ is an example. Minimax-MSE estimation for finite-dimensional Gaussian mean models is well-studied and the minimax-MSE weights $\bm{w}_{\text{MSE}}$ are simple to compute, although the resulting plug-in decision rule does not generally possess decision theoretic optimality in terms of the planner's objective function. To quantify the welfare cost of $\hat{\delta}_{\bm{w}_{\text{MSE}}}$, we compare the worst-case regrets of $\hat{\delta}_{\bm{w}_{\text{minimax}}}$ and $\hat{\delta}_{\bm{w}_{\text{MSE}}}$, and show that the worst-case regret of $\hat{\delta}_{\bm{w}_{\text{MSE}}}$ is worse than the minimax regret only up to a constant factor of 5.88, independent of the number of studies $K$ and the parameter space.
Our framework can accommodate a vector of observable characteristics $x_k$, $k=0,1,\dots,K$, where $x_k$ includes the characteristics of the treatment and demographics of the population featured in study $k$. Under a linear functional form specification, $\tau_k = \beta_0 + x_k'\beta$, $\beta \in \mathcal{B}$, common to standard meta-regression analysis (see, e.g., SJ89), we discuss those restrictions on $\mathcal{B}$ under which we can apply our minimax regret decision rule. For minimax regret to be bounded, an important constraint is boundedness of $\mathcal{B}$, and the bounds of $\mathcal{B}$ have to be explicitly specified to obtain the minimax regret rule. In reality, the planner may not be able to come up with reasonable bounds for $\mathcal{B}$. To offer a practical solution to this difficulty, we consider a data-driven way to specify the parameter space based on confidence sets for $(\tau_1, \dots, \tau_K)$.
We illustrate the use of our minimax regret treatment rule by using two empirical examples. In the first application, we analyze whether an active labor market program should be adopted using the meta-database appearing in card2017works. We consider a pool of 14 RCT studies of job training programs, covering 8 different countries (Argentina, Brazil, Colombia, Dominica, Jordan, Nicaragua, Sri Lanka, Turkey, and the United States). Based on the average treatment effect and standard error estimates in each of these studies, and the demographic characteristics of the studied populations, we calculate the minimax regret adoption decisions for several countries (Japan, the United Kingdom, and Peru) for which the corresponding experimental estimates are not available in the meta-database.
In the second application, we consider the drug approval decision for a COVID-19 medication called Remdesivir. Remdesivir is an antiviral medication that is known to be effective against Middle East Respiratory Syndrome (MERS) and Severe Acute Respiratory Syndrome (SARS), while its effectiveness against COVID-19 remains unknown due to conflicting evidence. Using the meta-database of randomized clinical trials for COVID-19 treatments provided by juul2020interventions, we calculate the minimax regret treatment choice for Remdesivir for some specified demographic groups.
The remainder of the paper is organized as follows. The next subsection reviews the related literature. Section 2 formulates the minimax regret decision problem and shows the main analytical result of the paper. In Section 3, we compare the minimax regret with the maximum regret of the decision rule based on the minimax-MSE aggregation rule. Section 3 also discusses a data-driven construction of the parameter space. Section 4 performs numerical analysis to compare the minimax-regret aggregation rules with the minimax-MSE and meta-OLS rules.
This paper contributes to the growing literature on statistical treatment choice and individualized treatment assignment rules initiated by Manski2000, manski2004statistical and Dehejia2005. Contributions to the current literature include, HiranoPorter2009, stoye2009minimax, stoye2012minimax, Chamberlain2011, BhattacharyaDupas2012, tetenov2012statistical, Kasy2014b,Kasy2018, kitagawa2018should, KT19, KW20, Russell20, KST21, MT17, AW20, Sakaguchi21, and Viviano21, among others. The problem of individualized treatment assignment rules has also been an area of active research in the fields of medical statistics and machine learning; see, for instance, Zadrozny03, BeygelzimerLangford09 QianMurphy2011, Zhao2012JASA, SJ15, Kallus_2020, to list but a few papers. The standard setting in the existing literature considers an optimal treatment assignment policy for the population from which a sample was drawn, rather than combining the pool of estimates from multiple studies performed on different populations.
There is a growing literature on how to inform policy using multiple pieces of evidence or extrapolation from one or multiple reference populations. Dehejiaetal21 considers the use of (quasi-)experimental evidence to study the decision of whether to experiment or to extrapolate, and, if applicable, where to conduct a new experiment. Manski18 analyzes decision-making for personalized risk assessment under the ecological inference setting where (partial) identification of a long regression is obtained by combining information on a short regression and the joint distribution among the regressors. The meta-analysis setting considered in this paper differs from the ecological inference setting in terms of the object to identify and the type of information provided by the available studies. Focusing on conditional cash transfer programs, Gechteretal19 runs multiple program evaluation methods on data obtained from Mexico to inform treatment assignment policies for Morocco, and empirically compare the welfare performances of these policies. Hotzetal05 and Dehejiaetal21 analyze how to predict the effects of future programs from past experimental evaluations by adjusting for differences in the distributions of observable characteristics. AO19 propose a method to conduct sensitivity analysis and to approximate external validity bias when the trial and target populations differ in the distribution of unobservables. Gechter16 considers bounding causal effects in a target population by restricting the dependence between the treated and control outcomes.
Meta-analysis for research synthesis has been actively studied in statistics and the resulting literature is vast; see, e.g., Borenstein09textbook for a textbook and CooperHandbook2019 for a handbook volume. In economics, existing applications of meta-analysis and meta-regression include CK95, Dehejia03, Bandieraetal17, card2017works, Meager19, Meager20, Imai20, and Vivalt20. See Stanley01 for a review. The common framework of meta-analysis introduces the population of studies and draws inference for the parameters thereof. As we argue in the Introduction via the quote from manski2020towards, the usefulness of the conventional framework of meta-analsyis is not obvious for informing the planner's policy decision. This paper follows and pushes forward the perspective of patient-centered meta-analysis. The methodological proposals in manski2020towards concern predicting treatment effects for the target population by intersecting the population identified sets for $\tau_0$ formed by extrapolation from each study, rather than explicitly taking into account sampling uncertainty due to finite sample size when considering the treatment choice decision.
In terms of the framework and analytical and computational challenges for obtaining minimax regret rules, this paper is most closely related to stoye2012minimax. In one of his baseline settings, stoye2012minimax considers Gaussian experiments for conditional average treatment effects with a scalar covariate $x \in \mathcal{X}$, and analyzes the properties of the minimax regret treatment rule under a restriction that the conditional average treatment effects depend on $x$ with bounded variation. Similar to stoye2012minimax, our framework allows (study-specific) covariates to constrain the parameter space for $(\tau_0, \tau_1, \dots, \tau_K)$, but there are two aspects in which our framework differs. First, the treatment assignment rules considered in stoye2012minimax are functions $\delta: \mathcal{X} \to \{0 ,1\}$, while we concern ourselves with the treatment choice at a particular covariate value $x_0$ in $\mathcal{X}$ that corresponds to the covariate value of the target population. This reduction of the treatment choice rule from a function of $x$ to a point significantly simplifies the analysis and computation of the minimax-regret rule. Second, the conditions we impose on the parameter space for feasible computation of the minimax regret rule is general and includes the bounded variation restriction considered in stoye2012minimax as a special case.
In another baseline setting of stoye2012minimax, he derives a minimax regret treatment choice rule under partially identified welfare, where the identified set has the known width but unknown location, and a Gaussian signal is available for it. Our setting is more complex than his setting due to multiple Gaussian signals with both the location and width of the identified set unknown. Contemporaneously or after the initial version of the current paper was circulated, there have been some important advances in the literature. In an similar setting to this paper, yata2021optimal obtained an analytical representation for an exact unconstrained minimax regret rule, which implements an fractional assignment when the identified set for ATE is wide relative to the variances of the bound estimates. montiel2023decision shows that minimax regret rules are not unique and considers refining the set of minimax regret rules. An alternative approach to handle uncertainty in the bound estimates and ambiguity within the identified set is to introduce multiple priors as in giacomini2021 and performs a gamma minimax decision rule. See giacomini_etal2021 for gamma minimax decision rules for set-identified models. Christensen2022 studies this line of approach for treatment choice and shows its optimality properties in limit experiments.
Viewing $\tau_0$ as the value of the regression equation at $x_0$ in a Gaussian regression model and considering standard estimation risk such as the mean squared errors for $\tau_0$, the problem is reduced to an interpolation or extrapolation exercise based on the Gaussian signals. As such, the minimax estimation and inference problem for $\tau_0$ is similar to the extrapolation issue in the regression discontinuity setting analyzed in KR18. Recent contributions regarding estimation and inference for Gaussian sequence models are made by Johnstone17 and AK18, AK20. These papers consider statistical losses for estimation and inference, but do not consider the welfare criterion for statistical treatment choice.
Suppose we have access to the publicized results of $K$ studies indexed by $k=1, \dots, K$. Each of the studies estimates the causal effect of a particular binary policy or treatment. We allow the details of the policy (implementation protocol, dosage, program contents, etc.) to differ across the studies. For $k = 1, \cdots, K$, let $\hat{\tau}_k$ denote the estimate of the policy effect reported in study $k$ and $\sigma_k$ denote the standard error of $\hat{\tau}_k$. For simplicity, we assume $\sigma_k$ is known, although, in practice, we can only construct a consistent estimator for $\sigma_k$. We solve for the finite sample minimax regret rule with known $(\sigma_k : k=1,\dots,K)$, recommending that in practice the rule is implemented with the true standard errors replaced by their consistent estimates.\footnote{Solving the decision problem with Gaussian signals with known variances and obtaining a feasible decision rule by plugging in consistent estimators for the variances are similar to the construction of an asymptotically optimal decision rule within the framework of Gaussian limit experiments. See HiranoPorter2020 for a recent review.}
We assume
where $\tau_k$ is the true policy effect of the population featured in study $k$. We allow $\tau_k$ to vary across the studies. The assumption that $\hat{\tau}_k$ follows a Gaussian sampling distribution for is reasonable if the reported estimator $\hat{\tau}_k$ is consistent and asymptotically normal and each study has a moderate to large sample.
Throughout this paper, we consider a planner who, upon observing data $\mathbf{D} \equiv \{\hat{\tau}_k\}_{k=1}^K$, must determine whether or not to adopt the policy in the target population given that its policy effect $\tau_0$ is unknown. Following manski2004statistical, Manski2007, stoye2009minimax, stoye2012minimax, and tetenov2012statistical, we focus on minimax regret criterion to solve this decision problem. To this end, we assume that the true parameters $\bm{\tau} \equiv (\tau_0, \tau_1, \dots , \tau_K)'$ are ex ante known to belong to the parameter space $\mathcal{T}$.
We impose the following restrictions on the parameter space $\mathcal{T}$:
The symmetry assumption rules out the imposition of a sign restriction on the causal effect parameters, i.e., $\tau_k \geq 0$ for some $k$. The condition of invariance to common constant addition (hereafter, shortened to invariance) implies that $\{(\tau_k - \tau_0) : \bm{\tau} \in \mathcal{T}, \tau_0 = t\}$ does not depend on $t$. We use this result to simplify derivation of a minimax regret treatment rule. It is worth noting that the invariance condition rules out the case in which the parameter space for some $\tau_k$, $k \in \{0,1,2, \dots, K\}$, is bounded. For instance, if the outcome is binary, the treatment effect on the outcome is bounded by $[-1,1]$ for all $\tau_k$, $k=0,1, \dots, K$.\footnote{In a similar setting to the one in this paper, ishihara2023bandwidth studies the treatment choice problem with $\tau_k$, $k \in \{0,1,2, \dots, K\}$ being bounded due to binary outcomes.} However, if the standard errors $(\sigma_k^2:k=1,2, \dots, K)$ and the variations among $(\tau_0, \tau_1, \dots, \tau_K)$ imposed on $\mathcal{T}$ (e.g., Lipschitz constants $C_{kl}$, $ 0 \leq k,l \leq K$, in Example (ref) below) are sufficiently small relative to the size of the supports, bounded support of $(\tau_0, \tau_1, \dots, \tau_K)$ is less of an issue because extreme values of $\bm{\tau}$ beyond its logical support are unlikely to correspond to a worst-case in terms of the regret.
The following two examples satisfy the parameter space constraints of Assumption (ref).
When Assumption (ref) does not hold in a given application, we can still formulate the optimization problem to derive the minimax regret treatment rule, although solving for this is accompanied by a substantial increase in computational complexity. In Remark (ref) below, we discuss a derivation of the minimax regret treatment rule that does not rely upon Assumption (ref).
Given a non-randomized treatment choice action $\delta \in \{0,1\}$, define the welfare attained by $\delta$ as
where $c_0$ is the per-person cost of the policy and $\mu_0$ is the average outcome that would be realized in the absence of the policy. An optimal treatment choice action given knowledge of $\tau_0$ and $c_0$ is
Let $\hat{\delta}(\mathbf{D}) \in \{0,1\}$ be a non-randomized statistical treatment rule that maps the meta-sample $\mathbf{D}$ to the binary decision of treatment choice in the target population. The welfare regret of $\hat{\delta}(\mathbf{D})$ is defined as
where $E_{\bm{\tau}}(\cdot)$ is the expectation with respect to the sampling distribution of $\mathbf{D}$ given the parameters $\bm{\tau}$. Hereafter, we normalize the cost of treatment to $c_0=0$, i.e., interpret $\tau_k$, $k=0, \dots, K$, as the average treatment effect net of the per-person treatment cost in the target population.
The minimax regret criterion selects a statistical treatment rule that minimizes maximum regret:
where $\mathcal{D}$ is a class of statistical treatment rules. We refer to $\hat{\delta}_{\text{minimax}}$ as a minimax regret rule. In the next subsection, we derive a minimax regret rule under the class of statistical treatment rules spanned by linear aggregation rules.
To gain analytical and computational tractability, we focus on the class of linear aggregation rules.
This class rules out nonlinear treatment rules that plug in an aggregation of $(\hat{\tau}_1, \dots, \hat{\tau}_K)$ in the manner of James-Stein shrinkage or empirical Bayes. Nevertheless, the class of linear aggregation rules contains many reasonable treatment rules. For example, plug-in rules $\mathbf{1} \left\{ \hat{\tau}_0(\mathbf{D}) \geq 0 \right\}$ based on linear estimators $\hat{\tau}_0(\mathbf{D})$ for $\tau_0$ belong to $\mathcal{D}_{\text{lin}}$. When study characteristics are included in the available covariates as in Example (ref), $\mathcal{D}_{\text{lin}}$ includes those rules that plug in fitted values based on parametric linear regression or nonparametric kernel regression. Furthermore, as shown in Remark (ref) above, hierarchical Bayes decision rules under the linear meta-regression specification with Gaussian priors yields the linear aggregation rule.
We consider the minimax regret rule among $\mathcal{D}_{\mathrm{lin}}$ whose corresponding weight vector solves \[ \bm{w}_{\text{minimax}}\ \in \ \text{arg} \min_{\bm{w}} \max_{\bm{\tau} \in \mathcal{T}} R(\bm{\tau}, \hat{\delta}_{\bm{w}}). \] To develop a computation method for $\bm{w}_{\text{minimax}}$, note from ((ref)) that $$ \sum_{k=1}^K w_k \hat{\tau}_k \ \sim \ \mathcal{N} \left( \sum_{k=1}^K w_k \tau_k, \sum_{k=1}^K w_k^2 \sigma_k^2 \right). $$ Hence, from ((ref)), the regret of $\hat{\delta}_{\bm{w}}$ can be written as
where the first equality follows from the normality of $\sum_{k=1}^K w_k \hat{\tau}_k$ and the second equality follows from $1-\Phi(a) = \Phi(-a)$. We then obtain $\bm{w}_{\text{minimax}}$ by minimizing the maximum regret $\max_{\bm{\tau} \in \mathcal{T}} R(\bm{\tau}, \hat{\delta}_{\bm{w}})$.
If the parameter space $\mathcal{T}$ satisfies Assumption 1, we can simplify the derivation of $\bm{w}_{\text{minimax}}$. Noting the symmetry of $\mathcal{T}$ from Assumption 1, we obtain
where the last equality follows from $\sum_{k=1}^K w_k = 1$. Let $s(\bm{w}) \equiv \sqrt{\sum_{k=1}^K w_k^2 \sigma_k^2}$ denote the standard deviation of $\sum_{k=1}^K w_k \hat{\tau}_k$. Then, we have
By Assumption 1, the term $\min_{\bm{\tau} \in \mathcal{T}, \tau_0 = t} \left\{ \sum_{k=1}^K w_k (\tau_k-\tau_0) \right\}$ does not depend on $t$. Hence, we obtain
Viewing $\sum_{k=1}^K w_k \hat{\tau}_k$ as an estimator for $\tau_0$, $b(\bm{w})$ and $s(\bm{w})$ can be interpreted as the maximum bias and the standard deviation of a linear estimator $\sum_{k=1}^K w_k \hat{\tau}_k$, respectively. Using these terms, we can express the maximum regret as
where $\eta(a) \equiv \max_{t \geq 0}\left\{ t \cdot \Phi(-t+a) \right\}$. Hence, we obtain the following theorem:
In view of Theorem 1, we can compute $\bm{w}_{\text{minimax}}$ using the following algorithm:
In step 1, we compute $b(\bm{w})$ by solving the optimization in $\bm{\tau}$. As shown below, there are many examples in which we can calculate $b(\bm{w})$ using linear programming. If so, $b(\bm{w})$ can be solved quickly and reliably even when $K$ is large. Furthermore, because $t \mapsto t \cdot \Phi(-t+a)$ is a smooth unimodal function, $\eta(a)$ is easy to compute.
Figure (ref) displays the shape of $\eta(a)$ as a function of $a$. From the proof of Lemma (ref) below, we find that $\eta(a)$ is strictly increasing and convex. In numerical simulations given in Section 4 below, we compute the optimization of step 2 using the R package “Rsolnp”. We find that this optimization step is quick and stable even when $K$ exceeds 100.
Even if the parameter space $\mathcal{T}$ satisfies Assumption 1, the minimax regret can be unbounded. For example, $\mathcal{T} = \mathbb{R}^{K+1}$ satisfies Assumption 1 but the maximum regret is unbounded with $b(\bm{w}) = +\infty$ for any $\bm{w}$. This is because $\lim_{a \rightarrow \infty} \eta(a) = + \infty$.
To have the maximum regret bounded, we need to impose a restriction that the difference between $\tau_k$ and $\tau_0$ is bounded for some $k$.
Assumption 2 means that there exists some study in the pool that provides some (partially) identifying information about $\tau_0$. This condition holds for the parameter space $\mathcal{T}_{\mathrm{meta}}$ of Example 1 if $\mathcal{B}$ is compact. Similarly, $\mathcal{T}_C$ satisfies Assumption 2. Theorem 2 then applies to these cases and guarantees that the minimax regret is bounded.
In this section, we compare $\bm{w}_{\text{minimax}}$ with other ways of forming the weights. First, we consider a minimax linear estimator of $\tau_0$ in terms of the mean squared errors (MSE). It is well known that the maximum MSE of $\sum_{k=1}^K w_k \hat{\tau}_k$ can be decomposed into the variance and the squared maximum bias:
Hence, the weights of the minimax MSE estimator are
We refer to $\hat{\delta}_{\bm{w}_{\text{MSE}}}$ as the minimax MSE rule.
To compare the minimax regret and MSE rules, we focus on the analytical properties of $\eta(a)$. tetenov2012statistical shows that $\eta(a)$ is a continuous, strictly increasing function and $\eta(0) \simeq 0.17$. Furthermore, in the proof of the following lemmas, we show that $\eta(a)$ is concave. We accordingly obtain the following upper and lower bounds on $\eta(a)$:
Relying on Theorem 1 and Lemmas 1--2, the next theorem bounds the maximum regret.
Theorem 3 provides lower and upper bounds on the maximum regret. These bounds show that the maximum regret is bounded from above and from below by $b(\bm{w}) + s(\bm{w})$ and $\sqrt{b^2(\bm{w}) + s^2(\bm{w})}$ up to some proportional factors, independently of the number of studies, $K$, and the dimension of $x_k$, $d_x$. The second set of inequalities imply that the minimax regret is equivalent to the minimax RMSE (root-MSE) up to a constant factor. In other words, minimax RMSE enables us to bound the minimax regret.
Furthermore, Theorem 1 and Lemma (ref) lead to the following comparison of the maximum regret between the minimax regret rule $\hat{\delta}_{\bm{w}_{\text{minimax}}}$ and the minimax MSE rule $\hat{\delta}_{\bm{w}_{\text{MSE}}}$.
Theorem 4 shows that the maximum regret of the minimax MSE rule is the same as the minimax regret up to a constant factor, independently of $K$ and $d_x$. Numerical simulations in Section 4 suggest that the maximum regret of $\hat{\delta}_{\bm{w}_{\text{MSE}}}$ can be about 40 percent greater than the minimax regret.
The proof of Lemma (ref) given in the Appendix shows that the minimax regret criterion places greater emphasis on the bias than on the variance compared with the minimax MSE criterion. To see this, consider the directional derivatives of the maximum regret and MSE. We fix $\bm{\theta} = (\theta_1, \cdots , \theta_K)'$ with $\sum_{k=1}^K \theta_k = 1$ and assume that $b(\bm{w})$ and $s(\bm{w})$ are directionally differentiable. We define
where $Q_{\bm{\theta}}(\bm{w})$ is the directional derivative of the maximum regret. Let $t^*(a)$ be the maximizer of $t \cdot \Phi(-t+a)$. Then, by the proof of Lemma (ref), we have
Here, the sign of $Q_{\bm{\theta}}(\bm{w})$ is determined by $$ s_{\bm{\theta}}'(\bm{w}) \left( t^*\left( \frac{b(\bm{w})}{s(\bm{w})} \right) - \frac{b(\bm{w})}{s(\bm{w})} \right) + b_{\bm{\theta}}'(\bm{w}), $$ where $t^*(a) - a$ is decreasing in $a$ as shown in the proof of Lemma (ref). Similarly, the sign of the directional derivative of the maximum MSE, $b^2(\bm{w})+s^2(\bm{w})$, is determined by $$ s_{\bm{\theta}}'(\bm{w}) \left( \frac{b(\bm{w})}{s(\bm{w})} \right)^{-1} + b_{\bm{\theta}}'(\bm{w}). $$ Suppose that $b_{\bm{\theta}}'(\bm{w}) < 0$ and $s_{\bm{\theta}}'(\bm{w}) > 0$, that is, we face the bias-variance tradeoff. Then, because numerical evaluation implies $t^*(a)-a < a^{-1}$ for $a \geq 0$, we obtain $$ s_{\bm{\theta}}'(\bm{w}) \left( t^*\left( \frac{b(\bm{w})}{s(\bm{w})} \right) - \frac{b(\bm{w})}{s(\bm{w})} \right) + b_{\bm{\theta}}'(\bm{w}) \ < \ s_{\bm{\theta}}'(\bm{w}) \left( \frac{b(\bm{w})}{s(\bm{w})} \right)^{-1} + b_{\bm{\theta}}'(\bm{w}). $$ When $\bm{w} = \bm{w}_{\mathrm{MSE}}$, the right-hand side must be zero. Hence, if $b_{\bm{\theta}}'(\bm{w}_{\mathrm{MSE}}) < 0$ and $s_{\bm{\theta}}'(\bm{w}_{\mathrm{MSE}}) > 0$, we conclude \[ Q_{\bm{\theta}}\left( \bm{w}_{\mathrm{MSE}} \right) \ < \ 0. \] That is, at the minimax MSE weights $\bm{w} = \bm{w}_{\mathrm{MSE}}$, locally perturbing the weight vector in the direction that reduces the bias and increases the variance improves the welfare regret. This implies that the minimax regret criterion places greater emphasis on the bias than on the variance compared with the minimax MSE criterion. In the numerical analysis of Section 4, we plot $\bm{w}_{minimax}$ and $\bm{w}_{\mathrm{MSE}}$ to illustrate the difference in their bias-variance balancing properties.
To illustrate the minimax regret rule among $\mathcal{D}_{\mathrm{lin}}$ and compare it with Bayes rules, consider the simple case where $K=2$ and $\mathcal{T} = \{ \bm{\tau}: |\tau_k - \tau_l| \leq C \ \text{for $k,l=0,1,2$ and $k \neq l$} \}$. Because $w_1 + w_2 = 1$, we can express $(w_1,w_2) = (w,1-w)$ and the maximum regret of $\hat{\delta}_{\bm{w}}$ as $s(w) \cdot \eta \left( b(w)/s(w) \right)$, where $b(w) = \max_{\bm{\tau} \in \mathcal{T}, \tau_0 = 0} \left\{ w \tau_1 + (1-w) \tau_2 \right\}$ and $s(w) = \sqrt{w^2 \sigma_1^2 + (1-w)^2 \sigma_2^2}$. First, we derive the analytical expression of $b(w)$. When $0 \leq w \leq 1$, we obtain $ b(w) = \max_{\bm{\tau} \in \mathcal{T}, \tau_0 = 0} \left\{ w \tau_1 + (1-w) \tau_2 \right\} = C$. For $w > 1$, we obtain $b(w) = \max_{\bm{\tau} \in \mathcal{T}, \tau_0 = 0} \left\{ w \tau_1 + (1-w) \tau_2 \right\} = C w$. Similarly, we have $b(w) = C (1-w)$ when $w < 0$. Hence, the maximum bias can be written as follows:
Next, we derive the minimax weight $w^{\ast} \in \text{arg} \min_{w} \left\{ s(w) \cdot \eta \left( b(w)/s(w) \right) \right\}$. From the proof of Lemma 2, let $t^*(a) \equiv \text{arg} \max_{t \geq 0} \left\{ t \cdot \Phi(-t+a) \right\}$ and obtain
where $t^{\ast}(a) - a$ is a strictly decreasing function. From numerical evaluation, we have $t^{\ast}(a) - a = 0$ when $a = a^{\ast} \risingdotseq 1.253$. Hence, $s \eta(C/s)$ is increasing when $s > C / a^{\ast} \risingdotseq 0.798 C$ and decreasing when $s < C / a^{\ast}$. In addition, $\sqrt{\frac{\sigma_1^2 \sigma_2^2}{\sigma_1^2 + \sigma_2^2}} \leq s(w) \leq \max\{\sigma_1, \sigma_2\}$ holds with the inequalities hold with equalities at some $0 \leq w \leq 1$. Because $b(w) > C$ for $w < 0$ or $w > 1$ and $\eta(a)$ is a strictly increasing function, the minimax weight $w^{\ast}$ satisfies the following conditions:
Therefore, if the dispersion of parameters $C$ is small compared to the standard deviations $\sigma_1$ and $\sigma_2$, the minimax weight attains the smallest variance, that is, $w^{\ast} = \frac{\sigma_2^2}{\sigma_1^2 + \sigma_2^2}$. On the other hand, if $C$ is large compared to $\sigma_1$ and $\sigma_2$, then the minimax regret criterion may favor rules with variance $s(w)$ larger than the minimum.\footnote{If we consider the randomized statistical treatment rules in ((ref)), for $C \geq a^{\ast} \sqrt{\frac{\sigma_1^2 \sigma_2^2}{\sigma_1^2 + \sigma_2^2}}$ the maximum regret is minimized when $w \in [0,1]$ and
Hence, if $C / a^{\ast} > \max\{\sigma_1, \sigma_2\}$, then $v$ must be positive to satisfy ((ref)). This implies that when the dispersion of parameters is large, the minimax regret criterion may favor randomized treatment rules over non-randomized treatment rules, and this observation is consistent with stoye2012minimax and yata2021optimal.}
To compare the minimax regret rule with Bayes rules, consider the following hierarchical Bayes model:
where $\tau_0$, $\epsilon_1$, and $\epsilon_2$ are mutually independent. Then the posterior distribution of $\tau_0$ can be written as follows:
As discussed in Remark 1, the Bayes optimal rule becomes
Note that the minimax regret weight $w^{\ast}$ and the weight of the Bayes optimal rule $w_{\text{HB}} = \frac{ \sigma_2^2 + \sigma_{\epsilon}^2}{\sigma_1^2 + \sigma_2^2 + 2 \sigma_{\epsilon}^2}$ satisfy $|w^{\ast}-1/2| > |w_{\text{HB}}-1/2|$ whenever $\sigma_{\epsilon}^2 > 0$, i.e, $w_{\text{HB}}$ shrinks $w^{\ast}$ toward $1/2$, and the degree of shrinkage is increasing in $\sigma_{\epsilon}^2$. The weights $w^{\ast}$ and $w_{\text{HB}}$ agree only in an extreme scenario of $\sigma_{\epsilon}^2 = 0$.
If we have perfect knowledge of $(\tau_1, \dots, \tau_K)$, i.e., $\sigma_k = 0$ for all $1 \leq k \leq K$, we can obtain the (true) identified set of $\tau_0$ based on the constraints on the parameter space $\mathcal{T}$. For instance, we construct the identified set of $\tau_0$ by intersecting multiple bounds for $\tau_0$, each of which is constructed by extrapolating from $\tau_k$, as considered in manski2020towards. We can then consider finding the minimax regret treatment rule given the true identified set of $\tau_0$ without any sampling uncertainty. We denote by $\delta^*_{IS}$ such a (non-randomized) minimax regret rule. As the sample size of each study increases, that is, as $\sigma_k \rightarrow 0$, should we expect the minimax regret rule $\hat{\delta}_{\bm{w}_{\text{minimax}}}$ we constructed in the previous section to converge to $\delta^*_{IS}$?
In what follows, we compare $\delta^*_{IS}$ with the limiting version of $\hat{\delta}_{\bm{w}_{\text{minimax}}}$, and show that $\hat{\delta}_{\bm{w}_{\text{minimax}}}$ does not necessarily converge to $\delta^*_{IS}$ as $\sigma_k \rightarrow 0$. We then consider an alternative class of treatment choice rules that converge to $\delta^*_{IS}$ as $\sigma_k \rightarrow 0$. These alternative treatment rules solve the minimax regret with a data-driven parameter space built upon confidence regions for $\bm{\tau}$. These rules, therefore, do not belong to the linear aggregation rules of Definition (ref). Moreover, their computation are not as simple as the linear minimax regret rule $\delta_{\bm{w}_{\text{minimax}}}$ obtained in the previous section, and we do not know if they coincide with any exact minimax regret (nonlinear) rule obtained for a data-independent parameter space. Nevertheless, we can show that such modified treatment rules converge to $\delta^*_{IS}$ as $\sigma_k \rightarrow 0$, which could be of theoretical interest.
First, we consider the minimax regret rule when $\sigma_k = 0$ for all $k = 1, \dots , K$. Then, $\tau_k = \hat{\tau}_k$ for all $k = 1, \dots , K$ and the identified set of $\tau_0$ is
In this case, the parameter space $\mathcal{T}$ projected for $\tau_0$ yields the identified set $IS_0$. Because there is no randomness in this problem, for a treatment rule $\delta$, the welfare regret of $\delta$ becomes $$ W(\delta^*)-W(\delta) \ = \
. $$ Hence, the minimax regret rule over $IS_0$ can be written as
where $\underline{\tau}_0 \equiv \inf\{\tau_0:\tau_0 \in IS_0\}$ and $\overline{\tau}_0 \equiv \sup\{\tau_0:\tau_0 \in IS_0\}$ are the smallest and largest values of the identified set of $\tau_0$, respectively. The rule $\delta_{IS}^*$ becomes 1 (or 0) when we have $\underline{\tau}_0 > 0$ (or $\overline{\tau}_0 < 0$), that is, all values of the identified set of $\tau_0$ are positive (or negative). When the identified set of $\tau_0$ contains both of positive and negative values, $\delta_{IS}^*$ becomes 1 (or 0) if the absolute value of $\overline{\tau}_0$ is larger (or smaller) than that of $\underline{\tau}_0$.\footnote{In this paper, we consider only non-randomized treatment rules. manski2011choosing shows that the minimax regret criterion always yields a randomized treatment rule when $\underline{\tau}_0 < 0 < \overline{\tau}_0$. He shows that the minimax randomized treatment rule randomly assigns a fraction $|\overline{\tau}_0|/(|\underline{\tau}_0|+|\overline{\tau}_0|)$ of the population to treatment 1 and the remaining $|\underline{\tau}_0|/(|\underline{\tau}_0|+|\overline{\tau}_0|)$ to treatment 0.}
Next, we consider the large sample properties of our minimax regret rule $\hat{\delta}_{\bm{w}_{\text{minimax}}}$, i.e., $\sigma_k \rightarrow 0$. In this case, we can show that the minimax regret criterion yields the treatment rule that minimizes the maximum bias. The proof of Lemma (ref) shows that $\eta(a)$ is strictly increasing and convex with its slope bounded from above by one. Hence, the slope of $\eta(a)$ converges to a positive constant $c \in (0,1]$ as $a \rightarrow +\infty$. This implies that when $a$ is large, $\eta(a)$ can be approximated by $d+c \cdot a$ for some $d$. As $\sigma_k \rightarrow 0$ for all $k$, we have $s(\bm{w}) \rightarrow 0$ for any $\bm{w}$. From Theorem 1, as $s(\bm{w}) \rightarrow 0$, we can approximate the maximum regret of $\hat{\delta}_{\bm{w}}$ by $c \cdot b(\bm{w})$. This implies that in large samples, the minimax regret rule becomes a treatment rule that minimizes the maximum bias $b(\bm{w})$.
To be specific, consider the case in which the parameter space is the class of Lipschitz vectors given in Example 2. Since we have $$ b(\bm{w}) \ = \ \max_{|\tau_k - \tau_l| \leq C \|x_k - x_l\|} \left\{ \sum_{k=1}^K w_k (\tau_k - \tau_0) \right\}, $$ as $\sigma_k \rightarrow 0$ for all $k$, the minimax regret rule converges to the rule that depends only on the closest study in terms of the metric on the covariate space, i.e., the weight of the closest study $w_{k^*}$ converges to $1$, where $k^*$ satisfies $\|x_{k^*}-x_0\| \leq \|x_k-x_0\|$ for all $k = 1, \cdots , K$. Hence, the minimax regret rule converges to $$ \hat{\delta}_{\bm{w}_{\text{minimax}}}(\mathbf{D}) \ = \ \mathbf{1}\{ \hat{\tau}_{k^*} \geq 0 \}, $$ and the decision of whether or not to introduce the policy is solely based on the closest study.
In this case, we can show that $\hat{\delta}_{\bm{w}_{\text{minimax}}}$ does not converge to $\delta^*_{IS}$ as $\sigma_k \rightarrow 0$. If the observed covariates are scalar, then the identified set of $\tau_0$ can be written as $$ \bigcap_{k = 1}^K \big[ \hat{\tau}_{k}-C|x_k-x_0|,\hat{\tau}_{k}+C|x_k-x_0| \big]. $$ Because our minimax regret rule $\hat{\delta}_{\bm{w}_{\text{minimax}}}$ uses only the closest study, it does not agree with $\delta^*_{IS}$ from ((ref)). In fact, it is possible that $\hat{\tau}_{k^*}$ is positive but that the absolute value of $\overline{\tau}_0$ is larger than that of $\underline{\tau}_0$.
To resolve such a disagreement, we propose a minimax treatment rule refined by a confidence region of $\bm{\tau}$. Data $\mathbf{D}$ provide some information about the parameter space $\mathcal{T}$. If there is not an a priori assumption available to constrain $\mathcal{T}$, we may want to exploit such in-sample information to refine the minimax regret rule.
For $\alpha \in (0,1)$, consider a subset $\hat{\mathcal{T}}(\alpha) \subset \mathcal{T}$ that depends on the data $\mathbf{D}$ and satisfies
where $P_{\bm{\tau}}$ is the sampling probability distribution of the data when the true parameter value is $\bm{\tau}$. $\hat{\mathcal{T}}(\alpha)$ is a confidence set for $\bm{\tau}$ with a coverage probability of at least $1-\alpha$. For example, the following hyper-rectangle satisfies condition ((ref)): \[ \hat{\mathcal{T}}_{\text{HR}}(\alpha) \ \equiv \ \left\{ \bm{\tau} \in \mathcal{T} : \tau_k \in [\hat{\tau}_k - \sigma_k \cdot z_{\alpha, K}, \hat{\tau}_k + \sigma_k \cdot z_{\alpha, K}] \ \text{for $k = 1, \cdots, K$} \right\}, \] where $z_{\alpha, K}$ is the value such that $P(|Z| \leq z_{\alpha,K}) = (1-\alpha)^{1/K}$ for a standard normal variable $Z$.
When the parameter space is $\mathcal{T}_{\text{meta}}$ and $d_x < K$, we can construct another confidence region of $\bm{\tau}$. Let $\hat{\beta}$ be the OLS estimator of $\beta$. Then, the following set satisfies the condition ((ref)): \[ \hat{\mathcal{T}}_{\text{meta}}(\alpha) \ \equiv \ \ \left\{ \bm{\tau} \in \mathcal{T}_{\text{meta}} : (\hat{\beta}-\beta)'S(\hat{\beta})^{-1}(\hat{\beta}-\beta) \leq \chi(\alpha,d_x) \right\}, \] where $S(\hat{\beta})$ is the variance matrix of $\hat{\beta}$ and $\chi(\alpha,d_x)$ is the $(1-\alpha)$-th quantile of the chi-square distribution with $d_x$ degrees of freedom.
By replacing the parameter space $\mathcal{T}$ with $\hat{\mathcal{T}}(\alpha)$, we can compute the refined minimax regret rule. We define
For the class of linear aggregation rules considered in the previous sections, $\bm{w}_{\text{minimax}}$ cannot depend on the data $\mathbf{D}$. In contrast, $\hat{\bm{w}}_{\text{minimax}}(\alpha)$ depends on the data through $\hat{\mathcal{T}}(\alpha)$. Hence, the refined minimax regret rule $\hat{\delta}_{\hat{\bm{w}}_{\text{minimax}}(\alpha)}$ becomes a non-linear aggregation rule. Because $\hat{\mathcal{T}}(\alpha)$ is contained in the parameter space $\mathcal{T}$, the refined minimax regret rule is less conservative than $\hat{\delta}_{\bm{w}_{\text{minimax}}}$. From ((ref)), for any $\bm{w}$, we obtain \[ R(\bm{\tau},\hat{\delta}_{\bm{w}}) \ \leq \ \max_{\bm{\tau} \in \hat{\mathcal{T}}(\alpha)} R(\bm{\tau},\hat{\delta}_{\bm{w}}) \ \ \text{with probability $1-\alpha$.} \] Hence, $\hat{\bm{w}}_{\text{minimax}}(\alpha)$ minimizes the worst-case regret over $\hat{\mathcal{T}}(\alpha)$, which is a valid upper bound on the true regret with probability $1-\alpha$.
Since $\hat{\mathcal{T}}(\alpha)$ may not satisfy Assumption 1, we cannot derive $\hat{\bm{w}}_{\text{minimax}}(\alpha)$ using Theorem 1. However, even if $\hat{\mathcal{T}}(\alpha)$ does not satisfy Assumption 1, we can calculate $\hat{\bm{w}}_{\text{minimax}}(\alpha)$ using ((ref)) in Remark (ref). When the parameter space is $\mathcal{T}_{\text{meta}}$, we can easily calculate the refined minimax regret rule using $\hat{\mathcal{T}}_{\text{meta}}(\alpha)$. If $(\hat{\beta}-\beta)'S(\hat{\beta})^{-1}(\hat{\beta}-\beta) \leq \chi(\alpha,d_x)$ implies $\beta \in \mathcal{B}$, then we have \[ \hat{\mathcal{T}}_{\text{meta}}(\alpha) \ \equiv \ \ \left\{ \bm{\tau}: \text{$\tau_k = \beta_0 + x_k'\beta$, $\beta_0 \in \mathbb{R}$, and $(\hat{\beta}-\beta)'S(\hat{\beta})^{-1}(\hat{\beta}-\beta) \leq \chi(\alpha,d_x)$ } \right\}. \] Then, for $t >0$, we obtain
Let $\beta^*$ be a maximizer of the above problem. Using the method of Lagrange multipliers, we find that $\beta^*$ satisfies
where $X_0(\bm{w}) \equiv \sum_{k=1}^K w_k (x_0 - x_k)$. Then, equations ((ref)) imply that
For $t < 0$, we can obtain similar results. Hence, $\tilde{b}(t,\bm{w})$ has the following closed form expression:
This result makes the computation of $\hat{\bm{w}}_{\text{minimax}}(\alpha)$ easier. In this case, because $\mathcal{S} = \mathbb{R}$, we obtain
Hence, in this case, it is not difficult to compute $\hat{\bm{w}}_{\text{minimax}}(\alpha)$.
As shown above, when $\sigma_k \rightarrow 0$ for all $k$, $\hat{\delta}_{\bm{w}_{\mathrm{minimax}}}$ does not converge to $\delta_{IS}^*$. However, we can show that the refined minimax regret rule using $\hat{\mathcal{T}}_{\mathrm{HR}}(\alpha)$ converges to $\delta_{IS}^*$. When $\sigma_k \rightarrow 0$ for all $k$, the hyper-rectangle confidence region $\hat{\mathcal{T}}_{\mathrm{HR}}(\alpha)$ projected for $\tau_0$ converges to the identified set $IS_0$ defined in ((ref)). Hence, in this case, if $\left\{ \sum_{k=1}^K w_k \hat{\tau}_k: \sum_{k=1}^K w_k = 1 \right\}$ includes both positive and negative values, that is, $\hat{\tau}_k$ is not the same for all $k$, then the refined minimax regret rule $\hat{\delta}_{\hat{\bm{w}}_{\mathrm{minimax}}(\alpha)}$ converges to $\delta_{IS}^*$.
To illustrate the results we obtain in the previous sections, we present some numerical analysis. Throughout this section, we set the study-specific covariates as equidistant grid points on $[0,1]$: $$ x_k = (k-1)/(K-1) \ \ \text{and} \ \ \sigma_k = 1 \ \ \text{for $k=1, \cdots , K$.} $$
We consider the following two parameter spaces:
where $C$ is a positive constant. For these two parameter spaces, we derive the minimax regret and minimax MSE rules and compare the maximum regrets of these two treatment rules.
In the case of the linear parameter space $\mathcal{T}_1$, one natural treatment rule is based on plugging in the OLS estimator, $$ \hat{\delta}_{\bm{w}_{\text{OLS}}} \ \equiv \ \mathbf{1} \left\{ \tilde{x}_0'\hat{b} \geq 0 \right\}, $$ where $\tilde{x}_0 \equiv (1,x_0)'$ and $\hat{b}$ is the OLS estimator of $(\beta_0, \beta)'$. Because $\hat{b}$ is linear with respect to $(\hat{\tau}_1, \dots , \hat{\tau}_K)'$, this rule can be expressed as a linear aggregation rule, $\sum_{k=1}^K w_{\text{OLS},k} \hat{\tau}_k$.
Another natural treatment rule is based on plugging in the hierarchical Bayes (HB) estimator defined in Remark (ref), $$ \hat{\delta}_{\bm{w}_{\text{HB}}} \ \equiv \ \mathbf{1} \left\{ \sum_{k=1}^K w_{\text{HB},k} \hat{\tau}_k \geq 0 \right\}, $$ where $\bm{w}_{\text{HB}} \equiv (w_{\text{HB},1}, \dots, w_{\text{HB},K})' \equiv \tilde{\bm{w}}_{\text{HB}} / \sum_{k=1}^K \tilde{w}_{\text{HB},k}$ and $\tilde{\bm{w}}_{\text{HB}} \equiv (\tilde{w}_{\text{HB},1}, \dots, \tilde{w}_{\text{HB},K})' \equiv \bm{\Sigma}^{-1} \bm{X} \left(\bm{\Sigma}_{\tilde{\beta}}^{-1} + \bm{X}' \bm{\Sigma}^{-1} \bm{X} \right)^{-1} \tilde{x}_0$. We set the prior variance matrix as \[ \bm{\Sigma}_{\tilde{\beta}} \ = \ \left(
\right). \] Because the state space $\mathcal{T}_1$ does not restrict $\beta_0$, we specify a diffuse prior for $\beta_0$. In contrast, because $\mathcal{T}_1$ assumes that $\beta \in [-C,C]$, we assume that $\beta$ is contained in $[-C,C]$ with prior probability $0.95$.
We calculate $\bm{w}_{\text{OLS}}$, $\bm{w}_{\text{HB}}$, $\bm{w}_{\text{MSE}}$, and $\bm{w}_{\text{minimax}}$ for $K = 30$ and $x_0 = 0.1$. Table 1 contains the results of this experiment for $C = 0.1$, $1.0$, and $2.0$. Table 1 shows the ratios of $b(\bm{w})$ and $s(\bm{w})$, and the ratios of the maximum regrets, that is, $$ \cfrac{\max_{\bm{\tau} \in \mathcal{T}} R(\bm{\tau},\hat{\delta}_{\bm{w}_{\text{OLS}}})}{\max_{\bm{\tau} \in \mathcal{T}} R(\bm{\tau},\hat{\delta}_{\bm{w}_{\text{minimax}}})}, \ \ \cfrac{\max_{\bm{\tau} \in \mathcal{T}} R(\bm{\tau},\hat{\delta}_{\bm{w}_{\text{HB}}})}{\max_{\bm{\tau} \in \mathcal{T}} R(\bm{\tau},\hat{\delta}_{\bm{w}_{\text{minimax}}})}, \ \ \text{and} \ \ \cfrac{\max_{\bm{\tau} \in \mathcal{T}} R(\bm{\tau},\hat{\delta}_{\bm{w}_{\text{MSE}}})}{\max_{\bm{\tau} \in \mathcal{T}} R(\bm{\tau},\hat{\delta}_{\bm{w}_{\text{minimax}}})}. $$ Because the OLS estimator is unbiased, the maximum bias of the OLS estimator is exactly zero. Hence, the ratio $b(\bm{w}_{\text{OLS}})/s(\bm{w}_{\text{OLS}})$ is exactly zero in all settings. For the HB rule, $b(\bm{w}_{\text{HB}})/s(\bm{w}_{\text{HB}})$ increases as $C$ increases. The ratio of $b(\bm{w}_{\text{minimax}})$ and $s(\bm{w}_{\text{minimax}})$ is smaller than that of $\bm{w}_{\text{MSE}}$ in all settings. This implies that the minimax regret criterion places more emphasis on the bias than the variance compared with the minimax MSE criterion. Table 1 shows that the maximum regret of the minimax MSE rule is about 40 percent greater than the minimax regret when $C=1.0$. When $C=0.1$, the maximum regrets of $\bm{w}_{\text{MSE}}$ and $\bm{w}_{\text{HB}}$ are close to the minimax regret. If $C$ is sufficiently large, $\bm{w}_{\text{minimax}}$ is almost the same as $\bm{w}_{\text{OLS}}$. Hence, when $C = 1.0$ or $2.0$, the maximum regret of $\hat{\delta}_{\bm{w}_{\text{OLS}}}$ is almost identical to the minimax regret. In contrast, when $C$ is small, $\bm{w}_{\text{minimax}}$ is quite different from $\bm{w}_{\text{OLS}}$ and the maximum regret of $\hat{\delta}_{\bm{w}_{\text{OLS}}}$ is about 30 percent greater than the minimax regret.
Next, we consider the Lipschitz parameter space $\mathcal{T}_2$. We calculate $\bm{w}_{\text{minimax}}$ and $\bm{w}_{\text{MSE}}$ for $K=30$ and $x_0 = 0.5$. Similar to $\mathcal{T}_1$, we consider the following hierarchical Bayesian model with the prior distribution, \[ \bm{\tau} \ = \ (\tau_0, \bm{\tau}_{-0})' \ \sim \ \mathcal{N} \left( \bm{0}, \bm{\Sigma}_{\bm{\tau}} \right), \] where $\bm{\tau}_{-0} \equiv (\tau_1, \dots, \tau_K)'$ and \[ \bm{\Sigma}_{\bm{\tau}} \ \equiv \ \left(
\right). \] We set the prior variance of $\tau_k$ as $10$ and the prior covariance of $\tau_k$ and $\tau_l$ as $10 \cdot \exp(-|x_k-x_l| / a)$ for some positive constant $a>0$. We choose a positive constant $a$ that satisfies $\frac{1}{K(K+1)/2} \sum_{k <l} P(|\tau_k - \tau_l| > C |x_k-x_l|) = 0.05$. Then, the posterior mean of $\tau_0$ is written as \[ E[\tau_0 | \hat{\bm{\tau}}] \ = \ \bm{\Sigma}_{\bm{\tau}, 12} \bm{\Sigma}_{\bm{\tau}, 22}^{-1} \left(\bm{\Sigma}_{\bm{\tau},22}^{-1} + \bm{\Sigma}^{-1} \right)^{-1} \bm{\Sigma}^{-1} \hat{\bm{\tau}}, \] which pins down the weights of the Bayes optimal decision rule $\bm{w}_{\text{HB}}$.
Table 2 shows the ratios of $b(\bm{w})$ and $s(\bm{w})$ and the ratios of the maximum regrets for $C = 0.1$, $1.0$, and $2.0$. The ratio of $b(\bm{w}_{\text{minimax}})$ and $s(\bm{w}_{\text{minimax}})$ is smaller than the ratios of $\bm{w}_{\text{HB}}$ and $\bm{w}_{\text{MSE}}$ in all settings. When $C$ is small, the maximum regret of the minimax MSE rule nearly attains the minimax regret. In contrast, when $C$ is large, the maximum regret of the minimax MSE rule is about 17 percent greater than the minimax regret. Similar to the minimax MSE rule, the maximum regret of the hierarchical Bayes rule nearly attains the minimax regret when $C$ is small. In addition, it is about 30 percent greater than the minimax regret when $C$ is large. Figure (ref) shows $\bm{w}_{\text{minimax}}$, $\bm{w}_{\text{MSE}}$, and $\bm{w}_{\text{HB}}$ for $C=1.0$. It shows that the minimax treatment rule is quite different from other treatment rules. The minimax MSE and Bayes criteria gives positive weights to most of the studies. In contrast, the minimax regret criterion yields weights that sharply concentrate around zero and rule out half of the studies.
\if0
\fi
We illustrate the use of our methods by means of two applications. The first application considers whether an active labor market policy should be adopted, and the second application considers whether a COVID-19 treatment should be approved.
We use the meta-database of card2017works, which contains the estimates from over 200 recent studies of active labor market programs including training, subsidized employment, and job search assistance. We focus on papers that analyze RCT data to assess the impact of job training on the employment rate. This criterion reduces the meta-sample to 14 RCT estimates ($K=14$) collected from 8 different countries: Argentina, Brazil, Colombia, Dominica, Jordan, Nicaragua, Turkey, and the United States. Table 3 lists the papers included in the meta-sample of this application.
To form a vector of study characteristics $x_k$, $k=0,1,2,\dots,K$, we use five covariates that characterize the country and the sub-population on which the RCT study was performed. These are a gender dummy (male only $=$ 0, female only $=$ 0, mixed $=$ 0.5), an age dummy (age $<$ 25 only $=$ 1, age $\geq$ 25 only $=$ 0, both $=$ 0.5), an OECD dummy, the (standardized) GDP growth rate, and the (standardized) unemployment rate in 2010. Table 4 shows the estimates, standard errors, and study characteristics in this meta-sample.
We derive the minimax regret and minimax MSE rules with the following parameter space: $$ \mathcal{T}_C \ \equiv \ \left\{ \bm{\tau} : |\tau_k - \tau_l| \leq C \|x_k-x_l\| \ \text{for $k,l = 0, 1, \cdots, K$} \right\}, $$ with a prespecified Lipschitz constant $C \geq 0$. To determine $C$, we perform leave-one-out cross-validation with the study-average welfare criterion to obtain $C = 0.025$.
We consider whether the training program should be adopted in the following three target populations:
\if0
\fi
Figures (ref)--(ref) plot $\bm{w}_{\text{minimax}}$, $\bm{w}_{\text{MSE}}$, and $\bm{w}_{\text{HB}}$ for the three different target populations. Similar to Section 4, the hierarchical Bayes rule uses the following prior: \[ \bm{\tau} \ \sim \ \mathcal{N}\left( \bm{0}, \bm{\Sigma}_{\bm{\tau}} \right), \] where $\bm{\Sigma}_{\bm{\tau}}[k,l] = \exp (-\|x_{k-1} - x_{l-1}\| / a)$ and we choose $a$ satisfying $\frac{1}{K(K+1)/2} \sum_{k < l} P(|\tau_k - \tau_l| > C \|x_{k} - x_{l}\|) = 0.05$. The horizontal axis measures the Euclidean distance between $x_k$ and $x_0$. The size of the plotted circle is proportional to the precision of the estimates, i.e., a smaller $\hat{\sigma}_k$ corresponds to a larger circle. The figures show that, overall, both $\bm{w}_{\text{minimax}}$ and $\bm{w}_{\text{MSE}}$ tend to put greater weight on those studies that are in similar in terms of their population characteristics. This tendency is more evident for the minimax regret weights $\bm{w}_{\text{minimax}}$ than for the minimax MSE weights $\bm{w}_{\text{MSE}}$.
We note that $\bm{w}_{\text{minimax}}$ differs from $\bm{w}_{\text{MSE}}$ for every target population. In all cases, the minimax regret criterion puts the most weight on the closest study. In contrast, the minimax MSE criterion can put the largest weight on a study that is not closest provided that it has a small standard error. For instance, in the case of Japan, the minimax regret weight of the closest study is more than 0.6 but the minimax MSE weight is about 0.3. These results reflect the different degrees of bias variance trade-off that the minimax regret and minimax-MSE weights aim to balance out, as discussed in Section 3.1.
Table 5 lists $\hat{\tau}_0(\bm{w}_{\text{minimax}})$, $\hat{\tau}_0(\bm{w}_{\text{MSE}})$, $\hat{\tau}_0(\bm{w}_{\text{HB}})$, the ratio of the maximum regrets, and the countries that were awarded minimax regret weights larger than $1/K = 1/14$. The table also shows that the minimax regret and minimax MSE rules select different decisions in some cases. For example, the average annual salary amongst Japanese women aged 25--29 years is approximately \$30,000; if the cost per person of adopting the policy is \$1,500 and individuals that start a new job work for one year, we could set $c_0 = 0.05$.\footnote{According to fairlie2015behind, the cost per person of implementing the program in the United States is \$1,321.} Then, the recommendation of the minimax regret criterion is to introduce the policy in Japan. However, the minimax MSE criterion does not recommend the introduction of the policy in Japan.
For all of the target populations, the maximum regret of the minimax MSE rule is more than 10 percent greater than that of the minimax regret. For Japan and the UK, the minimax regret aggregation rule puts the most weight on the estimates of the US. In contrast, for Peru, the minimax regret criterion puts most of the weight on one estimate obtained from Argentina.
We consider a drug approval decision for a COVID-19 treatment using the meta-database of randomized clinical trials provided by juul2020interventions. There is an urgent gloabl need for evidence-based treatment of COVID-19. To search for effective treatments, numerous randomized clinical trials have been conducted in different countries across different demographic groups. At the time of writing, evidence as to the efficacy of various proposed treatments is mixed with limited precision of trial estimates.
We focus on Remdesivir, an antiviral medication known to be effective against viruses in the coronavirus family, such as Middle East Respiratory Syndrome (MERS) and Severe Acute Respiratory Syndrome (SARS). The effectiveness of Remdesivir in fighting Covid-19, however, remains undetermined and is a source of much controversy due to the conflicting nature of the existing evidence; the U.S. Food and Drug Administration approved emergency use of Remdesivir for Covid-19 patients, while the World Health Organization recommends against its use.
The meta-database of juul2020interventions collates data from 33 RCT studies enrolling a total of 13,312 participants. Each study provides an estimate of the treatment effect compared with standard care or a placebo. Because we focus on the effects of Remdesivir on mortality rate, we use 6 estimates from 4 RCTs, beigel2020remdesivir, who2021repurposed, spinner2020effect, and wang2020remdesivir.
To form a vector of study characteristics, we include two covariates summarizing the average patients' characteristics in each study. They are the (standardized) mean (or median) age and the (standardized) proportion of female patients. Table 6 lists the estimates, standard errors, and study characteristics.\footnote{Studies 2a--2c report the subgroup treatment effect estimates for three age subgroups ($<$50, 50--69, $\leq$70). Because we do not have detailed information about the age of these subgroups, we suppose that the mean age of these subgroups are 45, 60, and 75, respectively.}
Similar to the previous section, we derive the minimax regret and minimax MSE rules with the following parameter space: $$ \mathcal{T}_C \ \equiv \ \left\{ \bm{\tau} : |\tau_k - \tau_l| \leq C \|x_k-x_l\| \ \text{for $k,l = 0, 1, \cdots, K$} \right\}, $$ with a prespecified Lipschitz constant $C \geq 0$. Because $K$ is small, leave-one-out cross-validation does not seem sensible. Hence, in this application, we set $C=0.01$ based on the WLS estimates. We consider hypothetical populations of interest whose characteristics range over 40--80 in terms of average age and 0.34--0.41 for the fraction of female patients.
Figure (ref) plots $\hat{\tau}_0(\bm{w}_{\text{minimax}})$ at each specified grid point in the space of characteristics of the target populations. In this figure, dark red means the value of $\hat{\tau}_0(\bm{w}_{\text{minimax}})$ is large and white means the value is small. Figure (ref) also plots covariate values of the meta-sample and the size of the plotted circle is proportional to the precision of the estimates, i.e., a smaller $\hat{\sigma}_k$ corresponds to a larger circle. This result implies that the minimax regret criterion recommends to treat Remdesivir for the age group greater than 50 years-old. In contrast, for some local populations below age 50 years-old, the minimax regret criterion recommends not to treat them with Remdesivir.
Motivated by the recently proposed paradigm of `patient-centered meta-analysis' (manski2020towards), this paper develops a method to aggregate available evidence and inform optimal treatment choice for a target population that is of interest to the planner. Building upon the framework of statistical decision theory and adopting the minimax regret criterion, we obtain a minimax regret treatment choice rule that is simple to implement in practice. The key steps of our analysis that deliver analytical and computational tractability are to constrain decision rules to the class of linear aggregation rules and to restrict the parameter space to a symmetric and invariant one (Assumption (ref)). These conditions for the parameter space are mild and hold in numerous contexts.
Several questions remain unanswered. First, when $\bm{\tau}$ is constrained to Lipschitz vectors while the Lipschitz constant $C$ is unknown, we do not know what is a theoretically justifiable data-driven way to select $C$. In the presented empirical applications, we selected $C$ by cross validation and WLS estimation without any analytical justification for this choice. Second, our framework assumes away any publication bias of published estimates despite a growing concern in the scientific community about this, and increasing interest in both how to detect, and correct for, any such bias (see, e.g., AK19). Third, other than the standard errors of the estimates, our framework does not offer any way to incorporate a measure of the credibility of reported estimates. Depending on how the data were sampled and what identifying assumptions the estimate relies on, the credibility of studies can vary greatly. How to incorporate a measure of the credibility of reported estimates beyond their standard errors remains an interesting open question. \setcounter{equation}{0}