EconBase
← Back to paper

Evidence Aggregation for Treatment Choice

The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.

105,609 characters

Evidence Aggregation for Treatment Choice




\author{Takuya Ishihara\footnote{Tohoku University, Graduate School of Economics and Management. Email: [email removed]} \and Toru Kitagawa\footnote{Brown University, Department of Economics. Email: toru\[email removed]}}

\title{\vspace{-2cm}Evidence Aggregation for Treatment Choice{\Large \thanks{
We would like to thank Tim Armstrong and participants of the 2018 Y-RISE meeting in the Cayman islands for inspiring discussion, which motivated us to initiate this project. We are grateful to Chuck Manski, Jonathan Roth, and J\"{o}rg Stoye for beneficial comments. We have also benefited from comments made by participants of the Cemmap econometrics lunch meeting, Kansai econometrics study group, and the annual meetings of the 2020 SEA, 2021 IAAE, and 2021 AMES.
Financial support from
the ESRC through the ESRC Centre for Microdata Methods and Practice (CeMMAP)
(grant number RES-589-28-0001) and the ERC (grant number 715940) is gratefully acknowledged. Takuya Ishihara gratefully acknowledges financial support from JSPS through the Overseas Challenge Program for Young Researchers (number 201980057).}}}

\date{\today}
\maketitle

\vspace{-1cm}

\begin{abstract}
Consider a planner who has limited knowledge of the policy's causal impact on a certain local population of interest due to a lack of data, but does have access to the publicized intervention studies performed for similar policies on different populations. How should the planner make use of and aggregate this existing evidence to make her policy decision? Following Manski (2020; Towards Credible Patient-Centered Meta-Analysis, \textit{Epidemiology}), we formulate the planner's problem as a statistical decision problem with a social welfare objective, and solve for an optimal aggregation rule under the minimax-regret criterion. We investigate the analytical properties, computational feasibility, and welfare regret performance of this rule. We apply the minimax regret decision rule to two settings: whether to enact an active labor market policy based on 14 randomized control trial studies; and whether to approve a drug (Remdesivir) for COVID-19 treatment using a meta-database of clinical trials.
\end{abstract}


\textbf{Keywords}: Meta-analysis, program evaluation, statistical decision theory, minimax regret.

\pagebreak


\section{Introduction}
An increasing number of policy-making authorities are  interested in making their policy decisions evidence-based. In evidence-based decision-making, it is crucial for a planner to acquire credible evidence of a policy's causal impact on the affected population. Obtaining credible evidence for better policy decision-making is, however, challenging in many contexts. For instance, although randomized control trials (RCTs) are considered to be ideal for obtaining evidence of the causal impact of a policy, conducting an RCT can be costly in terms of budget, time, or administrative resources. Moreover, ethical or legal constraints can prevent the use of an RCT in certain institutional environments. In contrast, observational data can be both more accessible and easier to collect, but the credibility of any resulting causal estimates is limited if the validity of these estimates relies on restrictive identifying assumptions. In  scenarios where the planner faces difficulties in collecting direct evidence, a practical alternative is to analyze the publicized results of intervention studies performed for similar policies on different populations. With this approach in mind, how should the planner make use of and aggregate existing evidence to reach her policy decision?

Statistical methodologies to aggregate evidence from multiple studies have been considered in the literature of meta-analysis and research synthesis. See, for instance, \citet{HO85} and a recent handbook volume, \citet{CooperHandbook2019}.  Since the seminal works of \citet{Rubin81} and \citet{DL86}, a common approach to aggregation of evidence is the hierarchical Bayesian approach, in which the typical objective of analysis is to infer hyper-parameters indexing the \textit{population of studies}. This framework of meta-analysis is useful for ``summarizing what has been learned and quantifying how results differ across the studies beyond the sampling error'' (\citet{DL15}). Its use is, however, limited when it comes to the planner's policy choice because the output of meta-analysis mainly concerns the population of studies rather than the particular population that is of interest to the planner. This point is made in \citet{manski2020towards}:

\begin{quote}
\textit{``Clinicians need to assess risks and choose treatments for populations
of patients, not population of studies. To express this
distinction succinctly, I will say that clinicians should want meta-analysis
to be patient-centered rather than study-centered.''}
\end{quote}

We pursue this paradigm of `patient-centered meta-analysis' to develop a method to aggregate existing studies for the purpose of making an optimal treatment decision on the local population that is of interest to the planner (hereafter, the target population). Building on the framework of statistical treatment choice proposed by \citet{Manski2000, manski2004statistical}, we formulate the planner's problem as a statistical decision problem \`{a} la \citet{Wald50}. The basic formulation of the decision problem analyzed in this paper is as follows. Let $\tau_0$ be the average welfare effect of introducing a new policy to the target population. There is no data from which the planner can directly infer $\tau_0$, but she does have access to the results of existing intervention or observational studies that are indexed by $k=1,  2, \dots, K$, $K \geq 1$. Each study $k$ reports a point estimate $\hat{\tau}_k$  for the average welfare effect $\tau_k$ in the study population, and an associated estimate $\hat{\sigma}_k$ of the standard error. We allow the study population to be different from the target population, so that the average welfare effects can differ, i.e., $\tau_k \neq \tau_{k'}$ for $k \neq k'$, $0 \leq k, k' \leq K$.
The planner's decision problem, which we solve in this paper, is  whether or not to adopt the new policy for the target population upon observing a meta-sample, $(\hat{\tau}_k, \hat{\sigma}_k)$, $k=1, \dots, K$. That is, the statistical treatment choice rule we consider in this paper is a function $\hat{\delta}$ that maps the meta-sample to the binary choice of whether to adopt the policy or not.

Following \citet{manski2004statistical, Manski2007}, \citet{stoye2009minimax, stoye2012minimax}, and \citet{tetenov2012statistical}, we apply the minimax regret criterion of \citet{Savage51} to obtain a minimax-regret treatment choice rule for the planner. We assume that the planner's objective function (social welfare function) is linear in $\tau_0$ and consider the class of \textit{non-randomized} statistical treatment choice rules that select the treatment based on the sign of linear aggregation of $(\hat{\tau}_k : k = 1, \dots, K)$:
\begin{equation*}
\left\{ \hat{\delta}_{\bm{w}} = 1\left\{ \sum_{k=1}^K w_k \hat{\tau}_k \geq 0 \right\}:\sum_{k=1}^K w_k=1\right\},
\end{equation*}
where $\bm{w}=(w_1, \dots, w_K)'$ is a vector of weights assigned to each estimate in the pool of studies, which does not depend on the data.
Restricting the feasible rules to non-randomized (non-fractional) ones can be attractive in the following contexts. First, in the real-world practice of treatment choice or drug approval decisions, non-randomized allocation of treatments are easier than randomized ones for policy authorities to administer. Second, non-fractional allocations are guaranteed to attain parity within the target population since either everyone or no one is treated. On the other hand, fractional rules do not attain the ex-post parity. The restriction to non-randomized rules, however, sacrifice the value of minimax regret since unconstrained minimax regret rules are known to be randomized for some realization of data as shown by \cite{stoye2012minimax, yata2021optimal}, and \cite{montiel2023decision}.

Assuming a Gaussian sampling distribution for $(\hat{\tau}_1, \dots, \hat{\tau}_K)$ with known variances and imposing certain symmetry and invariance conditions on the parameter space for $(\tau_0, \tau_1, \dots, \tau_K)$, we derive the aggregation weights   $\bm{w}_{\text{minimax}}$ leading to a minimax-regret treatment choice rule among the non-randomized rules.
Analytical characterization and computation of the exact minimax regret rule often become challenging in the context of statistical treatment choice.
Our approach to the planner's minimax regret aggregation rule, in contrast, overcomes these challenges by showing that some mild restrictions on the parameter space and the class of decision rules deliver analytically and computationally tractable minimax regret rules.

We assert that the perspective and tools of statistical decision theory are particularly appealing in the meta-analysis setting for the following reasons. First, if each study in the pool reports a consistent estimate using a sample of moderate to large size (e.g., the difference-in-means estimator for the average treatment effect) then, by its asymptotic normality, it is plausible to assume that $\hat{\tau}_k$ follows a Gaussian distribution centered at $\tau_k$. Hence, the standard and well-studied framework of Gaussian experiments fits well to the current meta-analysis setting. Second, it is common for the meta-sample to consist of only a small number of studies. In such instances, asymptotic analysis with $K \to \infty$ can be misleading, and deriving finite-$K$ optimal procedures, which statistical decision theory is particularly suitable for, is desirable.

As an alternative to the minimax regret treatment choice rule, one could consider using a \textit{plug-in} rule that chooses the treatment according to the sign of an estimate of $\tau_0$. The plug-in rule that uses a minimax mean squared error (MSE) optimal estimate of $\tau_0$ is an example. Minimax-MSE estimation for finite-dimensional Gaussian mean models is well-studied and the minimax-MSE weights $\bm{w}_{\text{MSE}}$ are simple to compute, although the resulting plug-in decision rule  does not generally possess decision theoretic optimality in terms of the planner's objective function. To quantify the welfare cost of $\hat{\delta}_{\bm{w}_{\text{MSE}}}$, we compare the worst-case regrets of $\hat{\delta}_{\bm{w}_{\text{minimax}}}$ and $\hat{\delta}_{\bm{w}_{\text{MSE}}}$, and show that the worst-case regret of $\hat{\delta}_{\bm{w}_{\text{MSE}}}$ is worse than the minimax regret only up to a constant factor of 5.88, independent of the number of studies $K$ and the parameter space.

Our framework can accommodate a vector of observable characteristics $x_k$, $k=0,1,\dots,K$, where $x_k$ includes the characteristics of the treatment and demographics of the population featured in study $k$. Under a linear functional form specification, $\tau_k = \beta_0 + x_k'\beta$, $\beta \in \mathcal{B}$, common to standard meta-regression analysis  (see, e.g., \citet{SJ89}), we discuss those restrictions on $\mathcal{B}$ under which we can apply our minimax regret decision rule. For minimax regret to be bounded, an important constraint is boundedness of $\mathcal{B}$, and the bounds of $\mathcal{B}$ have to be explicitly specified to obtain the minimax regret rule. In reality, the planner may not be able to come up with reasonable bounds for $\mathcal{B}$. To offer a practical solution to this difficulty, we consider a data-driven way to specify the parameter space based on confidence sets for $(\tau_1, \dots, \tau_K)$.

We illustrate the use of our minimax regret treatment rule by using two empirical examples. In the first application, we analyze whether an active labor market program should be adopted using the meta-database appearing in \citet{card2017works}. We consider a pool of 14 RCT studies of job training programs, covering 8 different countries (Argentina, Brazil, Colombia, Dominica, Jordan, Nicaragua, Sri Lanka, Turkey, and the United States). Based on the average treatment effect and standard error estimates in each of these studies, and the demographic characteristics of the studied populations, we calculate the minimax regret adoption decisions for several countries (Japan, the United Kingdom, and Peru) for which the corresponding experimental estimates are not available in the meta-database.

In the second application, we consider the drug approval decision for a COVID-19 medication called Remdesivir. Remdesivir is an antiviral medication that is known to be effective against Middle East Respiratory Syndrome (MERS) and Severe Acute Respiratory Syndrome (SARS), while its effectiveness against COVID-19 remains unknown due to conflicting evidence. Using the meta-database of randomized clinical trials for COVID-19 treatments provided by \citet{juul2020interventions}, we calculate the minimax regret treatment choice for Remdesivir for some specified demographic groups.

The remainder of the paper is organized as follows. The next subsection reviews the related literature. Section 2 formulates the minimax regret decision problem and shows the main analytical result of the paper. In Section 3, we compare the minimax regret with the maximum regret of the decision rule based on the minimax-MSE aggregation rule. Section 3 also discusses a data-driven construction of the parameter space. Section 4 performs numerical analysis to compare the minimax-regret aggregation rules with the minimax-MSE and meta-OLS rules.

\subsection{Related Literature}
This paper contributes to the growing literature on statistical treatment choice and individualized treatment assignment rules initiated by \citet{Manski2000, manski2004statistical} and \citet{Dehejia2005}. Contributions to the current literature include, \citet{HiranoPorter2009}, \citet{stoye2009minimax, stoye2012minimax}, \citet{Chamberlain2011}, \citet{BhattacharyaDupas2012}, \citet{tetenov2012statistical}, \citet{Kasy2014b,Kasy2018}, \citet{kitagawa2018should, KT19}, \citet{KW20}, \citet{Russell20}, \citet{KST21}, \citet{MT17}, \citet{AW20}, \citet{Sakaguchi21}, and \citet{Viviano21}, among others.
The problem of individualized treatment assignment rules has also been an area of active research in the fields of medical statistics and machine learning; see, for instance, \citet{Zadrozny03}, \citet{BeygelzimerLangford09} \citet{QianMurphy2011}, \citet{Zhao2012JASA}, \citet{SJ15}, \citet{Kallus_2020}, to list but a few papers. The standard setting in the existing literature considers an optimal treatment assignment policy for the population from which a sample was drawn, rather than combining the pool of estimates from multiple studies performed on different populations.

There is a growing literature on how to inform policy using multiple pieces of evidence or extrapolation from one or multiple reference populations. \citet{Dehejiaetal21} considers the use of (quasi-)experimental evidence to study the decision of whether to experiment or to extrapolate, and, if applicable, where to conduct a new experiment. \citet{Manski18} analyzes decision-making for personalized risk assessment under the ecological inference setting where (partial) identification of a long regression is obtained by combining information on a short regression and the joint distribution among the regressors. The meta-analysis setting considered in this paper differs from the ecological inference setting in terms of the object to identify and the type of information provided by the available studies. Focusing on conditional cash transfer programs, \citet{Gechteretal19} runs multiple program evaluation methods on data obtained from Mexico to inform treatment assignment policies for Morocco, and empirically compare the welfare performances of these policies. \citet{Hotzetal05} and \citet{Dehejiaetal21} analyze how to predict the effects of future programs from past experimental evaluations by adjusting for differences in the distributions of observable characteristics. \citet{AO19} propose a method to conduct sensitivity analysis and to approximate external validity bias when the trial and target populations differ in the distribution of unobservables. \citet{Gechter16} considers bounding causal effects in a target population by restricting the dependence between the treated and control outcomes.


Meta-analysis for research synthesis has been actively studied in statistics and the resulting literature is vast; see, e.g., \citet{Borenstein09textbook} for a textbook and \citet{CooperHandbook2019} for a handbook volume.
In economics, existing applications of meta-analysis and meta-regression include \citet{CK95}, \citet{Dehejia03}, \citet{Bandieraetal17}, \citet{card2017works}, \citet{Meager19, Meager20}, \citet{Imai20}, and \citet{Vivalt20}. See \citet{Stanley01} for a review.
The common framework of meta-analysis introduces the population of studies and draws inference for the parameters thereof. As we argue in the Introduction via the quote from \citet{manski2020towards}, the usefulness of the conventional framework of meta-analsyis is not obvious for informing the planner's policy decision.
This paper follows and pushes forward the perspective of patient-centered meta-analysis.
The methodological proposals in \citet{manski2020towards} concern predicting treatment effects for the target population by intersecting the population identified sets for $\tau_0$ formed by extrapolation from each study, rather than explicitly taking into account sampling uncertainty due to finite sample size when considering the treatment choice decision.

In terms of the framework and analytical and computational challenges for obtaining minimax regret rules, this paper is most closely related to \citet{stoye2012minimax}.
In one of his baseline settings, \citet{stoye2012minimax} considers Gaussian experiments for conditional average treatment effects with a scalar covariate $x \in \mathcal{X}$, and analyzes the properties of the minimax regret treatment rule under a restriction that the conditional average treatment effects depend on $x$ with bounded variation.
Similar to \citet{stoye2012minimax}, our framework allows (study-specific) covariates to constrain the parameter space for $(\tau_0, \tau_1, \dots, \tau_K)$, but there are two aspects in which our framework differs.
First, the treatment assignment rules considered in \citet{stoye2012minimax} are functions $\delta: \mathcal{X} \to \{0 ,1\}$, while we concern ourselves with the treatment choice at a particular covariate value $x_0$ in $\mathcal{X}$ that corresponds to the covariate value of the target population.
This reduction of the treatment choice rule from a function of $x$ to a point significantly simplifies the analysis and computation of the minimax-regret rule.
Second, the conditions we impose on the parameter space for feasible computation of the minimax regret rule is general and includes the bounded variation restriction considered in \citet{stoye2012minimax} as a special case.

In another baseline setting of \citet{stoye2012minimax}, he derives a minimax regret treatment choice rule under partially identified welfare, where the identified set has the known width but unknown location, and a Gaussian signal is available for it. Our setting is more complex than his setting due to multiple Gaussian signals with both the location and width of the identified set unknown. Contemporaneously or after the initial version of the current paper was circulated, there have been some important advances in the literature. In an similar setting to this paper, \citet{yata2021optimal} obtained an analytical representation for an exact unconstrained minimax regret rule, which implements an fractional assignment when the identified set for ATE is wide relative to the variances of the bound estimates. \citet{montiel2023decision} shows that minimax regret rules are not unique and considers refining the set of minimax regret rules. An alternative approach to handle uncertainty in the bound estimates and ambiguity within the identified set is to introduce multiple priors as in \citet{giacomini2021} and performs a gamma minimax decision rule. See \citet{giacomini_etal2021} for gamma minimax decision rules for set-identified models. \citet{Christensen2022} studies this line of approach for treatment choice and shows its optimality properties in limit experiments.

Viewing $\tau_0$ as the value of the regression equation at $x_0$ in a Gaussian regression model and considering standard estimation risk such as the mean squared errors for $\tau_0$, the problem is reduced to an interpolation or extrapolation exercise based on the Gaussian signals. As such, the minimax estimation and inference problem for $\tau_0$ is similar to the extrapolation issue in the regression discontinuity setting analyzed in \citet{KR18}. Recent contributions regarding estimation and inference for Gaussian sequence models are made by \citet{Johnstone17} and \citet{AK18, AK20}. These papers consider statistical losses for estimation and inference, but do not consider the welfare criterion for statistical treatment choice.

\section{Minimax regret treatment rule}

\subsection{Setting}

Suppose we have access to the publicized results of $K$ studies indexed by $k=1, \dots, K$. Each of the studies estimates the causal effect of a particular binary policy or treatment. We allow the details of the policy (implementation protocol, dosage, program contents, etc.) to differ across the studies. For $k = 1, \cdots, K$, let $\hat{\tau}_k$ denote the estimate of the policy effect reported in study $k$ and $\sigma_k$ denote the standard error of $\hat{\tau}_k$. For simplicity, we assume $\sigma_k$ is known, although, in practice, we can only construct a consistent estimator for $\sigma_k$. We solve for the finite sample minimax regret rule with known $(\sigma_k : k=1,\dots,K)$, recommending that in practice the rule is implemented with the true standard errors replaced by their consistent estimates.\footnote{Solving the decision problem with Gaussian signals with known variances and obtaining a feasible decision rule by plugging in consistent estimators for the variances are similar to the construction of an asymptotically optimal decision rule within the framework of Gaussian limit experiments. See \citet{HiranoPorter2020} for a recent review.}

We assume
\begin{equation}
\hat{\tau}_k \ \sim \ \mathcal{N}(\tau_k,\sigma_k^2), \ \ \ k = 1, \dots, K, \label{tau_hat}
\end{equation}
where $\tau_k$ is the true policy effect of the population featured in study $k$. We allow $\tau_k$ to vary across the studies. The assumption that $\hat{\tau}_k$ follows a Gaussian sampling distribution for is reasonable if the reported estimator $\hat{\tau}_k$ is consistent and asymptotically normal and each study has a moderate to large sample.

Throughout this paper, we consider a planner who, upon observing data $\mathbf{D} \equiv \{\hat{\tau}_k\}_{k=1}^K$,  must determine whether or not to adopt the policy in the target population given that its policy effect $\tau_0$ is unknown. Following \cite{manski2004statistical, Manski2007}, \citet{stoye2009minimax, stoye2012minimax}, and \citet{tetenov2012statistical}, we focus on minimax regret criterion to solve this decision problem. To this end, we assume that the true parameters $\bm{\tau} \equiv (\tau_0, \tau_1, \dots , \tau_K)'$ are \textit{ex ante} known to belong to the parameter space $\mathcal{T}$.

We impose the following restrictions on the parameter space $\mathcal{T}$:

\begin{Assumption} \label{parameter_space}
The parameter space $\mathcal{T}$ satisfies
\begin{enumerate}
\item Symmetry: $\bm{\tau} \in \mathcal{T} \ \Rightarrow \ -\bm{\tau} \in \mathcal{T},$ and
\item Invariance to common constant addition: $\bm{\tau} \in \mathcal{T} \ \Rightarrow \ \bm{\tau} + c \cdot (1, \dots , 1)' \in \mathcal{T}$ for any $c \in \mathbb{R}$.
\end{enumerate}
\end{Assumption}

The symmetry assumption rules out the imposition of a sign restriction on the causal effect parameters, i.e., $\tau_k \geq 0$ for some $k$.
The condition of invariance to common constant addition (hereafter, shortened to invariance) implies that $\{(\tau_k - \tau_0) : \bm{\tau} \in \mathcal{T}, \tau_0 = t\}$ does not depend on $t$.
We use this result to simplify derivation of a minimax regret treatment rule.
It is worth noting that the invariance condition rules out the case in which the parameter space for some $\tau_k$, $k \in \{0,1,2, \dots, K\}$, is bounded.
For instance, if the outcome is binary, the treatment effect on the outcome is bounded by $[-1,1]$ for all $\tau_k$, $k=0,1, \dots, K$.\footnote{In a similar setting to the one in this paper, \cite{ishihara2023bandwidth} studies the treatment choice problem with $\tau_k$, $k \in \{0,1,2, \dots, K\}$ being bounded due to binary outcomes.}
However, if the standard errors $(\sigma_k^2:k=1,2, \dots, K)$ and the variations among $(\tau_0, \tau_1, \dots, \tau_K)$  imposed on $\mathcal{T}$  (e.g., Lipschitz constants $C_{kl}$, $ 0 \leq k,l \leq K$, in Example \ref{ex:constant variations} below)  are sufficiently small relative to the size of the supports, bounded support of $(\tau_0, \tau_1, \dots, \tau_K)$ is less of an issue because extreme values of $\bm{\tau}$ beyond its logical support are unlikely to correspond to a worst-case in terms of the regret.

The following two examples satisfy the parameter space constraints of Assumption \ref{parameter_space}.

\begin{Example}[The space of $\bm{\tau}$ spanned by the meta-regressions] \label{Ex:meta_regression}

Suppose that for each study in the pool, $k=1, \dots, K$, we can construct a vector of study-specific observable characteristics $x_k \in \mathbb{R}^{d_x}$. For example, $x_k$ can contain the average characteristics of the individuals in the sample studied in the $k$-th study. It can also include the socioeconomic or demographic characteristics of the country that the sample was drawn from and the characteristics of the treatment studied. The target population has known covariate value $x_0 \in \mathbb{R}^{d_x}$, which shares the same interpretation and dimension as $x_k$, $k=1, \dots, K$.

In meta-regression analysis, $\tau_k$ is often specified as
$$
\tau_k = \beta_0 + x_k' \beta.
$$
Accordingly, we assume that the parameter space can be written as
\begin{equation}
\mathcal{T}_{\mathrm{meta}} \equiv \left\{ \bm{\tau} = (\tau_0, \tau_1, \dots, \tau_K) : \ \text{$\tau_k = \beta_0 + x_k' \beta$, $\beta_0 \in \mathbb{R}$, and $\beta \in \mathcal{B}$ for $ 0 \leq k \leq K$.} \right\}, \label{T_meta}
\end{equation}
where $\mathcal{B}$ is a compact subset of $\mathbb{R}^{d_x}$. As seen in Theorem 2, compactness of $\mathcal{B}$ implies that minimax regret is finite. If $\mathcal{B}$ satisfies $\beta \in \mathcal{B} \Rightarrow -\beta \in \mathcal{B}$, then $\mathcal{T}_{\mathrm{meta}}$ satisfies Assumption \ref{parameter_space}. We can allow study-specific intercepts without violating Assumption \ref{parameter_space} by viewing them as study-specific fixed dummy variables added to the covariate vector.
\end{Example}

\begin{Example}[The class of constant variations] \label{ex:constant variations}
Consider the following parameter space:
\begin{equation}
\mathcal{T}_C \equiv \left\{ \bm{\tau} : |\tau_k - \tau_l| \leq C_{kl}  \ \ \text{for $k, l = 0, 1, \cdots , K$.} \right\}, \label{T_C}
\end{equation}
where $\{C_{kl} : k, l = 0, 1, \cdots , K \}$ is a set of known positive constants. With study-specific covariates as introduced in Example \ref{Ex:meta_regression}, setting $C_{kl} = C \|x_k - x_l\|$, $C>0$, yields the class of Lipschitz vectors of $\bm{\tau}$. Clearly, for any $\{C_{kl} : k, l = 0, 1, \cdots , K \}$, the parameter space $\mathcal{T}_C$ satisfies Assumption \ref{parameter_space}. Assumption 1 in \citet{stoye2012minimax} corresponds to the case where  $C_{kl}$
is a common constant for any $0 \leq k,l \leq K$.
\end{Example}

When Assumption \ref{parameter_space} does not hold in a given application, we can still formulate the optimization problem to derive the minimax regret treatment rule, although solving for this is accompanied by a substantial increase in computational complexity. In Remark \ref{rem:without_A1} below, we discuss a derivation of the minimax regret treatment rule that does not rely upon Assumption \ref{parameter_space}.


\subsection{Welfare and regret}

Given a non-randomized treatment choice action $\delta \in \{0,1\}$, define the welfare attained by $\delta$ as
\begin{equation}
W(\delta) \ \equiv \  (\tau_0-c_0) \cdot \delta + \mu_0, \label{welfare}
\end{equation}
where $c_0$ is the per-person cost of the policy and $\mu_0$ is the average outcome that would be realized in the absence of the policy. An optimal treatment choice action given knowledge of $\tau_0$ and $c_0$ is
\begin{equation}
\delta^* \ \equiv \ \mathbf{1}\{ \tau_0 \geq c_0 \}. \label{optimal treatment rule}
\end{equation}

Let $\hat{\delta}(\mathbf{D}) \in \{0,1\}$ be a non-randomized statistical treatment rule that maps the meta-sample $\mathbf{D}$ to the binary decision of treatment choice in the target population. The welfare regret of $\hat{\delta}(\mathbf{D})$ is defined as
\begin{equation}
R(\bm{\tau} , \hat{\delta})  \equiv  E_{\bm{\tau}}\left[ W(\delta^*)-W(\hat{\delta}(\mathbf{D})) \right] =  (\tau_0-c_0) \left\{ \delta^* - E_{\bm{\tau}}[\hat{\delta}(\mathbf{D})] \right\}, \label{regret}
\end{equation}
where $E_{\bm{\tau}}(\cdot)$ is the expectation with respect to the sampling distribution of $\mathbf{D}$ given the parameters $\bm{\tau}$.
Hereafter, we normalize the cost of treatment to $c_0=0$, i.e., interpret $\tau_k$, $k=0, \dots, K$, as the average treatment effect net of the per-person treatment cost in the target population.

The minimax regret criterion selects a statistical treatment rule that minimizes maximum regret:
\begin{equation}
\hat{\delta}_{\text{minimax}} \ \in \ \text{arg} \min_{\hat{\delta} \in \mathcal{D}} \max_{\bm{\tau} \in \mathcal{T}} R(\bm{\tau}, \hat{\delta}), \nonumber
\end{equation}
where $\mathcal{D}$ is a class of statistical treatment rules. We refer to $\hat{\delta}_{\text{minimax}}$ as a minimax regret rule. In the next subsection, we derive a minimax regret rule under the class of statistical treatment rules spanned by linear aggregation rules.

\begin{Remark} \label{rem:Bayes}
A popular approach in the meta-analysis literature is to perform hierarchical Bayesian inference that yields a posterior distribution for each parameter in $\bm{\tau}$ and its hyperparameters. Once uncertainty for $\bm{\tau}$ is summarized by the posterior distribution, we can show that the Bayes optimal decision rule is determined by the posterior mean of $\tau_0$. For any prior $\pi$, the Bayes optimal decision rule $\hat{\delta}_{\pi}$ is defined as
$$
\hat{\delta}_{\pi} \ \in \ \mathrm{arg} \min_{\hat{\delta}} \left\{ \int R(\bm{\tau},\hat{\delta})\,\mathrm{d}\pi(\bm{\tau}) \right\}.
$$
We observe that
\begin{eqnarray}
\int R(\bm{\tau},\hat{\delta}) \,\mathrm{d}\pi(\bm{\tau}) &=& \int \tau_0 \left\{ \mathbf{1}\{\tau_0 \geq 0\} - E_{\tau}[\hat{\delta}(\mathbf{D})] \right\} \,\mathrm{d}\pi(\bm{\tau}) \nonumber \\
&=& E_{\mathbf{D}}\left[ E_{\pi}\left( \tau_0 \left\{ \mathbf{1}\{\tau_0 \geq 0\} - \hat{\delta}(\mathbf{D}) \right\} \Big| \mathbf{D} \right) \right] \nonumber \\
&=& E_{\mathbf{D}}\left[ E_{\pi}\left( \tau_0 \mathbf{1}\{\tau_0 \geq 0\}| \mathbf{D} \right) - E_{\pi}\left( \tau_0 | \mathbf{D} \right) \hat{\delta}(\mathbf{D})  \right], \nonumber
\end{eqnarray}
where $E_{\mathbf{D}}(\cdot)$ denotes the expectation with respect to the marginal distribution of $\mathbf{D}$ and $E_{\pi}(\cdot | \mathbf{D})$ denotes the posterior mean. This implies that
$$
\hat{\delta}_{\pi}(\mathbf{D}) \ = \ \mathbf{1}\left\{ E_{\pi}\left( \tau_0 | \mathbf{D} \right) \geq 0 \right\}.
$$

In Example 1, typical hierarchical Bayes approaches assume $\tilde{\beta} \equiv (\beta_0,\beta')' \sim \mathcal{N} \left( \bm{0},\bm{\Sigma}_{\tilde{\beta}} \right)$. Then, the posterior mean of $\tau_0 \equiv \beta_0 + x_0' \beta$ can be written as
\[
E[\tau_0 | \mathbf{D}] \ = \ \tilde{x}_0' \left(\bm{\Sigma}_{\tilde{\beta}}^{-1} + \bm{X}' \bm{\Sigma}^{-1} \bm{X} \right)^{-1} \bm{X}' \bm{\Sigma}^{-1} \hat{\bm{\tau}},
\]
where $\tilde{x}_k \equiv (1,x_k')'$, $\bm{X} \equiv (\tilde{x}_1, \dots, \tilde{x}_{K})'$, $\bm{\Sigma} \equiv diag \left( \sigma_1^2, \dots, \sigma_K^2 \right)$, and $\hat{\bm{\tau}} \equiv (\hat{\tau}_1, \dots, \hat{\tau}_K)'$. Hence, in this case, the Bayes optimal decision rule is a linear aggregation rule in the sense of Definition \ref{def:linear_rule} below. The weights of the Bayes optimal rule depend on the hyperparameters of the Gaussian prior for $\tilde{\beta}$ and the matrix of study characteristics$, \bm{X}$. If the mean of the prior distribution is not zero, then the posterior mean of $\tau_0$ involves a constant. In such a case, the Bayes optimal decision rule belongs to the class of extended linear aggregation rules considered in Remark \ref{rem:construle} below.
\end{Remark}

\subsection{The minimax regret rule among linear aggregation rules}

To gain analytical and computational tractability, we focus on the class of linear aggregation rules.

\begin{Definition}[Linear aggregation rules]\label{def:linear_rule}
The class of linear aggregation rules consists of non-randomized treatment choice rules, each of which chooses a treatment according to the sign of a linear aggregation of $(\hat{\tau}_1, \dots, \hat{\tau}_K)$:
\begin{equation}
\mathcal{D}_{\mathrm{lin}} \equiv \left\{ \hat{\delta}_{\bm{w}}(\mathbf{D}) = \mathbf{1}\left\{ \sum_{k=1}^K w_k \hat{\tau}_k \geq 0 \right\} : \sum_{k=1}^K w_k = 1 \right\}, \label{linear_aggregation_rule}
\end{equation}
where $\bm{w} = (w_1, \cdots, w_K)'$ does not depend on the data $\mathbf{D}$.
\end{Definition}

This class rules out nonlinear treatment rules that plug in an aggregation of $(\hat{\tau}_1, \dots, \hat{\tau}_K)$ in the manner of James-Stein shrinkage or empirical Bayes.
Nevertheless, the class of linear aggregation rules contains many reasonable treatment rules. For example, plug-in rules $\mathbf{1} \left\{ \hat{\tau}_0(\mathbf{D}) \geq 0 \right\}$ based on linear estimators $\hat{\tau}_0(\mathbf{D})$ for $\tau_0$ belong to $\mathcal{D}_{\text{lin}}$.
When study characteristics are included in the available covariates as in Example \ref{Ex:meta_regression}, $\mathcal{D}_{\text{lin}}$ includes those rules that plug in fitted values based on parametric linear regression or nonparametric kernel regression.
Furthermore, as shown in Remark \ref{rem:Bayes} above, hierarchical Bayes decision rules under the linear meta-regression specification with Gaussian priors yields the linear aggregation rule.

We consider the minimax regret rule among $\mathcal{D}_{\mathrm{lin}}$ whose corresponding weight vector solves
\[
\bm{w}_{\text{minimax}}\ \in \ \text{arg} \min_{\bm{w}} \max_{\bm{\tau} \in \mathcal{T}} R(\bm{\tau}, \hat{\delta}_{\bm{w}}).
\]
To develop a computation method for $\bm{w}_{\text{minimax}}$, note from (\ref{tau_hat}) that
$$
\sum_{k=1}^K w_k \hat{\tau}_k \ \sim \ \mathcal{N} \left( \sum_{k=1}^K w_k \tau_k, \sum_{k=1}^K w_k^2 \sigma_k^2 \right).
$$
Hence, from (\ref{regret}), the regret of $\hat{\delta}_{\bm{w}}$ can be written as
\begin{eqnarray}
R(\bm{\tau}, \hat{\delta}_{\bm{w}}) &=& \tau_0 \cdot \left[ \mathbf{1}\{\tau_0 \geq 0\} - \Phi\left( \frac{\sum_{k=1}^K w_k \tau_k}{\sqrt{\sum_{k=1}^K w_k^2 \sigma_k^2}}\right) \right] \nonumber \\
&=& | \tau_0 | \cdot \Phi \left( - sgn(\tau_0) \cdot \frac{\sum_{k=1}^K w_k \tau_k}{\sqrt{\sum_{k=1}^K w_k^2 \sigma_k^2}}\right), \label{linear regret}
\end{eqnarray}
where the first equality follows from the normality of $\sum_{k=1}^K w_k \hat{\tau}_k$ and the second equality follows from $1-\Phi(a) = \Phi(-a)$. We then obtain $\bm{w}_{\text{minimax}}$ by minimizing the maximum regret $\max_{\bm{\tau} \in \mathcal{T}} R(\bm{\tau}, \hat{\delta}_{\bm{w}})$.

If the parameter space $\mathcal{T}$ satisfies Assumption 1, we can simplify the derivation of $\bm{w}_{\text{minimax}}$. Noting the symmetry of $\mathcal{T}$ from Assumption 1, we obtain
\begin{eqnarray}
\max_{\bm{\tau} \in \mathcal{T}} R(\bm{\tau},\hat{\delta}_{\bm{w}}) &=& \max_{\bm{\tau} \in \mathcal{T}} \left\{ | \tau_0 | \cdot \Phi \left( - sgn(\tau_0) \cdot \frac{\sum_{k=1}^K w_k \tau_k}{\sqrt{\sum_{k=1}^K w_k^2 \sigma_k^2}}\right) \right\} \nonumber \\
&=& \max_{t \geq 0} \max_{\bm{\tau} \in \mathcal{T}, \tau_0 = t} \left\{ t \cdot \Phi \left( -  \frac{\sum_{k=1}^K w_k \tau_k}{\sqrt{\sum_{k=1}^K w_k^2 \sigma_k^2}}\right) \right\} \nonumber \\
&=& \max_{t \geq 0} \max_{\bm{\tau} \in \mathcal{T}, \tau_0 = t} \left\{ t \cdot \Phi \left( -  \frac{t+\sum_{k=1}^K w_k (\tau_k-t)}{\sqrt{\sum_{k=1}^K w_k^2 \sigma_k^2}}\right) \right\}, \nonumber
\end{eqnarray}
where the last equality follows from $\sum_{k=1}^K w_k = 1$. Let $s(\bm{w}) \equiv \sqrt{\sum_{k=1}^K w_k^2 \sigma_k^2}$ denote the standard deviation of $\sum_{k=1}^K w_k \hat{\tau}_k$. Then, we have
\begin{eqnarray}
\max_{\bm{\tau} \in \mathcal{T}} R(\bm{\tau},\hat{\delta}_{\bm{w}}) &=& \max_{t \geq 0} \max_{\bm{\tau} \in \mathcal{T}, \tau_0 = t} \left\{ t \cdot \Phi \left( - s^{-1}(\bm{w})\left( t+\sum_{k=1}^K w_k (\tau_k-t) \right) \right) \right\} \nonumber \\
&=& \max_{t \geq 0} \left\{ t \cdot \Phi \left( - s^{-1}(\bm{w})\left( t+\min_{\bm{\tau} \in \mathcal{T}, \tau_0 = t} \left\{ \sum_{k=1}^K w_k (\tau_k-\tau_0) \right\} \right) \right) \right\}. \nonumber
\end{eqnarray}

By Assumption 1, the term $\min_{\bm{\tau} \in \mathcal{T}, \tau_0 = t} \left\{ \sum_{k=1}^K w_k (\tau_k-\tau_0) \right\}$ does not depend on $t$. Hence, we obtain
\begin{equation}
b(\bm{w}) \ \equiv \ \max_{\bm{\tau} \in \mathcal{T}, \tau_0 = 0} \left\{ \sum_{k=1}^K w_k \tau_k \right\} \ = \ - \min_{\bm{\tau} \in \mathcal{T}, \tau_0 = t} \left\{ \sum_{k=1}^K w_k (\tau_k-\tau_0) \right\}. \nonumber
\end{equation}
Viewing $\sum_{k=1}^K w_k \hat{\tau}_k$ as an estimator for $\tau_0$, $b(\bm{w})$ and $s(\bm{w})$ can be interpreted as the maximum bias and the standard deviation of a linear estimator $\sum_{k=1}^K w_k \hat{\tau}_k$, respectively. Using these terms, we can express the maximum regret as
\begin{eqnarray}
\max_{\bm{\tau} \in \mathcal{T}} R(\bm{\tau},\hat{\delta}_{\bm{w}}) &=& \max_{t \geq 0} \left\{ t \cdot \Phi \left( - \frac{t}{s(\bm{w})} + \frac{b(\bm{w})}{s(\bm{w})} \right) \right\} \nonumber \\
&=& s(\bm{w}) \cdot \eta \left( \frac{b(\bm{w})}{s(\bm{w})} \right), \nonumber
\end{eqnarray}
where $\eta(a) \equiv \max_{t \geq 0}\left\{ t \cdot \Phi(-t+a) \right\}$. Hence, we obtain the following theorem:

\bigskip

\begin{Theorem}
Suppose that the parameter space $\mathcal{T}$ satisfies Assumption 1. Then, the minimax regret rule among $\mathcal{D}_{\mathrm{lin}}$ is obtained via the following optimization:
\begin{equation}
\bm{w}_{\mathrm{minimax}} \ \in \ \mathrm{arg} \min_{\bm{w}} \left\{ s(\bm{w}) \cdot \eta \left( \frac{b(\bm{w})}{s(\bm{w})} \right) \right\}. \label{Theorem 1}
\end{equation}
\end{Theorem}

In view of Theorem 1, we can compute $\bm{w}_{\text{minimax}}$ using the following algorithm: \vspace{0.05in}
\begin{itemize}
\item[\textbf{1.}] Fix $\bm{w}$ such that $\sum_{k=1}^K w_k = 1$. Calculate $s(\bm{w})$ and $b(\bm{w})$, and obtain the maximum regret $s(\bm{w}) \cdot \eta(b(\bm{w})/s(\bm{w}))$, where we approximate $\eta(\cdot)$ by a piecewise linear function.
\item[\textbf{2.}] Minimize the maximum regret $s(\bm{w}) \cdot \eta(b(\bm{w})/s(\bm{w}))$ subject to $\sum_{k=1}^K w_k = 1$. \vspace{0.05in}
\end{itemize}

\bigskip

In step 1, we compute $b(\bm{w})$ by solving the optimization in $\bm{\tau}$. As shown below, there are  many examples in which we can calculate $b(\bm{w})$ using linear programming.
If so, $b(\bm{w})$ can be solved quickly and reliably even when $K$ is large. Furthermore, because $t \mapsto t \cdot \Phi(-t+a)$ is a smooth unimodal function, $\eta(a)$ is easy to compute.

Figure \ref{fig:eta} displays the shape of $\eta(a)$ as a function of $a$. From the proof of Lemma \ref{lemma:2} below, we find that $\eta(a)$ is strictly increasing and convex. In numerical simulations given in Section 4 below, we compute the optimization of step 2 using the R package ``Rsolnp''. We find that this optimization step is quick and stable even when $K$ exceeds 100.
\begin{figure}[h]
\centering
\includegraphics[width=9cm]{eta.pdf}
\caption{Functional form of $\eta(a)$. The function $\eta(a)$ is strictly increasing and convex and $\eta(0)$ is approximately equal to 0.17.} \label{fig:eta}
\end{figure}

\begin{Remark}\label{rem:LP}
There are some important cases where we can calculate $b(\bm{w})$ using linear programming. For example, consider the parameter space $\mathcal{T}_{\text{meta}}$ of Example 1, where we have
$$
b(\bm{w}) \ = \ \max_{\beta \in \mathcal{B}} \left\{ \sum_{k=1}^K w_k(x_k-x_0)'\beta \right\}. \nonumber
$$
Hence, if $\mathcal{B}$ is a polyhedron, we can calculate $b(\bm{w})$ using linear programming.

If the parameter space is $\mathcal{T}_C$ as in Example 2, we have
$$
b(\bm{w}) \ = \ \max_{|\tau_k - \tau_l| \leq C_{kl}} \left\{ \sum_{k=1}^K w_k (\tau_k - \tau_0) \right\},
$$
where this maximization can again be solved using linear programming. As an alternative approach to compute the minimax regret rule in the current setting, Proposition 1 in \cite{montiel2023decision} presents a procedure based on a solution to a fixed point problem.
\end{Remark}

Even if the parameter space $\mathcal{T}$ satisfies Assumption 1, the minimax regret can be unbounded. For example, $\mathcal{T} = \mathbb{R}^{K+1}$ satisfies Assumption 1 but the maximum regret is unbounded with $b(\bm{w}) = +\infty$ for any $\bm{w}$. This is because $\lim_{a \rightarrow \infty} \eta(a) = + \infty$.

To have the maximum regret bounded, we need to impose a restriction that  the difference between $\tau_k$ and $\tau_0$ is bounded for some $k$.

\begin{Assumption}
There exists $M < \infty$ such that $\bm{\tau} \in \mathcal{T}$ implies $|\tau_k - \tau_0| < M$ for some $k$.
\end{Assumption}

\begin{Theorem}
Suppose that the parameter space $\mathcal{T}$ satisfies Assumptions 1 and 2. Then, the minimax regret is finite.
\end{Theorem}

Assumption 2 means that there exists some study in the pool that provides some (partially) identifying information about $\tau_0$. This condition holds for the parameter space $\mathcal{T}_{\mathrm{meta}}$ of Example 1 if $\mathcal{B}$ is compact. Similarly, $\mathcal{T}_C$ satisfies Assumption 2. Theorem 2 then applies to these cases and guarantees that the minimax regret is bounded.

\begin{Remark}\label{rem:construle}
We can add a constant term to equation (\ref{linear_aggregation_rule}), i.e., consider the following treatment rule:
\begin{eqnarray}
\hat{\delta}_{v,\bm{w}}(\mathbf{D}) \ = \ \mathbf{1} \left\{ v + \sum_{k=1}^K w_k \hat{\tau}_k \geq 0 \right\}, \label{linear constant rule}
\end{eqnarray}
where $v \in \mathbb{R}$. Then, the maximum regret of $\hat{\delta}_{v,\bm{w}}$ can be written as
$$
s(\bm{w}) \eta \left( s^{-1}(\bm{w})\{b(\bm{w})+|v|\} \right).
$$
Since $\eta(a)$ is strictly increasing in $a$, for $v \neq 0$, the maximum regret of $\hat{\delta}_{v,\bm{w}}$ is greater than that of $\hat{\delta}_{\bm{w}}$. Hence, the minimax regret rule sets $v = 0$, so that it is not necessary to consider treatment rules with an intercept like in (\ref{linear constant rule}).
\end{Remark}

\begin{Remark} \label{rem:without_A1}
If Assumption 1 is relaxed, the optimization problem for maximum regret can be expressed as
\begin{eqnarray}
\max_{\bm{\tau} \in \mathcal{T}} R(\bm{\tau},\hat{\delta}_{\bm{w}})
&=& \max_{\bm{\tau} \in \mathcal{T}} \left\{ | \tau_0 | \cdot \Phi \left( - sgn(\tau_0) \cdot \frac{\tau_0 + \sum_{k=1}^K w_k (\tau_k-\tau_0)}{s(\bm{w})}\right) \right\} \nonumber \\
&=& \max_{\bm{\tau} \in \mathcal{T}} \left\{ | \tau_0 | \cdot \Phi \left( \cdot \frac{- |\tau_0| - sgn(\tau_0) \cdot \sum_{k=1}^K w_k (\tau_k-\tau_0)}{s(\bm{w})}\right) \right\} \nonumber \\
&=& \max_{t \in \mathcal{S}} \left\{ | t | \cdot \Phi \left( \cdot \frac{- |t| - \min_{\tau \in \mathcal{T}, \tau_0 = t} \left\{ sgn(t) \cdot \sum_{k=1}^K w_k (\tau_k-t) \right\} }{s(\bm{w})}\right) \right\}, \nonumber
\end{eqnarray}
where $\mathcal{S} \equiv \{\tau_0: \bm{\tau} \in \mathcal{T} \} $. Hence, the maximum regret can be written as
\begin{equation}
\max_{\bm{\tau} \in \mathcal{T}} R(\bm{\tau},\hat{\delta}_{\bm{w}}) \ = \ \max_{t \in \mathcal{S}} \left\{ | t | \cdot \Phi \left( \cdot \frac{- |t| + \tilde{b}(t,\bm{w}) }{s(\bm{w})}\right) \right\}, \label{maximum_regret_without_Assumption_1}
\end{equation}
where
$$
\tilde{b}(t,\bm{w}) \equiv - \min_{\tau \in \mathcal{T}, \tau_0 = t} \left\{ sgn(t) \cdot \sum_{k=1}^K w_k (\tau_k-t) \right\}.
$$
The weights of the minimax regret rule therefore solve
$$
\bm{w}_{\mathrm{minimax}} \ \in \ \mathrm{arg}\min_{\bm{w}} \max_{t \in \mathcal{S}} \left\{ |t| \cdot \Phi \left( - \frac{|t| - \tilde{b}(t,\bm{w})}{s(\bm{w})} \right) \right\}.
$$
Since the parameter space $\mathcal{T}$ does not satisfy the second condition in Assumption \ref{parameter_space}, the maximum bias $\tilde{b}(t,\bm{w})$ may depend on $t$. This complicates computation of $\bm{w}_{\mathrm{minimax}}$.
\end{Remark}

\begin{Remark} \label{rem:randomized}
Our results can be extended to the following randomized statistical treatment rules:
\begin{equation}
\hat{\delta}_{v, \bm{w}} (\mathbf{D}, Z_v) \ = \ \mathbf{1} \left\{ \sum_{k=1}^K w_k \hat{\tau}_k + Z_v \geq 0 \right\}, \label{randomized_rule}
\end{equation}
where $v \geq 0$ and $Z_v \sim N(0,v^2)$ is independent of the data $\mathbf{D}$. Then, conditional on the data, this randomized statistical treatment rule (\ref{randomized_rule}) becomes 1 with probability $\Phi_v\left( \sum_{k=1}^K w_k \hat{\tau}_k \right)$, where $\Phi_v$ is the distribution function of $N(0,v^2)$. This rule becomes the linear aggregation rule when $v=0$. Because $\sum_{k=1}^K w_k \hat{\tau}_k + Z_v$ is normally distributed with mean $\sum_{k=1}^K w_k \tau_k$ and variance $\sum_{k=1}^K w_k^2 \sigma_k^2 + v^2$, we obtain
\begin{equation*}
R(\bm{\tau}, \hat{\delta}_{v, \bm{w}}) \ = \ | \tau_0 | \cdot \Phi \left( - sgn(\tau_0) \cdot \frac{\sum_{k=1}^K w_k \tau_k}{\sqrt{\sum_{k=1}^K w_k^2 \sigma_k^2 + v^2}}\right).
\end{equation*}
From the same discussion as in Theorem 1, we obtain
\begin{equation*}
\max_{\bm{\tau} \in \mathcal{T}} R(\bm{\tau}, \hat{\delta}_{v, \bm{w}}) \ = \ s(v,\bm{w}) \cdot \eta \left( \frac{b(\bm{w})}{s(v,\bm{w})} \right),
\end{equation*}
where $s(v,\bm{w}) \equiv \sqrt{\sum_{k=1}^K w_k^2 \sigma_k^2 + v^2}$. Therefore, optimization for finding a minimax regret rule among this class of randomized rules is similar to and a straightforward extension of the optimization for non-randomized rules shown in Theorem 1. Note that the class of randomized rules we are optimizing over contains the minimax regret rule shown in  \citet{yata2021optimal}.
\end{Remark}

\section{Extensions and discussion}

\subsection{Comparison with a minimax MSE rule}

In this section, we compare $\bm{w}_{\text{minimax}}$ with other ways of forming the weights. First, we consider a minimax linear estimator of $\tau_0$ in terms of the mean squared errors (MSE). It is well known that the maximum MSE of $\sum_{k=1}^K w_k \hat{\tau}_k$ can be decomposed into the variance and the squared maximum bias:
\begin{equation}
b^2(\bm{w}) + s^2(\bm{w}). \nonumber
\end{equation}
Hence, the weights of the minimax MSE estimator are
\begin{equation}
\bm{w}_{\text{MSE}} \ \in \ \text{arg} \min_{\bm{w}} \left\{ b^2(\bm{w}) + s^2(\bm{w}) \right\}. \label{MSE weight}
\end{equation}
We refer to $\hat{\delta}_{\bm{w}_{\text{MSE}}}$ as the minimax MSE rule.

To compare the minimax regret and MSE rules, we  focus on the analytical properties of $\eta(a)$. \cite{tetenov2012statistical} shows that $\eta(a)$ is a continuous, strictly increasing function and $\eta(0) \simeq 0.17$. Furthermore, in the proof of the following lemmas, we show that $\eta(a)$ is concave. We accordingly obtain the following upper and lower bounds on $\eta(a)$:

\begin{Lemma}
For any $v, a \geq 0$, we have
\begin{equation}
v\cdot\Phi(-v) + \Phi(-v)\cdot a \ \ \leq \ \ \eta(a) \ \ \leq \ \ \eta(0) + a. \label{Lemma 1}
\end{equation}
\end{Lemma}

\begin{Lemma} \label{lemma:2}
For any $a \geq 0$, we have
\begin{equation}
\eta(0) \sqrt{1+a^2} \ \ \leq \ \ \eta(a) \ \ \leq \ \ \sqrt{1+a^2}. \label{eq:Lemma2}
\end{equation}
\end{Lemma}

Relying on Theorem 1 and Lemmas 1--2, the next theorem bounds the maximum regret.

\bigskip

\begin{Theorem}
Suppose that the parameter space $\mathcal{T}$ satisfies Assumptions 1 and 2. Then, for any $\bm{w}$ and $v \geq 0$, we obtain
$$
\Phi(-v)\cdot s(\bm{w}) + v \cdot \Phi(-v)\cdot b(\bm{w}) \ \leq \ \max_{\bm{\tau} \in \mathcal{T}} R(\bm{\tau},\hat{\delta}_{\bm{w}}) \ \leq \ \eta(0)\cdot s(\bm{w}) + b(\bm{w}).
$$
In addition, we obtain
\begin{equation}
\eta(0) \cdot \sqrt{b^2(\bm{w}) + s^2(\bm{w})} \ \leq \ \max_{\bm{\tau} \in \mathcal{T}} R(\bm{\tau},\hat{\delta}_{\bm{w}}) \ \leq \ \sqrt{b^2(\bm{w}) + s^2(\bm{w})}. \nonumber \label{Theorem 3}
\end{equation}
\end{Theorem}

\bigskip

Theorem 3 provides lower and upper bounds on the maximum regret. These bounds show that the maximum regret is bounded from above and from below by $b(\bm{w}) + s(\bm{w})$ and $\sqrt{b^2(\bm{w}) + s^2(\bm{w})}$ up to some proportional factors, independently of the number of studies, $K$, and the dimension of $x_k$, $d_x$. The second set of inequalities imply that the minimax regret is equivalent to the minimax RMSE (root-MSE) up to a constant factor. In other words, minimax RMSE enables us to bound the minimax regret.

Furthermore, Theorem 1 and Lemma \ref{lemma:2} lead to the following comparison of the maximum regret between the minimax regret rule  $\hat{\delta}_{\bm{w}_{\text{minimax}}}$ and the minimax MSE rule $\hat{\delta}_{\bm{w}_{\text{MSE}}}$.

\bigskip

\begin{Theorem}
Suppose that the parameter space $\mathcal{T}$ satisfies Assumptions 1 and 2. Then, we have
\begin{equation}
1 \ \leq \ \cfrac{\max_{\bm{\tau} \in \mathcal{T}} R(\bm{\tau},\hat{\delta}_{\bm{w}_{\mathrm{MSE}}})}{\max_{\bm{\tau} \in \mathcal{T}} R(\bm{\tau},\hat{\delta}_{\bm{w}_{\mathrm{minimax}}})} \ \leq \frac{1}{\eta(0)} \ \simeq \ 5.88. \label{Theorem 4}
\end{equation}
\end{Theorem}

\bigskip

Theorem 4 shows that the maximum regret of the minimax MSE rule is the same as the minimax regret up to a constant factor, independently of $K$ and $d_x$. Numerical simulations in Section 4 suggest that the maximum regret of $\hat{\delta}_{\bm{w}_{\text{MSE}}}$ can be about 40 percent greater than the minimax regret.

The proof of Lemma \ref{lemma:2} given in the Appendix shows that the minimax regret criterion places greater emphasis on the bias than on the variance compared with the minimax MSE criterion. To see this, consider the directional derivatives of the maximum regret and MSE. We fix $\bm{\theta} = (\theta_1, \cdots , \theta_K)'$ with $\sum_{k=1}^K \theta_k = 1$ and assume that $b(\bm{w})$ and $s(\bm{w})$ are directionally differentiable. We define
\begin{eqnarray}
b_{\bm{\theta}}'(\bm{w}) &\equiv & \lim_{h \downarrow 0} \left( b\left((1-h)\cdot \bm{w}+h\cdot \bm{\theta} \right) - b(\bm{w}) \right)/h , \nonumber \\
s_{\bm{\theta}}'(\bm{w}) &\equiv & \lim_{h \downarrow 0} \left( s\left((1-h)\cdot \bm{w}+h\cdot \bm{\theta} \right) - s(\bm{w}) \right)/h , \nonumber \\
Q_{\bm{\theta}}(\bm{w}) & \equiv & \left( \max_{\bm{\tau} \in \mathcal{T}} R(\bm{\tau},\hat{\delta}_{(1-h) \cdot \bm{w}+h \cdot \bm{\theta}}) - \max_{\bm{\tau} \in \mathcal{T}} R(\bm{\tau},\hat{\delta}_{\bm{w}}) \right)/h, \nonumber
\end{eqnarray}
where $Q_{\bm{\theta}}(\bm{w})$ is the directional derivative of the maximum regret. Let $t^*(a)$ be the maximizer of $t \cdot \Phi(-t+a)$. Then, by the proof of Lemma \ref{lemma:2}, we have
\begin{eqnarray}
Q_{\bm{\theta}}(\bm{w}) &=& \Phi\left( -t^*\left( \frac{b(\bm{w})}{s(\bm{w})} \right) + \frac{b(\bm{w})}{s(\bm{w})} \right) \times \left\{ s_{\bm{\theta}}'(\bm{w}) \left( t^*\left( \frac{b(\bm{w})}{s(\bm{w})} \right) - \frac{b(\bm{w})}{s(\bm{w})} \right) + b_{\bm{\theta}}'(\bm{w}) \right\}. \nonumber
\end{eqnarray}
Here, the sign of $Q_{\bm{\theta}}(\bm{w})$ is determined by
$$
s_{\bm{\theta}}'(\bm{w}) \left( t^*\left( \frac{b(\bm{w})}{s(\bm{w})} \right) - \frac{b(\bm{w})}{s(\bm{w})} \right) + b_{\bm{\theta}}'(\bm{w}),
$$
where $t^*(a) - a$ is decreasing in $a$ as shown in the proof of Lemma \ref{lemma:2}. Similarly, the sign of the directional derivative of the maximum MSE, $b^2(\bm{w})+s^2(\bm{w})$, is determined by
$$
s_{\bm{\theta}}'(\bm{w}) \left( \frac{b(\bm{w})}{s(\bm{w})} \right)^{-1} + b_{\bm{\theta}}'(\bm{w}).
$$
Suppose that $b_{\bm{\theta}}'(\bm{w}) < 0$ and $s_{\bm{\theta}}'(\bm{w}) > 0$, that is, we face the bias-variance tradeoff. Then, because numerical evaluation implies $t^*(a)-a < a^{-1}$ for $a \geq 0$, we obtain
$$
s_{\bm{\theta}}'(\bm{w}) \left( t^*\left( \frac{b(\bm{w})}{s(\bm{w})} \right) - \frac{b(\bm{w})}{s(\bm{w})} \right) + b_{\bm{\theta}}'(\bm{w}) \ < \ s_{\bm{\theta}}'(\bm{w}) \left( \frac{b(\bm{w})}{s(\bm{w})} \right)^{-1} + b_{\bm{\theta}}'(\bm{w}).
$$
When $\bm{w} = \bm{w}_{\mathrm{MSE}}$, the right-hand side must be zero. Hence, if $b_{\bm{\theta}}'(\bm{w}_{\mathrm{MSE}}) < 0$ and $s_{\bm{\theta}}'(\bm{w}_{\mathrm{MSE}}) > 0$, we conclude
\[
Q_{\bm{\theta}}\left( \bm{w}_{\mathrm{MSE}} \right) \ < \ 0.
\]
That is, at the minimax MSE weights $\bm{w} = \bm{w}_{\mathrm{MSE}}$, locally perturbing the weight vector in the direction that reduces the bias and increases the variance improves the welfare regret. This implies that the minimax regret criterion places greater emphasis on the bias than on the variance compared with the minimax MSE criterion. In the numerical analysis of Section 4, we plot $\bm{w}_{minimax}$ and $\bm{w}_{\mathrm{MSE}}$ to illustrate the difference in their bias-variance balancing properties.

\subsection{Simple case: $K=2$}
To illustrate the minimax regret rule among $\mathcal{D}_{\mathrm{lin}}$ and compare it with Bayes rules, consider the simple case where $K=2$ and $\mathcal{T} = \{ \bm{\tau}: |\tau_k - \tau_l| \leq C \ \text{for $k,l=0,1,2$ and $k \neq l$} \}$.
Because $w_1 + w_2 = 1$, we can express $(w_1,w_2) = (w,1-w)$ and the maximum regret of $\hat{\delta}_{\bm{w}}$ as $s(w) \cdot \eta \left( b(w)/s(w) \right)$, where $b(w) = \max_{\bm{\tau} \in \mathcal{T}, \tau_0 = 0} \left\{ w \tau_1 + (1-w) \tau_2 \right\}$ and $s(w) = \sqrt{w^2 \sigma_1^2 + (1-w)^2 \sigma_2^2}$.
First, we derive the analytical expression of $b(w)$. When $0 \leq w \leq 1$, we obtain $ b(w) = \max_{\bm{\tau} \in \mathcal{T}, \tau_0 = 0} \left\{ w \tau_1 + (1-w) \tau_2 \right\} = C$. For $w > 1$, we obtain $b(w) =  \max_{\bm{\tau} \in \mathcal{T}, \tau_0 = 0} \left\{ w \tau_1 + (1-w) \tau_2 \right\} = C w$. Similarly, we have $b(w) = C (1-w)$ when $w < 0$. Hence, the maximum bias can be written as follows:
\begin{equation*}
b(w) \ = \ C \cdot \max\{ w, 1-w, 1 \}.
\end{equation*}

Next, we derive the minimax weight $w^{\ast} \in \text{arg} \min_{w} \left\{ s(w) \cdot \eta \left( b(w)/s(w) \right) \right\}$. From the proof of Lemma 2, let $t^*(a) \equiv \text{arg} \max_{t \geq 0} \left\{ t \cdot \Phi(-t+a) \right\}$ and obtain
\begin{eqnarray*}
\frac{\partial}{\partial s} \left\{ s \eta(C/s) \right\} &=& \eta\left( \frac{C}{s} \right) - \left( \frac{C}{s} \right) \eta'\left( \frac{C}{s} \right) \\
&=& \left\{ t^{\ast} \left( \frac{C}{s} \right) - \left( \frac{C}{s} \right) \right\} \cdot \Phi\left( - t^{\ast} \left( \frac{C}{s} \right) + \left( \frac{C}{s} \right)  \right),
\end{eqnarray*}
where $t^{\ast}(a) - a$ is a strictly decreasing function.
From numerical evaluation, we have $t^{\ast}(a) - a = 0$ when $a = a^{\ast} \risingdotseq 1.253$.
Hence, $s \eta(C/s)$ is increasing when $s > C / a^{\ast} \risingdotseq 0.798 C$ and decreasing when $s < C / a^{\ast}$.
In addition, $\sqrt{\frac{\sigma_1^2 \sigma_2^2}{\sigma_1^2 + \sigma_2^2}} \leq s(w) \leq \max\{\sigma_1, \sigma_2\}$ holds with the inequalities hold with equalities at some $0 \leq w \leq 1$. Because $b(w) > C$ for $w < 0$ or $w > 1$ and $\eta(a)$ is a strictly increasing function, the minimax weight $w^{\ast}$ satisfies the following conditions:
\begin{eqnarray*}
\text{if} \ C / a^{\ast} < \sqrt{\frac{\sigma_1^2 \sigma_2^2}{\sigma_1^2 + \sigma_2^2}} & \Rightarrow & s(w^{\ast}) = \sqrt{\frac{\sigma_1^2 \sigma_2^2}{\sigma_1^2 + \sigma_2^2}}, \\
\text{if} \ \sqrt{\frac{\sigma_1^2 \sigma_2^2}{\sigma_1^2 + \sigma_2^2}} \leq C / a^{\ast} \leq \max\{\sigma_1, \sigma_2\} & \Rightarrow & s(w^{\ast}) = C / a^{\ast}, \\
\text{if} \ C / a^{\ast} >  \max\{\sigma_1, \sigma_2\} & \Rightarrow & w^{\ast} \leq 0 \ \text{or} \ w^{\ast} \geq 1.
\end{eqnarray*}
Therefore, if the dispersion of parameters $C$ is small compared to the standard deviations $\sigma_1$ and $\sigma_2$, the minimax weight attains the smallest variance, that is, $w^{\ast} = \frac{\sigma_2^2}{\sigma_1^2 + \sigma_2^2}$. On the other hand, if $C$ is large compared to $\sigma_1$ and $\sigma_2$, then the minimax regret criterion may favor rules with variance $s(w)$ larger than the minimum.\footnote{If we consider the randomized statistical treatment rules in (\ref{randomized_rule}), for $C \geq a^{\ast} \sqrt{\frac{\sigma_1^2 \sigma_2^2}{\sigma_1^2 + \sigma_2^2}}$ the maximum regret is minimized when $w \in [0,1]$ and
\begin{equation}
s(v,w) \ \equiv \ \sqrt{w^2\sigma_1^2+(1-w)^2\sigma_2^2 + v^2} \ = \ C/a^{\ast}. \label{minimax_K=2_randomized}
\end{equation}
Hence, if $C / a^{\ast} >  \max\{\sigma_1, \sigma_2\}$, then $v$ must be positive to satisfy (\ref{minimax_K=2_randomized}). This implies that when the dispersion of parameters is large, the minimax regret criterion may favor randomized treatment rules over non-randomized treatment rules, and this observation is consistent with \cite{stoye2012minimax} and \cite{yata2021optimal}.}

To compare the minimax regret rule with Bayes rules, consider the following hierarchical Bayes model:
\begin{eqnarray*}
& & \tau_k = \tau_0 + \epsilon_k, \ \ k = 1,2, \\
& & \tau_0 \sim N(0, \sigma_{\tau}^2), \ \ \epsilon_k \sim N(0, \sigma_{\epsilon}^2),
\end{eqnarray*}
where $\tau_0$, $\epsilon_1$, and $\epsilon_2$ are mutually independent. Then the posterior distribution of $\tau_0$ can be written as follows:
\begin{eqnarray*}
\pi(\tau_0 | \hat{\tau}_1, \hat{\tau}_2)
& \propto & \exp \left[ - \frac{1}{2} \left\{ (\sigma_1^2 + \sigma_{\epsilon}^2)^{-1} + (\sigma_2^2 + \sigma_{\epsilon}^2)^{-1} + \sigma_{\tau}^{-2} \right\} \right. \\
& & \hspace{0.8in} \left. \times \left\{ \tau_0 - \frac{(\sigma_1^2 + \sigma_{\epsilon}^2)^{-1} \hat{\tau}_1 + (\sigma_2^2 + \sigma_{\epsilon}^2)^{-1} \hat{\tau}_2}{(\sigma_1^2 + \sigma_{\epsilon}^2)^{-1} + (\sigma_2^2 + \sigma_{\epsilon}^2)^{-1} + \sigma_{\tau}^{-2}} \right\} \right].
\end{eqnarray*}
As discussed in Remark 1, the Bayes optimal rule becomes
\begin{eqnarray*}
\bm{1}\{ E_{\pi}( \tau_0 | \mathbf{D}) \geq 0 \}
&=& \bm{1} \left\{ \left( \frac{\sigma_2^2 + \sigma_{\epsilon}^2}{\sigma_1^2 + \sigma_2^2 + 2 \sigma_{\epsilon}^2} \right) \hat{\tau}_1 + \left( \frac{\sigma_1^2 + \sigma_{\epsilon}^2}{\sigma_1^2 + \sigma_2^2 + 2 \sigma_{\epsilon}^2} \right)  \hat{\tau}_2 \geq 0 \right\}.
\end{eqnarray*}
Note that the minimax regret weight $w^{\ast}$ and the weight of the Bayes optimal rule $w_{\text{HB}} = \frac{ \sigma_2^2 + \sigma_{\epsilon}^2}{\sigma_1^2 + \sigma_2^2 + 2 \sigma_{\epsilon}^2}$ satisfy $|w^{\ast}-1/2| > |w_{\text{HB}}-1/2|$ whenever $\sigma_{\epsilon}^2 > 0$, i.e, $w_{\text{HB}}$ shrinks $w^{\ast}$ toward $1/2$, and the degree of shrinkage is increasing in $\sigma_{\epsilon}^2$. The weights $w^{\ast}$ and $w_{\text{HB}}$ agree only in an extreme scenario of $\sigma_{\epsilon}^2 = 0$.
\subsection{Consistency of the minimax treatment rule}

If we have perfect knowledge of $(\tau_1, \dots, \tau_K)$, i.e., $\sigma_k = 0$ for all $1 \leq k \leq K$, we can obtain the (true) identified set of $\tau_0$ based on the constraints on the parameter space $\mathcal{T}$. For instance, we construct the identified set of $\tau_0$ by intersecting multiple bounds for $\tau_0$, each of which is constructed by extrapolating from $\tau_k$, as considered in \citet{manski2020towards}. We can then consider finding the minimax regret treatment rule given the true identified set of $\tau_0$ without any sampling uncertainty. We denote by $\delta^*_{IS}$ such a (non-randomized) minimax regret rule. As the sample size of each study increases, that is, as $\sigma_k \rightarrow 0$, should we expect the minimax regret rule $\hat{\delta}_{\bm{w}_{\text{minimax}}}$ we constructed in the previous section to converge to $\delta^*_{IS}$?

In what follows, we compare $\delta^*_{IS}$ with the limiting version of $\hat{\delta}_{\bm{w}_{\text{minimax}}}$, and show that $\hat{\delta}_{\bm{w}_{\text{minimax}}}$ does \textit{not} necessarily converge to $\delta^*_{IS}$ as $\sigma_k \rightarrow 0$. We then consider an alternative class of treatment choice rules that converge to $\delta^*_{IS}$ as $\sigma_k \rightarrow 0$. These alternative treatment rules solve the minimax regret with a data-driven parameter space built upon confidence regions for $\bm{\tau}$.
These rules, therefore, do not belong to the linear aggregation rules of Definition \ref{def:linear_rule}.
Moreover, their computation are not as simple as the linear minimax regret rule $\delta_{\bm{w}_{\text{minimax}}}$ obtained in the previous section, and we do not know if they coincide with any exact minimax regret (nonlinear) rule obtained for a data-independent parameter space.
Nevertheless, we can show that such modified treatment rules converge to $\delta^*_{IS}$ as $\sigma_k \rightarrow 0$, which could be of theoretical interest.

First, we consider the minimax regret rule when $\sigma_k = 0$ for all $k = 1, \dots , K$. Then, $\tau_k = \hat{\tau}_k$ for all $k = 1, \dots , K$ and the identified set of $\tau_0$ is
\begin{equation}
IS_0 \ \equiv \ \left\{ \tau_0 : \text{$\bm{\tau} \in \mathcal{T}$ and $\tau_k = \hat{\tau}_k$ for all $k = 1, \cdots , K$} \right\}. \label{identified_set}
\end{equation}
In this case, the parameter space $\mathcal{T}$ projected for $\tau_0$ yields the identified set $IS_0$. Because there is no randomness in this problem, for a treatment rule $\delta$, the welfare regret of $\delta$ becomes
$$
W(\delta^*)-W(\delta) \ = \ \begin{cases}
    -\tau_0 \cdot \delta & \text{if $\tau_0 < 0$} \\
    \tau_0 \cdot (1-\delta) & \text{if $\tau_0 \geq 0$}
  \end{cases}.
$$
Hence, the minimax regret rule over $IS_0$ can be written as
\begin{equation}
\delta_{IS}^*(\mathbf{D}) \ \equiv \ \begin{cases}
    1 & \ \ \ \ \ \text{if $-\underline{\tau}_0 \leq \overline{\tau}_0$} \\
    0 & \ \ \ \ \ \text{if $-\underline{\tau}_0 > \overline{\tau}_0$}
  \end{cases}, \label{minimax_IS}
\end{equation}
where $\underline{\tau}_0 \equiv \inf\{\tau_0:\tau_0 \in IS_0\}$ and $\overline{\tau}_0 \equiv \sup\{\tau_0:\tau_0 \in IS_0\}$ are the smallest and largest values of the identified set of $\tau_0$, respectively. The rule $\delta_{IS}^*$ becomes 1 (or 0) when we have $\underline{\tau}_0 > 0$ (or $\overline{\tau}_0 < 0$), that is, all values of the identified set of $\tau_0$ are positive (or negative). When the identified set of $\tau_0$ contains both of positive and negative values, $\delta_{IS}^*$ becomes 1 (or 0) if the absolute value of $\overline{\tau}_0$ is larger (or smaller) than that of $\underline{\tau}_0$.\footnote{In this paper, we consider only non-randomized treatment rules. \cite{manski2011choosing} shows that the minimax regret criterion always yields a randomized treatment rule when $\underline{\tau}_0 < 0 < \overline{\tau}_0$. He shows that the minimax randomized treatment rule randomly assigns a fraction $|\overline{\tau}_0|/(|\underline{\tau}_0|+|\overline{\tau}_0|)$ of the population to treatment 1 and the remaining $|\underline{\tau}_0|/(|\underline{\tau}_0|+|\overline{\tau}_0|)$ to treatment 0.}

\bigskip

Next, we consider the large sample properties of our minimax regret rule $\hat{\delta}_{\bm{w}_{\text{minimax}}}$, i.e., $\sigma_k \rightarrow 0$. In this case, we can show that the minimax regret criterion yields the treatment rule that minimizes the maximum bias. The proof of Lemma \ref{lemma:2} shows that $\eta(a)$ is strictly increasing and convex with its slope bounded from above by one. Hence, the slope of $\eta(a)$ converges to a positive constant $c \in (0,1]$ as $a \rightarrow +\infty$. This implies that when $a$ is large, $\eta(a)$ can be approximated by $d+c \cdot a$ for some $d$. As $\sigma_k \rightarrow 0$ for all $k$, we have $s(\bm{w}) \rightarrow 0$ for any $\bm{w}$. From Theorem 1, as $s(\bm{w}) \rightarrow 0$, we can approximate the maximum regret of $\hat{\delta}_{\bm{w}}$ by $c \cdot b(\bm{w})$. This implies that in large samples, the minimax regret rule becomes a treatment rule that minimizes the maximum bias $b(\bm{w})$.

To be specific, consider the case in which the parameter space is the class of Lipschitz vectors given in Example 2. Since we have
$$
b(\bm{w}) \ = \ \max_{|\tau_k - \tau_l| \leq C \|x_k - x_l\|} \left\{ \sum_{k=1}^K w_k (\tau_k - \tau_0) \right\},
$$
as $\sigma_k \rightarrow 0$ for all $k$, the minimax regret rule converges to the rule that depends only on the closest study in terms of the metric on the covariate space, i.e., the weight of the closest study $w_{k^*}$ converges to $1$, where $k^*$ satisfies $\|x_{k^*}-x_0\| \leq \|x_k-x_0\|$ for all $k = 1, \cdots , K$. Hence, the minimax regret rule converges to
$$
\hat{\delta}_{\bm{w}_{\text{minimax}}}(\mathbf{D}) \ = \ \mathbf{1}\{ \hat{\tau}_{k^*} \geq 0 \},
$$
and the decision of whether or not to introduce the policy is solely based on the closest study.

In this case, we can show that $\hat{\delta}_{\bm{w}_{\text{minimax}}}$ does not converge to $\delta^*_{IS}$ as $\sigma_k \rightarrow 0$. If the observed covariates are scalar, then the identified set of $\tau_0$ can be written as
$$
\bigcap_{k = 1}^K \big[ \hat{\tau}_{k}-C|x_k-x_0|,\hat{\tau}_{k}+C|x_k-x_0| \big].
$$
Because our minimax regret rule $\hat{\delta}_{\bm{w}_{\text{minimax}}}$ uses only the closest study, it does not agree with $\delta^*_{IS}$ from (\ref{minimax_IS}). In fact, it is possible that $\hat{\tau}_{k^*}$ is positive but that the absolute value of $\overline{\tau}_0$ is larger than that of $\underline{\tau}_0$.

\bigskip

To resolve such a disagreement, we propose a minimax treatment rule refined by a confidence region of $\bm{\tau}$. Data $\mathbf{D}$ provide some information about the parameter space $\mathcal{T}$. If there is not an a priori assumption available to constrain $\mathcal{T}$, we may want to exploit such in-sample information to refine the minimax regret rule.

For $\alpha \in (0,1)$, consider a subset $\hat{\mathcal{T}}(\alpha) \subset \mathcal{T}$ that depends on the data $\mathbf{D}$ and satisfies
\begin{equation}
P_{\bm{\tau}}\left( \bm{\tau} \in \hat{\mathcal{T}}(\alpha) \right) \ \geq \ 1 - \alpha \ \ \text{for any $\bm{\tau} \in \mathcal{T}$,} \label{hat_Tau}
\end{equation}
where $P_{\bm{\tau}}$ is the sampling probability distribution of the data when the true parameter value is $\bm{\tau}$. $\hat{\mathcal{T}}(\alpha)$ is a confidence set for $\bm{\tau}$ with a coverage probability of at least $1-\alpha$. For example, the following hyper-rectangle satisfies condition (\ref{hat_Tau}):
\[
\hat{\mathcal{T}}_{\text{HR}}(\alpha) \ \equiv \ \left\{ \bm{\tau} \in \mathcal{T} : \tau_k \in [\hat{\tau}_k - \sigma_k \cdot z_{\alpha, K}, \hat{\tau}_k + \sigma_k \cdot z_{\alpha, K}] \ \text{for $k = 1, \cdots, K$} \right\},
\]
where $z_{\alpha, K}$ is the value such that $P(|Z| \leq z_{\alpha,K}) = (1-\alpha)^{1/K}$ for a standard normal variable $Z$.

When the parameter space is $\mathcal{T}_{\text{meta}}$ and $d_x < K$, we can construct another confidence region of $\bm{\tau}$. Let $\hat{\beta}$ be the OLS estimator of $\beta$. Then, the following set satisfies  the condition (\ref{hat_Tau}):
\[
\hat{\mathcal{T}}_{\text{meta}}(\alpha) \ \equiv \ \ \left\{ \bm{\tau} \in \mathcal{T}_{\text{meta}} : (\hat{\beta}-\beta)'S(\hat{\beta})^{-1}(\hat{\beta}-\beta) \leq \chi(\alpha,d_x) \right\},
\]
where $S(\hat{\beta})$ is the variance matrix of $\hat{\beta}$ and $\chi(\alpha,d_x)$ is the $(1-\alpha)$-th quantile of the chi-square distribution with $d_x$ degrees of freedom.

By replacing the parameter space $\mathcal{T}$ with $\hat{\mathcal{T}}(\alpha)$, we can compute the refined minimax regret rule. We define
\begin{equation}
\hat{\bm{w}}_{\text{minimax}}(\alpha) \ \in \ \text{arg} \min_{\bm{w}} \max_{\bm{\tau} \in \hat{\mathcal{T}}(\alpha)} R(\bm{\tau},\hat{\delta}_{\bm{w}}). \label{w_minimax_alpha}
\end{equation}
For the class of linear aggregation rules considered in the previous sections, $\bm{w}_{\text{minimax}}$ cannot depend on the data $\mathbf{D}$. In contrast, $\hat{\bm{w}}_{\text{minimax}}(\alpha)$ depends on the data through $\hat{\mathcal{T}}(\alpha)$. Hence, the refined minimax regret rule $\hat{\delta}_{\hat{\bm{w}}_{\text{minimax}}(\alpha)}$ becomes a non-linear aggregation rule. Because $\hat{\mathcal{T}}(\alpha)$ is contained in the parameter space $\mathcal{T}$, the refined minimax regret rule is less conservative than $\hat{\delta}_{\bm{w}_{\text{minimax}}}$. From (\ref{hat_Tau}), for any $\bm{w}$, we obtain
\[
R(\bm{\tau},\hat{\delta}_{\bm{w}}) \ \leq \ \max_{\bm{\tau} \in \hat{\mathcal{T}}(\alpha)} R(\bm{\tau},\hat{\delta}_{\bm{w}}) \ \ \text{with probability $1-\alpha$.}
\]
Hence, $\hat{\bm{w}}_{\text{minimax}}(\alpha)$ minimizes the worst-case regret over $\hat{\mathcal{T}}(\alpha)$, which is a valid upper bound on the true regret with probability $1-\alpha$.

Since $\hat{\mathcal{T}}(\alpha)$ may not satisfy Assumption 1, we cannot derive $\hat{\bm{w}}_{\text{minimax}}(\alpha)$ using Theorem 1. However, even if $\hat{\mathcal{T}}(\alpha)$ does not satisfy Assumption 1, we can calculate $\hat{\bm{w}}_{\text{minimax}}(\alpha)$ using (\ref{maximum_regret_without_Assumption_1}) in Remark \ref{rem:without_A1}. When the parameter space is $\mathcal{T}_{\text{meta}}$, we can easily calculate the refined minimax regret rule using $\hat{\mathcal{T}}_{\text{meta}}(\alpha)$. If $(\hat{\beta}-\beta)'S(\hat{\beta})^{-1}(\hat{\beta}-\beta) \leq \chi(\alpha,d_x)$ implies $\beta \in \mathcal{B}$, then we have
\[
\hat{\mathcal{T}}_{\text{meta}}(\alpha) \ \equiv \ \ \left\{ \bm{\tau}: \text{$\tau_k = \beta_0 + x_k'\beta$, $\beta_0 \in \mathbb{R}$, and $(\hat{\beta}-\beta)'S(\hat{\beta})^{-1}(\hat{\beta}-\beta) \leq \chi(\alpha,d_x)$ } \right\}.
\]
Then, for $t >0$, we obtain
\begin{eqnarray}
\tilde{b}(t,\bm{w}) &=& \max_{\bm{\tau} \in \hat{\mathcal{T}}_{\mathrm{meta}}(\alpha), \tau_0 = t} \left\{\sum_{k=1}^K w_k (\tau_0 -\tau_k) \right\} \nonumber \\
&=& \max_{(\hat{\beta}-\beta)'S(\hat{\beta})^{-1}(\hat{\beta}-\beta) \leq \chi(\alpha,d_x)} \left\{\sum_{k=1}^K w_k (x_0 - x_k)'\beta \right\}. \nonumber
\end{eqnarray}
Let $\beta^*$ be a maximizer of the above problem. Using the method of Lagrange multipliers, we find that $\beta^*$ satisfies
\begin{equation}
\arraycolsep=0pt
\begin{array}{l}
X_0(\bm{w}) - 2 \lambda S(\hat{\beta})^{-1}(\beta^* - \hat{\beta})
=0, \\
(\hat{\beta}-\beta^*)'S(\hat{\beta})^{-1}(\hat{\beta}-\beta^*) - \chi(\alpha,d_x)
=0,
\end{array} \label{FOC}
\end{equation}
where $X_0(\bm{w}) \equiv \sum_{k=1}^K w_k (x_0 - x_k)$. Then, equations (\ref{FOC}) imply that
\begin{equation}
\beta^* \ = \ \hat{\beta} + \frac{\sqrt{\chi(\alpha,d_x)}}{\sqrt{X_0(\bm{w})'S(\hat{\beta})X_0(\bm{w})}} S(\hat{\beta})X_0(\bm{w}). \nonumber
\end{equation}
For $t < 0$, we can obtain similar results. Hence, $\tilde{b}(t,\bm{w})$ has the following closed form expression:
\begin{equation}
\tilde{b}(t,\bm{w}) \ = \ \begin{cases}
  b^{+}_{\text{meta}}(\bm{w}) \ \equiv \ -X_0(\bm{w})'\hat{\beta} + \sqrt{ \chi(\alpha,d_x) \cdot X_0(\bm{w})'S(\hat{\beta})X_0(\bm{w})} & \text{if $t \geq 0$} \\
  b^{-}_{\text{meta}}(\bm{w}) \ \equiv \ X_0(\bm{w})'\hat{\beta} + \sqrt{ \chi(\alpha,d_x) \cdot X_0(\bm{w})'S(\hat{\beta})X_0(\bm{w})} & \text{if $t < 0$}
 \end{cases}. \nonumber
\end{equation}
This result makes the computation of $\hat{\bm{w}}_{\text{minimax}}(\alpha)$ easier. In this case, because $\mathcal{S} = \mathbb{R}$, we obtain
\begin{equation}
\hat{\bm{w}}_{\text{minimax}}(\alpha) \ \in \ \max \left\{ s(\bm{w}) \cdot \eta \left( \frac{b^{+}_{\mathrm{meta}}(\bm{w})}{s(\bm{w})} \right), s(\bm{w}) \cdot \eta \left( \frac{b^{-}_{\mathrm{meta}}(\bm{w})}{s(\bm{w})} \right) \right\}. \nonumber
\end{equation}
Hence, in this case, it is not difficult to compute $\hat{\bm{w}}_{\text{minimax}}(\alpha)$.

As shown above, when $\sigma_k \rightarrow 0$ for all $k$, $\hat{\delta}_{\bm{w}_{\mathrm{minimax}}}$ does not converge to $\delta_{IS}^*$. However, we can show that the refined minimax regret rule using $\hat{\mathcal{T}}_{\mathrm{HR}}(\alpha)$ converges to $\delta_{IS}^*$. When $\sigma_k \rightarrow 0$ for all $k$, the hyper-rectangle confidence region $\hat{\mathcal{T}}_{\mathrm{HR}}(\alpha)$ projected for $\tau_0$ converges to the identified set $IS_0$ defined in (\ref{identified_set}). Hence, in this case, if $\left\{ \sum_{k=1}^K w_k \hat{\tau}_k: \sum_{k=1}^K w_k = 1 \right\}$ includes both positive and negative values, that is, $\hat{\tau}_k$ is not the same for all $k$, then the refined minimax regret rule $\hat{\delta}_{\hat{\bm{w}}_{\mathrm{minimax}}(\alpha)}$ converges to $\delta_{IS}^*$.

\begin{Remark} \label{rem:IS_randomized}
We can show that some minimax regret rule does not converges to $\delta^{\ast}_{IS}$ even if we allow randomized rules. For example, consider $\mathcal{T} = \{ (\tau_0,\tau_1) \in \mathbb{R}^2 : |\tau_0 - \tau_1| \leq 1 \}$. Then, the maximum welfare regret with the sample analogue of the identified set plugged in is
\begin{eqnarray*}
\max_{\tau_0 \in IS_0} r(\tau_0,\delta) \ = \ \max\{ -(\hat{\tau}_1-1)\cdot \delta, (\hat{\tau}_1+1)\cdot (1-\delta) \},
\end{eqnarray*}
where $IS_0 = [\hat{\tau}_1 - 1, \hat{\tau}_1 + 1]$. Hence, the maximum welfare regret is minimized by the following treatment rule:
\[
\delta^{\ast}_{IS}(\hat{\tau}_1) \ = \ \begin{cases}
    1 & \text{if $\hat{\tau}_1 > 1$} \\
    \frac{\hat{\tau}_1 + 1}{2} & \text{if $-1 \leq \hat{\tau}_1 \leq 1$} \\
    1 & \text{if $\hat{\tau}_1 < -11$}
\end{cases}.
\]
On the other hand, \cite{stoye2012minimax} shows that when $\mathcal{T} = \{ (\tau_0,\tau_1) \in \mathbb{R}^2 : |\tau_0 - \tau_1| \leq 1 \}$, the minimax regret treatment rule is given by
\[
\hat{\delta}(\hat{\tau}_1) \ = \ \begin{cases}
    \bm{1}\{\hat{\tau}_1 \geq 0\} & \text{if $\sigma_1 \geq 2 \phi(0)$} \\
    \Phi\left( \hat{\tau}_1 / \sqrt{4 \phi(0)^2 - \sigma_1^2} \right) & \text{if $\sigma_1 < 2 \phi(0)$}
\end{cases}.
\]
\cite{yata2021optimal} also show the same result in more general settings. Therefore, as $\sigma_1 \to 0$, the minimax regret treatment rule converges to $\Phi\left( \hat{\tau}_1 / \sqrt{4 \phi(0)^2} \right)$. These results imply that the minimax regret treatment rule does not converge to $\delta^{\ast}_{IS}$ in this setting.\footnote{\cite{montiel2023decision} shows that the set of minimax regret optimal rules contains a piecewise linear rule and it converges to $\delta^{\ast}_{IS}$ as the variance of the signal approaches zero. Hence, there is a sequence of minimax regret treatment rules that converges to $\delta^{\ast}_{IS}$, but there is also a sequence of minimax regret treatment rules that does not converge to $\delta^{\ast}_{IS}$.}
\end{Remark}


\section{Numerical analysis}

To illustrate the results we obtain in the previous sections, we present some numerical analysis. Throughout this section, we set the study-specific covariates as equidistant grid points on $[0,1]$:
$$
x_k = (k-1)/(K-1) \ \ \text{and} \ \ \sigma_k = 1 \ \ \text{for $k=1, \cdots , K$.}
$$

We consider the following two parameter spaces:
\begin{eqnarray}
\mathcal{T}_1 &=& \{ \bm{\tau}: \tau_k = \beta_0 + \beta x_k,  \ \text{$\beta_0 \in \mathbb{R}$ and $\beta \in [-C,C]$}\}, \nonumber \\
\mathcal{T}_2 &=& \{ \bm{\tau}: |\tau_k - \tau_l | \leq C |x_k - x_l| \ \text{for any $k,l$} \}, \nonumber
\end{eqnarray}
where $C$ is a positive constant. For these two parameter spaces, we derive the minimax regret and minimax MSE rules and compare the maximum regrets of these two treatment rules.

In the case of the linear parameter space $\mathcal{T}_1$, one natural treatment rule is based on plugging in the OLS estimator,
$$
\hat{\delta}_{\bm{w}_{\text{OLS}}} \ \equiv \ \mathbf{1} \left\{ \tilde{x}_0'\hat{b} \geq 0 \right\},
$$
where $\tilde{x}_0 \equiv (1,x_0)'$ and $\hat{b}$ is the OLS estimator of $(\beta_0, \beta)'$. Because $\hat{b}$ is linear with respect to $(\hat{\tau}_1, \dots , \hat{\tau}_K)'$, this rule can be expressed as a linear aggregation rule, $\sum_{k=1}^K w_{\text{OLS},k} \hat{\tau}_k$.

Another natural treatment rule is based on plugging in the hierarchical Bayes (HB) estimator defined in Remark \ref{rem:Bayes},
$$
\hat{\delta}_{\bm{w}_{\text{HB}}} \ \equiv \ \mathbf{1} \left\{ \sum_{k=1}^K w_{\text{HB},k} \hat{\tau}_k \geq 0 \right\},
$$
where $\bm{w}_{\text{HB}} \equiv (w_{\text{HB},1}, \dots, w_{\text{HB},K})' \equiv \tilde{\bm{w}}_{\text{HB}} / \sum_{k=1}^K \tilde{w}_{\text{HB},k}$ and $\tilde{\bm{w}}_{\text{HB}} \equiv (\tilde{w}_{\text{HB},1}, \dots, \tilde{w}_{\text{HB},K})' \equiv \bm{\Sigma}^{-1} \bm{X} \left(\bm{\Sigma}_{\tilde{\beta}}^{-1} + \bm{X}' \bm{\Sigma}^{-1} \bm{X} \right)^{-1} \tilde{x}_0$. We set the prior variance matrix as
\[
\bm{\Sigma}_{\tilde{\beta}} \ = \ \left( \begin{array}{cc}
10 & 0 \\
0 & 1.96 \cdot C
\end{array} \right).
\]
Because the state space $\mathcal{T}_1$ does not restrict $\beta_0$, we specify a diffuse prior for $\beta_0$. In contrast, because $\mathcal{T}_1$ assumes that $\beta \in [-C,C]$, we assume that $\beta$ is contained in $[-C,C]$ with prior probability $0.95$.

We calculate $\bm{w}_{\text{OLS}}$, $\bm{w}_{\text{HB}}$, $\bm{w}_{\text{MSE}}$, and $\bm{w}_{\text{minimax}}$ for $K = 30$ and $x_0 = 0.1$. Table 1 contains the results of this experiment for $C = 0.1$,  $1.0$, and $2.0$. Table 1 shows the ratios of $b(\bm{w})$ and $s(\bm{w})$, and the ratios of the maximum regrets, that is,
$$
\cfrac{\max_{\bm{\tau} \in \mathcal{T}} R(\bm{\tau},\hat{\delta}_{\bm{w}_{\text{OLS}}})}{\max_{\bm{\tau} \in \mathcal{T}} R(\bm{\tau},\hat{\delta}_{\bm{w}_{\text{minimax}}})}, \ \ \cfrac{\max_{\bm{\tau} \in \mathcal{T}} R(\bm{\tau},\hat{\delta}_{\bm{w}_{\text{HB}}})}{\max_{\bm{\tau} \in \mathcal{T}} R(\bm{\tau},\hat{\delta}_{\bm{w}_{\text{minimax}}})}, \ \ \text{and} \ \ \cfrac{\max_{\bm{\tau} \in \mathcal{T}} R(\bm{\tau},\hat{\delta}_{\bm{w}_{\text{MSE}}})}{\max_{\bm{\tau} \in \mathcal{T}} R(\bm{\tau},\hat{\delta}_{\bm{w}_{\text{minimax}}})}.
$$
Because the OLS estimator is unbiased, the maximum bias of the OLS estimator is exactly zero. Hence, the ratio $b(\bm{w}_{\text{OLS}})/s(\bm{w}_{\text{OLS}})$ is exactly zero in all settings. For the HB rule, $b(\bm{w}_{\text{HB}})/s(\bm{w}_{\text{HB}})$ increases as $C$ increases. The ratio of $b(\bm{w}_{\text{minimax}})$ and $s(\bm{w}_{\text{minimax}})$ is smaller than that of $\bm{w}_{\text{MSE}}$ in all settings. This implies that the minimax regret criterion places more emphasis on the bias than the variance compared with the minimax MSE criterion. Table 1 shows that the maximum regret of the minimax MSE rule is about 40 percent greater than the minimax regret when $C=1.0$. When $C=0.1$, the maximum regrets of $\bm{w}_{\text{MSE}}$ and $\bm{w}_{\text{HB}}$ are close to the minimax regret. If $C$ is sufficiently large, $\bm{w}_{\text{minimax}}$ is almost the same as $\bm{w}_{\text{OLS}}$. Hence, when $C = 1.0$ or $2.0$, the maximum regret of $\hat{\delta}_{\bm{w}_{\text{OLS}}}$ is almost identical to the minimax regret. In contrast, when $C$ is small, $\bm{w}_{\text{minimax}}$ is quite different from $\bm{w}_{\text{OLS}}$ and the maximum regret of $\hat{\delta}_{\bm{w}_{\text{OLS}}}$ is about 30 percent greater than the minimax regret.
\begin{table}[H]
\caption{Results for the linear parameter space $\mathcal{T}_1$.}
\begin{center}
  \begin{tabular}{l c c c c c c c} \hline
      & $\frac{b(\bm{w}_{\text{OLS}})}{s(\bm{w}_{\text{OLS}})}$ & $\frac{b(\bm{w}_{\text{HB}})}{s(\bm{w}_{\text{HB}})}$ & $\frac{b(\bm{w}_{\text{MSE}})}{s(\bm{w}_{\text{MSE}})}$ & $\frac{b(\bm{w}_{\text{minimax}})}{s(\bm{w}_{\text{minimax}})}$ & ratio (OLS) & ratio (HB) & ratio (MSE) \rule[0mm]{0mm}{5mm} \vspace{1mm} \\ \hline \hline
$C=0.1$ & 0 & .131 & .213 & .170 & 1.30 & 1.01 & 1.02 \\
$C=1.0$ & 0 & .237 & .427 & .000 & 1.00 & 1.21 & 1.41 \\
$C=2.0$ & 0 & .248 & .237 & .000 & 1.00 & 1.29 & 1.28 \\  \hline
  \end{tabular}
  \end{center}
  \footnotesize{Note: The ratios that are shown in the final three columns are the ratios of the maximum regrets, $\max_{\bm{\tau} \in \mathcal{T}} R(\bm{\tau},\hat{\delta}_{\bm{w}}) / \max_{\bm{\tau} \in \mathcal{T}} R(\bm{\tau},\hat{\delta}_{\bm{w}_{\text{minimax}}})$, for $\bm{w} = \bm{w}_{\text{OLS}}$, $\bm{w}_{\text{HB}}$, and $\bm{w}_{\text{MSE}}$.}
\begin{center}
\end{center}
\end{table}

Next, we consider the Lipschitz parameter space $\mathcal{T}_2$. We calculate $\bm{w}_{\text{minimax}}$ and $\bm{w}_{\text{MSE}}$ for $K=30$ and $x_0 = 0.5$. Similar to $\mathcal{T}_1$, we consider the following hierarchical Bayesian model with the prior distribution,
\[
\bm{\tau} \ = \ (\tau_0, \bm{\tau}_{-0})' \ \sim \ \mathcal{N} \left( \bm{0}, \bm{\Sigma}_{\bm{\tau}} \right),
\]
where $\bm{\tau}_{-0} \equiv (\tau_1, \dots, \tau_K)'$ and
\[
\bm{\Sigma}_{\bm{\tau}} \ \equiv \ \left( \begin{array}{cc}
    \bm{\Sigma}_{\bm{\tau},11} & \bm{\Sigma}_{\bm{\tau},12}  \\
    \bm{\Sigma}_{\bm{\tau},21} & \bm{\Sigma}_{\bm{\tau},22}
\end{array} \right).
\]
We set the prior variance of $\tau_k$ as $10$ and the prior covariance of $\tau_k$ and $\tau_l$ as $10 \cdot \exp(-|x_k-x_l| / a)$ for some positive constant $a>0$. We choose a positive constant $a$ that satisfies $\frac{1}{K(K+1)/2} \sum_{k <l} P(|\tau_k - \tau_l| > C |x_k-x_l|) = 0.05$. Then, the posterior mean of $\tau_0$ is written as
\[
E[\tau_0 | \hat{\bm{\tau}}] \ = \ \bm{\Sigma}_{\bm{\tau}, 12} \bm{\Sigma}_{\bm{\tau}, 22}^{-1} \left(\bm{\Sigma}_{\bm{\tau},22}^{-1} + \bm{\Sigma}^{-1} \right)^{-1} \bm{\Sigma}^{-1} \hat{\bm{\tau}},
\]
which pins down the weights of the Bayes optimal decision rule $\bm{w}_{\text{HB}}$.

Table 2 shows the ratios of $b(\bm{w})$ and $s(\bm{w})$ and the ratios of the maximum regrets for $C = 0.1$,  $1.0$, and $2.0$. The ratio of $b(\bm{w}_{\text{minimax}})$ and $s(\bm{w}_{\text{minimax}})$ is smaller than the ratios of $\bm{w}_{\text{HB}}$ and $\bm{w}_{\text{MSE}}$ in all settings. When $C$ is small, the maximum regret of the minimax MSE rule nearly attains the minimax regret. In contrast, when $C$ is large, the maximum regret of the minimax MSE rule is about 17 percent greater than the minimax regret. Similar to the minimax MSE rule, the maximum regret of the hierarchical Bayes rule nearly attains the minimax regret when $C$ is small. In addition, it is about 30 percent greater than the minimax regret when $C$ is large. Figure \ref{fig:weights_Lip} shows $\bm{w}_{\text{minimax}}$, $\bm{w}_{\text{MSE}}$, and $\bm{w}_{\text{HB}}$ for $C=1.0$. It shows that the minimax treatment rule is quite different from other treatment rules. The minimax MSE and Bayes criteria gives positive weights to most of the studies. In contrast, the minimax regret criterion yields weights that sharply concentrate around zero and rule out half of the studies.
\begin{table}[H]
\caption{Results of the Lipschitz parameter space $\mathcal{T}_2$.}
\begin{center}
  \begin{tabular}{c c c c c c} \hline
      & $\frac{b(\bm{w}_{\text{HB}})}{s(\bm{w}_{\text{HB}})}$ & $\frac{b(\bm{w}_{\text{MSE}})}{s(\bm{w}_{\text{MSE}})}$ & $\frac{b(\bm{w}_{\text{minimax}})}{s(\bm{w}_{\text{minimax}})}$ & ratio (HB) & ratio (MSE) \rule[0mm]{0mm}{5mm} \vspace{1mm} \\ \hline \hline
$C=0.1$ & .141 & .141 & .131 & 1.01 & 1.01 \\
$C=1.0$ & .910 & .708 & .282 & 1.31 & 1.17 \\
$C=2.0$ & .822 & .710 & .284 & 1.28 & 1.17 \\  \hline
  \end{tabular}
  \end{center}
  \footnotesize{Note: The ratios that are shown in the final two columns are the ratios of the maximum regrets, $\max_{\bm{\tau} \in \mathcal{T}} R(\bm{\tau},\hat{\delta}_{\bm{w}}) / \max_{\bm{\tau} \in \mathcal{T}} R(\bm{\tau},\hat{\delta}_{\bm{w}_{\text{minimax}}})$, for $\bm{w} = \bm{w}_{\text{HB}}$ and $\bm{w}_{\text{MSE}}$.}
\end{table}

\begin{figure}[H]
\centering
\includegraphics[width=9cm]{weights_Lipschitz.pdf}
\caption{{\footnotesize The black, red, and blue dots denote $\bm{w}_{\text{minimax}}$, $\bm{w}_{\text{MSE}}$, and $\bm{w}_{\text{HB}}$ for $C=1.0$.}}\label{fig:weights_Lip}
\end{figure}

\if0
\begin{figure}[H]
\centering
\includegraphics[width=9cm]{w_minimax_Lipschitz_C=1.pdf}
\caption{The minimax regret weights, $\bm{w}_{\text{minimax}}$, for $C=1.0$.}
\end{figure}
\begin{figure}[H]
\centering
\includegraphics[width=9cm]{w_MSE_Lipschitz_C=1.pdf}
\caption{The minimax MSE weights, $\bm{w}_{\text{MSE}}$, for $C=1.0$.}
\end{figure}
\begin{figure}[H]
\centering
\includegraphics[width=9cm]{w_HB_Lipschitz_C=1.pdf}
\caption{The HB weights, $\bm{w}_{\text{HB}}$, for $C=1.0$.}
\end{figure}
\fi

\section{Empirical Applications}

We illustrate the use of our methods by means of two applications.
The first application considers whether an active labor market policy should be adopted, and the second application considers whether a COVID-19 treatment should be approved.

\subsection{Active Labor Market Policies}

We use the meta-database of \cite{card2017works}, which contains the estimates from over 200 recent studies of active labor market programs including training, subsidized employment, and job search assistance. We focus on papers that analyze RCT data to assess the impact of job training on the employment rate. This criterion reduces the meta-sample to 14 RCT estimates ($K=14$) collected from 8 different countries: Argentina, Brazil, Colombia, Dominica, Jordan, Nicaragua, Turkey, and the United States. Table 3 lists the papers included in the meta-sample of this application.

To form a vector of study characteristics $x_k$, $k=0,1,2,\dots,K$, we use five covariates that characterize the country and the sub-population on which the RCT study was performed. These are a gender dummy (male only $=$ 0, female only $=$ 0, mixed $=$ 0.5), an age dummy (age $<$ 25 only $=$ 1, age $\geq$ 25 only $=$ 0, both $=$ 0.5), an OECD dummy, the (standardized) GDP growth rate, and the (standardized) unemployment rate in 2010. Table 4 shows the estimates, standard errors, and study characteristics in this meta-sample.

We derive the minimax regret and minimax MSE rules with the following parameter space:
$$
\mathcal{T}_C \ \equiv \ \left\{ \bm{\tau} : |\tau_k - \tau_l| \leq C \|x_k-x_l\| \ \text{for $k,l = 0, 1, \cdots, K$}  \right\},
$$
with a prespecified Lipschitz constant $C \geq 0$. To determine $C$, we perform leave-one-out cross-validation with the study-average welfare criterion to obtain $C = 0.025$.


We consider whether the training program should be adopted in the following three target populations:
\begin{itemize}
\item Japan (in 2010): female, age $\geq 25$, OECD, $\Delta$GDP $=4.2$, unemp. $=5.1$
\item The UK (in 2010): female, age $\geq 25$, OECD, $\Delta$GDP $= 1.7$, unemp. $= 7.8$
\item Peru (in 2010): female, age $\geq 25$, not OECD, $\Delta$GDP $= 8.3$, unemp. $= 7.9$
\end{itemize}


\begin{table}[H]
\caption{The list of studies used in the job training program application.}
\begin{center}
\begin{tabular}{c l l}
  \hline
  No. & Paper & Country \\ \hline \hline
  1 & \cite{alzua2016long} & Argentina \\
  2 & \cite{attanasio2011subsidizing} & Colombia \\
  3 & \cite{calero2017can} & Brazil \\
  4 & \cite{fairlie2015behind} & United States \\
  5 & \cite{groh2012soft} & Jordan \\
  6 & \cite{hirshleifer2016impact} & Turkey \\
  7 & \cite{ibarraran2014life} & Dominica \\
  8 & \cite{macours2013demand} & Nicaragua \\
  \hline
\end{tabular}
\end{center}
\end{table}



\begin{table}[H]
\caption{Estimates, standard errors, and study characteristics.}
\begin{center}
\begin{tabular}{c c c c c c c c c}
  \hline
  No. & Country & Estimates & S.E. & Gender & Age & OECD & GDP & Unemployment \\ \hline \hline
  1a & Argentina & 0.245 & 0.073 & male & both & no & 9.02 & 7.5 \\
  1b & Argentina & 0.005 & 0.039 & female & both & no & 9.02 & 7.5 \\
  2a & Colombia & -0.025 & 0.022 & male & $<$ 25 & no & 4.71 & 11.3 \\
  2b & Colombia & 0.061 & 0.024 & female & $<$ 25 & no & 4.71 & 11.3 \\
  3 & Brazil  & 0.129 & 0.069 & both & $<$ 25 & no & 0.87 & 6.9 \\
  4 & US & 0.071 & 0.046 & both & both & yes & 3.31 & 5.6 \\
  5 & Jordan & 0.015 & 0.032 & female & $<$ 25 & no & 2.31 & 12.5 \\
  6a & Turkey & 0.069 & 0.029 & male & $\geq$ 25 & yes & 6.72 & 10.3 \\
  6b & Turkey & 0.023 & 0.020 & female & $\geq$ 25 & yes & 6.72 & 10.3 \\
  6c & Turkey & 0.011 & 0.025 & female & $<$ 25 & yes & 6.72 & 10.3 \\
  6d & Turkey & -0.021 & 0.030 & male & $<$ 25 & yes & 6.72 & 10.3 \\
  7a & Dominica & 0.007 & 0.024 & female & $<$ 25 & no & 5.49 & 13.8 \\
  7b & Dominica & -0.024 & 0.023 & male & $<$ 25 & no & 5.49 & 13.8 \\
  8 & Nicaragua & 0.038 & 0.020 & both & both & no & 4.36 & 5.5 \\
  \hline
\end{tabular}
\end{center}
\end{table}




\if0
\begin{figure}[H]
\centering
\includegraphics[width=12cm]{JPN_minimax.pdf}
\caption{{\footnotesize The minimax regret weights, $\bm{w}_{\text{minimax}}$, for Japan. The horizontal axis measures the Euclidean distance between $x_k$ and $x_0$, $\|x_k - x_0\|$. The size of the plotted circle is proportional to the precision of the estimates, i.e.,  a smaller $\hat{\sigma}_k$ corresponds to a larger circle. Estimates 4 and 8 receive positive weights.}}
\end{figure}
\begin{figure}[H]
\centering
\includegraphics[width=12cm]{JPN_mse.pdf}
\caption{{\footnotesize The minimax MSE weights, $\bm{w}_{\text{MSE}}$, for Japan. The horizontal axis measures the Euclidean distance between $x_k$ and $x_0$, $\|x_k - x_0\|$. The size of the plotted circle is proportional to the precision of the estimates, i.e.,  a smaller $\hat{\sigma}_k$ corresponds to a larger circle. Estimates 4 and 8 receive positive weights.}}
\end{figure}
\begin{figure}[H]
\centering
\includegraphics[width=12cm]{JPN_HB.pdf}
\caption{{\footnotesize The HB weights, $\bm{w}_{\text{HB}}$, for Japan. The horizontal axis measures the Euclidean distance between $x_k$ and $x_0$, $\|x_k - x_0\|$. The size of the plotted circle is proportional to the precision of the estimates, i.e.,  a smaller $\hat{\sigma}_k$ corresponds to a larger circle. }}
\end{figure}

\begin{figure}[H]
\centering
\includegraphics[width=12cm]{UK_minimax.pdf}
\caption{{\footnotesize The minimax regret weights, $\bm{w}_{\text{minimax}}$, for the UK. The horizontal axis measures the Euclidean distance between $x_k$ and $x_0$, $\|x_k - x_0\|$. The size of the plotted circle is proportional to the precision of the estimates, i.e.,  a smaller $\hat{\sigma}_k$ corresponds to a larger circle. Estimates 3 and 4 receive positive weights.}}
\end{figure}
\begin{figure}[H]
\centering
\includegraphics[width=12cm]{UK_mse.pdf}
\caption{{\footnotesize The minimax MSE weights, $\bm{w}_{\text{MSE}}$, for the UK. The horizontal axis measures the Euclidean distance between $x_k$ and $x_0$, that is, $\|x_k - x_0\|$. The size of the plotted circle is proportional to the precision of the estimates, i.e.,  a smaller $\hat{\sigma}_k$ corresponds to a larger circle. Estimates 3, 4, and 8 receive positive weights.}}
\end{figure}
\begin{figure}[H]
\centering
\includegraphics[width=12cm]{UK_HB.pdf}
\caption{{\footnotesize The HB weights, $\bm{w}_{\text{HB}}$, for the UK. The horizontal axis measures the Euclidean distance between $x_k$ and $x_0$, that is, $\|x_k - x_0\|$. The size of the plotted circle is proportional to the precision of the estimates, i.e.,  a smaller $\hat{\sigma}_k$ corresponds to a larger circle.}}
\end{figure}

\begin{figure}[H]
\centering
\includegraphics[width=12cm]{PER_minimax.pdf}
\caption{{\footnotesize The minimax regret weights, $\bm{w}_{\text{minimax}}$, for Peru. The horizontal axis measures the Euclidean distance between $x_k$ and $x_0$,  $\|x_k - x_0\|$. The size of the plotted circle is proportional to the precision of the estimates, i.e.,  a smaller $\hat{\sigma}_k$ corresponds to a larger circle. Estimates 1a and 1b receive positive weights.}}
\end{figure}
\begin{figure}[H]
\centering
\includegraphics[width=12cm]{PER_mse.pdf}
\caption{{\footnotesize The minimax MSE weights, $\bm{w}_{\text{MSE}}$, for Peru. The horizontal axis measures the Euclidean distance between $x_k$ and $x_0$,  $\|x_k - x_0\|$. The size of the plotted circle is proportional to the precision of the estimates, i.e.,  a smaller $\hat{\sigma}_k$ corresponds to a larger circle. Estimates 1a, 1b, and 6b receive positive weights.}}
\end{figure}
\begin{figure}[H]
\centering
\includegraphics[width=12cm]{PER_HB.pdf}
\caption{{\footnotesize The HB weights, $\bm{w}_{\text{HB}}$, for Peru. The horizontal axis measures the Euclidean distance between $x_k$ and $x_0$,  $\|x_k - x_0\|$. The size of the plotted circle is proportional to the precision of the estimates, i.e.,  a smaller $\hat{\sigma}_k$ corresponds to a larger circle. }}
\end{figure}
\fi

\begin{figure}[htbp]
    \begin{tabular}{ccc}
      \begin{minipage}[t]{0.45\hsize}
        \centering
        \includegraphics[width=6.3cm]{weights_JPN.pdf}
        \subcaption{{\footnotesize Japan.}}
        \label{fig:JPN}
      \end{minipage} &
      \begin{minipage}[t]{0.45\hsize}
        \centering
        \includegraphics[width=6.3cm]{weights_UK.pdf}
        \subcaption{{\footnotesize The UK.}}
        \label{fig:UK}
      \end{minipage} \\

      \begin{minipage}[t]{0.45\hsize}
        \centering
        \includegraphics[width=6.3cm]{weights_PER.pdf}
        \subcaption{{\footnotesize Peru.}}
        \label{fig:PER}
      \end{minipage} &

    \end{tabular}
    \caption{{\footnotesize The black, red, and blue circles denote $\bm{w}_{\text{minimax}}$, $\bm{w}_{\text{MSE}}$, and $\bm{w}_{\text{HB}}$, respectively. The horizontal axis measures the Euclidean distance between $x_k$ and $x_0$, $\|x_k - x_0\|$. The size of the plotted circle is proportional to the precision of the estimates, i.e.,  a smaller $\hat{\sigma}_k$ corresponds to a larger circle.}}\label{fig:weights_ALMP}
\end{figure}

\begin{table}[H]
\caption{The minimax regret, minimax MSE, and hierarchical Bayes rules.}
\begin{center}
\begin{tabular}{l c c c}
  \hline
  & Japan & UK & Peru \\ \hline \hline
  $\hat{\tau}_0$ with $\bm{w}_{\text{minmax}}$ & 0.059 & 0.078 & 0.018 \\
  $\hat{\tau}_0$ with $\bm{w}_{\text{MSE}}$ & 0.046 & 0.059 & 0.031 \\
  $\hat{\tau}_0$ with $\bm{w}_{\text{HB}}$ & 0.029 & 0.026 & 0.025 \\
  ratio (MSE) & 1.107 & 1.215 & 1.180 \\
  ratio (HB) & 3.131 & 2.580 & 3.356 \\
  $w_{\text{minmax},k} \geq 1/K$ & Nicaragua, US  & Brazil, US & Argentina \\ \hline
\end{tabular}
\end{center}
\end{table}

Figures \ref{fig:JPN}--\ref{fig:PER} plot $\bm{w}_{\text{minimax}}$, $\bm{w}_{\text{MSE}}$, and $\bm{w}_{\text{HB}}$ for the three different target populations. Similar to Section 4, the hierarchical Bayes rule uses the following prior:
\[
\bm{\tau} \ \sim \ \mathcal{N}\left( \bm{0}, \bm{\Sigma}_{\bm{\tau}} \right),
\]
where $\bm{\Sigma}_{\bm{\tau}}[k,l] = \exp (-\|x_{k-1} - x_{l-1}\| / a)$ and we choose $a$ satisfying $\frac{1}{K(K+1)/2} \sum_{k < l} P(|\tau_k - \tau_l| > C \|x_{k} - x_{l}\|) = 0.05$. The horizontal axis measures the Euclidean distance between $x_k$ and $x_0$. The size of the plotted circle is proportional to the precision of the estimates, i.e., a smaller $\hat{\sigma}_k$ corresponds to a larger circle. The figures show that, overall, both $\bm{w}_{\text{minimax}}$ and $\bm{w}_{\text{MSE}}$ tend to put greater weight on those studies that are in similar in terms of their population characteristics. This tendency is more evident for the minimax regret weights $\bm{w}_{\text{minimax}}$ than for the minimax MSE weights $\bm{w}_{\text{MSE}}$.


We note that $\bm{w}_{\text{minimax}}$ differs from $\bm{w}_{\text{MSE}}$ for every target population. In all cases, the minimax regret criterion puts the most weight on the closest study. In contrast, the minimax MSE criterion can put the largest weight on a study that is not closest provided that it has a small standard error. For instance, in the case of Japan, the minimax regret weight of the closest study is more than 0.6 but the minimax MSE weight is about 0.3. These results reflect the different degrees of bias variance trade-off that the minimax regret and minimax-MSE weights aim to balance out, as discussed in Section 3.1.

Table 5 lists $\hat{\tau}_0(\bm{w}_{\text{minimax}})$, $\hat{\tau}_0(\bm{w}_{\text{MSE}})$, $\hat{\tau}_0(\bm{w}_{\text{HB}})$, the ratio of the maximum regrets, and the countries that were awarded minimax regret weights larger than $1/K = 1/14$. The table also shows that the minimax regret and minimax MSE rules select different decisions in some cases. For example, the average annual salary amongst Japanese women aged 25--29 years is approximately \$30,000; if the cost per person of adopting the policy is \$1,500 and individuals that start a new job work for one year, we could set $c_0 = 0.05$.\footnote{According to \cite{fairlie2015behind}, the cost per person of implementing the program in the United States is \$1,321.} Then, the recommendation of the minimax regret criterion is to introduce the policy in Japan. However, the minimax MSE criterion does not recommend the introduction of the policy in Japan.

For all of the target populations, the maximum regret of the minimax MSE rule is more than 10 percent greater than that of the minimax regret. For Japan and the UK, the minimax regret aggregation rule puts the most weight on the estimates of the US. In contrast, for Peru, the minimax regret criterion puts most of the weight on one estimate obtained from Argentina.



\subsection{Treatments for COVID-19}

We consider a drug approval decision for a COVID-19 treatment using the meta-database of randomized clinical trials provided by \citet{juul2020interventions}. There is an urgent gloabl need for evidence-based treatment of COVID-19. To search for effective treatments, numerous randomized clinical trials have been conducted in different countries across different demographic groups. At the time of writing, evidence as to the efficacy of various proposed treatments is mixed with limited precision of trial estimates.

We focus on Remdesivir, an antiviral medication known to be effective against viruses in the coronavirus family, such as Middle East Respiratory Syndrome (MERS) and Severe Acute Respiratory Syndrome (SARS). The effectiveness of Remdesivir in fighting Covid-19, however, remains undetermined and is a source of much controversy due to the conflicting nature of the existing evidence; the U.S. Food and Drug Administration approved emergency use of Remdesivir for Covid-19 patients, while the World Health Organization recommends against its use.

The meta-database of \cite{juul2020interventions} collates data from 33 RCT studies enrolling a total of 13,312 participants. Each study provides an estimate of the treatment effect compared with standard care or a placebo. Because we focus on the effects of Remdesivir on mortality rate, we use 6 estimates from 4 RCTs, \cite{beigel2020remdesivir}, \cite{who2021repurposed}, \cite{spinner2020effect}, and \cite{wang2020remdesivir}.

To form a vector of study characteristics, we include two covariates summarizing the average patients' characteristics in each study. They are the (standardized) mean (or median) age and the (standardized) proportion of female patients. Table 6 lists the estimates, standard errors, and study characteristics.\footnote{Studies 2a--2c report the subgroup treatment effect estimates for three age subgroups ($<$50, 50--69, $\leq$70). Because we do not have detailed information about the age of these subgroups, we suppose that the mean age of these subgroups are 45, 60, and 75, respectively.}

Similar to the previous section, we derive the minimax regret and minimax MSE rules with the following parameter space:
$$
\mathcal{T}_C \ \equiv \ \left\{ \bm{\tau} : |\tau_k - \tau_l| \leq C \|x_k-x_l\| \ \text{for $k,l = 0, 1, \cdots, K$}  \right\},
$$
with a prespecified Lipschitz constant $C \geq 0$. Because $K$ is small, leave-one-out cross-validation does not seem sensible. Hence, in this application, we set $C=0.01$ based on the WLS estimates. We consider hypothetical populations of interest whose characteristics range over 40--80 in terms of average age and 0.34--0.41 for the fraction of female patients.

\begin{table}[H]
\caption{Estimates, standard errors, and study characteristics.}
\begin{center}
\begin{tabular}{c c c c c c}
  \hline
  No. & Paper & Estimates & S.E. & Age & Female (\%) \\ \hline \hline
  1  & \cite{beigel2020remdesivir} &  0.038 & 0.021 & 58.9 & 0.356 \\
  2a & \cite{who2021repurposed} & -0.002 & 0.011 & 45.0 & 0.371 \\
  2b & \cite{who2021repurposed} &  0.005 & 0.013 & 60.0 & 0.371 \\
  2c & \cite{who2021repurposed} &  0.005 & 0.024 & 75.0 & 0.371 \\
  3  & \cite{spinner2020effect} &  0.030 & 0.021 & 57.0 & 0.389 \\
  4  & \cite{wang2020remdesivir} & -0.011 & 0.047 & 65.3 & 0.407 \\
  \hline
\end{tabular}
\end{center}
\end{table}



Figure \ref{fig:covid} plots $\hat{\tau}_0(\bm{w}_{\text{minimax}})$ at each specified grid point in the space of characteristics of the target populations. In this figure, dark red means the value of $\hat{\tau}_0(\bm{w}_{\text{minimax}})$ is large and white means the value is small. Figure \ref{fig:covid} also plots covariate values of the meta-sample and the size of the plotted circle is proportional to the precision of the estimates, i.e.,  a smaller $\hat{\sigma}_k$ corresponds to a larger circle. This result implies that the minimax regret criterion recommends to treat Remdesivir for the age group greater than 50 years-old. In contrast, for some local populations below age 50 years-old, the minimax regret criterion recommends not to treat them with Remdesivir.

\begin{figure}[H]
\centering
\includegraphics[width=9.5cm]{heatmap_covid.pdf}
\caption{{\footnotesize Heatmap of $\hat{\tau}_0(\bm{w}_{\text{minimax}})$ based on the values of $\hat{\tau}_0(\bm{w}_{\text{minimax}})$ computed at each $x_0$. Dark red areas correspond to regions of $x_0$ such that $\hat{\tau}_0(\bm{w}_{\text{minimax}})$ is positive and large. The white area corresponds to the region of $x_0$ such that $\hat{\tau}_0(\bm{w}_{\text{minimax}})$ is negative or near zero. The grey line is the boundary that separates the regions of positive and negative $\hat{\tau}_0(\bm{w}_{\text{minimax}})$. We also plot the covariate values of the meta-sample with the sizes of the plotted circles being proportional to the precision of the estimates.}}\label{fig:covid}
\end{figure}

\section{Conclusion}
Motivated by the recently proposed paradigm of `patient-centered meta-analysis' (\citet{manski2020towards}), this paper develops a method to aggregate available evidence and inform optimal treatment choice for a target population that is of interest to the planner. Building upon the framework of statistical decision theory and adopting the minimax regret criterion, we obtain a minimax regret treatment choice rule that is simple to implement in practice. The key steps of our analysis that deliver analytical and computational tractability are to
constrain decision rules to the class of linear aggregation rules and to restrict the
parameter space to a symmetric and invariant one (Assumption \ref{parameter_space}). These conditions for the parameter space are mild and hold in numerous contexts.

Several questions remain unanswered. First, when $\bm{\tau}$ is constrained to Lipschitz vectors while the Lipschitz constant $C$ is unknown, we do not know what is a theoretically justifiable data-driven way to select $C$. In the presented empirical applications, we selected $C$ by cross validation and WLS estimation without any analytical justification for this choice. Second, our framework assumes away any publication bias of published estimates despite a growing concern in the scientific community about this, and increasing interest in both how to detect, and correct for, any such bias (see, e.g., \citet{AK19}). Third, other than the standard errors of the estimates, our framework does not offer any way to incorporate a measure of the credibility of reported estimates. Depending on how the data were sampled and what identifying assumptions the estimate relies on, the credibility of studies can vary greatly. How to incorporate a measure of the credibility of reported estimates beyond their standard errors remains an interesting open question.
\setcounter{equation}{0}