EconBase
← Back to paper

Optimal Decision Rules Under Partial Identification

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

223,683 characters · 30 sections · 135 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Optimal Decision Rules Under Partial Identification

abstractI consider a class of statistical decision problems in which the policymaker must decide between two policies to maximize social welfare (e.g., the population mean of an outcome) based on a finite sample. The framework introduced in this paper allows for various types of restrictions on the structural parameter (e.g., the smoothness of a conditional mean potential outcome function) and accommodates settings with partial identification of social welfare. As the main theoretical result, I derive a finite-sample optimal decision rule under the minimax regret criterion. This rule has a simple form, yet achieves optimality among all decision rules; no ad hoc restrictions are imposed on the class of decision rules. I apply my results to the problem of whether to change an eligibility cutoff in a regression discontinuity setup, and illustrate them in an empirical application to a school construction program in Burkina Faso. \begin{comment} I consider a class of statistical decision problems in which the policy maker must decide between two alternative policies to maximize social welfare based on a finite sample. The central assumption is that the underlying, possibly infinite-dimensional parameter, lies in a known convex set, potentially leading to partial identification of the welfare effect. An example of such restrictions is the smoothness of counterfactual outcome functions. As the main theoretical result, I derive a finite-sample, exact minimax regret decision rule within the class of all decision rules under normal errors with known variance. When the error distribution is unknown, I obtain a feasible decision rule that is asymptotically minimax regret. I apply my results to the problem of whether to change a policy eligibility cutoff in a regression discontinuity setup, and illustrate them in an empirical application to a school construction program in Burkina Faso. \end{comment} \\ \\ {\it Keywords:} Statistical decision theory, finite-sample minimax regret, partial identification, nonparametric regression models, regression discontinuity.

\sloppy

Introduction

A fundamental goal of empirical research in economics is to inform policy decisions. Evaluation of counterfactual policies often requires extrapolating from observables to unobservables. Without strong model restrictions such as functional form assumptions, the performance of each counterfactual policy may be only partially determined by observed data. In such situations, policy decision-making is challenging, since we have no clear understanding of which policy is the best. For example, a regression discontinuity (RD) design only credibly estimates the impact of treatment on individuals at the eligibility cutoff. Therefore, without restrictive assumptions such as constant treatment effects, whether to offer the treatment to those away from the cutoff is ambiguous. Even randomized controlled trials may provide only partial knowledge of the impact of a new intervention, as can happen if the experimental sample is an unrepresentative subset of the target population.

This paper studies the problem of using data to make policy decisions in settings in which social welfare under each policy may be only partially identified. Following the literature on statistical treatment choice Manski2004hetero, I formulate the policy decision problem as a statistical decision problem. The framework introduced in this paper allows for various types of restrictions on the structural parameter, which potentially leads to partial identification of social welfare. It builds on donoho1994's donoho1994 framework for optimal estimation in nonparametric regression models, which has recently been applied to estimation and inference on treatment-effect parameters armstrong2018optimal,Imbens2019RDD,kwon2020rd,Armstrong2021ATE,rambachan2023parallel,Chaisemartin2021aet. I extend the framework to study optimal policy choice in a wide range of empirical settings with partial identification. Examples of this paper's framework include treatment choice using experiments with imperfect internal or external validity Stoye2012minimax,ishihara2021meta; treatment choice using observational data under unconfoundedness with imperfect overlap; and policy adoption choice in difference-in-differences designs without exact parallel trends.

Specifically, in the setup described in Section (ref), the policymaker must decide between two policies, policy 1 and policy 0, to maximize social welfare. The difference in welfare between the two policies is given by $L(\theta)$, where $\theta$ is a possibly infinite-dimensional structural parameter that resides in a vector space $\mathbb{V}$, and $L:\mathbb{V}\rightarrow\mathbb{R}$ is a known linear function. If $\theta$ is known, it is optimal to choose policy 1 if $L(\theta)\ge 0$ and policy 0 if $L(\theta)<0$. The policymaker does not know $\theta$, but instead has access to a multivariate Gaussian sample $\boldsymbol Y=(Y_1,...,Y_n)'\in\mathbb{R}^n$ of the form

align*[align* omitted — 87 chars of source]

where $\boldsymbol{m}:\mathbb{V}\rightarrow\mathbb{R}^n$ is a known linear function and $\boldsymbol\Sigma$ is known. After observing $\boldsymbol{Y}$, the policymaker decides between policies 1 and 0. The main structural assumption is that $\theta$ belongs to a known set $\Theta\subset\mathbb{V}$ that is convex and centrosymmetric (i.e., $\theta\in\Theta$ implies $-\theta\in\Theta$), which encodes the policymaker's a priori knowledge of parameter restrictions. Depending on the restrictions, the welfare contrast $L(\theta)$ may or may not be point identified from the knowledge of the point-identified reduced-form parameter $\boldsymbol m(\theta)$.

As detailed in Section (ref), an example of this setup is the choice between assigning treatment to everyone in the population (policy 1) and assigning treatment to no one (policy 0). Suppose that the policymaker has access to data from a regression model $Y_i=f(x_i,d_i)+U_i$, where $f(x,d)$ represents the conditional mean potential outcome under treatment $d\in\{0,1\}$ given covariates $x$, and $\{(x_i,d_i)\}_{i=1}^n$ is treated as fixed. Suppose further that treating everyone is preferred to treating no one if the population average treatment effect is positive. This problem is a special case when we assume $(U_1,...,U_n)'\sim {\cal N}(\boldsymbol{0}, \boldsymbol\Sigma)$ and set $\boldsymbol Y=(Y_1,...,Y_n)'$, $\theta=f$, $\boldsymbol{m}(f)=(f(x_1,d_1),...,f(x_n,d_n))'$, and $L(f)=\int[f(x,1)-f(x,0)]dP_X$, where $P_X$ is the population distribution of the covariates and is assumed to be known. The parameter space $\Theta$ is a class of conditional mean potential outcome functions $f$ that satisfy, for example, some smoothness restrictions (e.g., bounds on derivatives or the linearity of a function), which leads to either point or partial identification of the average treatment effect $L(f)$.

The main contribution of this paper is to obtain a finite-sample optimal decision rule under the minimax regret criterion, which is a standard criterion used in the literature on statistical treatment choice Manski2004hetero,Manski2007missing,Stoye2009minimax,Stoye2012minimax,Kitagawa2018EWM. A decision rule is a mapping from the sample $\boldsymbol Y$ to the probability of choosing policy 1. The minimax regret criterion evaluates the performance of a decision rule based on its worst-case expected welfare loss, or {\it regret}, relative to the oracle welfare-maximizing policy $\mathbf{1}\{L(\theta)\ge 0\}$, where the worst-case scenario is considered over the parameter space $\Theta$. I derive a decision rule that minimizes the worst-case regret among all decision rules, with no functional-form restrictions imposed on the class of rules. When $\boldsymbol Y$ is non-Gaussian and/or its variance is unknown, a feasible version of this decision rule can be constructed by plugging in an estimated variance. Appendix (ref) provides conditions under which its maximum regret over a class of distributions of $\boldsymbol Y$ converges to that of a minimax regret rule as $n\rightarrow\infty$.

To solve the minimax regret problem, I use the {\it hardest one-dimensional subfamily} argument, which donoho1994 used to solve minimax affine estimation problems. The key idea is to search for the hardest one-dimensional subproblem---specifically, the one with the highest minimax risk among all subproblems with parameter spaces defined by one-dimensional linear subfamilies of the original parameter space $\Theta$. Then, verify that a minimax rule for the hardest one-dimensional subproblem is also minimax optimal for the original problem. Applying this strategy to the minimax regret problem is challenging due to the difference in the structure of the risk functions: Unlike standard risk functions for estimation, such as mean squared error (MSE), the regret cannot be decomposed into the bias and variance; instead, it can be decomposed into the probability of misidentifying the best policy and the potential welfare loss due to misidentification. To derive a minimax regret rule, I first characterize the hardest one-dimensional subproblem by optimizing a certain measure of the strength of the signal with respect to the best policy within subproblems (Lemma (ref)). The subproblem with the optimal level of the signal strength achieves the best balance between the probability of misidentification and the potential welfare loss. I then propose a specific minimax regret rule for the hardest subproblem, and prove its minimax regret optimality for the original problem (Theorem (ref)).

The results of this paper provide novel insights into how a minimax regret rule uses data to make decisions. First, the derived rule depends on the observations $Y_1,...,Y_n$ only through their weighted sum, $\sum_{i=1}^n{w}_iY_i$, although no such restrictions are imposed a priori. The weights ${w}_1,...,{w}_n$ can be calculated by solving a sequence of convex optimization problems, which is computationally and analytically tractable in leading examples.\footnote{Thus, this paper addresses a challenge in the application of statistical decision theory raised by Manski2020decision, who wrote (p. 2848): “The primary challenge to use of statistical decision theory is computational \ldots\ Future advances should continue to expand the scope of applications.”}

Second, the minimax regret rule is nonrandomized or randomized, depending on the strength of the parameter restrictions and the variance of $\boldsymbol{Y}$. Specifically, if the restrictions are strong or the variance of $\boldsymbol{Y}$ is large in a certain formal sense, the minimax regret rule is a nonrandomized threshold rule of the form $\mathbf{1}\left\{\sum_{i=1}^n{w}_iY_i\ge 0\right\}$. On the other hand, if the restrictions are weak or the variance of $\boldsymbol{Y}$ is small, the rule is a randomized threshold rule of the form $\mathbf{1}\left\{\sum_{i=1}^n{w}_iY_i+\xi\ge 0\right\}$, where $\xi$ is generated independent of $\boldsymbol{Y}$ according to a certain distribution. In the latter case, randomization plays the role of reducing the probability of misidentification under worst-case parameter values, and thereby leads to a reduction in worst-case regret. This result generalizes Stoye2012minimax's Stoye2012minimax from a specific univariate problem to a general class of multivariate problems.

Third, this paper sheds light on the connection between minimax regret treatment choice and minimax estimation. The weighted sum $\sum_{i=1}^n{w}_iY_i$ used by the minimax regret rule can be viewed as an estimator of the welfare contrast $L(\theta)$. In other words, the minimax regret rule can be viewed as a {\it plug-in} rule, which plugs the estimator $\sum_{i=1}^n{w}_iY_i$ (plus a random noise $\xi$ for the randomized rule) into the oracle optimal decision $\mathbf{1}\{L(\theta)\ge 0\}$. I show that this estimator is optimal in the sense of minimizing the worst-case squared bias (over $\Theta$) among all estimators of $L(\theta)$ subject to a certain bound on the variance (Theorem (ref)). Furthermore, I show that this estimator places more importance on bias than variance compared with the minimax affine MSE estimator in donoho1994 (Theorem (ref)).

Fourth, while this paper's main focus is on optimal rules under partial identification, my results are novel even under point identification for problems with restricted parameter spaces. When the welfare contrast $L(\theta)$ is point identified, the minimax regret rule is shown to always be a nonrandomized threshold rule, with its form depending on the strength and type of restrictions. For example, consider a linear regression model in which $\boldsymbol{m}(\theta)=\boldsymbol{X}\theta$ for some fixed $n\times k$ design matrix $\boldsymbol{X}$, and suppose $\Theta=\{\theta\in\mathbb{R}^k: \|\theta\|_p\le C\}$ for some known constants $C\ge 0$ and $p\ge 1$, where $\|\cdot\|_p$ denotes the $L_p$-norm. The results of this paper imply that the minimax regret rule makes decisions based on the sign of $L(\hat\theta)$, where $\hat\theta$ is an estimator for $\theta$ that resolves the bias-variance tradeoff in a certain way (e.g., the ridge estimator with the regularization parameter depending on $C$ for $p=2$). The result of hirano2009asymptotics applies to problems with no parameter restrictions (i.e., $\Theta=\mathbb{R}^k$), but does not apply to ones with restricted parameter spaces.

commentA leading example of this setup is a choice between two treatment assignment policies based on data generated by nonparametric regression models, including an RD model. A treatment assignment policy specifies who would receive treatment based on an individual's observable covariates. In this example, the parameter $\theta$ represents a conditional mean function of a counterfactual outcome given covariates and treatment. $m_i(\theta)$ is the value of the conditional mean function evaluated at individual $i$'s observed covariates and treatment. The welfare difference $L(\theta)$ corresponds to, for example, the average treatment effect for the subpopulation that would be affected by the switch from the status quo to a new policy. The parameter space $\Theta$ is a class of conditional mean functions that satisfy, for example, some smoothness restrictions (e.g., bounds on derivatives or the linearity of a function). The welfare difference $L(\theta)$ may or may not be point identified, depending on which function class the policy maker imposes.\footnote{In Section (ref), I will discuss what I mean by identification in this finite-sample setup.}
commentSpecifically, in a simplified case where $Y_i$'s are independent with homoskedastic variance $\sigma^2$, I show that the rule of the following form is optimal: \begin{align*} \delta^*(\boldsymbol Y)=\begin{cases}\mathbf{1}\left\{\sum_{i=1}{w}^*_iY_i\ge 0\right\} & if s^* < \sigma,\\ \mathbf{1}\left\{\sum_{i=1}{w}^*_iY_i+\xi\ge 0\right\} & if s^* \ge \sigma, \end{cases} \end{align*} where $\boldsymbol{w}^*=(w_1^*,...,w_n^*)'$ is a unit vector that depends on $(L,\boldsymbol{m},\Theta,\sigma)$, $s^*\ge 0$ depends on $(L,\boldsymbol{m},\Theta)$, and $\xi\sim {\cal N}(0,(s^*)^2-\sigma^2)$ is a noise independent of the sample used to induce randomization. This result provides several insights on the minimax regret rule. \begin{itemize} • The rule depends on the observations $Y_1,...,Y_n$ only through a linear combination of them, even though no such restrictions are imposed when solving the minimax problem. • A main component that determines the value of $s^*$ is the width of the identified set of the welfare contrast $L(\theta)$ when the reduced-form parameter $m_i(\theta)$ is zero for every $i$. When the width is sufficiently small relative to $\sigma$, • $s^*$ determines whether to randomize the policy choice. $s^*$ depends on the ratio of the two parameters: the length of the identified set of the welfare contrast $L(\theta)$ when the reduced-form parameter $m_i(\theta)$ is zero for every $i$, and the derivative of the upper bound on $L(\theta)$ with respect to the reduced-form parameter in a certain direction (specifically, direction of steepest ascent) (sensitivity of the upper bound to the reduced-form parameter). • How to think about randomized rules in practice? Having a randomized rule suggests that the policy effect is so uncertain, and potential expected regret is too large. To avoid this, the policy maker needs to impose more restrictions of the parameter, or consider a less radical policy as the new policy. \end{itemize} The main tool that I use to solve the minimax regret problem is what is called the {\it modulus of continuity}, defined in Section (ref). It was originally introduced by donoho1994 to characterize minimax affine nonrandomized estimation and inference procedures, such as a minimax affine mean squared error (MSE) estimator, for a linear functional of a nonparametric regression function. I show that the minimax {\it regret} problem over the class of {\it all} decision rules can be simplified into an optimization problem with respect to the modulus of continuity. The optimization problem is analytically and computationally tractable. \begin{itemize} • The resulting decision rule is simple and thus easy to compute. It makes a decision based on a linear function of $\boldsymbol Y$. The minimax regret rule may be randomized or nonrandomized, depending on the restrictions imposed on the parameter space. More specifically, the rule is nonrandomized if the identified set of the welfare contrast $L(\theta)$ is, in a certain formal sense, small relative to the variance of the sample $\boldsymbol{Y}$, including the case where $L(\theta)$ is point identified. Otherwise, it is a randomized rule, assigning a positive probability both to policies 1 and 0. • When the minimax regret rule is nonrandomized, it can be viewed as a rule that plugs a particular linear estimator of the welfare contrast $L(\theta)$ into the optimal decision $\mathbf{1}\{L(\theta)\ge 0\}$. I compare this linear estimator with a linear minimax MSE estimator of $L(\theta)$, which minimizes the maximum of the MSE over the parameter space within the class of all linear estimators. The two estimators are shown to be generally different, which suggests that the plug-in rule based on the linear minimax MSE estimator is not optimal under the minimax regret criterion. More precisely, the linear estimator used by the minimax regret rule places more importance on the bias than on the variance compared to the linear minimax MSE estimator. \end{itemize}
commentThis paper makes new contributions even under point identification in settings with restricted parameter spaces. When the welfare difference $L(\theta)$ is point identified, the minimax regret rule is insensitive to the choice of the restrictions imposed on the parameter space as long as the restrictions are weak enough. For example, consider linear regression models where $m_i(\theta)=x_i'\theta$, $x_i\in\mathbb{R}^k$ is unit $i$'s fixed regressors, and $\theta\in\Theta\subset\mathbb{R}^k$. The minimax regret rule bases decisions on the sign of $L(\hat\theta)$, where $\hat\theta$ is the best linear unbiased estimator of $\theta$, if the parameter space $\Theta$ is sufficiently large (e.g., if $\Theta=\mathbb{R}^k$). When the restrictions on $\Theta$ become strong enough, the minimax regret rule starts to use an estimator that optimally trades off the bias and variance.
commentThis paper's result provides the following practical implications on policy decisions. \begin{itemize} • Can we say the result highlights the importance of identification over sample size? How does the minimax regret change as $\sigma$ decreases? • How to think about randomized rules in practice? Having a randomized rule suggests that the policy effect is so uncertain, and potential regret is too large. To avoid this, the policy maker needs to impose more restrictions of the parameter, or consider a less radical policy as the new policy. • Any implications on the sample/research design? \end{itemize}
comment\begin{itemize} • As the main application, I consider the problem of eligibility cutoff in an RD setup, described briefly below and in detail in Section (ref). Suppose that, for each unit $i=1,...,n$, we observe a running variable $x_i\in\mathbb{R}$, a binary treatment status $d_i=\mathbf{1}\{x_i\ge c_0\}$, and an outcome $Y_i\in\mathbb{R}$, where $c_0\in\mathbb{R}$ is the status quo cutoff determining the eligibility for treatment. $Y_i$ is determined by $Y_i=f(x_i,d_i)+u_i$, where $f:\mathbb{R}\times\{0,1\}\rightarrow \mathbb{R}$ represents the conditional mean counterfactual outcome function given the running variable and treatment, and $u_i$'s are independent mean-zero errors. We are interested in using the data $\{(Y_i,x_i,d_i)\}_{i=1}^n$ to decide whether to change the cutoff from $c_0$ to a specific value $c_1$. The welfare under $c_a$, $a\in\{0,1\}$, is, for example, an average counterfactual outcome across individuals: $$ W_a(f)=\int[f(x,1)\mathbf{1}\{x\ge c_a\}+f(x,0)\mathbf{1}\{x< c_a\}]d\nu(x) $$ for some known measure $\nu$. It is optimal to choose the new cutoff $c_1$ if $W_1(f)\ge W_0(f)$ and to maintain the status quo if $W_1(f)< W_0(f)$. While such an optimal choice is infeasible as $f$ is unknown, we can make a decision using the data $\{(Y_i,x_i,d_i)\}_{i=1}^n$, which are informative about $f$. This setting fits into my framework by treating $\{(x_i,d_i)\}_{i=1}^n$ as fixed and setting $\boldsymbol Y=(Y_1,...,Y_n)'$, $\theta=f$, $\boldsymbol m(f)=(f(x_1,d_1),...,f(x_n,d_n))'$, and $L(f)=W_1(f)-W_0(f)$. The parameter space $\Theta$ is a class of functions $f$ that satisfy certain restrictions (e.g., bounds on derivatives or the linearity), specified by the policy maker so that the welfare contrast $L(f)$ is point or partially identified. \end{itemize}

I demonstrate the practical relevance of my framework through an application to the problem of eligibility cutoff choice. In Section (ref), I consider a situation in which the eligibility for treatment (e.g., social, educational, or welfare programs) is determined based on whether the value of an individual's characteristic exceeds a certain cutoff, as in RD setups. The policymaker wants to change the cutoff to a specific new value if the welfare effect of the cutoff change is positive. Here, the welfare effect is defined as the average treatment effect across units whose treatment status would be changed under the new cutoff. In this context, a decision rule maps the data collected under the status quo cutoff to the probability of changing the cutoff. My results can be used to find an optimal decision rule for this problem, under various restrictions on the conditional mean potential outcome function that enable extrapolation from one side of the cutoff to the other. For illustration, I assume that the function satisfies Lipschitz continuity with a known Lipschitz constant (i.e., a bound on the first derivatives), which leads to partial identification of the welfare effect. Applying my general results, I show that the minimax regret rule makes a decision based on the difference between a weighted average of observed outcomes for the treated units and that for the untreated units, with the weight vector determined by the choice of the Lipschitz constant. The rule can easily be computed by solving finite-dimensional convex optimization problems.

commentIn many policy domains, the eligibility for treatment is determined based on an individual's observable characteristics. One crucial policy question is whether we should change the eligibility criterion to achieve better outcomes. I provide a data-driven approach to answering this question. Specifically, I consider an RD setup and study the problem of whether to change the eligibility cutoff from the current value to a new value, using the sample data generated under the current eligibility criterion. Suppose that the welfare effect of the cutoff change is given by the average treatment effect across units whose treatment status would be changed under the new cutoff. One approach to point or partially identifying the welfare effect from the data under the status quo cutoff is to impose restrictions on the conditional mean potential outcome function, such as smoothness, to enable extrapolation across the cutoff. For illustration, I assume that the function satisfies Lipschitz continuity with a known Lipschitz constant (i.e., a bound on the first derivatives), leading to partial identification of the welfare effect. Applying my general results, I derive a minimax regret rule to decide whether or not to change the cutoff. The rule makes a decision based on the difference between a weighted average of outcomes for the treated units and that for the untreated units, with the weight vector determined by the choice of Lipschitz constant.

Finally, in Section (ref), I apply this rule to the Burkinab\'{e} Response to Improve Girls' Chances to Succeed (BRIGHT) program, a school construction program in Burkina Faso Kazianga2013bright. With the aim of improving educational outcomes in rural villages, the program constructed primary schools in 132 villages from 2005 to 2008. To allocate schools, the Ministry of Education first computed a score that summarized village characteristics for each of the nominated 293 villages, then selected the highest-ranking villages to receive a school. Consider a policymaker who uses data collected after the completion of this program to decide whether to scale it up. As a hypothetical policy question, I consider whether to construct schools in the top 20% of previously ineligible villages, using the enrollment rate as the welfare measure. I impose the Lipchitz constraint on the counterfactual enrollment rates across villages. To consider policy costs, I assume that it is optimal to scale up the program if its cost-effectiveness is better than that of a similar policy. Applying my theoretical results, I find that the minimax regret rule nonrandomly decides not to scale up the program for a plausible range of the Lipschitz constant.

comment\paragraph{Contributions and Related Work.} \begin{enumerate} • The first contribution is to develop a general framework for policy decisions allowing for partial identification. Extend the framework for minimax statistical inference donoho1994,low1995tradeoff,armstrong2018optimal to cover \begin{itemize} • policy decision problems (as opposed to estimation and inference) • minimax {\it regret} problems (as opposed to minimax MSE and CI length) \end{itemize} Recently, donoho1994's framework has been applied to estimation and inference on treatment effects in a variety of settings, including RD and difference-in-differences designs armstrong2018optimal,Imbens2019RDD,kwon2020rd,Armstrong2021ATE,Chaisemartin2021aet,rambachan2023parallel. My framework can be applied to the problem of treatment choice under these settings. Within the treatment choice literature, the framework is a generalization of two existing setups. First, hirano2009asymptotics \begin{itemize} • \end{itemize} Second, the univariate problems with two-dimensional parameters by Stoye, allowing for . \begin{itemize} • Some insights are preserved, but some are not. \end{itemize} Recently, christensen2022discrete ... \begin{itemize} • Discuss differences \end{itemize} How to discuss policy learning with partial id? • The second contribution is to derive a finite-sample minimax regret rule. This paper is the first to use the hardest one-dimensional approach to solve the problem. \begin{itemize} • Differences from donoho1994 The difference poses two challenges. (1) existence of the case where there exist an infinite number of minimax regret rules for the hardest one-dimensional subproblem; (2) need to exploit a specific feature of the risk to verify that a guessed rule is indeed minimax regret. The paper addresses two issues • After the initial version of this paper, kitagawa2023partial \end{itemize} This approach is in contrast to ishihara2021meta... At the high level, this paper addresses potential future direction raised by Manski2020decision. \end{enumerate} \begin{itemize} • Point out that the paper addresses potential future direction raised by Manski2020decision. • General Insights \begin{itemize} • For point identification, nonrandomized rule • For partial identification, depends on three factors \begin{itemize} • This is observed in Stoye's setting, but I generalize this in a large class of problems. • Maybe illustrate the stoye's setup and insight first in Section 2, and emphasize that \end{itemize} • Even for point identification, restricted parameter space \end{itemize} \end{itemize}

\paragraph{Related Literature.} This paper contributes to the literature on minimax regret statistical treatment choice under point identification hirano2009asymptotics,Stoye2009minimax,Stoye2012minimax,tetenov2012asymmetric and partial identification Manski2007missing,Stoye2012minimax. Stoye2012minimax derives a minimax regret rule in settings in which the experiment has imperfect validity, considering both Bernoulli and Gaussian models. My result generalizes Stoye2012minimax's Stoye2012minimax in Gaussian models (Proposition 7(iii)) by allowing for multivariate samples, three- or higher-dimensional parameters, and various forms of parameter spaces. Recently, ishihara2021meta consider the problem of deciding whether to introduce a new policy based on results from multiple external studies. They restrict attention to the class of nonrandomized threshold rules based on a weighted average of the sample and propose a way to numerically minimize the maximum regret. In contrast, I do not impose any restrictions on decision rules, and use the hardest one-dimensional subfamily argument to analytically derive a minimax regret rule. My approach does not involve numerically minimizing the maximum regret and, for some problems, offers a closed-form expression for a minimax regret rule. Since the initial version of this paper was circulated, there have been some advances in the literature. kitagawa2023partial,kitagawa2022nonlinear derive minimax fractional rules under squared welfare regret loss for both point and partial identification settings. olea2023partial point out the nonuniqueness of minimax regret rules for problems in which my minimax regret rule is randomized, and propose the least randomizing rule. Their work and mine are complementary; their analysis relies on the existence of a minimax regret rule based on a weighted sum of the sample, while I prove the existence of such a rule and derive its formula.

Broadly, this paper contributes to the literature on treatment choice and policy learning under partial identification, which has been growing in econometrics and statistics Manski2000ambiguity,manski2009diversified,Manski2010vaccine,manski2011are,manski2011partial,Manski2020decision,chamberlain2011bayesian,kasy2018taxation,russell2020policy,Mo2020robust,Kallus2020confounding,dadamo2022orthogonal,ben-michael2022safe,adjaho2023,christensen2022discrete,Kido2023locally.\footnote{An extensive literature examines the problem of learning treatment allocation policies that map an individual's covariates to a treatment. See Manski2004hetero,Dehejia2005decision,Stoye2009minimax,Stoye2012minimax,qian2011ind,Bhattacharya2012budget,Kitagawa2018EWM,kitagawa2021equal,Athey2021policy; and Mbakop2021penalized, among others, for point identification settings. My approach can be applied to partial identification settings in which the choice set consists of two treatment assignment policies. } Many recent studies, either implicitly or explicitly, consider the worst-case welfare (or welfare loss) over the identified set given the point-identified parameter as the loss function of a statistical decision problem. By contrast, this paper directly uses the welfare loss as the loss function without taking its worst-case value given the point-identified parameter, following the standard minimax regret criterion. Furthermore, they derive finite-sample bounds on the expected loss, its convergence rates, or the local asymptotic optimality of their proposed rules, while this paper derives a finite-sample exact optimal rule. Another difference, which is empirically relevant, is that many studies consider welfare parameters whose upper and lower bounds can be estimated at a parametric rate. In contrast, this paper covers welfare parameters whose bounds cannot be estimated at a parametric rate, such as the average treatment effect at a point of the running variable in a nonparametric RD setup.

The problem of eligibility cutoff choice considered in this paper is related to the literature on extrapolation away from the cutoff in RD designs, including rokkanen2015rd,angrist2015rd,dong2015rd,Bertanha2020rd,Bertanha2020many,Bennett2020rd; and Cattaneo2020multi. Unlike these papers, I explicitly consider the decision problem of whether to change the cutoff and derive an optimal decision rule. Recently, zhang2022rd consider the problem of learning cutoff-based policies under multi-cutoff designs and propose a maximin policy that is guaranteed to perform no worse than the existing policy.

commentThe Gaussian model used in this paper has been studied for the problem of minimax estimation and inference in nonparametric regression models. donoho1994 uses the modulus of continuity to characterize minimax affine nonrandomized estimation and inference procedures. I show that the modulus of continuity can be used for deriving a minimax regret decision rule within the class of all decision rules. While the derivation of my result and that of donoho1994's consist of similar steps, the proof of each step is significantly different. Recently, donoho1994's framework has been applied to estimation and inference on treatment effects in a variety of settings, including RD and difference-in-differences designs armstrong2018optimal,Imbens2019RDD,kwon2020rd,Armstrong2021ATE,Chaisemartin2021aet,rambachan2023parallel. My framework can be applied to the problem of treatment choice under these settings.
commentThe remainder of this paper is organized as follows. Section (ref) introduces the general setup and present examples. Section (ref) states the main results. Section (ref) discusses the relation to minimax estimation. Section (ref) applies the main results to eligibility cutoff choice. Section (ref) presents an empirical application. Section (ref) concludes. Appendix (ref) contains the proof and a technical discussion of the main theorem (Theorem (ref)). Appendices (ref), (ref), (ref), and (ref) contain auxiliary lemmas and the proofs of the other results, derivations for Section (ref), additional results for the general setup, and additional figures for the empirical application in Section (ref), respectively.

Problem Setup and Examples

In this section, I introduce my framework and provide examples to illustrate it. Section (ref) describes the statistical model that generates the data available to the policymaker. Section (ref) defines the policymaker's action set and the associated social welfare functions. Section (ref) introduces the minimax regret criterion as an optimality criterion for decision rules. Section (ref) provides a brief discussion of the relationship between this framework and existing ones. Finally, Section (ref) illustrates the framework using two examples.

Data-generating Model

Suppose that the policymaker observes a sample $\boldsymbol{Y}=(Y_1,...,Y_n)'\in \mathbb{R}^n$ of the form

align[align omitted — 99 chars of source]

where $\theta$ is an unknown parameter that lies in a known subset $\Theta$ of a vector space $\mathbb{V}$; $\boldsymbol{m}:\mathbb{V}\rightarrow \mathbb{R}^n$ is a known linear function; and $\boldsymbol\Sigma$ is a known, positive-definite $n\times n$ matrix. I allow $\theta$ to be an infinite-dimensional parameter such as a function.

The linearity of $\boldsymbol{m}$ is not necessarily restrictive. If we specify $\theta$ so that it contains each of the expected values of $Y_1,...,Y_n$ as its element, $\boldsymbol m$ is a function that extracts those expected values from $\theta$, which is linear in $\theta$.

This model allows the expected value of $\boldsymbol{Y}$ to depend on other observed variables such as covariates and treatment by treating them as fixed and subsuming them into $\boldsymbol m$ and $\boldsymbol \Sigma$. For example, a regression model with fixed regressors

align*[align* omitted — 97 chars of source]

is a special case in which $\boldsymbol{Y}=(Y_1,...,Y_n)'$, $\theta=f$, $\Theta$ is a class of functions, $\boldsymbol m(f)=(f(x_1),...,f(x_n))'$, and $\boldsymbol \Sigma={\rm diag}(\sigma^2(x_1),...,\sigma^2(x_n))$.

The normality of $\boldsymbol Y$ and the assumption of known variance are restrictive, but are often imposed to deliver finite-sample optimality results for statistical decision problems. In some cases, the normal model is motivated as an approximation to a finite-sample problem. Suppose that we observe an $n$-dimensional vector of statistics derived from the original data, which is an asymptotically normal estimator of its population counterpart. For example, the mean outcome difference between the treatment and control groups in a randomized experiment is a statistic that is asymptotically normal for the population mean difference. If we regard the $n$-dimensional vector of statistics as $\boldsymbol Y$, the normal model ((ref)) can be viewed as an asymptotic approximation. Also, in Appendix (ref), I consider an asymptotic framework in which the distribution of the error $\boldsymbol Y-\boldsymbol m(\theta)$ is unknown and the sample size $n$ goes to infinity. I propose a feasible decision rule and derive conditions under which its maximum regret over a class of distributions converges to that of a minimax regret rule as $n\rightarrow\infty$.

I assume that the parameter space $\Theta$ is convex and centrosymmetric (i.e., $\theta\in\Theta$ implies $-\theta\in\Theta$) throughout the paper. Typical parameter spaces considered in empirical analyses are convex. For example, in the regression model above, classes of functions with bounded derivatives are convex. The centrosymmetry simplifies the minimax analysis; see Remark (ref) in Section (ref) for the role of centrosymmetry. However, it rules out some shape restrictions. In the regression model above, the class of convex (or concave) functions is noncentrosymmetric.

comment\footnote{\textcolor{red}{On the other hand, in some cases, it is possible to impose the monotonicity of the regression function by normalizing the sample $\boldsymbol Y$ so that the new parameter space is centrosymmetric. Suppose, for example, that $\Theta=\{f\in{\cal F}_{\rm Lip}(C): f(x) \text{ is nondecreasing in } x\}$, where ${\cal F}_{\rm Lip}(C)=\{f:|f(x)-f(\tilde x)|\le C|x-\tilde x| \text{ for every } x, \tilde x\in\mathbb{R}\}$. ${\cal F}_{\rm Lip}(C)$ is centrosymmetric while $\Theta$ is not. It is easy to show that $\Theta=\{\tilde f+f_0:\tilde f\in {\cal F}_{\rm Lip}(C/2)\}$, where $f_0(x)=\frac{C}{2}x$ for all $x\in\mathbb{R}$. Therefore, the model $\boldsymbol Y\sim{\cal N}(\boldsymbol m(f),\boldsymbol \Sigma)$, $f\in\Theta$, is equivalent to the model $\tilde{\boldsymbol Y}\sim {\cal N}(\boldsymbol m(\tilde f), \boldsymbol \Sigma)$, $\tilde f\in{\cal F}_{\rm Lip}(C/2)$, where $\tilde{\boldsymbol Y}=\boldsymbol Y-\boldsymbol m(f_0)=(Y_1-f_0(x_1),...,Y_n-f_0(x_n))'$; the set of distributions of $\boldsymbol Y$ over $f\in\Theta$ is identical to the set of distributions of $\tilde{\boldsymbol Y}+\boldsymbol m(f_0)$ over $\tilde f\in{\cal F}_{\rm Lip}(C/2)$.}}

Policy Choice Problem

Now, suppose that the policymaker is interested in choosing between two policies, policy $1$ and policy $0$, to maximize social welfare. Suppose that the welfare resulting from implementing policy $a\in\{0,1\}$ under $\theta$ is $W_a(\theta)$, where $W_a:\mathbb{V}\rightarrow \mathbb{R}$ is a known function specified by the policymaker. The welfare contrast between policy $1$ and policy $0$ is given by $$ L(\theta)\coloneqq W_1(\theta) - W_0(\theta). $$ I assume that $L:\mathbb{V}\rightarrow \mathbb{R}$ is a linear function. The optimal policy under $\theta$ is policy 1 if $L(\theta)>0$, policy 0 if $L(\theta)<0$, and either if $L(\theta)=0$.

One example of a welfare criterion is a weighted average of an outcome across individuals. For example, suppose a policy could change the outcome of each individual. Suppose also that we specify $\theta=(f_1(\cdot),f_0(\cdot))$, where $f_a(x)$ represents the counterfactual mean outcome under policy $a$ across individuals whose observed covariates are $x$. The welfare under policy $a$ can be defined, for example, by the population mean outcome $W_a(\theta)=\int f_a(x)dP_X$, where $P_X$ is the probability measure of covariates and is assumed to be known. In this case, the welfare contrast $L(\theta)=\int [f_1(x)-f_0(x)]dP_X$ is linear in $\theta$. On the other hand, the linearity of $L$ may rule out welfare criteria that depend on the distribution of the counterfactual outcome other than the mean. See kitagawa2021equal for such welfare criteria.

Importantly, this framework allows for cases in which $L(\theta)$ is not point identified. Let ${\cal M}\coloneqq\{\boldsymbol{m}(\theta):\theta\in\Theta\}\subset\mathbb{R}^n$ denote the set of possible values of the reduced-form parameter $\boldsymbol{m}(\theta)$. The {\it identified set} of $L(\theta)$ when $\boldsymbol m(\theta)=\boldsymbol \mu\in \mathbb{R}^n$ is defined as $$ I(\boldsymbol \mu)\coloneqq \{L(\theta):\boldsymbol m(\theta)=\boldsymbol \mu, \theta\in \Theta\}. $$ $I(\boldsymbol \mu)$ may contain multiple elements for some or all $\boldsymbol \mu\in{\cal M}$. If $I(\boldsymbol \mu)$ contains both positive and negative values, the superior policy is ambiguous even without sampling uncertainty.

commentThe policy maker may postulate that $\theta$ lies in a known parameter space $\Theta$. This assumption may lead to partial identification of $L(\theta)$ in the following sense. Suppose we observe $\boldsymbol m(\theta)=\mathbb{E}_\theta[\boldsymbol Y]$ without sampling error and the observed value is $\boldsymbol \mu \in\mathbb{R}^n$. Given the parameter space $\Theta$, the set of values $L(\theta)$ consistent with a particular value $\boldsymbol \mu$ of $\boldsymbol{m}(\theta)$ is given by $$ {\cal S}(\boldsymbol \mu;\Theta)\coloneqq\{L(\theta): \theta\in \Theta, \boldsymbol{m}(\theta)=\boldsymbol \mu\}. $$ I say that $L(\theta)$ is {\it identified} if ${\cal S}(\boldsymbol \mu;\Theta)$ is singleton for every possible value $\boldsymbol \mu$ of $\boldsymbol{m}(\theta)$. I say that $L(\theta)$ is {\it partially identified} if ${\cal S}(\boldsymbol \mu;\Theta)$ is a proper subset of $\mathbb{R}$ for every $\boldsymbol \mu$, and ${\cal S}(\boldsymbol \mu;\Theta)$ contains two or more elements for some $\boldsymbol \mu$. If either ${\cal S}(\boldsymbol \mu;\Theta)\subset R_{+}$ or ${\cal S}(\boldsymbol \mu;\Theta)\subset R_{-}$, the optimal policy is unambiguous. Otherwise, the policy maker faces ambiguity even conditional on the knowledge of $\boldsymbol{m}(\theta)$. Whether

Optimality Criterion

This paper's goal is to provide an optimal decision rule for using data to make a policy decision. A (randomized) {\it decision rule} is a measurable function $\delta:\mathbb{R}^n\rightarrow [0,1]$, where $\delta(\boldsymbol{y})$ represents the probability of choosing policy $1$ when the realization of the sample $\boldsymbol{Y}$ is $\boldsymbol y$.\footnote{In contexts in which fractional treatment allocations, which assign treatment to a fraction of individuals in a population, are permitted, we can also interpret $\delta$ as a fractional rule. Here, $\delta(\boldsymbol y)$ represents the fraction of individuals to whom we would assign treatment. If the welfare of treating a fraction $a\in [0,1]$ is defined as $W_a(\theta)=W_0(\theta)+a(W_1(\theta)-W_0(\theta))$, the regret of a fractional rule is equal to that of a randomized rule.} I consider the minimax regret criterion as an optimality criterion for decision rules. To introduce it, I first define the {\it welfare regret loss} for policy choice $a\in\{0,1\}$ under $\theta$ as

align*[align* omitted — 211 chars of source]

The welfare regret loss $l(a,\theta)$ is the difference in welfare between the optimal policy and policy $a$ under $\theta$. If the policymaker chooses the superior policy, they do not incur any loss; otherwise, they incur a loss of the absolute value of the welfare contrast $L(\theta)$.

The {\it risk} or {\it regret} of decision rule $\delta$ under $\theta$ is the expected welfare regret loss $$ R(\delta,\theta)\coloneqq

casesL(\theta)(1-\mathbb{E}_\theta[\delta(\boldsymbol{Y})]) & if L(\theta)\ge 0,\\ -L(\theta)\mathbb{E}_\theta[\delta(\boldsymbol{Y})] & if L(\theta)< 0,

$$ where $\mathbb{E}_\theta$ denotes the expectation taken with respect to $\boldsymbol{Y}$ under $\theta$. Using the regret as a performance measure allows one to consider not only the error probabilities ($1-\mathbb{E}_\theta[\delta(\boldsymbol{Y})]$ or $\mathbb{E}_\theta[\delta(\boldsymbol{Y})]$) but also the potential welfare loss ($|L(\theta)|$).

The regret of a decision rule can vary with $\theta$ over the parameter space $\Theta$. Generally, no rule uniformly dominates all other rules. The minimax regret criterion aggregates the regret over $\Theta$ by considering the {\it maximum} or {\it worst-case regret}, defined as $\sup_{\theta\in\Theta}R(\delta,\theta)$.

commentI consider two classes of decision rules. The first is the set of all possible decision rules ${\cal D}$. The second is a class of nonrandomized and randomized decision rules based on linear estimators: ${\cal D}_{\rm lin}\coloneqq {\cal D}_{\rm lin, nonrandom}\cup{\cal D}_{\rm lin, random}$, where \begin{align*} {\cal D}_{\rm lin,nonrandom}&=\left\{\delta_{\boldsymbol{w}}(\boldsymbol{y})=\Id\left\{\sum_{i=1}^n w_i y_i\ge 0\right\}:\boldsymbol{w}\in\mathbb{R}^n\right\},\\ {\cal D}_{\rm lin, random}&=\left\{\delta_{\boldsymbol{w},\sigma_\xi}(\boldsymbol{y})=\Phi\left(\frac{\sum_{i=1}^n w_i y_i}{\sigma_\xi}\right):\boldsymbol{w}\in\mathbb{R}^n, \sigma_\xi>0\right\}, \end{align*} where $\Phi$ is the cumulative distribution function of a standard normal variable. In the setup of the above example, ${\cal D}_{\rm lin,nonrandom}$ includes decision rules that make a decision based on the sign of local linear estimators, MSE-optimal linear estimators, etc. Decision rules in ${\cal D}_{\rm lin, random}$ can be thought of as two-step rules, where we add a noise $\xi\sim{\cal N}(0,\sigma_\xi^2)$ to $\sum_{i=1}^nw_iy_i$, where $\sigma_\xi>0$, and then make a decision based on the sign of $\sum_{i=1}^nw_iy_i+\xi$: $$ \delta_{\boldsymbol{w},\sigma_\xi}(\boldsymbol{y})={\rm Pr}_{\xi\sim{\cal N}(0,\sigma_\xi^2)}\left(\sum_{i=1}^nw_iy_i+\xi\ge0\right). $$

Let ${\cal R}(\Theta)$ denote the {\it minimax risk} or {\it minimax regret}, defined as

align[align omitted — 120 chars of source]

where ${\cal D}$ denotes the set of all decision rules. Under the minimax regret criterion, we aim to derive a {\it minimax regret} decision rule $\delta^*$, which satisfies $\sup_{\theta\in\Theta}R(\delta^*,\theta)={\cal R}(\Theta)$.

To further understand the minimax regret criterion, note that the above minimax problem can equivalently be written as $\inf_{\delta\in{\cal D}}\sup_{\theta\in\Theta}\left(\max_{a\in\{0,1\}}W_a(\theta)-U(\delta,\theta)\right)$, where $U(\delta,\theta)\coloneqq W_1(\theta)\mathbb{E}_\theta[\delta(\boldsymbol{Y})]+W_0(\theta)(1-\mathbb{E}_\theta[\delta(\boldsymbol{Y})])$ is the expected welfare of decision rule $\delta$ under $\theta$. Thus, the minimax regret criterion optimizes the uniform closeness of the expected welfare to the maximum attainable welfare. When the minimax risk ${\cal R}(\Theta)$ is small, a minimax regret rule uniformly achieves near-optimal expected welfare across all parameter values.

\sloppy

Alternative optimality criteria include the maximin criterion, which maximizes the worst-case expected welfare. Specifically, one aims to solve $\sup_{\delta\in{\cal D}}\inf_{\theta\in\Theta}U(\delta,\theta)$. It has been pointed out that the maximin criterion is unreasonably pessimistic and can lead to pathological decision rules Manski2004hetero,Stoye2009minimax; see olea2023partial for such a result in the setting of this paper.\footnote{For example, suppose the policymaker perfectly knows that the welfare of the status quo policy ($a=0$) is given by $W_0(\theta)=w_0$ for some constant $w_0$. If the welfare of the new policy ($a=1$) can be less than $w_0$ under at least one parameter value (i.e., $\inf_{\theta\in\Theta}W_1(\theta)<w_0$), then the decision rule that always maintains the status quo regardless of the data (i.e., $\delta(\boldsymbol{Y})=0$) is optimal under the maximin criterion.} By contrast, the minimax regret criterion optimizes the worst-case expected welfare relative to what is achievable at a given parameter value, which often leads to nontrivial decision rules. Another alternative is the Bayes criterion, which optimizes the average risk over a prior. That is, one aims to solve $\sup_{\delta\in{\cal D}}\int U(\delta,\theta)d\pi(\theta)$ (or equivalently $\inf_{\delta\in{\cal D}}\int R(\delta,\theta)d\pi(\theta)$), where $\pi$ is a prior on $\Theta$. See, for example, chamberlain2011bayesian and kasy2018taxation for Bayesian treatment choice. This paper focuses on the minimax regret criterion, following prior treatment choice studies Manski2004hetero,Manski2007missing,hirano2009asymptotics,Stoye2009minimax,Stoye2012minimax,tetenov2012asymmetric,Kitagawa2018EWM. Studying optimal rules under the Bayes or other possible criteria is beyond the scope of this paper.

Relation to Existing Frameworks

Here, I discuss the relationship between the above framework and existing ones. For minimax regret treatment choice, the framework in this paper generalizes the univariate Gaussian problems with two-dimensional parameters in Stoye2012minimax to accommodate multivariate samples, parameters of three or higher dimensions, and various types of parameter restrictions. It also generalizes the limiting version of minimax regret problems under parametric models studied by hirano2009asymptotics to accommodate partially identified welfare contrasts and restricted parameter spaces. In the setting described in Section (ref), ishihara2021meta derive a minimax regret rule within the class of nonrandomized threshold rules based on a weighted average of the sample, namely $\delta(\boldsymbol{Y})=\mathbf{1}\{\sum_{i=1}^nw_iY_i\ge 0\}$ with $\sum_{i=1}^nw_i=1$. Their result does not require the convexity of the parameter space, but instead assumes that the parameter space is invariant to the addition of vectors of ones, which excludes bounded parameter spaces.\footnote{A generalization of their invariance assumption to the general setup of this paper is as follows: There exists $\iota\in\Theta$ such that $L(\iota)=1$ and $\theta+c\iota\in\Theta$ for all $\theta\in\Theta$ and $c\in\mathbb{R}$. Under this condition and the centrosymmetry of $\Theta$, it is possible to extend their approach to derive a minimax regret rule among rules of the form $\delta(\boldsymbol{Y})=\mathbf{1}\{\boldsymbol{w}'\boldsymbol{Y}\ge 0\}$ with $\boldsymbol{w}\in\mathbb{R}^n$ and $\boldsymbol{w}'\boldsymbol{m}(\iota)=1$, in the general setup of this paper.} The setting of ishihara2021meta with any convex parameter space, whether bounded or unbounded, is a special case of my setting.

donoho1994 studies the optimal estimation of a linear functional of $\theta$ in a more general version of this paper's model, which allows for infinite-dimensional Gaussian models and noncentrosymmetric parameter spaces. donoho1994 derives minimax estimators and confidence intervals within the class of affine procedures, using squared error, absolute error, or the length of a fixed-length two-sided confidence interval as a loss function. Using the framework of donoho1994, armstrong2018optimal provide a one-sided confidence interval that minimizes the maximum $\beta$th quantile of the excess length among all one-sided confidence intervals of a given confidence level. In contrast to these studies, this paper focuses on a binary decision problem under welfare regret loss. Section (ref) discusses the connection between minimax estimation and minimax regret treatment choice in detail.

Examples

I illustrate my framework using two examples. The first is a special case of ishihara2021meta's ishihara2021meta setup, which I use as a running example to illustrate theoretical results in Section (ref). The second is policy choice using observational data under unconfoundedness.

comment\begin{example} One of the simplest examples of my framework is where the policy maker must decide whether to treat members of a population based on an estimator for the average treatment effect from a randomized experiment. In the notation of my framework, we observe a scalar estimator $Y\sim {\cal N}(m(\theta),\sigma^2)$, $\theta$ represents the average treatment effect, $\Theta=\mathbb{R}$, and $m(\theta)=L(\theta)=\theta$. The results from hirano2009asymptotics and tetenov2012asymmetric show that $\delta^*(Y)=\mathbf{1}\{Y\ge 0\}$ is a minimax regret rule. Stoye2012minimax considers an extended setup where the experiment may have limited validity because of selective noncompliance or because the treatment population is different from the sampling population. He formalizes such situations as a decision problem with partial identification. In the notation of my framework, we observe a scalar sample $Y\sim {\cal N}(m(\theta),\sigma^2)$, $\theta=(\theta_1,\theta_2)'\in\mathbb{R}^2$, $m(\theta)=\theta_1$, $\Theta=\{(\theta_1,\theta_2)'\in [-1,1]^2:\theta_2\in [a\theta_1-b, a\theta_1+b]\}$ for some known constants $a\in (0,1]$ and $b>0$, and $L(\theta)=\theta_2$. Here, $\theta_1$ and $\theta_2$ represent the average treatment effects for the sampling and treatment populations, respectively, $Y$ is an estimator for $\theta_1$ from an experiment, and $[a\theta_1-b, a\theta_1+b]\cap [-1,1]$ is the identified set of $\theta_2$ given $\theta_1$. Stoye2012minimax derives a minimax regret rule within the class of all decision rules (see Section (ref) for its expression). My framework generalizes this setup in three ways: (1) the sample can be multidimensional; (2) the parameter can be three or higher dimensional, even infinite dimensional; and (3) flexible forms of the parameter space are allowed. \end{example}
example[Evidence Aggregation ishihara2021meta] Consider a policymaker who is interested in deciding whether to introduce a new policy to a specific local population based on causal evidence of similar policies implemented in other populations. We observe an $n$-dimensional sample $\boldsymbol Y\sim {\cal N}(\boldsymbol m(\theta), \boldsymbol\Sigma)$, where $\theta=(\theta_1,...,\theta_n,\theta_{n+1})'\in \Theta\subset \mathbb{R}^{n+1}$, $\boldsymbol m(\theta)=(\theta_1,...,\theta_n)'$, and $\boldsymbol{\Sigma}={\rm diag}(\sigma_1^2,...,\sigma_n^2)$. The welfare contrast is given by $L(\theta)=\theta_{n+1}$. Here, $\theta_{n+1}$ is the average welfare effect of a new policy on the target population; $\theta_1,...,\theta_n$ are the average welfare effects on $n$ study populations; and $Y_1,...,Y_n$ are estimators for $\theta_1,...,\theta_n$. $\Theta$ imposes restrictions on the differences between $\theta_i$'s, so that $\theta_{n+1}$ is point or partially identified from $\boldsymbol m(\theta)$. To illustrate my results in Section (ref), I focus on a simple case in which $n=2$ and $\sigma_1^2=\sigma_2^2=\sigma^2$ for some $\sigma^2>0$. Also, I specify $$ \Theta=\{\theta\in\mathbb{R}^3: |\theta_1-\theta_3|\le C_1, |\theta_2-\theta_3|\le C_2\} $$ for some known constants $C_1,C_2\ge 0$. For $i=1,2$, $C_i$ reflects prior knowledge about how similar study population $i$ and the target population are in terms of the average welfare effect. In this example, ${\cal M}=\{\boldsymbol{m}(\theta):\theta\in\Theta\}=\{(\theta_1,\theta_2)'\in\mathbb{R}^2:|\theta_1-\theta_2|\le C_1+C_2\}$. The identified set of $L(\theta)$ when $\boldsymbol{m}(\theta)=\boldsymbol{\mu}\in {\cal M}$ is given by the following intersection bounds: $$ I(\boldsymbol{\mu})=\{\theta_3:\boldsymbol{m}(\theta)=\boldsymbol{\mu},\theta\in\Theta\}=\left[\max\{\mu_1-C_1,\mu_2-C_2\},\min\{\mu_1+C_1,\mu_2+C_2\}\right]. $$ Note that the upper bound (and the lower bound) is not differentiable with respect to $\boldsymbol{\mu}$ if $\mu_1+C_1=\mu_2+C_2$ (and $\mu_1-C_1=\mu_2-C_2$, respectively). My framework covers problems with nondifferentiable upper and lower bounds on the welfare contrast.
example[Choice of Treatment Assignment Policy under Unconfoundedness] Consider a policymaker interested in choosing who should be treated based on an individual’s observable covariates. Suppose each member $i$ of the population is characterized by potential outcomes $Y_i(1)$ and $Y_i(0)$ with and without treatment; a vector of covariates $X_i\in{\cal X}\subset\mathbb{R}^k$; and a treatment indicator $D_i\in\{0,1\}$. The policymaker observes a random sample $\{(Y_i,X_i,D_i)\}_{i=1}^n$, where $Y_i=Y_i(1)D_i+Y_i(0)(1-D_i)$ is the realized outcome. Let $f(x,d)=\mathbb{E}[Y_i(d)|X_i=x]$ and $\sigma^2(x,d)={\rm Var}(Y_i(d)|X_i=x)$ for $(x,d)\in {\cal X}\times\{0,1\}$. Assume that the unconfoundedness condition, $(Y_i(1),Y_i(0))\indep D_i|X_i$, holds, so that $f(x,d)=\mathbb{E}[Y_i|X_i=x,D_i=d]$ and $\sigma^2(x,d)={\rm Var}(Y_i|X_i=x,D_i=d)$ for $x\in{\cal X}_d$ and $d\in\{0,1\}$, where ${\cal X}_d$ denotes the support of $X_i$ conditional on $D_i=d$. On the other hand, I do not impose the overlap condition, $0<\mathbb{P}(D_i=1|X_i)<1$, so the conditional average treatment effect $f(x,1)-f(x,0)$ is generally not point identified without additional structure. I condition on the realized values $\{(x_i,d_i)\}_{i=1}^n$ of $\{(X_i,D_i)\}_{i=1}^n$ to obtain a regression model with fixed regressors $$ Y_i=f(x_i,d_i)+U_i, $$ where $U_i$ is independent across $i$, $\mathbb{E}[U_i]=0$, and ${\rm Var}(U_i)=\sigma^2(x_i,d_i)$. This model fits into my framework by assuming $U_i$ is normal and setting $\boldsymbol Y=(Y_1,...,Y_n)'$, $\theta=f$, $\boldsymbol m(f)=(f(x_1,d_1),...,f(x_n,d_n))'$, and $\boldsymbol \Sigma={\rm diag}(\sigma^2(x_1,d_1),...,\sigma^2(x_n,d_n))$. Now, suppose the policymaker must decide between two treatment assignment policies, $\pi_1$ and $\pi_0$, where each policy $\pi_a:{\cal X}\rightarrow[0,1]$, $a\in\{0,1\}$, specifies the probability of assigning treatment to individuals with covariates $x\in{\cal X}$. For example, $\pi_0$ may be the status quo policy that generates the treatment assignment of the units in the sample (i.e., $\pi_0(X_i)=\mathbb{P}(D_i=1|X_i)$), and $\pi_1$ may be a new policy; in this case, if $\pi_0$ is a deterministic policy (i.e., $\pi_0(x)\in\{0,1\}$ for all $x$), then there is no overlap between the supports of covariates for the treated and untreated groups in the sampling population (i.e., ${\cal X}_1\cap{\cal X}_0=\varnothing$). Alternatively, $\pi_0$ may assign treatment to no one (i.e., $\pi_0(x)=0$ for all $x$), and $\pi_1$ may assign treatment to everyone (i.e., $\pi_1(x)=1$ for all $x$). Suppose the welfare under policy $a\in\{0,1\}$ is an average of the conditional mean potential outcome across different values of covariates $$ W_a(f) = \int [f(x,1)\pi_a(x)+f(x,0)(1-\pi_a(x))]d\nu(x) $$ for some known measure $\nu$. The welfare contrast between the two policies is $$ L(f)=W_1(f)-W_0(f)=\int (\pi_1(x)-\pi_0(x))[f(x,1)-f(x,0)]d\nu(x). $$ To point or partially identify $L(f)$ under imperfect overlap, we need to impose restrictions on $f$. Suppose that $f\in{\cal F}$, where ${\cal F}$ is a known convex and centrosymmetric function class and plays the role of the parameter space $\Theta$. Possible function classes include the class of functions with a known bound on derivatives. In Section (ref), I consider a special case of this setting, in which $x$ is scalar, $\pi_0$ is the status quo cutoff-based policy that generates the data, and $\pi_1$ is a new cutoff-based policy. I derive a minimax regret rule under the assumption that $f$ belongs to the Lipschitz class with a known Lipschitz constant.

Minimax Regret Rules in General Setup

commentThis section proceeds in the following way. In Section (ref), I state assumptions and present the formula for a minimax regret rule. The focus of Section (ref) is on succinctly describing the main result and illustrating it through a simple example. In Section (ref), I discuss the approach to deriving the minimax regret rule, through which I provide technical insights and interpretations of and intuitions for the minimax regret rule. In Section XX, I summarize and discuss implications.

In this section, I solve the minimax regret problem by using the hardest one-dimensional subfamily argument, which donoho1994 used to solve minimax affine estimation problems. This approach consists of three steps. The first step is to solve one-dimensional subproblems, in which the parameter space is restricted to a one-dimensional linear bounded subfamily. The second step is to search for the hardest one-dimensional subproblem, defined as the one with the highest minimax risk. The final step is to show that a minimax rule for the hardest one-dimensional subproblem is also minimax optimal for the original problem.

A key distinction between minimax regret treatment choice and minimax affine estimation lies in the structure of their risk functions. For estimation, standard risk functions such as mean squared error (MSE) can be decomposed into bias and variance. In contrast, the regret can be decomposed into the error probability and the potential welfare loss. Consequently, substantially different arguments are required for each of the above three steps.

I normalize $\boldsymbol\Sigma=\sigma ^2\boldsymbol I_n$ for some $\sigma>0$ throughout this section, where $\boldsymbol I_n$ is the identity matrix. This normalization is without loss of generality, since $\boldsymbol \Sigma$ is known.\footnote{Specifically, let $\tilde{\boldsymbol Y}=\boldsymbol{\Sigma}^{-1/2}\boldsymbol{Y}$ and $\tilde{\boldsymbol{m}}(\theta)=\boldsymbol{\Sigma}^{-1/2}\boldsymbol{m}(\theta)$ so that $\tilde{\boldsymbol Y}\sim {\cal N}(\tilde{\boldsymbol{m}}(\theta), \boldsymbol I_n)$. For any rule $\delta(\boldsymbol{Y})$, its regret in the problem with $(L,\boldsymbol{m},\Theta, \boldsymbol{\Sigma})$ is the same as the regret of the rule $\tilde \delta(\tilde{\boldsymbol{Y}})$ in the problem with $(L,\tilde{\boldsymbol{m}},\Theta, \boldsymbol{I}_n)$, where $\tilde \delta(\tilde{\boldsymbol{Y}})=\delta(\boldsymbol{\Sigma}^{1/2}\tilde{\boldsymbol{Y}})=\delta(\boldsymbol{Y})$.} I use the following assumption to derive a minimax regret rule.

assumption(i) $L:\mathbb{V}\rightarrow\mathbb{R}$ and $\boldsymbol m:\mathbb{V}\rightarrow\mathbb{R}^n$ are linear; (ii) $\Theta$ is a nonempty, convex, and centrosymmetric subset of $\mathbb{V}$; (iii) $L(\theta)\neq 0$ for at least one $\theta\in\Theta$; (iv) $\sup I(\boldsymbol{0})<\infty$.

The first two conditions are introduced in Section (ref). It is straightforward to see that ${\cal M}=\{\boldsymbol{m}(\theta):\theta\in\Theta\}$ is a nonempty, convex, and centrosymmetric subset of $\mathbb{R}^n$ under these two conditions. These conditions also imply the following relationship between the lower and upper bounds on the welfare contrast: $\sup I(\boldsymbol{\mu})=-\inf I(-\boldsymbol{\mu})$ for all $\boldsymbol{\mu}\in {\cal M}$. By this symmetry, it is sufficient to focus on $\sup I(\boldsymbol{\mu})$ in the analysis below. The third condition excludes the trivial case in which $L(\theta)=0$ for any $\theta\in\Theta$ and hence the worst-case regret of any decision rule is zero. The fourth condition is also necessary to obtain nontrivial results: If $\sup I(\boldsymbol{0})=\infty$, the regret of any decision rule is unbounded on $\{\theta\in\Theta:\boldsymbol m(\theta)=\boldsymbol 0\}$, and hence the worst-case regret of any rule is infinity.

In the following, I first solve one-dimensional subproblems in Section (ref). Next, I characterize the hardest one-dimensional subproblem in Section (ref) and present a minimax regret rule for the original problem in Section (ref).

Minimax Regret Rules for One-dimensional Subproblems

First, I consider one-dimensional subproblems. To define a one-dimensional subproblem, take any $\bar\theta\in \Theta$ such that $L(\bar\theta)\ge 0$. A {\it one-dimensional subfamily}, denoted by $[-\bar\theta,\bar\theta]$, is defined as the set of all convex combinations of $\bar\theta$ and $-\bar\theta$: $$ [-\bar\theta,\bar\theta]\coloneqq\{\theta\in \mathbb{V}:\theta=\lambda\bar\theta,\lambda\in [-1,1]\}. $$ $[-\bar\theta,\bar\theta]$ is a subset of $\Theta$, since $\Theta$ is convex and centrosymmetric. Given $\bar\theta$, $[-\bar\theta,\bar\theta]$ can be viewed as a one-dimensional parameter space with a scalar parameter $\lambda\in[-1,1]$. A {\it one-dimensional subproblem} is the problem of finding a minimax regret rule for $[-\bar\theta,\bar\theta]$, whose maximum regret over $[-\bar\theta,\bar\theta]$ equals ${\cal R}([-\bar\theta,\bar\theta])$, where ${\cal R}([-\bar\theta,\bar\theta])=\inf_{\delta\in{\cal D}}\sup_{\theta\in [-\bar\theta,\bar\theta]}R(\delta,\theta)$.

The following result derives minimax regret rules for one-dimensional subproblems. Let $\|\cdot\|$ denote the Euclidean norm and $\Phi$ denote the cumulative distribution function of a standard normal random variable.

lemma[Minimax Regret Rules for One-dimensional Subproblems] Suppose Assumption (ref) holds, and consider a one-dimensional subproblem for $[-\bar\theta,\bar\theta]$, where $\bar \theta\in\Theta$ and $L(\bar\theta)\ge 0$. Then, the following holds. \begin{enumerate}[label=(\roman*)] • If $\boldsymbol{m}(\bar\theta)\neq \boldsymbol{0}$, then the decision rule $ \delta^*(\boldsymbol{Y})=\mathbf{1}\left\{\boldsymbol{m}(\bar \theta)'\boldsymbol{Y}\ge 0\right\} $ is minimax regret. • If $\boldsymbol{m}(\bar\theta)=\boldsymbol{0}$, then any decision rule $\delta^*$ such that $ \mathbb{E}[\delta^*(\boldsymbol{Y})]=1/2, $ where $\boldsymbol Y\sim {\cal N}(\boldsymbol 0,\sigma^2 \boldsymbol I_n)$, is minimax regret. • The minimax risk is given by \begin{align*} {\cal R}([-\bar \theta,\bar \theta]) =\begin{cases} L(\bar\theta)\Phi\left(-\frac{\|\boldsymbol m(\bar\theta)\|}{\sigma}\right) & if \|\boldsymbol m(\bar\theta)\|\le \tau^*\sigma,\\ \tau^*\sigma \frac{L(\bar\theta)}{\|\boldsymbol m(\bar\theta)\|}\Phi\left(-\tau^*\right) & if \|\boldsymbol m(\bar\theta)\|> \tau^*\sigma, \end{cases} \end{align*} where $\tau^*\in \arg\max_{t\ge 0}t\Phi(-t)$, which is unique ($\tau^*\approx 0.752$). \end{enumerate}
proofSee Appendix (ref).

Lemma (ref) provides minimax regret rules for subproblem $[-\bar \theta,\bar \theta]$ separately for the following two cases: (i) $\boldsymbol{m}(\bar\theta)\neq \boldsymbol{0}$ and (ii) $\boldsymbol{m}(\bar\theta)= \boldsymbol{0}$. In case (i), a linear threshold rule based on $\boldsymbol{m}(\bar \theta)'\boldsymbol{Y}$ is minimax regret. On the other hand, in case (ii), any decision rule that chooses each policy with probability one-half over the distribution of $\boldsymbol Y\sim {\cal N}(\boldsymbol 0,\sigma^2 \boldsymbol I_n)$ is minimax regret. For example, a linear threshold rule $\delta(\boldsymbol{Y})=\mathbf{1}\left\{\boldsymbol{w}'\boldsymbol{Y}\ge 0\right\}$ for any $\boldsymbol{w}\neq \boldsymbol{0}$, a probit-like randomized rule $\delta(\boldsymbol{Y})=\Phi(\boldsymbol{w}'\boldsymbol{Y})$ for any $\boldsymbol{w}\neq \boldsymbol{0}$, and a data-independent randomized rule $\delta(\boldsymbol{Y})=1/2$ are all minimax regret. There exist infinitely many minimax regret rules for this case.

To gain intuition for Lemma (ref), I outline the derivation for case (i) and provide the proof for case (ii) when $L(\bar\theta)>0$. \paragraph{Case (i): $L(\bar\theta)>0$ and $\boldsymbol{m}(\bar\theta)\neq \boldsymbol{0}$.} Under $\theta=\lambda\bar\theta$, $\boldsymbol{Y}\sim {\cal N}(\lambda \boldsymbol{m}(\bar\theta),\sigma^2\boldsymbol{I}_n)$ and the welfare constrast is $\lambda L(\bar\theta)$ by the linearity of $\boldsymbol{m}$ and $L$. Viewing $\lambda\in [-1,1]$ as the underlying parameter of $[-\bar\theta,\bar\theta]=\{\theta\in \mathbb{V}:\theta=\lambda\bar\theta,\lambda\in [-1,1]\}$, one can show that the scalar statistic $T(\boldsymbol{Y})=\frac{\boldsymbol m(\bar\theta)'\boldsymbol Y}{\|\boldsymbol m(\bar\theta)\|^2}\sim {\cal N}\left(\lambda,\frac{\sigma^2}{\|\boldsymbol m(\bar\theta)\|^2}\right)$ is a sufficient statistic of $\boldsymbol Y$ for $\lambda$. Since the class of decision rules that only depend on a sufficient statistic is essentially complete\footnote{A class ${\cal C}$ of decision rules is {\it essentially complete} if, for any decision rule $\delta\notin{\cal C}$, there is a decision rule $\delta'\in{\cal C}$ such that $R(\delta,\theta)\ge R(\delta',\theta)$ for all $\theta\in\Theta$.} Berger1985book, it is justified to restrict one's attention to rules that depend on $\boldsymbol{Y}$ only through $T(\boldsymbol{Y})\in\mathbb{R}$.

With this restricted class of rules, the minimax regret problem for $[-\bar \theta,\bar \theta]$ is equivalent to a {\it univariate} problem in which we observe a univariate sample $T\sim {\cal N}\left(\lambda,\frac{\sigma^2}{\|\boldsymbol m(\bar\theta)\|^2}\right)$ and the welfare contrast is $\lambda L(\bar\theta)$ for $\lambda\in [-1,1]$. By a mild extension of the results of hirano2009asymptotics and tetenov2012asymmetric for univariate problems with unbounded parameter spaces to ones with bounded parameter spaces, the simple threshold rule $\delta(T)=\mathbf{1}\left\{T\ge 0\right\}$ is minimax regret for this univariate problem. Consequently, $\delta^*(\boldsymbol Y)=\mathbf{1}\{T(\boldsymbol{Y})\ge 0\}=\mathbf{1}\left\{\boldsymbol{m}(\bar \theta)'\boldsymbol{Y}\ge 0\right\}$ is minimax regret for the original one-dimensional subproblem $[-\bar \theta,\bar \theta]$.

For the minimax risk, a simple calculation shows that the regret of $\delta^*$ under $\theta=\lambda\bar\theta$ is

align*[align* omitted — 151 chars of source]
commentobserve that $$ \mathbb{E}_{\lambda\bar\theta}[\delta^*(\boldsymbol{Y})]=\mathbb{P}_{\lambda\bar\theta}\left(T(\boldsymbol{Y})\ge 0\right)=1-\Phi\left(-\lambda\|\boldsymbol{m}(\bar \theta)\|/\sigma\right)=\Phi\left(\lambda\|\boldsymbol{m}(\bar \theta)\|/\sigma\right). $$ The regret of $\delta^*$ under $\theta=\lambda\bar\theta$ is given by \begin{align*} R(\delta^*,\lambda\bar\theta)&=\lambda L(\bar\theta)(1-\mathbb{E}_{\lambda\bar\theta}[\delta^*(\boldsymbol Y)])\mathbf{1}\{\lambda L(\bar\theta)\ge 0\}+(-\lambda L(\bar\theta))\mathbb{E}_{\lambda\bar\theta}[\delta^*(\boldsymbol Y)]\mathbf{1}\{\lambda L(\bar\theta)< 0\}\\ &=\lambda L(\bar\theta) \Phi\left(-\lambda\|\boldsymbol{m}(\bar \theta)\|/\sigma\right)\mathbf{1}\{\lambda \ge 0\}+(-\lambda L(\bar\theta))\Phi\left(\lambda\|\boldsymbol{m}(\bar \theta)\|/\sigma\right)\mathbf{1}\{\lambda< 0\}\\ &=|\lambda| L(\bar\theta) \cdot \Phi\left(-|\lambda|\cdot \|\boldsymbol{m}(\bar \theta)\|/\sigma\right). \end{align*}

The first factor $|\lambda|L(\bar\theta)$ is the welfare loss when $\delta^*$ chooses the inferior policy under $\lambda\bar\theta$, which is increasing in $|\lambda|$. The second factor $\Phi\left(-|\lambda|\cdot \|\boldsymbol{m}(\bar \theta)\|/\sigma\right)$ is the probability of choosing the inferior policy, which decreases in $|\lambda|$. The regret $R(\delta^*,\lambda\bar\theta)$ is shown to be a bimodal function of $\lambda$ symmetric around zero, globally maximized at $\lambda\in \{-\tau^*\sigma/\|\boldsymbol{m}(\bar \theta)\|,\tau^*\sigma/\|\boldsymbol{m}(\bar \theta)\|\}$. Maximizing this function over $\lambda\in [-1,1]$ yields the minimax risk in Lemma (ref)(ref).

\paragraph{Case (ii): $L(\bar\theta)>0$ and $\boldsymbol{m}(\bar\theta)= \boldsymbol{0}$.} By the linearity of $\boldsymbol{m}$, $\boldsymbol{m}(\theta)=\boldsymbol{0}$ for any $\theta\in [-\bar\theta,\bar\theta]$. For a given rule $\delta$, the probability of choosing policy 1, $\mathbb{E}_{\theta}[\delta(\boldsymbol Y)]$, is constant over $\theta\in [-\bar\theta,\bar\theta]$, so that

align*[align* omitted — 188 chars of source]

where $\boldsymbol Y\sim {\cal N}(\boldsymbol 0,\sigma^2 \boldsymbol I_n)$. Consequently, the maximum regret can be calculated as follows: $$ \sup_{\theta\in[-\bar\theta,\bar\theta]}R(\delta,\theta)=

casesL(\bar \theta) (1-\mathbb{E}[\delta(\boldsymbol{Y})]) (at $\theta=\bar\theta$)& if \mathbb{E}[\delta(\boldsymbol{Y})]<1/2,\\ L(\bar \theta) /2 \quad\quad\quad\quad (at $\theta\in\{-\bar\theta,\bar\theta\}$)& if \mathbb{E}[\delta(\boldsymbol{Y})]=1/2,\\ L(\bar \theta)\mathbb{E}[\delta(\boldsymbol{Y})] \quad\quad (at $\theta=-\bar\theta$)& if \mathbb{E}[\delta(\boldsymbol{Y})]>1/2.

$$ Thus, any rule $\delta^*$ with $\mathbb{E}[\delta^*(\boldsymbol{Y})]=1/2$ is minimax regret, and the minimax risk is $L(\bar \theta)/2$.

comment\paragraph{Case (ii): $\boldsymbol{m}(\bar\theta)= \boldsymbol{0}$.} If $\boldsymbol{m}(\bar\theta)= \boldsymbol{0}$, then $\boldsymbol Y\sim {\cal N}(\boldsymbol 0,\sigma^2 \boldsymbol I_n)$ under $\theta=\lambda\bar\theta$ for any $\lambda\in[-1,1]$, and therefore, the sample $\boldsymbol Y$ is uninformative about $\lambda$. For a given decision rule $\delta$, the choice probability $\mathbb{E}_{\lambda\bar\theta}[\delta(\boldsymbol Y)]$ is constant across $\lambda\in[-1,1]$. The regret of $\delta$ under $\theta=\lambda\bar\theta$ is \begin{align*} R(\delta,\lambda\bar\theta) &=\lambda L(\bar\theta)(1-\mathbb{E}[\delta(\boldsymbol Y)])\mathbf{1}\{\lambda\ge 0\}+(-\lambda) L(\bar\theta)\mathbb{E}[\delta(\boldsymbol Y)]\mathbf{1}\{\lambda< 0\}, \end{align*} where $\mathbb{E}[\delta(\boldsymbol Y)]$ denotes the constant choice probability. Consequently, the maximum regret is attained at one or both of the boundary points of the parameter space, and can be calculated as follows: $$ \sup_{\lambda\in[-1,1]}R(\delta,\lambda\bar\theta)=\begin{cases} L(\bar \theta) (1-\mathbb{E}[\delta(\boldsymbol{Y})]) ~~\text{(at $\lambda=1$)}&\text{ if } \mathbb{E}[\delta(\boldsymbol{Y})]<1/2,\\ L(\bar \theta) /2 \quad\quad\quad\quad\hspace{0.62em}~~~\text{(at $\lambda\in\{-1,1\}$)}&\text{ if } \mathbb{E}[\delta(\boldsymbol{Y})]=1/2,\\ L(\bar \theta)\mathbb{E}[\delta(\boldsymbol{Y})] \quad\quad~~~\hspace{.2em}\text{(at $\lambda=-1$)}&\text{ if } \mathbb{E}[\delta(\boldsymbol{Y}^*)]>1/2. \end{cases} $$ Thus, any decision rule $\delta^*$ such that $\mathbb{E}[\delta^*(\boldsymbol{Y})]=1/2$ is minimax regret, and the minimax risk is given by $L(\bar \theta)/2$.
comment\begin{itemize} • To gain more intuition, note that for one-dimensional subproblems $[-\bar \theta,\bar \theta]$ where $L(\bar\theta)>0$ and $\boldsymbol{m}(\bar\theta)=\boldsymbol{0}$, the welfare contrast $L(\theta)$ is {\it necessarily} partially identified from the knowledge that $\boldsymbol{m}(\theta)=\boldsymbol{0}$. Indeed, the identified set of $L(\theta)$ is a nonempty bounded interval: $$ \{L(\theta):\boldsymbol{m}(\theta)=\boldsymbol{0}, \theta\in [-\bar \theta,\bar \theta]\}=\{L(\theta):\theta\in [-\bar \theta,\bar \theta]\}=[-L(\bar\theta),L(\bar\theta)]. $$ • Intuitively, if $\mathbb{E}[\delta^*(\boldsymbol{Y}^*)]\neq \frac{1}{2}$, \end{itemize}
commentthe decision rule $\bar\delta(\boldsymbol Y)=\mathbf{1}\{\boldsymbol m(\bar\theta)'\boldsymbol Y\ge 0\}$ is shown to be minimax regret for the subproblem $[-\bar \theta,\bar \theta]$. Computing the maximum regret $\sup_{\theta\in[-\bar \theta,\bar \theta]}R(\bar\delta,\theta)$ yields the above display. If $L(\bar\theta)\ge 0$ and $\|\boldsymbol m(\bar\theta)\|=0$ (i.e., the sample $\boldsymbol Y$ is uninformative), any decision rule $\bar\delta$ such that $\mathbb{E}_{\boldsymbol Y\sim {\cal N}(\boldsymbol 0,\sigma^2 \boldsymbol I_n)}[\bar\delta(\boldsymbol{Y})]=\frac{1}{2}$ is minimax regret for the subproblem. The maximum regret of such decision rules is $L(\bar\theta)/2$.

Hardest One-dimensional Subproblem

commentAs shown in Lemmas (ref) and (ref) in Appendix (ref), the minimax risk for the one-dimensional subproblem $[-\bar\theta,\bar\theta]$ with $L(\bar\theta)\ge 0$ is given by \begin{align*} {\cal R}(\sigma;[-\bar \theta,\bar \theta]) =\begin{cases} L(\bar\theta)\Phi\left(-\frac{\|\boldsymbol m(\bar\theta)\|}{\sigma}\right) & if \|\boldsymbol m(\bar\theta)\|\le \tau^*\sigma,\\ \tau^*\sigma \frac{L(\bar\theta)}{\|\boldsymbol m(\bar\theta)\|}\Phi\left(-\tau^*\right) & if \|\boldsymbol m(\bar\theta)\|> \tau^*\sigma. \end{cases} \end{align*} For example, if $L(\bar\theta)\ge 0$ and $\|\boldsymbol m(\bar\theta)\|>0$ (i.e., the sample $\boldsymbol Y$ is informative), the decision rule $\bar\delta(\boldsymbol Y)=\mathbf{1}\{\boldsymbol m(\bar\theta)'\boldsymbol Y\ge 0\}$ is shown to be minimax regret for the subproblem $[-\bar \theta,\bar \theta]$. Computing the maximum regret $\sup_{\theta\in[-\bar \theta,\bar \theta]}R(\bar\delta,\theta)$ yields the above display. If $L(\bar\theta)\ge 0$ and $\|\boldsymbol m(\bar\theta)\|=0$ (i.e., the sample $\boldsymbol Y$ is uninformative), any decision rule $\bar\delta$ such that $\mathbb{E}_{\boldsymbol Y\sim {\cal N}(\boldsymbol 0,\sigma^2 \boldsymbol I_n)}[\bar\delta(\boldsymbol{Y})]=\frac{1}{2}$ is minimax regret for the subproblem. The maximum regret of such decision rules is $L(\bar\theta)/2$.

Now, I search for the {\it hardest one-dimensional subfamily} $[-\bar{\theta}^*,\bar{\theta}^*]\subset\Theta$, which satisfies $ {\cal R}([-\bar{\theta}^*,\bar{\theta}^*])=\sup_{\bar\theta\in\Theta:L(\bar\theta)\ge 0}{\cal R}([-\bar \theta,\bar \theta])$. The key to characterizing the hardest one-dimensional subfamily is the {\it modulus of continuity}, defined as

align[align omitted — 149 chars of source]

The modulus of continuity and its variants have been used in constructing minimax estimators and confidence intervals on linear functionals in Gaussian models donoho1991,donoho1994,low1995tradeoff,armstrong2018optimal.\footnote{donoho1994 defines the modulus of continuity as $\tilde \omega(\epsilon)= \sup\{|L(\theta)-L(\tilde\theta)|: \|\boldsymbol{m}(\theta-\tilde\theta)\|\le \epsilon,\theta,\tilde\theta\in\Theta\}$. If $\Theta$ is convex and centrosymmetric, the relationship $\tilde\omega(\epsilon)=2\omega(\epsilon/2)$ holds.} Under Assumption (ref), $\omega(\epsilon)$ is the value of a convex optimization problem. If $\Theta$ is closed, this problem typically has a solution. \footnote{See donoho1994 for sufficient conditions for the existence of a solution.} In addition, $\omega(\cdot)$ is nonnegative and nondecreasing by construction. Furthermore, the modulus of continuity has the following properties.

lemmaUnder Assumption (ref), $\omega(\epsilon)<\infty$ for all $\epsilon\ge 0$, and $\omega(\cdot)$ is concave and continuous on $[0,\infty)$ and right differentiable at $0$.
proofSee Lemma (ref) in Appendix (ref).
commentI impose the following restriction to characterize the hardest one-dimensional subproblem in terms of the modulus of continuity. Recall ${\cal M}=\{\boldsymbol{m}(\theta)\in\mathbb{R}^n:\theta\in\Theta\}$. \begin{assumption} $\boldsymbol{0}$ is an interior point of ${\cal M}$, that is, $\{\boldsymbol{\mu}\in\mathbb{R}^n:\|\boldsymbol{\mu}\|<\epsilon\}\subset {\cal M}$ for some $\epsilon>0$. \end{assumption} Assumption (ref) requires that the mean vector $\boldsymbol{m}(\theta)$ take any value in a sufficiently small ball centered at $\boldsymbol{0}$. This condition holds as long as ${\cal M}\subset \mathbb{R}^n$ contains a set of $n$ linearly independent vectors. The running example satisfies Assumption (ref) as ${\cal M}=\mathbb{R}^2$. On the other hand, if the dimension of $\theta$ is lower than $n$, or if there is a linear restriction between the elements of $\theta$, then ${\cal M}$ is a subset of a lower-dimensional linear subspace of $\mathbb{R}^n$, and the condition is not satisfied. For example, suppose that $\Theta=\{\theta=(\theta_1,\theta_2)'\in\mathbb{R}^2:\theta_2=a\theta_1\}$ for some constant $a\neq 0$ and $\boldsymbol{m}(\theta)=(\theta_1,\theta_2)'$. In this case, ${\cal M}=\Theta$, which is a one-dimensional subspace of $\mathbb{R}^2$, and $\boldsymbol{0}$ is not an interior point of ${\cal M}$. Yet, it is possible to accommodate such a case under a relaxed version of Assumption (ref) that assumes that $\boldsymbol{0}$ is an interior point of ${\cal M}$ relative to the linear subspace spanned by ${\cal M}$, not the full Euclidean space $\mathbb{R}^n$. See Appendix XX for details. \textcolor{red}{***Any example that violates even this weaker assumption?***}
commentThe fifth condition assumes that $\boldsymbol{0}$ is in the interior of ${\cal M}$ relative to ${\rm aff}({\cal M})$, the linear subspace spanned by ${\cal M}$. A sufficient condition is that $\boldsymbol{0}$ is an interior point of ${\cal M}$, that is, $B_\epsilon(\boldsymbol{0})\subset {\cal M}$ for some $\epsilon>0$. Typically, ${\cal M}$ contains $n$ linearly independent vectors. In this case, ${\rm aff}({\cal M})=\mathbb{R}^n$, and the fifth condition simply requires that $\boldsymbol{0}$ be an interior point of ${\cal M}$. If the dimension of $\theta$ is lower than $n$, or if there is a linear restriction between the elements of $\theta$, then ${\rm aff}({\cal M})$ may be a lower-dimensional subspace of $\mathbb{R}^n$. For example, suppose that $\boldsymbol{m}(\theta)=\theta$ and $\Theta=\{\theta\in\mathbb{R}^2:\theta_1=\theta_2\}$. In this case, ${\cal M}={\rm aff}({\cal M})=\Theta$, which is a one-dimensional subspace of $\mathbb{R}^2$. While $\boldsymbol{0}$ is not an interior point of ${\cal M}$, it is a relative interior point of ${\cal M}$ as $B_\epsilon(\boldsymbol{0})\cap {\rm aff}({\cal M})=\{\theta\in\mathbb{R}^2:\theta_1=\theta_2,-\epsilon<\theta_1<\epsilon\}\subset {\cal M}$. I use this condition to guarantee nonemptyness and boundedness of the superdifferential of $\bar I(\cdot)$ at $\boldsymbol{0}$, defined below.
commentIn Section (ref), I show that a minimax regret rule for the hardest one-dimensional subproblem is minimax regret also for the full problem. The key to characterizing the hardest one-dimensional subproblem is the {\it modulus of continuity}, defined as $$ \omega(\epsilon)\coloneqq\omega(\epsilon;L,\boldsymbol{m},\Theta)\coloneqq \sup\{L(\theta): \|\boldsymbol{m}(\theta)\|\le \epsilon,\theta\in\Theta\},~~~\epsilon\ge 0, $$ where $\|\cdot\|$ is the Euclidean norm. The modulus of continuity and its variants have been used in constructing minimax optimal estimators and confidence intervals on linear functionals in Gaussian models donoho1994,low1995tradeoff,cai2004adaptive,armstrong2018optimal.\footnote{donoho1994 defines the modulus of continuity as $\tilde \omega(\epsilon)= \sup\{|L(\theta)-L(\tilde\theta)|: \|\boldsymbol{m}(\theta-\tilde\theta)\|\le \epsilon,\theta,\tilde\theta\in\Theta\}$. If $\Theta$ is convex and centrosymmetric, the relationship $\tilde\omega(\epsilon)=2\omega(\epsilon/2)$ holds.} Below, I suppress the arguments $L,\boldsymbol{m}$, and $\Theta$ if they are clear from the context. Before discussing how the modulus of continuity can be used to characterize the hardest one-dimensional subproblem, I state its properties and make a technical assumption. First, $\omega(\epsilon)$ is the value of a convex optimization problem as the objective function $L$ is linear and the constrained set $\{\theta\in\Theta: \|\boldsymbol{m}(\theta)\|\le \epsilon\}$ is convex by the linearity of $\boldsymbol{m}$ and convexity of $\Theta$. If $\Theta$ is closed, this optimization problem typically has a solution, that is, there exists $\theta_{\epsilon}\in\Theta$ such that $L(\theta_\epsilon)=\omega(\epsilon)$ and $\|\boldsymbol{m}(\theta_\epsilon)\|\le \epsilon$.\footnote{See donoho1994 for sufficient conditions for the existence of a solution.} By construction, $\omega(\epsilon)$ is nonnegative and nondecreasing in $\epsilon$. Furthermore, $\omega(\epsilon)$ is concave in $\epsilon$. \footnote{See, for example, donoho1994 and armstrong2018optimal.} I assume the following condition for the modulus of continuity at zero. \begin{assumption}[Modulus at Zero] $\omega(0)<\infty$, and $\omega(\cdot)$ is continuous at $\epsilon=0$. \end{assumption} Under Assumption (ref), it can be shown that $\omega(\epsilon)<\infty$ for every $\epsilon\ge 0$ and $\omega(\cdot)$ is continuous at every $\epsilon\ge 0$: see Lemma (ref) in Appendix (ref). Substantively, Assumption (ref) rules out cases where the welfare contrast is not even partially identified. To see this, observe that $ \omega(0)= \sup\{L(\theta): \boldsymbol{m}(\theta)=\boldsymbol 0,\theta\in\Theta\} $ by definition. The convexity and centrosymmetry of $\Theta$ implies that $$ {\rm cl}(I(\boldsymbol 0))=[-\omega(0),\omega(0)]. $$ Therefore, the identified set of $L(\theta)$ is bounded if and only if $\omega(0)<\infty$.

The following result shows that the hardest one-dimensional subfamily can be obtained by solving an optimization problem that involves the modulus of continuity. Define the right derivative of $\omega(\cdot)$ at $0$ as $\omega'(0)\coloneqq\lim_{\epsilon\downarrow 0}\frac{\omega(\epsilon)-\omega(0)}{\epsilon}$. Let $\phi$ denote the probability density function of a standard normal random variable. For $\epsilon\ge 0$, I say that {\it $\theta_\epsilon\in \Theta$ attains the modulus of continuity at $\epsilon$} if $L(\theta_\epsilon)=\omega(\epsilon)$ and $\|\boldsymbol{m}(\theta_\epsilon)\|\le \epsilon$.

lemma[Hardest One-dimensional Subproblem] Under Assumption (ref), the following holds. \begin{enumerate}[label=(\roman*)] • The largest minimax risk among one-dimensional subproblems is given by \begin{align} \sup_{\bar\theta\in\Theta:L(\bar\theta)\ge 0}{\cal R}([-\bar \theta,\bar \theta])=\sup_{\epsilon\in[0,\tau^*\sigma]}\omega(\epsilon)\Phi(-\epsilon/\sigma). \end{align} • There exists a unique solution $\epsilon^*$ to the right-hand side of (ref). Furthermore, $\epsilon^*>0$ if and only if $\sigma\omega'(0)>2\phi(0)\omega(0)$. • If there exists $\theta_{\epsilon^*}\in\Theta$ that attains the modulus of continuity at $\epsilon^*$, then $[-\theta_{\epsilon^*},\theta_{\epsilon^*}]$ is the hardest one-dimensional subfamily. That is, $$ {\cal R}([-\theta_{\epsilon^*},\theta_{\epsilon^*}])=\sup_{\bar\theta\in\Theta:L(\bar\theta)\ge 0}{\cal R}([-\bar \theta,\bar \theta]). $$ Furthermore, $\boldsymbol{m}(\theta_{\epsilon^*})$ does not depend on the choice of $\theta_{\epsilon^*}$ among potentially multiple $\theta$'s that attain the modulus of continuity at $\epsilon^*$, and $\|\boldsymbol{m}(\theta_{\epsilon^*})\|=\epsilon^*$. \end{enumerate}
proofSee Appendix (ref).

To provide intuition for this result, consider a convenient case in which $\omega(\epsilon)=\sup_{\theta\in\Theta:\|\boldsymbol{m}(\theta)\|= \epsilon} L(\theta)$ for each $\epsilon\ge 0$. In other words, suppose, for illustration, that the supremum remains the same if the inequality constraint $\|\boldsymbol{m}(\theta)\|\le \epsilon$ is replaced by the equality constraint $\|\boldsymbol{m}(\theta)\|= \epsilon$. In this case, the largest minimax risk among one-dimensional subproblems can be calculated as follows:

align*[align* omitted — 809 chars of source]

where the second equality uses Lemma (ref)(ref) and the third equality uses the assumption that $\omega(\epsilon)=\sup_{\theta\in\Theta:\|\boldsymbol{m}(\theta)\|= \epsilon} L(\theta)$. To simplify the last expression, note that $\frac{\omega(\epsilon)}{\epsilon}$ is continuous and nonincreasing on $(0,\infty)$ by the concavity of $\omega(\cdot)$, and hence $\sup_{\epsilon>\tau^*\sigma}\frac{\tau^*\sigma\omega(\epsilon)}{\epsilon}\Phi(-\tau^*)=\omega(\tau^*\sigma)\Phi(-\tau^*)\le \sup_{\epsilon\in[0,\tau^*\sigma]}\omega(\epsilon)\Phi(-\epsilon/\sigma)$. As a result, $$ \sup_{\bar\theta\in\Theta:L(\bar\theta)\ge 0}{\cal R}([-\bar \theta,\bar \theta])=\sup_{\epsilon\in[0,\tau^*\sigma]}\omega(\epsilon)\Phi(-\epsilon/\sigma). $$

In light of the above derivation, the problem of searching for the hardest one-dimensional subfamily can be interpreted as the following problem by an adversarial Nature. Nature optimizes $\epsilon\in[0,\tau^*\sigma]$, which represents a level of the strength of the signal provided by the sample $\boldsymbol{Y}$ within subfamily $[-\bar\theta,\bar\theta]$. The signal strength is measured by $\|\boldsymbol m(\bar\theta)\|$: The larger $\|\boldsymbol m(\bar\theta)\|$ is, the more information the sample $\boldsymbol Y$ provides about the sign of $L(\theta)$, and the smaller the error probability $\Phi(-\|\boldsymbol m(\bar\theta)\|/\sigma)$ is. The modulus of continuity $\omega(\epsilon)= \sup\{L(\bar \theta): \|\boldsymbol{m}(\bar \theta)\|= \epsilon,\bar \theta\in\Theta\}$ then represents the maximum potential welfare loss (i.e., $|L(\bar \theta)|$) among subfamilies subject to a given level of signal strength. Nature finally searches for the best level $\epsilon^*\in\arg\max_{\epsilon\in[0,\tau^*\sigma]}\omega(\epsilon)\Phi(-\epsilon/\sigma)$, which optimizes the balance between the error probability and maximum potential welfare loss. If $\epsilon^*>0$, the sample $\boldsymbol{Y}$ is informative within the hardest subfamily $[-\theta_{\epsilon^*},\theta_{\epsilon^*}]$. On the other hand, if $\epsilon^*=0$, the sample $\boldsymbol{Y}$ is uninformative within the hardest subfamily.

Lemma (ref)(ref) implies that $\epsilon^*>0$ if and only if $\omega'(0)/\omega(0)$ or $\sigma$ is sufficiently large. This is intuitive from the perspective of Nature's problem of optimizing the signal strength described above. The larger $\omega'(0)/\omega(0)(=\left.\frac{\partial}{\partial\epsilon}\log\omega(\epsilon)\right\vert_{\epsilon=0})$, the larger the percentage increase in maximum potential welfare loss associated with an increase in the signal level from $0$, and thus the greater Nature's incentive to choose a nonzero level of signal strength. Similarly, the larger $\sigma$ is, the noisier the sample $\boldsymbol{Y}$ becomes, leading to a smaller decrease in the error probability associated with an increase in the signal level, and consequently, a greater incentive for Nature to choose a nonzero signal strength.

remark[Role of the Centrosymmetry of $\Theta$] In Lemma (ref), I derive the hardest one-dimensional subfamily among centrosymmetric one-dimensional subfamilies $[-\bar\theta,\bar\theta]$. Under the centrosymmetry of the original parameter space $\Theta$, one can show that this subfamily is also hardest among all one-dimensional subfamilies of the form $[\bar\theta_0,\bar\theta_1]=\{(1-\lambda)\bar\theta_0+\lambda\bar\theta_1:\lambda\in [0,1]\}\subset\Theta$, including noncentrosymmetric ones. If $\Theta$ is noncentrosymmetric, it is also necessary to consider noncentrosymmetric subfamilies to characterize the hardest one. Although it is possible to solve noncentrosymmetric subproblems, their minimax risks do not take a simple form as in Lemma (ref).\footnote{More specifically, the minimax risk of a subproblem $[\bar\theta_0,\bar\theta_1]$ depends on both $\bar\theta_0$ and $\bar\theta_1$ even when $\bar\theta_1-\bar\theta_0$ is held fixed. This sharply contrasts with the minimax affine estimation problems studied by donoho1994, in which the minimax risk of a subproblem $[\bar\theta_0,\bar\theta_1]$ depends on $\bar\theta_0$ and $\bar\theta_1$ only through $\bar\theta_1-\bar\theta_0$ (in particular $L(\bar\theta_1-\bar\theta_0)$ and $\boldsymbol{m}(\bar\theta_1-\bar\theta_0)$).} Consequently, it is not straightforward to apply a similar argument to the one above to characterize the hardest one-dimensional subfamily. I leave the extension to noncentrosymmetric parameter spaces to future work.

I illustrate the results in Lemma (ref) using Example (ref). \setcounter{example}{0}

example[Continued] Without loss of generality, I assume $C_1\ge C_2$. The modulus of continuity is given by $\omega(\epsilon)=\sup\{\theta_3:\theta\in\mathbb{R}^3, |\theta_1-\theta_3|\le C_1, |\theta_2-\theta_3|\le C_2, \theta_1^2+\theta_2^2\le\epsilon^2\}$. Solving the optimization problem on the right-hand side yields $$ \omega(\epsilon)=\begin{cases} \epsilon+C_2 ~~ &\text{ if } 0\le \epsilon\le C_1-C_2,\\ \frac{1}{2}\left(\left(2\epsilon^2-(C_1-C_2)^2\right)^{1/2}+C_1+C_2\right) ~~ &\text{ if } \epsilon> C_1-C_2, \end{cases} $$ and the solution $\theta_\epsilon=(\theta_{\epsilon,1},\theta_{\epsilon,2},\theta_{\epsilon,3})'$ is given by \begin{align} \theta_\epsilon=\begin{cases} (0,\epsilon,\omega(\epsilon))' & if 0\le \epsilon\le C_1-C_2,\\ (\theta_{\epsilon,1},\theta_{\epsilon,1}+C_1-C_2,\omega(\epsilon))' & if \epsilon> C_1-C_2, \end{cases} \end{align} where $\theta_{\epsilon,1}=\frac{1}{2}\left(\left(2\epsilon^2-(C_1-C_2)^2\right)^{1/2}-C_1+C_2\right)$ if $\epsilon> C_1-C_2$. By Lemma (ref), the hardest subfamily is $[-\theta_{\epsilon^*},\theta_{\epsilon^*}]$, where $\epsilon^*\in \arg\max_{\epsilon\in[0,\tau^*\sigma]}\omega(\epsilon)\Phi(-\epsilon/\sigma)$. If $C_1>C_2$, then $\omega(0)=C_2$ and $\omega'(0)=1$, and therefore the condition $\sigma\omega'(0)>2\phi(0)\omega(0)$ corresponds to $\sigma>2\phi(0)C_2$. If $C_1=C_2$, then $\omega(0)=C_2$ and $\omega'(0)=\frac{1}{\sqrt{2}}$, and therefore the condition corresponds to $\frac{\sigma}{\sqrt{2}}>2\phi(0)C_2$. In both cases, $\epsilon^*>0$ if and only if $C_2$ is small relative to $\sigma$.
comment\begin{itemize} • Additionally, define $\bar I(\boldsymbol{\mu})$ as the upper bound on the welfare contrast when the reduced-form parameter is $\boldsymbol{\mu}$: $$ \bar I(\boldsymbol{\mu})\coloneqq \sup I(\boldsymbol{\mu}),~~~\boldsymbol{\mu}\in \mathbb{R}^n, $$ where $\bar I(\boldsymbol{\mu})=-\infty$ when $I(\boldsymbol{\mu})$ is empty. The modulus of continuity $\omega(\epsilon)$ can be equivalently written as \begin{align} \omega(\epsilon)=\sup_{\boldsymbol{\mu}\in{\cal M}:\|\boldsymbol{\mu}\|\le\epsilon}\bar I(\boldsymbol{\mu}). \end{align} In other words, $\omega(\epsilon)$ is the supremum of the upper bounds of the identified sets of the welfare contrast over the reduced-form parameter value $\boldsymbol{\mu}\in{\cal M}$ within a ball of radius $\epsilon$ around $\boldsymbol 0$. In particular, we have $\omega(0)=\sup_{\boldsymbol{\mu}\in{\cal M}:\|\boldsymbol{\mu}\|\le 0}\bar I(\boldsymbol{\mu})=\bar I(\boldsymbol{0})$. That is, $\omega(0)$ is the upper bound on the identified set of the welfare constrast when $\boldsymbol{m}(\theta)=\boldsymbol{0}$. \end{itemize}

Main Result: Minimax Regret Rule for the Full Problem

In the final step, I propose a minimax regret rule for the hardest one-dimensional subproblem $[-\theta_{\epsilon^*},\theta_{\epsilon^*}]$ and verify that it is indeed minimax regret for the full problem. For the case in which $\|\boldsymbol{m}(\theta_{\epsilon^*})\|=\epsilon^*>0$, I consider the rule $\delta(\boldsymbol{Y})=\mathbf{1}\left\{\boldsymbol{m}(\theta_{\epsilon^*})'\boldsymbol{Y}\ge 0\right\}$, which is minimax regret for $[-\theta_{\epsilon^*},\theta_{\epsilon^*}]$ by Lemma (ref). On the other hand, for the case in which $\|\boldsymbol{m}(\theta_{\epsilon^*})\|=\epsilon^*=0$, there exist infinitely many minimax regret rules for $[-\theta_{\epsilon^*},\theta_{\epsilon^*}]$. Among them, I consider a rule that can be viewed as a continuous extension of the rule $\delta(\boldsymbol{Y})=\mathbf{1}\left\{\boldsymbol{m}(\theta_{\epsilon^*})'\boldsymbol{Y}\ge 0\right\}$ from the case with $\epsilon^*>0$ to the case with $\epsilon^*=0$.

To construct the rule, first define $\bar I(\boldsymbol{\mu})$ as the upper bound on the welfare contrast when the reduced-form parameter is $\boldsymbol{\mu}$: $$ \bar I(\boldsymbol{\mu})\coloneqq \sup I(\boldsymbol{\mu})=\sup\{L(\theta):\boldsymbol m(\theta)=\boldsymbol \mu, \theta\in \Theta\},~~~\boldsymbol{\mu}\in \mathbb{R}^n, $$ where I use the convention that $\sup I(\boldsymbol{\mu})=-\infty$ when $I(\boldsymbol{\mu})$ is empty. Next, consider a sequence $\{\boldsymbol{\mu}_\epsilon\}_{\epsilon\in (0,\bar\epsilon)}$ indexed by positive real numbers $\epsilon\in (0,\bar\epsilon)$ with some $\bar\epsilon>0$ such that

align[align omitted — 181 chars of source]

Lastly, define

align[align omitted — 123 chars of source]

In Lemma (ref) in Appendix (ref), I show that the following holds under Assumption (ref). First, there exists a sequence $\{\boldsymbol{\mu}_\epsilon\}_{\epsilon\in (0,\bar\epsilon)}$ that satisfies (ref). Second, if $\omega'(0)>0$, the limit $\boldsymbol{w^*}$ exists and does not depend on the choice of $\{\boldsymbol{\mu}_\epsilon\}_{\epsilon\in (0,\bar\epsilon)}$ among potentially multiple sequences that satisfy (ref); that is, $\boldsymbol{w^*}$ is uniquely defined. Third, $\boldsymbol{w}^*$ is the direction in which the directional derivative of $\bar I(\cdot)$ at $\boldsymbol{0}$ is maximized among unit vectors. In other words, $\boldsymbol w^*$ is the “least favorable” direction, in the sense that the upper bound on the welfare contrast increases the most if the reduced-form parameter $\boldsymbol{m}(\theta)$ is changed from $\boldsymbol{0}$ in the direction $\boldsymbol w^*$.

The vector $\boldsymbol{w}^*$ relates to the modulus of continuity $\omega(\cdot)$ as follows: $\omega(\epsilon)=\sup_{\theta\in\Theta:\|\boldsymbol m(\theta)\|\le\epsilon}L(\theta)=\sup_{\boldsymbol{\mu}\in{\cal M}:\|\boldsymbol{\mu}\|\le\epsilon}\bar I(\boldsymbol{\mu})$; and if $\theta_\epsilon\in\Theta$ attains the modulus of continuity at $\epsilon$, then setting $\boldsymbol \mu_\epsilon=\boldsymbol{m}(\theta_\epsilon)$ satisfies (ref), and $\boldsymbol{w}^*=\lim_{\epsilon\downarrow 0}\frac{\boldsymbol{m}(\theta_\epsilon)}{\epsilon}$.

The following result derives a minimax regret rule for the original problem, which covers both the case in which $\epsilon^*>0$ (i.e., $\sigma\omega'(0)> 2\phi(0)\omega(0)$) and the case in which $\epsilon^*=0$ (i.e., $\sigma\omega'(0)\le 2\phi(0)\omega(0)$). The proof and a technical discussion are provided in Appendix (ref).

theorem[Minimax Regret Rule for the Full Problem] Suppose that Assumption (ref) holds. Let $\epsilon^*\in\arg\max_{\epsilon\in[0,\tau^*\sigma]}\omega(\epsilon)\Phi(-\epsilon/\sigma)$ and suppose that there exists $\theta_{\epsilon^*}\in\Theta$ that attains the modulus of continuity at $\epsilon^*$. Then, the following decision rule is minimax regret: \begin{align} \delta^*(\boldsymbol Y)=\begin{cases}\mathbf{1}\left\{\boldsymbol{m}(\theta_{\epsilon^*})'\boldsymbol{Y}\ge 0\right\} & if \omega'(0)>0, 2\phi(0)\frac{\omega(0)}{\omega'(0)} < \sigma,\\ \mathbf{1}\left\{(\boldsymbol{w}^*)'\boldsymbol{Y}\ge 0\right\} & if \omega'(0)>0, 2\phi(0)\frac{\omega(0)}{\omega'(0)} = \sigma,\\ \Phi\left(\dfrac{(\boldsymbol{w}^*)'\boldsymbol{Y}}{((2\phi(0)\omega(0)/\omega'(0))^2-\sigma^2)^{1/2}}\right)& if \omega'(0)>0, 2\phi(0)\frac{\omega(0)}{\omega'(0)} > \sigma,\\ 1/2 & if \omega'(0)=0. \end{cases} \end{align} Furthermore, the minimax risk is given by $$ {\cal R}(\Theta)=R(\delta^*,-\theta_{\epsilon^*})=R(\delta^*,\theta_{\epsilon^*})=\omega(\epsilon^*)\Phi(-\epsilon^*/\sigma). $$

Theorem (ref) yields the following implications. First, the minimax regret rule in Theorem (ref) depends on the sample only through a weighted sum, $\boldsymbol{m}(\theta_{\epsilon^*})'\boldsymbol{Y}$ or $(\boldsymbol{w}^*)'\boldsymbol{Y}$, even though no such restrictions are imposed. The weights can be calculated by solving convex optimization problems; see Section (ref) for a computational procedure.

Second, the rule is nonrandomized or randomized, depending on whether the condition $2\phi(0)\frac{\omega(0)}{\omega'(0)}\le \sigma$ holds. This condition is related to the strength of the identifying restrictions. Under Assumption (ref), the closure of $I(\boldsymbol{0})$ is shown to be $[-\omega(0),\omega(0)]$.\footnote{Since $L$ and $\boldsymbol m$ are linear and $\Theta$ is centrosymmetric, $-\omega(0)=\inf\{L(\theta):\boldsymbol m(\theta)=\boldsymbol 0, \theta\in\Theta\}$. Moreover, for any $\alpha\in (-\omega(0),\omega(0))$, we can find $\theta\in\Theta$ such that $L(\theta)=\alpha$ and $\boldsymbol m(\theta)=\boldsymbol 0$ by the linearity of $L$ and $\boldsymbol m$ and the convexity of $\Theta$.} We can thus interpret $\omega(0)$ as half the length of the identified set of $L(\theta)$ when $\boldsymbol m(\theta)=\boldsymbol 0$.

If $L(\theta)$ is point identified, the length of the identified set is zero, so $2\phi(0)\frac{\omega(0)}{\omega'(0)}\le \sigma$. Therefore, the minimax regret rule is always nonrandomized under point identification. Even if $L(\theta)$ is not point identified, when the identified set is small relative to the noise level $\sigma$ (holding $\omega'(0)$ fixed), the condition $2\phi(0)\frac{\omega(0)}{\omega'(0)}\le\sigma$ holds and the rule is nonrandomized. On the other hand, when the identified set is large relative to $\sigma$, the minimax rule is randomized.

commentit is useful to consider the problem of finding worst-case parameter values for a generic decision rule $\delta$. The worst-case regret is attained at the parameter values that optimally trade off the potential welfare loss and the probability of incurring loss. Suppose that $\omega(0)$ is large relative to $\sigma$ and that one uses a nonrandomized rule. \textcolor{red}{Since $\sigma$ is small, $\boldsymbol Y$ does not vary much across repeated samples, which makes their choice based on the nonrandomized rule close to deterministic under a given parameter value. By exploiting it, it is easy to find a value of $\theta$ under which the inferior policy is chosen with a high probability. If $\omega(0)$ is large enough, such choice of $\theta$ is not likely associated with a small welfare loss, leading to a large expected welfare loss of the decision rule.} One can avoid this by randomizing their decisions; randomization makes their choice less predictable and protects against the exploitation of predictable choices.

The randomized rule can equivalently be written as $\delta^*(\boldsymbol{Y})=\mathbb{P}\left((\boldsymbol{w}^*)'\boldsymbol{Y}+\xi\ge 0|\boldsymbol{Y}\right)$, where $\xi|\boldsymbol{Y}\sim {\cal N}(0,(2\phi(0)\omega(0)/\omega'(0))^2-\sigma^2)$. This rule can be implemented by first adding an independent noise $\xi$ to a scalar statistic $(\boldsymbol{w}^*)'\boldsymbol{Y}$ and then making a decision according to the sign of $(\boldsymbol{w}^*)'\boldsymbol{Y}+\xi$. This addition artificially increases the standard deviation of $(\boldsymbol{w}^*)'\boldsymbol{Y}$ from $\sigma$ to $2\phi(0)\omega(0)/\omega'(0)$, which is the threshold at which we switch from a nonrandomized rule to a randomized rule. The larger $\omega(0)$ is, the larger the variance of $\xi$ is and the more dependent the choice is on the noise. As a result, given any realization of $\boldsymbol Y$, the probabilities of choosing policy 1 and policy 0 approach $1/2$ as $\omega(0)$ increases, which suggests that the decisions become more mixed if we impose weaker restrictions on $\Theta$.

I apply Theorem (ref) to derive a minimax regret rule for Example (ref). \setcounter{example}{0}

example[Continued] First, consider the case in which $C_1>C_2$. In this case, for any $\epsilon\in (0,C_1-C_2)$, $\omega(\epsilon)=\epsilon+C_2$ and $\theta_\epsilon=\left(0,\epsilon,\omega(\epsilon)\right)'$. Let $\boldsymbol{\mu}_\epsilon=\left(0,\epsilon\right)'$, which satisfies (ref) for every $\epsilon\in (0,C_1-C_2)$, so that $\boldsymbol{w}^*=\lim_{\epsilon\downarrow 0}\frac{\boldsymbol \mu_\epsilon}{\epsilon}=\left(0,1\right)'$. By Theorem (ref), the following decision rule is minimax regret: \begin{align*} \delta^*(\boldsymbol Y)=\begin{cases}\mathbf{1}\left\{\theta_{\epsilon^*,1}Y_1+\theta_{\epsilon^*,2}Y_2\ge 0\right\} & if 2\phi(0)C_2 < \sigma,\\ \mathbf{1}\left\{Y_2\ge 0\right\} & if 2\phi(0)C_2 = \sigma,\\ \Phi\left(\dfrac{Y_2}{((2\phi(0)C_2)^2-\sigma^2)^{1/2}}\right)& if 2\phi(0)C_2 > \sigma, \end{cases} \end{align*} where $\epsilon^*\in\arg\max_{\epsilon\in[0,\tau^*\sigma]}\omega(\epsilon)\Phi(-\epsilon/\sigma)$ and $\theta_\epsilon$ is given by (ref). If $C_2$ is sufficiently small relative to $\sigma$, the rule is nonrandomized. Whether it gives a nonzero weight to both $Y_1$ and $Y_2$ depends on $(C_1,C_2)$: If $C_1$ is sufficiently close to $C_2$ so that $C_1-C_2<\epsilon^*$, then $\theta_{\epsilon^*,1}=\frac{1}{2}\left(\left(2(\epsilon^*)^2-(C_1-C_2)^2\right)^{1/2}-C_1+C_2\right)>0$; otherwise, $\theta_{\epsilon^*,1}=0$ even if the rule is nonrandomized. If $C_2$ is sufficiently large so that $2\phi(0)C_2 > \sigma$, the rule is randomized. Next, consider the case in which $C_1=C_2$. In this case, for any $\epsilon>0$, $\omega(\epsilon)=\frac{1}{\sqrt{2}}\epsilon+C_2$ and $\theta_\epsilon=\left(\frac{\epsilon}{\sqrt{2}},\frac{\epsilon}{\sqrt{2}},\omega(\epsilon)\right)'$. Let $\boldsymbol{\mu}_\epsilon=\left(\frac{\epsilon}{\sqrt{2}},\frac{\epsilon}{\sqrt{2}}\right)'$, which satisfies (ref) for every $\epsilon>0$, so that $\boldsymbol{w}^*=\lim_{\epsilon\downarrow 0}\frac{\boldsymbol \mu_\epsilon}{\epsilon}=\left(\frac{1}{\sqrt{2}},\frac{1}{\sqrt{2}}\right)'$. By Theorem (ref), the following decision rule is minimax regret: \begin{align*} \delta^*(\boldsymbol Y)=\begin{cases}\mathbf{1}\left\{\frac{1}{\sqrt{2}}(Y_1+Y_2)\ge 0\right\} & if 2\sqrt{2}\phi(0)C_2 \le \sigma,\\ \Phi\left(\dfrac{\frac{1}{\sqrt{2}}(Y_1+Y_2)}{((2\sqrt{2}\phi(0)C_2)^2-\sigma^2)^{1/2}}\right)& if 2\sqrt{2}\phi(0)C_2 > \sigma. \end{cases} \end{align*} The rule always gives equal weights to $Y_1$ and $Y_2$, unlike in the case in which $C_1>C_2$.

Below, I discuss the role of randomization in Section (ref) and the relationship between Theorem (ref) and existing results in Section (ref). Section (ref) introduces the proof strategy. Finally, Section (ref) provides the procedure for computing the minimax regret rule.

commentTheorem (ref) contains Proposition 7(iii) of Stoye2012minimax as a special case. He considers a simple setup with a specific form of partial identification, given in Section (ref).\footnote{Assumption (ref)(ref)(ref) and (ref) (the differentiability of $\rho$ and $\omega$) do not hold for some choices of constants $a$ and $b$. For such cases, the minimax regret rule is derived by Theorems (ref) and (ref) in Appendix (ref), which do not require the differentiability.} He shows that the following rule is minimax regret if $b<1$:\footnote{Stoye2012minimax also covers the case where $b\ge 1$. In this case, Assumption (ref)(ref)(ref) does not hold; $\theta^*=(0,1)'$ attains the modulus of continuity at $\epsilon=0$, but there exists no $\theta\in\Theta$ such that $L(\theta)=\theta_2\neq 0$ and $\theta^*+c\theta\in\Theta$ for any $c$ in a neighborhood of zero. Theorem (ref) in Appendix (ref) covers this case.} $$ \delta^*(Y)=\begin{cases} \mathbf{1}\{Y\ge 0\} ~~ & \text{if } \sigma\ge 2\phi(0)\frac{b}{a},\\ \Phi\left(\dfrac{Y}{\left((2\phi(0)b/a)^2-\sigma^2\right)^{1/2}}\right) ~~ & \text{if } \sigma< 2\phi(0)\frac{b}{a}. \end{cases} $$ The condition $\sigma\ge 2\phi(0)\frac{b}{a}$ is equivalent to $\sigma \ge 2\phi(0)\frac{\omega(0)}{\omega'(0)}$ since $\omega(\epsilon)=\sup\{\theta_2:\theta_1\in [-\epsilon,\epsilon],(\theta_1,\theta_2)'\in\Theta\}=\min\{a\epsilon+b,1\}$. Note that the nonrandomized minimax regret rule $\delta^*(Y)=\mathbf{1}\{Y\ge 0\}$ is insensitive to any of $\sigma$, $a$, and $b$ as long as $\sigma\ge 2\phi(0)\frac{b}{a}$. Theorem (ref) confirms that a minimax regret rule can be both nonrandomized and randomized even in much more general setups. At the same time, Theorem (ref) suggests that the nonrandomized minimax regret rule $\mathbf{1}\left\{\boldsymbol{m}(\theta_{\epsilon^*})'\boldsymbol{Y}\ge 0\right\}$ may be sensitive to $\sigma$ and $\Theta$, since $\boldsymbol{m}(\theta_{\epsilon^*})$ depends on them. Therefore, the robustness of the nonrandomized minimax regret rule to the error variance and to the parameter space is not a general property.

The Role of Randomization

Randomization plays the role of reducing the probability of choosing the inferior policy under parameter values at which the error probability of an original rule exceeds one-half. If the maximum regret of the original rule is attained at such parameter values, randomization can lead to a reduction in maximum regret.

To illustrate this, it is useful to compare the nonrandomized rule $\delta^*_{\rm NR}(\boldsymbol{Y})=\mathbf{1}\left\{(\boldsymbol{w}^*)'\boldsymbol{Y}\ge 0\right\}$ with the randomized rule $\delta^*$ in Theorem (ref). In Appendix (ref), I show that, under a mild condition on the local behavior of $\bar I(\boldsymbol{\mu})$ around $\boldsymbol{0}$, if $2\phi(0)\frac{\omega(0)}{\omega'(0)}>\sigma$, there exists $\theta^*\in\Theta$ such that $L(\theta^*)\ge 0$ (i.e., the optimal policy is policy 1),

align[align omitted — 365 chars of source]

Condition (ref) means that the error probability of $\delta^*_{\rm NR}$ exceeds one-half under $\theta^*$. Condition (ref) says that the regret of $\delta^*_{\rm NR}$ under $\theta^*$ exceeds the minimax risk over $\Theta$, and therefore $\delta^*_{\rm NR}$ is not minimax regret.

Now, consider the randomized rule $\delta^*$ in Theorem (ref). Under any $\theta^*$ that satisfies (ref) and (ref), randomization reduces the error probability and therefore regret of $\delta^*_{\rm NR}$. Specifically, a simple calculation shows that if $2\phi(0)\frac{\omega(0)}{\omega'(0)}>\sigma$, $$ 1-\mathbb{E}_{\theta^*}[\delta^*(\boldsymbol{Y})]=\Phi\left(-\frac{(\boldsymbol{w}^*)'\boldsymbol{m}(\theta^*)}{2\phi(0)\omega(0)/\omega'(0)}\right)<\Phi\left(-\frac{(\boldsymbol{w}^*)'\boldsymbol{m}(\theta^*)}{\sigma}\right)=1-\mathbb{E}_{\theta^*}[\delta^*_{\rm NR}(\boldsymbol{Y})], $$ and hence $R(\delta^*,\theta^*)<R(\delta^*_{\rm NR},\theta^*)$. Although randomization may increase regret under other values of $\theta\in\Theta$, it turns out that $\delta^*$ achieves a smaller maximum regret over $\Theta$ than $\delta^*_{\rm NR}$.

commentFirst, one can show that the maximum regret of $\delta^*_{\rm NR}$ over $\Theta$ is given by $$ \sup_{\theta\in\Theta}R(\delta^*_{\rm NR},\theta)=\sup_{\boldsymbol{\mu}\in{\cal M}:\bar I(\boldsymbol{\mu})\ge 0}\bar I(\boldsymbol{\mu})\Phi\left(-(\boldsymbol{w}^*)'\boldsymbol{\mu}/\sigma\right), $$ using an argument similar to the one used in Appendix (ref). Here, $\bar I(\boldsymbol{\mu})$ and $\Phi\left(-(\boldsymbol{w}^*)'\boldsymbol{\mu}/\sigma\right)$ represent the maximum potential welfare loss and the error probability, respectively, when $\boldsymbol{m}(\theta)=\boldsymbol{\mu}$ and $L(\theta)\ge 0$. The product of these two represents the maximum regret of $\delta^*_{\rm NR}$ over $\Theta_1(\boldsymbol{\mu})\coloneqq\{\theta\in\Theta:\boldsymbol{m}(\theta)=\boldsymbol{\mu},L(\theta)\ge 0\}$; that is, $\sup_{\theta\in\Theta_1(\boldsymbol{\mu})}R(\delta^*_{\rm NR},\theta)=\bar I(\boldsymbol{\mu})\Phi\left(-(\boldsymbol{w}^*)'\boldsymbol{\mu}/\sigma\right)$. In Appendix (ref), I show that, under a mild condition on the local behavior of $\bar I(\boldsymbol{\mu})$ around $\boldsymbol{0}$, if $2\phi(0)\frac{\omega(0)}{\omega'(0)}>\sigma$, there exists $\boldsymbol{\mu}^*\in{\cal M}$ such that \begin{align} \sup_{\theta\in\Theta_1(\boldsymbol{\mu}^*)}R(\delta^*_{\rm NR},\theta)=\bar I(\boldsymbol{\mu}^*)\Phi\left(-(\boldsymbol{w}^*)'\boldsymbol{\mu}^*/\sigma\right)>\mathcal{R}(\Theta) and \Phi\left(-(\boldsymbol{w}^*)'\boldsymbol{\mu}^*/\sigma\right)>1/2. \end{align} The first condition says that the maximum regret of $\delta^*_{\rm NR}$ over $\Theta_1(\boldsymbol{\mu}^*)\subset\Theta$ exceeds the minimax risk over $\Theta$, and therefore $\delta^*_{\rm NR}$ is not minimax regret. The second means that the error probability of $\delta^*_{\rm NR}$ exceeds one-half when $\boldsymbol{m}(\theta)=\boldsymbol{\mu}^*$. Now, consider the randomized rule $\delta^*$ in Theorem (ref). As shown in Appendix (ref), $$ \sup_{\theta\in\Theta_1(\boldsymbol{\mu})}R(\delta^*_{\rm NR},\theta)=\bar I(\boldsymbol{\mu})\Phi\left(-(\boldsymbol{w}^*)'\boldsymbol{\mu}/s^*\right), $$ where $s^*=2\phi(0)\frac{\omega(0)}{\omega'(0)}>\sigma$. For any $\boldsymbol{\mu}^*\in{\cal M}$ that satisfies (ref), randomization reduces the error probability and therefore the maximum regret of $\delta^*_{\rm NR}$ over $\Theta_1(\boldsymbol{\mu}^*)$; that is, $\bar I(\boldsymbol{\mu}^*)\Phi\left(-(\boldsymbol{w}^*)'\boldsymbol{\mu}^*/\sigma\right)>\bar I(\boldsymbol{\mu}^*)\Phi\left(-(\boldsymbol{w}^*)'\boldsymbol{\mu}^*/s^*\right)$. While randomization may increase the maximum regret over $\Theta_1(\boldsymbol{\mu})$ for some other $\boldsymbol{\mu}\in{\cal M}$, it turns out that $\delta^*$ achieves a smaller maximum regret over the entire parameter space $\Theta$ than $\delta^*_{\rm NR}$.
remark[Approaches to Avoiding Randomization] In practice, randomization may not be permitted in some cases due to ethical or legislative constraints. An informal way of avoiding randomization is to reconsider the problem specification to reach the regime in which the minimax regret rule in Theorem (ref) is nonrandomized. This can be done by imposing more restrictions on the parameter space and/or respecifying policies 1 and 0. Note that this is not a post hoc analysis, since the problem specification and the associated decision rule are still determined before observing the data. A more formal way is to consider using the recently proposed least randomizing minimax regret rule by olea2023partial, which is a piecewise linear function of $(\boldsymbol{w}^*)'\boldsymbol{Y}$. This approach reduces the range of data realizations for which the decision is randomized. However, it does not lead to the complete elimination of randomization.
comment\footnote{For understanding why the policymaker should randomize their decisions in this case, it is useful to first consider the worst-case parameter values for the nonrandomized rule $\delta(\boldsymbol{Y})=\mathbf{1}\left\{(\boldsymbol{w}^*)'\boldsymbol{Y}\ge 0\right\}$. Its worst-case regret is attained at the parameter values that optimally trade off the probability of misidentifying the best policy and the potential welfare loss. If $\sigma$ is small, $\boldsymbol Y$ does not vary much across repeated samples. Therefore, $\delta(\boldsymbol{Y})$ almost deterministically chooses the inferior policy under parameter values where the expected value of $(\boldsymbol{w}^*)'\boldsymbol{Y}$ is negative but the welfare contrast $L(\theta)$ is positive, or vice versa. If, additionally, $\omega(0)/\omega'(0)$ is large, the worst-case regret is attained at such parameter values: the potential welfare loss $|L(\theta)|$ can also be moderately large, since the upper bound on the welfare contrast ($\omega(0)$) is large and not very sensitive to a change in the expected value of $\boldsymbol{Y}$ due to small $\omega'(0)$. The policymaker can avoid this worst-case scenario, where they almost deterministically choose the inferior policy, by randomizing their decisions. Randomization reduces the misidentification probability sufficiently to offset the increase in maximum potential welfare loss, thereby leading to a reduction in worst-case regret.}

Relation to Existing Results

Theorem (ref) generalizes Stoye2012minimax's Stoye2012minimax result from univariate problems to multivariate problems. For the case with randomization, the use of a probit-like rule in Theorem (ref) is inspired by Stoye2012minimax's Stoye2012minimax for univariate problems. A novelty of my result is to use the hardest one-dimensional subfamily argument to construct a certain scalar statistic, $\boldsymbol{m}(\theta_{\epsilon^*})'\boldsymbol{Y}$ or $(\boldsymbol{w}^*)'\boldsymbol{Y}$, of the multivariate sample $\boldsymbol{Y}$ such that a threshold rule based on the statistic (plus a noise for the case with randomization) is minimax regret. This sharply contrasts with univariate problems, in which the only natural choice of the statistic is the univariate sample $Y$ itself.

In the setting described in Example (ref), ishihara2021meta propose a way of numerically computing a minimax regret rule within the class of decision rules of the form $\delta(\boldsymbol{Y})=\mathbf{1}\{\sum_{i=1}^nw_iY_i\ge 0\}$ with $\sum_{i=1}^nw_i=1$. An application of Theorem (ref) shows that this restricted class contains an unconstrained minimax regret rule when $2\phi(0)\frac{\omega(0)}{\omega'(0)}\le \sigma$ and may not when $2\phi(0)\frac{\omega(0)}{\omega'(0)}>\sigma$.

Proof Strategy

I briefly describe the approach to proving Theorem (ref). First note that lower and upper bounds on the minimax risk for the full problem are given by $$ \sup_{\theta\in [-\theta_{\epsilon^*},\theta_{\epsilon^*}]}R(\delta^*,\theta)={\cal R}([-\theta_{\epsilon^*},\theta_{\epsilon^*}])\le {\cal R}(\Theta)\le \sup_{\theta\in\Theta}R(\delta^*,\theta), $$ where the equality follows from the fact that $\delta^*$ is minimax regret for the subproblem $[-\theta_{\epsilon^*},\theta_{\epsilon^*}]$ by Lemma (ref), and the two inequalities hold by the definition of the minimax risk. To prove that $\delta^*$ is also minimax regret for the full problem, it suffices to show that the above lower and upper bounds coincide. More specifically, I show that the maximum regret of $\delta^*$ over $\Theta$ is attained at $-\theta_{\epsilon^*}$ and $\theta_{\epsilon^*}$; that is,

align[align omitted — 138 chars of source]
commentThat is, \begin{align} \sup_{\theta\in [-\theta_{\epsilon^*},\theta_{\epsilon^*}]}R(\delta^*,\theta)={\cal R}([-\theta_{\epsilon^*},\theta_{\epsilon^*}]). \end{align} To show that this rule is also minimax regret for the full problem, I show that its maximum regret over $\Theta$ is attained at $-\theta_{\epsilon^*}$ and $\theta_{\epsilon^*}$. That is, \begin{align} R(\delta^*,-\theta_{\epsilon^*})=R(\delta^*,\theta_{\epsilon^*})=\sup_{\theta\in\Theta}R(\delta^*,\theta). \end{align} (ref) and (ref) imply that $\sup_{\theta\in\Theta}R(\delta^*,\theta)={\cal R}([-\theta_{\epsilon^*},\theta_{\epsilon^*}])$. However, the full problem is no less difficult than the subproblem by definition, so that $$ \sup_{\theta\in\Theta}R(\delta^*,\theta)\ge {\cal R}(\Theta)\ge {\cal R}([-\theta_{\epsilon^*},\theta_{\epsilon^*}]). $$ It follows that $\sup_{\theta\in\Theta}R(\delta^*,\theta)={\cal R}(\Theta)$, and therefore, $\delta^*$ is also minimax regret for the full problem.

To show this, I first express the maximum regret of $\delta^*$ over $\Theta$ as $$ \sup_{\theta\in\Theta}R(\delta^*,\theta)=

cases\sup_{\boldsymbol{\mu}\in {\cal M}:\bar I(\boldsymbol{\mu})\ge 0}\bar I(\boldsymbol{\mu})\Phi\left(-\dfrac{\boldsymbol{m}(\theta_{\epsilon^*})'\boldsymbol{\mu}}{\sigma\epsilon^*}\right) & if \epsilon^*>0,\\ \sup_{\boldsymbol{\mu}\in{\cal M}:\bar I(\boldsymbol{\mu})\ge 0}\bar I(\boldsymbol{\mu})\Phi\left(-\dfrac{(\boldsymbol{w}^*)'\boldsymbol{\mu}}{2\phi(0)\omega(0)/\omega'(0)}\right)& if \epsilon^*=0.

$$ Here, the objective function on the right-hand side represents the maximum regret of $\delta^*$ over $\{\theta\in\Theta:\boldsymbol{m}(\theta)=\boldsymbol{\mu},L(\theta)\ge 0\}$. I then show that the supremum is attained at $\boldsymbol{\mu}=\boldsymbol{m}(\theta_{\epsilon^*})$, which proves (ref).

The arguments used to show (ref) are substantially different from the arguments used by donoho1994, who shows a counterpart of (ref) for minimax affine estimation problems. donoho1994's donoho1994 arguments rely on the following property of the maximum risk (such as the maximum MSE) of an affine estimator: The maximum risk and maximum squared bias are attained at the same parameter values. Since the maximum regret does not have this property, the arguments by donoho1994 cannot be applied to show (ref) for the minimax regret problem.

Computation

I conclude this section by summarizing the procedure for computing the minimax regret rule for a given problem $(L,\boldsymbol{m},\Theta,\sigma^2\boldsymbol{I}_n)$. For a general variance $\boldsymbol{\Sigma}$, the minimax regret rule can be obtained by applying the procedure after normalizing $\boldsymbol{Y}$ and $(L,\boldsymbol{m},\Theta,\boldsymbol{\Sigma})$ to $\boldsymbol{\Sigma}^{-1/2}\boldsymbol{Y}$ and $(L,\boldsymbol{\Sigma}^{-1/2}\boldsymbol{m},\Theta,\boldsymbol{I}_n)$, respectively. The procedure consists of the following steps:

enumerate• Compute $\epsilon^*\in \arg\max_{\epsilon\in[0,\tau^*\sigma]}\omega(\epsilon)\Phi(-\epsilon/\sigma)$.\footnote{For numerical optimization, possible algorithms include grid search or ternary search for finding a maximum of a unimodal function. The bisection method can also be used if a closed-form expression for $\omega'(\cdot)$ is available; see Lemma (ref) in Appendix (ref) for the differentiability of $\omega(\cdot)$ and Appendix (ref) for the bisection method.} For a given $\epsilon\ge 0$, $\omega(\epsilon)$ is computed by solving the convex optimization problem (ref). Even if $\theta$ is infinite dimensional, the problem may be reduced to a finite-dimensional problem, as illustrated in Section (ref). If the upper bound $\bar I(\boldsymbol{\mu})$ is analytically tractable, $\omega(\epsilon)$ can also be computed by solving the convex optimization problem (ref). • If $\epsilon^*>0$ (i.e., $\sigma\omega'(0)> 2\phi(0)\omega(0)$), then solve (ref) for $\epsilon=\epsilon^*$ to calculate $\theta_{\epsilon^*}$ and compute $\boldsymbol{m}(\theta_{\epsilon^*})$. Alternatively, $\boldsymbol{m}(\theta_{\epsilon^*})$ can be directly computed as a solution to (ref) for $\epsilon=\epsilon^*$. • If $\epsilon^*=0$ (i.e., $\sigma\omega'(0)\le 2\phi(0)\omega(0)$), then compute $\boldsymbol{w}^*$, $\omega(0)$, and $\omega'(0)$. $\boldsymbol{w}^*$ can be computed by using the definition given by (ref) or one of the characterizations (ref) and (ref) in Appendix (ref). $\omega'(0)$ can be analytically computed if an analytical expression of $\omega(\epsilon)$ is available for any sufficiently small $\epsilon\ge 0$. Alternatively, $\omega'(0)$ can be computed using the characterization in (ref) in Appendix (ref). • Compute the decision rule given by (ref) in Theorem (ref).

Relation to Optimal Estimation

In this section, I study the relationship between optimal treatment choice and optimal estimation. Given an estimator $\hat L$ of the welfare contrast $L(\theta)$, a decision rule can be constructed by plugging the estimator into the oracle optimal decision $\mathbf{1}\{L(\theta)\ge 0\}$: $\delta(\boldsymbol{Y})=\mathbf{1}\{\hat L(\boldsymbol{Y})\ge 0\}$. Such a rule is called a {\it plug-in} rule. For example, a plug-in rule can be constructed by using an estimator of $L(\theta)$ that is optimal under some standard criterion for estimation, such as minimax MSE optimality. The minimax regret rule in Theorem (ref) can also be viewed as a plug-in rule, which uses a (possibly randomized) estimator. A natural question is whether the estimator used in the minimax regret rule is optimal in a certain sense. Another question is how the estimator used in the minimax regret rule differs from optimal estimators under standard criteria. This section aims to answer these two questions.

As a preliminary step, Section (ref) presents a class of estimators that optimally trade off bias and variance in the estimation of $L(\theta)$. Using the results, Section (ref) discusses an interpretation of the minimax regret rule as a plug-in rule based on an estimator that satisfies a certain optimality. In Section (ref), I compare this estimator with a minimax affine MSE estimator, which is an existing optimal estimator in the setting of this paper.

Throughout Section (ref), I normalize $\boldsymbol\Sigma=\sigma ^2\boldsymbol I_n$ for some $\sigma>0$ as in Section (ref). I further assume that $\omega(\cdot)$ is differentiable on $(0,\infty)$ to simplify the presentation and proof of the results. The results can be modified to allow for nondifferentiability of $\omega(\cdot)$ by using the superdifferentials of $\omega(\cdot)$, which exist by the concavity of $\omega(\cdot)$. I also note that $\omega(\cdot)$ is differentiable in Example (ref) and for eligibility cutoff choice in Section (ref). See Lemma (ref) in Appendix (ref) for a sufficient condition for the differentiability.

Optimal Bias and Variance Tradeoff

As a basis for the discussion in Sections (ref) and (ref), I introduce a class of estimators that optimally trade off bias and variance. Let ${\cal C}$ denote the class of all (nonrandomized) estimators for $L(\theta)$ (i.e., measurable functions from $\mathbb{R}^n$ to $\mathbb{R}$). For estimator $\tilde L\in {\cal C}$, let ${\rm Bias}(\tilde L,\theta)\coloneqq\mathbb{E}_\theta[\tilde L(\boldsymbol{Y})]-L(\theta)$ and $\Var(\tilde L,\theta)\coloneqq\mathbb{E}_\theta[(\tilde L(\boldsymbol{Y})-\mathbb{E}_\theta[\tilde L(\boldsymbol{Y})])^2]$. For scalar $V\ge 0$, let ${\cal C}(V)$ denote the class of estimators with the maximum variance over $\Theta$ less than or equal to $V$: ${\cal C}(V)\coloneqq\{\tilde L\in {\cal C}:\sup_{\theta\in\Theta}\Var(\tilde L,\theta)\le V\}$. Consider the following minimax problem: $$ \inf_{\tilde L\in{\cal C}(V)}\sup_{\theta\in\Theta}{\rm Bias}(\tilde L,\theta)^2. $$ Solving this problem yields a class of estimators indexed by $V$ that optimally trade off the maximum squared bias and the maximum variance.

low1995tradeoff derives estimators that achieve minimax optimality in the above sense for infinite-dimensional Gaussian models. In Theorem (ref) in Appendix (ref), I extend the result of low1995tradeoff to the multivariate Gaussian models in this paper. The following is a corollary of Theorem (ref), which translates a class of optimal estimators indexed by $V$ in Theorem (ref) into a class of optimal estimators indexed by $\epsilon\ge 0$. For $\epsilon\ge 0$, define

align*[align* omitted — 419 chars of source]

assuming that $\omega(\cdot)$ is differentiable on $(0,\infty)$ and $\theta_{\epsilon}\in\Theta$ attains the modulus of continuity at $\epsilon$ with $\|\boldsymbol m(\theta_{\epsilon})\|=\epsilon$.

theorem[Minimax Optimality of $\hat L_\epsilon$] Suppose that Assumption (ref) holds; that $\omega'(0)>0$; that $\omega(\cdot)$ is differentiable on $(0,\infty)$; and that for each $\epsilon> 0$, $\theta_{\epsilon}\in\Theta$ attains the modulus of continuity at $\epsilon$ with $\|\boldsymbol m(\theta_{\epsilon})\|=\epsilon$. Then, the following holds. \begin{enumerate}[label=(\roman*)] • For each $\epsilon\ge 0$, $\hat L_\epsilon$ has a constant variance of $V_\epsilon$: $\Var(\hat L_\epsilon,\theta)=V_\epsilon$ for all $\theta\in\Theta$. • For each $\epsilon\ge 0$, $\hat L_\epsilon$ minimizes the maximum squared bias among all estimators with the maximum variance less than or equal to $V_\epsilon$: $$\sup_{\theta\in\Theta}{\rm Bias}(\hat L_{\epsilon},\theta)^2=\inf_{\tilde L\in {\cal C}(V_\epsilon)}\sup_{\theta\in\Theta}{\rm Bias}(\tilde L,\theta)^2.$$ Furthermore, $\hat L_{0}$ minimizes the maximum squared bias among all estimators: $$\sup_{\theta\in\Theta}{\rm Bias}(\hat L_{0},\theta)^2=\inf_{\tilde L\in {\cal C}}\sup_{\theta\in\Theta}{\rm Bias}(\tilde L,\theta)^2.$$ • As $\epsilon$ increases, the maximum squared bias of $\hat L_\epsilon$ weakly increases and the variance of $\hat L_\epsilon$ weakly decreases. \end{enumerate}
proofSee Appendix (ref).

Theorem (ref) shows that the linear estimator $\hat L_{\epsilon}$ minimizes the maximum squared bias among all estimators (including nonlinear ones) with variance bounded by $V_\epsilon$. Furthermore, Theorem (ref) shows that the linear estimator $\hat L_{0}$ achieves the minimum maximum squared bias among all estimators. In contrast, low1995tradeoff does not provide a minimax squared bias estimator when variance constraints are absent in infinite-dimensional Gaussian models.\footnote{Specifically, Theorem 2 in low1995tradeoff does not provide a minimax squared bias estimator for the range of variance bound $V$ for which $0$ is the unique maximizer of $(\omega(\epsilon)-\epsilon\sqrt{V}/\sigma)$ over $\epsilon\ge 0$.}

commentThe following result extends the results to the multivariate Gaussian models of this paper. For each $\epsilon>0$, assuming that $\omega(\cdot)$ is differentiable at $\epsilon$ and $\theta_{\epsilon}\in\Theta$ attains the modulus of continuity at $\epsilon$ with $\|\boldsymbol m(\theta_{\epsilon})\|=\epsilon$, define \begin{align*} \hat L_{\epsilon}(\boldsymbol{Y})\coloneqq\begin{cases} \omega'(0)(\boldsymbol{w}^*)'\boldsymbol{Y} & if \epsilon=0,\\ \omega'(\epsilon)\frac{\boldsymbol{m}(\theta_{\epsilon})'}{\|\boldsymbol{m}(\theta_{\epsilon})\|}\boldsymbol{Y} & if \epsilon>0. \end{cases} \end{align*} \begin{theorem}[Optimal Bias-Variance Tradeoff] Under Assumption (ref), the following holds. \begin{enumerate}[label=(\roman*)] • Suppose that $\omega'(0)>0$ and that $\omega(\cdot)$ is differentiable on $(0,\infty)$. Let $V\ge 0$ and suppose that there exists $\epsilon_V\in\arg\max_{\epsilon\ge 0}(\omega(\epsilon)-\epsilon\sqrt{V}/\sigma)$ and that $\theta_{\epsilon_V}\in\Theta$ attains the modulus of continuity at $\epsilon_V$ with $\|\boldsymbol m(\theta_{\epsilon_V})\|=\epsilon_V$. Then, $\hat L_{\epsilon_V}$ has constant variance of $V$ or less: \begin{align*} \Var(\hat L_{\epsilon_V},\theta)=\begin{cases} (\sigma\omega'(0))^2 \le V & if \epsilon_V=0,\\ (\sigma\omega'(\epsilon_V))^2=V & if \epsilon_V>0. \end{cases} \end{align*} Furthermore, $\hat L_{\epsilon_V}$ minimizes the maximum squared bias among all estimators with the maximum variance less than or equal to $V$: $$\sup_{\theta\in\Theta}{\rm Bias}(\hat L_{\epsilon_V},\theta)^2=\inf_{\tilde L\in {\cal C}(V)}\sup_{\theta\in\Theta}{\rm Bias}(\tilde L,\theta)^2=(\omega(\epsilon_V)-\epsilon_V\omega'(\epsilon_V))^2.$$ • The estimator $\hat L_{0}(\boldsymbol{Y})=\omega'(0)(\boldsymbol{w}^*)'\boldsymbol{Y}$ minimizes the maximum squared bias among all estimators: $$\sup_{\theta\in\Theta}{\rm Bias}(\hat L_{0},\theta)^2=\inf_{\tilde L\in {\cal C}}\sup_{\theta\in\Theta}{\rm Bias}(\tilde L,\theta)^2=\omega(0)^2.$$ \end{enumerate} \end{theorem}
comment\begin{theorem} Suppose that Assumption (ref) holds, that $\omega'(0)>0$, and that $\omega(\cdot)$ is differentiable on $(0,\infty)$. Let $V\ge 0$ and suppose that there exists $\epsilon_V\in\arg\max_{\epsilon\ge 0}(\omega(\epsilon)-\epsilon\sqrt{V}/\sigma)$ and that $\theta_{\epsilon_V}\in\Theta$ attains the modulus of continuity at $\epsilon_V$ with $\|\boldsymbol m(\theta_{\epsilon_V})\|=\epsilon_V$. Then, the estimator \begin{align*} \hat L_{\epsilon_V}(\boldsymbol{Y})=\begin{cases} \omega'(0)(\boldsymbol{w}^*)'\boldsymbol{Y} & if \epsilon_V=0,\\ \omega'(\epsilon_V)\frac{\boldsymbol{m}(\theta_{\epsilon_V})'}{\|\boldsymbol{m}(\theta_{\epsilon_V})\|}\boldsymbol{Y} & if \epsilon_V>0 \end{cases} \end{align*} satisfies the following: \begin{enumerate}[label=(\roman*)] • $\Var(\hat L_{\epsilon_V},\theta)=(\sigma\omega'(\epsilon_V))^2=V$ for all $\theta\in\Theta$; and • $\sup_{\theta\in\Theta}{\rm Bias}(\hat L_{\epsilon_V},\theta)^2=\inf_{\tilde L\in {\cal C}(V)}\sup_{\theta\in\Theta}{\rm Bias}(\tilde L,\theta)^2=(\omega(\epsilon_V)-\omega'(\epsilon_V)\sqrt{V}/\sigma)^2$. \end{enumerate} Furthermore, the estimator $\hat L_{0}(\boldsymbol{Y})=\omega'(0)(\boldsymbol{w}^*)'\boldsymbol{Y}$ minimizes maximum squared bias among all estimators without constraints on the variance: specifically, $\sup_{\theta\in\Theta}{\rm Bias}(\hat L_{0},\theta)^2=\inf_{\tilde L\in {\cal C}}\sup_{\theta\in\Theta}{\rm Bias}(\tilde L,\theta)^2=\omega(0)^2$. \end{theorem}
comment\begin{theorem} Suppose that Assumption (ref) holds and that $\omega'(0)>0$. Let $\epsilon\ge 0$ and suppose that $\theta_\epsilon\in\Theta$ attains the modulus of continuity at $\epsilon$ with $\|\boldsymbol m(\theta_\epsilon)\|=\epsilon$ and that, if $\epsilon>0$, $\omega(\cdot)$ is differentiable at $\epsilon$. Then, the estimator \begin{align*} \hat L_\epsilon(\boldsymbol{Y})=\begin{cases} \omega'(0)(\boldsymbol{w}^*)'\boldsymbol{Y} & if \epsilon=0,\\ \omega'(\epsilon)\frac{\boldsymbol{m}(\theta_\epsilon)'}{\|\boldsymbol{m}(\theta_\epsilon)\|}\boldsymbol{Y} & if \epsilon>0 \end{cases} \end{align*} satisfies the following: \begin{enumerate}[label=(\roman*)] • $\Var(\hat L_\epsilon,\theta)=(\sigma\omega'(\epsilon))^2$ for all $\theta\in\Theta$; • $\sup_{\theta\in\Theta}{\rm Bias}(\hat L_\epsilon,\theta)^2={\rm Bias}(\tilde L_\epsilon,-\theta_\epsilon)^2={\rm Bias}(\tilde L_\epsilon,\theta_\epsilon)^2=\left(\omega(\epsilon)-\omega'(\epsilon)\epsilon\right)^2$; • $\sup_{\theta\in\Theta}{\rm Bias}(\hat L_\epsilon,\theta)^2=\inf_{\tilde L\in {\cal C}(V_\epsilon)}\sup_{\theta\in\Theta}{\rm Bias}(\tilde L,\theta)^2$, where $V_\epsilon\coloneqq (\sigma\omega'(\epsilon))^2$; and • $\sup_{\theta\in\Theta}{\rm Bias}(\hat L_\epsilon,\theta)^2=\inf_{\tilde L\in {\cal C}}\sup_{\theta\in\Theta}{\rm Bias}(\tilde L,\theta)^2$ for $\epsilon=0$. \end{enumerate} \end{theorem}
comment\begin{proof} See Appendix (ref). \end{proof} Theorem (ref)(ref) shows that the linear estimator $\hat L_{\epsilon_V}$ minimizes the maximum squared bias among all estimators (including nonlinear ones) subject to the constraint that the maximum variance is smaller than or equal to $V$. Theorem (ref)(ref) is a modification of Theorem 2 of low1995tradeoff from the infinite-dimensional Gaussian models to the multivariate Gaussian models. While low1995tradeoff does not provide an optimal estimator for the range of $V$ for which $\epsilon_V=0$ is the unique maximizer of $(\omega(\epsilon)-\epsilon\sqrt{V}/\sigma)$, the above result covers such a range of $V$. Specifically, if $V> (\sigma\omega'(0))^2$, then $\epsilon_V=0$ is shown to be the unique maximizer of $(\omega(\epsilon)-\epsilon\sqrt{V}/\sigma)$ since $\omega(\cdot)$ is concave. Theorem (ref)(ref) then implies that for any variance bound $V> (\sigma\omega'(0))^2$, $\hat L_{0}(\boldsymbol{Y})=\omega'(0)(\boldsymbol{w}^*)'\boldsymbol{Y}$ is an optimal estimator. Note that the minimum maximum squared bias is $\omega(0)^2$ in this range of $V$, which indicates that no further bias reduction is possible at the cost of increasing variance more. Indeed, Theorem (ref)(ref) shows that $\hat L_{0}$ also achieves the minimum maximum squared bias among all estimators even if there is no constraint on variance. For the discussion in Sections (ref) and (ref), it is useful to translate the optimality property of $\hat L_{\epsilon_V}$ for a given $V$ into the optimality property of $\hat L_\epsilon$ for a given $\epsilon$. \begin{corollary}[Minimax Optimality of $\hat L_\epsilon$] Suppose that Assumption (ref) holds, that $\omega'(0)>0$, that $\omega(\cdot)$ is differentiable on $(0,\infty)$, and that for each $\epsilon\ge 0$, $\theta_{\epsilon}\in\Theta$ attains the modulus of continuity at $\epsilon$ with $\|\boldsymbol m(\theta_{\epsilon})\|=\epsilon$. Let $V_\epsilon\coloneqq (\sigma\omega'(\epsilon))^2$. Then, for each $\epsilon\ge 0$, $\hat L_\epsilon$ minimizes the maximum squared bias among estimators with the maximum variance less than or equal to $V_\epsilon$ (and among all estimators when $\epsilon=0$). As $\epsilon$ increases, the maximum squared bias of $\hat L_\epsilon$ weakly increases and the variance of $\hat L_\epsilon$ weakly decreases. \end{corollary} \begin{proof} Since $\omega(\cdot)$ is concave and differentiable, $\epsilon\in\arg\max_{\epsilon\ge 0}(\omega(\epsilon)-\epsilon\sqrt{V_{\epsilon}}/\sigma)$ for each $\epsilon\ge 0$. The statements then follow from Theorem (ref). \end{proof}

Interpreting the Minimax Regret Rule as a Plug-in Rule

Theorem (ref) provides an interpretation of the minimax regret rule $\delta^*$ in Theorem (ref). If $2\phi(0)\frac{\omega(0)}{\omega'(0)}<\sigma$, the minimax regret rule is given by $\delta^*(\boldsymbol Y)=\mathbf{1}\{\boldsymbol{m}(\theta_{\epsilon^*})'\boldsymbol Y\ge 0\}=\mathbf{1}\{\hat L_{\epsilon^*}(\boldsymbol Y)\ge 0\}$, where $\epsilon^*\in\arg\max_{\epsilon\in [0,\tau^*\sigma]}\omega(\epsilon)\Phi(-\epsilon/\sigma)$. This corresponds to a plug-in rule based on the linear estimator $\hat L_{\epsilon^*}$, which minimizes the maximum squared bias among all estimators with variance bounded by $V_{\epsilon^*}$. In contrast, if $2\phi(0)\frac{\omega(0)}{\omega'(0)}>\sigma$, the minimax regret rule is given by $\delta^*(\boldsymbol Y)=\Phi\left(\frac{(\boldsymbol{w}^*)'\boldsymbol{Y}}{((2\phi(0)\omega(0)/\omega'(0))^2-\sigma^2)^{1/2}}\right)=\Phi\left(\frac{\hat L_0(\boldsymbol{Y})}{((2\phi(0)\omega(0))^2-(\sigma\omega'(0))^2)^{1/2}}\right)$. Equivalently, this can be written as $\delta^*(\boldsymbol Y)=\mathbb{P}\left(\hat L_0(\boldsymbol{Y})+\xi\ge 0|\boldsymbol{Y}\right)$, where $\xi|\boldsymbol{Y}\sim {\cal N}(0,(2\phi(0)\omega(0))^2-(\sigma\omega'(0))^2)$. Thus, this rule can be interpreted as a plug-in rule based on the {\it randomized} estimator $\hat L_0(\boldsymbol{Y})+\xi$ for $L(\theta)$. This estimator is constructed by adding a mean-zero Gaussian noise to the linear estimator $\hat L_0$, which minimizes the maximum squared bias among all estimators.

The above interpretation highlights a key distinction between minimax regret treatment choice and minimax estimation. The fact that a plug-in rule based on a randomized estimator can be minimax regret suggests that increasing the variance of an estimator while holding the bias constant may reduce the maximum regret of the resulting plug-in rule. This means that the maximum regret cannot be expressed as an increasing function of the maximum squared bias and the variance.\footnote{This observation aligns with a result from ishihara2021meta, who show that in their setting, the maximum regret of a linear threshold rule $\delta(\boldsymbol{Y})=\mathbf{1}\{\boldsymbol{w}'\boldsymbol{Y}\ge 0\}$ with $\sum_{i=1}^nw_i=1$ can be written as a function of the maximum absolute bias and the variance of $\boldsymbol{w}'\boldsymbol{Y}$, though the function is not monotonic in variance.} This observation sharply contrasts with certain performance measures used in minimax affine estimation. For example, the maximum MSE of an affine estimator $\tilde L(\boldsymbol{Y})=c+\boldsymbol{w}'\boldsymbol{Y}$ is given by $\sup_{\theta\in \Theta}\mathbb{E}_\theta[(\tilde L(\boldsymbol{Y})-L(\theta))^2]=\sup_{\theta\in\Theta}{\rm Bias}(\tilde L,\theta)^2+\sigma^2\|\boldsymbol{w}\|^2$, which is increasing in both the maximum squared bias and the variance.

commentFor minimax estimation and inference, the class of linear estimators $\{\hat L_\epsilon\}_{\epsilon\ge 0}$ can be used to find optimal procedures under other performance criteria beyond the optimality in the sense of Theorem (ref). First, armstrong2018optimal provide a one-sided confidence interval based on $\hat L_\epsilon$, which minimizes the maximum $\beta$th quantile of the excess length among all one-sided confidence intervals of a given confidence level. Second, suppose we evaluate {\it affine} estimators by a performance measure that is an increasing function of both the maximum squared bias and the variance. For example, the maximum MSE of an affine estimator $\tilde L(\boldsymbol{Y})=c+\boldsymbol{w}'\boldsymbol{Y}$ is given by $$ \sup_{\theta\in \Theta}\mathbb{E}_\theta[(\tilde L(\boldsymbol{Y})-L(\theta))^2]=\sup_{\theta\in\Theta}\left({\rm Bias}(\tilde L,\theta)^2+\Var(\tilde L,\theta)\right)=\left(\sup_{\theta\in\Theta}{\rm Bias}(\tilde L,\theta)^2\right)+\sigma^2\|\boldsymbol{w}\|^2, $$ which is increasing in both the maximum squared bias and the variance. Other performance measures that have this property include the maximum mean absolute error and the length of fixed-length two-sided confidence intervals centered at an affine estimator of a given confidence level donoho1994,low1995tradeoff. In view of Theorem (ref), minimizing each of these performance measures over affine estimators is equivalent to minimizing that of $\hat L_\epsilon$ over $\epsilon\ge 0$. In contrast, the maximum regret $\sup_{\theta\in\Theta}R(\delta,\theta)$ cannot be written as an increasing function of both the maximum squared bias and the variance, even for decision rules based on affine estimators. This can be seen by the result of Theorem (ref): if $2\phi(0)\frac{\omega(0)}{\omega'(0)}>\sigma$, artificially increasing the variance of $(\boldsymbol{w}^*)'\boldsymbol{Y}$ without changing its expected value (and hence the bias) reduces the maximum regret of the resulting threshold rule. Also, ishihara2021meta show in their setting that the maximum regret of a linear threshold rule $\delta(\boldsymbol{Y})=\mathbf{1}\{\boldsymbol{w}'\boldsymbol{Y}\ge 0\}$ with $\sum_{i=1}^nw_i=1$ can be written as a function of the maximum absolute bias and variance of $\boldsymbol{w}'\boldsymbol{Y}$, but the function is not monotonic in variance. Therefore, Theorem (ref) cannot be used to provide an alternative characterization of minimax regret rules; it simply provides a minimax optimality of the estimator used by the minimax regret rule derived in Theorem (ref).

Comparison with Minimax Affine MSE Estimator

To shed further light on the distinction between minimax regret treatment choice and minimax estimation, I compare the linear estimator used by the minimax regret rule with a minimax affine MSE estimator. To introduce the latter, let ${\cal C}_{\rm affine}$ denote the class of all affine estimators of $L(\theta)$: ${\cal C}_{\rm affine}\coloneqq\{\tilde L\in {\cal C}:\tilde L(\boldsymbol{Y})=c+\boldsymbol{w}'\boldsymbol{Y}, c\in\mathbb{R},\boldsymbol{w}\in\mathbb{R}^n\}$. An estimator $\hat L\in {\cal C}_{\rm affine}$ is {\it minimax affine MSE} if $\sup_{\theta\in \Theta}\mathbb{E}_\theta[(\hat L(\boldsymbol{Y})-L(\theta))^2]=\inf_{\tilde L\in {\cal C}_{\rm affine}}\sup_{\theta\in \Theta}\mathbb{E}_\theta[(\tilde L(\boldsymbol{Y})-L(\theta))^2]$.

Given the results in Theorem (ref), one approach to finding a minimax affine MSE estimator is to minimize the maximum MSE of $\hat{L}_\epsilon$ over $\epsilon \geq 0$. However, for comparison with the minimax regret rule, the following theorem instead relies on an alternative characterization of the minimax affine MSE estimator based on donoho1994's donoho1994 approach.

theorem[Comparison with Minimax Affine MSE Estimator] Suppose that Assumption (ref) holds, that $\omega'(0)>0$, and that $\omega(\cdot)$ is differentiable on $(0,\infty)$. Also, suppose that there exists $\epsilon_{\rm MSE}\in\arg\max_{\epsilon\ge 0}\frac{\sigma^2}{\epsilon^2+\sigma^2}\omega(\epsilon)^2$ and that $\theta_{\epsilon_{\rm MSE}}\in\Theta$ attains the modulus of continuity at $\epsilon_{\rm MSE}$ with $\|\boldsymbol m(\theta_{\epsilon_{\rm MSE}})\|=\epsilon_{\rm MSE}$. Then, the following holds. \begin{enumerate}[label=(\roman*)] • $\epsilon_{\rm MSE}>0$. Furthermore, $\hat L_{\epsilon_{\rm MSE}}(\boldsymbol{Y})=\omega'(\epsilon_{\rm MSE})\frac{\boldsymbol{m}(\theta_{\epsilon_{\rm MSE}})'}{\|\boldsymbol{m}(\theta_{\epsilon_{\rm MSE}})\|}\boldsymbol{Y}$ is a minimax affine MSE estimator. • Let $\epsilon^*\in\arg\max_{\epsilon\in [0,\tau^*\sigma]}\omega(\epsilon)\Phi(-\epsilon/\sigma)$. Then, $\epsilon_{\rm MSE}>\epsilon^*$. \end{enumerate}
proofSee Appendix (ref).

Theorem (ref)(ref) follows from the results in donoho1994. The optimization problem $\max_{\epsilon\ge 0}\frac{\sigma^2}{\epsilon^2+\sigma^2}\omega(\epsilon)^2$ corresponds to calculating the largest minimax affine MSE over all one-dimensional subfamilies. The estimator $\hat L_{\epsilon_{\rm MSE}}$ is shown to be minimax affine MSE for the hardest one-dimensional subproblem $[-\theta_{\epsilon_{\rm MSE}},\theta_{\epsilon_{\rm MSE}}]$ and also for the full problem $\Theta$. Furthermore, the optimal level $\epsilon_{\rm MSE}$ is always positive, regardless of the degree of partial identification or the noise level $\sigma$. This implies that, in the hardest one-dimensional subproblem for the minimax affine MSE problem, the sample $\boldsymbol{Y}$ is informative about the sign of $L(\theta)$, unlike in the minimax regret problem, where $\boldsymbol{Y}$ can be completely uninformative.

Theorem (ref)(ref) compares the optimal level of $\epsilon$ for the minimax affine MSE and minimax regret problems. Together with Theorem (ref), this result implies that $\hat L_{\epsilon^*}$ has a maximum squared bias that is no larger and a variance that is no smaller than those of $\hat L_{\epsilon_{\rm MSE}}$. \footnote{ishihara2021meta derive a related result using a different argument in their setting.}

commentI conclude this section with the following technical remark highlighting the distinction in the proof strategy for the minimax regret and minimax affine MSE problems. \begin{remark} For both of the minimax regret and minimax affine MSE problems, the primal challenge is to find a procedure that (i) is minimax for some subfamily and (ii) achieves maximum risk over $\Theta$ within the subfamily (see (ref) and (ref) for the minimax regret problem). For the minimax affine MSE problem, donoho1994 constructs two classes of affine estimators indexed by $\epsilon>0$, $\{\tilde L_\epsilon\}_{\epsilon>0}$ and $\{\hat L_\epsilon\}_{\epsilon>0}$, where $\tilde L_\epsilon(\boldsymbol{Y})=\frac{\omega(\epsilon)}{\epsilon^2+\sigma^2}\boldsymbol{m}(\theta_{\epsilon})'\boldsymbol{Y}$ and $\hat L_\epsilon(\boldsymbol{Y})=\omega'(\epsilon)\frac{\boldsymbol{m}(\theta_{\epsilon})'}{\|\boldsymbol{m}(\theta_{\epsilon})\|}\boldsymbol{Y}$. For each $\epsilon>0$, $\tilde L_\epsilon$ is minimax affine for $[-\theta_\epsilon,\theta_\epsilon]$ while $\hat L_\epsilon$ achieves the maximum squared bias and hence the maximum MSE over $\Theta$ within $[-\theta_\epsilon,\theta_\epsilon]$. If one can find $\epsilon_{\rm MSE}$ such that $\tilde L_{\epsilon_{\rm MSE}}=\hat L_{\epsilon_{\rm MSE}}$, then $\hat L_{\epsilon_{\rm MSE}}$ satisfies conditions (i) and (ii) and is minimax affine for the original problem. This strategy cannot be applied to the minimax regret problem. If, for example, one considers two threshold rules, one based on $\tilde L_\epsilon$ and another based on $\hat L_\epsilon$, then both rules are the same as $\delta(\boldsymbol{Y})=\mathbf{1}\{\boldsymbol{m}(\theta_{\epsilon})'\boldsymbol{Y}\ge 0\}$. This rule is minimax regret for $[-\theta_\epsilon,\theta_\epsilon]$ but does not achieve maximum regret over $\Theta$ within $[-\theta_\epsilon,\theta_\epsilon]$ in general. Moreover, for the minimax regret problem, the set of minimax rules for the hardest one-dimensional subproblem substantially differs for the case where the sample $\boldsymbol{Y}$ is informative ($\epsilon^*>0$) and the case where $\boldsymbol{Y}$ is uninformative ($\epsilon^*=0$). As a result, separate constructions of a rule and separate arguments to prove that it satisfies condition (ii) for the two cases are required. \textcolor{red}{***Discuss the strategy***} \end{remark}
commentTheorem (ref) implies that the minimax regret rule and some existing minimax estimators and confidence intervals can be calculated through similar optimization, with the bias-variance tradeoff resolved in a different way depending on the performance criterion. To see this, I assume $\omega(\cdot)$ is differentiable on $(0,\infty)$ with $\omega'(\epsilon)>0$ for all $\epsilon>0$ for simplicity\footnote{See Lemma (ref) in Appendix (ref) for a sufficient condition for the differentiability.} and refer to the following result from low1995tradeoff (see also armstrong2018optimal for a brief review): the optimal bias-variance frontier in the estimation of $L(\theta)$ can be traced out by a class of linear estimators $\{\hat L_\epsilon(\boldsymbol Y)\}_{\epsilon>0}$ of the form $ \hat L_\epsilon(\boldsymbol Y)=\boldsymbol{w}_\epsilon'\boldsymbol Y $, where $\boldsymbol{w}_\epsilon=\frac{\omega'(\epsilon)}{\epsilon}\boldsymbol{m}(\theta_{\epsilon})$. Here, $\theta_{\epsilon}$ attains the modulus of continuity at $\epsilon$ with $\|\boldsymbol m(\theta_{\epsilon})\|=\epsilon$. Specifically, for each $\epsilon>0$, $\hat L_\epsilon(\boldsymbol Y)$ minimizes the maximum bias among all linear estimators with variance bounded by $\Var(\hat L_\epsilon(\boldsymbol Y))=(\sigma\omega'(\epsilon))^2$: $$ \boldsymbol{w}_\epsilon\in \arg\min_{\boldsymbol w\in \mathbb{R}^n}\overline{{\rm Bias}}_{\Theta}(\boldsymbol w'\boldsymbol Y) ~~s.t.~~ \Var(\boldsymbol w'\boldsymbol Y)\le (\sigma\omega'(\epsilon))^2, $$ where $\overline{{\rm Bias}}_{\Theta}(\boldsymbol w'\boldsymbol Y)=\sup_{\theta\in\Theta}\mathbb{E}_\theta[\boldsymbol w'\boldsymbol Y-L(\theta)]$ is the maximum bias of $\boldsymbol w'\boldsymbol Y$ over $\Theta$. As $\epsilon$ increases, the maximum bias $\overline{{\rm Bias}}_{\Theta}(\hat L_\epsilon(\boldsymbol Y))$ increases and the variance $\Var(\hat L_\epsilon(\boldsymbol Y))=(\sigma\omega'(\epsilon))^2$ decreases. Consequently, the optimal weights can be found by optimizing $\epsilon>0$ for a given performance criterion, such as the MSE or mean absolute deviation of affine estimators, or the length of fixed-length two-sided confidence intervals centered at an affine estimator donoho1994. For example, let $\hat L_{\rm MSE}(\boldsymbol Y)=\boldsymbol{w}_{\rm MSE}'\boldsymbol{Y}$ be a {\it linear minimax MSE estimator} of $L(\theta)$, where $$ \boldsymbol{w}_{\rm MSE} \in\arg\min_{\boldsymbol w\in\mathbb{R}^n}\sup_{\theta\in\Theta}\mathbb{E}_\theta[(\boldsymbol{w}'\boldsymbol{Y}-L(\theta))^2]. $$ The optimal weights are then given by $\boldsymbol{w}_{\rm MSE}=\frac{\omega'(\epsilon_{\rm MSE})}{\epsilon_{\rm MSE}}\boldsymbol{m}(\theta_{\epsilon_{\rm MSE}})$, where $\epsilon_{\rm MSE}$ minimizes the worst-case MSE $\overline{{\rm Bias}}_{\Theta}(\hat L_\epsilon(\boldsymbol Y))^2+\Var(\hat L_\epsilon(\boldsymbol Y))$ over $\epsilon>0$. Note that the class of linear estimators $\{\hat L_\epsilon(\boldsymbol Y)\}_{\epsilon>0}$ can also be used to construct one-sided confidence intervals that minimize the maximum $\beta$th quantile of the excess length among all one-sided confidence intervals, as shown by armstrong2018optimal. To connect these existing results to the minimax regret rule, note that if $2\phi(0)\frac{\omega(0)}{\omega'(0)} < \sigma$, the minimax regret rule is given by $\delta^*(\boldsymbol Y)=\mathbf{1}\{\boldsymbol{m}(\theta_{\epsilon^*})'\boldsymbol Y\ge 0\}=\mathbf{1}\{\hat L_{\epsilon^*}(\boldsymbol Y)\ge 0\}$, where $\epsilon^*\in\arg\max_{\epsilon\in[0,\tau^*\sigma]}\omega(\epsilon)\Phi(-\epsilon/\sigma)$. Therefore, the rule uses the linear estimator $\hat L_{\epsilon^*}(\boldsymbol Y)$ that resolves the bias-variance tradeoff in a certain way. The following result compares $\hat L_{\epsilon^*}(\boldsymbol Y)$ with the linear minimax MSE estimator $\hat L_{\rm MSE}(\boldsymbol Y)$. The result uses the alternative characterization of $\epsilon_{\rm MSE}$ as a solution to $\frac{\epsilon^2}{\epsilon^2+\sigma^2}=\frac{\omega'(\epsilon)\epsilon}{\omega(\epsilon)}$, which is derived by donoho1994. \begin{proposition}[Comparison with Linear Minimax MSE Estimator] Suppose Assumption (ref) holds, $\omega(\cdot)$ is differentiable on $(0,\infty)$, and there exists $\epsilon_{\rm MSE}>0$ such that $\frac{\epsilon_{\rm MSE}^2}{\epsilon_{\rm MSE}^2+\sigma^2}=\frac{\omega'(\epsilon_{\rm MSE})\epsilon_{\rm MSE}}{\omega(\epsilon_{\rm MSE})}$. Let $\epsilon^*\in\arg\max_{\epsilon\in [0,\tau^*\sigma]}\omega(\epsilon)\Phi(-\epsilon/\sigma)$. Then, $\epsilon^*< \epsilon_{\rm MSE}$. \end{proposition} \begin{proof} See Appendix (ref). \end{proof} Proposition (ref) implies that $\hat L_{\epsilon^*}(\boldsymbol Y)$ has a smaller maximum bias and a larger variance than $\hat L_{\rm MSE}(\boldsymbol Y)$. In other words, the minimax regret rule places more importance on the bias than on the variance compared with the linear minimax MSE estimator. ishihara2021meta use a different analytical approach to derive such a finding in their setup. Lastly, I note that this paper's minimax regret problem poses three unique challenges in deriving optimal rules, compared to minimax affine estimation and inference problems studied by donoho1994. First, donoho1994 restricts attention to affine estimators and confidence intervals centered at an affine estimator, while I do not restrict attention to the class of threshold rules based on an affine estimator. Second, in my problem, it is possible that $\boldsymbol{m}(\theta)=\boldsymbol{0}$ for all $\theta$ within the hardest one-dimensional, and the set of minimax rules for this case is different from that for the case where $\boldsymbol{m}(\theta)$ varies within the hardest one-dimensional. (For the MSE problem, $0$ is the essentially unique minimax estimator for such one-dimensional problem.) As a result, separate constructions of a rule and separate arguments to prove its optimality for the two cases are required. Third, the risk functions considered in estimation and inference (e.g., the MSE) can be expressed in terms of bias and variance. This feature makes it easy to find the worst-case parameter values for an affine estimator: the maximum risk and maximum bias of an affine estimator are attained at the same parameter values, since the variance of an affine estimator is invariant to the parameter under known variance. This property cannot be exploited for my minimax regret problem, since the regret does not admit the bias and variance decomposition. Instead, I make use of the decomposition of the regret into the error probablity and potential welfare loss to find the worst-case parameter values for the proposed decision rule and verify its optimality.
commentThe construction of the scalar statistic involves optimization similar to the one for minimax estimation and inference. Here, I compare the nonrandomized minimax regret rule with a plug-in rule based on a linear minimax mean squared error (MSE) estimator.\footnote{In Appendix (ref), I also compare the minimax regret rule with a hypothesis testing rule that chooses policy 1 if a hypothesis that supports policy 0 is rejected.} To define the alternative rule, let $\hat L_{\rm MSE}(\boldsymbol Y)=\boldsymbol{w}_{\rm MSE}'\boldsymbol{Y}$ be a {\it linear minimax MSE estimator} of $L(\theta)$, where $$ \boldsymbol{w}_{\rm MSE} \in\arg\min_{\boldsymbol w\in\mathbb{R}^n}\sup_{\theta\in\Theta}\mathbb{E}_\theta[(\boldsymbol{w}'\boldsymbol{Y}-L(\theta))^2]. $$ $\hat L_{\rm MSE}(\boldsymbol Y)$ has the smallest maximum MSE within the class of linear estimators. I define the {\it plug-in MSE rule} as $\delta_{\rm MSE}(\boldsymbol Y)=\mathbf{1}\{\hat L_{\rm MSE}(\boldsymbol Y)\ge 0\}$, which makes a choice according to the sign of the linear minimax MSE estimator of $L(\theta)$. donoho1994 characterizes $\hat L_{\rm MSE}(\boldsymbol Y)$ using the modulus of continuity. Let $\epsilon_{\rm MSE}>0$ solve $\frac{\epsilon^2}{\epsilon^2+\sigma^2}=\frac{\omega'(\epsilon)\epsilon}{\omega(\epsilon)}$. The linear minimax MSE estimator is then given by $\hat L_{\rm MSE}(\boldsymbol Y)=\frac{\omega'(\epsilon_{\rm MSE})}{\epsilon_{\rm MSE}}\boldsymbol{m}(\theta_{\epsilon_{\rm MSE}})'\boldsymbol Y$, where $\theta_{\epsilon_{\rm MSE}}$ attains the modulus of continuity at $\epsilon_{\rm MSE}$ with $\|\boldsymbol m(\theta_{\epsilon_{\rm MSE}})\|=\epsilon_{\rm MSE}$. The plug-in MSE rule is $\delta_{\rm MSE}(\boldsymbol Y)=\mathbf{1}\{\boldsymbol{m}(\theta_{\epsilon_{\rm MSE}})'\boldsymbol Y\ge 0\}$. Recall that the minimax regret rule is $\delta^*(\boldsymbol Y)=\mathbf{1}\{\boldsymbol{m}(\theta_{\epsilon^*})'\boldsymbol Y\ge 0\}$ if $\sigma > 2\phi(0)\frac{\omega(0)}{\omega'(0)}$, where $\epsilon^*$ solves $\max_{\epsilon\in [0,\tau^*\sigma]}\omega(\epsilon)\Phi(-\epsilon/\sigma)$. \begin{proposition}[Comparison with Plug-in MSE Rule] Suppose that $\omega(\cdot)$ is differentiable with $\omega'(0)>0$, and let $\epsilon_{\rm MSE}>0$ solve $\frac{\epsilon^2}{\epsilon^2+\sigma^2}=\frac{\omega'(\epsilon)\epsilon}{\omega(\epsilon)}$ and $\epsilon^*$ solve $\max_{\epsilon\in [0,\tau^*\sigma]}\omega(\epsilon)\Phi(-\epsilon/\sigma)$. Then, $\epsilon^*< \epsilon_{\rm MSE}$. \end{proposition} \begin{proof} See Appendix (ref). \end{proof} Since $\delta^*(\boldsymbol Y)=\mathbf{1}\{\boldsymbol{m}(\theta_{\epsilon^*})'\boldsymbol Y\ge 0\}=\mathbf{1}\{\hat L_{\epsilon^*}(\boldsymbol Y)\ge 0\}$, the minimax regret rule $\delta^*(\boldsymbol Y)$ can be viewed as a rule that makes a choice according to the sign of the linear estimator $\hat L_{\epsilon^*}(\boldsymbol Y)$. Proposition (ref) implies that the corresponding linear estimator $\hat L_{\epsilon^*}(\boldsymbol Y)$ places more importance on the bias than on the variance compared with the linear minimax MSE estimator $\hat L_{\rm MSE}(\boldsymbol Y)$. This result suggests that the plug-in MSE rule is not necessarily optimal under the minimax regret criterion.

Application to Eligibility Cutoff Choice

In many policy domains, the eligibility for treatment is determined based on an individual's observable characteristics. In this section, I demonstrate how my framework can be valuable for using data collected under the status quo eligibility criterion to decide whether to change it to a new one. This approach does not require conducting a randomized experiment that directly evaluates the performance of the status quo and new criteria.

Setup

Consider the following special case of Example (ref). For each unit $i=1,..,n$, we observe a fixed running variable $x_i\in\mathbb{R}$, a binary treatment status $d_i\in\{0,1\}$, and an outcome $Y_i\in\mathbb{R}$. The eligibility for treatment is determined based on whether the running variable exceeds a specific cutoff $c_0\in \mathbb{R}$, so that $d_i=\mathbf{1}\{x_i\ge c_0\}$. Suppose

align*[align* omitted — 107 chars of source]

where $f:\mathbb{R}\times\{0,1\}\rightarrow \mathbb{R}$ is an unknown function and plays the role of the parameter $\theta$ and $\sigma^2(x_i,d_i)>0$. We interpret $f(x,d)$ as the conditional mean potential outcome under treatment status $d\in\{0,1\}$ given $x$. \sloppy We can write the model in a vector form $\boldsymbol Y \sim{\cal N}(\boldsymbol m(f), \boldsymbol\Sigma)$, where $\boldsymbol Y=(Y_1,...,Y_n)'$, $\boldsymbol m(f)=(f(x_1,d_1),...,f(x_n,d_n))'$, and $\boldsymbol\Sigma={\rm diag}(\sigma^2(x_1,d_1),...,\sigma^2(x_n,d_n))$.

Now, suppose we are interested in changing the cutoff from $c_0$ to a specific value $c_1$. For illustration, assume $c_1<c_0$. The welfare under the cutoff $c_a$, $a\in\{0,1\}$, is given by $$ W_a(f) = \int [f(x,1)\mathbf{1}\{x\ge c_a\}+f(x,0)\mathbf{1}\{x< c_a\}]d\nu(x) $$ for some known measure $\nu$. An implicit assumption made here is that $f$ is invariant to the cutoff change. For an illustration of the results, I use an empirical measure as $\nu$, for which the welfare is the unweighted sample average: $ W_a(f) = \frac{1}{n}\sum_{i=1}^n[f(x_i,1)\mathbf{1}\{x_i\ge c_a\}+f(x_i,0)\mathbf{1}\{x_i< c_a\}]. $ I define the welfare contrast between the two cutoffs as $$ L(f)=\frac{n}{\tilde n}(W_1(f)-W_0(f))=\frac{1}{\tilde n}\sum_{i=1}^n\mathbf{1}\{c_1\le x_i< c_0\}[f(x_i,1)-f(x_i,0)], $$ where $\tilde n=\sum_{i=1}^n\mathbf{1}\{c_1\le x_i<c_0\}$ denotes the number of units between the two cutoffs $c_1$ and $c_0$, whose treatment status would be changed if the cutoff were changed. I scale $W_1(f)-W_0(f)$ by $n/\tilde n$ so that $L(f)$ represents the sample average treatment effect for these units. This scaling does not change the form of a minimax regret rule. Figure (ref) presents an example of this setting. Panel (a) shows an example of the conditional mean potential outcome function $f$. In this example, the welfare contrast is given by $L(f)=\frac{1}{2}\sum_{i\in\{2,3\}}[f(x_i,1)-f(x_i,0)]$, which is the sample average treatment effect for $x_2$ and $x_3$.

figure[figure omitted — 1,004 chars of source]

To point or partially identify $L(f)$, suppose that $f\in{\cal F}$ for some function class ${\cal F}$. Here, I focus on the {\it Lipchitz class} $$ {\cal F}_{\rm Lip}(C)=\{f:|f(x,d)-f(\tilde x,d)|\le C|x-\tilde x| \text{ for every } x, \tilde x\in\mathbb{R} \text{ and } d\in\{0,1\}\}. $$ The Lipschitz constraint bounds the maximum possible change in $f(x,d)$ in response to a shift in $x$ by one unit. I assume the Lipschitz constant is common for $f(\cdot,1)$ and $f(\cdot,0)$ for simplicity; it is possible to impose two separate constants. Other possible function classes include the class of functions with a known bound on the second derivative, as used by Imbens2019RDD for inference in RD designs.

To illustrate how the Lipschitz constraint allows one to partially identify $L(f)$, I present the upper bound on $L(f)$. The lower bound can be obtained analogously. Let ${\cal M}=\{\boldsymbol m(f):f\in {\cal F}_{\rm Lip}(C)\}$ and $x_{+,{\rm min}}=\min\{x_i:x_i\ge c_0\}$ be the value of $x$ of the treated unit closest to the original cutoff $c_0$. The upper bound on $L(f)$ when $\boldsymbol m(f)=\boldsymbol{\mu}\in {\cal M}$ is given by

align*[align* omitted — 400 chars of source]

where I define $\mu_{+,\min}=\mu_i$ for the unit $i$ with $x_i=x_{+,\min}$. The second equality holds, since $f(x_i,0)=\mu_i$ for any $i$ with $x_i<c_0$ (i.e., $d_i=0$) and $f(x_i,1)=\mu_i$ for any $i$ with $x_i\ge c_0$ (i.e., $d_i=1$). The last equality holds, since the upper bound on $f(x_i,1)$ for any unit $i$ with $x_i<c_0$ is shown to be $\mu_{+,\min}+C(x_{+,{\rm min}} - x_i)$ under the Lipschitz constraint. The upper bound $\bar I(\boldsymbol{\mu})$ increases with the Lipschitz constant $C$ and weakly increases with the size of cutoff change $|c_1-c_0|$ (holding $c_0$ fixed). In the example presented in Figure (ref), $x_4$ is the treated unit closest to the original cutoff $c_0$. The two dashed lines in Panel (b) indicate the upper and lower bounds on the function $f(x,1)$ on the range of $x<x_4$. In this example, the upper bound on $L(f)$ is given by $\bar I(\boldsymbol{\mu})=\frac{1}{2}\sum_{i\in\{2,3\}}[\mu_4+C(x_4 - x_i)-\mu_i]$.

comment\paragraph{Identified Set of $L(f)$.} Imposing $f\in{\cal F}_{\rm Lip}(C)$ is not strong enough to uniquely determine $L(f)$ from a given value of the reduced-form parameter $\boldsymbol m(f)=(f(x_1,d_1),...,f(x_n,d_n))'$. Nevertheless, it produces an informative identified set of $L(f)$ $$ \{L(f):f(x_i,d_i)=\mu_i, i=1,...,n, f\in {\cal F}_{\rm Lip}(C)\} $$ from the knowledge that $\boldsymbol m(f)=\boldsymbol \mu$, since it gives finite upper and lower bounds on $f(x,d)$ for every $(x,d)\in\mathbb{R}\times\{0,1\}$.\footnote{The upper bound on $f(x,d)$ is $\min_{i:d_i=d}(\mu_i+C|x_i-x|)$. The lower bound on $f(x,d)$ is $\max_{i:d_i=d}(\mu_i-C|x_i-x|)$.} \textcolor{red}{Figure XX illustrates the upper and lower bounds on $f(x,d)$.} The larger the cutoff change is, the larger the identified set of $L(f)$ is. Also, as $C$ increases, the identified set of $L(f)$ becomes larger.

Note that we are not interested in the identified set per se, but are interested in using the sample $\boldsymbol{Y}$ to choose between the two cutoffs given the specified function class ${\cal F}$. For a decision rule $\delta:\mathbb{R}^n\rightarrow[0,1]$, $\delta(\boldsymbol y)\in [0,1]$ represents the probability of changing the cutoff from $c_0$ to $c_1$ when the realized sample is $\boldsymbol y$. Alternatively, we can interpret $\delta(\boldsymbol y)$ as the fraction of individuals to whom we would assign treatment within the units between $c_1$ and $c_0$. In Section (ref), I derive a minimax regret rule when the welfare is the sample average outcome and ${\cal F}={\cal F}_{\rm Lip}(C)$. The form of the rule depends on the empirical distribution of $x_i$, the two cutoffs $c_0$ and $c_1$, the Lipschitz constant $C$, and the conditional variance $\sigma^2(x_i,d_i)$, all of which are treated as known. In practice, the policymaker must specify $C$ and $\sigma^2(x_i,d_i)$ to implement the rule. In Section (ref), I provide practical guidance on how to specify them.

Minimax Regret Rule

To apply the results in Section (ref), I normalize $\tilde{\boldsymbol Y}=\boldsymbol\Sigma^{-1/2}\boldsymbol Y=(Y_1/\sigma(x_1,d_1),...,Y_n/\sigma(x_n,d_n))'$ and $\tilde{\boldsymbol m}(f)=\boldsymbol\Sigma^{-1/2}\boldsymbol m(f)=(f(x_1,d_1)/\sigma(x_1,d_1),...,f(x_n,d_n)/\sigma(x_n,d_n))'$, so that $\tilde{\boldsymbol Y} \sim{\cal N}(\tilde{\boldsymbol m}(f), \boldsymbol I_n)$. Then, $\omega(\epsilon)=\sup\{L(f): \|\tilde {\boldsymbol m}(f)\|\le \epsilon,f\in{\cal F}_{\rm Lip}(C)\}$ is the value of

align[align omitted — 213 chars of source]

The unknown parameter $f$ is infinite dimensional, but the objective and the norm constraint $\sum_{i=1}^n\frac{f(x_i,d_i)^2}{\sigma^2(x_i,d_i)}\le\epsilon^2$ depend on $f$ only through its values at $(x_1,0),...,(x_n,0),(x_1,1),...,(x_n,1)$. By a slight modification of Theorem 2.2 in Armstrong2021ATE, this optimization problem can be reduced to the following problem:

align[align omitted — 332 chars of source]

A solution to ((ref)) exists, since the objective function is continuous and the set of the vectors of $2n$ unknowns that satisfy the constraints is closed and bounded. Once we find a solution $(f(x_i,0),f(x_i,1))_{i=1,...,n}$, we can always find a function $f\in{\cal F}_{\rm Lip}(C)$ that interpolates the points $(x_i,f(x_i,0)),(x_i,f(x_i,1))$, $i=1,...,n$ BELIAKOV2006lipschitz, which is a solution to the original problem ((ref)). Problem ((ref)) is a finite-dimensional convex optimization problem with $2n$ unknowns, one quadratic and $2n(n-1)$ linear constraints, and a linear objective function, and can be solved using off-the-shelf convex optimization packages.\footnote{In the empirical application in Section (ref), I use CVXPY, a Python-embedded modeling language for convex optimization problems diamond2016cvxpy,agrawal2018rewriting.}

The following result derives a minimax regret rule.

proposition[Minimax Regret Rule for Eligibility Cutoff Choice] Consider the setup in Section (ref) with $L(f)=\frac{1}{\tilde n}\sum_{i:c_1\le x_i< c_0}[f(x_i,1)-f(x_i,0)]$ and ${\cal F}={\cal F}_{\rm Lip}(C)$. For simplicity, suppose $x_i\neq x_j$ for any $i\neq j$. Let $\omega(\epsilon)$ be the value of ((ref)) for $\epsilon\ge 0$, $\epsilon^*\in\arg\max_{\epsilon\in[0,\tau^*]}\omega(\epsilon)\Phi(-\epsilon)$, and $(f_{\epsilon^*}(x_i,0),f_{\epsilon^*}(x_i,1))_{i=1,...,n}$ solve ((ref)) for $\epsilon=\epsilon^*$. Then, the following decision rule is minimax regret: \begin{align*} \delta^*(\boldsymbol Y)=\begin{cases} \mathbf{1}\left\{\sum_{i=1}^nf_{\epsilon^*}(x_i,d_i)Y_i/\sigma^2(x_i,d_i)\ge 0\right\} & if s^*<\bar\sigma,\\ \mathbf{1}\left\{Y_{+,\min}-\frac{1}{\tilde n}\sum_{i:c_1\le x_i<c_0}Y_i\ge 0\right\} & if s^*=\bar\sigma,\\ \Phi\left(\dfrac{Y_{+,\min}-\frac{1}{\tilde n}\sum_{i:c_1\le x_i<c_0}Y_i}{((s^*)^2-\bar\sigma^2)^{1/2}}\right) & if s^*>\bar\sigma, \end{cases} \end{align*} where $x_{+,{\rm min}}=\min\{x_i:x_i\ge c_0\}$, $Y_{+,\min}=Y_i$ for the unit $i$ with $x_i=x_{+,\rm min}$, $\bar\sigma=(\sigma^2(x_{+,{\rm min}},1)+\frac{1}{\tilde n^2}\sum_{i:c_1\le x_i<c_0}\sigma^2(x_i,0))^{1/2}$, and $s^*=2\phi(0)C\frac{1}{\tilde n}\sum_{i:c_1\le x_i<c_0}[x_{+,{\rm min}} - x_i]$. Furthermore, the maximum regret of $\delta^*$ over ${\cal F}_{\rm Lip}(C)$ is attained at $-f^*$ and $f^*$, where $f^*$ is any function in ${\cal F}_{\rm Lip}(C)$ that interpolates the points $(x_i,f_{\epsilon^*}(x_i,0)),(x_i,f_{\epsilon^*}(x_i,1))$, $i=1,...,n$.
proofIn Appendix (ref), I derive a solution to ((ref)) for any sufficiently small $\epsilon\ge0$, which provides closed-form expressions for $\omega(0)$, $\omega'(0)$, and $\boldsymbol{w}^*$. The results then follow from an application of Theorem (ref) and simple calculations.
commentI now apply Theorem (ref) to derive the minimax regret rule. To simplify the exposition, I assume that $x_i\neq x_j$ for any $i\neq j$, $i,j=1,...,n$, in what follows. Let $\tilde n=\sum_{i=1}^n\mathbf{1}\{c_1\le x_i < c_0\}$ denote the number of units whose treatment status would be changed if the cutoff were changed. Recall $x_{+,{\rm min}}=\min\{x_i:x_i\ge c_0\}$, and let $\sigma_{+,\rm{\min}}^2=\sigma^2(x_{+,{\rm min}},1)$ and $\bar\sigma=(\tilde n^2\sigma_{+,\rm{\min}}^2+\sum_{i:c_1\le x_i<c_0}\sigma^2(x_i,0))^{1/2}$. In Appendix (ref), I derive a closed-form solution to the problem ((ref)) for any sufficiently small $\epsilon\ge0$, which provides the following expressions for $\omega(0)$, $\omega'(0)$, and $\boldsymbol{w}^*=(w_1^*,...,w_n^*)'$: $$ \omega(0)=\frac{C}{n}\sum_{i=1}^n\mathbf{1}\{c_1\le x_i < c_0\}[x_{+,{\rm min}} - x_i],~~~\omega'(0)=\frac{\bar\sigma}{n}, $$ and \begin{align} w_i^*=\begin{cases} 0 & if x_i<c_1 or x_i>x_{+,{\rm min}},\\ -\dfrac{\sigma(x_i,0)}{\bar\sigma} & if c_1\le x_i<c_0,\\ \dfrac{\tilde n\sigma_{+,\rm{\min}}}{\bar\sigma} & if x_i=x_{+,{\rm min}}. \end{cases} \end{align} Let $\epsilon^*\in\arg\max_{\epsilon\in[0,\tau^*]}\omega(\epsilon)\Phi(-\epsilon)$ and $(f_{\epsilon^*}(x_i,0),f_{\epsilon^*}(x_i,1))_{i=1,...,n}$ denote a solution to the problem ((ref)) for $\epsilon=\epsilon^*$. By Theorem (ref), the following rule is minimax regret: \begin{align*} \delta^*(\boldsymbol Y)=\begin{cases} \mathbf{1}\left\{\sum_{i=1}^nf_{\epsilon^*}(x_i,d_i)Y_i/\sigma^2(x_i,d_i)\ge 0\right\} & if s^*<1,\\ \mathbf{1}\left\{\sum_{i=1}^nw_i^*Y_i/\sigma(x_i,d_i)\ge 0\right\} & if s^*=1,\\ \Phi\left(\dfrac{\sum_{i=1}^nw_i^*Y_i/\sigma(x_i,d_i)}{((s^*)^2-1)^{1/2}}\right) &\text{ if } s^*>1, \end{cases} \end{align*} where \begin{align*} s^* \coloneqq 2\phi(0)\frac{\omega(0)}{\omega'(0)}=2\phi(0)C\sum_{i=1}^n\mathbf{1}\{c_1\le x_i < c_0\}[x_{+,{\rm min}} - x_i]/\bar\sigma. \end{align*}

The minimax regret rule is randomized or nonrandomized, depending on $s^*$ and $\bar\sigma$. $s^*$ is increasing in the Lipschitz constant $C$ and nondecreasing in the size of cutoff change $|c_1-c_0|$. $\bar\sigma$ is the standard deviation of $Y_{+,\min}-\frac{1}{\tilde n}\sum_{i:c_1\le x_i<c_0}Y_i$, and therefore increases with the variances of the treated unit closest to the status quo cutoff and the untreated units between the two cutoffs. If the Lipschitz constant $C$ or the cutoff change is large relative to the variances so that $s^*>\bar\sigma$, the minimax regret rule is a randomized rule based on $Y_{+,\min}-\frac{1}{\tilde n}\sum_{i:c_1\le x_i<c_0}Y_i$. In view of the results in Section (ref), $Y_{+,\min}-\frac{1}{\tilde n}\sum_{i:c_1\le x_i<c_0}Y_i$ can be interpreted as an estimator of $L(f)$ that minimizes the maximum squared bias over ${\cal F}_{\rm Lip}(C)$ among all estimators. As $C$ or the cutoff change increases, $(s^*)^2-\bar\sigma^2$ increases, and the decision is more randomized given the realization of the estimator.

On the other hand, if $s^*<\bar\sigma$, the minimax regret rule is a nonrandomized rule based on $\sum_{i=1}^nf_{\epsilon^*}(x_i,d_i)Y_i/\sigma^2(x_i,d_i)$. Using the results in Appendix (ref), we can show that if $s^*$ is marginally below $\bar\sigma$, $\sum_{i=1}^nf_{\epsilon^*}(x_i,d_i)Y_i/\sigma^2(x_i,d_i)$ is proportional to $Y_{+,\min}-\frac{1}{\tilde n}\sum_{i:c_1\le x_i<c_0}Y_i$. As the Lipschitz constant $C$ or cutoff change decreases or the variances increase so that $s^*$ becomes sufficiently smaller than $\bar\sigma$, the weighted sum assigns nonzero weights to some of the units with $x_i<c_1$ or $x_i>x_{+,\rm min}$. In Section (ref), I numerically examine the relationship between the weights and the choice of $C$ in the empirical application; see Figure (ref).

commentTo understand how the rule differs across different values of the Lipschitz constant $C$, note that suppose first that the magnitude of $C$ is moderate so that $s^*$ is marginally smaller than $1$. In this case, $\epsilon^*>0$ tends to be sufficiently small, which implies that $\frac{f_{\epsilon^*}(x_i,d_i)/\sigma(x_i,d_i)}{\epsilon^*}=w_i^*$ by Proposition (ref) in Appendix (ref). The minimax regret rule for this case is given by \begin{align*} \delta^*(\boldsymbol Y) &=\mathbf{1}\left\{\sum_{i=1}^nw_i^*Y_i/\sigma(x_i,d_i)\ge 0\right\}=\mathbf{1}\left\{Y_{+,\min}-\frac{1}{\tilde n}\sum_{i:c_1\le x_i<c_0}Y_i\ge 0\right\}, \end{align*} where I define $Y_{+,\min}=Y_i$ for the unit $i$ with $x_i=x_{+,\rm min}$. $Y_{+,\min}-\frac{1}{\tilde n}\sum_{i:c_1\le x_i<c_0}Y_i$ is the difference between the outcome of the treated unit closest to the status quo cutoff $c_0$ and the mean outcome across the untreated units between the two cutoffs $c_0$ and $c_1$. This difference can be interpreted as an estimator of the effect of the cutoff change. The outcomes of the other units are not used to construct the estimator. On the other hand, if the Lipschitz constant $C$ is small enough so that $s^*$ is substantially smaller than $1$, nonzero weights may be assigned to some of the other units, that is, $f_{\epsilon^*}(x_i,d_i)$ may be nonzero for some of the units with $x_i<c_1$ or $x_i>x_{+,\rm min}$. If the Lipschitz constant $C$ is large enough so that $s^*>1$, the minimax regret rule is a randomized rule based on $Y_{+,\min}-\frac{1}{\tilde n}\sum_{i:c_1\le x_i<c_0}Y_i$. Figure (ref) in Section (ref) illustrates how the weights differ across different values of $C$ in the empirical application. Whether the minimax regret rule is randomized or not depends not only on the Lipschitz constant $C$ but also on the cutoffs $c_0$ and $c_1$ and $\bar\sigma=(\tilde n^2\sigma_{+,\rm{\min}}^2+\sum_{i:c_1\le x_i<c_0}\sigma^2(x_i,0))^{1/2}$. To investigate their relationships, suppose that $\sigma^2(x_i,d_i)=\sigma^2$ for all $i$ for some $\sigma^2>0$ for simplicity. In this situation, $\bar\sigma=(\tilde n^2+\tilde n)^{1/2}\sigma$, and \begin{align*} s^* &=\frac{2\phi(0)C\frac{1}{\tilde n}\sum_{i=1}^n\mathbf{1}\{c_1\le x_i < c_0\}[x_{+,{\rm min}} - x_i]}{\left(1+1/\tilde n\right)^{1/2}\sigma}. \end{align*} $s^*$ is nonincreasing in $c_1$ since $\frac{1}{\tilde n}\sum_{i=1}^n\mathbf{1}\{c_1\le x_i < c_0\}[x_{+,{\rm min}} - x_i]$ and $\tilde n$ are nonincreasing in $c_1$.\footnote{Whether $s^*$ is increasing in $c_0$ or not depends on the empirical distribution of $x_i$.} Furthermore, $s^*$ is decreasing in $\sigma$. Therefore, the minimax regret rule is nonrandomized when $c_1$ is large (i.e., when the cutoff change $c_0-c_1$ is small) or $\sigma$ is large. The minimax regret rule is randomized otherwise.
commentand consider an asymptotic setup where $\tilde n\rightarrow\infty$ and $\frac{1}{\tilde n}\sum_{i=1}^n\mathbf{1}\{c_1\le x_i < c_0\}[x_{+,{\rm min}} - x_i]\rightarrow E_{P_X}[c_0-X|c_1\le X<c_0]$ for some probability measure $P_X$ as $n\rightarrow\infty$. In this situation, $\bar\sigma=(\tilde n^2+\tilde n)^{1/2}\sigma$, and \begin{align*} \sigma^* &=\frac{2\phi(0)C\frac{1}{\tilde n}\sum_{i=1}^n\mathbf{1}\{c_1\le x_i < c_0\}[x_{+,{\rm min}} - x_i]}{\left(1+1/\tilde n\right)^{1/2}\sigma}\rightarrow \frac{2\phi(0)C E_{P_X}[c_0-X|c_1\le X<c_0]}{\sigma} \end{align*} as $n\rightarrow\infty$. Note that $E_{P_X}[c_0-X|c_1\le X<c_0]$ is decreasing in $c_1$. Therefore, the minimax regret rule is nonrandomized when $c_1$ is large enough (i.e., the cutoff change $c_0-c_1$ is small enough) or $\sigma$ is large. The minimax regret rule is randomized otherwise.

Practical Issues

commentHere, I summarize the procedure for computing the minimax regret rule and discuss practical issues. Given the conditional variances $\sigma^2(x_i,d_i)$, $i=1,...,n$ and the Lipschitz constant $C$, the minimax regret rule is computed as follows. \begin{enumerate} • Compute $\sigma^*$ using the closed-form expression ((ref)). • If $1>\sigma^*$, find $\epsilon^*\in\arg\max_{\epsilon\in[0,\tau^*]}\omega(\epsilon)\Phi(-\epsilon)$ and compute $f_{\epsilon^*}$ that attains the modulus of continuity at $\epsilon^*$. For each $\epsilon\ge 0$, $\omega(\epsilon)$ is computed by solving the convex optimization problem ((ref)).\footnote{In the empirical application in Section (ref), I solve the convex optimization problem using CVXPY, a Python-embedded modeling language for convex optimization problems diamond2016cvxpy,agrawal2018rewriting.} An efficient method for computing $\epsilon^*$ is provided in Appendix (ref). • If $1\le \sigma^*$, compute $\boldsymbol w^*$ using the closed-form expression ((ref)). • Construct the decision rule according to ((ref)). \end{enumerate}

In practice, the conditional variance $\sigma^2(x_i,d_i)$ is unknown. A feasible version of the minimax regret rule is obtained by using a consistent estimator in place of the true $\sigma^2(x_i,d_i)$. The conditional variance can be estimated, for example, by applying a local linear regression to the squared residuals fan1998variance or by the nearest-neighbor variance estimator abadie2006matching. In the case in which unit $i$ represents a group of individuals and $Y_i$ is the sample mean outcome within group $i$, as in the empirical application in Section (ref), it is natural to use the conventional standard error of the sample mean as $\sigma(x_i,d_i)$.

Implementation of the minimax regret rule requires choosing the Lipschitz constant $C$. In principle, it is not possible to choose $C$ that applies to both sides of the cutoff $c_0$ in a data-driven way, since we only observe outcomes either under treatment or under no treatment on each side. It is, however, possible to estimate a lower bound on $C$. If $f\in{\cal F}_{\rm Lip}(C)$ is differentiable, a lower bound on $C$ is given by $\max\left\{\max_{\tilde x\ge c_0}\left\vert\frac{\partial f(\tilde x,1)}{\partial x}\right\vert,\max_{\tilde x< c_0}\left\vert\frac{\partial f(\tilde x,0)}{\partial x}\right\vert\right\}$, since $\left\vert\frac{\partial f(\tilde x,d)}{\partial x}\right\vert\le C$ for all $\tilde x$ and $d$. To estimate the lower bound, we could estimate the derivatives $\frac{\partial f(\tilde x,1)}{\partial x}$ for $\tilde x\ge c_0$ and $\frac{\partial f(\tilde x,0)}{\partial x}$ for $\tilde x< c_0$ by a local polynomial regression and then take the maximum of their absolute values. In practice, I suggest choosing the initial value of $C$ by estimating the lower bound or using application-specific knowledge, and considering a range of plausible values of $C$ to conduct a sensitivity analysis.

comment\section{Additional Implications of the Main Result} In this section, I present two implications of Theorem (ref). First, I discuss the difference between the minimax regret rule and a plug-in decision rule based on a linear minimax mean squared error (MSE) estimator. Second, I investigate the properties of the minimax regret rule when the welfare difference is point identified.
commentHence, I obtain the following result. \begin{corollary} Suppose that Assumptions (ref)-(ref) hold and that $\{c\tilde\theta^*:0\le c\le \tau^*\sigma\}\subset\Theta$. Then, the decision rule $\delta^*$ such that $$ \delta^*(\boldsymbol{y})=\mathbf{1}\left\{\boldsymbol{m}(\tilde\theta^*)'\boldsymbol{y}\ge 0\right\}, ~~~\boldsymbol{y}\in\mathbb{R}^n, $$ is minimax regret. The minimax risk is ${\cal R}(\sigma)=\tilde\omega \tau^*\sigma\Phi(-\tau^*)$. \end{corollary}

Empirical Application

I now illustrate my approach in an empirical application to consider whether to scale up the BRIGHT program in Burkina Faso.

Background and Data

With the aim of improving children's---especially girls'---educational outcomes in rural villages, the BRIGHT program constructed well-resourced village-based schools with three classrooms for grades 1 to 3 in 132 villages from 47 departments\footnote{Departments are the third-level administrative divisions of Burkina Faso, below regions and provinces.} during the period 2005 to 2008. The Ministry of Education determined the villages in which schools would be built through the following process. First, 293 villages were nominated based on low school enrollment rates. Second, the Ministry administered a survey in each village and assigned each village a score using a set formula. The formula attached a large weight to the estimated number of children to be served from the nominated and neighboring villages, giving additional weight to girls. The Ministry then ranked villages within each department and selected the top half of the villages to receive a school. For further details on the BRIGHT program and allocation process, see levy2009bright and Kazianga2013bright.

Since the school allocation was determined at department level, the cutoff score for program eligibility differed across departments. Following Kazianga2013bright, I define the {\it relative score} as the score for each village minus the cutoff score for the department the village belongs to. As a result, a village is eligible for the program when the relative score is larger than zero. Kazianga2013bright use the relative score as a running variable and evaluate the causal effect of the program on educational outcomes using an RD design.

I use the replication data for Kazianga2013bright's Kazianga2013bright results Kazianga2019data and consider whether we should expand the program. The dataset contains survey results on 30 households from 287 nominated villages, for a total sample of 23,282 children between the ages of 5 and 12. The survey was conducted in 2008---namely, 2.5 years after the start of the program. Table (ref) in Appendix (ref) reports summary statistics on child educational outcomes and characteristics.

I consider school enrollment as the target outcome. Since the score and eligibility are determined at village level, I use the village-level mean outcome---namely, the enrollment rate for each village. This setting fits into the setup in Section (ref), where $i$ represents a village, $Y_i$ is the sample enrollment rate of village $i$, $d_i$ is program eligibility, and $x_i$ is the relative score. The original cutoff is $c_0=0$; that is, $d_i=\mathbf{1}\{x_i\ge 0\}$. The parameter is a function $f:\mathbb{R}\times\{0,1\}\rightarrow \mathbb{R}$, where $f(x,d)$ represents the counterfactual enrollment rate conditional on the relative score if the eligibility status were set to $d\in\{0,1\}$. Since $Y_i$ is a village-level sample mean, it is plausible to assume that $Y_i$ is approximately normally distributed. I use the conventional standard error of the sample mean as the standard deviation of $Y_i$.\footnote{The sample enrollment rate is zero in 21 out of 287 villages. I exclude these villages from the analysis, since the standard error of $Y_i$ is zero.}

Hypothetical Policy Choice Problem

Suppose we are evaluating the program to decide whether to scale it up. Specifically, consider the following decision problem. The counterfactual policy is to build BRIGHT schools in previously ineligible villages whose relative scores are in the top 20%, which corresponds to lowering the cutoff from $0$ to $-0.256$; in Section (ref), I examine the sensitivity of the result to the choice of the new cutoff. I use the average enrollment rate across villages as the welfare criterion, so that the welfare effect of this policy relative to the status quo is $$ L(f)=\frac{1}{\tilde n}\sum_{i=1}^n\mathbf{1}\{-0.256\le x_i<0\}[f(x_i,1)-f(x_i,0)], $$ where $\tilde n=\sum_{i=1}^n\mathbf{1}\{-0.256\le x_i<0\}$ is the number of villages that would receive a school under the new policy. Some assumptions underlying this choice of $L(f)$ are: (i) we consider the set of villages in the sample rather than a new set of villages, and (ii) the counterfactual enrollment rate function $f$ remains constant over time between the period when the BRIGHT program was implemented and the period when the program is expanded.\footnote{Another underlying assumption is no spillover effects. The plausibility of this assumption can be indirectly verified, for example, based on how isolated each village is from other villages.} In principle, these assumptions can potentially be relaxed by a suitable choice of $L(f)$ and the function class ${\cal F}$, while I focus on the above choice of $L(f)$ for a simple illustration of my approach.

When deciding whether to implement the policy, it is important to consider the benefit relative to the cost. Kazianga2013bright provide an estimate of the cost of constructing a BRIGHT school, which is \$4,758 per village.\footnote{I assume that the cost is known and constant across villages. If village-level cost data are available, my framework allows for unknown and heterogeneous costs by introducing the cost model on top of the outcome model.} To incorporate the cost in the decision problem, suppose that the policymaker cares about the cost-effectiveness of this new policy relative to similar programs. Cost-effectiveness is defined as the ratio of the policy cost to the increase in the enrollment. I assume that it is optimal to implement the policy if its cost-effectiveness is smaller than that of a benchmark policy, denoted by $CE_0$---that is, $$ \frac{\text{\$4,758}}{416\cdot\frac{1}{\tilde n}\sum_{i=1}^n\mathbf{1}\{-0.256\le x_i<0\}[f(x_i,1)-f(x_i,0)]}\le CE_0, $$ where $416$ is the number of children per village. The denominator represents the increase in the average enrollment across villages that would receive a BRIGHT school under the new policy. For concreteness, I set the benchmark cost-effectiveness to $\$83.77$, which is the cost-effectiveness of a school construction program in Indonesia duflo2001school,Kazianga2013bright.\footnote{The cost per village and the cost-effectiveness of a school construction program in Indonesia are found in Tables A18 and A20, respectively, in Online Appendix of Kazianga2013bright. I compute the number of children per village by dividing the total enrollment by the enrollment rate reported in Table A17 in Online Appendix of Kazianga2013bright.} The above condition with $CE_0=\$83.77$ is equivalent to $$ \frac{1}{\tilde n}\sum_{i=1}^n\mathbf{1}\{-0.256\le x_i<0\}[f(x_i,1)-0.137-f(x_i,0)]\ge 0. $$ My method can be used to consider this decision problem by setting the outcome to $Y_i-0.137d_i$, where $0.137$ can be viewed as the policy cost measured in the unit of the enrollment rate. Alternatively, $0.137$ can be viewed as the effect on the enrollment rate of implementing a benchmark policy with the same cost as the new policy. I present the results for this scenario as well as for a benchmark scenario in which we ignore the policy cost.

I implement my method assuming that the counterfactual outcome function $f$ belongs to the Lipschitz class ${\cal F}_{\rm Lip}(C)$.\footnote{It is possible to incorporate the natural bound of $[0,1]$ on enrollment rates by setting the outcome to $Y_i-0.5$ and assuming $f(x,d)\in [-0.5,0.5]$ for all $(x,d)$ in addition to the Lipschitz constraint. I find that this adjustment does not change the minimax regret rule and its maximum regret for the range of the Lipschitz constant $C$ considered in this analysis.} The Lipschitz constant $C$ represents the maximum possible change in the enrollment rate in response to a one-unit change in the relative score. While the relative score is computed based on multiple village-level characteristics, it is largely based on the estimated number of students to be served. All other characteristics being equal, a one-unit increase in the relative score corresponds to around 100 additional children in the village, where the average number of children per village is 416. To obtain a reasonable range of $C$, I estimate a lower bound on $C$ using the method described in Section (ref), which yields the lower bound estimate of $0.149$.\footnote{I estimate $\frac{\partial f(x,0)}{\partial x}$ at $x\in\{-2.5,-2.45,...,-0.05\}$ and $\frac{\partial f(x,1)}{\partial x}$ at $x\in\{0.05,0.1,...,2.5\}$ by local quadratic regression and take the maximum of their absolute values. For local quadratic regression, I use the MSE-optimal bandwidth selection procedure of calonic2018bian, which can be implemented by R package “nprobust.”} I present the results for $C\in\{0.05,0.1,...,0.95,1\}$ and examine their sensitivity to the choice of $C$.

Results

Figure (ref) plots $\delta^*(\boldsymbol Y)$, the probability of choosing the new policy computed by the minimax regret rule, against the Lipschitz constant $C$. When $C< 0.6$, the minimax regret rule is nonrandomized. It chooses the new policy in the no-cost scenario and maintains the status quo in the scenario in which the policy cost is $0.137$. When $C\ge 0.6$, on the other hand, the minimax regret rule is randomized. The decisions become more mixed as $C$ increases. Given that the estimate of the lower bound on $C$ is 0.149, the minimax regret rule is nonrandomized when $C$ is less than four times the estimated lower bound. Under this reasonable range of $C$, the optimal decision is the same in each scenario.

figure[figure omitted — 757 chars of source]
figure[figure omitted — 1,132 chars of source]

If the minimax regret rule is nonrandomized, the rule is of the form $\delta^*(\boldsymbol Y)=\mathbf{1}\{\sum_{i=1}^nw_iY_i\ge 0\}$ for some weights $w_i$'s. Panels (a) and (b) of Figure (ref) plot the weight $w_i$ attached to each village against the relative score $x_i$ for $C=0.1$ and $C=0.5$, respectively. In the plots, the size of circles is proportional to the inverse of the standard error of the enrollment rate $Y_i$. For both $C=0.1$ and $C=0.5$, a few treated units just above the original cutoff (the solid vertical line) receive a positive weight, the untreated units between the original cutoff and the new cutoff (the dashed vertical line) receive a negative weight, and no other units receive any weight. When $C=0.1$, the weight tends to be larger for units with a smaller standard error. When $C=0.5$, a positive weight is attached only to the treated unit closest to the original cutoff. Also, the weights on the untreated units between the two cutoffs are almost identical. This situation corresponds to the minimax regret rule of the form $\delta^*(\boldsymbol Y)=\mathbf{1}\left\{Y_{+,\min}-\frac{1}{\tilde n}\sum_{i:c_1\le x_i<c_0}Y_i\ge 0\right\}$ discussed in Section (ref).

Comparison with Alternative Rules

I compare the minimax regret rule with several plug-in decision rules of the form $\delta(\boldsymbol{Y})=\mathbf{1}\{\hat L(\boldsymbol{Y})\ge 0\}$, where $\hat L(\boldsymbol{Y})$ is an estimator of the policy effect $L(f)$. I consider the following three estimators of $L(f)$. (i) The minimax affine MSE estimator donoho1994, described in Section (ref), under the Lipschitz class ${\cal F}_{\rm Lip}(C)$. (ii) The minimax affine MSE estimator under the additional assumption of constant conditional treatment effects. In other words, I construct the estimator assuming that $ {\cal F}=\{f\in{\cal F}_{\rm Lip}(C): f(x,1)-f(x,0)=f(\tilde x,1)-f(\tilde x,0) ~\text{for all}~x,\tilde x\} $. This estimation corresponds to first nonparametrically estimating the average treatment effect at the original cutoff and then extrapolating the effects on the units between the two cutoffs by the constant effects assumption. (iii) The polynomial regression estimator Kazianga2013bright.\footnote{Kazianga2013bright estimate the treatment effect at the cutoff, not the effect on units away from the cutoff. They apply global polynomial regression RD estimators to child-level data. } Given the degree of polynomial $p$, I first estimate the model $f(x,d)=\alpha_0+\alpha_1 x+\cdots+\alpha_p x^p +\beta_0 d+\beta_1 d\cdot x +\cdots +\beta_pd\cdot x^p$ by weighted least squares regression using $1/\sigma^2(x_i,d_i)$ as the weight. I then estimate $L(f)$ by $\frac{1}{\tilde n}\sum_{i=1}^n\mathbf{1}\{-0.256\le x_i<0\}[\hat f(x_i,1)-\hat f(x_i,0)]$, where $\hat f$ is the estimated function. This estimator relies on the functional form of $f$ to extrapolate $f(x_i,1)$ for the untreated units.

figure[figure omitted — 1,239 chars of source]

Panel (a) of Figure (ref) reports the estimated policy effects from the minimax affine MSE estimators with and without constant conditional treatment effects. Overall, these two estimators exhibit a similar pattern. While the estimated policy effects are larger than the policy cost when $C$ is close to zero, they are smaller than the policy cost when $C$ is moderate or large. For $C\ge 0.15$, the resulting decisions about whether to choose the new policy are the same as the decision made by the minimax regret rule until $C$ reaches 0.6, where the minimax regret rule starts to randomize. In contrast, the estimated policy effects from the polynomial regression estimators of degrees $1$ to $5$ exceed the policy cost, as reported in Panel (b) of Figure (ref). The estimates appear to be close to the simple mean outcome difference between eligible and ineligible villages that can be computed from Table (ref) in Appendix (ref). The resulting decisions differ from the decision made by the minimax regret rule.\footnote{The estimators presented here can be written as $\sum_{i=1}^nw_iY_i$ for some weights $w_i$'s. See Figure (ref) in Appendix (ref) for the plots of these weights. While the minimax affine MSE estimators attach weights to units just above the original cutoff and to units between the two cutoffs, polynomial regression estimators even attach weights to units further from the cutoffs.}

The above decisions are computed from a particular realization of the sample. To assess the ex ante performance of different decision rules, I compute the maximum regret of these rules when the true function class is ${\cal F}_{\rm Lip}(C)$.\footnote{I compute the maximum regret of the minimax regret rule using the formula in Theorem (ref). For the other rules, I adapt the approach of ishihara2021meta to numerically calculate the maximum regret in this setup.} Panel (a) of Figure (ref) reports the result for the minimax regret rule and the plug-in rules based on the minimax affine MSE estimators with and without constant conditional treatment effects.\footnote{ The result for the plug-in rules based on polynomial regression estimators is omitted, since these rules turn out to have significantly larger maximum regret than the other rules.} The maximum regret of the plug-in MSE rule with constant conditional treatment effects is much larger than that of the other two, especially when the Lipschitz constant $C$ is large. The plug-in MSE rule without constant conditional treatment effects performs slightly worse than the minimax regret rule. The ratio of the maximum regret between the two rules is maximized at $C=0.6$, where the minimax regret rule starts to randomize. The maximized ratio is about 1.233.

figure[figure omitted — 1,418 chars of source]

Sensitivity Analysis

I conduct several sensitivity analyses to assess how the results depend on the problem specification. First, I examine the sensitivity of the decision from the minimax regret rule to the choice of the new eligibility cutoff. Figure (ref) in Appendix (ref) reports the results when the new policy builds schools in the top 10% or 30% of previously ineligible villages instead of the top 20%. As predicted by the result in Section (ref), the minimax regret rule switches from a nonrandomized rule to a randomized rule at a smaller Lipschitz constant $C$ when the fraction of the target villages is larger. When the fraction is 30%, the rule is nonrandomized and suggests that the new policy is not cost-effective as long as $C$ is less than $0.4$.

Second, I examine the sensitivity of the decision to the policy cost. I find that the minimax regret rule nonrandomly decides to maintain the status quo as long as the cost exceeds $0.10$, for the range of $C$ between its estimated lower bound of 0.149 and 0.6.

So far, I have constructed decision rules assuming that the Lipschitz constant $C$ is known, which is a crucial assumption in my theoretical analysis. To assess the sensitivity of the performance to misspecification of $C$, I construct decision rules assuming $C=0.3$ and then compute their maximum regret when the true value of $C$ lies in $\{0.05,0.1,...,0.95,1\}$. Panel (b) of Figure (ref) reports the result. The solid line indicates the “oracle” maximum regret, which can be achieved if we correctly specify $C$. The result shows that the plug-in MSE rule without constant conditional treatment effects performs slightly better than the minimax regret rule when the true $C$ is close to zero. On the other hand, the minimax regret rule outperforms the plug-in MSE rule with nonnegligible differences for any value of the true $C$ greater than 0.3. The result suggests that the minimax regret rule is more robust to misspecification of $C$ toward zero than the plug-in MSE rule.\footnote{The potential superiority of the minimax regret rule seems consistent with the theoretical results in the following way. As shown in Section (ref), when the true value of $C$ is large, the oracle minimax regret rule only uses the treated units just above the original cutoff and the untreated units between the original and new cutoffs (see Panel (b) of Figure (ref)). If the specified $C$ is smaller than the true value, the resulting minimax regret rule is closer to the oracle rule than the plug-in MSE rule, since the minimax regret rule places more importance on the bias than the minimax affine MSE estimator, as discussed in Section (ref). Therefore, it is expected that the minimax regret rule performs better than the plug-in MSE rule under misspecification of $C$ toward zero.}

Conclusion

This paper derives an optimal decision rule for a large class of policy decision problems. The framework introduced in this paper allows for infinite-dimensional parameters, various forms of parameter restrictions, and partial identification of social welfare. I illustrate my approach through an application to the problem of eligibility cutoff choice in an RD setup.

Another promising application of this framework lies in policy adoption decisions using a difference-in-differences design. Specifically, consider a group of units that have experienced a policy change and another group that has not. Suppose the policymaker needs to decide whether to implement the new policy for the latter group. The average policy effect on this group may not be point identified if either the parallel trends assumption is violated or the policy effect varies between the two groups. My framework can be applied to this problem by imposing a set of restrictions on the degree of the violation of parallel trends and the amount of heterogeneity in policy effects.

Future research may explore several theoretical directions. First, one of the crucial assumptions in my approach is knowledge of the parameter space, such as the smoothness parameters of a function class. It would be interesting to investigate the possibility of adaptation over a collection of parameter spaces---that is, achieving near-optimal worst-case regret simultaneously over multiple parameter spaces, as has been studied for estimation and inference problems. Second, my approach only covers a binary choice problem. It would be challenging but both theoretically and practically important to extend the analysis to a multiple or continuous policy space.

\singlespacing \onehalfspacing

\pagenumbering{arabic} \setcounter{footnote}{0}

center[center omitted — 175 chars of source]

Appendix (ref) contains the proof and a discussion of Theorem (ref). Appendix (ref) contains auxiliary lemmas and proofs of Lemmas (ref), (ref), and (ref) and Theorems (ref) and (ref). Appendix (ref) contains derivations and a computational procedure for Section (ref). Appendix (ref) presents asymptotic properties of a feasible decision rule in the case with unknown error distribution. Appendix (ref) contains additional results for the empirical application in Section (ref).