Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
223,683 characters · 30 sections · 135 citation commands
Optimal Decision Rules Under Partial Identification
\sloppy
A fundamental goal of empirical research in economics is to inform policy decisions. Evaluation of counterfactual policies often requires extrapolating from observables to unobservables. Without strong model restrictions such as functional form assumptions, the performance of each counterfactual policy may be only partially determined by observed data. In such situations, policy decision-making is challenging, since we have no clear understanding of which policy is the best. For example, a regression discontinuity (RD) design only credibly estimates the impact of treatment on individuals at the eligibility cutoff. Therefore, without restrictive assumptions such as constant treatment effects, whether to offer the treatment to those away from the cutoff is ambiguous. Even randomized controlled trials may provide only partial knowledge of the impact of a new intervention, as can happen if the experimental sample is an unrepresentative subset of the target population.
This paper studies the problem of using data to make policy decisions in settings in which social welfare under each policy may be only partially identified. Following the literature on statistical treatment choice Manski2004hetero, I formulate the policy decision problem as a statistical decision problem. The framework introduced in this paper allows for various types of restrictions on the structural parameter, which potentially leads to partial identification of social welfare. It builds on donoho1994's donoho1994 framework for optimal estimation in nonparametric regression models, which has recently been applied to estimation and inference on treatment-effect parameters armstrong2018optimal,Imbens2019RDD,kwon2020rd,Armstrong2021ATE,rambachan2023parallel,Chaisemartin2021aet. I extend the framework to study optimal policy choice in a wide range of empirical settings with partial identification. Examples of this paper's framework include treatment choice using experiments with imperfect internal or external validity Stoye2012minimax,ishihara2021meta; treatment choice using observational data under unconfoundedness with imperfect overlap; and policy adoption choice in difference-in-differences designs without exact parallel trends.
Specifically, in the setup described in Section (ref), the policymaker must decide between two policies, policy 1 and policy 0, to maximize social welfare. The difference in welfare between the two policies is given by $L(\theta)$, where $\theta$ is a possibly infinite-dimensional structural parameter that resides in a vector space $\mathbb{V}$, and $L:\mathbb{V}\rightarrow\mathbb{R}$ is a known linear function. If $\theta$ is known, it is optimal to choose policy 1 if $L(\theta)\ge 0$ and policy 0 if $L(\theta)<0$. The policymaker does not know $\theta$, but instead has access to a multivariate Gaussian sample $\boldsymbol Y=(Y_1,...,Y_n)'\in\mathbb{R}^n$ of the form
where $\boldsymbol{m}:\mathbb{V}\rightarrow\mathbb{R}^n$ is a known linear function and $\boldsymbol\Sigma$ is known. After observing $\boldsymbol{Y}$, the policymaker decides between policies 1 and 0. The main structural assumption is that $\theta$ belongs to a known set $\Theta\subset\mathbb{V}$ that is convex and centrosymmetric (i.e., $\theta\in\Theta$ implies $-\theta\in\Theta$), which encodes the policymaker's a priori knowledge of parameter restrictions. Depending on the restrictions, the welfare contrast $L(\theta)$ may or may not be point identified from the knowledge of the point-identified reduced-form parameter $\boldsymbol m(\theta)$.
As detailed in Section (ref), an example of this setup is the choice between assigning treatment to everyone in the population (policy 1) and assigning treatment to no one (policy 0). Suppose that the policymaker has access to data from a regression model $Y_i=f(x_i,d_i)+U_i$, where $f(x,d)$ represents the conditional mean potential outcome under treatment $d\in\{0,1\}$ given covariates $x$, and $\{(x_i,d_i)\}_{i=1}^n$ is treated as fixed. Suppose further that treating everyone is preferred to treating no one if the population average treatment effect is positive. This problem is a special case when we assume $(U_1,...,U_n)'\sim {\cal N}(\boldsymbol{0}, \boldsymbol\Sigma)$ and set $\boldsymbol Y=(Y_1,...,Y_n)'$, $\theta=f$, $\boldsymbol{m}(f)=(f(x_1,d_1),...,f(x_n,d_n))'$, and $L(f)=\int[f(x,1)-f(x,0)]dP_X$, where $P_X$ is the population distribution of the covariates and is assumed to be known. The parameter space $\Theta$ is a class of conditional mean potential outcome functions $f$ that satisfy, for example, some smoothness restrictions (e.g., bounds on derivatives or the linearity of a function), which leads to either point or partial identification of the average treatment effect $L(f)$.
The main contribution of this paper is to obtain a finite-sample optimal decision rule under the minimax regret criterion, which is a standard criterion used in the literature on statistical treatment choice Manski2004hetero,Manski2007missing,Stoye2009minimax,Stoye2012minimax,Kitagawa2018EWM. A decision rule is a mapping from the sample $\boldsymbol Y$ to the probability of choosing policy 1. The minimax regret criterion evaluates the performance of a decision rule based on its worst-case expected welfare loss, or {\it regret}, relative to the oracle welfare-maximizing policy $\mathbf{1}\{L(\theta)\ge 0\}$, where the worst-case scenario is considered over the parameter space $\Theta$. I derive a decision rule that minimizes the worst-case regret among all decision rules, with no functional-form restrictions imposed on the class of rules. When $\boldsymbol Y$ is non-Gaussian and/or its variance is unknown, a feasible version of this decision rule can be constructed by plugging in an estimated variance. Appendix (ref) provides conditions under which its maximum regret over a class of distributions of $\boldsymbol Y$ converges to that of a minimax regret rule as $n\rightarrow\infty$.
To solve the minimax regret problem, I use the {\it hardest one-dimensional subfamily} argument, which donoho1994 used to solve minimax affine estimation problems. The key idea is to search for the hardest one-dimensional subproblem---specifically, the one with the highest minimax risk among all subproblems with parameter spaces defined by one-dimensional linear subfamilies of the original parameter space $\Theta$. Then, verify that a minimax rule for the hardest one-dimensional subproblem is also minimax optimal for the original problem. Applying this strategy to the minimax regret problem is challenging due to the difference in the structure of the risk functions: Unlike standard risk functions for estimation, such as mean squared error (MSE), the regret cannot be decomposed into the bias and variance; instead, it can be decomposed into the probability of misidentifying the best policy and the potential welfare loss due to misidentification. To derive a minimax regret rule, I first characterize the hardest one-dimensional subproblem by optimizing a certain measure of the strength of the signal with respect to the best policy within subproblems (Lemma (ref)). The subproblem with the optimal level of the signal strength achieves the best balance between the probability of misidentification and the potential welfare loss. I then propose a specific minimax regret rule for the hardest subproblem, and prove its minimax regret optimality for the original problem (Theorem (ref)).
The results of this paper provide novel insights into how a minimax regret rule uses data to make decisions. First, the derived rule depends on the observations $Y_1,...,Y_n$ only through their weighted sum, $\sum_{i=1}^n{w}_iY_i$, although no such restrictions are imposed a priori. The weights ${w}_1,...,{w}_n$ can be calculated by solving a sequence of convex optimization problems, which is computationally and analytically tractable in leading examples.\footnote{Thus, this paper addresses a challenge in the application of statistical decision theory raised by Manski2020decision, who wrote (p. 2848): “The primary challenge to use of statistical decision theory is computational \ldots\ Future advances should continue to expand the scope of applications.”}
Second, the minimax regret rule is nonrandomized or randomized, depending on the strength of the parameter restrictions and the variance of $\boldsymbol{Y}$. Specifically, if the restrictions are strong or the variance of $\boldsymbol{Y}$ is large in a certain formal sense, the minimax regret rule is a nonrandomized threshold rule of the form $\mathbf{1}\left\{\sum_{i=1}^n{w}_iY_i\ge 0\right\}$. On the other hand, if the restrictions are weak or the variance of $\boldsymbol{Y}$ is small, the rule is a randomized threshold rule of the form $\mathbf{1}\left\{\sum_{i=1}^n{w}_iY_i+\xi\ge 0\right\}$, where $\xi$ is generated independent of $\boldsymbol{Y}$ according to a certain distribution. In the latter case, randomization plays the role of reducing the probability of misidentification under worst-case parameter values, and thereby leads to a reduction in worst-case regret. This result generalizes Stoye2012minimax's Stoye2012minimax from a specific univariate problem to a general class of multivariate problems.
Third, this paper sheds light on the connection between minimax regret treatment choice and minimax estimation. The weighted sum $\sum_{i=1}^n{w}_iY_i$ used by the minimax regret rule can be viewed as an estimator of the welfare contrast $L(\theta)$. In other words, the minimax regret rule can be viewed as a {\it plug-in} rule, which plugs the estimator $\sum_{i=1}^n{w}_iY_i$ (plus a random noise $\xi$ for the randomized rule) into the oracle optimal decision $\mathbf{1}\{L(\theta)\ge 0\}$. I show that this estimator is optimal in the sense of minimizing the worst-case squared bias (over $\Theta$) among all estimators of $L(\theta)$ subject to a certain bound on the variance (Theorem (ref)). Furthermore, I show that this estimator places more importance on bias than variance compared with the minimax affine MSE estimator in donoho1994 (Theorem (ref)).
Fourth, while this paper's main focus is on optimal rules under partial identification, my results are novel even under point identification for problems with restricted parameter spaces. When the welfare contrast $L(\theta)$ is point identified, the minimax regret rule is shown to always be a nonrandomized threshold rule, with its form depending on the strength and type of restrictions. For example, consider a linear regression model in which $\boldsymbol{m}(\theta)=\boldsymbol{X}\theta$ for some fixed $n\times k$ design matrix $\boldsymbol{X}$, and suppose $\Theta=\{\theta\in\mathbb{R}^k: \|\theta\|_p\le C\}$ for some known constants $C\ge 0$ and $p\ge 1$, where $\|\cdot\|_p$ denotes the $L_p$-norm. The results of this paper imply that the minimax regret rule makes decisions based on the sign of $L(\hat\theta)$, where $\hat\theta$ is an estimator for $\theta$ that resolves the bias-variance tradeoff in a certain way (e.g., the ridge estimator with the regularization parameter depending on $C$ for $p=2$). The result of hirano2009asymptotics applies to problems with no parameter restrictions (i.e., $\Theta=\mathbb{R}^k$), but does not apply to ones with restricted parameter spaces.
I demonstrate the practical relevance of my framework through an application to the problem of eligibility cutoff choice. In Section (ref), I consider a situation in which the eligibility for treatment (e.g., social, educational, or welfare programs) is determined based on whether the value of an individual's characteristic exceeds a certain cutoff, as in RD setups. The policymaker wants to change the cutoff to a specific new value if the welfare effect of the cutoff change is positive. Here, the welfare effect is defined as the average treatment effect across units whose treatment status would be changed under the new cutoff. In this context, a decision rule maps the data collected under the status quo cutoff to the probability of changing the cutoff. My results can be used to find an optimal decision rule for this problem, under various restrictions on the conditional mean potential outcome function that enable extrapolation from one side of the cutoff to the other. For illustration, I assume that the function satisfies Lipschitz continuity with a known Lipschitz constant (i.e., a bound on the first derivatives), which leads to partial identification of the welfare effect. Applying my general results, I show that the minimax regret rule makes a decision based on the difference between a weighted average of observed outcomes for the treated units and that for the untreated units, with the weight vector determined by the choice of the Lipschitz constant. The rule can easily be computed by solving finite-dimensional convex optimization problems.
Finally, in Section (ref), I apply this rule to the Burkinab\'{e} Response to Improve Girls' Chances to Succeed (BRIGHT) program, a school construction program in Burkina Faso Kazianga2013bright. With the aim of improving educational outcomes in rural villages, the program constructed primary schools in 132 villages from 2005 to 2008. To allocate schools, the Ministry of Education first computed a score that summarized village characteristics for each of the nominated 293 villages, then selected the highest-ranking villages to receive a school. Consider a policymaker who uses data collected after the completion of this program to decide whether to scale it up. As a hypothetical policy question, I consider whether to construct schools in the top 20% of previously ineligible villages, using the enrollment rate as the welfare measure. I impose the Lipchitz constraint on the counterfactual enrollment rates across villages. To consider policy costs, I assume that it is optimal to scale up the program if its cost-effectiveness is better than that of a similar policy. Applying my theoretical results, I find that the minimax regret rule nonrandomly decides not to scale up the program for a plausible range of the Lipschitz constant.
\paragraph{Related Literature.} This paper contributes to the literature on minimax regret statistical treatment choice under point identification hirano2009asymptotics,Stoye2009minimax,Stoye2012minimax,tetenov2012asymmetric and partial identification Manski2007missing,Stoye2012minimax. Stoye2012minimax derives a minimax regret rule in settings in which the experiment has imperfect validity, considering both Bernoulli and Gaussian models. My result generalizes Stoye2012minimax's Stoye2012minimax in Gaussian models (Proposition 7(iii)) by allowing for multivariate samples, three- or higher-dimensional parameters, and various forms of parameter spaces. Recently, ishihara2021meta consider the problem of deciding whether to introduce a new policy based on results from multiple external studies. They restrict attention to the class of nonrandomized threshold rules based on a weighted average of the sample and propose a way to numerically minimize the maximum regret. In contrast, I do not impose any restrictions on decision rules, and use the hardest one-dimensional subfamily argument to analytically derive a minimax regret rule. My approach does not involve numerically minimizing the maximum regret and, for some problems, offers a closed-form expression for a minimax regret rule. Since the initial version of this paper was circulated, there have been some advances in the literature. kitagawa2023partial,kitagawa2022nonlinear derive minimax fractional rules under squared welfare regret loss for both point and partial identification settings. olea2023partial point out the nonuniqueness of minimax regret rules for problems in which my minimax regret rule is randomized, and propose the least randomizing rule. Their work and mine are complementary; their analysis relies on the existence of a minimax regret rule based on a weighted sum of the sample, while I prove the existence of such a rule and derive its formula.
Broadly, this paper contributes to the literature on treatment choice and policy learning under partial identification, which has been growing in econometrics and statistics Manski2000ambiguity,manski2009diversified,Manski2010vaccine,manski2011are,manski2011partial,Manski2020decision,chamberlain2011bayesian,kasy2018taxation,russell2020policy,Mo2020robust,Kallus2020confounding,dadamo2022orthogonal,ben-michael2022safe,adjaho2023,christensen2022discrete,Kido2023locally.\footnote{An extensive literature examines the problem of learning treatment allocation policies that map an individual's covariates to a treatment. See Manski2004hetero,Dehejia2005decision,Stoye2009minimax,Stoye2012minimax,qian2011ind,Bhattacharya2012budget,Kitagawa2018EWM,kitagawa2021equal,Athey2021policy; and Mbakop2021penalized, among others, for point identification settings. My approach can be applied to partial identification settings in which the choice set consists of two treatment assignment policies. } Many recent studies, either implicitly or explicitly, consider the worst-case welfare (or welfare loss) over the identified set given the point-identified parameter as the loss function of a statistical decision problem. By contrast, this paper directly uses the welfare loss as the loss function without taking its worst-case value given the point-identified parameter, following the standard minimax regret criterion. Furthermore, they derive finite-sample bounds on the expected loss, its convergence rates, or the local asymptotic optimality of their proposed rules, while this paper derives a finite-sample exact optimal rule. Another difference, which is empirically relevant, is that many studies consider welfare parameters whose upper and lower bounds can be estimated at a parametric rate. In contrast, this paper covers welfare parameters whose bounds cannot be estimated at a parametric rate, such as the average treatment effect at a point of the running variable in a nonparametric RD setup.
The problem of eligibility cutoff choice considered in this paper is related to the literature on extrapolation away from the cutoff in RD designs, including rokkanen2015rd,angrist2015rd,dong2015rd,Bertanha2020rd,Bertanha2020many,Bennett2020rd; and Cattaneo2020multi. Unlike these papers, I explicitly consider the decision problem of whether to change the cutoff and derive an optimal decision rule. Recently, zhang2022rd consider the problem of learning cutoff-based policies under multi-cutoff designs and propose a maximin policy that is guaranteed to perform no worse than the existing policy.
In this section, I introduce my framework and provide examples to illustrate it. Section (ref) describes the statistical model that generates the data available to the policymaker. Section (ref) defines the policymaker's action set and the associated social welfare functions. Section (ref) introduces the minimax regret criterion as an optimality criterion for decision rules. Section (ref) provides a brief discussion of the relationship between this framework and existing ones. Finally, Section (ref) illustrates the framework using two examples.
Suppose that the policymaker observes a sample $\boldsymbol{Y}=(Y_1,...,Y_n)'\in \mathbb{R}^n$ of the form
where $\theta$ is an unknown parameter that lies in a known subset $\Theta$ of a vector space $\mathbb{V}$; $\boldsymbol{m}:\mathbb{V}\rightarrow \mathbb{R}^n$ is a known linear function; and $\boldsymbol\Sigma$ is a known, positive-definite $n\times n$ matrix. I allow $\theta$ to be an infinite-dimensional parameter such as a function.
The linearity of $\boldsymbol{m}$ is not necessarily restrictive. If we specify $\theta$ so that it contains each of the expected values of $Y_1,...,Y_n$ as its element, $\boldsymbol m$ is a function that extracts those expected values from $\theta$, which is linear in $\theta$.
This model allows the expected value of $\boldsymbol{Y}$ to depend on other observed variables such as covariates and treatment by treating them as fixed and subsuming them into $\boldsymbol m$ and $\boldsymbol \Sigma$. For example, a regression model with fixed regressors
is a special case in which $\boldsymbol{Y}=(Y_1,...,Y_n)'$, $\theta=f$, $\Theta$ is a class of functions, $\boldsymbol m(f)=(f(x_1),...,f(x_n))'$, and $\boldsymbol \Sigma={\rm diag}(\sigma^2(x_1),...,\sigma^2(x_n))$.
The normality of $\boldsymbol Y$ and the assumption of known variance are restrictive, but are often imposed to deliver finite-sample optimality results for statistical decision problems. In some cases, the normal model is motivated as an approximation to a finite-sample problem. Suppose that we observe an $n$-dimensional vector of statistics derived from the original data, which is an asymptotically normal estimator of its population counterpart. For example, the mean outcome difference between the treatment and control groups in a randomized experiment is a statistic that is asymptotically normal for the population mean difference. If we regard the $n$-dimensional vector of statistics as $\boldsymbol Y$, the normal model ((ref)) can be viewed as an asymptotic approximation. Also, in Appendix (ref), I consider an asymptotic framework in which the distribution of the error $\boldsymbol Y-\boldsymbol m(\theta)$ is unknown and the sample size $n$ goes to infinity. I propose a feasible decision rule and derive conditions under which its maximum regret over a class of distributions converges to that of a minimax regret rule as $n\rightarrow\infty$.
I assume that the parameter space $\Theta$ is convex and centrosymmetric (i.e., $\theta\in\Theta$ implies $-\theta\in\Theta$) throughout the paper. Typical parameter spaces considered in empirical analyses are convex. For example, in the regression model above, classes of functions with bounded derivatives are convex. The centrosymmetry simplifies the minimax analysis; see Remark (ref) in Section (ref) for the role of centrosymmetry. However, it rules out some shape restrictions. In the regression model above, the class of convex (or concave) functions is noncentrosymmetric.
Now, suppose that the policymaker is interested in choosing between two policies, policy $1$ and policy $0$, to maximize social welfare. Suppose that the welfare resulting from implementing policy $a\in\{0,1\}$ under $\theta$ is $W_a(\theta)$, where $W_a:\mathbb{V}\rightarrow \mathbb{R}$ is a known function specified by the policymaker. The welfare contrast between policy $1$ and policy $0$ is given by $$ L(\theta)\coloneqq W_1(\theta) - W_0(\theta). $$ I assume that $L:\mathbb{V}\rightarrow \mathbb{R}$ is a linear function. The optimal policy under $\theta$ is policy 1 if $L(\theta)>0$, policy 0 if $L(\theta)<0$, and either if $L(\theta)=0$.
One example of a welfare criterion is a weighted average of an outcome across individuals. For example, suppose a policy could change the outcome of each individual. Suppose also that we specify $\theta=(f_1(\cdot),f_0(\cdot))$, where $f_a(x)$ represents the counterfactual mean outcome under policy $a$ across individuals whose observed covariates are $x$. The welfare under policy $a$ can be defined, for example, by the population mean outcome $W_a(\theta)=\int f_a(x)dP_X$, where $P_X$ is the probability measure of covariates and is assumed to be known. In this case, the welfare contrast $L(\theta)=\int [f_1(x)-f_0(x)]dP_X$ is linear in $\theta$. On the other hand, the linearity of $L$ may rule out welfare criteria that depend on the distribution of the counterfactual outcome other than the mean. See kitagawa2021equal for such welfare criteria.
Importantly, this framework allows for cases in which $L(\theta)$ is not point identified. Let ${\cal M}\coloneqq\{\boldsymbol{m}(\theta):\theta\in\Theta\}\subset\mathbb{R}^n$ denote the set of possible values of the reduced-form parameter $\boldsymbol{m}(\theta)$. The {\it identified set} of $L(\theta)$ when $\boldsymbol m(\theta)=\boldsymbol \mu\in \mathbb{R}^n$ is defined as $$ I(\boldsymbol \mu)\coloneqq \{L(\theta):\boldsymbol m(\theta)=\boldsymbol \mu, \theta\in \Theta\}. $$ $I(\boldsymbol \mu)$ may contain multiple elements for some or all $\boldsymbol \mu\in{\cal M}$. If $I(\boldsymbol \mu)$ contains both positive and negative values, the superior policy is ambiguous even without sampling uncertainty.
This paper's goal is to provide an optimal decision rule for using data to make a policy decision. A (randomized) {\it decision rule} is a measurable function $\delta:\mathbb{R}^n\rightarrow [0,1]$, where $\delta(\boldsymbol{y})$ represents the probability of choosing policy $1$ when the realization of the sample $\boldsymbol{Y}$ is $\boldsymbol y$.\footnote{In contexts in which fractional treatment allocations, which assign treatment to a fraction of individuals in a population, are permitted, we can also interpret $\delta$ as a fractional rule. Here, $\delta(\boldsymbol y)$ represents the fraction of individuals to whom we would assign treatment. If the welfare of treating a fraction $a\in [0,1]$ is defined as $W_a(\theta)=W_0(\theta)+a(W_1(\theta)-W_0(\theta))$, the regret of a fractional rule is equal to that of a randomized rule.} I consider the minimax regret criterion as an optimality criterion for decision rules. To introduce it, I first define the {\it welfare regret loss} for policy choice $a\in\{0,1\}$ under $\theta$ as
The welfare regret loss $l(a,\theta)$ is the difference in welfare between the optimal policy and policy $a$ under $\theta$. If the policymaker chooses the superior policy, they do not incur any loss; otherwise, they incur a loss of the absolute value of the welfare contrast $L(\theta)$.
The {\it risk} or {\it regret} of decision rule $\delta$ under $\theta$ is the expected welfare regret loss $$ R(\delta,\theta)\coloneqq
$$ where $\mathbb{E}_\theta$ denotes the expectation taken with respect to $\boldsymbol{Y}$ under $\theta$. Using the regret as a performance measure allows one to consider not only the error probabilities ($1-\mathbb{E}_\theta[\delta(\boldsymbol{Y})]$ or $\mathbb{E}_\theta[\delta(\boldsymbol{Y})]$) but also the potential welfare loss ($|L(\theta)|$).
The regret of a decision rule can vary with $\theta$ over the parameter space $\Theta$. Generally, no rule uniformly dominates all other rules. The minimax regret criterion aggregates the regret over $\Theta$ by considering the {\it maximum} or {\it worst-case regret}, defined as $\sup_{\theta\in\Theta}R(\delta,\theta)$.
Let ${\cal R}(\Theta)$ denote the {\it minimax risk} or {\it minimax regret}, defined as
where ${\cal D}$ denotes the set of all decision rules. Under the minimax regret criterion, we aim to derive a {\it minimax regret} decision rule $\delta^*$, which satisfies $\sup_{\theta\in\Theta}R(\delta^*,\theta)={\cal R}(\Theta)$.
To further understand the minimax regret criterion, note that the above minimax problem can equivalently be written as $\inf_{\delta\in{\cal D}}\sup_{\theta\in\Theta}\left(\max_{a\in\{0,1\}}W_a(\theta)-U(\delta,\theta)\right)$, where $U(\delta,\theta)\coloneqq W_1(\theta)\mathbb{E}_\theta[\delta(\boldsymbol{Y})]+W_0(\theta)(1-\mathbb{E}_\theta[\delta(\boldsymbol{Y})])$ is the expected welfare of decision rule $\delta$ under $\theta$. Thus, the minimax regret criterion optimizes the uniform closeness of the expected welfare to the maximum attainable welfare. When the minimax risk ${\cal R}(\Theta)$ is small, a minimax regret rule uniformly achieves near-optimal expected welfare across all parameter values.
\sloppy
Alternative optimality criteria include the maximin criterion, which maximizes the worst-case expected welfare. Specifically, one aims to solve $\sup_{\delta\in{\cal D}}\inf_{\theta\in\Theta}U(\delta,\theta)$. It has been pointed out that the maximin criterion is unreasonably pessimistic and can lead to pathological decision rules Manski2004hetero,Stoye2009minimax; see olea2023partial for such a result in the setting of this paper.\footnote{For example, suppose the policymaker perfectly knows that the welfare of the status quo policy ($a=0$) is given by $W_0(\theta)=w_0$ for some constant $w_0$. If the welfare of the new policy ($a=1$) can be less than $w_0$ under at least one parameter value (i.e., $\inf_{\theta\in\Theta}W_1(\theta)<w_0$), then the decision rule that always maintains the status quo regardless of the data (i.e., $\delta(\boldsymbol{Y})=0$) is optimal under the maximin criterion.} By contrast, the minimax regret criterion optimizes the worst-case expected welfare relative to what is achievable at a given parameter value, which often leads to nontrivial decision rules. Another alternative is the Bayes criterion, which optimizes the average risk over a prior. That is, one aims to solve $\sup_{\delta\in{\cal D}}\int U(\delta,\theta)d\pi(\theta)$ (or equivalently $\inf_{\delta\in{\cal D}}\int R(\delta,\theta)d\pi(\theta)$), where $\pi$ is a prior on $\Theta$. See, for example, chamberlain2011bayesian and kasy2018taxation for Bayesian treatment choice. This paper focuses on the minimax regret criterion, following prior treatment choice studies Manski2004hetero,Manski2007missing,hirano2009asymptotics,Stoye2009minimax,Stoye2012minimax,tetenov2012asymmetric,Kitagawa2018EWM. Studying optimal rules under the Bayes or other possible criteria is beyond the scope of this paper.
Here, I discuss the relationship between the above framework and existing ones. For minimax regret treatment choice, the framework in this paper generalizes the univariate Gaussian problems with two-dimensional parameters in Stoye2012minimax to accommodate multivariate samples, parameters of three or higher dimensions, and various types of parameter restrictions. It also generalizes the limiting version of minimax regret problems under parametric models studied by hirano2009asymptotics to accommodate partially identified welfare contrasts and restricted parameter spaces. In the setting described in Section (ref), ishihara2021meta derive a minimax regret rule within the class of nonrandomized threshold rules based on a weighted average of the sample, namely $\delta(\boldsymbol{Y})=\mathbf{1}\{\sum_{i=1}^nw_iY_i\ge 0\}$ with $\sum_{i=1}^nw_i=1$. Their result does not require the convexity of the parameter space, but instead assumes that the parameter space is invariant to the addition of vectors of ones, which excludes bounded parameter spaces.\footnote{A generalization of their invariance assumption to the general setup of this paper is as follows: There exists $\iota\in\Theta$ such that $L(\iota)=1$ and $\theta+c\iota\in\Theta$ for all $\theta\in\Theta$ and $c\in\mathbb{R}$. Under this condition and the centrosymmetry of $\Theta$, it is possible to extend their approach to derive a minimax regret rule among rules of the form $\delta(\boldsymbol{Y})=\mathbf{1}\{\boldsymbol{w}'\boldsymbol{Y}\ge 0\}$ with $\boldsymbol{w}\in\mathbb{R}^n$ and $\boldsymbol{w}'\boldsymbol{m}(\iota)=1$, in the general setup of this paper.} The setting of ishihara2021meta with any convex parameter space, whether bounded or unbounded, is a special case of my setting.
donoho1994 studies the optimal estimation of a linear functional of $\theta$ in a more general version of this paper's model, which allows for infinite-dimensional Gaussian models and noncentrosymmetric parameter spaces. donoho1994 derives minimax estimators and confidence intervals within the class of affine procedures, using squared error, absolute error, or the length of a fixed-length two-sided confidence interval as a loss function. Using the framework of donoho1994, armstrong2018optimal provide a one-sided confidence interval that minimizes the maximum $\beta$th quantile of the excess length among all one-sided confidence intervals of a given confidence level. In contrast to these studies, this paper focuses on a binary decision problem under welfare regret loss. Section (ref) discusses the connection between minimax estimation and minimax regret treatment choice in detail.
I illustrate my framework using two examples. The first is a special case of ishihara2021meta's ishihara2021meta setup, which I use as a running example to illustrate theoretical results in Section (ref). The second is policy choice using observational data under unconfoundedness.
In this section, I solve the minimax regret problem by using the hardest one-dimensional subfamily argument, which donoho1994 used to solve minimax affine estimation problems. This approach consists of three steps. The first step is to solve one-dimensional subproblems, in which the parameter space is restricted to a one-dimensional linear bounded subfamily. The second step is to search for the hardest one-dimensional subproblem, defined as the one with the highest minimax risk. The final step is to show that a minimax rule for the hardest one-dimensional subproblem is also minimax optimal for the original problem.
A key distinction between minimax regret treatment choice and minimax affine estimation lies in the structure of their risk functions. For estimation, standard risk functions such as mean squared error (MSE) can be decomposed into bias and variance. In contrast, the regret can be decomposed into the error probability and the potential welfare loss. Consequently, substantially different arguments are required for each of the above three steps.
I normalize $\boldsymbol\Sigma=\sigma ^2\boldsymbol I_n$ for some $\sigma>0$ throughout this section, where $\boldsymbol I_n$ is the identity matrix. This normalization is without loss of generality, since $\boldsymbol \Sigma$ is known.\footnote{Specifically, let $\tilde{\boldsymbol Y}=\boldsymbol{\Sigma}^{-1/2}\boldsymbol{Y}$ and $\tilde{\boldsymbol{m}}(\theta)=\boldsymbol{\Sigma}^{-1/2}\boldsymbol{m}(\theta)$ so that $\tilde{\boldsymbol Y}\sim {\cal N}(\tilde{\boldsymbol{m}}(\theta), \boldsymbol I_n)$. For any rule $\delta(\boldsymbol{Y})$, its regret in the problem with $(L,\boldsymbol{m},\Theta, \boldsymbol{\Sigma})$ is the same as the regret of the rule $\tilde \delta(\tilde{\boldsymbol{Y}})$ in the problem with $(L,\tilde{\boldsymbol{m}},\Theta, \boldsymbol{I}_n)$, where $\tilde \delta(\tilde{\boldsymbol{Y}})=\delta(\boldsymbol{\Sigma}^{1/2}\tilde{\boldsymbol{Y}})=\delta(\boldsymbol{Y})$.} I use the following assumption to derive a minimax regret rule.
The first two conditions are introduced in Section (ref). It is straightforward to see that ${\cal M}=\{\boldsymbol{m}(\theta):\theta\in\Theta\}$ is a nonempty, convex, and centrosymmetric subset of $\mathbb{R}^n$ under these two conditions. These conditions also imply the following relationship between the lower and upper bounds on the welfare contrast: $\sup I(\boldsymbol{\mu})=-\inf I(-\boldsymbol{\mu})$ for all $\boldsymbol{\mu}\in {\cal M}$. By this symmetry, it is sufficient to focus on $\sup I(\boldsymbol{\mu})$ in the analysis below. The third condition excludes the trivial case in which $L(\theta)=0$ for any $\theta\in\Theta$ and hence the worst-case regret of any decision rule is zero. The fourth condition is also necessary to obtain nontrivial results: If $\sup I(\boldsymbol{0})=\infty$, the regret of any decision rule is unbounded on $\{\theta\in\Theta:\boldsymbol m(\theta)=\boldsymbol 0\}$, and hence the worst-case regret of any rule is infinity.
In the following, I first solve one-dimensional subproblems in Section (ref). Next, I characterize the hardest one-dimensional subproblem in Section (ref) and present a minimax regret rule for the original problem in Section (ref).
First, I consider one-dimensional subproblems. To define a one-dimensional subproblem, take any $\bar\theta\in \Theta$ such that $L(\bar\theta)\ge 0$. A {\it one-dimensional subfamily}, denoted by $[-\bar\theta,\bar\theta]$, is defined as the set of all convex combinations of $\bar\theta$ and $-\bar\theta$: $$ [-\bar\theta,\bar\theta]\coloneqq\{\theta\in \mathbb{V}:\theta=\lambda\bar\theta,\lambda\in [-1,1]\}. $$ $[-\bar\theta,\bar\theta]$ is a subset of $\Theta$, since $\Theta$ is convex and centrosymmetric. Given $\bar\theta$, $[-\bar\theta,\bar\theta]$ can be viewed as a one-dimensional parameter space with a scalar parameter $\lambda\in[-1,1]$. A {\it one-dimensional subproblem} is the problem of finding a minimax regret rule for $[-\bar\theta,\bar\theta]$, whose maximum regret over $[-\bar\theta,\bar\theta]$ equals ${\cal R}([-\bar\theta,\bar\theta])$, where ${\cal R}([-\bar\theta,\bar\theta])=\inf_{\delta\in{\cal D}}\sup_{\theta\in [-\bar\theta,\bar\theta]}R(\delta,\theta)$.
The following result derives minimax regret rules for one-dimensional subproblems. Let $\|\cdot\|$ denote the Euclidean norm and $\Phi$ denote the cumulative distribution function of a standard normal random variable.
Lemma (ref) provides minimax regret rules for subproblem $[-\bar \theta,\bar \theta]$ separately for the following two cases: (i) $\boldsymbol{m}(\bar\theta)\neq \boldsymbol{0}$ and (ii) $\boldsymbol{m}(\bar\theta)= \boldsymbol{0}$. In case (i), a linear threshold rule based on $\boldsymbol{m}(\bar \theta)'\boldsymbol{Y}$ is minimax regret. On the other hand, in case (ii), any decision rule that chooses each policy with probability one-half over the distribution of $\boldsymbol Y\sim {\cal N}(\boldsymbol 0,\sigma^2 \boldsymbol I_n)$ is minimax regret. For example, a linear threshold rule $\delta(\boldsymbol{Y})=\mathbf{1}\left\{\boldsymbol{w}'\boldsymbol{Y}\ge 0\right\}$ for any $\boldsymbol{w}\neq \boldsymbol{0}$, a probit-like randomized rule $\delta(\boldsymbol{Y})=\Phi(\boldsymbol{w}'\boldsymbol{Y})$ for any $\boldsymbol{w}\neq \boldsymbol{0}$, and a data-independent randomized rule $\delta(\boldsymbol{Y})=1/2$ are all minimax regret. There exist infinitely many minimax regret rules for this case.
To gain intuition for Lemma (ref), I outline the derivation for case (i) and provide the proof for case (ii) when $L(\bar\theta)>0$. \paragraph{Case (i): $L(\bar\theta)>0$ and $\boldsymbol{m}(\bar\theta)\neq \boldsymbol{0}$.} Under $\theta=\lambda\bar\theta$, $\boldsymbol{Y}\sim {\cal N}(\lambda \boldsymbol{m}(\bar\theta),\sigma^2\boldsymbol{I}_n)$ and the welfare constrast is $\lambda L(\bar\theta)$ by the linearity of $\boldsymbol{m}$ and $L$. Viewing $\lambda\in [-1,1]$ as the underlying parameter of $[-\bar\theta,\bar\theta]=\{\theta\in \mathbb{V}:\theta=\lambda\bar\theta,\lambda\in [-1,1]\}$, one can show that the scalar statistic $T(\boldsymbol{Y})=\frac{\boldsymbol m(\bar\theta)'\boldsymbol Y}{\|\boldsymbol m(\bar\theta)\|^2}\sim {\cal N}\left(\lambda,\frac{\sigma^2}{\|\boldsymbol m(\bar\theta)\|^2}\right)$ is a sufficient statistic of $\boldsymbol Y$ for $\lambda$. Since the class of decision rules that only depend on a sufficient statistic is essentially complete\footnote{A class ${\cal C}$ of decision rules is {\it essentially complete} if, for any decision rule $\delta\notin{\cal C}$, there is a decision rule $\delta'\in{\cal C}$ such that $R(\delta,\theta)\ge R(\delta',\theta)$ for all $\theta\in\Theta$.} Berger1985book, it is justified to restrict one's attention to rules that depend on $\boldsymbol{Y}$ only through $T(\boldsymbol{Y})\in\mathbb{R}$.
With this restricted class of rules, the minimax regret problem for $[-\bar \theta,\bar \theta]$ is equivalent to a {\it univariate} problem in which we observe a univariate sample $T\sim {\cal N}\left(\lambda,\frac{\sigma^2}{\|\boldsymbol m(\bar\theta)\|^2}\right)$ and the welfare contrast is $\lambda L(\bar\theta)$ for $\lambda\in [-1,1]$. By a mild extension of the results of hirano2009asymptotics and tetenov2012asymmetric for univariate problems with unbounded parameter spaces to ones with bounded parameter spaces, the simple threshold rule $\delta(T)=\mathbf{1}\left\{T\ge 0\right\}$ is minimax regret for this univariate problem. Consequently, $\delta^*(\boldsymbol Y)=\mathbf{1}\{T(\boldsymbol{Y})\ge 0\}=\mathbf{1}\left\{\boldsymbol{m}(\bar \theta)'\boldsymbol{Y}\ge 0\right\}$ is minimax regret for the original one-dimensional subproblem $[-\bar \theta,\bar \theta]$.
For the minimax risk, a simple calculation shows that the regret of $\delta^*$ under $\theta=\lambda\bar\theta$ is
The first factor $|\lambda|L(\bar\theta)$ is the welfare loss when $\delta^*$ chooses the inferior policy under $\lambda\bar\theta$, which is increasing in $|\lambda|$. The second factor $\Phi\left(-|\lambda|\cdot \|\boldsymbol{m}(\bar \theta)\|/\sigma\right)$ is the probability of choosing the inferior policy, which decreases in $|\lambda|$. The regret $R(\delta^*,\lambda\bar\theta)$ is shown to be a bimodal function of $\lambda$ symmetric around zero, globally maximized at $\lambda\in \{-\tau^*\sigma/\|\boldsymbol{m}(\bar \theta)\|,\tau^*\sigma/\|\boldsymbol{m}(\bar \theta)\|\}$. Maximizing this function over $\lambda\in [-1,1]$ yields the minimax risk in Lemma (ref)(ref).
\paragraph{Case (ii): $L(\bar\theta)>0$ and $\boldsymbol{m}(\bar\theta)= \boldsymbol{0}$.} By the linearity of $\boldsymbol{m}$, $\boldsymbol{m}(\theta)=\boldsymbol{0}$ for any $\theta\in [-\bar\theta,\bar\theta]$. For a given rule $\delta$, the probability of choosing policy 1, $\mathbb{E}_{\theta}[\delta(\boldsymbol Y)]$, is constant over $\theta\in [-\bar\theta,\bar\theta]$, so that
where $\boldsymbol Y\sim {\cal N}(\boldsymbol 0,\sigma^2 \boldsymbol I_n)$. Consequently, the maximum regret can be calculated as follows: $$ \sup_{\theta\in[-\bar\theta,\bar\theta]}R(\delta,\theta)=
$$ Thus, any rule $\delta^*$ with $\mathbb{E}[\delta^*(\boldsymbol{Y})]=1/2$ is minimax regret, and the minimax risk is $L(\bar \theta)/2$.
Now, I search for the {\it hardest one-dimensional subfamily} $[-\bar{\theta}^*,\bar{\theta}^*]\subset\Theta$, which satisfies $ {\cal R}([-\bar{\theta}^*,\bar{\theta}^*])=\sup_{\bar\theta\in\Theta:L(\bar\theta)\ge 0}{\cal R}([-\bar \theta,\bar \theta])$. The key to characterizing the hardest one-dimensional subfamily is the {\it modulus of continuity}, defined as
The modulus of continuity and its variants have been used in constructing minimax estimators and confidence intervals on linear functionals in Gaussian models donoho1991,donoho1994,low1995tradeoff,armstrong2018optimal.\footnote{donoho1994 defines the modulus of continuity as $\tilde \omega(\epsilon)= \sup\{|L(\theta)-L(\tilde\theta)|: \|\boldsymbol{m}(\theta-\tilde\theta)\|\le \epsilon,\theta,\tilde\theta\in\Theta\}$. If $\Theta$ is convex and centrosymmetric, the relationship $\tilde\omega(\epsilon)=2\omega(\epsilon/2)$ holds.} Under Assumption (ref), $\omega(\epsilon)$ is the value of a convex optimization problem. If $\Theta$ is closed, this problem typically has a solution. \footnote{See donoho1994 for sufficient conditions for the existence of a solution.} In addition, $\omega(\cdot)$ is nonnegative and nondecreasing by construction. Furthermore, the modulus of continuity has the following properties.
The following result shows that the hardest one-dimensional subfamily can be obtained by solving an optimization problem that involves the modulus of continuity. Define the right derivative of $\omega(\cdot)$ at $0$ as $\omega'(0)\coloneqq\lim_{\epsilon\downarrow 0}\frac{\omega(\epsilon)-\omega(0)}{\epsilon}$. Let $\phi$ denote the probability density function of a standard normal random variable. For $\epsilon\ge 0$, I say that {\it $\theta_\epsilon\in \Theta$ attains the modulus of continuity at $\epsilon$} if $L(\theta_\epsilon)=\omega(\epsilon)$ and $\|\boldsymbol{m}(\theta_\epsilon)\|\le \epsilon$.
To provide intuition for this result, consider a convenient case in which $\omega(\epsilon)=\sup_{\theta\in\Theta:\|\boldsymbol{m}(\theta)\|= \epsilon} L(\theta)$ for each $\epsilon\ge 0$. In other words, suppose, for illustration, that the supremum remains the same if the inequality constraint $\|\boldsymbol{m}(\theta)\|\le \epsilon$ is replaced by the equality constraint $\|\boldsymbol{m}(\theta)\|= \epsilon$. In this case, the largest minimax risk among one-dimensional subproblems can be calculated as follows:
where the second equality uses Lemma (ref)(ref) and the third equality uses the assumption that $\omega(\epsilon)=\sup_{\theta\in\Theta:\|\boldsymbol{m}(\theta)\|= \epsilon} L(\theta)$. To simplify the last expression, note that $\frac{\omega(\epsilon)}{\epsilon}$ is continuous and nonincreasing on $(0,\infty)$ by the concavity of $\omega(\cdot)$, and hence $\sup_{\epsilon>\tau^*\sigma}\frac{\tau^*\sigma\omega(\epsilon)}{\epsilon}\Phi(-\tau^*)=\omega(\tau^*\sigma)\Phi(-\tau^*)\le \sup_{\epsilon\in[0,\tau^*\sigma]}\omega(\epsilon)\Phi(-\epsilon/\sigma)$. As a result, $$ \sup_{\bar\theta\in\Theta:L(\bar\theta)\ge 0}{\cal R}([-\bar \theta,\bar \theta])=\sup_{\epsilon\in[0,\tau^*\sigma]}\omega(\epsilon)\Phi(-\epsilon/\sigma). $$
In light of the above derivation, the problem of searching for the hardest one-dimensional subfamily can be interpreted as the following problem by an adversarial Nature. Nature optimizes $\epsilon\in[0,\tau^*\sigma]$, which represents a level of the strength of the signal provided by the sample $\boldsymbol{Y}$ within subfamily $[-\bar\theta,\bar\theta]$. The signal strength is measured by $\|\boldsymbol m(\bar\theta)\|$: The larger $\|\boldsymbol m(\bar\theta)\|$ is, the more information the sample $\boldsymbol Y$ provides about the sign of $L(\theta)$, and the smaller the error probability $\Phi(-\|\boldsymbol m(\bar\theta)\|/\sigma)$ is. The modulus of continuity $\omega(\epsilon)= \sup\{L(\bar \theta): \|\boldsymbol{m}(\bar \theta)\|= \epsilon,\bar \theta\in\Theta\}$ then represents the maximum potential welfare loss (i.e., $|L(\bar \theta)|$) among subfamilies subject to a given level of signal strength. Nature finally searches for the best level $\epsilon^*\in\arg\max_{\epsilon\in[0,\tau^*\sigma]}\omega(\epsilon)\Phi(-\epsilon/\sigma)$, which optimizes the balance between the error probability and maximum potential welfare loss. If $\epsilon^*>0$, the sample $\boldsymbol{Y}$ is informative within the hardest subfamily $[-\theta_{\epsilon^*},\theta_{\epsilon^*}]$. On the other hand, if $\epsilon^*=0$, the sample $\boldsymbol{Y}$ is uninformative within the hardest subfamily.
Lemma (ref)(ref) implies that $\epsilon^*>0$ if and only if $\omega'(0)/\omega(0)$ or $\sigma$ is sufficiently large. This is intuitive from the perspective of Nature's problem of optimizing the signal strength described above. The larger $\omega'(0)/\omega(0)(=\left.\frac{\partial}{\partial\epsilon}\log\omega(\epsilon)\right\vert_{\epsilon=0})$, the larger the percentage increase in maximum potential welfare loss associated with an increase in the signal level from $0$, and thus the greater Nature's incentive to choose a nonzero level of signal strength. Similarly, the larger $\sigma$ is, the noisier the sample $\boldsymbol{Y}$ becomes, leading to a smaller decrease in the error probability associated with an increase in the signal level, and consequently, a greater incentive for Nature to choose a nonzero signal strength.
I illustrate the results in Lemma (ref) using Example (ref). \setcounter{example}{0}
In the final step, I propose a minimax regret rule for the hardest one-dimensional subproblem $[-\theta_{\epsilon^*},\theta_{\epsilon^*}]$ and verify that it is indeed minimax regret for the full problem. For the case in which $\|\boldsymbol{m}(\theta_{\epsilon^*})\|=\epsilon^*>0$, I consider the rule $\delta(\boldsymbol{Y})=\mathbf{1}\left\{\boldsymbol{m}(\theta_{\epsilon^*})'\boldsymbol{Y}\ge 0\right\}$, which is minimax regret for $[-\theta_{\epsilon^*},\theta_{\epsilon^*}]$ by Lemma (ref). On the other hand, for the case in which $\|\boldsymbol{m}(\theta_{\epsilon^*})\|=\epsilon^*=0$, there exist infinitely many minimax regret rules for $[-\theta_{\epsilon^*},\theta_{\epsilon^*}]$. Among them, I consider a rule that can be viewed as a continuous extension of the rule $\delta(\boldsymbol{Y})=\mathbf{1}\left\{\boldsymbol{m}(\theta_{\epsilon^*})'\boldsymbol{Y}\ge 0\right\}$ from the case with $\epsilon^*>0$ to the case with $\epsilon^*=0$.
To construct the rule, first define $\bar I(\boldsymbol{\mu})$ as the upper bound on the welfare contrast when the reduced-form parameter is $\boldsymbol{\mu}$: $$ \bar I(\boldsymbol{\mu})\coloneqq \sup I(\boldsymbol{\mu})=\sup\{L(\theta):\boldsymbol m(\theta)=\boldsymbol \mu, \theta\in \Theta\},~~~\boldsymbol{\mu}\in \mathbb{R}^n, $$ where I use the convention that $\sup I(\boldsymbol{\mu})=-\infty$ when $I(\boldsymbol{\mu})$ is empty. Next, consider a sequence $\{\boldsymbol{\mu}_\epsilon\}_{\epsilon\in (0,\bar\epsilon)}$ indexed by positive real numbers $\epsilon\in (0,\bar\epsilon)$ with some $\bar\epsilon>0$ such that
Lastly, define
In Lemma (ref) in Appendix (ref), I show that the following holds under Assumption (ref). First, there exists a sequence $\{\boldsymbol{\mu}_\epsilon\}_{\epsilon\in (0,\bar\epsilon)}$ that satisfies (ref). Second, if $\omega'(0)>0$, the limit $\boldsymbol{w^*}$ exists and does not depend on the choice of $\{\boldsymbol{\mu}_\epsilon\}_{\epsilon\in (0,\bar\epsilon)}$ among potentially multiple sequences that satisfy (ref); that is, $\boldsymbol{w^*}$ is uniquely defined. Third, $\boldsymbol{w}^*$ is the direction in which the directional derivative of $\bar I(\cdot)$ at $\boldsymbol{0}$ is maximized among unit vectors. In other words, $\boldsymbol w^*$ is the “least favorable” direction, in the sense that the upper bound on the welfare contrast increases the most if the reduced-form parameter $\boldsymbol{m}(\theta)$ is changed from $\boldsymbol{0}$ in the direction $\boldsymbol w^*$.
The vector $\boldsymbol{w}^*$ relates to the modulus of continuity $\omega(\cdot)$ as follows: $\omega(\epsilon)=\sup_{\theta\in\Theta:\|\boldsymbol m(\theta)\|\le\epsilon}L(\theta)=\sup_{\boldsymbol{\mu}\in{\cal M}:\|\boldsymbol{\mu}\|\le\epsilon}\bar I(\boldsymbol{\mu})$; and if $\theta_\epsilon\in\Theta$ attains the modulus of continuity at $\epsilon$, then setting $\boldsymbol \mu_\epsilon=\boldsymbol{m}(\theta_\epsilon)$ satisfies (ref), and $\boldsymbol{w}^*=\lim_{\epsilon\downarrow 0}\frac{\boldsymbol{m}(\theta_\epsilon)}{\epsilon}$.
The following result derives a minimax regret rule for the original problem, which covers both the case in which $\epsilon^*>0$ (i.e., $\sigma\omega'(0)> 2\phi(0)\omega(0)$) and the case in which $\epsilon^*=0$ (i.e., $\sigma\omega'(0)\le 2\phi(0)\omega(0)$). The proof and a technical discussion are provided in Appendix (ref).
Theorem (ref) yields the following implications. First, the minimax regret rule in Theorem (ref) depends on the sample only through a weighted sum, $\boldsymbol{m}(\theta_{\epsilon^*})'\boldsymbol{Y}$ or $(\boldsymbol{w}^*)'\boldsymbol{Y}$, even though no such restrictions are imposed. The weights can be calculated by solving convex optimization problems; see Section (ref) for a computational procedure.
Second, the rule is nonrandomized or randomized, depending on whether the condition $2\phi(0)\frac{\omega(0)}{\omega'(0)}\le \sigma$ holds. This condition is related to the strength of the identifying restrictions. Under Assumption (ref), the closure of $I(\boldsymbol{0})$ is shown to be $[-\omega(0),\omega(0)]$.\footnote{Since $L$ and $\boldsymbol m$ are linear and $\Theta$ is centrosymmetric, $-\omega(0)=\inf\{L(\theta):\boldsymbol m(\theta)=\boldsymbol 0, \theta\in\Theta\}$. Moreover, for any $\alpha\in (-\omega(0),\omega(0))$, we can find $\theta\in\Theta$ such that $L(\theta)=\alpha$ and $\boldsymbol m(\theta)=\boldsymbol 0$ by the linearity of $L$ and $\boldsymbol m$ and the convexity of $\Theta$.} We can thus interpret $\omega(0)$ as half the length of the identified set of $L(\theta)$ when $\boldsymbol m(\theta)=\boldsymbol 0$.
If $L(\theta)$ is point identified, the length of the identified set is zero, so $2\phi(0)\frac{\omega(0)}{\omega'(0)}\le \sigma$. Therefore, the minimax regret rule is always nonrandomized under point identification. Even if $L(\theta)$ is not point identified, when the identified set is small relative to the noise level $\sigma$ (holding $\omega'(0)$ fixed), the condition $2\phi(0)\frac{\omega(0)}{\omega'(0)}\le\sigma$ holds and the rule is nonrandomized. On the other hand, when the identified set is large relative to $\sigma$, the minimax rule is randomized.
The randomized rule can equivalently be written as $\delta^*(\boldsymbol{Y})=\mathbb{P}\left((\boldsymbol{w}^*)'\boldsymbol{Y}+\xi\ge 0|\boldsymbol{Y}\right)$, where $\xi|\boldsymbol{Y}\sim {\cal N}(0,(2\phi(0)\omega(0)/\omega'(0))^2-\sigma^2)$. This rule can be implemented by first adding an independent noise $\xi$ to a scalar statistic $(\boldsymbol{w}^*)'\boldsymbol{Y}$ and then making a decision according to the sign of $(\boldsymbol{w}^*)'\boldsymbol{Y}+\xi$. This addition artificially increases the standard deviation of $(\boldsymbol{w}^*)'\boldsymbol{Y}$ from $\sigma$ to $2\phi(0)\omega(0)/\omega'(0)$, which is the threshold at which we switch from a nonrandomized rule to a randomized rule. The larger $\omega(0)$ is, the larger the variance of $\xi$ is and the more dependent the choice is on the noise. As a result, given any realization of $\boldsymbol Y$, the probabilities of choosing policy 1 and policy 0 approach $1/2$ as $\omega(0)$ increases, which suggests that the decisions become more mixed if we impose weaker restrictions on $\Theta$.
I apply Theorem (ref) to derive a minimax regret rule for Example (ref). \setcounter{example}{0}
Below, I discuss the role of randomization in Section (ref) and the relationship between Theorem (ref) and existing results in Section (ref). Section (ref) introduces the proof strategy. Finally, Section (ref) provides the procedure for computing the minimax regret rule.
Randomization plays the role of reducing the probability of choosing the inferior policy under parameter values at which the error probability of an original rule exceeds one-half. If the maximum regret of the original rule is attained at such parameter values, randomization can lead to a reduction in maximum regret.
To illustrate this, it is useful to compare the nonrandomized rule $\delta^*_{\rm NR}(\boldsymbol{Y})=\mathbf{1}\left\{(\boldsymbol{w}^*)'\boldsymbol{Y}\ge 0\right\}$ with the randomized rule $\delta^*$ in Theorem (ref). In Appendix (ref), I show that, under a mild condition on the local behavior of $\bar I(\boldsymbol{\mu})$ around $\boldsymbol{0}$, if $2\phi(0)\frac{\omega(0)}{\omega'(0)}>\sigma$, there exists $\theta^*\in\Theta$ such that $L(\theta^*)\ge 0$ (i.e., the optimal policy is policy 1),
Condition (ref) means that the error probability of $\delta^*_{\rm NR}$ exceeds one-half under $\theta^*$. Condition (ref) says that the regret of $\delta^*_{\rm NR}$ under $\theta^*$ exceeds the minimax risk over $\Theta$, and therefore $\delta^*_{\rm NR}$ is not minimax regret.
Now, consider the randomized rule $\delta^*$ in Theorem (ref). Under any $\theta^*$ that satisfies (ref) and (ref), randomization reduces the error probability and therefore regret of $\delta^*_{\rm NR}$. Specifically, a simple calculation shows that if $2\phi(0)\frac{\omega(0)}{\omega'(0)}>\sigma$, $$ 1-\mathbb{E}_{\theta^*}[\delta^*(\boldsymbol{Y})]=\Phi\left(-\frac{(\boldsymbol{w}^*)'\boldsymbol{m}(\theta^*)}{2\phi(0)\omega(0)/\omega'(0)}\right)<\Phi\left(-\frac{(\boldsymbol{w}^*)'\boldsymbol{m}(\theta^*)}{\sigma}\right)=1-\mathbb{E}_{\theta^*}[\delta^*_{\rm NR}(\boldsymbol{Y})], $$ and hence $R(\delta^*,\theta^*)<R(\delta^*_{\rm NR},\theta^*)$. Although randomization may increase regret under other values of $\theta\in\Theta$, it turns out that $\delta^*$ achieves a smaller maximum regret over $\Theta$ than $\delta^*_{\rm NR}$.
Theorem (ref) generalizes Stoye2012minimax's Stoye2012minimax result from univariate problems to multivariate problems. For the case with randomization, the use of a probit-like rule in Theorem (ref) is inspired by Stoye2012minimax's Stoye2012minimax for univariate problems. A novelty of my result is to use the hardest one-dimensional subfamily argument to construct a certain scalar statistic, $\boldsymbol{m}(\theta_{\epsilon^*})'\boldsymbol{Y}$ or $(\boldsymbol{w}^*)'\boldsymbol{Y}$, of the multivariate sample $\boldsymbol{Y}$ such that a threshold rule based on the statistic (plus a noise for the case with randomization) is minimax regret. This sharply contrasts with univariate problems, in which the only natural choice of the statistic is the univariate sample $Y$ itself.
In the setting described in Example (ref), ishihara2021meta propose a way of numerically computing a minimax regret rule within the class of decision rules of the form $\delta(\boldsymbol{Y})=\mathbf{1}\{\sum_{i=1}^nw_iY_i\ge 0\}$ with $\sum_{i=1}^nw_i=1$. An application of Theorem (ref) shows that this restricted class contains an unconstrained minimax regret rule when $2\phi(0)\frac{\omega(0)}{\omega'(0)}\le \sigma$ and may not when $2\phi(0)\frac{\omega(0)}{\omega'(0)}>\sigma$.
I briefly describe the approach to proving Theorem (ref). First note that lower and upper bounds on the minimax risk for the full problem are given by $$ \sup_{\theta\in [-\theta_{\epsilon^*},\theta_{\epsilon^*}]}R(\delta^*,\theta)={\cal R}([-\theta_{\epsilon^*},\theta_{\epsilon^*}])\le {\cal R}(\Theta)\le \sup_{\theta\in\Theta}R(\delta^*,\theta), $$ where the equality follows from the fact that $\delta^*$ is minimax regret for the subproblem $[-\theta_{\epsilon^*},\theta_{\epsilon^*}]$ by Lemma (ref), and the two inequalities hold by the definition of the minimax risk. To prove that $\delta^*$ is also minimax regret for the full problem, it suffices to show that the above lower and upper bounds coincide. More specifically, I show that the maximum regret of $\delta^*$ over $\Theta$ is attained at $-\theta_{\epsilon^*}$ and $\theta_{\epsilon^*}$; that is,
To show this, I first express the maximum regret of $\delta^*$ over $\Theta$ as $$ \sup_{\theta\in\Theta}R(\delta^*,\theta)=
$$ Here, the objective function on the right-hand side represents the maximum regret of $\delta^*$ over $\{\theta\in\Theta:\boldsymbol{m}(\theta)=\boldsymbol{\mu},L(\theta)\ge 0\}$. I then show that the supremum is attained at $\boldsymbol{\mu}=\boldsymbol{m}(\theta_{\epsilon^*})$, which proves (ref).
The arguments used to show (ref) are substantially different from the arguments used by donoho1994, who shows a counterpart of (ref) for minimax affine estimation problems. donoho1994's donoho1994 arguments rely on the following property of the maximum risk (such as the maximum MSE) of an affine estimator: The maximum risk and maximum squared bias are attained at the same parameter values. Since the maximum regret does not have this property, the arguments by donoho1994 cannot be applied to show (ref) for the minimax regret problem.
I conclude this section by summarizing the procedure for computing the minimax regret rule for a given problem $(L,\boldsymbol{m},\Theta,\sigma^2\boldsymbol{I}_n)$. For a general variance $\boldsymbol{\Sigma}$, the minimax regret rule can be obtained by applying the procedure after normalizing $\boldsymbol{Y}$ and $(L,\boldsymbol{m},\Theta,\boldsymbol{\Sigma})$ to $\boldsymbol{\Sigma}^{-1/2}\boldsymbol{Y}$ and $(L,\boldsymbol{\Sigma}^{-1/2}\boldsymbol{m},\Theta,\boldsymbol{I}_n)$, respectively. The procedure consists of the following steps:
In this section, I study the relationship between optimal treatment choice and optimal estimation. Given an estimator $\hat L$ of the welfare contrast $L(\theta)$, a decision rule can be constructed by plugging the estimator into the oracle optimal decision $\mathbf{1}\{L(\theta)\ge 0\}$: $\delta(\boldsymbol{Y})=\mathbf{1}\{\hat L(\boldsymbol{Y})\ge 0\}$. Such a rule is called a {\it plug-in} rule. For example, a plug-in rule can be constructed by using an estimator of $L(\theta)$ that is optimal under some standard criterion for estimation, such as minimax MSE optimality. The minimax regret rule in Theorem (ref) can also be viewed as a plug-in rule, which uses a (possibly randomized) estimator. A natural question is whether the estimator used in the minimax regret rule is optimal in a certain sense. Another question is how the estimator used in the minimax regret rule differs from optimal estimators under standard criteria. This section aims to answer these two questions.
As a preliminary step, Section (ref) presents a class of estimators that optimally trade off bias and variance in the estimation of $L(\theta)$. Using the results, Section (ref) discusses an interpretation of the minimax regret rule as a plug-in rule based on an estimator that satisfies a certain optimality. In Section (ref), I compare this estimator with a minimax affine MSE estimator, which is an existing optimal estimator in the setting of this paper.
Throughout Section (ref), I normalize $\boldsymbol\Sigma=\sigma ^2\boldsymbol I_n$ for some $\sigma>0$ as in Section (ref). I further assume that $\omega(\cdot)$ is differentiable on $(0,\infty)$ to simplify the presentation and proof of the results. The results can be modified to allow for nondifferentiability of $\omega(\cdot)$ by using the superdifferentials of $\omega(\cdot)$, which exist by the concavity of $\omega(\cdot)$. I also note that $\omega(\cdot)$ is differentiable in Example (ref) and for eligibility cutoff choice in Section (ref). See Lemma (ref) in Appendix (ref) for a sufficient condition for the differentiability.
As a basis for the discussion in Sections (ref) and (ref), I introduce a class of estimators that optimally trade off bias and variance. Let ${\cal C}$ denote the class of all (nonrandomized) estimators for $L(\theta)$ (i.e., measurable functions from $\mathbb{R}^n$ to $\mathbb{R}$). For estimator $\tilde L\in {\cal C}$, let ${\rm Bias}(\tilde L,\theta)\coloneqq\mathbb{E}_\theta[\tilde L(\boldsymbol{Y})]-L(\theta)$ and $\Var(\tilde L,\theta)\coloneqq\mathbb{E}_\theta[(\tilde L(\boldsymbol{Y})-\mathbb{E}_\theta[\tilde L(\boldsymbol{Y})])^2]$. For scalar $V\ge 0$, let ${\cal C}(V)$ denote the class of estimators with the maximum variance over $\Theta$ less than or equal to $V$: ${\cal C}(V)\coloneqq\{\tilde L\in {\cal C}:\sup_{\theta\in\Theta}\Var(\tilde L,\theta)\le V\}$. Consider the following minimax problem: $$ \inf_{\tilde L\in{\cal C}(V)}\sup_{\theta\in\Theta}{\rm Bias}(\tilde L,\theta)^2. $$ Solving this problem yields a class of estimators indexed by $V$ that optimally trade off the maximum squared bias and the maximum variance.
low1995tradeoff derives estimators that achieve minimax optimality in the above sense for infinite-dimensional Gaussian models. In Theorem (ref) in Appendix (ref), I extend the result of low1995tradeoff to the multivariate Gaussian models in this paper. The following is a corollary of Theorem (ref), which translates a class of optimal estimators indexed by $V$ in Theorem (ref) into a class of optimal estimators indexed by $\epsilon\ge 0$. For $\epsilon\ge 0$, define
assuming that $\omega(\cdot)$ is differentiable on $(0,\infty)$ and $\theta_{\epsilon}\in\Theta$ attains the modulus of continuity at $\epsilon$ with $\|\boldsymbol m(\theta_{\epsilon})\|=\epsilon$.
Theorem (ref) shows that the linear estimator $\hat L_{\epsilon}$ minimizes the maximum squared bias among all estimators (including nonlinear ones) with variance bounded by $V_\epsilon$. Furthermore, Theorem (ref) shows that the linear estimator $\hat L_{0}$ achieves the minimum maximum squared bias among all estimators. In contrast, low1995tradeoff does not provide a minimax squared bias estimator when variance constraints are absent in infinite-dimensional Gaussian models.\footnote{Specifically, Theorem 2 in low1995tradeoff does not provide a minimax squared bias estimator for the range of variance bound $V$ for which $0$ is the unique maximizer of $(\omega(\epsilon)-\epsilon\sqrt{V}/\sigma)$ over $\epsilon\ge 0$.}
Theorem (ref) provides an interpretation of the minimax regret rule $\delta^*$ in Theorem (ref). If $2\phi(0)\frac{\omega(0)}{\omega'(0)}<\sigma$, the minimax regret rule is given by $\delta^*(\boldsymbol Y)=\mathbf{1}\{\boldsymbol{m}(\theta_{\epsilon^*})'\boldsymbol Y\ge 0\}=\mathbf{1}\{\hat L_{\epsilon^*}(\boldsymbol Y)\ge 0\}$, where $\epsilon^*\in\arg\max_{\epsilon\in [0,\tau^*\sigma]}\omega(\epsilon)\Phi(-\epsilon/\sigma)$. This corresponds to a plug-in rule based on the linear estimator $\hat L_{\epsilon^*}$, which minimizes the maximum squared bias among all estimators with variance bounded by $V_{\epsilon^*}$. In contrast, if $2\phi(0)\frac{\omega(0)}{\omega'(0)}>\sigma$, the minimax regret rule is given by $\delta^*(\boldsymbol Y)=\Phi\left(\frac{(\boldsymbol{w}^*)'\boldsymbol{Y}}{((2\phi(0)\omega(0)/\omega'(0))^2-\sigma^2)^{1/2}}\right)=\Phi\left(\frac{\hat L_0(\boldsymbol{Y})}{((2\phi(0)\omega(0))^2-(\sigma\omega'(0))^2)^{1/2}}\right)$. Equivalently, this can be written as $\delta^*(\boldsymbol Y)=\mathbb{P}\left(\hat L_0(\boldsymbol{Y})+\xi\ge 0|\boldsymbol{Y}\right)$, where $\xi|\boldsymbol{Y}\sim {\cal N}(0,(2\phi(0)\omega(0))^2-(\sigma\omega'(0))^2)$. Thus, this rule can be interpreted as a plug-in rule based on the {\it randomized} estimator $\hat L_0(\boldsymbol{Y})+\xi$ for $L(\theta)$. This estimator is constructed by adding a mean-zero Gaussian noise to the linear estimator $\hat L_0$, which minimizes the maximum squared bias among all estimators.
The above interpretation highlights a key distinction between minimax regret treatment choice and minimax estimation. The fact that a plug-in rule based on a randomized estimator can be minimax regret suggests that increasing the variance of an estimator while holding the bias constant may reduce the maximum regret of the resulting plug-in rule. This means that the maximum regret cannot be expressed as an increasing function of the maximum squared bias and the variance.\footnote{This observation aligns with a result from ishihara2021meta, who show that in their setting, the maximum regret of a linear threshold rule $\delta(\boldsymbol{Y})=\mathbf{1}\{\boldsymbol{w}'\boldsymbol{Y}\ge 0\}$ with $\sum_{i=1}^nw_i=1$ can be written as a function of the maximum absolute bias and the variance of $\boldsymbol{w}'\boldsymbol{Y}$, though the function is not monotonic in variance.} This observation sharply contrasts with certain performance measures used in minimax affine estimation. For example, the maximum MSE of an affine estimator $\tilde L(\boldsymbol{Y})=c+\boldsymbol{w}'\boldsymbol{Y}$ is given by $\sup_{\theta\in \Theta}\mathbb{E}_\theta[(\tilde L(\boldsymbol{Y})-L(\theta))^2]=\sup_{\theta\in\Theta}{\rm Bias}(\tilde L,\theta)^2+\sigma^2\|\boldsymbol{w}\|^2$, which is increasing in both the maximum squared bias and the variance.
To shed further light on the distinction between minimax regret treatment choice and minimax estimation, I compare the linear estimator used by the minimax regret rule with a minimax affine MSE estimator. To introduce the latter, let ${\cal C}_{\rm affine}$ denote the class of all affine estimators of $L(\theta)$: ${\cal C}_{\rm affine}\coloneqq\{\tilde L\in {\cal C}:\tilde L(\boldsymbol{Y})=c+\boldsymbol{w}'\boldsymbol{Y}, c\in\mathbb{R},\boldsymbol{w}\in\mathbb{R}^n\}$. An estimator $\hat L\in {\cal C}_{\rm affine}$ is {\it minimax affine MSE} if $\sup_{\theta\in \Theta}\mathbb{E}_\theta[(\hat L(\boldsymbol{Y})-L(\theta))^2]=\inf_{\tilde L\in {\cal C}_{\rm affine}}\sup_{\theta\in \Theta}\mathbb{E}_\theta[(\tilde L(\boldsymbol{Y})-L(\theta))^2]$.
Given the results in Theorem (ref), one approach to finding a minimax affine MSE estimator is to minimize the maximum MSE of $\hat{L}_\epsilon$ over $\epsilon \geq 0$. However, for comparison with the minimax regret rule, the following theorem instead relies on an alternative characterization of the minimax affine MSE estimator based on donoho1994's donoho1994 approach.
Theorem (ref)(ref) follows from the results in donoho1994. The optimization problem $\max_{\epsilon\ge 0}\frac{\sigma^2}{\epsilon^2+\sigma^2}\omega(\epsilon)^2$ corresponds to calculating the largest minimax affine MSE over all one-dimensional subfamilies. The estimator $\hat L_{\epsilon_{\rm MSE}}$ is shown to be minimax affine MSE for the hardest one-dimensional subproblem $[-\theta_{\epsilon_{\rm MSE}},\theta_{\epsilon_{\rm MSE}}]$ and also for the full problem $\Theta$. Furthermore, the optimal level $\epsilon_{\rm MSE}$ is always positive, regardless of the degree of partial identification or the noise level $\sigma$. This implies that, in the hardest one-dimensional subproblem for the minimax affine MSE problem, the sample $\boldsymbol{Y}$ is informative about the sign of $L(\theta)$, unlike in the minimax regret problem, where $\boldsymbol{Y}$ can be completely uninformative.
Theorem (ref)(ref) compares the optimal level of $\epsilon$ for the minimax affine MSE and minimax regret problems. Together with Theorem (ref), this result implies that $\hat L_{\epsilon^*}$ has a maximum squared bias that is no larger and a variance that is no smaller than those of $\hat L_{\epsilon_{\rm MSE}}$. \footnote{ishihara2021meta derive a related result using a different argument in their setting.}
In many policy domains, the eligibility for treatment is determined based on an individual's observable characteristics. In this section, I demonstrate how my framework can be valuable for using data collected under the status quo eligibility criterion to decide whether to change it to a new one. This approach does not require conducting a randomized experiment that directly evaluates the performance of the status quo and new criteria.
Consider the following special case of Example (ref). For each unit $i=1,..,n$, we observe a fixed running variable $x_i\in\mathbb{R}$, a binary treatment status $d_i\in\{0,1\}$, and an outcome $Y_i\in\mathbb{R}$. The eligibility for treatment is determined based on whether the running variable exceeds a specific cutoff $c_0\in \mathbb{R}$, so that $d_i=\mathbf{1}\{x_i\ge c_0\}$. Suppose
where $f:\mathbb{R}\times\{0,1\}\rightarrow \mathbb{R}$ is an unknown function and plays the role of the parameter $\theta$ and $\sigma^2(x_i,d_i)>0$. We interpret $f(x,d)$ as the conditional mean potential outcome under treatment status $d\in\{0,1\}$ given $x$. \sloppy We can write the model in a vector form $\boldsymbol Y \sim{\cal N}(\boldsymbol m(f), \boldsymbol\Sigma)$, where $\boldsymbol Y=(Y_1,...,Y_n)'$, $\boldsymbol m(f)=(f(x_1,d_1),...,f(x_n,d_n))'$, and $\boldsymbol\Sigma={\rm diag}(\sigma^2(x_1,d_1),...,\sigma^2(x_n,d_n))$.
Now, suppose we are interested in changing the cutoff from $c_0$ to a specific value $c_1$. For illustration, assume $c_1<c_0$. The welfare under the cutoff $c_a$, $a\in\{0,1\}$, is given by $$ W_a(f) = \int [f(x,1)\mathbf{1}\{x\ge c_a\}+f(x,0)\mathbf{1}\{x< c_a\}]d\nu(x) $$ for some known measure $\nu$. An implicit assumption made here is that $f$ is invariant to the cutoff change. For an illustration of the results, I use an empirical measure as $\nu$, for which the welfare is the unweighted sample average: $ W_a(f) = \frac{1}{n}\sum_{i=1}^n[f(x_i,1)\mathbf{1}\{x_i\ge c_a\}+f(x_i,0)\mathbf{1}\{x_i< c_a\}]. $ I define the welfare contrast between the two cutoffs as $$ L(f)=\frac{n}{\tilde n}(W_1(f)-W_0(f))=\frac{1}{\tilde n}\sum_{i=1}^n\mathbf{1}\{c_1\le x_i< c_0\}[f(x_i,1)-f(x_i,0)], $$ where $\tilde n=\sum_{i=1}^n\mathbf{1}\{c_1\le x_i<c_0\}$ denotes the number of units between the two cutoffs $c_1$ and $c_0$, whose treatment status would be changed if the cutoff were changed. I scale $W_1(f)-W_0(f)$ by $n/\tilde n$ so that $L(f)$ represents the sample average treatment effect for these units. This scaling does not change the form of a minimax regret rule. Figure (ref) presents an example of this setting. Panel (a) shows an example of the conditional mean potential outcome function $f$. In this example, the welfare contrast is given by $L(f)=\frac{1}{2}\sum_{i\in\{2,3\}}[f(x_i,1)-f(x_i,0)]$, which is the sample average treatment effect for $x_2$ and $x_3$.
To point or partially identify $L(f)$, suppose that $f\in{\cal F}$ for some function class ${\cal F}$. Here, I focus on the {\it Lipchitz class} $$ {\cal F}_{\rm Lip}(C)=\{f:|f(x,d)-f(\tilde x,d)|\le C|x-\tilde x| \text{ for every } x, \tilde x\in\mathbb{R} \text{ and } d\in\{0,1\}\}. $$ The Lipschitz constraint bounds the maximum possible change in $f(x,d)$ in response to a shift in $x$ by one unit. I assume the Lipschitz constant is common for $f(\cdot,1)$ and $f(\cdot,0)$ for simplicity; it is possible to impose two separate constants. Other possible function classes include the class of functions with a known bound on the second derivative, as used by Imbens2019RDD for inference in RD designs.
To illustrate how the Lipschitz constraint allows one to partially identify $L(f)$, I present the upper bound on $L(f)$. The lower bound can be obtained analogously. Let ${\cal M}=\{\boldsymbol m(f):f\in {\cal F}_{\rm Lip}(C)\}$ and $x_{+,{\rm min}}=\min\{x_i:x_i\ge c_0\}$ be the value of $x$ of the treated unit closest to the original cutoff $c_0$. The upper bound on $L(f)$ when $\boldsymbol m(f)=\boldsymbol{\mu}\in {\cal M}$ is given by
where I define $\mu_{+,\min}=\mu_i$ for the unit $i$ with $x_i=x_{+,\min}$. The second equality holds, since $f(x_i,0)=\mu_i$ for any $i$ with $x_i<c_0$ (i.e., $d_i=0$) and $f(x_i,1)=\mu_i$ for any $i$ with $x_i\ge c_0$ (i.e., $d_i=1$). The last equality holds, since the upper bound on $f(x_i,1)$ for any unit $i$ with $x_i<c_0$ is shown to be $\mu_{+,\min}+C(x_{+,{\rm min}} - x_i)$ under the Lipschitz constraint. The upper bound $\bar I(\boldsymbol{\mu})$ increases with the Lipschitz constant $C$ and weakly increases with the size of cutoff change $|c_1-c_0|$ (holding $c_0$ fixed). In the example presented in Figure (ref), $x_4$ is the treated unit closest to the original cutoff $c_0$. The two dashed lines in Panel (b) indicate the upper and lower bounds on the function $f(x,1)$ on the range of $x<x_4$. In this example, the upper bound on $L(f)$ is given by $\bar I(\boldsymbol{\mu})=\frac{1}{2}\sum_{i\in\{2,3\}}[\mu_4+C(x_4 - x_i)-\mu_i]$.
Note that we are not interested in the identified set per se, but are interested in using the sample $\boldsymbol{Y}$ to choose between the two cutoffs given the specified function class ${\cal F}$. For a decision rule $\delta:\mathbb{R}^n\rightarrow[0,1]$, $\delta(\boldsymbol y)\in [0,1]$ represents the probability of changing the cutoff from $c_0$ to $c_1$ when the realized sample is $\boldsymbol y$. Alternatively, we can interpret $\delta(\boldsymbol y)$ as the fraction of individuals to whom we would assign treatment within the units between $c_1$ and $c_0$. In Section (ref), I derive a minimax regret rule when the welfare is the sample average outcome and ${\cal F}={\cal F}_{\rm Lip}(C)$. The form of the rule depends on the empirical distribution of $x_i$, the two cutoffs $c_0$ and $c_1$, the Lipschitz constant $C$, and the conditional variance $\sigma^2(x_i,d_i)$, all of which are treated as known. In practice, the policymaker must specify $C$ and $\sigma^2(x_i,d_i)$ to implement the rule. In Section (ref), I provide practical guidance on how to specify them.
To apply the results in Section (ref), I normalize $\tilde{\boldsymbol Y}=\boldsymbol\Sigma^{-1/2}\boldsymbol Y=(Y_1/\sigma(x_1,d_1),...,Y_n/\sigma(x_n,d_n))'$ and $\tilde{\boldsymbol m}(f)=\boldsymbol\Sigma^{-1/2}\boldsymbol m(f)=(f(x_1,d_1)/\sigma(x_1,d_1),...,f(x_n,d_n)/\sigma(x_n,d_n))'$, so that $\tilde{\boldsymbol Y} \sim{\cal N}(\tilde{\boldsymbol m}(f), \boldsymbol I_n)$. Then, $\omega(\epsilon)=\sup\{L(f): \|\tilde {\boldsymbol m}(f)\|\le \epsilon,f\in{\cal F}_{\rm Lip}(C)\}$ is the value of
The unknown parameter $f$ is infinite dimensional, but the objective and the norm constraint $\sum_{i=1}^n\frac{f(x_i,d_i)^2}{\sigma^2(x_i,d_i)}\le\epsilon^2$ depend on $f$ only through its values at $(x_1,0),...,(x_n,0),(x_1,1),...,(x_n,1)$. By a slight modification of Theorem 2.2 in Armstrong2021ATE, this optimization problem can be reduced to the following problem:
A solution to ((ref)) exists, since the objective function is continuous and the set of the vectors of $2n$ unknowns that satisfy the constraints is closed and bounded. Once we find a solution $(f(x_i,0),f(x_i,1))_{i=1,...,n}$, we can always find a function $f\in{\cal F}_{\rm Lip}(C)$ that interpolates the points $(x_i,f(x_i,0)),(x_i,f(x_i,1))$, $i=1,...,n$ BELIAKOV2006lipschitz, which is a solution to the original problem ((ref)). Problem ((ref)) is a finite-dimensional convex optimization problem with $2n$ unknowns, one quadratic and $2n(n-1)$ linear constraints, and a linear objective function, and can be solved using off-the-shelf convex optimization packages.\footnote{In the empirical application in Section (ref), I use CVXPY, a Python-embedded modeling language for convex optimization problems diamond2016cvxpy,agrawal2018rewriting.}
The following result derives a minimax regret rule.
The minimax regret rule is randomized or nonrandomized, depending on $s^*$ and $\bar\sigma$. $s^*$ is increasing in the Lipschitz constant $C$ and nondecreasing in the size of cutoff change $|c_1-c_0|$. $\bar\sigma$ is the standard deviation of $Y_{+,\min}-\frac{1}{\tilde n}\sum_{i:c_1\le x_i<c_0}Y_i$, and therefore increases with the variances of the treated unit closest to the status quo cutoff and the untreated units between the two cutoffs. If the Lipschitz constant $C$ or the cutoff change is large relative to the variances so that $s^*>\bar\sigma$, the minimax regret rule is a randomized rule based on $Y_{+,\min}-\frac{1}{\tilde n}\sum_{i:c_1\le x_i<c_0}Y_i$. In view of the results in Section (ref), $Y_{+,\min}-\frac{1}{\tilde n}\sum_{i:c_1\le x_i<c_0}Y_i$ can be interpreted as an estimator of $L(f)$ that minimizes the maximum squared bias over ${\cal F}_{\rm Lip}(C)$ among all estimators. As $C$ or the cutoff change increases, $(s^*)^2-\bar\sigma^2$ increases, and the decision is more randomized given the realization of the estimator.
On the other hand, if $s^*<\bar\sigma$, the minimax regret rule is a nonrandomized rule based on $\sum_{i=1}^nf_{\epsilon^*}(x_i,d_i)Y_i/\sigma^2(x_i,d_i)$. Using the results in Appendix (ref), we can show that if $s^*$ is marginally below $\bar\sigma$, $\sum_{i=1}^nf_{\epsilon^*}(x_i,d_i)Y_i/\sigma^2(x_i,d_i)$ is proportional to $Y_{+,\min}-\frac{1}{\tilde n}\sum_{i:c_1\le x_i<c_0}Y_i$. As the Lipschitz constant $C$ or cutoff change decreases or the variances increase so that $s^*$ becomes sufficiently smaller than $\bar\sigma$, the weighted sum assigns nonzero weights to some of the units with $x_i<c_1$ or $x_i>x_{+,\rm min}$. In Section (ref), I numerically examine the relationship between the weights and the choice of $C$ in the empirical application; see Figure (ref).
In practice, the conditional variance $\sigma^2(x_i,d_i)$ is unknown. A feasible version of the minimax regret rule is obtained by using a consistent estimator in place of the true $\sigma^2(x_i,d_i)$. The conditional variance can be estimated, for example, by applying a local linear regression to the squared residuals fan1998variance or by the nearest-neighbor variance estimator abadie2006matching. In the case in which unit $i$ represents a group of individuals and $Y_i$ is the sample mean outcome within group $i$, as in the empirical application in Section (ref), it is natural to use the conventional standard error of the sample mean as $\sigma(x_i,d_i)$.
Implementation of the minimax regret rule requires choosing the Lipschitz constant $C$. In principle, it is not possible to choose $C$ that applies to both sides of the cutoff $c_0$ in a data-driven way, since we only observe outcomes either under treatment or under no treatment on each side. It is, however, possible to estimate a lower bound on $C$. If $f\in{\cal F}_{\rm Lip}(C)$ is differentiable, a lower bound on $C$ is given by $\max\left\{\max_{\tilde x\ge c_0}\left\vert\frac{\partial f(\tilde x,1)}{\partial x}\right\vert,\max_{\tilde x< c_0}\left\vert\frac{\partial f(\tilde x,0)}{\partial x}\right\vert\right\}$, since $\left\vert\frac{\partial f(\tilde x,d)}{\partial x}\right\vert\le C$ for all $\tilde x$ and $d$. To estimate the lower bound, we could estimate the derivatives $\frac{\partial f(\tilde x,1)}{\partial x}$ for $\tilde x\ge c_0$ and $\frac{\partial f(\tilde x,0)}{\partial x}$ for $\tilde x< c_0$ by a local polynomial regression and then take the maximum of their absolute values. In practice, I suggest choosing the initial value of $C$ by estimating the lower bound or using application-specific knowledge, and considering a range of plausible values of $C$ to conduct a sensitivity analysis.
I now illustrate my approach in an empirical application to consider whether to scale up the BRIGHT program in Burkina Faso.
With the aim of improving children's---especially girls'---educational outcomes in rural villages, the BRIGHT program constructed well-resourced village-based schools with three classrooms for grades 1 to 3 in 132 villages from 47 departments\footnote{Departments are the third-level administrative divisions of Burkina Faso, below regions and provinces.} during the period 2005 to 2008. The Ministry of Education determined the villages in which schools would be built through the following process. First, 293 villages were nominated based on low school enrollment rates. Second, the Ministry administered a survey in each village and assigned each village a score using a set formula. The formula attached a large weight to the estimated number of children to be served from the nominated and neighboring villages, giving additional weight to girls. The Ministry then ranked villages within each department and selected the top half of the villages to receive a school. For further details on the BRIGHT program and allocation process, see levy2009bright and Kazianga2013bright.
Since the school allocation was determined at department level, the cutoff score for program eligibility differed across departments. Following Kazianga2013bright, I define the {\it relative score} as the score for each village minus the cutoff score for the department the village belongs to. As a result, a village is eligible for the program when the relative score is larger than zero. Kazianga2013bright use the relative score as a running variable and evaluate the causal effect of the program on educational outcomes using an RD design.
I use the replication data for Kazianga2013bright's Kazianga2013bright results Kazianga2019data and consider whether we should expand the program. The dataset contains survey results on 30 households from 287 nominated villages, for a total sample of 23,282 children between the ages of 5 and 12. The survey was conducted in 2008---namely, 2.5 years after the start of the program. Table (ref) in Appendix (ref) reports summary statistics on child educational outcomes and characteristics.
I consider school enrollment as the target outcome. Since the score and eligibility are determined at village level, I use the village-level mean outcome---namely, the enrollment rate for each village. This setting fits into the setup in Section (ref), where $i$ represents a village, $Y_i$ is the sample enrollment rate of village $i$, $d_i$ is program eligibility, and $x_i$ is the relative score. The original cutoff is $c_0=0$; that is, $d_i=\mathbf{1}\{x_i\ge 0\}$. The parameter is a function $f:\mathbb{R}\times\{0,1\}\rightarrow \mathbb{R}$, where $f(x,d)$ represents the counterfactual enrollment rate conditional on the relative score if the eligibility status were set to $d\in\{0,1\}$. Since $Y_i$ is a village-level sample mean, it is plausible to assume that $Y_i$ is approximately normally distributed. I use the conventional standard error of the sample mean as the standard deviation of $Y_i$.\footnote{The sample enrollment rate is zero in 21 out of 287 villages. I exclude these villages from the analysis, since the standard error of $Y_i$ is zero.}
Suppose we are evaluating the program to decide whether to scale it up. Specifically, consider the following decision problem. The counterfactual policy is to build BRIGHT schools in previously ineligible villages whose relative scores are in the top 20%, which corresponds to lowering the cutoff from $0$ to $-0.256$; in Section (ref), I examine the sensitivity of the result to the choice of the new cutoff. I use the average enrollment rate across villages as the welfare criterion, so that the welfare effect of this policy relative to the status quo is $$ L(f)=\frac{1}{\tilde n}\sum_{i=1}^n\mathbf{1}\{-0.256\le x_i<0\}[f(x_i,1)-f(x_i,0)], $$ where $\tilde n=\sum_{i=1}^n\mathbf{1}\{-0.256\le x_i<0\}$ is the number of villages that would receive a school under the new policy. Some assumptions underlying this choice of $L(f)$ are: (i) we consider the set of villages in the sample rather than a new set of villages, and (ii) the counterfactual enrollment rate function $f$ remains constant over time between the period when the BRIGHT program was implemented and the period when the program is expanded.\footnote{Another underlying assumption is no spillover effects. The plausibility of this assumption can be indirectly verified, for example, based on how isolated each village is from other villages.} In principle, these assumptions can potentially be relaxed by a suitable choice of $L(f)$ and the function class ${\cal F}$, while I focus on the above choice of $L(f)$ for a simple illustration of my approach.
When deciding whether to implement the policy, it is important to consider the benefit relative to the cost. Kazianga2013bright provide an estimate of the cost of constructing a BRIGHT school, which is \$4,758 per village.\footnote{I assume that the cost is known and constant across villages. If village-level cost data are available, my framework allows for unknown and heterogeneous costs by introducing the cost model on top of the outcome model.} To incorporate the cost in the decision problem, suppose that the policymaker cares about the cost-effectiveness of this new policy relative to similar programs. Cost-effectiveness is defined as the ratio of the policy cost to the increase in the enrollment. I assume that it is optimal to implement the policy if its cost-effectiveness is smaller than that of a benchmark policy, denoted by $CE_0$---that is, $$ \frac{\text{\$4,758}}{416\cdot\frac{1}{\tilde n}\sum_{i=1}^n\mathbf{1}\{-0.256\le x_i<0\}[f(x_i,1)-f(x_i,0)]}\le CE_0, $$ where $416$ is the number of children per village. The denominator represents the increase in the average enrollment across villages that would receive a BRIGHT school under the new policy. For concreteness, I set the benchmark cost-effectiveness to $\$83.77$, which is the cost-effectiveness of a school construction program in Indonesia duflo2001school,Kazianga2013bright.\footnote{The cost per village and the cost-effectiveness of a school construction program in Indonesia are found in Tables A18 and A20, respectively, in Online Appendix of Kazianga2013bright. I compute the number of children per village by dividing the total enrollment by the enrollment rate reported in Table A17 in Online Appendix of Kazianga2013bright.} The above condition with $CE_0=\$83.77$ is equivalent to $$ \frac{1}{\tilde n}\sum_{i=1}^n\mathbf{1}\{-0.256\le x_i<0\}[f(x_i,1)-0.137-f(x_i,0)]\ge 0. $$ My method can be used to consider this decision problem by setting the outcome to $Y_i-0.137d_i$, where $0.137$ can be viewed as the policy cost measured in the unit of the enrollment rate. Alternatively, $0.137$ can be viewed as the effect on the enrollment rate of implementing a benchmark policy with the same cost as the new policy. I present the results for this scenario as well as for a benchmark scenario in which we ignore the policy cost.
I implement my method assuming that the counterfactual outcome function $f$ belongs to the Lipschitz class ${\cal F}_{\rm Lip}(C)$.\footnote{It is possible to incorporate the natural bound of $[0,1]$ on enrollment rates by setting the outcome to $Y_i-0.5$ and assuming $f(x,d)\in [-0.5,0.5]$ for all $(x,d)$ in addition to the Lipschitz constraint. I find that this adjustment does not change the minimax regret rule and its maximum regret for the range of the Lipschitz constant $C$ considered in this analysis.} The Lipschitz constant $C$ represents the maximum possible change in the enrollment rate in response to a one-unit change in the relative score. While the relative score is computed based on multiple village-level characteristics, it is largely based on the estimated number of students to be served. All other characteristics being equal, a one-unit increase in the relative score corresponds to around 100 additional children in the village, where the average number of children per village is 416. To obtain a reasonable range of $C$, I estimate a lower bound on $C$ using the method described in Section (ref), which yields the lower bound estimate of $0.149$.\footnote{I estimate $\frac{\partial f(x,0)}{\partial x}$ at $x\in\{-2.5,-2.45,...,-0.05\}$ and $\frac{\partial f(x,1)}{\partial x}$ at $x\in\{0.05,0.1,...,2.5\}$ by local quadratic regression and take the maximum of their absolute values. For local quadratic regression, I use the MSE-optimal bandwidth selection procedure of calonic2018bian, which can be implemented by R package “nprobust.”} I present the results for $C\in\{0.05,0.1,...,0.95,1\}$ and examine their sensitivity to the choice of $C$.
Figure (ref) plots $\delta^*(\boldsymbol Y)$, the probability of choosing the new policy computed by the minimax regret rule, against the Lipschitz constant $C$. When $C< 0.6$, the minimax regret rule is nonrandomized. It chooses the new policy in the no-cost scenario and maintains the status quo in the scenario in which the policy cost is $0.137$. When $C\ge 0.6$, on the other hand, the minimax regret rule is randomized. The decisions become more mixed as $C$ increases. Given that the estimate of the lower bound on $C$ is 0.149, the minimax regret rule is nonrandomized when $C$ is less than four times the estimated lower bound. Under this reasonable range of $C$, the optimal decision is the same in each scenario.
If the minimax regret rule is nonrandomized, the rule is of the form $\delta^*(\boldsymbol Y)=\mathbf{1}\{\sum_{i=1}^nw_iY_i\ge 0\}$ for some weights $w_i$'s. Panels (a) and (b) of Figure (ref) plot the weight $w_i$ attached to each village against the relative score $x_i$ for $C=0.1$ and $C=0.5$, respectively. In the plots, the size of circles is proportional to the inverse of the standard error of the enrollment rate $Y_i$. For both $C=0.1$ and $C=0.5$, a few treated units just above the original cutoff (the solid vertical line) receive a positive weight, the untreated units between the original cutoff and the new cutoff (the dashed vertical line) receive a negative weight, and no other units receive any weight. When $C=0.1$, the weight tends to be larger for units with a smaller standard error. When $C=0.5$, a positive weight is attached only to the treated unit closest to the original cutoff. Also, the weights on the untreated units between the two cutoffs are almost identical. This situation corresponds to the minimax regret rule of the form $\delta^*(\boldsymbol Y)=\mathbf{1}\left\{Y_{+,\min}-\frac{1}{\tilde n}\sum_{i:c_1\le x_i<c_0}Y_i\ge 0\right\}$ discussed in Section (ref).
I compare the minimax regret rule with several plug-in decision rules of the form $\delta(\boldsymbol{Y})=\mathbf{1}\{\hat L(\boldsymbol{Y})\ge 0\}$, where $\hat L(\boldsymbol{Y})$ is an estimator of the policy effect $L(f)$. I consider the following three estimators of $L(f)$. (i) The minimax affine MSE estimator donoho1994, described in Section (ref), under the Lipschitz class ${\cal F}_{\rm Lip}(C)$. (ii) The minimax affine MSE estimator under the additional assumption of constant conditional treatment effects. In other words, I construct the estimator assuming that $ {\cal F}=\{f\in{\cal F}_{\rm Lip}(C): f(x,1)-f(x,0)=f(\tilde x,1)-f(\tilde x,0) ~\text{for all}~x,\tilde x\} $. This estimation corresponds to first nonparametrically estimating the average treatment effect at the original cutoff and then extrapolating the effects on the units between the two cutoffs by the constant effects assumption. (iii) The polynomial regression estimator Kazianga2013bright.\footnote{Kazianga2013bright estimate the treatment effect at the cutoff, not the effect on units away from the cutoff. They apply global polynomial regression RD estimators to child-level data. } Given the degree of polynomial $p$, I first estimate the model $f(x,d)=\alpha_0+\alpha_1 x+\cdots+\alpha_p x^p +\beta_0 d+\beta_1 d\cdot x +\cdots +\beta_pd\cdot x^p$ by weighted least squares regression using $1/\sigma^2(x_i,d_i)$ as the weight. I then estimate $L(f)$ by $\frac{1}{\tilde n}\sum_{i=1}^n\mathbf{1}\{-0.256\le x_i<0\}[\hat f(x_i,1)-\hat f(x_i,0)]$, where $\hat f$ is the estimated function. This estimator relies on the functional form of $f$ to extrapolate $f(x_i,1)$ for the untreated units.
Panel (a) of Figure (ref) reports the estimated policy effects from the minimax affine MSE estimators with and without constant conditional treatment effects. Overall, these two estimators exhibit a similar pattern. While the estimated policy effects are larger than the policy cost when $C$ is close to zero, they are smaller than the policy cost when $C$ is moderate or large. For $C\ge 0.15$, the resulting decisions about whether to choose the new policy are the same as the decision made by the minimax regret rule until $C$ reaches 0.6, where the minimax regret rule starts to randomize. In contrast, the estimated policy effects from the polynomial regression estimators of degrees $1$ to $5$ exceed the policy cost, as reported in Panel (b) of Figure (ref). The estimates appear to be close to the simple mean outcome difference between eligible and ineligible villages that can be computed from Table (ref) in Appendix (ref). The resulting decisions differ from the decision made by the minimax regret rule.\footnote{The estimators presented here can be written as $\sum_{i=1}^nw_iY_i$ for some weights $w_i$'s. See Figure (ref) in Appendix (ref) for the plots of these weights. While the minimax affine MSE estimators attach weights to units just above the original cutoff and to units between the two cutoffs, polynomial regression estimators even attach weights to units further from the cutoffs.}
The above decisions are computed from a particular realization of the sample. To assess the ex ante performance of different decision rules, I compute the maximum regret of these rules when the true function class is ${\cal F}_{\rm Lip}(C)$.\footnote{I compute the maximum regret of the minimax regret rule using the formula in Theorem (ref). For the other rules, I adapt the approach of ishihara2021meta to numerically calculate the maximum regret in this setup.} Panel (a) of Figure (ref) reports the result for the minimax regret rule and the plug-in rules based on the minimax affine MSE estimators with and without constant conditional treatment effects.\footnote{ The result for the plug-in rules based on polynomial regression estimators is omitted, since these rules turn out to have significantly larger maximum regret than the other rules.} The maximum regret of the plug-in MSE rule with constant conditional treatment effects is much larger than that of the other two, especially when the Lipschitz constant $C$ is large. The plug-in MSE rule without constant conditional treatment effects performs slightly worse than the minimax regret rule. The ratio of the maximum regret between the two rules is maximized at $C=0.6$, where the minimax regret rule starts to randomize. The maximized ratio is about 1.233.
I conduct several sensitivity analyses to assess how the results depend on the problem specification. First, I examine the sensitivity of the decision from the minimax regret rule to the choice of the new eligibility cutoff. Figure (ref) in Appendix (ref) reports the results when the new policy builds schools in the top 10% or 30% of previously ineligible villages instead of the top 20%. As predicted by the result in Section (ref), the minimax regret rule switches from a nonrandomized rule to a randomized rule at a smaller Lipschitz constant $C$ when the fraction of the target villages is larger. When the fraction is 30%, the rule is nonrandomized and suggests that the new policy is not cost-effective as long as $C$ is less than $0.4$.
Second, I examine the sensitivity of the decision to the policy cost. I find that the minimax regret rule nonrandomly decides to maintain the status quo as long as the cost exceeds $0.10$, for the range of $C$ between its estimated lower bound of 0.149 and 0.6.
So far, I have constructed decision rules assuming that the Lipschitz constant $C$ is known, which is a crucial assumption in my theoretical analysis. To assess the sensitivity of the performance to misspecification of $C$, I construct decision rules assuming $C=0.3$ and then compute their maximum regret when the true value of $C$ lies in $\{0.05,0.1,...,0.95,1\}$. Panel (b) of Figure (ref) reports the result. The solid line indicates the “oracle” maximum regret, which can be achieved if we correctly specify $C$. The result shows that the plug-in MSE rule without constant conditional treatment effects performs slightly better than the minimax regret rule when the true $C$ is close to zero. On the other hand, the minimax regret rule outperforms the plug-in MSE rule with nonnegligible differences for any value of the true $C$ greater than 0.3. The result suggests that the minimax regret rule is more robust to misspecification of $C$ toward zero than the plug-in MSE rule.\footnote{The potential superiority of the minimax regret rule seems consistent with the theoretical results in the following way. As shown in Section (ref), when the true value of $C$ is large, the oracle minimax regret rule only uses the treated units just above the original cutoff and the untreated units between the original and new cutoffs (see Panel (b) of Figure (ref)). If the specified $C$ is smaller than the true value, the resulting minimax regret rule is closer to the oracle rule than the plug-in MSE rule, since the minimax regret rule places more importance on the bias than the minimax affine MSE estimator, as discussed in Section (ref). Therefore, it is expected that the minimax regret rule performs better than the plug-in MSE rule under misspecification of $C$ toward zero.}
This paper derives an optimal decision rule for a large class of policy decision problems. The framework introduced in this paper allows for infinite-dimensional parameters, various forms of parameter restrictions, and partial identification of social welfare. I illustrate my approach through an application to the problem of eligibility cutoff choice in an RD setup.
Another promising application of this framework lies in policy adoption decisions using a difference-in-differences design. Specifically, consider a group of units that have experienced a policy change and another group that has not. Suppose the policymaker needs to decide whether to implement the new policy for the latter group. The average policy effect on this group may not be point identified if either the parallel trends assumption is violated or the policy effect varies between the two groups. My framework can be applied to this problem by imposing a set of restrictions on the degree of the violation of parallel trends and the amount of heterogeneity in policy effects.
Future research may explore several theoretical directions. First, one of the crucial assumptions in my approach is knowledge of the parameter space, such as the smoothness parameters of a function class. It would be interesting to investigate the possibility of adaptation over a collection of parameter spaces---that is, achieving near-optimal worst-case regret simultaneously over multiple parameter spaces, as has been studied for estimation and inference problems. Second, my approach only covers a binary choice problem. It would be challenging but both theoretically and practically important to extend the analysis to a multiple or continuous policy space.
\singlespacing \onehalfspacing
\pagenumbering{arabic} \setcounter{footnote}{0}
Appendix (ref) contains the proof and a discussion of Theorem (ref). Appendix (ref) contains auxiliary lemmas and proofs of Lemmas (ref), (ref), and (ref) and Theorems (ref) and (ref). Appendix (ref) contains derivations and a computational procedure for Section (ref). Appendix (ref) presents asymptotic properties of a feasible decision rule in the case with unknown error distribution. Appendix (ref) contains additional results for the empirical application in Section (ref).