Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
87,948 characters · 18 sections · 100 citation commands
Treatment Choice with Nonlinear Regret
\sloppy
\newgeometry{verbose,tmargin=1in,bmargin=1in,lmargin=1.25in,rmargin=1.25in,footskip=1cm}
\restoregeometry
\onehalfspacing
Evidence-based policy making using randomized control trial data is becoming increasingly common in various fields of economics. How should we use data to inform an optimal policy decision in terms of social welfare? Building on the framework of statistical decision theory as laid out in Wald50, the literature on statistical treatment choice initiated by manski2004statistical analyzes how to use data to inform a welfare optimal policy. Following Savage51 and manski2004statistical, researchers often focus on the average of welfare regret across the sampled data (called expected regret) and obtain an optimal decision rule by minimizing a worst-case expected regret.
When it comes to the ranking of different statistical decision rules, once we eliminate those that are stochastically dominated, it becomes less obvious how we should compare decision rules that do not stochastically dominate each other. Focusing on the expected regret, as suggested by manski2004statistical, provides a natural starting point.\footnote{There is, however, no compelling argument why we should limit our attention to the mean of regret, as has been acknowledged by manski2014quantile and manski2021econometrics. } In general, regardless of whether we consider a Bayes or minimax criterion, optimal decision rules defined in terms of their expected regret are singleton rules, i.e., given a sample, optimal decision rules either treat everyone, or no-one in the population. As an artificial example, suppose that the outcome of interest is $+1$ or $-1$ (success or failure) and imagine that we observe 100 successes and 99 failures (the status quo is zero for everyone). The empirical success (ES) rule, which is asymptotically optimal in terms of the mean of regret, suggests that everyone in the entire population should be treated. If there is a swing of one outcome from $+1$ to $-1$ though, then the same ES rule now dictates that no-one should be treated. Such high sensitivity and aggressiveness of treatment decisions with respect to sampling uncertainty can incur a large welfare loss due to random sampling errors, especially when the sample size is small. Axiomatically, the mean regret paradigm fails to capture some important and empirically relevant factors in decision making, e.g., a decision maker may be worried about large welfare loss in particularly bad states of the world.
To address these concerns in the mean regret framework, this paper proposes a novel, simple and tractable approach to treatment choice with finite data by optimizing a nonlinear transformation of welfare regret. In what follows, we let $g(\cdot)$ be a nonlinear transformation of the regret. We assess the performance of each treatment rule via the expected value of the transformed regret loss that it delivers. In the spirit of Wald50, this average nonlinear regret over realizations of the sampling process becomes the risk function. We refer to this risk as a nonlinear regret risk. Due to the nonlinearity of $g(\cdot)$, information relating to other moments of the regret distribution is encoded in the risk function. For example, when $g(r)=r^2$, the associated risk function is the sum of the squared expected regret and the variance of regret, penalizing decision rules that lead to a high variance of regret. We refer to this nonlinear regret risk as mean square regret.\footnote{As shown in Remark (ref), the variance regret is equivalent to the variance of welfare. Thus, compared to the expected regret criterion, our mean square regret can also be viewed as penalizing rules with a high variance of welfare.} We stress that our nonlinear regret approach is not an ad-hoc modification of the existing mean regret paradigm. In fact, it finds counterparts in decision theory from hayashi2008regret and stoye2011axioms who build a rich axiomatic model for a class of regret-driven choices, including mean square regret and many other nonlinear regret criteria (corresponding to what hayashi2008regret calls regret aversion). These criteria can better accommodate decision maker's aversion to large welfare loss in bad states of the world. We discuss the connection of our approach and the results of hayashi2008regret more in detail in Section (ref).
This shift of criterion towards a nonlinear transformation of regret changes optimal rules drastically. We show that, for many nonlinear transformations and a large class of distributions, singleton rules are either incomplete or inadmissible, offering novel decision-theoretic justifications for implementing fractional treatment assignment rules. Our approach also stands out in terms of the tractability, simplicity and interpretability of the associated optimal rules. We are able to provide general formula on Bayes and minimax optimal rules based on nonlinear regret risks. For mean square regret, we derive closed-form solutions for both Bayes and minimax optimal decision rules, not only in Gaussian finite samples with known variance but also asymptotically. These optimal rules are fractional and surprisingly easy to calculate and interpret. For example, the minimax optimal treatment assignment fraction has the following logistic structure: \[ \frac{\exp\left(2\cdot1.23\cdot\text{t-statistic}\right)}{\exp\left(2\cdot1.23\cdot\text{t-statistic}\right)+1}, \] where the $t$-statistic is for the average treatment effect estimated from experimental data, and which coincides with the posterior probability-matching assignment under the least favorable prior.\footnote{The posterior probability-matching assignment, known as the Thompson sampling algorithm thompson1933likelihood, possesses a desirable exploration-exploitation property in bandit problems. Our results show that the posterior probability-matching assignment can be justified in terms of minimax mean square regret, even in the static treatment choice problem where the exploration motive does not exist.} For example, our asymptotically minimax optimal rule for the previous artificial example would only allocate 54% of the population to the treatment, dropping to 46% if one outcome switches. Due to their fractional nature, our rules are useful even beyond the treatment decision paradigm: researchers may conveniently interpret our rule as a summary statistic that quantifies the strength of evidence in support of the treatment versus control. See Section (ref) for further discussions on this matter.
Given a nonlinear regret risk and a prior for the underlying potential outcome distributions, we obtain the Bayes optimal rules. Consistent with our incompleteness and inadmissibility results, Bayes optimal rules are, in general, also fractional rules. For mean square regret, we show that the Bayes optimal rule is a tilted posterior-probability matching rule, where the probability of random assignment corresponds to the posterior probability tilted by a certain weighting term. In a special case where the prior for the average treatment effect is supported only on two symmetric points, the tilting term is nullified and the Bayes optimal rule boils down to the Thompson-sampling type posterior-probability matching rule. For the minimax optimal rule in a Gaussian experiment with known variance, we can show that a least favorable prior is supported on two symmetric points. Hence, the minimax optimal rule follows the posterior-probability matching assignment rule, and is a logistic transformation of the sample mean. This minimax mean square regret rule is easy to compute and tuning- or hyper-parameter free.
Imagine the outcome of interest now follows a normal distribution $N(1,1)$ with unit mean and unit variance, whereas the status quo is zero for everyone. In this scenario, the infeasible optimal rule is to treat everyone and the regret of any decision rule is supported on $[0,1]$. Suppose the planner observes one observation from the $N(1,1)$ distribution and needs to make a treatment choice. The ES rule is optimal in terms of expected regret, but could be far from ideal in terms of other features of the regret and welfare distribution. In fact, if the planner adopted ES rule, then there would be a mass of 16% probability that she ended up with the largest possible regret of one (and the smallest possible welfare of zero). In contrast with the mean regret criterion commonly used in the literature, our mean square regret criterion penalizes rules with large variance of the regret distribution (and equivalently, large variance of welfare). If, instead, the planner implemented our proposed minimax rule, she could avert such high chance of welfare loss: the probability of incurring a regret larger than 0.95 (and a welfare smaller than 0.05) is only 1.4%. Also see Figure (ref) for a comparison of the distributions of the regret and welfare for ES rule and our proposed mean square regret minimax optimal rule.
We demonstrate the usefulness of our approach in two applications. First, using our general theory, we derive a minimax optimal rule in a normal regression model with binary treatment, a specification frequently used by many practitioners. Second, in practice, the planner often has a preference for singleton rules, and calculates a sufficient sample size for their randomized experiment based on these singleton rules. We show that implementing these singleton rules can lead to a large efficiency loss in terms of mean square regret.
Following HiranoPorter2009, HiranoPorter2020, we extend our finite sample results to a large sample setting by engaging with the limit experiments framework introduced by le2012asymptotic. Even when potential outcome distributions are non-Gaussian but belong to a regular parametric class, we can obtain a Gaussian limit experiment with known variance. Therefore, we can apply our results from a finite sample Gaussian experiment to a limit experiment and find feasible and asymptotically optimal rules with some efficient estimator of the parameters. Interestingly, in the limit experiment, the Bayes optimal rule under the mean square regret remains different from the minimax optimal rule, although the resulting mean square regret is quantitatively similar between the two rules. This is in contrast with the linear regret risk, for which it is known that the Bayes optimal and minimax optimal rules in the limit experiment are the same empirical success rule.
Our justification for implementing a fractional treatment assignment rule differs from those given in the existing literature so far, which all use the conventional expected regret as the criterion. See, for example, manski20092009 for a detailed review of fractional rules with standard regret under ambiguity and other non-standard settings, including nonlinear welfare, interacting treatments, learning and other non-cooperative aspects. More specifically, when the linear welfare is partially identified, Manski2000,manski2005social,manski2007identification, Manski2007 shows that minimax regret optimal rules are fractional even with the true knowledge of the identified set.\footnote{There are two approaches to go without assuming the true knowledge of the identified set. One approach is to plug in an estimate of the identified set, treating it as if it is the true object manski2013public,cassidy2019tuberculosis, manski2021probabilistic. The other approach is to directly consider finite sample minimax regret optimal rules, which can be also fractional stoye2012minimax,yata2021,manski2022identification if model ambiguity is sufficiently large compared to statistical uncertainty.} manski2007admissible and manski20092009 justify fractional rules via a nonlinear welfare in a point-identified setting. kock2022functional,kock2023treatment show that optimal treatment rules may be fractional if the decision maker targets a functional of the outcome distribution that is not quasi-convex. Fractional rules also arise when agents response with strategic behavior munro2020learning. Our results justify fractional rules in a standard setting without model ambiguity, nonlinear welfare or other strategic aspects.
In a series of papers, Manski and Tetenov have explored optimal treatment rules in frameworks that go beyond the classical paradigm of the statistical decision theory laid out by Wald50. manski1988ordinal,manski2011actualist argues to maximize a functional of the welfare distribution that at least weakly respects stochastic dominance. manski2014quantile consider the performance of a statistical treatment rule measured in terms of quantiles of the welfare. Our approach is distinct from the approaches taken by the aforementioned papers. In particular, we select treatment rules based on the distributions of their regret. Motivated by the risk aversion of policy makers, manski2007admissible consider a concave and monotone transformation of welfare measured in terms of a binary outcome, and define regret in terms of the transformed welfare. That is, manski2007admissible take a concave transformation of the welfare and keep the associated regret linear, while our approach looks at a nonlinear transformation of regret, keeping the welfare linear. These two approaches share a similar motivation: they advocate that decision makers may wish to take other aspects of the regret distribution into consideration when ranking decision rules. However, the strength of our approach lies in its technical tractability and practical implementability. See Online Appendix (ref) for further discussions on these matters.
The literature on the treatment choice problem has become an area of active research since the pioneering works of Manski2000, manski2002treatment,manski2004statistical and Dehejia2005 introduced a decision theoretic framework to the problem. When the welfare is point-identified, minimax regret treatment choice rules for finite samples are derived, in different settings, by schlag2006eleven, stoye2009minimax, tetenov2012statistical, masten2023minimax, and chen2024note. manski2014quantile and guggenberger2024minimax study treatment choice problems when quantiles are of interest. HiranoPorter2009,HiranoPorter2020 introduce an asymptotic framework to analyze treatment rules with limit experiments. When the welfare is partially identified, Manski2000,manski2005social,manski2007identification, Manski2007, manski20092009 analyzes the treatment choice given the knowledge of the identified set. Treatment allocation analyses without the knowledge of the identified set but with finite sample data include stoye2012minimax, christensen2020, ishihara2021, yata2021, manski2022identification, ishihara2023bandwidth and montielolea2023decision. Chamberlain2011 investigates a Bayesian approach to treatment choice, and christensen2020 and Giacomini2021 discuss a robust Bayesian approach.
There is a growing literature on learning in the context of individualized treatment rules that map an individual's observable characteristics to a treatment. See manski2004statistical, BhattacharyaDupas2012, kitagawa2018should, KT21, MT17, AW17, adjaho2022externally, han2023optimal, and cui2023individualized, among others, for analyses in different settings. Our analysis does not incorporate individuals' observable covariates. Since the nonlinear regret risk aggregates the conditional nonlinear regret risk additively, it is straightforward to incorporate observable discrete covariates into our analysis, i.e., an optimal individualized fractional assignment rule that applies an optimal fractional assignment rule to each subpopulation of individuals sharing the same covariate value.
The rest of the paper is organised as follows. Section (ref) introduces our setup. Section (ref) studies the admissibility and completeness of decision rules with nonlinear regret risk. Section (ref) presents finite sample results on Bayes and minimax optimal decision rules. In Section (ref), we discuss the axiomatic foundation of our criteria and the interpretation of our rules as a measure of strength of evidence. Section (ref) extends our results to the limit experiment framework and derives asymptotically optimal decision rules. Section (ref) applies our theory to an example of treatment choice in a normal regression model and an example of sufficient sample size calculation in randomized control trials. Section (ref) concludes. Proofs and lemmas are reserved for the Appendix.
Consider the assignment of a binary treatment $D \in \left\{ 1,0\right\}$ to an infinitely large population of individuals whose treatment effects can be heterogeneous. Let $Y(1)$ be the potential outcome when $D=1$ (with treatment) and $Y(0)$ be the potential outcome when $D=0$ (no treatment). Denote by $P\in \mathcal{P}$ the joint distribution of $(Y(1),Y(0))$, where $\mathcal{P}$ is a set of distributions under consideration. Define $\mu_{1} := \mathbb{E}[Y(1)]$ and $\mu_{0} := \mathbb{E}[Y(0)]$ as the means of the potential outcomes $Y(1)$ and $Y(0)$ under the distribution $P$. We assume that the welfare of the planner is determined by the mean outcome in the population. Defining the population average treatment effect as $\tau:=\tau (P):=\mu_{1}-\mu_{0}$, the infeasible optimal treatment policy is as follows: allocate $D=1$ to each individual in the population if $\tau\geq0$ and allocate everyone $D=0$ otherwise.
We also assume the decision problem is nontrivial in the sense that the image of the mapping $\tau(P),P\in\mathcal{P}$ contains both positive and negative values. Since the sign of $\tau$ is unknown, the planner collects an experimental sample of the observed outcomes of $n$ units randomly drawn from the population $P$, and the experimental design is known to the planner. The experiment generates a random vector $Z_{n}:=\left\{ Y_{i},D_{i}\right\} _{i=1}^{n} \in \mathbf{Z}_{n}$, where $Y_{i}$ is the observed outcome of unit $i$, $D_{i}$ is the treatment status of unit $i$, and $\mathbf{Z}_{n}$ is the sampling space. Let $P^{n}$ be the sampling distribution of $Z_{n}$, which depends on $P$ as well as the known experimental design.\footnote{For example, for a randomized control trial with a known treatment probability $0<\pi<1$, the joint likelihood of $Z_n$ is written as $P^n(Z_n)=\prod_{i=1}^{n}(\pi f_1(Y_i(1)))^{D_i}((1-\pi)f_0(Y_i(0)))^{1-D_i}$, where $f_1$ and $f_0$ are the marginal densities of $Y(1)$ and $Y(0)$, respectively. Our setup also accommodates other experimental designs. The derivation of $P^n$ is analogous but may be more involved if the experimental design is complicated. Also, we use the uppercase letter $Z_n$ to denote a random vector and use lowercase letter $z_n$ to denote a realized value of $Z_n$.} After observing data $Z_{n}$, the planner chooses a statistical treatment rule $\hat{\delta}$ that maps $Z_{n} \in \mathbf{Z}_{n}$ to a real number between 0 and 1, i.e., \[ \hat{\delta}:\mathbf{Z}_{n}\to[0,1], \] where $\hat{\delta}(z_{n})$ is the proportion of the population receiving the treatment. Denote by $\mathcal{D}$ the set of all statistical decision rules under consideration.
Applying the statistical treatment rule $\hat{\delta}$ to the population yields a welfare of \[ W(\hat{\delta}):=W(\hat{\delta},P):=\mu_{1}\hat{\delta}+\mu_{0}(1-\hat{\delta}) \] to the planner. The infeasible optimal treatment policy that maximizes welfare is $\delta^{*}:=\mathbf{1}\{\tau\geq0\}$. Following Savage51 and manski2004statistical, we define the regret of $\hat{\delta}$ as its welfare compared to the welfare of $\delta^{*}$, i.e., \[ Reg(\hat{\delta}):=Reg(\hat{\delta},P):=\tau[1\{\tau\geq0\}-\hat{\delta}]. \] Since $Reg(\hat{\delta})$ is a random object that depends on realizations of the random vector $Z_{n}$, manski2004statistical follows Wald50 in measuring the performance of $\hat{\delta}$ using its risk, i.e., the expected regret across realizations of the sampling process: \[ R(\hat{\delta},P):=\mathbb{E}_{P^{n}}[Reg(\hat{\delta})]:=\int_{z_{n}\in \mathbf{Z}_{n}} Reg(\hat{\delta}(z_{n}))dP^{n}(z_{n}), \] where $\mathbb{E}_{P^{n}}$ denotes the expectation with respect to $P^{n}$.
The risk criterion $R(\hat{\delta},P)$ ranks treatment rules according to their mean regret. We, instead, consider a planner whose assessment of the performance of statistical treatment rules depends not only on the mean of regret but also on some other features of the regret distribution. To take other features of the regret distribution into consideration, we look at the nonlinear transformation of regret: \[ g(Reg(\hat{\delta})), \] where $g:\mathbb{R}^{+}\rightarrow\mathbb{R}^{+}$ is some nonlinear function. The planner's preference over statistical decision rules $\hat{\delta}$ is measured by the expected value of $g(Reg(\hat{\delta}))$ with respect to realizations of $Z_n$:
We refer to the criterion $R_g(\hat{\delta},P)$ as the nonlinear regret risk.\footnote{The conventional mean regret risk function is a special case of our approach by taking $g$ to be linear.} Due to the nonlinearity of $g(\cdot)$, $R_g(\hat{\delta},P)$ depends not only on the mean but also on other features of the regret distribution, including its higher-order moments. For instance, if we specify the quadratic function $g(r)=r^2$, the squared regret is \[ (Reg(\hat{\delta}))^2=\tau^{2}[1\{\tau\geq0\}-\hat{\delta}]^{2}. \] This squared regret constitutes the new loss function, and we can evaluate the performance of $\hat{\delta}$ via mean square regret: \[ R_{sq}(\hat{\delta},P):=\tau^{2}\mathbb{E}_{P^{n}}[1\{\tau\geq0\}-\hat{\delta}]^{2}. \]
Viewing the nonlinear regret risk $R_g(\hat{\delta},P)$ defined in ((ref)) as the risk criterion within Wald's framework of statistical decision theory, we introduce the following standard definition of admissibility of a statistical treatment rule and the essential completeness of a class of rules.
As in standard statistical decision theory, admissibility defined through nonlinear regret is a minimal requirement that a desirable statistical treatment rule should satisfy. The notion of an essentially complete class simplifies the task of finding a good decision rule, as there is no need to consider rules outside this class.
Assumption G puts mild restrictions on the shape of the nonlinear transformation. Together with Assumption G, our first theorem shows that, in terms of nonlinear regret risk, a wide class of singleton rules are not essentially complete, implying that fractional rules cannot be eliminated from consideration in general. Recall for each $P\in\mathcal{P}$, $P^n$ denotes the distribution of data $Z_n$ and depends on $P$ while keeping the experimental design intact.\footnote{For example, for a randomized control trial with treatment assignment probability $\pi$, we have that for $P_*\in\mathcal{P}$, $P^n_*(Z_n)=\prod_{i=1}^{n}(\pi f_{1}^*(Y_i(1)))^{D_i}((1-\pi)f^*_{0}(Y_i(0)))^{1-D_i}$.} For an event $A$, write $\mathbb{P}_{P^{n}}(A):=\int\mathbf{1}\{z_{n}\in A\}dP^{n}(z_{n})$ as the probability of event $A$ under the sampling distribution $P^n$.
The incompleteness result in Theorem $\ref{thm:incomplete}$ is general, without relying on parametric specifications for $P$ or functional form restrictions on the decision rules. The class of singleton rules considered in Theorem $\ref{thm:incomplete}$ is also large, comprising both “non-degenerate singleton rules” ($\mathcal{D}_S^1$) and “degenerate singleton rules” ($\mathcal{D}_S^2$ and $\mathcal{D}_S^3$) under some pair of distributions, one with positive average treatment effect ($P_1$), and the other negative ($P_2$).\footnote{The assumption that the pair share the same absolute value of treatment effect is not essential. The proof still goes through as long as we have $\tau(P_1) = c_1 > 0$ and $\tau(P_2) = - c_2 < 0$ for some $c_1,c_2>0$.} Intuitively, any singleton rule with a positive probability of making mistakes under such pair of distributions is contained in $\mathcal{D}_S^1$. For the same pair of distributions, $\mathcal{D}_S^2$ would be the class of rules that treats everyone in the whole population with probability one, e.g., $\hat{\delta}_S=1$, which treats everyone and ignores data. In contrast, $\mathcal{D}_S^3$ is the class of rules that treats no one in the population with probability one under $P_1$ and $P_2$, and contains the trivial rule $\hat{\delta}_S=0$.
The results of Theorem $\ref{thm:incomplete}$ contrast sharply with the known result on the essentially completeness of singleton rules in more standard formulations of the treatment choice problem, where the (negative) expected welfare corresponds to the risk criterion in Wald's framework of statistical decision theory. For hypothesis testing problems with monotone likelihood ratio distributions, karlin1956theory show that the class of singleton threshold rules is essentially complete. As exploited in HiranoPorter2009 and tetenov2012statistical, the essential completeness of singleton threshold rules carries over to the treatment choice problem, implying that decision makers can discard fractional rules when in search of a good rule. Applying Theorem (ref) to monotone likelihood ratio distributions, we can conclude that the same class of singleton threshold rules are no longer essentially complete when it comes to many nonlinear regret risk criteria. Therefore, fractional rules cannot be eliminated from the consideration. Theorem (ref) is established with the following important observation:
Lemma (ref) reveals that if we focus on a binary class of distributions $\{P_1,P_2\} \subset \mathcal{P}$ with opposite treatment effects, then the class of nondegenerate singleton rules $\mathcal{D}_S^1$ are dominated. Theorem (ref) utilizes this dominance result in the restricted class of distributions $\{P_1,P_2\}$ to show that $\mathcal{D}_S$ is not essentially complete in the unrestricted (possibly infinite) class of distributions $\mathcal{P}$.
In general, proving the inadmissibility of a rule (in $\mathcal{P}$) is often more demanding, as one needs to find a dominating fractional rule not only for $\{P_1,P_2\}$ but all $P\in\mathcal{P}$. The next theorem confirms that a large set of singleton threshold rules are indeed inadmissible, if we impose stronger shape restrictions on $g$ and more distributional assumptions for the data.
Theorem (ref) is again in sharp contrast with the existing literature. For the same one-parameter exponential family distributions, results from karlin1956theory imply that singleton rules of the form (ref) are in fact admissible for the decision criterion of mean regret. Our results offer an opposite conclusion, demonstrating that the ranking of statistical treatment rules crucially depends on which features of the regret distribution the decision maker wishes to exploit. The homogeneity assumption on the shape of $g$ is a particularly convenient one that enables us to explicitly pin down a dominating rule in the form of (ref) over $\hat{\delta}_{t}$, but is not required and can be considerably weakened. This class of homogeneous functions also coincides with what hayashi2008regret calls regret aversion.\footnote{Our proof is constructive, and requires boundedness of the parameter space to pin down explicitly a fractional rule that dominates (ref) for all values of $\tau$ in the parameter space.}
We measure the performance of a rule $\hat{\delta}\in \mathcal{D}$ by its nonlinear regret risk $R_{g}(\hat{\delta},P)$, which depends on the true unknown $P$. In this section we look at two optimality criteria and derive general results on optimal rules for these criteria. We illustrate the usefulness of our results using specific parametric models.
We now characterize the Bayes optimal rule for the Bayes nonlinear risk. It turns out that under mild restrictions on the nonlinear transformation $g$, the associated Bayes optimal rule is also fractional. To proceed, let $\pi(P|z_{n})$ be the posterior distribution of $P$ given a prior $\pi$ and $Z_{n}=z_{n}$.
We now provide a simple example for which we derive the finite sample Bayes optimal rule with respect to a flat prior. This example also sheds some light on the form of the Bayes optimal rule in large samples, which is discussed in Section (ref).
Proposition (ref) is a direct application of Theorem (ref). Since the prior is flat, the `posterior density' is proportional to the likelihood ((ref)). The form of the Bayes optimal rule then follows ((ref)). The Bayes optimal rule $\hat{\delta}_{\pi_{f}}$ is a product of two terms. The first term, $\Phi(\bar{Y}_{1})$, is the posterior probability that the treatment effect is positive given the uninformative prior, and corresponds to the posterior probability matching rule. The second term, $( 1+\bar{Y}_{1}\cdot \Psi(\bar{Y}_{1}))$, adjusts the first term upwards if $\Bar{Y}_{1}>0$, and adjusts it downwards if $\Bar{Y}_{1}<0$ (note that $\Psi(x) > 0$). Therefore, this Bayes optimal rule tilts the posterior probability matching rule and assigns treatment with a probability closer to zero or one. Also see Table (ref) and Figure (ref) for the magnitudes of the probability assignment of the Bayes optimal rule and posterior probability matching rule with respect to the uniform prior.
As an alternative to Bayes rule, this section studies minimax optimal rule for nonlinear regret risk.\footnote{As shown by Savage51,manski2004statistical, maximin welfare criterion can be ultra-pessimistic and often leads to an optimal rule that always treats no-one in the population. Such ultra-pessimism carries over to maximin criterion applied to a nonlinear transformation of welfare. Echoing previous findings on minimax expected regret stoye2009minimax,tetenov2012statistical, our results show that the ultra-pessimism also does not occur for the minimax criterion applied to many nonlinear regret risks.}
The following proposition characterizes the minimax optimal rule as a Bayes rule under a least favorable prior.
Proposition (ref) is a direct result of lehmann2006theory. Using Proposition (ref), we can attempt to find the minimax optimal rule by adopting a `guess-and-verify' approach: guess a least favorable prior and derive its associated Bayes optimal rule; verify that the resulting Bayes nonlinear regret risk equals the worst frequentist nonlinear regret risk of the Bayes optimal rule. In general, it can still be difficult to guess the least favorable distribution. However, in many parametric models, the support of the least favorable distribution is often discrete and finite, or the minimax optimal rule has a constant frequentist risk across its parameter space. See, for example, kempthorne1987numerical. This greatly simplifies the problem. We now demonstrate the minimax optimal rule for Example (ref).
hayashi2008regret axiomatizes a class of regret-driven choices, including mean square regret and many other nonlinear transformations of regret. We briefly discuss how the results of hayashi2008regret provide a microeconomic justification of our approach. Let $\mathcal{S}:= \mathbf{Z}_n \times \Theta$ be the product space of sample $z_n \in \mathbf{Z}_n$ and the space of parameters indexing the distribution of the sample $\theta \in \Theta$. Let $\mathcal{D}$ be the set of statistical treatment choice rules $\hat{\delta} : \mathbf{Z}_n \to [0,1]$ and $\mathcal{D}^{\ast}$ be the extended set of treatment choice rules $\delta : \mathcal{S} \to [0,1]$ which includes the infeasible oracle rule $\delta^{\ast} = 1\{ \tau \geq 0 \}$.\footnote{Therefore, the set of statistical decision rules $\mathcal{D}$ available to a decision maker is allowed to be not as rich as $\mathcal{D}^*$.} We can view $\mathcal{S}$ as the state space, $\delta$ as acts, and $\mathcal{D}$ and $\mathcal{D}^*$ as menus in decision theory. To facilitate translation of our notation to that used in hayashi2008regret, write $W(\delta,\theta):=W(\delta)$, viewed as the utility of an act. For nonlinear transformation $g(x) = x^{\alpha}$, a nonlinear regret criterion with prior $\pi$ for $\theta$ writes as
The Bayes optimal nonlinear regret rule solves the following minimization
We can view the minimax optimal nonlinear regret rule as a solution to the following minimization:
where $\Pi$ denotes the set of probability distributions on $\Theta$.
Let $\left|\mathcal{S}\right|$ be the cardinality of $\mathcal{S}$, assumed to be finite in hayashi2008regret. hayashi2008regret builds a general axiomatic model where the choice of a decision maker is represented by the following minimization:
where $\varPhi:\mathbb{R}_{+}^{\left|\mathcal{S}\right|}\rightarrow\mathbb{R}_{+}$ is a homothetic function. The function $\varPhi$ is an aggregator that collects the decision maker's regret in different states of the world. Note both ((ref)) and ((ref)) may be viewed as special cases of ((ref)) (subject to caveats discussed below). While hayashi2008regret also discussed ((ref)), the minimax nonlinear regret criterion ((ref)) has not been considered elsewhere in the literature to the best of our knowledge. The case of $\alpha>1$ is called regret aversion in hayashi2008regret.
Axiomatic results in decision theory, like ((ref)), focus on decision making without sample data. Our criteria ((ref)) and ((ref)) are tailored for decision making with sample data. We recognize two caveats when interpreting ((ref)) and ((ref)) as special cases of ((ref)). First, in ((ref)) and ((ref)), the menu used to calculate regret, $\mathcal{D}^*$, includes the infeasible oracle rule and is allowed to be different from the actual menu $\mathcal{D}$ available to the decision maker. Second, in ((ref)) and ((ref)) the utility function $W$ is usually state dependent while in ((ref)) the utility function is state independent. Fully reconciling the differences between decision theory and statistical decision theory is beyond the scope of this paper. manski2021econometrics wrote: “As in decisions without sample data, there is no clearly best way to choose among admissible statistical decision functions (SDFs). Statistical decision theory has mainly studied the same criteria as has decision theory without sample data.” In view of manski2021econometrics, our nonlinear regret criteria can find their counterparts in decision theory without sample data from hayashi2008regret.
The treatment probability of our suggested minimax optimal rule is always between zero and one. As such, our rule can be naturally viewed as a summary statistic that measures the strength of evidence in favor of treatment versus control. Therefore, the usefulness of our approach does not hinge on the decision theoretic framework. One does not have to literally make a treatment decision based on $\hat{\delta}^{*}$. Instead, empirical researchers may view $\hat{\delta}^{*}$ as a degree of confidence gathered from data about the performance of the treatment in terms of welfare. Given finite sample from a single phase experiment, a larger value of $\hat{\delta}^{*}$ means we are more in favor of the treatment, while a smaller value of $\hat{\delta}^{*}$ signals less evidence supporting implementing the treatment. In contrast, in the standard mean regret paradigm, optimal rules are singleton and not fractional. Hence, it is not possible for applied researchers to solicit a measure of evidence strength from an optimal decision rule. Consider a scenario where $\Bar{Y}_{1}$ is only slightly larger than zero. The empirical success rule would dictate everyone in the population to be treated, even though we may think that the evidence reflected from data in favor of the treatment is not entirely strong.
Viewing $\hat{\delta}^{*}$ as a measure of the strength of evidence in a binary treatment setup, applied researchers may report $\hat{\delta}^{*}$ as an alternative summary statistic to the widely used P value. Despite its popularity, P value is known to be unfit for a measure of support for its hypothesis schervish1996p. Consider the setup in Example (ref) again. The P value for one-sided hypotheses $\mathbb{H}_{0}:\tau=0, \text{ v.s. } \mathbb{H}_{1}: \tau>0$ is $1-\Phi(\Bar{Y}_{1})$. Since $\Phi(\Bar{Y}_{1})$ is in fact the posterior probability with respect to the flat prior, reporting the P value corresponds to reporting the posterior probability under a flat prior. However, reporting a posterior probability under a specific prior is not necessarily associated with any optimality criterion. Different from the P value, $\hat{\delta}^{*}$ is an optimal treatment fraction under our mean square regret criterion. At the same time, $\hat{\delta}^{*}$ is also a posterior probability under a least favorable prior. Note given $\Bar{Y}_{1}>0(<0)$, $\hat{\delta}^{*}$ is quantitatively larger (smaller) than the P value. Therefore, reporting the P value would be more conservative than reporting $\hat{\delta}^{*}$ in our mean square regret framework. In Section (ref), we discuss how to calculate $\hat{\delta}^{*}$ in a normal regression model with binary treatment.
In this section we derive asymptotically optimal rules via the limit experiment framework le2012asymptotic, following the approach taken by HiranoPorter2009. We first consider a local parametrization of the statistical model $P$ so that, in large samples, the treatment choice problem is equivalent to a simpler problem in a Gaussian limit experiment. Then, we examine and normalize our nonlinear regret in the limit, and find the corresponding optimal treatment rule. A feasible and asymptotically optimal treatment rule also follows if there exists an efficient estimator of the parameters in the original statistical model $P$. For a review, see HiranoPorter2020.
For simplicity, we focus on regular parametric models of $P\in\mathcal{P}$ with mean square regret $R_{sq}$. Semiparametric models and other nonlinear regret criteria can also be considered, albeit necessitating more technical analysis. Without loss of generality, consider a case where the distribution of $Y(0)$ is known and the mean of $Y(0)$ is zero. Suppose now the distribution of $Y(1)$, denoted by $P$, is parameterized by a finite dimensional parameter $\theta\in\Theta\subseteq\mathbb{R}^{k}$. Hence, the population average treatment effect is \[ \tau(\theta)=\int zdP_{\theta}(z). \]
Data $Z_{n}=\{Z_{i}\}_{i=1}^{n}$ is independently and identically drawn from $P_{\theta}$. In particular, $Z_{i}\sim P_{\theta}$, where $Z_{i}\in\mathbf{Z}$ and $\mathbf{Z}$ is the support of $Z_{i}$. We now imagine a sequence of experiments $\mathcal{E}_{n}:=\{P_{\theta}^{n},\theta\in\Theta\}$ in which the sample size $n$ grows. Let $\theta_{0}\in\Theta$ satisfy $\tau(\theta_{0})=0$. We consider a sequence of local alternative parameters of the form $\theta_{0}+\frac{h}{\sqrt{n}}$, $h\in\mathbb{R}^{k}$, the most challenging case in which to determine the optimal treatment rule, even in large samples.
Assumption DQM is a standard assumption in the limit experiment framework van2000asymptotic. The function $s$ can usually be interpreted as the derivative of the loglikelihood function so that $I_{0}$ is the Fisher information under $P_{\theta_{0}}$.
Compared to mean regret criterion, our mean square regret additionally depends on the second moment of decision rules. Thus, Assumption C assumes convergence of both first and second moments of decision rules, differing from HiranoPorter2009, who only look at convergence of the first moment of decision rules. Under Assumptions DQM and C, we first establish the following result that allows us to simplify the original treatment problem to a Gaussian experiment in large samples.
Proposition (ref) is a special case of van2000asymptotic applied to the mean square regret setup, following HiranoPorter2009. To use Proposition (ref), note for any treatment rule $\hat{\delta}_{n}$ in the experiments $\mathcal{E}_{n}$, the mean square regret is \[ \mathbb{E}_{P_{\theta_{0}+\frac{h}{\sqrt{n}}}^{n}}\left[\tau\left(\theta_{0}+\frac{h}{\sqrt{n}}\right)^{2}\left(1\left\{ \tau\left(\theta_{0}+\frac{h}{\sqrt{n}}\right)\geq0\right\} -\hat{\delta}_{n}\right)^{2}\right], \] which depends on $\hat{\delta}_{n}$ only through $\mathbb{E}_{P_{\theta_{0}+\frac{h}{\sqrt{n}}}^{n}}[\hat{\delta}_{n}]$ and $\mathbb{E}_{P_{\theta_{0}+\frac{h}{\sqrt{n}}}^{n}}[\hat{\delta}_{n}^{2}]$, to which we can apply Proposition (ref). Thus, in terms of the mean square regret, any converging sequence of treatment rules is matched by some treatment rule in a simpler Gaussian experiment with unknown mean $h$ and known variance $I_{0}^{-1}$.
Let $\dot{\tau}$ be the partial derivative of $\tau(\theta)$ at $\theta_{0}$. Since $\tau\left(\theta_{0}\right)=0$, it follows that $\sqrt{n}\tau\left(\theta_{0}+\frac{h}{\sqrt{n}}\right)\rightarrow \dot{\tau}^{\prime}h$ as $n\rightarrow\infty$. Thus, for any rule $\delta$,
and $n\left[Reg\left(\delta,\left(\theta_{0}+\frac{h}{\sqrt{n}}\right)\right)\right]^{2}\rightarrow\left(Reg_{\infty}(\delta,h)\right)^{2}$ as $n\rightarrow\infty$. Hence, normalizing by $n$, for any converging rule $\hat{\delta}_{n}$ in the sense of Proposition (ref), we define the corresponding limit mean square regret as
With ((ref)) as the mean square regret in the limit experiment, we can apply our finite sample results in Section (ref) and derive a feasible and asymptotically optimal treatment rule via an efficient estimator of the parameters.
We first present results in terms of minimax optimality. Denote $\overset{h}{\rightsquigarrow}$ as convergence in distribution under the sequence of probability measures $P_{\theta_{0}+\frac{h}{\sqrt{n}}}^{n}$. Define $\sigma_{\tau}:=\sqrt{\dot{\tau}^{\prime}I_{0}^{-1}\dot{\tau}}$ to be the standard deviation of $\dot{\tau}^{\prime}\varDelta$, where $\varDelta\sim N(h,I_{0}^{-1})$.
Theorem (ref) extends our finite sample results to a large sample setting. Given a regular parametric model, the maximum likelihood estimator (MLE) usually satisfies ((ref)). Thus, Theorem (ref) suggests a simple way to construct an asymptotically minimax optimal rule in terms of mean square regret: estimate the parameters of $P_{\theta}$ via MLE, calculate a $t$-statistic for the mean, and then carry out a simple logit transformation for the $t$-statistic. This rule is always fractional and very easy to implement for practitioners. We expect that our result can also be extended to regular semiparametric models.
Next, we derive a feasible rule that is locally asymptotically Bayes optimal. Let $\pi(\theta)$ be a positive and continuous prior density on $\Theta$ (slightly abusing notation). For a treatment rule $\hat{\delta}_{n}$ that satisfies Assumption C, the normalized Bayes mean square regret is
We define the Bayes mean square regret in the limit experiment when $n\rightarrow\infty$ as \[ r_{sq}^{\infty}(\hat{\delta}):=\pi(\theta_{0})\int R_{sq}^{\infty}(\hat{\delta},h)dh. \] That is, as the Bayes mean square regret with respect to an uninformative prior. Then we can apply Theorem (ref) to derive the Bayes optimal rule for the limit experiment. Given an MLE estimate of the parameters in $P_{\theta}$, Theorem (ref) further implies that a feasible and asymptotically optimal Bayes rule also follows with a simple transformation of the $t$-statistic for the mean.
In the limit, the Bayes optimal rule is a tilted posterior probability matching rule with respect to the uninformative prior. Compared to the posterior probability matching rule, the Bayes optimal rule assigns treatment with a probability closer to zero or one. Compared to the limit minimax optimal rule, the Bayes optimal rule also assigns treatment with a probability close to zero or one. This contrasts with the case of linear regret risk, where it is known that the Bayes optimal and minimax optimal rules are the same empirical success rule. See Figure (ref) and Table (ref) for various rules in a Gaussian limit experiment with unit variance. It can be seen that all three fractional rules approach one as $\bar{Y}_{1}$ gets large. For sufficiently large positive values of $\bar{Y}_{1}$ (e.g., 2.33), the Bayes and minimax optimal rules are to effectively treat everyone. Even with a modest value of $\bar{Y}_{1} = 0.84$, the Bayes optimal rule recommends a probability of treatment of 0.94, which is quite high when compared with the corresponding probability of $0.8$ recommended by the posterior probability matching rule. Figures (ref), (ref) and (ref) present the mean square regret, mean regret and standard deviation of regret of the optimal rules in the same Gaussian limit experiment with unit variance. We make several observations: firstly, although they admit different forms, our Bayes optimal and minimax optimal rules in the limit experiment exhibit a similar performance in terms of the mean square regret (Figure (ref)); secondly, the ES rule is minimax optimal in terms of mean regret (Figure (ref)), but its excessive variance (Figure (ref)) in those states where mean regret is high implies that it is not optimal in terms of mean square regret.
Consider the following normal regression model frequently used by applied researchers:
where $Y$ is the outcome variable, $D$ is the binary treatment and $X$ is a vector of covariates (including the intercept). Suppose the treatment effect is homogeneous. Then, the parameter $\tau\in\mathbb{R}$ is the population average treatment effect. Let $\theta=(\tau,\beta^{\prime})^{\prime}$ and $Z=(D,X^{\prime})^{\prime}$. ((ref)) implies that the conditional density of $Y$ given $Z$ follows the parametric form \[ f(y|z)=\frac{1}{\sqrt{2\pi\sigma^{2}}}\exp\left(-\frac{1}{2\sigma^{2}}(y-\theta^{\prime}z)^{2}\right). \]
For now, assume the variance term $\sigma^{2}$ is known to focus on the finite-sample analysis. Given a random sample $\left\{ (Y_{i},Z_{i}^{\prime})^{\prime}\right\} _{i=1}^{n}$, the MLE estimator for $\theta$ is
the usual OLS estimator. Let $\hat{\tau}:=\hat{\theta}_{1}$, the first entry of $\hat{\theta}$. It follows by standard algebra that \[ \hat{\tau}\mid Z_{1},\ldots,Z_{n}\sim N\left(\tau,\sigma^{2}\left[\left(\sum_{i=1}^{n}Z_{i}Z_{i}^{\prime}\right)^{-1}\right]_{11}\right), \] where $\left[M\right]_{ij}$ denotes the $(i,j)$th entry of matrix $M$. By Theorem (ref), the finite sample minimax optimal rule is \[ \hat{\delta}^{*}=\frac{\exp\left(\frac{2\tau^{*}\hat{\tau}}{\sqrt{\sigma^{2}\left[\left(\sum_{i=1}^{n}Z_{i}Z_{i}^{\prime}\right)^{-1}\right]_{11}}}\right)}{\exp\left(\frac{2\tau^{*}\hat{\tau}}{\sqrt{\sigma^{2}\left[\left(\sum_{i=1}^{n}Z_{i}Z_{i}^{\prime}\right)^{-1}\right]_{11}}}\right)+1}, \] where $\tau^{*}$ is defined in Theorem (ref). Even if $\sigma^{2}$ is unknown, the MLE estimator for $\left(\theta^{\prime},\sigma^{2}\right)^{\prime}$ is $\left(\hat{\theta}^{\prime},\hat{\sigma}^{2}\right)^{\prime}$, where $\hat{\theta}$ is defined in ((ref)), and \[ \hat{\sigma}^{2}=\frac{1}{n}\sum_{i=1}^{n}\left(Y_{i}-Z_{i}^{\prime}\hat{\theta}\right)^{2}. \] Applying Theorem (ref), we may find a feasible asymptotically minimax optimal treatment rule as \[ \hat{\delta}_{F}^{*}=\frac{\exp\left(\frac{2\tau^{*}\hat{\tau}}{\sqrt{\hat{\sigma}^{2}\left[\left(\sum_{i=1}^{n}Z_{i}Z_{i}^{\prime}\right)^{-1}\right]_{11}}}\right)}{\exp\left(\frac{2\tau^{*}\hat{\tau}}{\sqrt{\hat{\sigma}^{2}\left[\left(\sum_{i=1}^{n}Z_{i}Z_{i}^{\prime}\right)^{-1}\right]_{11}}}\right)+1}. \]
Practitioners may report $\hat{\delta}_{F}^{*}$ as an alternative to the P value associated with $\tau$. Also see Section (ref) for more discussions on this issue.
In practice, the planner often has a preference for singleton rules like the empirical success (ES) rule or the hypothesis testing (HT) rule, and calculates what is a sufficient sample size based on these singleton rules. In this section we discuss the implications for the efficiency loss in terms of mean square regret if singleton rules were implemented instead of our proposed minimax optimal rules. Compared to our minimax optimal rule, these singleton rules often require significantly more data and thus are much less efficient. A similar discussion can be had for the Bayes optimal rule, but we omit this for brevity.
Consider the Gaussian experiment in Example (ref), but suppose now $\bar{Y}_{1}\sim N(\tau,\frac{\sigma^{2}}{n})$ is the sample average calculated from experimental data with a sample size of $n$ and known variance $\sigma^{2}>0$. In this case the minimax optimal rule in terms of mean square regret is \[ \hat{\delta}^{*}(\bar{Y}_{1})=\frac{\exp(2\tau^{*}\frac{\sqrt{n}}{\sigma}\bar{Y}_{1})}{\exp(2\tau^{*}\frac{\sqrt{n}}{\sigma}\bar{Y}_{1})+1}, \] where $\tau^{*}$ solves ((ref)). Given each $\varepsilon>0$, we can select $n$ such that \[ \sqrt{\sup_{\tau\in[0,\infty)} R_{sq}(\hat{\delta}^{*},P_{\tau})}\leq\varepsilon, \] i.e., the square root of the worst case mean square regret does not exceed $\varepsilon$. The worst case mean square regret can be calculated as
where $R_{sq}^{*}(1)\approx0.1199$ is the worst case mean square regret of the minimax optimal rule in Example (ref). Thus, the worst case mean square regret shrinks to zero at a rate of $\frac{1}{n}$. In practice, we can choose $\varepsilon$ to be proportional to $\sigma$, e.g., $0.01\sigma$, so that the square root of the worst case mean square regret does not exceed 1% of the standard deviation.
manski2016sufficient choose a sufficient sample size for the ES rule via the $\varepsilon-$optimal approach: a policy $\hat{\delta}$ is $\varepsilon-$optimal if, for all states of the world, \[ W(\delta^{*})-\mathbb{E}_{P^{n}}[W(\hat{\delta})]\leq\varepsilon, \] where $\delta^{*}$ is the infeasible optimal treatment rule or, equivalently,
for all states of the world. Given our Gaussian experiment $\bar{Y}_{1}\sim N(\tau,\frac{\sigma^{2}}{n})$, the worst case mean regret of the ES rule $\widehat{\delta}_{ES}=\mathbf{1}\{\bar{Y}_{1}\geq0\}$ can be calculated exactly as \[ \sup_{\tau\in[0,\infty)}\tau\left(1-\Phi\left(\frac{\sqrt{n}\tau}{\sigma}\right)\right)=\frac{\sigma}{\sqrt{n}}\sup_{\tau\in[0,\infty)}\tau\left(1-\Phi\left(\tau\right)\right)=0.1700\frac{\sigma}{\sqrt{n}}. \] If the planner has a preference for the ES rule and decides to choose the sample size so that ((ref)) holds with some $\varepsilon>0$, then the sample size should be at least \[ n_{ES}=0.0289\frac{\sigma^{2}}{\varepsilon^{2}}. \] The worst case mean square regret of the ES rule, however, is
where $R_{sq}^{ES}(1)=\sup_{\tau\in[0,\infty)}\tau^{2}\mathbb{E}_{\bar{Y}_{1}\sim N(\tau,1)}\left[\left(1-\mathbf{1}\{\bar{Y}_{1}\geq0\}\right)^{2}\right]\approx0.1657$. Hence, at $n_{ES}$, the worst case mean square regret of $\widehat{\delta}_{ES}$ is $\frac{\sigma^{2}}{n_{ES}}0.1657=5.7355\frac{\varepsilon^{2}}{\sigma^{2}}$. If, instead, the planner uses our minimax optimal rule, she only needs a sample size of $n^{*}=0.0209\frac{\sigma^{2}}{\varepsilon^{2}}$ for the worst case mean square regret not to exceed $5.7355\frac{\varepsilon^{2}}{\sigma^{2}}$. Thus, to guarantee the same worst case mean square regret, the ES rule requires nearly 40% more observations than our minimax optimal rule.
Practitioners who prefer the HT rule often select sample size by balancing Type I and II errors. In the Gaussian experiment $\bar{Y}_{1}\sim N(\tau,\frac{\sigma^{2}}{n})$, if the planner uses a size $\alpha$ ($<0.5$) HT rule \[ \hat{\delta}_{HT}=\mathbf{1}\left\{ \frac{\sqrt{n}\bar{Y}_{1}}{\sigma}\geq z_{(1-\alpha)}\right\} , \] where $z_{(1-\alpha)}$ is the $(1-\alpha)$ quantile of a standard normal, then it is common for her to select sample size so that the power of the test is at least $\beta$ ($>0.5$), i.e., under the alternative $\tau>0$, the probability of rejection is \[ \text{Pr}\left\{ \frac{\bar{Y}_{1}-\tau}{\frac{\sigma}{\sqrt{n}}}>z_{(1-\alpha)}-\frac{\tau}{\frac{\sigma}{\sqrt{n}}}\right\} =\beta. \] Then the sample size should be at least \[ n_{HT}=\frac{\sigma^{2}}{\tau^{2}}(z_{(1-\alpha)}-z_{(1-\beta)})^{2}. \] At this $n_{HT}$, we can also calculate the worst case mean square regret of the HT rule, which is approximately $\frac{\tau^{2}}{(z_{(1-\alpha)}-z_{(1-\beta)})^{2}}1.4458.$ However, at this $n_{HT}$, the worst case mean square regret of our minimax rule is only $0.1199\frac{\sigma^{2}}{n_{HT}}=0.1199\frac{\tau^{2}}{(z_{(1-\alpha)}-z_{(1-\beta)})^{2}}$. That is to say, with the same sample size $n_{HT}$, our minimax optimal rule guarantees that the worst case mean square regret is only around 8.3% of the corresponding value for the HT rule. Equivalently, to guarantee the same worst case mean square regret, the HT rule requires around 11 times more observations than our minimax optimal rule.
Our paper proposes a novel approach to measure the performance of statistical decision rules by considering a nonlinear transformation of regret. Such a shift of criterion can incorporate other features of the regret distribution (e.g., second- or higher-order moments and tail probabilities) into the decision-making process, and yields optimal rules that are drastically different from the existing literature. For a large class of nonlinear transformations, optimal rules are fractional, allocating only a proportion of the population to the treatment. For the mean square regret criterion, we also derive Bayes optimal and minimax optimal rules both for finite Gaussian samples and in asymptotic limit experiments. These rules have a simple and insightful form, and can be calculated easily by practitioners. As an extension, kitagawa2023treatment apply our mean square regret criterion to study optimal treatment choice problems when the welfare is only partially identified, and find that the fractional nature and the fundamental form of the minimax optimal rules found in our paper is preserved.
Our approach suggests that decision makers may display regret aversion, a notion related to but different from ambiguity aversion klibanoff2005smooth,denti2022model. In particular, our nonlinear regret criteria can find their counterparts in decision theory from the work of hayashi2008regret and thus are justified in terms of its microeconomic foundation.
Since our rules are always fractional, they naturally provide a degree of confidence in the performance of the treatment versus control. In that sense, our rule is useful for practitioners even outside the treatment choice paradigm. Implementing our rules also has the additional benefit of getting more data from randomized experiments that can be helpful for the inference of treatment effect, which would not be possible if singleton rules were implemented.