Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
65,957 characters · 7 sections · 0 citation commands
Minimax regret treatment rules with finite samples when a quantile is the object of interest
$ \and $
$ \and $
$}
Consider a setup in which a decision maker (DM) is informed about the population by a finite sample drawn from the population and based on that sample has to decide whether or not to apply a certain treatment or whether to randomize treatment assignment. The DM could be a policymaker who applies a treatment to the entire population or a person who is applying the treatment to an individual (e.g., herself). As examples of the latter setup, think of an individual who books a hotel accommodation on an internet platform after observing a certain number of ratings or a medical doctor who picks a treatment for a patient after observing a certain number of outcomes on the treatments.
This paper is adding to a short but growing literature on finding treatment rules, i.e. measurable mappings from the sample to the unit interval, that have finite-sample optimality properties. In most of the literature, the focus is on the expected outcome under the chosen treatment rule.\footnote{See e.g. Manski (2004), Manski and Tetenov (2007), Stoye (2007, 2009, 2012), Tetenov (2012), Masten (2023), Montiel Olea, Qiu, and Stoye (2023), Yata (2023), Kitagawa, Lee, and Qiu (2024), Chen and Guggenberger (2025), and additional references in these papers. Hirano and Porter (2009), Kitagawa and Tetenov (2018), and Christensen, Moon, and Schorfheide (2023) and many other references therein also use the minimax regret criterion but consider an asymptotic, rather than a finite sample, framework. This literature is inspired by the classical work of Wald (1950). In a recent paper Manski and Tetenov (2023) study several potential features of the state-dependent distribution of loss (rather than just its expectation) that a decision rule generates across potential samples.} Given that typically there is no treatment rule that is uniformly best over all possible joint distributions for $(Y_{0},Y_{1}),$ where $Y_{0}$ and $Y_{1}$ denote random variables for outcomes without and with treatment, respectively, one has to resort to other criteria for optimality. One option is to consider a prior over the space of joint distributions and maximize expected outcome for this particular prior. Another option is to focus on admissible treatment rules but that criterion typically does not single out an individual treatment rule, see Manski and Tetenov (2023) and Montiel Olea, Qiu, and Stoye (2023). Alternatively, one might consider finding a treatment rule that maximizes minimal expected outcome where the minimum is taken over all joint distributions of $(Y_{0},Y_{1}).$ However, if there exists a distribution that assigns the minimal possible values in the shared domains of $Y_{0}$ and $Y_{1}$ with probability one then any treatment rule is going to be optimal according to this criterion and therefore also the \textquotedblleft max-min\textquotedblright\ approach is not pinning down a unique rule. For that reason, instead, the so-called \textquotedblleft minimax regret\textquotedblright\ criterion is often adopted that determines treatment rules that minimize the maximal regret where regret is defined as the difference between the largest expected outcome that could be achieved for any treatment rule and the expected outcome under the treatment rule under consideration, and the maximum is taken with respect to all possible distributions for $(Y_{0},Y_{1}).$ Stoye (2009) derives minimax regret rules in finite samples for the case of two treatments under various sampling schemes, namely matched pairs, random sampling, and testing an innovation, and furthermore provides near-uniqueness results.
In this project we are interested in a setup where rather than expected outcome, the DM is concerned about the $\alpha$-quantile of the outcome (for a given $\alpha\in\lbrack0,1]$). Wang et al. (2018) consider robust estimation of the quantile-optimal treatment regime and provide arguments as to why focusing on a quantile rather than the mean may be sensible in certain applications. In fact, in many applications the tail of the outcome distribution is at the center of interest. For example, when evaluating job training programs to improve earnings the focus is often for earnings in the lower tail and likewise in survival analysis (e.g., survival time of cancer patients) the lower tail is of key importance. Wang et al. (2018, Section 2) provide numerical evidence where the mean-optimal treatment is detrimental for patients in the lower tail. In Economics, one might be interested in median income (case $\alpha=1/2)$ or a minimal education achievement for school kids, \textquotedblleft no child left behind\textquotedblright\ (case $\alpha=0$). See Manski (1988) who studies the \textquotedblleft quantile utility\textquotedblright\ model whose predictions, unlike the expected utility model, are invariant under ordinal transformations of utility. Subsequently, Rostek (2010) axiomatizes quantile preference and, recently, Manski and Tetenov (2023) suggest considering various deviations from mean loss including quantiles. Also see Chambers (2009).\footnote{Related (but in a non-finite-sample setup) Qi, Cui, Liu, and Pang (2019) and Qi, Pang, and Liu (2023) consider optimal decision rules based on the conditional value at risk (CVaR) measure. Quantile preferences have been attracting growing interest in the literature, e.g. De Castro, Galvao, and Ota (2026) consider a model in which an economic agent maximizes the discounted value of a stream of future $\alpha$-quantile utilities.}
As the main contribution of this paper, we derive minimax regret treatment rules $\delta$ in finite samples when an $\alpha$-quantile of the outcome distribution is the focus of interest and outcomes $Y_{0}$ and $Y_{1}$ take values in the unit interval. Somewhat surprisingly, we show that under various sampling schemes all treatment rules are minimax regret, namely in the case i) when the sample consists of a fixed number of treated and a fixed number of untreated units (that is, unbalanced panels are allowed for) and in the case ii) under random assignment with arbitrary treatment assignment probability in the sample equal to $p\in(0,1).$ In both cases i) and ii) maximal regret equals 1 for any treatment rule.
On the other hand, in the case iii) \textquotedblleft testing an innovation\textquotedblright, that is, the case where only data on $Y_{1}$ is observed and the $\alpha$-quantile $q_{s,\alpha}(Y_{0})$ of $Y_{0}$ is known, where the subscript $s$ denotes the joint distribution of $(Y_{0},Y_{1})$, if $q_{s,\alpha}(Y_{0})>1/2$ then $\delta\equiv0$ is the unique minimax rule, if $q_{s,\alpha}(Y_{0})<1/2$ then $\delta\equiv1$ is a minimax rule, and finally, if $q_{s,\alpha}(Y_{0})=1/2$ then any treatment rule $\delta$ is minimax regret; in each case, for minimax treatment rules $\delta$ we obtain the formula $\min\{q_{s,\alpha}(Y_{0}),1-q_{s,\alpha}(Y_{0})\}$ for the maximal regret.
Obviously, like for the case where expected outcomes are the focus, also here the max-min criterion is not informative. Namely, as long as the joint distribution of $(Y_{0},Y_{1})$ can be chosen such that $q_{s,\alpha} (Y_{0})=q_{s,\alpha}(Y_{1})=0,$ that is, the worst possible outcome under the given restriction that $Y_{0},Y_{1}\in\lbrack0,1],$ any treatment rule would be max-min optimal. In contrast though, to the case where expected outcomes are the focus, for quantiles also minimax regret is not informative for sampling designs i) and ii) and also for iii) when $q_{s,\alpha}(Y_{0})=1/2.$
The typical strategy in the extant literature for determining minimax regret rules when the focus is on expected outcomes is via a Nash equilibrium approach in a fictitious zero sum game in which the DM plays against an antagonistic nature whose payoff equals regret. To establish that a particular treatment rule $\delta$ is minimax regret one attempts to guess a \textquotedblleft state of nature\textquotedblright\ $s$ , allowed to be mixed strategy (over a finite number of states), called a least favorable distribution, for which the pair $(\delta,s)$ constitutes a Nash equilibrium. Existence of such a pair $(\delta,s)$ implies that $\delta$ is indeed a minimax regret rule, see for example Berger (1985), Stoye (2009), and Chen and Guggenberger (2025). Often, in a first step, one restricts outcomes to be Bernoulli and finds a minimax rule in this simplified setup and then, in a second step, uses the so-called coarsening approach to tackle the general case, see e.g., Cucconi (1968), Gupta and Hande (1992), and Schlag (2003, 2006). When nature picks a state of the world trying to inflict high regret it faces the trade-off that on the one hand a high differential between expected outcomes with and without treatment is needed but on the other hand the more different the distributions of $Y_{0}$ and $Y_{1}$ are the easier the DM can tell them apart using the sample information.
The proof structure for the main results in this paper differs from the one just described. Namely for cases i)\ and ii) we show that for any treatment rule $\delta$ one can find a state of nature $s=s_{\delta}$ such that $s$ inflicts the highest possible regret, namely 1. That insight is sufficient to establish that all rules are minimax regret and that maximal regret equals 1 for all treatment rules. It is noteworthy and remarkable that, despite the trade-off just described, nature is powerful enough to inflict maximal regret on the DM. Even more surprisingly, the conclusion under i) and ii) continue to be true even if nature is restricted to only Bernoulli distributions. The proof for part iii) relies on the main insight that if $\delta\neq0$ then there exists a state of nature such that regret equals the known $\alpha $-quantile of $Y_{0}.$ For that to be true it is sufficient for nature to have discrete distributions, supported on $N+1$ points, at its disposal.
We start off with the case with no covariates and then allow for a discrete-valued covariate in the model. In the latter case, a treatment rule maps a sample onto treatment probabilities for each of the $K$ possible outcomes of the covariate. Not surprisingly, given the results without a covariate, we find that in cases i), ii), and case iii) with known quantile equal to 1/2, again, each treatment rule is minimax regret, while in case iii) when the known quantile is different from 1/2, no-data rules are minimax regret.
We include a small finite sample simulation study in the case of \textquotedblleft testing an innovation\textquotedblright\ where we simulate regret of various treatment rules, namely the empirical success rule and several no-data rules. The study corroborates our theoretical findings about minimax regret treatment rules and the formulas for maximal regret that we derive.
The remainder of the paper is organized as follows. Section 2 introduces the theoretical setup for our notion of a minimax regret rule with quantiles and contains analytical results when no covariates are included. In Subsection 2.1 we present various approaches for how quantiles could be incorporated into the Waldean framework of statistical decision theory and juxtapose them with our approach. Subsection 2.2 derives minimax regret treatment rules under various sampling schemes and also provides a brief discussion of minimax regret treatment rules when certain restrictions are imposed on the states of nature. Subsection 2.3 simulates the performance of various treatment rules in finite samples in the case of testing an innovation. Finally, Section 3 derives minimax regret treatment rules in the case where covariates are included in the model. All proofs are given in the Appendix.
A decision maker has to decide whether or not to assign treatment after being informed about the population by a finite sample.\footnote{Alternatively, rather than framing the options as "treatment" and "no treatment", one could frame the setup as a choice between two treatments.} The setup is very similar to the one in Stoye (2009) with one key modification. Namely, rather than focusing on mean outcomes, here we are concerned with the $\alpha$-quantile of the outcome distribution.
For most parts of the paper potential outcomes $Y_{0}$ and $Y_{1}$ for untreated/treated individuals are restricted to the unit interval
where $S$ is assumed known to the DM. At first, different members of the population are all identical to the decision maker. Later, we will consider the case where a covariate is included and treatment assignment can be made conditional on the realization of the covariate.
By $\mathbb{S}$ we denote the set of \textquotedblleft states of the world\textquotedblright\ $s$ where an $s$ denotes a possible joint probability distribution over the potential outcomes for $Y_{0}$ and $Y_{1}.$ Unless otherwise stated $\mathbb{S}$ is unrestricted and contains all possible joint distributions for $Y_{0}$ and $Y_{1}$ on $S^{2}.$ Upon observing a sample $w_{N}=(t,y)$ of size $N$ of treatment statuses $t=(t_{1},...,t_{N})$ and outcomes $y=(y_{1},...,y_{N})$ where the $i$-th component of $y,$ $y_{i},$ is an independent realization of $Y_{t_{i}}$ and $t_{i}$ denotes the treatment received by individual $i,$ the task for the decision maker is to choose
which denotes the probability with which treatment is assigned.\footnote{We do not index $t$ and $y$ by $N$ because it would make the notation too cumbersome. } Namely, which treatment the DM assigns is determined as an independent draw of the Bernoulli random variable $B=B(\delta(w_{N}))\in\{0,1\}$ that equals $1$ with probability $\delta(w_{N})$ and is assumed to be independent of all other random objects. Therefore, the setup here allows for randomized treatment rules.
Let $\alpha\in\lbrack0,1].$ Given a state of the world $s$ and a statistical treatment rule $\delta$ the objective function for the DM is
where $Y_{B(\delta(w_{N}))}$ denotes random outcomes generated when the treatment rule $\delta$ is used, and $q_{s,\alpha}(Y_{B(\delta(w_{N}))})$ denotes the $\alpha$-quantile of $Y_{B(\delta(w_{N}))}$ when the state of the world is $s\in\mathbb{S}.$ In particular, when the treatment rule $\delta(w_{N})$ equals 0 (or 1) then with probability 1 $Y_{B(\delta(w_{N}))}$ equals $Y_{0}$ (or $Y_{1}$).
By definition, an $\alpha$-quantile of a scalar valued random variable $X$ with domain $D$ is any number $q\in D$ that satisfies
Clearly then, any $\alpha$-quantile of $Y_{0},$ $Y_{1}$ and $Y_{B(\delta (w_{N}))}$ is an element of [0,1]. In general, this definition does not lead to a unique $\alpha$-quantile. The definition allows for the case where a quantile has a non-zero point mass. To be explicit in cases where there is non-uniqueness, we make the following definition that hinges on a choice $r\in\lbrack0,1].$
Definition $\alpha$-quantile: Let $Q$ denote the set of all $\alpha$-quantiles $q$. For $\alpha\in(0,1)$ we define the $\alpha$-quantile as
When $\alpha=0$ we use $\sup Q$ (that is, we use $r=1$) and when $\alpha=1$ we use $\inf Q$ (that is, we use $r=0).\medskip$
From now on, by $q_{s,\alpha}(Y_{B(\delta(w_{N}))})$ we denote the $\alpha $-quantile of $Y_{B(\delta(w_{N}))}$ when the state of the world is $s.$ For simplicity of notation, unless needed for clarity, we do not index that expression by $r.$
By definition, regret of the treatment rule $\delta$ for a given distribution $s$ of $(Y_{0},Y_{1})$ equals
where $\mathbb{D}$ denotes the set of all possible treatment rules. In words, regret equals the difference between the highest $\alpha$-quantile that could have been achieved for any treatment rule $d\in\mathbb{D}$ and the $\alpha $-quantile obtained for the particular treatment rule $\delta$ for the given state of nature $s.$ If $q_{s,\alpha}(Y_{1})\geq q_{s,\alpha}(Y_{0})$ ($q_{s,\alpha}(Y_{1})<q_{s,\alpha}(Y_{0}))$ and $d^{\ast}=1$ ($d^{\ast}=0$) is an element of $\mathbb{D}$ then $\sup_{d\in\mathbb{D}}u(d,s)$ is taken on by $d^{\ast}$ and equals $q_{s,\alpha}(Y_{1})$ ($q_{s,\alpha}(Y_{0})$), see Lemma 3(ii) in the Appendix. If restrictions are imposed on $\mathbb{D}$ then $\sup_{d\in\mathbb{D}}u(d,s)$ may not be taken on by any element in $\mathbb{D}\mathfrak{.}$ Again, without restrictions, $1(q_{s,\alpha} (Y_{1})\geq q_{s,\alpha}(Y_{0}))$ is an infeasible optimal rule$,$ where $1(\cdot)$ denotes the indicator function.
In this paper, we focus on minimax regret treatment rules. By definition, if it exists, such a rule satisfies
In contrast, a maximin treatment rule, if it exists, satisfies
As discussed already elsewhere (see e.g., Manski (2004), Stoye (2009)) the maximin criterion may lead to the uninformative result that all $\delta\in\mathbb{D}$ are maximizers. That occurs e.g., if for a particular $s^{+}\in\mathbb{S}\mathfrak{,}$ $u(\delta,s^{+})$ does not depend on $\delta$ and takes on its smallest possible value, $u(\delta,s^{+})=\inf_{s\in \mathbb{S}}u(\delta,s).$ That situation also occurs in our setup where the objective is to maximize the quantile of the outcome distribution, namely when $s^{+}\in\mathbb{S}$ is chosen such that $q_{s,\alpha}(Y_{0})=q_{s,\alpha }(Y_{1})=0.$ In contrast to the setup where expected outcomes are the objective, as we will establish next, it turns out that for quantiles, depending on the particular sampling design, also the minimax regret criterion may be uninformative.
In this subsection\footnote{This section is inspired by the constructive comments of an anonymous referee.} we further discuss the proposed criterion introduced in ((ref)) and juxtapose it with other possible approaches of how quantiles could be incorporated into a Waldean framework of statistical decision theory.
In a series of pathbreaking contributions Wald (1945, 1947, 1950) introduced a framework for statistical decision theory whose main components are a statistical model, an action space, statistical decision rules, a welfare and an expected welfare function, and finally an evaluation criterion. Instead of welfare and expected welfare the presentation could be framed in terms of loss (defined as negative welfare) and risk (defined as expected loss).\ See Hirano (2025) for a comprehensive review that also includes contemporary developments.
In the particular context considered here where the DM\ needs to pick treatment 0 or 1 after observing the i.i.d. sample $w_{N}$ a minimax regret rule in the Waldean formulation solves
where $\mu_{t}=E_{s}Y_{t}$ for $t=0,1$ and $E_{s}$ denotes the expectation operator under $s.$ It is easily shown that $s$ matters only via $(\mu_{0} ,\mu_{1}).$ By the law of iterated expectations,
overall uncertainty can be separated into sampling uncertainty through $w_{N},$ (potential) randomness through the treatment assignment $B,$ and randomness through the potential outcome variables $(Y_{0},Y_{1} )$; risk is defined as the average loss over sampling uncertainty.
To adapt the Waldean framework to one that is based on the notion of $\alpha $-quantile rather than expectation, we suggest replacing expectations by $\alpha$-quantiles in the formulation ((ref)), that is, we suggest solving ((ref)) which is
Given there is no equivalent to the law of iterated expectations for quantiles, by doing so, one loses the separation of the regret criterion into loss without sampling uncertainty and sampling uncertainty. Our proposed criterion aggregates joint uncertainty (sampling, treatment assignment, and outcome) before taking the quantile.\footnote{Note that $q_{s,\alpha }(Y_{B(\delta(w_{N}))}|w_{N})$ does not in general equal $\delta (w_{N})q_{s,\alpha}(Y_{1})+(1-\delta(w_{N}))q_{s,\alpha}(Y_{0})$ in cases where $\delta(w_{N})\in(0,1).$}
Instead, to maintain such separation, one could first define regret loss given a realized sample $w_{N}$ as
and then define an alternative regret to ((ref)) as the average of regret loss over the sampling distribution $w_{N}$
The first term in ((ref)) is a benchmark attainable if one could pick the better treatment in terms of quantile outcome without sampling uncertainty; the second term is the $\alpha$-quantile of $Y_{B(\delta(w_{N} ))}$ conditional on the realized sample $w_{N}$. This criterion may be misaligned when the DM cares about tails in the distribution of $q_{s,\alpha }(Y_{B(\delta(w_{N}))}|w_{N})$ over realized samples rather than the average over sampling uncertainty. As yet another alternative, one could consider a regret function defined as the $\beta$-quantile (for some $\beta$ in the unit interval) of the loss function in ((ref)), that is
By the monotone transform identity of quantiles it follows that
and the interpretation of the criterion is not straightforward.
Manski and Tetenov (2023, Section 6.1) suggest an entire class of alternative approaches that also maintain separation. Starting with an arbitrary loss function (for instance negative welfare) for a given action by the DM (i.e. in our setup, a choice of treatment 0 or 1) and DGP\ $s,$ the suggestion is to consider for example an $\alpha$-quantile of the loss function over the sampling distribution. In our setup, such an approach could be formulated as
One issue with implementation of the various criteria is that analytical formulas are not generally available.\footnote{Recently suggested numerical procedures for the implementation of minimax rules by Aradillas Fern\'{a}ndez et al. (2025) and Guggenberger and Huang (2025) might be applicable also for these scenarios.}
We do not have a strong opinion about which one of the above criteria is generally preferable. However, we find our criterion very natural in situations where a DM applies treatment to an individual (e.g., herself) after observing the sample. Consider, for example, an individual who books a hotel accommodation on an internet platform, picking one of two options after observing a certain number of ratings on each (or picks one of two restaurants after having observed a number of ratings for each on the internet). The outcome is the rating the DM assigns to the hotel she booked, which could be interpreted as a proxy for the welfare the DM received from staying at the hotel. Another example is a medical doctor who picks one of two treatments for a patient after observing a certain number of outcomes on the two treatments.
In that type of example, when repeating the exercise, every observed outcome combines the randomness of the sample, (potential) randomness of treatment assignment, and the randomness of $(Y_{0},Y_{1}).$ If the DM is concerned about quantiles of the outcome distribution it seems that the criterion proposed in ((ref)) is a natural choice.
Recall that a sample of size $N$ has the general structure $w_{N}=(t,y).$ We consider three different sample designs, namely:
(i) Fixed number of untreated/treated units with $N_{0}\in\{0,...,N\}$ i.i.d. observations of $Y_{0}$ and $N_{1}:=N-N_{0}$ i.i.d. observations of $Y_{1}$ for some $N\in\mathbb{N}\cup\{0\}$. Here, the notation for a sample can be simplified by dropping $t.$ We write $w_{N}=(y(0)^{\prime} ,y(1)^{\prime})^{\prime}\in$ $[0,1]^{N}$ where $y(0)\in\lbrack0,1]^{N_{0}}$ contains the $N_{0}$ observations on untreated units and $y(1)\in \lbrack0,1]^{N_{1}}$ contains the $N_{1}$ observations on the treated units. A treatment rule $\delta\in\mathbb{D}$ is then any mapping $[0,1]^{N} \rightarrow\lbrack0,1].$ It assigns a treatment probability $\delta(w_{N} )\in\lbrack0,1]$ after observing the sample.
(ii) Random assignment with $N\in\mathbb{N}\cup\{0\}$ i.i.d. observations, where in the sample, the treatment probability equals $p\in(0,1).$\footnote{Note that $p\in\{0,1\}$ leads back to design (i) with $N_{0}$ or $N_{1}$ equal to $N.$} A treatment rule $\delta\in\mathbb{D}$ is any mapping $\{0,1\}^{N}\times\lbrack0,1]^{N}\rightarrow\lbrack0,1].$ The rule $\delta$ assigns a treatment probability $\delta(w_{N})\in\lbrack0,1]$ after observing a sample $w_{N}=(t,y)$ of treatment statuses $t=(t_{1},...,t_{N})$ and realizations $y=(y_{1},...,y_{N}),$ where the $i$-th component of $y,$ $y_{i},$ is an independent realization of $Y_{t_{i}}$.
(iii) Testing an innovation, is the case where aspects of the distribution of $Y_{0}$ are known to the DM; in particular, we assume that the $\alpha$-quantile of $Y_{0}$ is known (but nature can pick arbitrary distributions for $Y_{0}$ subject to that restriction). That is, in this case the set $\mathbb{S}$ consists of all joint distributions $s$ for $(Y_{0} ,Y_{1})$ with the restriction that the $\alpha$-quantile of the marginal for $Y_{0}$ equals a certain value $q_{\alpha}(Y_{0}).$ In this case, $N\in\mathbb{N}\cup\{0\}$ i.i.d. observations of $Y_{1}$ are observed. Here again the notation for a sample can be simplified. We write $w_{N} =y(1)\in\lbrack0,1]^{N}$ where $y(1)$ contains the $N_{1}$ observations of the treated units. A treatment rule $\delta\in\mathbb{D}$ is any mapping $[0,1]^{N}\rightarrow\lbrack0,1].$ The rule $\delta$ assigns a treatment probability $\delta(w_{N})\in\lbrack0,1]$ after observing a sample $w_{N}=(y_{1,1},...,y_{1,N})$ of $N$ independent realizations $y_{1,i},$ $i=1,...,N$ of $Y_{1}.$
The designs above nest the ones in Stoye (2009). In contrast to Stoye (2009), in design (i) we allow for arbitrary numbers of treated and untreated units rather than $N/2$ units each as in \textquotedblleft matched pairs\textquotedblright\ and in design (ii) the treatment probability is any fixed number $p\in(0,1)$ rather than necessarily $p=.5$.
The following statement provides the analogue to Proposition 1 in Stoye (2009) and is the main contribution of this paper.
Comments. 1. Proposition (ref) establishes that the minimax regret criterion when applied to $\alpha$-quantiles does not favor data-driven rules. In fact, in the \textquotedblleft testing an innovation\textquotedblright case data-driven rules are strictly dominated by $\delta^{0}\equiv0$ when $q_{\alpha}(Y_{0})>1/2$ and weakly dominated by $\delta^{1}\equiv1$ when $q_{\alpha}(Y_{0})<1/2.$ When $q_{\alpha}(Y_{0})=1/2$ all treatment rules are minimax. This is in stark contrast to the results in Stoye (2009) about minimax regret treatment rules when the focus is on mean outcomes. Namely Stoye (2009) shows that e.g., in the case of binary outcomes where the sample is obtained as a matched pair, to be minimax regret optimal, the treatment that has more successes in the sample must be chosen with probability one.
2. To provide intuition of the result in (i)-(ii) assume $\alpha\in(0,1)$ and consider first the simplest case where the sample size $N$ is 0. In that case, a treatment rule $\delta$ is simply an element in $[0,1]$ that denotes the probability of assigning treatment. For given $\delta$, the objective for nature is to find a distribution for $(Y_{0},Y_{1})$ such that the $\alpha $-quantiles of the marginals have maximal distance (that is distance 1) and such that the $\alpha$-quantile of $Y_{B(\delta)}$ is zero. Let $\delta>0$ first$.$ Assume nature picks $Y_{0}$ and $Y_{1}$ as independent Bernoulli
and, consequently, $P(Y_{0}=1)=1-(\alpha-\varepsilon).$ Thus $q_{s,\alpha }(Y_{1})=0$ and $q_{s,\alpha}(Y_{0})=1$ and
Thus, $P(Y_{B(\delta)}=0)>\alpha$ iff $\delta(1-\alpha)>\varepsilon(1-\delta)$ which holds for $\varepsilon$ small enough. Thus
for all $\delta\in(0,1].$ If instead $\delta=0$ then nature chooses $P(Y_{1}=1)=1$ and $P(Y_{0}=0)=1$ which leads to regret of 1.
This result may be surprising at first. For example, if the DM tries to be completely balanced and picks $\delta=1/2$ (which seems reasonable in the no data case) nature can inflict regret equal to one by picking $P(Y_{1}=0)=1$ and e.g., $P(Y_{0}=0)=\alpha-\min\{\alpha/2,(1-\alpha)/2\}.$ This is in stark contrast to the case considered in Stoye (2009) where the DM cares about expected outcome under $\delta,$ that is, $\mu_{0}(1-\delta)+\mu_{1}\delta,$ where $\mu_{t}$ for $t=1,2$ denotes the expectation of $Y_{t}$ under $s.$ Regret for given $\delta$ and $s$ (which only matters via $(\mu_{0},\mu_{1}))$ then equals $\max\{\mu_{0},\mu_{1}\}-(\mu_{0}(1-\delta)+\mu_{1}\delta).$ When the DM picks a $\delta$ with $\delta\geq1/2$ then the maximal regret nature can inflict equals $\delta$ obtained for any $s$ with $E_{s}Y_{0}=1$ and $E_{s}Y_{1}=0.$ Therefore, the DM's minimax regret choice is $\delta=1/2.$ What explains the different results with quantiles and expectations? What drives the results with quantiles is the discontinuity of the $\alpha $-quantile of a random variable $X$ with respect to the cdf $F_{X}$ of $X.$ In the above construction the random variables $Y_{0}$ and $Y_{B(\delta)}$ have cdfs that are uniformly \textquotedblleft very close\textquotedblright\ yet their $\alpha$-quantiles differ by 1. E.g. take again $\delta=1/2$ and consider $\alpha=.99.$ Then both $Y_{0}$ and $Y_{B(\delta)}$ are Bernoulli with $P(Y_{0}=0)=.99-.005=.985$ and $P(Y_{B(\delta)}=0)=1/2+1/2\cdot .985=.9925.$ Thus, their cdfs are almost identical but their $.99$-quantiles differ by 1. Instead the expectations of these two random variables are very close.
Surprisingly, the intuition of the no-data example generalizes to cases with arbitrary sample size $N>0.$ Again, for any given treatment rule $\delta \in\mathbb{D}$ one constructs an $s_{\delta}\in\mathbb{S}$ for which $\max\{q_{s_{\delta},\alpha}(Y_{0}),q_{s_{\delta},\alpha}(Y_{1})\}=1$ and $u(\delta,s_{\delta})=0$ and thus $R(\delta,s_{\delta})=1.$ Again $s_{\delta }\in\mathbb{S}$ can be chosen such that $Y_{0}$ and $Y_{1}$ are independent and both have Bernoulli distributions. That is, $s_{\delta}$ is then fully described by the two parameters $P_{s_{\delta}}(Y_{0}=0)$ and $P_{s_{\delta} }(Y_{1}=0).$ Denote by $0^{N_{0}}$ an $N_{0}$-dimensional column vector of zeros.
In case (i) of Proposition (ref), if $\delta$ is such that $\delta((0^{N_{0}\prime},v^{\prime})^{\prime})=1$ for all $v\in\{0,1\}^{N_{1}}$ then nature attempts punishing the DM for always using treatment 1 when no successes are observed for treatment 0, by choosing $P_{s_{\delta}}(Y_{1}=0)=1$ and by choosing $P(Y_{0}=0)$ \textquotedblleft big enough\textquotedblright\ that only zeros are observed for the untreated individuals \textquotedblleft sufficiently often\textquotedblright \ guaranteeing $Y_{B(\delta(w))}$ has $\alpha$-quantile 0, but small enough that $Y_{0}$ has $\alpha$-quantile 1. The proof shows that this can indeed be achieved by picking $P(Y_{0}=0)=\alpha-\varepsilon$ for some small enough $\varepsilon>0.$
If on the other hand $\delta$ is such that $\delta((0^{N_{0}\prime},v^{\prime })^{\prime})<1$ for at least one $v\in\{0,1\}^{N_{1}}$ then nature attempts punishing the DM for not using treatment 1 often enough when no successes are observed for treatment 0, by choosing $P_{s_{\delta}}(Y_{0}=0)=1$ and by choosing $P(Y_{1}=0)$ small enough that $Y_{1}$ has $\alpha$-quantile 1 but \textquotedblleft big enough\textquotedblright\ that $Y_{B(\delta(w))}$ has $\alpha$-quantile 0. The proof shows that this can indeed be achieved by picking $P(Y_{0}=0)=\alpha-\varepsilon$ for some small enough $\varepsilon>0.$
A surprising fact about the proof is that exploiting properties of $\delta$ on the set of samples $\{(0^{N_{0}\prime},v^{\prime})^{\prime},$ $v\in \{0,1\}^{N_{1}}\}$ alone gives nature enough leverage to inflict maximal regret.
3. In part (iii) of Proposition (ref), in the case $q_{\alpha}(Y_{0})<1/2$ we could not rule out that there are other minimax regret rules besides $\delta^{1}\equiv1$ when $N>0.$ When $N=0,$ it is obvious that $\delta^{1}$ is the unique minimax rule. A minimax regret treatment rule, if it is not unique, may be inadmissible. In the case of \textquotedblleft testing an innovation\textquotedblright\ $\delta^{0}\equiv0$ (when $q_{\alpha}(Y_{0})>1/2)$ is admissible, but we have not determined whether that is true for $\delta^{1}$ (when $q_{\alpha}(Y_{0})<1/2)$ and for $\delta^{.5}\equiv1/2$ (when $q_{\alpha}(Y_{0})=1/2).$
4. In case $Y_{0},Y_{1}\in(0,1)$ rather than $Y_{0},Y_{1}\in\lbrack0,1],$ $\max_{s\in\mathbb{S}}R(\delta,s)$ does not generally exist when $\mathbb{S}$ denotes the set of all joint probability distribution over the potential outcomes for $Y_{0},Y_{1}\in(0,1)$. E.g. in cases (i) and (ii) nature choosing only distributions with support on $[\varepsilon,1-\varepsilon]$ for some small $\varepsilon>0$ it can generate regret of $1-2\varepsilon$ and therefore $\sup_{s\in\mathbb{S}}R(\delta,s)=1$ for the unrestricted space of distributions for $Y_{0},Y_{1}\in(0,1).$ This result can be proven using the exact same proof technique as for Proposition (ref). Therefore, if minimax rules are defined with $\sup_{s\in\mathbb{S}}R(\delta,s)$ (as we do in ((ref))) rather than $\max_{s\in\mathbb{S}}R(\delta,s)$ the results in Proposition (ref)(i)-(ii) continue to hold.
5. If $Y_{0},Y_{1}\in\lbrack0,\infty)$ rather than $Y_{0},Y_{1}\in \lbrack0,1],$ then $\max_{s\in\mathbb{S}}R(\delta,s)$ does not typically exist when $\mathbb{S}$ denotes the set of all joint probability distributions over the potential outcomes for $Y_{0},Y_{1}\in\lbrack0,\infty)$. E.g. in cases (i) and (ii) nature choosing only distributions with support on $[0,M]$ for $M>0$ it can generate regret of $M$ and therefore $\sup_{s\in\mathbb{S}} R(\delta,s)=\infty$ when $\mathbb{S}$ denotes the set of all joint distributions for $Y_{0},Y_{1}\in\lbrack0,\infty).$ This result can be proven using the exact same proof technique as for Proposition (ref). Therefore, if minimax rules are defined with $\sup_{s\in\mathbb{S}}R(\delta,s)$ (as we do in ((ref))) rather than $\max_{s\in\mathbb{S}}R(\delta,s)$ then for $\alpha\in(0,1)$ any treatment rule is minimax regret optimal in cases (i)-(ii).
Given the limit experiment results in Hirano and Porter (2009) it would also be interesting to study the case where potential outcomes are restricted to be normally distributed. However, in that case, one would need to employ other proof techniques than the ones used in the current paper.
Restrictions on nature's action space
In what follows we impose various restrictions on nature's action space and study the implications on the results obtained in Proposition (ref). For simplicity assume $\alpha\in(0,1).$
Comments. 1. Corollary (ref)(a) is a direct corollary from the proof of Proposition (ref). In the proof of Proposition (ref)(i)-(ii) only Bernoulli distributions are used for nature while in part (iii) only discrete distributions are used.
2. Corollary (ref) (b) considers the case where nature is restricted to continuous distributions. We have seen in Proposition (ref) that in cases (i)-(ii) any treatment rule $\delta\in\mathbb{D}$ is minimax and $\max_{s\in\mathbb{S}}R(\delta,s)=1.$ Because without pointmasses it is impossible for a random variable to have $\alpha$-quantile equal to 0 it follows that $R(\delta,s)$ is always strictly smaller than 1 for any pair $(\delta,s)$ when $s$ is continuous with respect to Lebesgue measure. The main construction in the proof of Proposition (ref) still goes through when one considers a sequence of continuous distributions that converge to the Bernoulli distributions that are used in that proof. As a technical detail it is important to use $\sup_{s\in\mathbb{S}}R(\delta,s)$ rather than $\max _{s\in\mathbb{S}}R(\delta,s)$ in the definition of regret, because $\max _{s\in\mathbb{S}}R(\delta,s)$ would not exist in the case considered here. One can establish that $\sup_{s\in\mathbb{S}}R(\delta,s)=1$ for any treatment rule $\delta\in\mathbb{D}\mathfrak{.}$
3. In the case of \textquotedblleft testing an innovation\textquotedblright a restriction on nature's action space occurs if one assumes that the entire distribution of $Y_{0}$ is known, not just its $\alpha$-quantile. Namely, in that case the set $\mathbb{S}$ consists of all possible distributions $s$ for $Y_{1}\in\lbrack0,1]$ (while the distribution of $Y_{0}$ is given)$.$ The analysis of that problem is more difficult. Denote by $F_{Y_{0}}$ and $q_{\alpha}(Y_{0})$ the cdf and $\alpha$-quantile of $Y_{0},$ respectively.
Take $r=0$ in ((ref)) and assume $F_{Y_{0} }(q_{\alpha}(Y_{0}))>\alpha$. Consider the case $N=0$ in which case a treatment rule $\delta\in\lbrack0,1]$ denotes the treatment probability. Under these assumptions the following statements hold.
Denote by $q(\delta)$ the smallest $q\in\lbrack0,q_{\alpha}(Y_{0})]$ such that
Then, when $q_{\alpha}(Y_{0})<1/2,$ $\delta^{1}\equiv1$ is the only minimax regret rule and $\max_{s\in\mathbb{S}}R(\delta^{1},s)=q_{\alpha}(Y_{0}).$ When $q_{\alpha}(Y_{0})=1/2$ any $\delta\in\lbrack0,1]$ is minimax regret with maximal regret equal to $q_{\alpha}(Y_{0}).$ Finally, when $q_{\alpha} (Y_{0})>1/2,$ $\delta\in\lbrack0,1]$ is minimax regret iff
Given that $q_{\alpha}(Y_{0})-q(\delta)$ is a weakly increasing function in $\delta,$ if we denote by $\delta^{\ast}\in\lbrack0,1]$ an intersection of $q_{\alpha}(Y_{0})-q(\delta)$ and $1-q_{\alpha}(Y_{0})$ then any $\delta \in\lbrack0,\delta^{\ast}]$ is minimax regret and for those $\delta$ we have $\max_{s\in\mathbb{S}}R(\delta,s)=1-q_{\alpha}(Y_{0})$.
In the Appendix, we give a proof of the statements made above. The results for the case $q_{\alpha}(Y_{0})>1/2$ are partly in contrast to Proposition (ref)(iii) where there is a unique minimax regret rule. However, maximal regret is the same here as in Proposition (ref)(iii). We have not yet generalized the results to arbitrary sample sizes $N>0$ and to the case $F_{Y_{0}}(q_{\alpha}(Y_{0}))=\alpha.$
For the case of \textquotedblleft testing an innovation\textquotedblright \ Proposition (ref)(iii) establishes that the minimax regret criterion when applied to $\alpha$-quantiles does not favor data-driven rules. In this section we conduct a simulation experiment to juxtapose the pointwise (in $s$) regret of the data-driven empirical success rule $\delta^{ES},$ defined by
where $q_{\alpha}(Y_{1},w)$ denotes the $\alpha$-sample quantile of $Y_{1}$ for the sample $w,$ with the regret of the minimax regret rule $\delta ^{1}\equiv1$ (in the case when $q_{\alpha}(Y_{0})\leq1/2),$ $\delta^{.5} \equiv1/2$ (in the case when $q_{\alpha}(Y_{0})=1/2),$ and the minimax regret rule $\delta^{0}\equiv0$ (in the case when $q_{\alpha}(Y_{0})\geq1/2)$.
Obviously, when simulating regret we cannot possibly include all states of nature $s\in\mathbb{S}$. For the simulation experiment, we create a \textquotedblleft sufficiently rich\textquotedblright\ subset of distributions for $(Y_{0},Y_{1}).$ Namely, for some $n,w\in\mathbb{N}$, we consider the set of states of nature $\mathbb{S}^{E}=\mathbb{S}^{E}(n,w)$ that consists of all discrete distributions $s$ for $Y_{1}$ supported on a grid
with probabilities
being of the form $i/w,$ $i=0,...,w;$ while for the distribution of $Y_{0}$ we consider two choices, namely
I) $Y_{0}\in\{0,q_{\alpha}(Y_{0})\}$ with $P(Y_{0}=0)=\alpha-\varepsilon$ and $P(Y_{0}=q_{\alpha}(Y_{0}))=1-(\alpha-\varepsilon)$ for $\varepsilon=.000001$ and
II) $Y_{0}$ being continuously distributed on [0,1] with density $f(x)$ equal to $\alpha/q_{\alpha}(Y_{0})$ for $x\leq q_{\alpha}(Y_{0})$ and equal to $(1-\alpha)/(1-q_{\alpha}(Y_{0}))$ otherwise. Note that in both I) and II) the $\alpha$-quantile of $Y_{0},$ $q_{s,\alpha}(Y_{0}),$ equals $q_{\alpha} (Y_{0}).$ We report results for all choices of $\alpha\in\{.1,.5,.9\}$, $q_{\alpha}(Y_{0})\in\{.1,.5,.9\},$ sample size $N=30,$ and $(n,w)=(6,12).$ The latter leads to $S^{E}$ having cardinality 18564.
For the given choices of $n,w,$ and $N$ and for each choice of $\alpha$ and $q_{\alpha}(Y_{0})$, for each state of nature $s\in\mathbb{S}^{E}(n,w)$ and one (of the two possible) choice of distribution for $Y_{0}$ we simulate regret for the four treatment rules $\delta^{ES},$ $\delta^{1},$ $\delta ^{.5},$ and $\delta^{0}$ by generating $R=100K$ samples of size $N$ by drawing i.i.d. observations of the distribution of $Y_{1}.$ We use $r=0$ when simulating $\alpha$-quantiles of the outcome distribution under the various treatment rules. We analytically calculate $\alpha$-quantiles for $Y_{1}$ for $Y_{B(\delta^{.5})}$, and likewise use the true $\alpha$-quantile $q_{\alpha }(Y_{0})$ of $Y_{0}$ when calculating regret. For a given treatment rule $\delta$ and a given state of nature $s$, regret is calculated as $R(\delta,s)=\max\{q_{\alpha}(Y_{0}),q_{s,\alpha}(Y_{1})\}-q_{s,\alpha }(Y_{B(\delta(w))}).$
We compare the treatment rules along several dimensions, namely
a) mean regret over all 18564 states of nature $s\in\mathbb{S}^{E}(n,w)$,
b) maximal regret over all states of nature $s\in\mathbb{S}^{E}(n,w)$,
c) minimal regret over all states of nature $s\in\mathbb{S}^{E}(n,w)$, and
d) the proportion of $s\in\mathbb{S}^{E}(n,w)$ for which regret for the empirical success rule $\delta^{ES}$ is smaller than regret for each one of its three competitors.
Just for clarity, in a) for each treatment rule we sum up its regret over all 18564 states of nature $s\in\mathbb{S}^{E}(n,w)$ and then report that sum divided by 18564. For a given $s\in\mathbb{S}^{E}(n,w)$, in our simulations, we interpret regret of $\delta^{ES}$ as smaller than regret of another rule, say $\delta^{0},$ if the simulated regret of $\delta^{ES}$ is smaller than the simulated regret of the rule $\delta^{0}$ minus a threshold of $\xi.$ If instead the simulated regret of $\delta^{ES}$ falls into the interval [(simulated regret of $\delta^{0})-\xi,$(simulated regret of $\delta^{0})+\xi $] we record regrets of the two rules as equal for that state of nature. We take $\xi=.00000001$ below. Similarly, when programming the empirical success rule in ((ref)) the event \textquotedblleft$q_{\alpha}(Y_{1},w)=q_{\alpha }(Y_{0})\textquotedblright$ is implemented as the simulated $\alpha$-quantile of $Y_{1}$ falling into the interval $[q_{\alpha}(Y_{0})-\xi,q_{\alpha} (Y_{0})+\xi].$
All reported results are rounded to the second digit after the comma and so a reported zero could be as large as .004.
In the tables below we do not include results for c) because those results turn out to be equal to zero for all treatment rules and all designs except for $\delta^{.5}$ when $q_{\alpha}(Y_{0})\in\{.1,.9\}$ in which case minimal regret equals .067 for all choices of $\alpha$ and both choices of distributions for $Y_{0}.$
TABLE I provides results for maximal and mean regret over all 18564 states of nature $s\in S^{E}(6,12)$ for the four different treatment rules. Results for rules that are minimax regret in a given setting are reported in bold.
We first discuss results for maximal regret. The treatment rules that are known to be minimax regret for unrestricted $\mathbb{S}$ also have smallest maximal regret over $\mathbb{S}^{E}(n,w)$ (relative to the other treatment rules considered here) for all cases except for case II) with $\alpha=.9,$ when $q_{\alpha}(Y_{0})=.1$ (in which case $\delta^{ES}$ has smaller maximal regret than the optimal rule $\delta^{1})$, when $q_{\alpha}(Y_{0})=.5$ (in which case it is known that all four rules are optimal but in finite samples for $\mathbb{S}^{E}(n,w)$ again $\delta^{ES}$ does best), and finally when $q_{\alpha}(Y_{0})=.1$ (in which case $\delta^{ES}$ has smaller maximal regret than the optimal rule $\delta^{0})$. The explanation is of course that $\mathbb{S}^{E}(n,w)$ does not contain those states of nature that would inflict the highest regret on $\delta^{ES}$ (like the particular Bernoulli and discrete distributions for $Y_{0}$ and $Y_{1},$ respectively, that are used in the proof of Proposition (ref)(iii)). Furthermore, reported finite sample maximal regret over $\mathbb{S}^{E}(n,w)$ for the optimal rules matches the theoretical value $\min\{q_{\alpha} (Y_{0}),1-q_{\alpha}(Y_{0})\}$ reported in Proposition (ref)(iii) except for the case II) with $\alpha=.9$ when $q_{\alpha}(Y_{0})=.5$ where the optimal rule $\delta^{ES}$ has maximal regret smaller than .5 (which occurs, again, because $\mathbb{S}^{E}(n,w)$ is not rich enough). An open question from Proposition (ref)(iii) is whether for $q_{\alpha }(Y_{0})<.5$ other minimax regret rules besides $\delta^{1}$ might exist. The simulations for $q_{\alpha}(Y_{0})=.1$ are compatible with the possibility that when $\alpha=.5$ or $.9$ also $\delta^{ES}$ might be minimax regret.
We next discuss results for mean regret. In most cases where $q_{\alpha} (Y_{0})\in\{.1,.9\}$ in which case either $\delta^{0}$ or $\delta^{1}$ are minimax regret, their mean performance is also best (or very close to best) among the four treatment rules. On the other hand, when $q_{\alpha}(Y_{0})=.5$ (in which case all four treatment rules are minimax regret according to Proposition (ref)(iii)) we see huge difference in mean performance across the four treatment rules for a given case I) or II) and $\alpha,$ but also huge differences in performance for a given treatment rule and $\alpha$ across cases I) and II). With respect to the former point, in case I) when $\alpha=.1$ mean regret for the four rules are in the interval [0,.43]. With regards to the latter point, e.g., for $\delta^{ES}$ when $\alpha=.1$ mean regret equals .43 and .06, in cases I) and II) respectively. Differences across cases I) and II) are also often huge for other quantiles. E.g. again for $\delta^{ES}$ when $q_{\alpha}(Y_{0})=.9$ and $\alpha=.1$ mean regret equals .9 and 0, in cases I) and II) respectively. (Recall that we round results to the second digit. When $q_{\alpha}(Y_{0})=.9$ and $\alpha=.1$ there are not many distributions for $Y_{1}$ in $\mathbb{S} ^{E}(n,w)$ that have an $\alpha$-quantile that exceeds .9.)
We next discuss the results for exercise d) contained in TABLE II where we report the proportion of the 18564 states of nature $s\in S^{E}(6,12)$ for which the regret of $\delta^{ES}$ is smaller than the regret of $\delta$ where consider all $\delta$ from the set of no-data rules $\{\delta^{1},\delta ^{.5},\delta^{0}\}$. The results indicate that for many states $s$, choices of $q_{\alpha}(Y_{0})$, $\alpha,$ and the distribution for $Y_{0},$ $R(\delta^{ES},s)$ and $R(\delta,s)$ are very close which leads to numerically unstable results. To deal with the instability we introduce the buffer $\xi$ as explained above and $Prop(R(\delta^{ES},s)<R(\delta,s))$ and $Prop(R(\delta ^{ES},s)\leq R(\delta,s))$ in TABLE II\ represent the proportion of states $s$ for which $R(\delta^{ES},s)<R(\delta,s)-\xi$ and $R(\delta^{ES},s)<R(\delta ,s)+\xi,$ respectively. Those two proportions can be drastically different, suggesting that in many scenarios regret for $\delta^{ES}$ and $\delta$ are (virtually) identical. According to the measure $Prop(R(\delta^{ES},s)\leq R(\delta,s)),$ maybe somewhat surprisingly given the results from TABLE I, $\delta^{ES}$ is to be preferred over the other three rules in all scenarios considered in Case II) (except the case $\alpha=q_{\alpha}(Y_{0})=.9$ for $\delta^{0})$ and in the majority of scenarios in Case I) (except compared to $\delta^{0}$ when $q_{\alpha}(Y_{0})=.5,\alpha=.1$ and except for most cases with $q_{\alpha}(Y_{0})=.9$ and all three rules$).$ Quite often $Prop(R(\delta ^{ES},s)\leq R(\delta,s))$ is reported as higher than 95%, but the improvement in regret for many states of nature is minuscule. For example compared to $\delta^{.5}$ in Case I) with $\alpha=q_{\alpha}(Y_{0})=.1$ only in 33.3% of the cases $Prop(R(\delta^{ES},s)<R(\delta,s))$ while for 100% of the cases $Prop(R(\delta^{ES},s)\leq R(\delta,s))$ implying that the regret of $\delta^{ES}$ in 66.4% of the cases is at most $\xi$ smaller than the regret of $\delta^{.5}$.
Next, as in Stoye (2009) we next allow for a discrete covariate $X\in \mathcal{X}=\{x_{1},...,x_{K}\}$ that is observed both in the sample and in the treatment data. Outcomes $Y_{t,x}$ now carry a double subindex to indicate the treatment status $t\in\{0,1\}$ and the value of the covariate $x\in\mathcal{X}.$ It is assumed that $x_{k}$ occurs with positive probability for each $k=1,...,K.$ Denote by $F_{X}$ the distribution of $X.$ A state of the world $s\in\mathbb{S}$ now represents a joint distribution for $(Y_{t,x})_{t\in\{0,1\},\text{ }x\in\mathcal{X}}.$
A sample $w=w_{N}$ now consists of realizations $(t_{i},x_{i},y_{i})$ for $i=1,...,N$ of $(T,X,Y_{T,X}),$ where $y_{i}$ is a realization of $Y_{t_{i},x_{i}}$ and we consider again sampling designs (i)-(iii) from Section (ref). It is assumed that $x_{i},$ $i=1,...,n,$ are i.i.d. and $F_{X}$ is independent of $s$ and $T$.
In design (i), \textquotedblleft fixed number of treated/untreated units\textquotedblright, there are $N_{0}$ observations $(0,x_{i},y_{i})$ with $y_{i}$ being i.i.d. draws of $Y_{0,x_{i}},$ $i=1,...,N_{0}$ and $N_{1}=N-N_{0}$ observations $(1,x_{i},y_{i})$ with $y_{i}$ being i.i.d. draws of $Y_{1,x_{i}},$ $i=N_{0}+1,...,N.$ Note that there are not typically equally many treated and untreated observations for each covariate $x_{k},$ $k=1,...,K.$ In fact, there may be zero observations altogether in the sample for a given $x_{k}.$
In design (ii), \textquotedblleft random sampling\textquotedblright, the observations $(t_{i},x_{i},y_{i})$ for $i=1,...,N$ are i.i.d. realizations of $(T,X,Y_{T,X})$ with $P(T_{i}=1)=p\in(0,1)$ with $T,X$, and $s$ being independent of each other$.$
In design (iii), \textquotedblleft testing an innovation\textquotedblright, $(1,x_{i},y_{i})$ for $i=1,...,N$ are observed where $y_{i}$ is an independent realization of $Y_{1,x_{i}},$ for $i=1,...,N.$
In design (iii), we assume the DM knows the $\alpha$-quantile, denoted by $q_{\alpha}(Y_{0,X}),$ of $Y_{0,X}$ and nature can choose from joint distributions $s\in\mathbb{S}$ for $(Y_{t,x})_{t\in\{0,1\},\text{ } x\in\mathcal{X}}$ such that the $\alpha$-quantile of $Y_{0,X}$ equals $q_{\alpha}(Y_{0,X}).$
A treatment rule $\delta$ maps a sample $w_{N}$ onto treatment probabilities for each $x_{k}$ for $k=1,...,K,$ that is $\delta(w_{N})\in\lbrack0,1]^{K},$ where the $k$-th component of $\delta(w_{N})$ indicates the treatment probability for individuals with covariate $x_{k}$ for $k=1,...,K.$ Denote the $k$-th component of $\delta(w_{N})$ by $\delta_{x_{k}}(w_{N})$ for $k=1,...,K.$ With some abuse of notation, for each design (i)-(iii) we denote the set of all treatment rules by the same symbols $\mathbb{D}$ (even though it means something different for different designs)$.$
The object of interest is
the $\alpha$-quantile of the outcome distribution. With that definition, regret is then defined formally analogously to the setup without a covariate, namely, $R(\delta,s)=\sup_{d\in\mathbb{D}}u(d,s)-u(\delta,s).$ Alternatively, one could focus on the $\alpha$-quantile of the outcome distribution for a particular covariate only, $x_{1}$ say, $u_{x_{1}}(\delta,s)=q_{s,F_{X} ,\alpha}(Y_{B(\delta_{x_{1}}(w_{N})),x_{1}})$ and obtain analogous results to the ones in Corollary (ref) below.
If the space of probability distributions for nature $\mathbb{S}$ is unrestricted and thus equals the space of all distributions for $(Y_{t,x} )_{t\in\{0,1\},\text{ }x\in\mathcal{X}}$ one can show in designs (i)-(ii) (as an implication of the proof of Proposition (ref)) that for every treatment rule $\delta$ maximal risk over $\mathbb{S}$ equals 1, $\max_{s\in\mathbb{S} }R(\delta,s)=1$. This result then immediately implies that every treatment rule is minimax regret. The corollary (to Proposition (ref)) that follows gives a stronger result; it shows that maximal risk continues to be 1 even if certain restrictions are imposed on $\mathbb{S}$. For simplicity of the presentation assume $\alpha\in(0,1).$
Comments: 1. Various variants could be considered in design (iii). E.g. one could instead assume that the joint distribution $(Y_{0,x} )_{x\in\mathcal{X}}$ is known (in which case nature only chooses a joint distribution for $(Y_{1,x})_{x\in\mathcal{X}})$ or one could assume that the DM knows the vector $(q_{\alpha}(Y_{0,x_{1}}),...,q_{\alpha}(Y_{0,x_{K}}))$ of $\alpha$-quantiles of all the marginal distributions $(Y_{0,x})_{x\in \mathcal{X}}$ and nature can choose from joint distributions $s\in\mathbb{S}$ for $(Y_{t,x})_{t\in\{0,1\},\text{ }x\in\mathcal{X}}$ such that all the marginals $(Y_{0,x})_{x\in\mathcal{X}}$ have the required $\alpha$-quantiles.
2. One could also consider alternative sampling designs (i)-(iii), where in the sampling stage, rather than being randomly assigned, the values of the covariate $X$ are assigned deterministically, as is done for treatment status in design (i). This could be referred to as \textquotedblleft stratified sampling.\textquotedblright
For brevity we do not explicitly deal with these variations.
We derive minimax regret treatment rules in finite samples when an $\alpha $-quantile of the outcome distribution is the focus of interest. We establish that when the sample i) consists of a fixed number of untreated/treated units or ii) is generated via random treatment assignment then all treatment rules are minimax regret and therefore the minimax regret criterion is not helpful in singling out a recommended treatment rule. Given that the same shortcoming applies to the max-min criterion, an important question concerns finding a meaningful criterion in this setup based on which an optimal treatment rule should be chosen. The idea from Montiel Olea, Qiu, and Stoye (2023) to look for rules that randomize \textquotedblleft the least\textquotedblright\ in a situation where there are multiple minimax regret rules would not lead to a unique rule in our setup because in cases i) and ii)\ both $\delta^{0}$ and $\delta^{1}$ are minimax regret and never randomize. Given all these facts, it then seems reasonable to simply adopt a rule that is minimax regret optimal when regret is based on the notion of expected welfare (in particular, such a rule is optimal according to criterion ((ref))).
We also establish that when iii) the sample consists of only realizations from the treated population while the $\alpha$-quantile of the untreated population is known, never treating is the unique minimax rule if the known quantile exceeds .5, while always treating is a minimax rule if that quantile is strictly smaller than .5, and finally, if the quantile equals .5, then any treatment rule is minimax regret. It follows that based on the minimax regret criterion, rules that take the data into consideration are never strictly preferred.
An interesting but quite difficult extension that we are currently investigating concerns applying the minimax regret criterion in finite samples to conditional value at risk, defined as $S_{s,\alpha}(Y):=\alpha^{-1} E_{s}[Y1(Y\leq q_{s,\alpha}(Y))]$ for a random variable $Y$ rather than to the $\alpha$-quantile of $Y$.\footnote{Or, using the more general definition $S_{s,\alpha}(Y):=\sup_{\gamma\in R}\{\gamma-\alpha^{-1}E_{s}[\gamma-Y]_{+} \}$, see Qi, Pang, and Liu (2023) were, $[Y]_{+}=\max\{Y,0\}$ denotes the positive part of $Y$.}