Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
116,375 characters · 30 sections · 89 citation commands
Optimal Decision Rules when Payoffs are Partially Identified
\onehalfspacing
\thispagestyle{empty}
\setcounter{page}{1}
Many important policy decisions involve discrete choices. Examples include whether or not to treat an aggregate population or large sub-population, firm or worker decisions at the extensive margin, and pricing policies when, in practice, prices must be expressed in whole currency units. Suppose a decision maker must choose a policy from a discrete choice set. The decision maker has data that may be used to bound, but not point identify, the payoffs associated with some choices. How should they proceed?
In this paper, we propose an approach for making asymptotically optimal discrete statistical (i.e., data-driven) decisions when the payoffs associated with some choices are only partially identified. We assume the decision maker observes data which may be used to learn about a vector of parameters $\mu$. The decision maker then chooses a policy from a finite set. The distribution of payoffs associated with the different policies depends on a parameter $\theta$. A key assumption underlying the analysis in this paper is that $\theta$ is possibly set-identified, but the parameters $\mu$ may be used to deduce restrictions on $\theta$. The decision maker therefore confronts both ambiguity (the payoff distribution is not point-identified) and statistical uncertainty ($\mu$ must be estimated from the data).
We propose a theory of optimal statistical decision making in this setting, building on a line of research going back to Manski2000. We depart from this body of work by combining a minimax approach to handle the ambiguity that arises from partial identification of $\theta$ given $\mu$ and average (or integrated) risk minimization for $\mu$. This asymmetric treatment of parameters is in the spirit of the generalized Bayes-minimax principle of Hurwicz1951. The resulting optimal decision rules are “robust” in the sense that they minimize maximum risk or regret over the identified set $\Theta_0(\mu)$ for $\theta$ conditional on $\mu$, and use the data to learn efficiently about features of $\mu$ germane to the choice problem.
Our optimal decision rules can be implemented very easily in realistic empirical settings. All we require is that, for every choice, the maximum risk (or regret) over $\theta \in \Theta_0(\mu)$ conditional on $\mu$ can be computed. The maximum risk is averaged across a bootstrap distribution for an efficient estimator $\hat \mu$ of $\mu$, a posterior distribution for $\mu$ in parametric models, or a quasi-posterior based on a limited-information criterion for $\mu$ in semiparametric models. Our optimal decision is then simply to choose whatever choice has smallest average maximum risk. We refer to this rule as a bootstrap decision when the maximum risk is averaged over a bootstrap distribution, a Bayes decision when averaged over a posterior,\footnote{These are not Bayes decision rules in the usual sense, but rather hybrid Bayes-minimax rules which, as we explain below, also have an interpretation as a type of robust Bayes (or conditional $\Gamma$-minimax) rules. For simplicity, we refer to them as Bayes decisions and drop the robust/minimax qualifiers.} and a quasi-Bayes decision when averaged over a quasi-posterior. Despite its simplicity, we provide a formal asymptotic optimality theory to justify our approach.
As running example, we show how to implement optimal decisions in the context of a treatment assignment problem under partial identification of the average treatment effect (ATE). In an empirical illustration, taken from IK, a decision maker decides whether or not to adopt a job-training program based on several RCT estimates from other studies and their standard errors. The identified set $\Theta_0(\mu)$ is defined through intersection bounds derived by extrapolating multiple studies. In addition to treatment assignment problems, we discuss how our framework can be used for optimal pricing decisions in an environment with rich unobserved heterogeneity, where revealed preference arguments may be used to derive bounds on demand responses under counterfactual prices. In both contexts, the maximum risk of different choices is available in closed form or can be computed by solving a standard optimization problem, e.g., a linear program.
To elaborate on practicality a little, consider the two competing paradigms: Bayes and minimax. In our setting, the usual Bayes decision rule requires specifying a prior on the parameter space for $(\theta,\mu)$, computing the posterior for $(\theta,\mu)$ having observed the data, then choosing the decision that minimizes posterior risk. A common criticism of Bayes decisions (see, e.g., Manski2021Haavelmo) is that they are only justified if one can elicit a credible subjective prior, which can be difficult in practice, and that the resulting decision will depend, to some extent, on the decision maker's choice of prior. While this is true under both point- and partial identification, the problem is more severe under the latter because the prior for $\theta$ is not updated by the data (e.g., MoonSchorfheide2012). Thus, even asymptotically, the decision will depend on the prior for $\theta$. By contrast, our (quasi-)Bayesian rules only require specifying a prior for $\mu$ and our decisions are asymptotically independent of this choice of prior. Moreover, our bootstrap rules sidestep the choice of a prior altogether.
Minimax decisions minimize the maximum risk over the parameter space for $(\theta,\mu)$. While this may be desirable, it is often not feasible. To derive minimax decisions one usually has to make strong assumptions on the data-generating process and restrict the dependence of payoffs on parameters to be of a very simple form. Indeed, we are not aware of any work deriving a minimax treatment rule under partial identification even in the simplest case of a binary outcome and binary treatment when randomization is not permitted and bounds on the ATE must be estimated from data. Algorithms such as that of Chamberlain2000 may be used to compute approximate minimax decisions, but the performance gap between these and the true minimax decision can be difficult to quantify. Adopting a minimax approach with respect to $\theta$ and a Bayes approach with respect to $\mu$ lends a great deal of tractability, allowing us to derive asymptotically optimal rules for a very broad class of empirically relevant settings where minimax rules are intractable.
We develop an asymptotic optimality theory for decisions based on parametric and semiparametric models. For parametric models, our optimality criterion extends the asymptotic average risk criterion introduced by HiranoPorter2009 for point-identified settings to partially identified settings. We further extend this notion to semiparametric models via a least favorable parametric submodel. Both of these extensions represent new contributions to the literature on asymptotic optimality for statistical decision rules. Our main results show formally that the proposed (quasi-)Bayesian rule is asymptotically optimal. Moreover, any decision rule that is asymptotically equivalent to the (quasi-)Bayes rule is optimal as well. This includes an implementation that replaces averaging under a (quasi-)posterior by averaging under a bootstrap approximation of the sampling distribution of an efficient estimator $\hat \mu$ of $\mu$, in cases in which these two distributions are asymptotically equivalent.
Importantly, we show asymptotic equivalence to the (quasi-)Bayes or bootstrap rules is necessary for optimality: any decision whose asymptotic behavior is different from these is sub-optimal. It follows that “plug-in” rules, which plug an efficient estimator $\hat \mu$ into the oracle decision rule if $\mu$ were known, can perform sub-optimally.\footnote{Manski2021Haavelmo,Manski2021 refers to plug-in rules as “as-if” optimization, because estimators $\hat \mu$ are treated as if they are the true parameters.} Manski2021Haavelmo,Manski2021 shows in numerical experiments that plug-in rules may perform poorly under finite-sample minimax regret criteria. Our necessity result provides a complementary and quite general theoretical explanation for why plug-in rules may perform poorly, albeit under a different (but related) optimality criterion.\footnote{For the intuition, consider a treatment assignment problem under partial identification of the ATE. Oracle rules depend on a robust welfare contrast $b(\mu)$ formed from bounds on the ATE. The bounds are non-smooth functions of $\mu$ in many empirically relevant settings reviewed in Section (ref). This non-smoothness leads to a failure of the $\delta$-method that breaks the asymptotic equivalence between plug-in rules, which depend on $b(\hat \mu)$, and optimal rules, which depend on the average of $b(\cdot)$ across a bootstrap or posterior distribution.} In our empirical application to the adoption of a job-training program, our optimal decision produces different treatment recommendations than the plug-in rule for some sub-populations.
On the technical side, partial identification often leads to non-differentiability of key components of the maximum risk function in $\mu$. We address this challenge by extending arguments in HiranoPorter2009 for models with smooth welfare contrasts to settings with directional (but not full) differentiability. An important technical building-block is the derivation of the asymptotic distribution of the (quasi-)posterior mean of directionally differentiable functions in both parametric and semiparametric models (see Propositions (ref) and (ref) in Appendix (ref)). These results are of independent interest. While KitagawaOleaPayneVelez2020 derived the asymptotic behavior of the posterior distribution of directionally differentiable functions, we instead characterize the large-sample (frequentist) distribution of the posterior mean of such functions. Our results offer a novel contribution to the literature on asymptotics for non-smooth functions Dumbgen1993,FangSantos2019. We also introduce the concept of $\sigma$-optimality to handle settings in which the average (under Lebesgue measure) maximum risk is infinite at the optimal decision.
Our paper complements prior work on treatment assignment under partial identification, including Manski2000,Manski2007treatment,Manski2020Meta,Manski2021Haavelmo,Manski2021, Chamberlain2011, Stoye2012, Russell2020, IK, and Yata2021. Except for Chamberlain2011, these works seek decision rules that are optimal under finite-sample minimax regret criteria. We depart from these works in two respects. First, our asymptotic framework allows us to relax restrictive parametric assumptions and accommodate a much broader class of data-generating processes, including semiparametric models. Our optimality results apply to settings where the decision maker cannot confidently assert that the data are drawn from a given parametric model, or where bounds on the ATE are estimated using a vector of moments or summary statistics (e.g. regression or IV estimates from observational studies) whose finite-sample distribution is unknown.\footnote{Finite-sample results are sometimes developed for this case assuming Gaussianity of the statistics, by arguing that the studies are sufficiently large that the sampling distribution of the statistics is approximately normal. In these cases, it seems logically consistent to use a large-sample optimality criterion.} Second, except for Chamberlain2011, these works use optimality criteria that (in our notation) are minimax over $(\theta,\mu)$, whereas our criterion is minimax over the partially-identified parameter $\theta$ and averages over the point-identified parameter $\mu$, reflecting the asymmetric parameterization of the problem.
Our approach is related to the multiple priors framework of GilboaSchmeidler1989 and a particular form of robust Bayes decision making, called conditional $\Gamma$-minimax; see DasGuptaStudden1989, BetroRuggeri1992, and the recent survey by GiacominiKitagawaRead2021. We show how, under specific informational assumptions, our decision rule arises from a two-player zero-sum game between a decision maker and an adversarial nature. In the conditional $\Gamma$-minimax version of this game, nature chooses a prior for $\theta \in \Theta_0(\mu)$ conditional on the reduced-form parameter $\mu$ and the data that are available for estimation of $\mu$.
In the optimal pricing application we consider an environment in which a monopolist is choosing to set a price but faces ambiguity about the true distribution of demand. As in BergemannSchlag2011, we assume the decision maker has a preference for robustness. However, we use revealed preference demand theory to derive bounds on demand, rather than assuming true demand is in a neighborhood of a pre-specified distribution. Moreover, in our setting the decision maker must estimate the bounds from data. This application builds on prior work on revealed-preference demand theory, including Blundelletal2007,BlundellBrowningCrawford2008, Blundelletal2014,Blundelletal2017, HoderleinStoye2015, Manski2007,Manski2014, and KitamuraStoye2018,KitamuraStoye2019. These works are primarily concerned with testing rationality or deriving bounds on demand. We instead focus on using bounds to solve an optimal pricing problem, and using demand data efficiently in that context.
The remainder of the paper is structured as follows. Section (ref) presents an application to treatment assignment based on extrapolating meta-analyses. We first discuss the oracle decision based on a known $\mu$, then describe our proposed optimal decision rule in this context, and finally provide an empirical illustration. Section (ref) outlines our general framework for optimal decision making, discusses Bayesian and bootstrap implementations, and provides comparisons to minimax and conditional $\Gamma$-minimax decisions. The large sample optimality theory is presented in Section (ref) for parametric models and Section (ref) for semiparametric models. Further applications to other treatment assignment problems and optimal pricing are discussed in Sections (ref) and (ref), respectively. Finally, Section (ref) concludes. Proofs and additional results and lemmas are relegated to the Appendix.
To set the stage, we will present a specific treatment assignment application in Section (ref), describe the proposed decision rule in Sections (ref) and (ref), and implement it numerically for the application in Section (ref).
Suppose a social planner is choosing whether or not to introduce a treatment in a target population. The social planner does not know the target population's ATE, defined as $\theta = \mathbb{E}[Y]$, where $Y:=Y_1-Y_0$ is the treatment effect associated with any individual, $Y_1$ is the outcome under treatment and $Y_0$ the outcome in the absence of treatment. Instead, the social planner observes a meta-analysis consisting of estimates $\hat{\mu}_k$ of the ATE $\mu_k$ in several related populations $k=1,\ldots,K$, their standard errors $s_k$, characteristics $x_k$ of the related populations, and characteristics $x_0$ of the target population.
For concreteness, we revisit Example 2 of IK. The authors consider a subset of $K = 14$ studies form the database of CardKluveWeber2017, each of which is an RCT looking at the impact of job training programs on employment. Studies are implemented in a number of different countries and in groups that differ by characteristics $x_k$ consisting of gender (males, females, or both), age (youths, adults, or both), OECD membership status, GDP growth (standardized) and unemployment (standardized). We consider the hypothetical question of whether to roll out a job-training program in two populations: German male youths in 2010 and German female youths in 2010 (with GDP growth of 3.48% and an unemployment rate of 9.45%).
To connect the RCT studies to the ATE $\theta$ associated with the target population, suppose that treatment effects are Lipschitz in the distance between size-specific covariates:
where $C$ is a pre-specified constant. We refer to $\mu := (\mu_1,\ldots,\mu_k)$ as the reduced-form parameter. It is assumed to be consistently estimable from the data and its estimate is denoted by $\hat{\mu}$. We obtain the following lower ($L$) and upper ($U$) bounds for $\theta$:
The bounds in ((ref)) define an identified set for $\theta$ as a function of $\mu$:
For now we will assume that $\mu$ is known. Following Manski2000,Manski2004, it is common to derive treatment rules under a utilitarian social welfare function that is linear in the target population's ATE. Interpreting negative utility as loss, we write \[ l(d,\theta) = -d \theta, \] where $d \in \{0,1\}$ indicates treatment. The loss function is minimized by the treatment decision $\mathbb{I}[ \theta \ge 0]$, where $\mathbb{I}[\cdot]$ is the indicator function that is equal to one if its argument is true and zero otherwise. Following Manski2000,Manski2004, we subsequently use the regret criterion
Failure to treat ($d=0$) incurs zero regret when $\theta <0 $, otherwise the regret is $\theta$. Similarly, treating ($d=1$) incurs zero regret when $\theta \geq 0$, otherwise the regret is $-\theta$.
If the social planner knows $\mu$, they can choose the decision that minimizes maximum regret over $\Theta_0(\mu)$, defined as \[ R(d,\mu) := \sup_{\theta \in \Theta_0(\mu)} r(d,\theta). \] In our specific application we obtain
where $(a)_+ := \max\{a,0\}$ and $(a)_- := \min\{a,0\}$. This leads to the oracle decision
that assigns treatment if the robust welfare contrast \[ b(\mu) := \left( b_U(\mu) \right)_+ + \left( b_L(\mu) \right)_- \] is non-negative, and non-treatment otherwise.
Notice from plugging ((ref)) into ((ref)) that the max and min operators make $R(0,\mu)$ and $R(1,\mu)$ only directionally differentiable with respect to $\mu$. Directional differentiability of the maximum risk function is a generic feature of all the applications discussed in this paper. It generates technical challenges that we tackle in our large sample analysis and distinguishes our analysis from that in HiranoPorter2009.
The oracle rule is an infeasible first-best as it requires knowledge of the true $\mu$. In practice $\mu$ is unknown and any practical rule must also confront sampling uncertainty about $\mu$. One option is to simply plug $\hat \mu$ into ((ref)), which leads to the plug-in rule
However, this rule does not account for the precision of $\hat \mu$. Manski2021 explores how different estimates of parameters affect the performance of plug-in rules.
We develop an optimality theory in Sections (ref) to (ref) below to guide decision-making in situations such as these where a payoff-relevant parameter $\theta$ is partially identified and its identified set $\Theta_0(\mu)$ is indexed by a parameter $\mu$ which can be estimated from data. In the context of the treatment assignment application, the procedure amounts to replacing $b(\mu)$ in ((ref)) by its average $\bar b_n$ under a posterior for $\mu$ (in parametric models) or quasi-posterior for $\mu$ (in semiparametric models):
We show formally in Sections (ref) and (ref) that this decision is more efficient than the plug-in rule, as it takes into account the sampling uncertainty in $\hat \mu$.
We use a limited-information quasi-posterior distribution for $\mu$ that interprets the sampling distribution $\hat{\mu}_k \stackrel{\mbox{\tiny approx}}{\sim} N(\mu_k,s_k^2)$ as a quasi-likelihood. The estimates $\hat \mu_k$ are independent as they come from independent RCTs. Following DoksumLo1990, Kim2002, and Mueller2013, we combine the $N(\mu, \hat \Sigma)$ quasi-likelihood for $\hat \mu$ with a flat prior for $\mu$ to obtain a $N(\hat \mu, \hat \Sigma)$ quasi-posterior for $\mu$, where $\hat \Sigma$ denotes a diagonal matrix with $(s_k^2)_{k=1}^K$ down its diagonal. The quasi-posterior mean of $b(\mu)$, denoted by $\bar{b}_n$, is computed by drawing independent $\mu \sim N(\hat \mu,\hat \Sigma)$, computing $b(\mu)$, then averaging across a large number of draws. The efficient decision from ((ref)) is then to treat if $\bar b_n$ is non-negative, and not treat otherwise.
We compute the optimal decision as described above using data from IK. We increase the Lipschitz constant in ((ref)) from $C = 0.025$ in their paper to $C = 0.25$ so that the bounds on $\theta$ are non-empty. Figures (ref) and (ref) plot the quasi-posterior distribution of $b(\mu)$ under $\mu \sim N(\hat \mu,\hat \Sigma)$ for males and females, respectively. Both figures display the mean $\bar b_n$ whose sign determines the optimal treatment decision $\delta^*_n$ in ((ref)). We also display the plug-in value $b(\hat \mu)$ whose sign determines the plug-in rule $\delta^{plug}_n$ in ((ref)).
In Figure (ref) (male youths) we see that the distribution has a pronounced right-skew. Recall that the lower bound is $\max_{1 \leq k \leq K} (\mu_k - C \|x_0 - x_k\|)$. When evaluated at $\hat \mu$, the largest value of $\hat \mu_k - C \|x_0 - x_k\|$ is $-0.3190$ (corresponding to a US study) and the second-largest is $-0.3298$ (corresponding to a Brazilian study). As the maximum is not well separated relative to sampling uncertainty (the average $s_k$ across the studies is $0.034$), the distribution of the lower bound is right-skewed because it behaves like a maximum of two Gaussians. The upper bound is less skewed because the minimum (corresponding to a Colombian study) is better separated. As a consequence of this asymmetry, the mean $\bar b_n$ is positive and the optimal decision is to treat. By contrast, the plug-in value $b(\hat \mu)$ is negative so the plug-in decision is to not treat. In Section (ref) we show this difference between the two rules, where the optimal rule predicts treatment more aggressively than the plug-in rule, is as predicted by our optimality theory.
The situation is different for female youths (Figure (ref)). Here the minima and maxima are relatively better separated, so the distribution of $b(\mu)$ for $\mu \sim N(\hat \mu,\hat \Sigma)$ is close to Gaussian. In consequence, the mean $\bar b_n$ and the plug-in value $b(\hat \mu)$ are almost identical, and the optimal and plug-in decisions are both to treat.
The treatment assignment application in the previous section imposed a lot of structure which we will now relax. We begin by stating the general decision problem in Section (ref). Many empirical examples are covered by our framework: see Sections (ref) and (ref) for applications to treatment assignment and optimal pricing and CMS2020forecast for an application to forecasting. We develop our concept of optimal decision rules in Section (ref). Bayesian and bootstrap implementations are presented in Sections (ref) and (ref). Finally, in Section (ref) we discuss a two-player zero-sum game representation of our setup and its relationship to minimax and robust Bayesian decision making.
A decision maker (DM) observes $X^n \sim P_{n,\mu}$ taking values in a sample space $\mc X^n$. Often $X^n$ may be a sample $(X_1,\ldots,X_n)$ of size $n$ with joint distribution $P_{n,\mu}$. But $X^n$ may also be a vector of estimators (e.g., OLS or IV estimators) with sampling distribution $P_{n,\mu}$. The distribution $P_{n,\mu}$ is indexed by a reduced-form parameter $\mu \in \mc M$.
After observing $X^n$, the DM chooses an action $d$ in the action space $\mc D = \{0,1,\ldots,D\}$. A statistical decision rule $\delta_n : \mc X^n \to \mc D$ maps realizations of the data into actions.\footnote{We consider nonrandomized rules as opposed to randomized rules where the action space is the set of all distributions over $\{0,1,\ldots,D\}$.} The DM's utility $u(d,Y,\theta,\mu)$ from choosing $d$ depends on a random variable $Y \sim G_\theta$ independent of $X^n$, a structural parameter $\theta \in \Theta$, and $\mu$. In a treatment assignment problem, $X^n$ may represent data or summary statistics from an experiment, $Y$ may represent the treatment effect for an individual sampled at random from the remaining population (hence independent of $X^n$), and $\mathbb{E}_\theta[Y]$ may represent the average treatment effect. Both $G_\theta$ and $\mathbb{E}_\theta[\,\cdot\,]$ can be conditional on covariates but we suppress this to simplify notation.
We interpret negative utility as loss and write $l(d,Y,\theta,\mu) \equiv -u(d,Y,\theta,\mu)$. The DM incurs this loss once $Y$ is realized. Here it is helpful to draw a distinction between the ex-post problem that the DM faces after observing $X^n$ but before $Y$ is realized, and the ex-ante problem the DM faces before $X^n$ is realized. In the ex-post problem, the DM faces risk \[ r(d,\theta,\mu) := \mathbb{E}_\theta[l(d,Y,\theta,\mu)] \] from choosing $d \in \mc D$, where the expectation is taken with respect to $Y \sim G_\theta$.\footnote{This notation nests welfare regret by setting $l(d,y,\theta,\mu) = \max_{d' \in \mc D} \; \mathbb{E}_\theta [W(d',Y,\theta,\mu)] - W(d,y,\theta,\mu)$ for a welfare function $W$. Hence, we do not distinguish between “risk” and “regret” in what follows.} We refer to $r$ as “risk” because it involves taking an expectation with respect to the random outcome $Y$ in the ex-post problem. One can equally view $r$ as a loss function in the DM's ex-ante problem of choosing a decision rule $\delta_n$ before $X^n$ is observed.
An important aspect of this paper is that we allow the payoff-relevant parameter $\theta$ to be set-identified. That is, the most that the DM can infer about $\theta$ is that $\theta \in \Theta_0(\mu) \subseteq \Theta$, where $\Theta_0(\cdot)$ is a known set-valued mapping from $\mathcal M$ to $\Theta$. While $X^n$ can be used to learn about $\mu$, it contains no information about the identity of the true $\theta$ within $\Theta_0(\mu)$.
\@startsection{paragraph}{4}{\z@} {0pt \@plus1ex \@minus.2ex} {-1em} {\normalfont}{Example: Treatment Assignment.}
Consider the setup in Section (ref): $d \in \{0,1\}$ is a binary treatment indicator, $\theta$ is the ATE, and $r(d,\theta,\mu) \equiv r(d,\theta)$ is given in display ((ref)). The DM has bounds $b_L(\mu)$ and $b_U(\mu)$ on $\theta$ as a function of $\mu$ and observes data or summary statistics $X^n \sim P_{n,\mu}$ that may be used to learn $\mu$. In the application from Section (ref), the statistics were the vector of ATE estimators $X^n = (\hat \mu_1,\ldots,\hat \mu_k)$, $P_{n,\mu}$ is their sampling distribution, the reduced-form parameter is the vector of ATEs $\mu = (\mu_1,\ldots,\mu_k)$, and $b_L(\mu)$ and $b_U(\mu)$ are given in ((ref)). We review other methods for constructing bounds in different empirical settings in Section (ref). The identified set is $\Theta_0(\mu) = [b_L(\mu) ,b_U(\mu)]$. $\square$
Our setup features two types of parameters: a point-identified parameter $\mu$ and a set-identified parameter $\theta$. We adopt a minimax approach to handle the ambiguity that arises from the partial identification of $\theta$ conditional on $\mu$ and average (or integrated) risk minimization to estimate $\mu$. Our asymmetric treatment of parameters is in the spirit of Hurwicz1951; we discuss the connection in more detail in Section (ref). We first consider the case in which $\mu$ is known and then extend the analysis to the case of unknown $\mu$, which will lead us to the definition of our optimality concept.
\@startsection{paragraph}{4}{\z@} {0pt \@plus1ex \@minus.2ex} {-1em} {\normalfont}{Known $\boldsymbol \mu$.} If the DM knows $\mu$, then the data are irrelevant and the germane problem is the ex-post problem in which $d$ is chosen. To handle the ambiguity about $\theta \in \Theta_0(\mu)$, the DM evaluates actions by their maximum risk
Define \[ \delta^o(\mu) = \operatorname*{argmin}_{d \in \mc D} R(d,\mu), \] if the argmin is unique, otherwise $\delta^o(\mu)$ is chosen randomly from $\operatorname*{argmin}_{d \in \mc D} R(d,\mu)$. This choice can be interpreted as the equilibrium in a two-player zero-sum game in which an adversarial nature chooses $\theta$ in response to the DM's choice of $d$ to minimize $\mathbb{E}_\theta[u(d,Y,\theta,\mu)]$. We call $\delta^o(\mu)$ the oracle decision: it represents the DM's optimal choice if $\mu$ was known. The oracle decision is infeasible in any practical application because $\mu$ will need to be estimated. Nevertheless, it serves as a useful benchmark because $R(\delta_n(X^n),\mu) \geq R(\delta^o(\mu), \mu) \equiv \min_{d \in \mc D} R(d,\mu)$ for any data-dependent decision rule $\delta_n(X^n)$.
\@startsection{paragraph}{4}{\z@} {0pt \@plus1ex \@minus.2ex} {-1em} {\normalfont}{Unknown $\boldsymbol \mu$.} Here the DM uses the data $X^n$ to learn about $\mu$ before choosing $d \in \mc D$. We treat $R(d,\mu)$ as a loss function for the DM's ex-ante problem. For a decision rule $\delta_n$, the DM incurs ex-ante risk
where $\mathbb{E}_\mu[\,\cdot\,]$ denotes expectation with respect to $X^n \sim P_{n,\mu}$. Our goal is to construct sequences $\{\delta_n\}$ of decision rules that use the data efficiently, so that criterion ((ref)) is minimized in large samples over a range of data-generating processes.
To simplify exposition, we consider parametric models in the rest of this section and defer discussion of semiparametric models to Section (ref). Following HiranoPorter2009, we work in a local asymptotic framework centered at a fixed $\mu_0 \in \mathcal M \subseteq \mathbb R^K$ and parameterized by a local parameter $h = \sqrt n(\mu - \mu_0)$ ranging over $\mathbb R^K$. As $h$ is not consistently estimable, this asymptotic framework is designed to mimic the finite-sample problem faced by the researcher, where $\mu$ is not known with certainty. We let $P_{n,h}$ denote the distribution of $X^n$ when $\mu = \mu_0 + h/\sqrt n$ and let $\mathbb{E}_{n,h}$ denote expectation under $P_{n,h}$.
We rank decision rules by their ex-ante risk in excess of the oracle. We scale excess risk by $\sqrt n$ so that the large-sample limit is not degenerate. This leads to the criterion
As there is no a priori reason to assign more weight to one local parameter than another, we integrate with respect to Lebesgue measure over $h$ to arrive at the criterion
We rank sequences of decision rules by criterion ((ref)). For now we assume $\mc R(\{\delta_n\};\mu_0)$ is finite for at least one sequence of decisions $\{\delta_n\}$, so that criterion ((ref)) provides a meaningful ranking. Section (ref) discusses extensions to settings where this criterion is infinite and the improper Lebesgue prior on $h$ is approximated by a sequence of proper priors.
Before introducing our definition of optimality, we first define a class $\mathbb D$ of sequences of decision rules for which the above criteria are well defined. Let \[ \underline{\mathcal D}_{\mu_0} = \operatorname*{argmin}_{d \in \mathcal D} R(d,\mu_0) \] denote the set of optimal choices at $\mu_0$. Note that $\underline{\cal D}_{\mu_0}$ will be a non-singleton at values of $\mu_0$ where there are multiple choices that minimize $R(d,\mu_0)$.
Definition (ref)(ref) requires that the sequence of random variables $\{\delta_n(X^n)\}$ converges in distribution along $\{P_{n,h}\}$. Definition (ref)(ref) is a minor technical condition requiring $\delta_n$ to choose among $\underline{\mc D}_{\mu_0}$ with high probability under small perturbations of $\mu$ around $\mu_0$. For the intuition, by the maximum theorem we know that optimal actions under $\mu$ near $\mu_0$ should belong to $\underline{\mc D}_{\mu_0}$. If $d_n$ chooses actions outside this set, then it makes mistakes: it chooses actions that should be easy to identify as sub-optimal. Condition (ref) requires the probability of these mistakes vanishes sufficiently fast. Bayes decisions satisfy this condition: their mistake probabilities vanish at rate $n^{-1}$ (see Lemma (ref) in Appendix (ref)).
We refer to sequences $\{\delta_n\}$ as optimal if they satisfy the following definition:
This is a non-standard problem and there are different notions of optimality that may lead to different rankings over sequences of decision rules, as we discuss in Section (ref) below. An attractive feature of our optimality criterion is that the DM only needs to be able to compute $R(d,\mu)$ in order to implement optimal decisions. This makes our approach broadly applicable including in scenarios when $R(d,\mu)$ has no closed-form expression, such as those in Sections (ref) and (ref). In the next subsections we present Bayesian and bootstrap implementations.
Integrating criterion ((ref)) with respect to a prior $\pi$ yields the integrated maximum risk criterion
After observing $X^n$, the DM can form a posterior $\pi_n(\mu) = \pi_n(\mu|X^n)$ for $\mu$. Standard arguments (e.g. Wald1950, Wald1950, Chapter 5.1) imply that ((ref)) can be minimized by minimizing the posterior maximum risk
with respect to $d \in \mc D$ for almost every realization of the data. This leads to the Bayes rule
if the argmin is unique, otherwise $\delta_n^*(X^n;\pi)$ is chosen randomly from $\operatorname*{argmin}_{d \in \mc D} \bar R_n(d)$. We include $\pi$ as an argument to indicate that the decision depends on $\pi$ in finite samples.
Bayes decisions $\delta_n^*(X^n;\pi)$ are optimal under criterion ((ref)) for any prior $\pi$ whose density is positive, bounded, and continuous (Theorem (ref)). To see the connection between ((ref)) and ((ref)), use a change-of-variables to express the prior for $h$ as $\propto \pi(\mu_0+ h/\sqrt{n})$. This prior becomes uniform as $n \to \infty$, which is the weight function underlying ((ref)).
\@startsection{paragraph}{4}{\z@} {0pt \@plus1ex \@minus.2ex} {-1em} {\normalfont}{Example: Treatment Assignment.} Recall from Section (ref) that the maximum risks of $d \in \{0,1\}$ are the positive and negative parts of the bounds on the ATE: \[
\] Averaging $R(1,\mu)$ and $R(0,\mu)$ across $\pi_n$ yields \[
\] We have $\bar R_n(1) \leq \bar R_n(0)$ if and only if $\bar b_n := \int b(\mu) \, d \pi_n(\mu) \geq 0$ with $b(\mu) := (b_U(\mu))_+ + (b_L(\mu))_-$. Hence, the Bayes rule is $\delta_n^*(X^n;\pi) = 12 I\left[ \bar b_n \geq 0\right]$.\footnote{This rule deterministically chooses treatment when there is a tie (i.e., when $\bar b_n = 0$). Any (possibly randomized) tie-breaking rule leads to decision which is optimal under our optimality criteria.} We discuss implementation in a number of different empirical settings in Section (ref). $\square$
Let $\hat \mu$ denote an efficient estimator of $\mu$ and let $\mathbb{E}_n^*$ denote expectation with respect to the bootstrap version $\hat \mu^*$ of $\hat \mu$ conditional on $X^n$. Define the bootstrap average maximum risk \[ R_n^*(d) = \mathbb{E}_n^* \left[ R(d,\hat \mu^*) \right] . \] The bootstrap rule is \[ \delta_n^{**}(X^n) = \operatorname*{argmin}_{d \in \mc D} R_n^*(d) \] if the argmin is unique (with a random selection from the argmin otherwise). We show in Theorem (ref) below that as long as $\delta_n^{**}$ behaves asymptotically like $\delta_n^*$ it inherits the optimality property of the Bayes decision.
\@startsection{paragraph}{4}{\z@} {0pt \@plus1ex \@minus.2ex} {-1em} {\normalfont}{Example: Treatment assignment.} Here we have \[
\] Therefore, the bootstrap rule is $\delta_n^{**}(X^n) = 12 I \left[ \mathbb{E}_n^* \left[ b(\hat \mu^*) \right] \geq 0\right]$. $\square$
The first distinguishing feature of our analysis is that we separate two groups of parameters, $\theta$ and $\mu$, and apply the minimax reasoning only to $\theta$ because only it is partially identified. The approach of treating groups of parameters differently dates back to Hurwicz1951. He argued that one might consider multiple priors for some parameters and referred to it as a generalized Bayes-minimax principle. He provided a two-parameter example in which the marginal prior for one of the parameters, $\mu$ in our notation, is fixed, whereas a family of priors is considered for the conditional distribution of the second parameter $\theta$ given $\mu$.
Recall that $\delta_n^*(\,\cdot\,;\pi)$ minimizes criterion ((ref)) over all decision rules $\delta_n : \mc X^n \to \mc D$. By the definition of $R(d,\mu)$ and the discussion in Section (ref), we see that $\delta_n^*(x^n;\pi)$ solves
for almost every realization $x^n$ of $X^n$, where $r(d,\theta,\mu) = \mathbb{E}_\theta[-u(d,Y,\theta,\mu)]$. Unlike the usual minimax framework, we allow the adversary (“nature”) to choose $\theta$ conditional on $X^n$. This is the second distinguishing feature of our analysis. It creates a more adversarial setting than usual, but it is disciplined by restricting nature's choice to the identified set $\Theta_0(\mu)$ rather than the whole parameter space $\Theta$.
In this regard, our approach is closely related to the derivation of conditional $\Gamma$-minimax decision rules in the Bayesian literature, e.g., DasGuptaStudden1989, BetroRuggeri1992, and GiacominiKitagawaRead2021. One difference is that we keep the marginal prior distribution of $\mu$ fixed, which alleviates concerns about the conservativeness of the approach. Suppose one combines the unique prior $\pi$ for $\mu$ with a family $\Lambda$ of conditional priors $\lambda := \{\lambda(\cdot|\mu): \mu \in \mc M\}$ for $\theta$, where $\lambda(\theta|\mu)$ has support contained in $\Theta_0(\mu)$ for each $\mu \in \mc M$. Then one can define $\Gamma = \{ \pi \otimes \lambda : \lambda \in \Lambda\}$, where each $\gamma \in \Gamma$ is a prior over $(\theta,\mu)$.\footnote{We use the notation $\pi \otimes \lambda$ to denote the conditional probability measures $\lambda = \{\lambda(\cdot|\mu): \mu \in \mc M\}$ integrated against a marginal probability measure $\pi$ for $\mu$. Thus, for $\gamma = \pi \otimes \lambda$ and $g : \Theta \times \mc M \to 12 R$, we define $\int_{\Theta \times \mc M} g(\theta,\mu) d \gamma(\theta,\mu) = \int_{\mc M} \int_\Theta g(\theta,\mu) \, d \lambda(\theta|\mu) \, d \pi(\mu)$.} The conditional $\Gamma$-minimax decision solves the following min-max problem for almost every realization $x^n$ of $X^n$:
where $\gamma_n(\theta,\mu|x^n)$ is the posterior for $(\theta,\mu)$ given $X^n = x^n$ for the prior $\gamma \in \Gamma$. This is a robust Bayes criterion for the DM's ex-post problem faced after observing $X^n$. Accordingly, $X^n$ is treated as fixed: the DM and nature know $X^n$ when choosing their actions.
In responding to the DM's choice of $d$ in ((ref)), nature has to choose a set of conditional priors $\{\lambda(\theta|\mu) : \mu \in {\cal M}\}$ where each $\lambda(\cdot|\mu)$ has support contained in $\Theta_0(\mu)$. For each $\mu \in {\cal M}$, define $\theta_*(\mu;d) \in \mathrm{arg}\sup_{\theta \in \Theta_0({\mu})} r(d,\theta,\mu)$, assuming that the arg sup is non-empty, and let $\lambda_*(\theta|\mu,d)$ be a point mass at $\theta_*({\mu};d)$. In combination, $\pi$ and $\{\lambda_*(\theta|\mu,d) : \mu \in \mc M\}$ define a joint prior $\gamma_* = \pi \otimes \lambda_*$ over $(\theta,\mu)$. Then, by construction, the posterior risk of $d$ under the $\gamma_*$ prior equals the posterior maximum risk \[ \int_{\Theta \times \mc M} r(d,\theta,\mu) \, d \gamma_{*n}(\theta,\mu|x^n) \label{eq:risk.lambdastar} = \int_{ \mc M} \bigg( \sup_{\theta \in \Theta_0(\mu)} r(d,\theta,\mu) \bigg) \, d \pi_n(\mu|x^n). \] Because the supremum over $\theta \in \Theta_0(\mu)$ is weakly larger than the average under any $\lambda(\theta|\mu)$, we can deduce that, provided $\gamma_* \in \Gamma$,
Combining ((ref)) and ((ref)), we see that the decision rules that are considered optimal under Definition (ref) can be viewed as solutions to a conditional $\Gamma$-minimax problem.
As is well known, criterion ((ref)) can lead to different optimal decision than the worst-case Bayes risk (or unconditional $\Gamma$-minimax) criterion for the DM's ex-ante problem faced before $X^n$ is realized:
where the infimum is over all measurable $\delta_n : \mc X^n \to \mc D$; see GiacominiKitagawaRead2021 and AFMQT for discussions. Thus, decision rules that are optimal under criterion ((ref)) may not align with optimal actions in the conditional $\Gamma$-minimax problem ((ref)). Further, optimal decisions under criterion ((ref)) are, in general, intractable (see, e.g., GiacominiKitagawaRead2021).
Our proposal---namely, to use criterion ((ref)) for the ex-ante problem---circumvents both of these problems. Optimal decisions under criterion ((ref)) are tractable, align with optimal actions under the conditional $\Gamma$-minimax criterion ((ref)), and are asymptotically efficient in a frequentist sense that we formalize in Section (ref).
Finally, we note our approach departs from the textbook Wald approach in two regards. The Wald approach would seek a minimax decision rule that solves \[ \inf_{\delta_n} \sup_{\mu \in \mc M} \sup_{\theta \in \Theta_0(\mu)} \mathbb{E}_\mu\left[ r(\delta_n(X^n), \theta, \mu) \right]. \] Comparing with criterion ((ref)), we see that our approach moves the supremum over $\theta$ inside the expectation, potentially making the criterion more conservative. But as we discussed above, this conservativeness is constrained by the identified set $\Theta_0(\mu)$. Moreover, this conservativeness is offset, at least in part, by the fact that we average across a prior $\pi$ for $\mu$ rather than taking a supremum over $\mu$. These departures seem reasonable given the asymmetric nature of the identification problem. They also buy tractability, allowing the derivation of optimal decisions in a much broader class of problems than the standard Wald approach.
We now present the main optimality results for parametric models. We first outline the assumptions in Section (ref). Section (ref) presents two main results. Theorem (ref) establishes that Bayes decision rules $\delta_n^*(\,\cdot\,;\pi)$, defined in ((ref)), are optimal. Theorem (ref) shows that any decision whose asymptotic behavior is different from $\delta_n^*(\,\cdot\,;\pi)$ is sub-optimal. In particular, we prove in Section (ref) that plug-in decisions $\delta_n^{plug}$ are sub-optimal when the maximum risk $R(d,\mu)$ is a non-smooth function of $\mu$. As we documented in the empirical illustration in Section (ref), this finding can have important implications for treatment assignment under partial identification. Finally, in Section (ref) we provide a refinement of our analysis for settings in which the average maximum risk at the optimal decision not finite.
We first place some assumptions on the risk functions. Say $f : \mathcal M \to \mathbb R^k$ is directionally differentiable at $\mu_0$ if there is a continuous map $ \dot f_{\mu_0}[\,\cdot\,] : \mathbb R^K \to \mathbb R^k$ such that \[ \lim_{n \to \infty} \frac{f(\mu_0 + t_n h_n) - f(\mu_0)}{t_n} = \dot f_{\mu_0}[h] \] for all sequences $\{t_n\} \subset \mathbb R_+$ and $\{h_n\} \subset \mathbb R^K$ with $t_n \downarrow 0$ and $h_n \to h \in \mathbb R^K$. If so, we refer to $ \dot f_{\mu_0}[\,\cdot\,]$ as the directional derivative of $f$ at $\mu_0$. Note that $ \dot f_{\mu_0}[\,\cdot\,]$ is positively homogeneous of degree one but not necessarily linear. If $\dot f_{\mu_0}[\,\cdot\,]$ is linear, then $f$ is (fully) differentiable at $\mu_0$.
We note that $R(d,\,\cdot\,)$ may be directionally but not fully differentiable in many empirically relevant cases of treatment assignment under partial identification (see Sections (ref) and (ref)). We only require directional differentiability to hold when $|\underline{\mathcal D}_{\mu_0}| > 1$. In this case, optimal actions can differ for different $\mu$ close to $\mu_0$ and the problem of distinguishing between optimal actions remains nontrivial as the sample size increases. In these scenarios, we use directional differentiability to derive asymptotic approximations for different decision rules.
We also place some regularity conditions on the statistical model $\mathcal P = \{P_\mu : \mu \in \mathcal M\}$. We present assumptions for a random sample to simplify exposition --- so $X^n = (X_1,\ldots,X_n)$ where each $X_i$ is an independent draw from $P_\mu$ --- though it is straightforward to extend our theory to weakly dependent data. We assume implicitly that each $P_\mu$ admits a density $p_\mu$ with respect to a common dominating measure $\nu$. Say that $\mathcal P$ is differentiable in quadratic mean (DQM) at $\mu$ if there exists a vector of measurable functions $\dot \ell_\mu$ such that \[ \int \left( \sqrt{p_{\mu + h}} - \sqrt{p_{\mu}} - \frac 12 h^T \dot \ell_\mu \sqrt{p_\mu} \right)^2 d\nu = o \left( \|h\|^2 \right) \] as $h \to 0$. Under DQM, the Fisher information matrix is $I_\mu := \int \dot \ell_\mu \dot \ell_\mu^T \, d P_\mu$. Let $D_{KL}(p_{\mu}\|p_{\mu'}) = \int p_\mu \log(p_{\mu}/p_{\mu'}) d\nu$ if $P_\mu(p_{\mu'}(X) = 0) = 0$ and $D_{KL}(p_\mu\|p_{\mu'}) = +\infty$ if $P_\mu(p_{\mu'}(X) = 0) > 0$. Say that $\mathcal P$ is locally quadratic for $D_{KL}$ if for all $\mu_0 \in \mathcal M$, \[ D_{KL}(p_\mu\|p_{\mu'}) \leq 2(\mu - \mu')^T I_{\mu_0}(\mu - \mu') \] holds for all $\mu,\mu'$ in a neighborhood of $\mu_0$. Following ClarkeBarron, we say that $\mathcal P$ is sound if weak convergence of $P_\mu$ to $P_{\mu'}$ is equivalent to convergence of $\mu$ to $\mu'$.
Assumptions (ref)(ref)-(ref) are similar to the conditions for parametric models in HiranoPorter2009. DQM and local quadraticity hold under standard smoothness and integrability conditions. These conditions rule out certain “irregular” models, such as those with parameter-dependent support and parameters on the boundary. Assumption (ref)(ref) holds under conditions similar to those used to establish asymptotic normality of maximum likelihood estimators. Soundness is a minimal identifiability condition. Following Schwartz1965, posterior consistency is typically established in the weak topology. Soundness converts this to consistency for the parameter $\mu$. This condition trivially holds: (i) in exponential families ClarkeBarron; or (ii) if $\mu$ is identified (i.e., $P_\mu = P_{\mu'}$ if and only if $\mu = \mu'$), $\mathcal M$ is precompact, and $\mu \mapsto P_\mu$ is weakly continuous on the closure of $\mathcal M$.\footnote{Alternatively, one could assume the existence of uniformly consistent tests of $H_0: \mu = \mu_0$ against $H_1: \|\mu - \mu_0\| \geq \epsilon$ Schwartz1965,vanderVaart1998. Soundness is a weak sufficient condition for the existence of such tests ClarkeBarron. It also allows us to control the error rate of tests along $\{P_{n,h}\}$, for which the usual classical testing condition seemed inadequate.}
\@startsection{paragraph}{4}{\z@} {0pt \@plus1ex \@minus.2ex} {-1em} {\normalfont}{Example: Treatment Assignment.}
Consider Assumption 1. In the empirical application, the lower bound for male youths behaved approximately like the maximum of two independent Gaussian random variables (corresponding to estimates from US and Brazilian RCTs) while the upper bound behaved approximately like a third independent Gaussian random variable (corresponding to the estimate from a Colombian RCT). We can mimic this setting with the following stylized example. Let
As the bounds from the three studies are roughly the same magnitude, let $\mu_0 = (-c,-c,c)$ for $c > 0$. Then $R(0,\mu) = (\mu_3)_+$ and $R(1,\mu) = -(\mu_1 \vee \mu_2)_-$, and so $R(0,\mu_0) = R(1,\mu_0) = c$ and $\underline{\mc D}_{\mu_0} = \{0,1\}$. We deduce that $R(d,\mu)$ is bounded (provided ${\cal M} \subset \mathbb R^3$ is bounded) and continuous in $\mu$. Here $R(0,\mu)$ is smooth at $\mu_0$ with $\dot R_{0,\mu_0}[h] = h_3$. However, $R(1,\mu)$ is only directionally differentiable at $\mu_0$ with $\dot R_{1,\mu_0}[h] = -(h_1 \vee h_2)$. Overall, $R(d,\mu)$ satisfies Assumption (ref).
Assumption (ref) places restrictions on the family of probability distributions $\{P_\mu \, : \, \mu \in {\cal M}\}$. One formalization of the application described in Section (ref) is to regard the published estimates, rather than the observations upon which these estimates are based, as data. So $X^n = \hat{\mu}$. Using the Normality assumption underlying our empirical illustration, $P_\mu$ is a multivariate Normal distribution with mean $\mu$ and a diagonal covariance matrix with elements $s^2_k$. The Normal family clearly satisfies Assumption (ref). In Section (ref) we will relax the strict Normality assumption and cast the paper in a semiparametric framework. $\square$
We first introduce some terminology. Let $\Pi$ denote the class of all priors on $\mc M$ with positive, continuous, and bounded Lebesgue density. We say that $\{\delta_n\},\{\delta_n'\} \in 12 D$ are asymptotically equivalent if \[ \lim_{n \to \infty} P_{n,h}(\delta_n(X^n) = d) = \lim_{n \to \infty} P_{n,h}(\delta_n'(X^n) = d) \] for all $d \in \mc D$, $h \in 12 R^K$, and $\mu_0 \in \mc M$. Proposition (ref) in Appendix (ref) shows that $\sqrt n (\bar R_n(d) - R(d,\mu_0))$ converges in distribution along $P_{n,h}$ to $\mathbb{E}^*[\dot R_{d,\mu_0}[Z^* + Z]|Z ]$, where $\dot R_{d,\mu_0}$ is the directional derivative of $R(d,\mu)$ at $\mu_0$, $Z \sim N(h,I_{\mu_0}^{-1})$ represents the asymptotic behavior of an efficient estimator of $\mu$, $Z^* \sim N(0,I_{\mu_0}^{-1})$ independently of $Z$ represents asymptotic uncertainty about the local parameter $h$, and $\mathbb{E}^*$ denotes expectation with respect to $Z^*$, which integrates over the uncertainty in the local parameter. We say there are no first-order ties if for each $\mu_0 \in \mc M$ the minimizer of $d \mapsto \mathbb{E}^*[\dot R_{d,\mu_0}[Z^* + z] ]$ over $\underline{\mc D}_{\mu_0}$ is unique for almost every $z$. This condition trivially holds if $\underline{\mc D}_{\mu_0}$ is a singleton. To interpret the condition when $|\underline{\mc D}_{\mu_0}| > 1$, suppose that $R$ is differentiable at $\mu_0$, which is analogous to the case of smooth welfare contrasts studied in HiranoPorter2009. Then $\mathbb{E}^*[\dot R_{d,\mu_0}[Z^* + z] ] = g_{d,\mu_0}^T z$ with $g_{d,\mu_0} = \frac{\partial R(d,\mu_0)}{\partial \mu}$. In that case, a sufficient condition for no first-order ties is that the derivatives are different: $g_{d,\mu_0} \neq g_{d',\mu_0}$ for $d,d' \in \underline{\mc D}_{\mu_0}$.
Our first main result establishes optimality of Bayes decisions $\{\delta_n^*(\,\cdot\,;\pi)\}$ for $\pi \in \Pi$.
According to Theorem (ref), Bayes decisions with priors $\pi \in \Pi$ are asymptotically equivalent and optimal. All such Bayes decisions are asymptotically independent of $\pi$. Part (ref) implies that any decision that is asymptotically equivalent to a Bayes decision under a prior $\pi \in \Pi$ is optimal. Optimality of the bootstrap implementation discussed in Section (ref) follows from (ref) under suitable regularity conditions.
We now show that asymptotic equivalence to a Bayes decision is necessary for optimality. Say $\{\delta_n\},\{\delta_n'\} \in 12 D$ fail to be asymptotically equivalent at $\mu_0$ if \[ \lim_{n \to \infty} P_{n,h}(\delta_n(X^n) = d) \neq \lim_{n \to \infty} P_{n,h}(\delta_n'(X^n) = d) \] for some $h \in \mathbb R^K$ and some $d \in {\cal D}$.
In the next subsection we discuss further implications of Theorem (ref) with respect to decision rules based on plugging-in efficient estimators of $\mu$.
A natural alternative to $\delta_n^*(X^n; \pi)$ is to plug in an efficient estimator $\hat \mu = \hat \mu(X^n)$ of $\mu$ into the oracle decision rule, yielding the “plug-in” rule $\delta^{plug}_n(X^n) = \delta^o(\hat \mu)$. Manski2021Haavelmo,Manski2021 refers to this approach as “as-if” optimization: the estimator $\hat \mu$ is treated “as if” it is the true parameter.\footnote{This approach also has connections to anticipated utility Kreps1998,CogleySargent2008.} We show that asymptotic equivalence of the Bayes and plug-in rules can fail when the maximum risk does not depend smoothly on $\mu$; if so, Theorem (ref) implies plug-in rules are sub-optimal. We illustrate this difference within the context of the application to treatment assignment from Section (ref).
To understand when asymptotic equivalence holds, we turn to Lemma (ref) in Appendix (ref), which characterizes the asymptotic behavior of sequences of decision rules. Lemma (ref) implies \[ \lim_{n \to \infty} P_{n,h} \left( \delta_n^*(X^n;\pi) = d \right) =
\] where $\mathbb P_h$ is the probability measure of $Z \sim N(h,I_{\mu_0}^{-1})$. One can similarly show
Asymptotic equivalence of the Bayes and plug-in rules holds when the right-hand side probabilities in the above two displays are equal.
Suppose $R(d,\mu)$ is fully differentiable at $\mu_0$ for all $d \in \underline{\mc D}_{\mu_0}$. Then each $\dot R_{d,\mu_0}[\,\cdot\,]$ is linear, and \[ \mathbb E^* \left[ \left. \dot R_{d,\mu_0}[Z^* + Z] \right| Z \right] = \dot R_{d,\mu_0}[ \mathbb E^* \left[Z^* \right]+ Z] \equiv \dot R_{d,\mu_0}[Z] . \] Hence, the plug-in and Bayes rules are asymptotically equivalent, and therefore optimal by Theorem (ref). Formally:
Corollary (ref) is consistent with Theorem 3.2 of HiranoPorter2009, which shows plug-in rules are optimal when welfare contrasts depend smoothly on a point-identified, regularly estimable parameter.
If $R(d,\mu)$ is only directionally differentiable at $\mu_0$, then $\dot R_{d,\mu_0}[\,\cdot\,]$ is not linear and the above reasoning no longer applies. In this case, asymptotic equivalence of the Bayes and plug-in rules cannot be guaranteed. If it fails, Theorem (ref) implies the plug-in rule is sub-optimal.
\@startsection{paragraph}{4}{\z@} {0pt \@plus1ex \@minus.2ex} {-1em} {\normalfont}{Example: Treatment Assignment.}
We continue the calculations for the stylized example introduced in Section (ref). By Jensen's inequality, \[
\] Hence, asymptotic equivalence fails and the plug-in rule is sub-optimal: the Bayes rule recommends treatment more aggressively than the plug-in rule. This is reflected in the empirical application, where the optimal rule recommends treatment but the plug-in rule does not. $\square$
Criterion $\mc R(\{\delta_n\};\mu_0)$ in ((ref)) is formed by integrating $\mc R(\{\delta_n\};\mu_0, h)$ with respect to Lebesgue measure on $\mathbb R^K$. This raises the possibility that $\mc R(\{\delta_n\};\mu_0) = +\infty$ for all $\{\delta_n\} \in 12 D$, in which case it does not produce a meaningful ranking. We first note this is not possible if $K = 1$:
If $K > 1$, however, we may have $\mc R(\{\delta_n\};\mu_0) = +\infty$ for all $\{\delta_n\} \in 12 D$. For instance, one could take the $K = 1$ case and augment the parameter space with redundant parameters. To handle cases where criterion ((ref)) is infinite, we approximate the improper Lebesgue prior on $h$ by a sequence of proper priors. We used Lebesgue measure on $\mathbb R^K$ in criterion ((ref)) because there is no a priori reason to view one local parameter as more likely than another. A similar effect is achieved for a $N(0,\sigma I)$ prior with large $\sigma$. However, for any finite $\sigma$ the integrated risk under this proper prior is also finite because $R(d,\cdot)$ is bounded according to Assumption (ref)(i). Let $\pi_\sigma$ denote the $N(0,\sigma I)$ density. Define
Lemma (ref) implies $\mc R_\sigma(\{\delta_n\};\mu_0)$ is finite for all $\{\delta_n\} \in 12 D$. Analogously to Definition (ref), we define the concept of $\sigma$-optimality:
As we are using the large-$\sigma$ limit of the $N(0,\sigma I)$ prior to approximate Lebesgue measure, we are really interested in the behavior of $\sigma$-optimal decisions as $\sigma \to \infty$. The following result shows Bayes and $\sigma$-optimal decisions lead to the same choices in large samples.
The proof of Theorem (ref) shows that Bayes and $\sigma$-optimal decisions are identical in large samples, for almost every realization of the data, provided $\sigma$ is sufficiently large. We omit the details here to avoid introducing additional technicalities, and instead illustrate the idea within the context of our running example.\footnote{Here is some intuition: in many Bayesian settings posterior distributions of parameters are proper even under improper priors. But integrated risk is typically only finite under a proper prior. As long as the likelihood function asymptotically dominates the prior, posteriors, and decisions derived from them, under an improper prior are very close to those obtained from a prior with a large variance.}
\@startsection{paragraph}{4}{\z@} {0pt \@plus1ex \@minus.2ex} {-1em} {\normalfont}{Example: Treatment Assignment.} We continue with our discussion of the stylized example from ((ref)) that mimics the empirical application. We have $\sqrt n(\hat \mu - \mu) \to_d Z \sim N(h,\Sigma)$ along $P_{n,h}$, where $\Sigma = \mathrm{diag}(s_1,s_2,s_3)$, say, since the RCTs are independent. Lemma (ref) in Appendix (ref) implies that asymptotically, any Bayes decision behaves like \[ \delta^*_\infty(z) = 12 I[ \mathbb{E}^*[(Z_1^* + z_1) \vee (Z_2^* + z_2)] + z_3 \geq 0 ], \] for almost every realization $z$ of $Z$, where $\mathbb{E}^*$ denotes expectation taken with respect to $Z^* \sim N(0,\Sigma)$. The expectation can be computed in closed-form (see, e.g., NKotz) to give
and $\Phi$ and $\phi$ are the standard normal CDF and PDF, respectively. Similarly, Lemma (ref) in Appendix (ref) implies that asymptotically, the $\sigma$-optimal decision behaves like
The difference between the two arises because in the $\sigma$-optimal decision the posterior mean is shrunk from $z$ to $(I + \sigma^{-1}\Sigma)^{-1}z$ and the posterior variance is shrunk from $\Sigma$ to $(\Sigma^{-1} + \sigma^{-1} I)^{-1}$. There are no first-order ties because the set of $z$ values for which $f(z;s_1,s_2) = 0$ has measure zero. For any $z$ outside this negligible set, we have $\mathrm{sign}(f(z;s_1,s_2)) = \mathrm{sign}(f_\sigma(z;s_1,s_2))$ for $\sigma$ sufficiently large. Hence, for almost every realization $z$ of $Z$, there is a minimal value of $\sigma$ above which the Bayes and $\sigma$-optimal decisions are identical. $\square$
This section extends our approach to semiparametric models, which is relevant in many empirical contexts. For example, the data may not follow a specific parametric model, or $X^n$ may be a vector of summary statistics with unknown finite-sample distribution, as in the empirical application in Section (ref). We present the model in Section (ref) then generalize our optimality concept in Section (ref). Section (ref) describes a quasi-Bayesian implementation of optimal decisions. Section (ref) presents the main optimality results. Theorem (ref) shows that quasi-Bayes decision rules are optimal, while Theorem (ref) shows that any decision whose asymptotic behavior differs from quasi-Bayes rules is sub-optimal.
Let $X^n \sim P_{n,(\mu,\eta)}$ with $\mu \in \mc M \subseteq 12 R^K$ and $\eta \in \mc H$, an infinite-dimensional space. In a GMM model, $\mu$ is a finite-dimensional parameter vector, $\eta$ is the marginal distribution of each observation $X_i \in \mc X$, and $\mc H = \{\eta \in 12 M(\mc X) : \int g(x,\mu) \, d \eta(x) = 0$ for some $\mu \in \mc M\}$ for a vector of moment functions $g$, where $12 M(\mc X)$ is the set of all probability measures on $\mc X$. We again assume that, given $\mu$, the structural parameter $\theta$ takes values in a set $\Theta_0(\mu)$. Therefore, the set of payoff distributions is indexed only by the parametric component $\mu$ and the nonparametric component $\eta$ is a nuisance parameter. For instance, $\mu$ may be a vector of population moments used to construct bounds on $\theta$. The nuisance parameter $\eta$ represents other features of the distribution of $X^n$ that are irrelevant for the DM's decision problem.
Our optimality criterion ((ref)) integrates the excess maximum risk of $\delta_n(X^n)$ relative to the oracle using Lebesgue measure on the local perturbations $h \in 12 R^K$ of $\mu_0$. This approach does not extend to perturbations of $(\mu_0,\eta_0)$ due to measure-theoretic complications in infinite-dimensional spaces. We therefore form our optimality criterion using local perturbations of $\mu_0$ within a least favorable submodel in which $X^n$ carries the least information about $\mu$ of all parametric submodels. The problem of parameter estimation in the least favorable submodel is asymptotically equivalent to the problem of estimating $\mu$ in the full semiparametric model.
To simplify exposition, we present the following discussion and results within the context of a random sample --- so $X^n = (X_1,\ldots,X_n)$ where each $X_i$ is an independent draw from $P_{(\mu,\eta)}$ --- though it is straightforward to extend our theory to weakly dependent data. We say that $\mc P = \{P_{(\mu,\eta)} : \mu \in \mc M, \eta \in \mc H\}$ has a least favorable submodel at $(\mu,\eta)$ if there exists an open neighborhood $\mc M_{(\mu,\eta)}$ of $\mu$ and a map $t \mapsto \eta_t$ from $\mc M_{(\mu,\eta)}$ into $\mc H$ such that $\{P_{\beta(t)} : t \in \mc M_{(\mu,\eta)}\}$ with $\beta(t) =(t,\eta_t)$ have densities $p_{\beta(t)}$ with respect to a common dominating measure $\nu$ that satisfy the DQM condition \[ \int \left( \sqrt{p_{\beta(\mu + h)}} - \sqrt{p_{\beta(\mu)}} - \frac 12 h^T \dot \ell_{(\mu,\eta)}\sqrt{p_{\beta(\mu + h)}} \right)^2 d \nu = o(\|h\|^2) \] as $h \to 0$, where $\dot \ell_{(\mu,\eta)} : \mc X \to 12 R^K$ is the efficient score for $\mu$ and $I_{(\mu,\eta)} := \int \dot \ell_{(\mu,\eta)}\dot \ell_{(\mu,\eta)}^T dP_{(\mu,\eta)}$ is the semiparametric information bound. In other words, there is a regular parametric submodel $\{P_{\beta(t)}: t \in \mc M_{(\mu,\eta)}\}$ whose information matrix at $t = \mu$ is the semiparametric information bound at $(\mu,\eta)$. The path $t \mapsto \eta_t$ and dominating measure $\nu$ can depend on $(\mu,\eta)$, but we suppress this to simplify notation. We note that our approach only requires the least favorable model to exist: the researcher doesn't need to derive it in order to implement optimal decisions. The least favorable model also needn't be unique, but any such model satisfying the above conditions will lead to the same optimality criterion.
Our optimality criterion is analogous to the parametric case. For each $(\mu_0,\eta_0) \in \mc M \times \mc H$, we reparametrize the least favorable submodel $\{P_{\beta(t)} : t \in \mc M_{(\mu_0,\eta_0)}\}$ using $t = \mu_0 + h/\sqrt n$ for $h \in \mathbb R^K$. Let $P_{n,h}$ denote the distribution of $X^n$ under $P_{\beta(\mu_0 + h/\sqrt n)}$ with $\beta(\mu_0 + h/\sqrt n) = (\mu_0 + h/\sqrt n,\eta_{\mu_0 + h/\sqrt n})$ and $\mathbb{E}_{n,h}$ denote expectation under $P_{n,h}$. Define
and let
The class $\mathbb D$ is defined analogously to the parametric case (cf. Definition (ref)). Recall that $\underline{\mathcal D}_{\mu_0} = \operatorname*{argmin}_{d \in \mathcal D} R(d,\mu_0)$ denotes the set of optimal choices at $\mu_0$.
Finally, we say a sequence of decision rules $\{\delta_n\}$ is optimal if it minimizes the criterion $\mc R(\,\cdot\,;(\mu_0,\eta_0))$ in ((ref)) over $12 D$ for all $(\mu_0,\eta_0) \in \mc M \times \mc H$ (cf. Definition (ref)):
If $\mc R(\{\delta_n\};(\mu_0,\eta_0))$ is infinite for all $\{\delta_n\} \in \mathbb D$, then a similar approach to Section (ref) can be followed whereby the Lebesgue prior on $h$ is approximated by a sequence of proper priors.
Optimal decisions are formed similarly to the parametric case, but we replace the posterior with a quasi-posterior formed from a limited-information criterion for $\mu$. Following DoksumLo1990, Kim2002, and Mueller2013, consider a limited information $N(\hat \mu, (n \hat I)^{-1})$ quasi-likelihood for $\mu$, where $\hat \mu$ is an efficient estimator of $\mu$ and $\hat I^{-1}$ is a consistent estimator of its asymptotic variance, namely $I_{(\mu,\eta)}^{-1}$. Combining the quasi-likelihood with a prior $\pi$ on $\mc M$ yields the quasi-posterior
The quasi-posterior maximum risk $\bar R_n(d)$ is calculated by averaging $R(d,\mu)$ across the quasi-posterior, as in ((ref)). The quasi-Bayes decision $\delta_n^*(X^n;\pi)$ is chosen to minimize the quasi-posterior maximum risk $\bar R_n(d)$, as in ((ref)).
Unlike the parametric case, here the optimal decision cannot be justified on the basis of robust Bayes analysis. A formal Bayesian approach would require specifying a prior on $\mc M \times \mc H$ then forming a marginal posterior for $\mu$.\footnote{See, for instance, the Bayesian exponentially tilted empirical likelihood approach of Schennach2005 or the Bayesian GMM approaches of Shin2015 and Walker2025. Our approach is computationally simple and avoids the delicate issue of specifying priors in infinite-dimensional parameter spaces.}
\@startsection{paragraph}{4}{\z@} {0pt \@plus1ex \@minus.2ex} {-1em} {\normalfont}{Example: Treatment Assignment.} In Section (ref) we rationalized the empirical illustration in Section (ref) through the assumption that the vector $\hat{\mu}$ is exactly Normally distributed. The semiparametric extension allows us to relax this assumption. Treating the estimates $\hat{\mu}$ as data, the semiparametric model allows for more general distributions of the form $\hat{\mu} \sim P_{n,(\mu,\eta)}$, satisfying the moment restriction $\int (\hat{\mu} - \mu) d P_{n,(\mu,\eta)} = 0$. The density in ((ref)) is identical to the one implied by the Normal model in Section (ref), but the interpretation changes from an exact posterior to a quasi-posterior. $\square$
We first state and discuss regularity conditions, then present our optimality results. Let $\hat \lambda_{\min}$ and $\hat \lambda_{\max}$ denote the smallest and largest eigenvalues of $\hat I$. Let $\overset{P_{n,(\mu_0,\eta_0)}}{\to}$ denote convergence in probability under $\{P_{n,(\mu_0,\eta_0)}\}$. Let $\overset{P_{n,h}}{\rightsquigarrow}$ denote convergence in distribution under $\{P_{n,h}\}$.
Assumptions (ref)(ref)-(ref) are analogous to Assumptions (ref)(ref)-(ref). We do not require versions of Assumptions (ref)(ref) and (ref)(ref) because the quasi-likelihood here has a particular quadratic structure. But we need to ensure that the estimators $\hat \mu$ and $\hat I$ used in the quasi-likelihood are sufficiently well behaved, which is the role of Assumptions (ref)(ref) and (ref)(ref). These may be verified for a wide variety of estimators $\hat \mu$ and $\hat I$ under suitable regularity conditions that cover the RCT estimators $\hat{\mu}$ of the empirical illustration in Section (ref).
\@startsection{paragraph}{4}{\z@} {0pt \@plus1ex \@minus.2ex} {-1em} {\normalfont}{Example: Treatment Assignment.} Each subpopulation corresponds to an independent randomized experiment. In each subpopulation $k$ we observe a random sample of $(X_i,Y_i)$ of size $n$ drawn from a distribution $\eta_k$,\footnote{It is straightforward to extend the analysis to have different sample sizes $n_k$ in each subpopulation, provided the $n_k$ all grow at the same rate.} where $X_i$ is a binary treatment indicator and $Y_i = X_i Y_{i1} + (1-X_i) Y_{i0}$ with $Y_{i1}$ and $Y_{i0}$ denoting the treated and untreated outcomes for individual $i$. Recall $\mu = (\mu_{k})_{k=1}^K$ and $\hat \mu = (\hat \mu_k)_{k=1}^K$, where we estimate the subpopulation-$k$ ATE $\mu_k$ by, for example, using the slope coefficient $\hat \mu_k$ in an OLS regression of $Y_i$ on $X_i$. Let $s_k$ denote the heteroskedasticity-robust standard error of $\hat \mu_k$. Then $\hat I$ is given by $(n\hat I)^{-1} = \mathrm{diag}(s_1^2,\ldots,s_K^2)$. The nuisance parameter is the joint distribution $\eta = \times_{k=1}^K \eta_k$ across the independent subpopulations. We have $P_{(\mu,\eta)} \equiv P_\eta = \eta$. We similarly drop dependence of $\mu$ in quantities such as the score and information matrix that follow.
Each $\hat \mu_k$ is semiparametrically efficient Hahn1998 with influence function \[ \psi_{k}(x,y) = \frac{x}{p_k}(y - \bar \mu_{1,k}) - \frac{1-x}{1-p_k}(y - \bar \mu_{0,k}), \] where $p_k$, $\bar \mu_{1,k}$, and $\bar \mu_{0,k}$ represent the means of $X_i$, $Y_{i1}$, and $Y_{i0}$ in subpopulation $k$ and are available from the moments of $X$, $XY$ and $(1-X)Y$ under $\eta_k$. We construct a least favorable submodel as follows. For each subpopulation $k$, fix any $\eta_{0,k}$ with $0 < p_k < 1$ and $0 < \int \psi_{k}^2 d \eta_{0,k} < \infty$. Let \[ \dot \ell_{k,\eta_0}(x,y) = \frac{\psi_k(x,y)}{\int \psi_k^2 d \eta_{0,k}}. \] Without confusion we also let $\eta_{0,k}$ denote the density of $\eta_{0,k}$ with respect to some dominating measure $\nu$. We define a smooth parametric model $\eta_{t_k,k}$ passing through $\eta_{0,k}$ at $t_k = \mu_{0,k}$ with score $\dot \ell_{k,\eta_0}$ by \[ \eta_{t_k,k}(x,y) = \eta_{0,k}(x,y) \frac{v((t_k - \mu_{0,k})\dot \ell_{k,(\mu_0,\eta_0)}(x,y))}{\int v((t_k - \mu_{0,k})\dot \ell_{k,(\mu_0,\eta_0)}(x,y)) d \eta_{0,k}(x,y)}, \] where $v(u) = 2(1+e^{-2u})^{-1}$ satisfies $v(0) = v'(0) = 1$ vanderVaart1998. Letting $\eta_t = \times_{k=1}^K \eta_{t_k,k}$ for $t = (t_k)_{k=1}^K$, we have a smooth parametric family passing through $\eta_0$ at $t = \mu_0$ whose score $\dot \ell_{\eta_0} = (\dot \ell_{k,\eta_0})_{k=1}^K$ satisfies $\int \dot \ell_{\eta_0}\dot \ell_{\eta_0}^T d \eta_0 = \mathrm{diag}(\sigma_1^{-2},\ldots,\sigma_K^{-2}) \equiv I_{\eta_0}^{-1}$, where $\sigma_k^2 = \int \psi_k^2 d \eta_{0,k}$ is the (efficient) asymptotic variance of $\hat \mu_k$.
By Le Cam's third lemma vanderVaart1998, for any $h \in \mathbb R^K$ we may deduce that $\sqrt n (\hat \mu - \mu_0)$ converges in distribution to a $N(h,I_{\eta_0}^{-1})$ random vector under $\{P_{n,h}\}$, where $P_{n,h}$ is the product measure formed by drawing $n$ copies of $(X,Y)$ under $\eta_{\mu_{0,k} + h_k/\sqrt n,k}$ for each subpopulation. This verifies Assumption (ref)(ref). Assumption (ref)(ref) holds by standard consistency arguments for heteroskedasticity-robust standard errors. It is also straightforward to verify, for instance by using suitable concentration inequalities, that Assumption (ref)(ref) holds. $\square$
We now present analogues of Theorems (ref) and (ref) for the semiparametric case. Recall that $\Pi$ denotes the class of all priors on $\mc M$ with positive, continuous, and bounded Lebesgue density. Similar to the parametric case, we say that $\{\delta_n\},\{\delta_n'\} \in 12 D$ are asymptotically equivalent if \[ \lim_{n \to \infty} P_{n,h}(\delta_n(X^n) = d) = \lim_{n \to \infty} P_{n,h}(\delta_n'(X^n) = d) \] for all $d \in \mc D$, $h \in 12 R^K$, and $(\mu_0,\eta_0) \in \mc M \times \mc H$.
It follows from Theorem (ref) that quasi-Bayes decisions with priors $\pi \in \Pi$ are asymptotically equivalent and optimal. Moreover, any decision that is asymptotically equivalent to a quasi-Bayes decision under a prior $\pi \in \Pi$ is optimal. In particular, bootstrap rules will be asymptotically equivalent to quasi-Bayes rules (and hence optimal) provided the bootstrap distribution for an efficient estimator $\hat \mu$ of $\mu$ and the quasi-posterior are sufficiently close.
Similar to the parametric case, we say $\{\delta_n\},\{\delta_n'\} \in 12 D$ fail to be asymptotically equivalent at $(\mu_0,\eta_0)$ if \[ \lim_{n \to \infty} P_{n,h}(\delta_n(X^n) = d) \neq \lim_{n \to \infty} P_{n,h}(\delta_n'(X^n) = d) \] for some $h \in \mathbb R^K$ and some $d \in {\cal D}$.
As in the parametric case, asymptotic equivalence to $\{\delta_n^*(\,\cdot\,;\pi)\}$ for some $\pi \in \Pi$ is necessary for optimality whenever there are no first-order ties. Thus, as discussed in Section (ref), plug-in rules may fail to be optimal in settings where the maximum risks $R(d,\mu)$ is only directionally differentiable at $\mu_0$. This includes the empirical application from Section (ref) and the further examples to treatment assignment and optimal pricing that we discuss in the next two sections.
This section expands on our running example of treatment assignment under partial identification. We first review several empirically relevant approaches for constructing bounds on the ATE in Section (ref). Each of these constructions leads to bounds $b_L(\mu)$ and $b_U(\mu)$ that will in general be only directionally differentiable in a vector of reduced-form parameters $\mu$. We then discuss implementation of the optimal decision rules in these settings in Section (ref). As far as we are aware, ours is the first work to propose optimal decisions for these realistic empirical settings.
\@startsection{paragraph}{4}{\z@} {0pt \@plus1ex \@minus.2ex} {-1em} {\normalfont}{Intersection Bounds.} This setting generalizes the empirical application from Section (ref). Suppose $X^n$ consists of data from $K$ observational studies. In each study $k$ we can consistently estimate lower and upper bounds $b_{L,k}(\mu_k)$ and $b_{U,k}(\mu_k)$ on the ATE as a function of population moments $\mu_k$. We then obtain the intersection bounds \[ b_L(\mu) = \max_{1 \leq k \leq K} b_{L,k}(\mu_k) \,, \quad \quad b_U(\mu) = \min_{1 \leq k \leq K} b_{U,k}(\mu_k) , \] with $\mu = (\mu_k)_{k=1}^K$. While the bounds $b_{L,k}$ and $b_{U,k}$ may themselves be smooth in $\mu_k$, the presence of the $\min$ and $\max$ operations makes the intersection bounds $b_L(\mu)$ and $b_U(\mu)$ only directionally differentiable in $\mu$.
\@startsection{paragraph}{4}{\z@} {0pt \@plus1ex \@minus.2ex} {-1em} {\normalfont}{Bounds via IV-like Estimands.} MST present an approach for bounding the ATE and other causal effects using IV-like estimands from observational studies. Suppose treatment is determined by $D = 12 I[U \leq v(Z)]$ where $U \sim \mathrm{Uniform}(0,1)$ and $Z = (X,Z_0)$ collects control variables $X$ and instrumental variables $Z_0$. According to HV1999,HV2005, the ATE may be expressed as a functional $\Gamma_0(m)$, where $m = (m_0,m_1)$ are the marginal treatment response (MTR) functions \[ m_d(u,x) = \mathbb{E}[ Y_d | U = u, X = x] \,, \quad d \in \{0,1\}, \] and \[ \Gamma_0(m) = \mathbb{E} \left[ \int_0^1 m_1(u,X) \, du - \int_0^1 m_0(u,X) \, du \right]. \]
MST show the MTR functions, and hence the identified set for the ATE, can be disciplined if we know the value of certain IV-like estimands. For ease of exposition, consider a single IV estimand \[ \beta_{IV} = \frac{\mathrm{Cov}(Y,Z_0)}{\mathrm{Cov}(D,Z_0)} \,, \] resulting from using $Z_0$ as an instrument for treatment status dummy $D$ in the observational study. The IV estimand may be expressed as $\beta_{IV} = \Gamma_\beta(m)$ where \[ \Gamma_\beta(m) = \mathbb{E} \left[ \int_0^1 m_0(u,X) s(0,Z_0) 12 I[u > p(z_0)] \, du + \int_0^1 m_1(u,X) s(1,Z_0) 12 I[u \leq p(z_0)]\, du \right] \] with $s(d,z) = \frac{z_0 - \mathbb{E}[Z_0]}{\mathrm{Cov}(Z_0,D)}$ and where $p(z_0) = \mathbb{E}[D|Z_0 = z_0]$ is the propensity score. In this case, the bounds of MST are \[ b_L(\mu) = \inf_{m \in \mc S : \Gamma_\beta(m) = \beta_{IV}} \Gamma_0(m) , \quad \quad b_U(\mu) = \sup_{m \in \mc S : \Gamma_\beta(m) = \beta_{IV}} \Gamma_0(m) , \] where $\mc S$ is a class of functions and $\mu = (\mathrm{Cov}(Y,Z_0), \mathrm{Cov}(D, Z_0), \mathbb{E}[Z_0], p)$, which is finite-dimensional if $Z_0$ has finite support (e.g. binary $Z_0$). They show that $b_L(\mu)$ and $b_U(\mu)$ may be expressed as the optimal values of linear programs parameterized by $\mu$. It is known from MilgromSegal2002 (see also Mills1956 and Williams1963 for linear programs) that the value of optimization problems may only be directionally differentiable in parameters.
\@startsection{paragraph}{4}{\z@} {0pt \@plus1ex \@minus.2ex} {-1em} {\normalfont}{Non-separable Panel Data Models.} Suppose the outcome for individual $i$ at date $t$ is of the form $Y_{it} = g(X_{it}, \alpha_i, \varepsilon_{it})$ where $X_{it}$ is a vector of covariates, $\alpha_i$ is a latent individual effect, and $\varepsilon_{it}$ is a vector of disturbances, which are independent across individuals and time. Consider an intervention that changes in covariates from $x^0$ to $x^1$. The ATE associated with the intervention is \[ \int \left( g(x^1,\alpha,\varepsilon) - g(x^0,\alpha,\varepsilon) \right) d Q(\alpha,\varepsilon), \] where $Q$ is a distribution over $(\alpha,\varepsilon)$. When outcomes and covariates are discrete, parametric restrictions on the distribution of $\varepsilon$ and functional form restrictions on $g$ are generally insufficient to point identify the ATE without parametric restrictions on the distribution of $\alpha$. A leading example is dynamic panel data models in which $X_{it}$ collects lagged values of a discrete outcome $Y_{it}$---see, e.g., HonoreTamer2006 and Torgovitsky2019. Building on HonoreTamer2006, Chernozhukovetal2013 and TorgovitskyPies2019 derive bounds on the ATE without parametric assumptions on $Q$. Their bounds may be expressed as the value of optimization problems (linear programs) parameterized by a finite-dimensional vector of choice probabilities $\mu$. As in the previous example, $b_L(\mu)$ and $b_U(\mu)$ may therefore only be directionally differentiable in $\mu$.
The first two examples are semiparametric and we can take $X^n$ to be a vector of estimators $\hat \mu$ of $\mu$. Assuming $\hat \mu$ is asymptotically efficient and the DM has available a consistent estimator $\hat I^{-1}$ of the asymptotic variance of $\hat \mu$, then the optimal decision can be implemented based on a $N(\hat \mu, (n \hat I)^{-1})$ quasi-posterior as in ((ref)).
In the third example with discrete outcomes and covariates, the distribution $P_{n,\mu}$ of the data $((Y_{it},X_{it})_{t=1}^T)_{i=1}^n$ can be identified with a multinomial distribution parameterized by $\mu$. In this case our parametric theory applies and the optimal decision can easily be implemented using either the bootstrap or Bayesian methods.
Our methods can also be applied to make optimal pricing decisions in models with rich unobserved heterogeneity using revealed-preference demand theory. Section (ref) provides the model and empirical setting. Section (ref) presents techniques for computing sharp bounds on functionals of counterfactual demand using linear programming. Finally, we discuss how to implement our methods in Section (ref).
The DM observes repeated cross sections of individual demands $X^n = \big( X_{b,1},\ldots,X_{b,n} \big)_{b=1}^{B}$, where each $X_{b,i} \in 12 R^m$ is the demand of individual $i$ for $m$ goods under prices $q_b$: \[ X_{b,i} = \mathrm{arg}\max_{x \in \mc B_b} u_i(x) , \] where $\mc B_b = \{x \in 12 R^m : x'q_b = 1\}$ is the budget set (expenditure is normalized to one). Individuals are heterogeneous in their utility functions $u_i$. We assume the demand system is rationalized by a random utility model with a probability distribution $F$ over utility functions $u$. The demand under $q_b$ of a randomly selected individual may therefore be interpreted as stochastic. The mass $p_b(s)$ of individuals whose demand is in a set $s \subset \mc B_b$ at price $q_b$ is
(see, e.g., KitamuraStoye2018, henceforth KS18).
The DM wishes to choose a new price vector $q_d$ for $d \in \mc D = \mc O \cup \mc C$. The price vectors $\{q_d : d \in \mc O\}$ are a subset of the observed prices $q_1,\ldots,q_B$, while each $\{q_d : d \in \mc C\}$ is a set of counterfactual price vectors. In principle, the set $\mc C$ of new price vectors could be large, representing prices rounded to nearest currency units or tax rates rounded to the nearest percentage. Let $w_d(X_d)$ represent a functional of demand (e.g., revenue or welfare) under prices $q_d$. The DM's goal is to choose $d \in \mc D$ to maximize the average of $w_d(X_d)$. For $d \in \mc O$ the average demand $\mathbb{E}[w_d(X_d)]$ is identified from observed choice behavior under $q_d$. However, the observed choice behavior is only sufficient to bound, but not point-identify, $\mathbb{E}[w_d(X_d)]$ for counterfactual prices $q_d$, $d \in \mc C$.
KitamuraStoye2019 present a general approach using linear programming to bound functionals of counterfactual demand. We introduce their approach with an example. Consider Figure (ref). There are two goods, and the DM has observed the demand for two price vectors $q_1$ and $q_2$, which generate the budget sets ${\cal B}_1$ and ${\cal B}_2$. The counterfactual price vector is $q_0$ which generates the counterfactual budget set $\mc B_0$.
The budget lines in Figure (ref) are divided into segments $s_{jb}$, called “patches” in KS18. Patch $s_{11}$ is the segment of $\mathcal{B}_1$ from the $y$-axis to the intersection of $\mathcal{B}_1$ and $\mathcal{B}_2$, patch $s_{21}$ is the segment of $\mathcal{B}_1$ between $\mathcal{B}_2$ and $\mathcal{B}_0$, and so forth.\footnote{Like KS18, we suppose for simplicity that the distribution of demand is continuous so we can disregard the “intersection patches” formed at the intersections of budget lines.} There are 9 potential combinations of patches consumers may choose from among ${\cal B}_1$ and ${\cal B}_2$: $(s_{11},s_{12})$, $(s_{11},s_{22})$, $(s_{11},s_{32})$, $(s_{21},s_{12})$, $(s_{21},s_{22})$, $(s_{21},s_{32})$, $(s_{31},s_{12})$, $(s_{31},s_{22})$, $(s_{31},s_{32})$. Each combination corresponds to a consumer type. By revealed preference we know a consumer will never choose $(s_{21},s_{12})$ or $(s_{31},s_{12})$. This leaves a total of 7 rational types of consumer. Let \[ A = \left[
\right]. \] The rows of $A$ correspond to $s_{11}, s_{21}, s_{31}, s_{12}, s_{22}, s_{32}$ and the columns of $A$ correspond to the 7 rational types. Let $p = (p_1(s_{11}),p_1(s_{21}), p_1(s_{31}), p_2(s_{12}), p_2(s_{22}), p_2(s_{32}))$ collect the corresponding choice probabilities. KS18 showed that the demand system $p$ is rationalizable if and only if \[ p = A f, \] for some $f \in \Delta^6$, the unit simplex in $12 R^7$, representing the probabilities of the $7$ rational types. These probabilities constrain the distribution $F$ of utilities.
Now consider choice behavior on the counterfactual budget set $\mc B_0$. Each of the 7 rational types may choose a counterfactual demand in patch $s_{10}$, $s_{20}$, or $s_{30}$, for a total of 21 potential types. By revealed preference, a consumer who chose $s_{31}$ must choose $s_{20}$ or $s_{30}$, and a consumer who chose $s_{32}$ must choose $s_{30}$. This leaves a total of 16 rational types. We may represent the system as
where $p^* = (p_0(s_{10}),p_0(s_{20}),p_0(s_{30}))$ collects the counterfactual choice probabilities for patches $s_{10}$, $s_{20}$, and $s_{30}$, $f^* \in \Delta^{15}$ collects the probabilities of observing each rational type, and \[ A^* = \left[
\right] , \] where the rows correspond to choosing $s_{11}, s_{21}, s_{31}, s_{12}, s_{22}, s_{32}, s_{10}, s_{20}, s_{30}$. Partition $A^*$ as \[ A^* = \left[
\right], \] where $A^*_{obs}$ collects the first $6$ rows of $A^*$ (corresponding to the observed budget sets) while $A^*_{unobs}$ collects the rows of $A^*$ corresponding to the patches $s_{10}$, $s_{20}$, and $s_{30}$.
Following KitamuraStoye2019, we may deduce sharp bounds on $\mathbb{E}[w_0(X_0)]$ as follows. Let $(\underline w_1, \underline w_2, \underline w_3)$ and $(\overline w_1, \overline w_2, \overline w_3)$ denote row vectors which collect the smallest and largest values of $w_0(x)$ for $x$ in $s_{10}$, $s_{20}$ and $s_{30}$. Then the bounds on $\mathbb{E}[w_0(X_0)]$ are \[
\]
More generally, the preceding argument applies with a collection of $B$ observed budget sets and multiple goods. In that case, representation ((ref)) holds for observed choice probabilities $p$ of patches on the $B$ observed budget sets and counterfactual choice probabilities $p^*$ of patches on the counterfactual budget set $\mc B_d$ generated by $q_d$ for $d \in \mc C$. Suppose there are $J$ patches across $\mc B_d$ and $T$ rational types. Then letting row vectors $\underline w_d$ and $\overline w_d$ collect the smallest and largest values of $w_d(x)$ for $x$ in each of the $J$ patches, sharp bounds on $\mathbb{E}[w_d(X_d)]$ are
where we have partitioned the $A^*$ matrix analogously to the simple 3 budget example.
As in BergemannSchlag2011, we assume the DM has a preference for robustness. That is, the DM wishes to choose $d$ to minimize maximum risk, where the maximum is taking over the identified set of counterfactual demand responses.
The DM's problem maps into our framework as follows. For each $d \in \mc O$, the value $\mathbb{E}[w_d(X_d)]$ is identified from observed choice behavior under $q_d$. The remaining reduced-form parameters are the patch probabilities $p$. Thus, $\mu = (p,(\mathbb{E}[w_d(X_d)])_{d \in \mc O})$. Here $\theta = F$ and $\Theta_0(\mu)$ is all $F$ that rationalize the patch probabilities $p$ consistent with revealed preference. The vector $Y = (w_d(X_d))_{d \in \mc C}$ collects functionals of demand under the counterfactual prices. Each $ \theta \equiv F \in \Theta_0(\mu)$ induces a distribution $G_\theta$ for $Y$. Although we suppressed it in the previous subsections, we now write $\mathbb{E}_\theta[w_d(X_d)]$ for $d \in \mc C$ to denote that the average counterfactual demand functional depends on the structural parameter $\theta$.
For $d \in \mc O$, the DM incurs risk \[ r(d,\theta,\mu) = -\mathbb{E}[w_d(X_d)], \] which is point identified from $\mu \equiv (p,(\mathbb{E}[w_d(X_d)])_{d \in \mc O})$. For $d \in \mc C$, the DM incurs risk \[ r(d,\theta,\mu) = -\mathbb{E}_\theta[w_d(X_d)], \] which is set-identified and may be bounded as described in the previous subsection. The maximum risk is \[ R(d,\mu) =
\] where $w_{d,L}(p)$ was defined in ((ref)).
This is a semiparametric model and our implementation follows the steps described in Section (ref). Partition $\mu = (\mu_1,\ldots,\mu_B)$ where $\mu_b$ collects the parameters corresponding to budget set $\mc B_b$. As we have assumed that each of the $B$ observed budgets is sampled in a repeated cross section, we can estimate each $\mu_b$ by just-identified GMM. As we observe repeated cross sections under $B$ different price vectors, we can simply take $\eta = \times_{b=1}^B \eta_b$, where $\eta_b$ is the distribution of $X_{b,i}$. We can then form a quasi-posterior $\pi_n(\mu|X^n)$ based on a limited-information Gaussian quasi-likelihood as in ((ref)). For $d \in \mc O$, the expected value $\mathbb{E}[w_d(X_d)]$ is an element of $\mu$ and so $\bar R_n(d)$ is simply its quasi-posterior mean: \[ \bar R_n(d) = -\int \mathbb{E}[w_d(X_d)] \, d \pi_n(\mu)\,. \] For $d \in \mc C$, the posterior maximum risk is \[ \bar R_n(d) = -\int w_{d,L}(p) \, d \pi_n(\mu)\,. \] This may be computed by sampling $p$ from the quasi-posterior, solving the linear program ((ref)) defining $w_{d,L}(p)$ for each draw of $p$, then taking the average across draws.
The optimal decision minimizes $\bar R_n(d)$ for $d \in \mathcal D = \mc O \cup \mc C$. As the value of a linear program is typically directionally differentiable, the asymptotic distribution of the posterior mean $\bar R_n(d)$ for $d \in \mc C$ may be different from that of the corresponding plug-in values $-w_{d,L}(\hat p)$. In this case, the optimal decision $\delta_n^*$ may outperform the plug-in rule.
We derived optimal statistical decision rules for discrete choice problems when payoffs depend on a set-identified parameter $\theta$ and the decision maker can use a point-identified parameter $\mu$ to deduce restrictions on $\theta$. Our notion of optimality combines a minimax approach to handle the ambiguity from partial identification of $\theta$ given $\mu$ with average risk minimization for $\mu$. In many empirically relevant applications, the maximum risk depends non-smoothly on $\mu$, making plug-in rules sub-optimal. We provided detailed applications to optimal treatment choice under partial identification and optimal pricing with rich unobserved heterogeneity. Our asymptotic approach is well suited for empirical settings in which the derivation of finite-sample optimal rules is intractable. While continuous decisions fall outside the scope of our theory and analysis, it would be interesting to study asymptotic efficiency in the continuous case as well.
\setcounter{page}{1}