EconBase
← Back to paper

Posterior and Likelihood Sensitivity in Bayesian Distributionally Robust Optimization

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

37,046 characters · 13 sections · 16 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Posterior and Likelihood Sensitivity in Bayesian Distributionally Robust Optimization

abstractWe introduce the notion of worst-case posterior and worst-case likelihood sensitivity. These measure, respectively, the sensitivity of the expected cost to worst-case perturbations of the posterior distribution and worst-case perturbations of the likelihood of a Bayesian model. Each defines a quantitative measure of robustness. A decision maker concerned about the sensitivity of the out-of-sample expected cost to deviations from her assumptions will want a decision for which both sensitivities are small. We derive posterior and likelihood sensitivities for uncertainty sets defined in terms of deviation measures. Posterior sensitivity vanishes when the posterior variance shrinks to zero, which occurs when parameter uncertainty is eliminated from learning. Parameter learning does not eliminate likelihood sensitivity. A distributionally robust formulation of a Bayesian optimization problem makes a near-Pareto-optimal tradeoff between performance (expected cost) and robustness (posterior and likelihood sensitivity).

{\bf Key words:} Bayesian distributionally robust optimization, worst-case sensitivity, model uncertainty, posterior sensitivity, likelihood sensitivity, robustness--performance tradeoff.

Introduction

Bayesian models provide a natural framework for learning from data. A Bayesian model requires two inputs from the decision maker (DM), a likelihood function $L^\theta(\cdot) \equiv L(\cdot|\theta)$ and a prior distribution $\rho$ for the uncertain parameter $\theta$. Together they characterize the joint distribution of the uncertain parameter $\theta$ and the random variable $Y$. The prior on $\theta$ is updated using observed samples of $Y$. The resulting posterior induces the predictive distribution which is used to compute the expected cost under a decision $x$. A DM concerned about misspecification would like a small expected cost that is insensitive to deviations from the predictive distribution.

In this paper, we introduce the notions of posterior and likelihood sensitivity. Posterior sensitivity is the sensitivity of the expected cost to worst-case deviations from the posterior while likelihood sensitivity defines an analogous notion for the likelihood function. They can be interpreted as quantitative measures of robustness. A DM concerned about the impact of misspecification on the out-of-sample performance would like a decision with small expected cost and low posterior and likelihood sensitivity. Adapting ideas from gotoh2018robust,gotoh2021calibration,gotoh2026sensitivity for robust empirical optimization problems, we show that a broad class of distributionally robust Bayesian optimization problems can be approximated by a “regularized" Bayesian problem where the regularizer is a weighted sum of posterior and likelihood sensitivities. It follows that approximately Pareto optimal tradeoffs between performance (expected cost) and robustness (posterior and likelihood sensitivity) can be obtained by solving a Bayesian Distributionally Robust Optimization (DRO) problem.

We derive expressions for posterior and likelihood sensitivity, which depend on the uncertainty sets for posterior and likelihood, and show that posterior sensitivity vanishes when the posterior distribution degenerates to a point mass, but not so for the likelihood. In other words, improvements in parameter estimates from updating the prior with data can reduce posterior, but not likelihood, sensitivity. As illustrated in gotoh2026sensitivity for empirical optimization, sensitivity measures can guide the selection of uncertainty sets, identify when the “price of robustness" is high, and guide system redesign to reduce this cost.

The outline of the paper is as follows. We briefly review the literature in Section (ref). We introduce the nominal model (the predictive distribution resulting from the posterior and likelihood) in Section (ref) and worst-case posterior and worst-case likelihood sensitivity in Section (ref). Examples where uncertainty sets for the posterior and likelihood are smooth $\phi$-divergence and formed using different $\phi$-divergences are considered. We show in Section (ref) that Bayesian DRO problems map out a nearly Pareto-optimal tradeoff between expected cost, posterior sensitivity and likelihood sensitivity. We consider a simple pricing application in Section (ref). We adopt the “constrained" formulation of the Bayesian DRO problem, where alternative posterior and likelihood distributions are determined by a constraint on the deviation from the nominal. A generalization to the penalty formulation of the Bayesian model can be found in the Appendix.

Literature review

This paper is related to three streams of work: approximations of DRO as a regularized nominal problem, robustness measures induced by worst-case sensitivity, and Bayesian DRO.

Several papers bertsimas2018,blanchet2019,Duchi2017variance,esfahani2018data,gao2023distributionally, gao2024wasserstein,kuhn2019wasserstein show that DRO can be approximated by a regularized empirical optimization problem. gotoh2026sensitivity takes this as the starting point, shows that the regularizer can be interpreted as a quantitative measure of robustness (worst-case sensitivity), and explores the implications for robust modeling including uncertainty set selection (family and size) and system design. The present paper extends this perspective to single stage Bayesian problems.

A Bayesian model consists of a likelihood function and a prior distribution; the predictive distribution (formed from the posterior and likelihood) is the nominal model. Generalizing ideas from gotoh2026sensitivity, we introduce worst-case posterior and worst-case likelihood sensitivity so that a Bayesian decision maker can quantify the fragility of her out-of-sample expected cost to the components of the predictive distribution. We show how each sensitivity measure depends on the uncertainty sets for the posterior and likelihood and that a distributionally robust Bayesian problem maps out a tradeoff between performance (expected cost) and robustness (posterior and likelihood sensitivities). Our results show that posterior sensitivity vanishes when there is no parameter uncertainty, but not so for the likelihood. shapiro2023bayesian develops Bayesian DRO and, among other results, derive worst-case likelihood sensitivity with a Kullback--Leibler uncertainty set. We complement this by introducing posterior sensitivity, and considering uncertainty sets for deviation measures beyond Kullback--Leibler divergence.

There is a related body of work in Bayesian statistics that studies the sensitivity of posterior-based inferences to perturbations of the (subjective) prior and likelihood (see for example berger1986Bayesian,gustafson1985Bayesian,Lavine1991Bayesian). A key difference is the focus on decision making and its role in our model (i.e., we explicitly include the objective function); in particular, how decisions can be chosen when posterior or likelihood sensitivity is large. By showing that the DRO objective can be decomposed into expected cost plue a weighted sum of posterior sensitivity and likelihood sensitivity, we provide a systematic approach to choosing decisions that balance performance and sensitivity to misspecification.

Nominal model

Let $(\theta, Y)$ be random variables with a joint distribution ${\mathbb P}({\rm d}\theta, {\rm d}y)$. We denote the marginal distribution ${\mathbb P}({\rm d}\theta) = \rho({\rm d}\theta)$ and the conditional distribution ${\mathbb P}({\rm d}y|\theta) = L^\theta({\rm d}y)$ so

eqnarray*[eqnarray* omitted — 94 chars of source]

Bayesian learning models belong to this class: ${\mathcal L} = \{L^{\theta}\vert \theta\in\Theta\}$ is a parameterized family of likelihoods and $\rho$ is the posterior distribution of the uncertain parameter $\theta$. The posterior is obtained by updating prior $\rho_0(d\theta)$ using data $Y_1, \cdots, Y_m$ and Bayes' rule. The likelihood function $L^\theta$ is a modeling choice, while the posterior depends on the user-specified prior $\rho_0(d\theta)$ and likelihood $L^{\theta}(dy)$. We are concerned about the sensitivity of the expected cost to deviations from ${\mathbb P}({\rm d}\theta, {\rm d}y)$.

Another important class is mixture models with populations indexed by $\theta\in\{\theta_1, \cdots, \theta_n\}$ and population distribution ${\mathbb P}({\rm d}y|\theta=\theta_i) = L^{\theta_i}({\rm d}y)$. For concreteness, we refer to the marginal distribution $\rho(d\theta)$ as the posterior and the conditional distribution $L^\theta({\rm d}y)$ as the likelihood, fully recognizing that it is more general than a Bayesian learning model.

Let $f$ be a cost function of the variables $(\theta, Y)$. The expected cost is

eqnarray*[eqnarray* omitted — 136 chars of source]

The expected cost depends on the joint distribution. We are concerned about the sensitivity of the expected cost relative to worst-case deviations from $\mathbb P$.

Worst-case sensitivity

We use ideas from gotoh2026sensitivity to define worst-case sensitivity with respect to the posterior and the likelihood.

Uncertainty in the posterior and likelihood

We begin by defining deviations from the posterior $\rho({\rm d}\theta)$ and likelihood ${\mathbb P}({\rm d}y | \theta)$.

Let $\eta$ be a distribution on $\Theta$ and ${\rm d}( \eta | \rho)$ be a deviation measure of $\eta$ from $\rho$. For every $\theta\in\Theta$ let ${\mathbb Q}^\theta$ be an alternative to $L^\theta$ for the distribution of $Y$ and ${\rm d}({\mathbb Q}^\theta| L^\theta)$ a measure of deviation of ${\mathbb Q}^\theta$ from $L^\theta$. It will occasionally be convenient to write ${\mathcal L} = \{ L^\theta| \theta\in\Theta\}$ when we talk about the entire family of distributions.

Let

eqnarray[eqnarray omitted — 266 chars of source]

be uncertainty sets for the posterior and likelihood distributions. We use the same notation ${\rm d}(\cdot|\cdot)$ for the deviation measure in each uncertainty set. However, different deviation measures can be used for the posterior and likelihood (see Section (ref)).

Consider the worst-case expected cost

eqnarray[eqnarray omitted — 226 chars of source]

The inner expectation is conditional on $\theta$; we consider worst-case deviations ${\mathbb Q}^\theta$ from the nominal likelihood $L^\theta$ for each $\theta$. The outer expectation considers worst-case deviations $\eta$ from the posterior $\rho$ so changes the weights on $\theta$.

Regularization

Consider first the inner problem. Since the worst-case expected cost is increasing in the size of the uncertainty set $\delta$, we can write for every $\theta\in\Theta$

eqnarray[eqnarray omitted — 261 chars of source]

where ${\mathcal A}_{L^\theta}\big(\delta; f(\theta, Y)\big)$, the ambiguity cost, is non-negative and non-decreasing in $\delta$.

Similarly, we can write the outer problem

eqnarray[eqnarray omitted — 223 chars of source]

where the ambiguity cost ${\mathcal A}_{\rho}\big(\varepsilon; \Psi(\theta)\big)$ is again non-negative and increasing in $\varepsilon$.

Substituting (ref) into (ref) it follows that the worst-case expected cost (ref) can be written

eqnarray*[eqnarray* omitted — 185 chars of source]

with ambiguity cost

eqnarray*[eqnarray* omitted — 319 chars of source]

We make the following assumptions about the ambiguity costs. The functions $g(\varepsilon)$ and $g(\delta)$ for the two ambiguity costs need not be the same.

assumptionThe ambiguity costs ${\mathcal A}_{L^\theta}\big(\delta; f(\theta, Y)\big)$ and ${\mathcal A}_{\rho}\big(\varepsilon; \Psi(\theta)\big)$ are such that \begin{enumerate} • there is a non-decreasing function $g(\delta)$ such that $g(\delta)\rightarrow 0$ when $\delta\rightarrow 0$ and \begin{eqnarray*} {\mathcal A}_{L^\theta}\big(\delta; f(\theta, Y)\big) = g(\delta) {\mathcal S}_{L^\theta}(f(\theta, Y))+ o(g(\delta)) \end{eqnarray*} for every $\theta\in{\Theta}$; • there is a non-decreasing function $g(\varepsilon)$ such that $g(\varepsilon)\rightarrow 0$ when $\varepsilon\rightarrow 0$ and \begin{eqnarray*} {\mathcal A}_{\rho}\big(\varepsilon; \Psi(\theta)\big) = g(\varepsilon) {\mathcal S}_\rho(\Psi(\theta)) + o(g(\varepsilon)). \end{eqnarray*} • ${\mathcal S}_{\rho}(\Psi)$ is continuous in $\Psi$. \end{enumerate}

Intuitively, $\Psi(\theta)$ is the worst-case expected cost when the uncertainty set is $\{{\mathcal L}^\theta(\delta)\}$ and

eqnarray*[eqnarray* omitted — 224 chars of source]

is the increase in expected cost relative to the nominal. When Assumption (ref) holds, ${\mathcal S}_{L^\theta}(f(\theta, Y))$ can be interpreted as the sensitivity of the expected cost under worst-case perturbations; it depends on the deviation measure ${\rm d}({\mathbb Q} | L^\theta)$ that defines the uncertainty set. For example, ${\mathcal S}_{L^\theta}(f(\theta, Y))$ is the standard deviation of the cost $\sigma_{L^\theta} (f(\theta, Y))$ when ${\rm d}({\mathbb Q}|L^\theta)$ is smooth $\phi$-divergence. Sensitivity measures induced by other measures of deviation can be found in gotoh2026sensitivity and references cited therein (see also Section (ref)).

Likewise, ${\mathbb E}_\rho \big\{\Psi(\theta) \big\}$ is the expected value of $\Psi(\theta)$ when $\theta$ has distribution $\rho$ and ${\mathcal S}_{\rho}(\Psi)$ is the sensitivity of the expected cost ${\mathbb E}_\rho \big\{\Psi(\theta) \big\}$ under worst-case perturbations of $\rho$.

Uncertainty in posterior and likelihood contribute to increases in the total expected cost

eqnarray*[eqnarray* omitted — 339 chars of source]

Our goal is to understand how the sensitivity of the total cost depends on the posterior and likelihood.

The following result gives an expansion of the value function when $\varepsilon$ and $\delta$ are small.

theoremSuppose that Assumption (ref) holds. Then \begin{eqnarray*} V(\epsilon, \delta) & = & {\mathbb E}_\rho\Big\{{\mathbb E}_{L^\theta}[f(\theta, Y)] \Big\} + g(\delta){\mathbb E}_\rho\Big\{{\mathcal S}_{L^\theta}(f(\theta, Y))\Big\} + g(\varepsilon) {\mathcal S}_\rho\Big({\mathbb E}_{L^\theta}(f(\theta, Y))\Big) \\ [5pt] & & + o(g(\delta)) + o(g({\varepsilon})). \end{eqnarray*}
proofIt follows from (ref) and Assumption (ref) that \begin{eqnarray*} \Psi(\theta) &= & {\mathbb E}_{L^\theta}[f(\theta, Y)] + {\mathcal A}_{L^\theta}\big(\delta; f(\theta, Y)\big) \\ & = & {\mathbb E}_{L^\theta}[f(\theta, Y)] + g(\delta) {\mathcal S}_{L^\theta}(f(\theta, Y))+ o(g(\delta)). \end{eqnarray*} From (ref) and Assumption (ref) we have \begin{align*} V(\varepsilon, \delta) & = \max_{\eta \in {\mathcal P}(\varepsilon)} {\mathbb E}_\eta \big\{\Psi(\theta) \big\} \\[8pt] & = \max_{\eta \in {\mathcal P}(\varepsilon)} {\mathbb E}_\eta\Big\{{\mathbb E}_{L^\theta}[f(\theta, Y)] + g(\delta) {\mathcal S}_{L^\theta}(f(\theta, Y))+ o(g(\delta))\Big\} \\[8pt] & = {\mathbb E}_\rho\Big\{{\mathbb E}_{L^\theta}[f(\theta, Y)]+ g(\delta) {\mathcal S}_{L^\theta}(f(\theta, Y))+ o(g(\delta)) \Big\} \\[8pt] & \qquad + g(\varepsilon){\mathcal S}_\rho\Big({\mathbb E}_{L^\theta}[f(\theta, Y)] + g(\delta) {\mathcal S}_{L^\theta}(f(\theta, Y))+ o(g(\delta))\Big) + o(g(\varepsilon)) \\[8pt] & = {\mathbb E}_\rho\Big\{{\mathbb E}_{L^\theta}[f(\theta, Y)] \Big\} + g(\delta){\mathbb E}_\rho\Big\{{\mathcal S}_{L^\theta}(f(\theta, Y))\Big\} + g(\varepsilon) {\mathcal S}_\rho\Big({\mathbb E}_{L^\theta}(f(\theta, Y))\Big) \\[8pt] & \qquad + g(\varepsilon) \left\{{\mathcal S}_\rho\Big({\mathbb E}_{L^\theta}[f(\theta, Y)] + g(\delta) {\mathcal S}_{L^\theta}(f(\theta, Y)) + o(g(\delta)) \Big) - {\mathcal S}_\rho\Big({\mathbb E}_{L^\theta}(f(\theta, Y))\Big)\right\} \\[8pt] & \qquad + o(g(\delta)) + o(g({\varepsilon})) \\[8pt] & = {\mathbb E}_\rho\Big\{{\mathbb E}_{L^\theta}[f(\theta, Y)] \Big\} + g(\delta){\mathbb E}_\rho\Big\{{\mathcal S}_{L^\theta}(f(\theta, Y))\Big\} + g(\varepsilon) {\mathcal S}_\rho\Big({\mathbb E}_{L^\theta}(f(\theta, Y))\Big) \\[8pt] & \qquad + o(g(\delta)) + o(g({\varepsilon})) \end{align*} where the last equality follows from the continuity of ${\mathcal S}_\rho(\Psi)$.

Worst-case sensitivity

Given likelihood function $L^\theta(y)$ ($\theta\in\Theta$) and posterior $\rho$, worst-case posterior sensitivity is the sensitivity of the expected cost under worst-case deviations from the posterior defined by the uncertainty set ${\mathcal P}(\varepsilon)$:

eqnarray*[eqnarray* omitted — 241 chars of source]

Given the nominal posterior $\rho$, likelihood $L^\theta(y)$, and uncertainty set ${\mathcal L}^\theta(\delta)$ for deviations from the likelihood, worst-case likelihood sensitivity is the sensitivity of the expected cost with respect to worst-case deviations from the nominal likelihood

eqnarray*[eqnarray* omitted — 244 chars of source]

Note that posterior (likelihood) sensitivity ${\mathcal S}_\rho$ depends on the deviation measure ${\rm d}(\eta|\rho)$ that defines the uncertainty set ${\mathcal P}(\varepsilon)$, and likelihood sensitivity ${\mathcal S}_{L^\theta}$ on the deviation measure ${\rm d}({\mathbb Q}|L^\theta)$ that controls deviations from the nominal likelihood.

In general, ${\mathcal S}_\rho$ is a generalized measure of deviation gotoh2026sensitivity,rockafellar2006generalized so posterior sensitivity is zero when $\rho$ is a point mass (i.e., there is no uncertainty about $\theta$) or when $g(\theta) = {\mathbb E}_{L^\theta}[f(\theta, Y)]$ is constant on the support of $\rho$. Likelihood sensitivity is zero only when, for $\rho$-almost every $\theta$, the sensitivity of the cost under $L^\theta$ is zero. For the deviation measures considered here, this occurs when $f(\theta, y)$ is constant in $y$ on the support of $L^\theta$.

The following example considers posterior and likelihood sensitivity when deviation measures are smooth $\phi$-divergence.

Smooth $\phi$-divergence

Suppose that $\eta$ is absolutely continuous with respect to $\rho$ and ${\mathbb Q}^\theta$ is absolutely continuous with respect to $L^\theta$ for every $\theta\in\Theta$. Let $\frac{{\rm d} \eta}{{\rm d}\rho}$ denote the Radon--Nikodym derivative of $\eta$ with respect to $\rho$, and for every $\theta\in\Theta$, $\frac{{\rm d}Q^\theta}{{\rm d}L^\theta}$ be the Radon--Nikodym derivative of $Q^{\theta}$ with respect to $L^\theta$.

For the likelihood function, $\phi$-divergence of ${\mathbb Q}^\theta$ with respect to $L^\theta$, for every $\theta\in{\Theta}$, is defined by

eqnarray*[eqnarray* omitted — 162 chars of source]

In the case of smooth $\phi$-divergence, Assumption (ref) holds and the inner problem gotoh2026sensitivity

eqnarray*[eqnarray* omitted — 248 chars of source]

For the outer problem we measure the deviation of $\eta$ from $\rho$ using $\phi$-divergence:

eqnarray*[eqnarray* omitted — 123 chars of source]

and it follows that

eqnarray*[eqnarray* omitted — 217 chars of source]

In particular, $g(\delta) = \sqrt{\delta}$ and

eqnarray*[eqnarray* omitted — 108 chars of source]

for the inner problem, where ${\sigma}_{L^\theta}(f(\theta, Y)) = \sqrt{ {\mathbb V}_{L^\theta}(f(\theta, Y))}$ is the standard deviation of the cost under the distribution $L^\theta$. For the outer problem, $g(\varepsilon) = \sqrt{\varepsilon}$ and

eqnarray*[eqnarray* omitted — 103 chars of source]

Note that the regularizer in the expansion of the worst-case expected costs are given by the standard deviation because we are using smooth $\phi$-divergence as the deviation measure. Other uncertainty sets give different regularizers gotoh2026sensitivity.

It now follows from Theorem (ref) that the worst-case expected cost

eqnarray*[eqnarray* omitted — 360 chars of source]

Worst-case posterior sensitivity is

eqnarray[eqnarray omitted — 159 chars of source]

and worst-case likelihood sensitivity is

eqnarray[eqnarray omitted — 177 chars of source]

Both sensitivity measures depend on the posterior and objective function, but quite differently. Posterior sensitivity is the standard deviation of the conditional expectation $g(\theta) = {\mathbb E}_{L^\theta}[f(\theta, Y)]$ and will be small if $g(\theta)$ is “flat" on the support of $\rho$, even if the standard deviation of $\theta$ is large. It vanishes when the posterior variance shrinks (e.g.) because of learning. Likelihood sensitivity depends on the variance of the cost given $\theta$. It typically will not vanish even if $\theta$ is known because there is still uncertainty about the distribution $L^\theta$ of $Y$.

The paper shapiro2023bayesian derives worst-case likelihood sensitivity with a Kullback--Leibler uncertainty set (the inner “max" in (ref)) but not the posterior. We complement this by introducing posterior sensitivity and by showing how the worst-case objective can be approximated by expected cost, posterior sensitivity, and likelihood sensitivity.

Mix and match

As noted in the discussion following (ref), there is no reason to restrict ourselves to smooth $\phi$-divergence for the deviation measure or to have the same deviation measure for the posterior and likelihood. For example, if instead of smooth $\phi$-divergence, we use

eqnarray*[eqnarray* omitted — 93 chars of source]

to define the uncertainty set for the likelihood function, we have gotoh2026sensitivity\footnote{Given $\theta$, $\mathrm{CVaR}_{L^\theta,\alpha}(f(\theta, Y))$ is the conditional-value-at-risk of the cost $f(\theta, Y)$ at level $\alpha$ when $Y$ has distribution $L^\theta$:

eqnarray*[eqnarray* omitted — 178 chars of source]

}

eqnarray*[eqnarray* omitted — 129 chars of source]

(WCS for other deviation measures are also derived in gotoh2026sensitivity). We choose this uncertainty set and deviation measure if we are concerned about misspecification of the right tail of the cost distribution. If smooth $\phi$-divergence is used for the posterior and this alternative is used for the likelihood, we have

align*[align* omitted — 281 chars of source]

Distributionally Robust Optimization

Performance--sensitivity tradeoffs in DRO

Theorem (ref) shows that the worst-case expected cost can be locally expanded as

eqnarray[eqnarray omitted — 555 chars of source]

This representation is related to results showing the relationship between DRO and a regularized nominal problem. For empirical optimization problems, gotoh2026sensitivity shows that the regularization term is not just a mathematical object that connects the nominal and worst-case problems, but has the physical interpretation as worst-case sensitivity. (ref) shows an analogous result for a Bayesian DRO problem, where the regularizer is now a weighted sum of posterior sensitivity ${\mathcal S}_\rho({\mathbb E}_{L^\theta}[f(x, \theta, Y)])$ and likelihood sensitivity ${\mathbb E}_\rho ({\mathcal S}_{L^\theta}(f(x, \theta, Y)))$, and varying the uncertainty set sizes maps out a tradeoff between performance (expected cost) and robustness (posterior and likelihood sensitivity).

Experiment

Demand $D(p)$ is normal with standard deviation $\sigma$ and expected value ${\mathbb E}[D(p)] = a + b p$ which depends on the price $p$. We assume that $\sigma$ is known and constant whereas $a$ and $b$ are constant but unknown to the DM. At the time of the decision, the decision maker has a posterior on $(a, b)$ which we assume to be normal. (This could have been obtained by updating an initial prior using data and Bayes' rule.) The DM's revenue when price is $p$ and observed demand is $D(p)$ is $p\cdot D(p)$. The expected revenue under the posterior is ${\mathbb E}_\rho [p \, D(p)] = p \cdot ({\mathbb E}_\rho[a] + {\mathbb E}_\rho[b]\cdot p)$; the nominal DM chooses price to maximize this quantity.

We assume a modified $\chi^2$--uncertainty set for the posterior and likelihood, i.e., $\phi(z) = \frac{1}{2}(z-1)^2$. It follows that posterior and likelihood sensitivities are given by

align[align omitted — 184 chars of source]

For this experiment, under the baseline posterior, $a$ has mean $5$ and standard deviation $0.75$, and $b$ has mean $-0.5$ and standard deviation $2.5$; we assume $a$ and $b$ are independent under the posterior. We set $\sigma = 3$. For the nominal problem the optimal price is $p =\$5$; the optimal expected reward is $\$12.50$.

Figure (ref) shows reward--posterior sensitivity and reward--likelihood sensitivity frontiers when the posterior variance for $a$ and $b$ is 0, $1/4$ of the baseline, the baseline, and double the baseline. This simulates the effect of learning, which reduces the posterior variance, on posterior and likelihood sensitivity. We see in the first plot that posterior sensitivity is increasing in the posterior variance, quantifying the sensitivity of the expected reward to the posterior. Note that posterior sensitivity is $0$ when there is no parameter uncertainty. The DRO solution makes a tradeoff between expected reward and worst-case posterior sensitivity. For this particular example, DRO is quite effective when the posterior variance is large ($2\times$baseline) as substantial reduction in the posterior sensitivity from its value under the nominal model can be achieved with minimal impact on expected cost. It is less effective at reducing posterior sensitivity when the posterior variance is small ($0.25 \times$baseline). However, the frontier shows that this is not really necessary as the posterior sensitivity is already relatively small.

figure[figure omitted — 874 chars of source]

From the lower plot, we see that the mean--likelihood sensitivity frontiers lie on top of one another (the small deviations are from numerical errors). This is not a general property but a quirk of this example; it would not be the case if $\sigma$ was also uncertain\footnote{ It is easy to see that

eqnarray*[eqnarray* omitted — 198 chars of source]

so the expected reward, given likelihood sensitivity ${\mathcal S}_{\mathcal L}$, does not depend on the posterior variance. It follows that the reward--likelihood sensitivity frontiers do not depend on the posterior variance. It can also be shown that the reward--posterior sensitivity frontier does not depend on the likelihood variance $\sigma^2$ (again, a quirk of the problem).}. The main message, however, is that we tradeoff between expected reward, posterior sensitivity, and likelihood sensitivity as the size of the uncertainty set $\varepsilon=\delta$ changes, with DRO prices being mapped in Figure (ref).

figure[figure omitted — 252 chars of source]

We can see from (ref) that the likelihood sensitivity does not vanish when the posterior variance is $0$, so long as $p\neq 0$. This is a general property: while more data shrinks the posterior and eliminates posterior sensitivity (as observed in Figure (ref)), uncertainty about the distribution of demand associated with the likelihood model remains unless $\sigma=0$.

figure[figure omitted — 719 chars of source]

Figure (ref) shows performance--robustness frontiers when the posterior variance is held fixed and the variance $\sigma^2$ of the likelihood function is varied. The reward--likelihood sensitivity frontier is now increasing in $\sigma^2$ whereas the reward--posterior frontier remains unchanged.

Conclusion

Bayesian models account for parameter uncertainty by specifying a prior on the uncertain parameters. However, the expected cost can still be sensitive to deviations from the likelihood and the posterior (obtained by updating the user-specified prior using Bayes' rule and data) which may result in poor out-of-sample performance. Worst-case posterior sensitivity and worst-case likelihood sensitivity measure the sensitivity of the expected cost to worst-case deviations from the posterior and the likelihood, respectively. They are measures of robustness and a decision maker, concerned about the fragility of her assumptions, may prefer a decision for which both robustness measures are small. Worst-case posterior sensitivity vanishes when the posterior concentrates to a single point, which can occur when parameter uncertainty is reduced from learning. This is not the case with likelihood sensitivity. This is clear in the case of smooth $\phi$-divergence, where posterior sensitivity is the posterior standard deviation of the conditional expectation of the cost whereas likelihood sensitivity is the posterior expected value of its conditional standard deviation. A nearly Pareto-optimal tradeoff between performance (expected cost) and robustness (worst-case posterior and likelihood sensitivity) can be mapped out by solving a Distributionally Robust Bayesian Optimization problem.

\paragraph{\bf Acknowledgments} Jun-ya Gotoh is supported by the MEXT Grant-in-Aid 24K01113. Michael Kim is supported by the Natural Sciences and Engineering Research Council (NSERC) Discovery Grant RGPIN-2015-04019. Andrew Lim is supported by the Ministry of Education, Singapore, under its 2024 Academic Research Fund Tier 2 grant call (Award ref: MOE-T2EP20224-0018).