EconBase
← Back to paper

Robust decision-making under risk and ambiguity

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

86,305 characters · 13 sections · 83 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Robust decision-making under risk and ambiguity

abstractEconomists often estimate economic models on data and use the point estimates as a stand-in for the truth when studying the model's implications for optimal decision-making. This practice ignores model ambiguity, exposes the decision problem to misspecification, and ultimately leads to post-decision disappointment. Using statistical decision theory, we develop a framework to explore, evaluate, and optimize robust decision rules that explicitly account for estimation uncertainty. We show how to operationalize our analysis by studying robust decisions in a stochastic dynamic investment model in which a decision-maker directly accounts for uncertainty in the model's transition dynamics.\\
tabular[tabular omitted — 159 chars of source]

\setcounter{page}{1} \thispagestyle{empty}

\FloatBarrier

Introduction

Decision-makers often confront uncertainties when determining their course of action. For example, individuals save to cover uncertain medical expenses in old age French.2014. Firms set prices in an uncertain competitive environment Ilut.2020, and policy-makers face uncertainties about future costs and benefits when voting on climate change mitigation efforts Barnett.2020. We consider the situation in which a decision-maker posits a collection of economic models to inform his decision-making process. Each model formalizes the relevant objectives and trade-offs involved and provides an implicit rule for optimal decisions. Uncertainty is limited to risk for a given model, as the model induces a unique probability distribution over possible future outcomes. However, a decision-maker also faces model ambiguity as the true model within the collection remains uncertain Arrow.1951,Knight.1921.\\

It is the standard practice in economics to estimate models on data and use the point estimates as a stand-in for the truth when studying the model's implications and optimal decision-making.\footnote{See examples in labor economics Adda.2017,Blundell.2012, industrial organization Hortacsu.2019,Igami.2017, and international trade Bagwell.2021,Eaton.2011.} This approach ignores model ambiguity, resulting from the remaining parametric uncertainty after the estimation, and opens the door for the misspecification of the decision problem. As-if decisions, decisions that are optimal if the point estimates used to inform decisions are correct Manski.2021, often turn out to be very sensitive to misspecification Smith.2006. This danger creates the need for robust decisions that perform well over a whole range of different models instead of as-if decisions that perform best for one particular model. However, increasing the robustness of decisions, often measured by a performance guarantee under a worst-case scenario, reduces performance in all other cases. Striking a balance between the two objectives is challenging.\\

We solve this trade-off and determine the optimal level of robustness by combining insights from statistical decision theory Berger.2010 with data-driven robust optimization Bertsimas.2018. A core concept in statistical decision theory is a statistical decision function (SDF) that provides a procedure to map all available data into decisions. At the same time, the literature on data-driven robust optimization provides us with precisely such procedures for decision-making with varying levels of robustness against misspecification of the decision problem. Our main contribution is interpreting these procedures as SDFs and evaluating their performance with the toolkit of statistical decision theory. This insight allows us to systematically determine the optimal level of robustness. In doing so, we bring together and extend research in economics and operations research by using econometric models in complex decision problems Bertsimas.2006,Manski.2021.\\

In our application, we revisit \citetalias{Rust.1987} seminal bus replacement problem. Model ambiguity is particularly consequential in dynamic models where the impact of erroneous decisions accumulates over time Mannor.2007.\footnote{\citetalias{Rust.1987} model serves as a computational illustration in a variety of settings. See for example Christensen.2019, Iskhakov.2016, Reich.2018, and Su.2012a.} In the model, the manager Harold Zurcher implements a maintenance plan for a fleet of buses that maximizes his expected discounted utility. He faces uncertainty about the future mileage utilization of the buses but has data on past utilization available to inform his decisions. While \citetalias{Rust.1987} original goal was to describe the investment behavior of Harold Zurcher, our analysis is normative. We are interested in how a generic decision-maker should make decisions in this instance.\\

The bus replacement problem is typically modeled as a standard Markov decision problem (MDP), and the point estimates for the mileage utilization are treated as-if they correspond to the true parameters. The solution of the MDP is an as-if decision rule that is optimal given the estimates. This approach ignores model ambiguity. From the perspective of statistical decision theory, an MDP is just one particular example of an SDF suitable for analyzing the bus replacement problem. We, on the other hand, consider a whole class of SDFs called robust Markov decision problems (RMDP) Ben-Tal.2009. RMDPs generalize the standard MDP, as they consider a whole set of distributions for the transition dynamics collected in an ambiguity set. The solution of an RMDP is a robust decision rule that is optimal under a worst-case scenario for all mileage utilization distributions inside the ambiguity set. We follow the literature and construct the ambiguity set so that it contains all distributions we cannot reject with a certain level of confidence $\omega \in [0, 1]$ around the point estimates under any possible realization of the data Ben-Tal.2013. The size of the ambiguity set is a choice by the decision-maker and determines the level of robustness. Given the realization of the data, the robust decision rule based on the solution of an RMDP is always conditional on the specified level of robustness. Each choice of $\omega$ defines a different RMDP, and applying the toolkit of statistical decision theory allows us to determine the optimal level of robustness $\omega^*$ within the whole class of SDFs.\\

To do so, we compare the performance of RMDPs with varying levels of robustness under different decision-theoretic criteria. We consider the situation before any data on mileage utilization is available and implement an ex-ante decision-theoretic analysis. We explore the performance of robust decision rules over the whole probability simplex and are thus able to determine the optimal level of robustness. Throughout, we compare robust and as-if decision rules, as the standard Markov decision problem remains one SDF within the broader class we consider.\\

Figure (ref) stresses the point that each RMDP is a different SDF that characterizes robust decisions for any realization of the data. Here, for example, we consider two RMPDs with different levels of robustness -- $\omega_1$ and $\omega_2$ -- that, once data realizes, lead to different decision rules. This situation creates the need to compare their performance under alternative decision-theoretic criteria.

figure[figure omitted — 1,573 chars of source]

\FloatBarrier Our insight to evaluate robust decisions using statistical decision theory applies to the whole literature on data-driven robust optimization. There exists a growing number of applications of data-driven robust decision-making in a variety of settings, including portfolio decisions Jin.2020,Zymler.2013, elective admission to hospitals He.2019,Meng.2015, the timing of medical interventions Goh.2018,Kaufman.2017, and managing the production of renewable energy Alismail.2018,Samuelson.2017.\footnote{See the recent surveys by Bertsimas.2018, Keith.2021, and Rahimian.2019 for numerous additional examples.} \\

Despite its broad field of application, the existing literature only offers limited guidance on choosing the optimal level of robustness. At the most basic level, the recommendations range from simply advocating a high level of robustness Ben-Tal.2013 to choosing a level of robustness that ensures a pre-specified worst-case performance Brown.2012. These approaches ignore the trade-off between a performance guarantee under a worst-case scenario and reduced performance in all other cases. Most recently, and much closer to our approach, Gotoh.2021 put the robustness trade-off front and center. Adopting ideas from the machine learning literature, they calibrate the level of robustness by trading off the mean and variance in the out-of-sample performance of a robust decision rule. However, their approach restricts attention to the neighborhood of a realized point estimate. Thus, their analysis is ex-post as it does not aggregate performance over all possible realizations of the data.\\

At the same time, in econometrics, there is a burgeoning interest in assessing the sensitivity of findings to model or moment misspecification.\footnote{See for example Andrews.2020, Andrews.2017, Armstrong.2021, Bonhomme.2020, Chernozhukov.2020, Christensen.2019, and Honore.2020.} Our work is related to Jorgensen.2021, who develops a measure to assess the sensitivity of results by fixing a subset of parameters of a model before the estimation of the remaining parameters. Our approach differs as we directly incorporate model ambiguity in the design of the decision-making process and assess the performance of a decision rule under misspecification of the decision environment. As such, our focus on ambiguity faced by decision-makers about the model draws inspiration from the research program summarized in Hansen.2016 that tackles similar concerns with a theoretical focus. We complement recent work by Saghafian.2018, who works in a setting similar to ours, but does not use statistical decision theory to determine the optimal robust decision rule. In ongoing work, Eisenhauer.2021 use statistical decision theory to structure policy decisions in light of uncertainty about counterfactual policy predictions due to the remaining model ambiguity after the estimation of a model. While they conduct an ex-post evaluation of alternative policy proposals using decision-theoretic criteria, we perform a proper ex-ante analysis of competing decision rules. We evaluate each rule's performance under all possible parameterizations of the model and directly account for the model ambiguity in their construction. In addition, we contribute to the work on optimal treatment allocation started in Manski.2004 and Manski.2009, which characterizes the structure of optimal statistical decision functions and provides (asymptotic) bounds on their performance Hirano.2009,Stoye.2009,Tetenov.2012,Stoye.2012b,Kitigawa.2018.\\

The structure of the remaining analysis is as follows. In Section (ref), we present statistical decision theory as our framework to compare decision rules. We then set up a canonical model of a data-driven robust Markov decision problem in Section (ref) and outline the decision-theoretic determination of the optimal level of robustness. Section (ref) presents our analysis of the robust bus replacement problem. Section (ref) concludes.

\FloatBarrier

Statistical decision theory

We now show how to compare as-if decision-making to its robust alternatives using statistical decision theory. We first review the basic setting and then turn to a classic urn example to illustrate some key points.

Decision problem

We study a decision problem in which the consequence $c \in \mathcal{C}$ of various alternative actions $a\in\mathcal{A}$ depend on the parameterization $\theta\in \Theta$ of an economic model. A consequence function $\rho: \mathcal{A} \times \Theta \mapsto \mathcal{C}$ details the consequence of action $a$ under parameters $\theta$:

align*[align* omitted — 34 chars of source]

A decision-maker ranks consequences according to a utility function $u: \mathcal{C} \mapsto \mathbb{R}$, where higher values are more desirable. The structure of the decision problem $(\mathcal{A}, \Theta, \mathcal{C}, \rho, u)$ is known, but the true parameterization $\theta_0$ is uncertain. As a result, the consequences of a particular action are ambiguous. An observed sample of data $\psi \in \Psi$, however, provides a signal about the true parameters, as $P_{\theta}$ -- the sampling distribution of $\psi$ -- differs by $\theta$. A statistical decision function (SDF) $\delta: \Psi \mapsto \mathcal{A}$ is a procedure that determines an action for each possible realization of the sample.\\

In our application, we study the bus replacement problem with unknown future mileage utilizations. The decision problem is dynamic, so a decision-maker acts by committing to a plan that specifies whether to maintain or replace a bus in any possible future scenario. The consequences of executing a particular plan are a stream of maintenance costs, aggregated by its discounted sum of utilities. A plan's total utility is determined by the true distribution of the bus mileage utilization. The optimal decision rule based on an RMDP depends on the observed sample of past mileage transitions as the sample informs the construction of the ambiguity set. So, each RMDP is one example of an SDF for the bus replacement problem. We consider many RMDPs with varying levels of robustness and thus analyze a whole class of SDFs.\\

Statistical decision theory provides the framework to compare the performance of alternative decision functions $\delta \in \Gamma$. The utility achieved by any $\delta$ is a random variable before realizing $\psi$. Thus, Wald.1950 suggests measuring the performance of $\delta$ at all possible parametrizations $\theta$ by computing the expected utility with respect to its induced sampling distribution $P_{\theta}$:

align*[align* omitted — 161 chars of source]

In general, no single decision function yields the highest expected utility for all possible parameterizations. In this case, determining the best decision function $\delta^*$ is not straightforward. Still, decision theory proposes various criteria Gilboa.2009,Marinacci.2015 to aggregate the performance of a decision function at all possible parameterizations. At the most fundamental level, any decision function is admissible if another function does not exist, whose expected utility is always at least as high. In most cases, several decision functions are admissible, and thus additional optimality criteria are needed. Our analysis explores three of the most common decision criteria: (1) maximin, (2) minimax regret, and (3) subjective Bayes.\\

Following the maximin decision criterium Wald.1950, we determine the optimal decision function by computing the minimum expected performance for each decision function over all points in the parameter space. We then choose the one with the highest minimum performance. Stated concisely,

align*[align* omitted — 167 chars of source]

For the minimax regret criterion Niehans.1948, we compute the maximum regret for each decision function over all points in the parameter space. The regret of choosing a decision function for any realization of $\theta$ is the difference between the maximum possible performance, where the true parameterization informs the decision, and its actual performance. We then select the decision function with the lowest maximum regret. Thus, the minimax regret criterion solves:

align*[align* omitted — 272 chars of source]

Subjective Bayes Savage.1954 requires a subjective probability distribution $f_{\theta}$ over the parameter space. Then, we select the decision function with the highest expected subjective utility:

align*[align* omitted — 170 chars of source]

Urn example

We now illustrate the key ideas that allow us to compare as-if and robust decision-making using statistical decision theory with an urn example. As in our empirical application, we study a whole class of statistical decision functions. We first compare the performance of two distinct alternatives and then determine the optimal function within the class.\\

We consider an urn with black $b$ and white $w$ balls where the true share of black balls $\theta_0$ is unknown. In this example, the action constitutes a guess $\tilde{\theta}$ about $\theta_0$ after drawing a fixed number of $n$ balls at random with replacement. The parameter and action space both correspond to the unit interval $\Theta = \mathcal{A} = [0, 1]$. \\

If the guess matches the true share, we receive a payment of one. On the other hand, the payment is reduced by the squared error in case of a discrepancy. Thus, the consequence function takes the following form:

align*[align* omitted — 81 chars of source]

Going forward, we assume a linear utility function and directly refer to the monetary consequences of a guess as its utility. The sample space is $\Psi = \{b, w\}^n$ where a sequence $(b, w, b, \hdots, b)$ of length $n$ is a typical realization of $\psi$. The observed number of black balls $r$ among the $n$ draws in a given sample $\psi$ provides a signal about $\theta_0$. The sampling distribution for the possible number of black balls $R$ takes the form of a probability mass function (PMF):

align*[align* omitted — 115 chars of source]

Any function $\delta: \{b, w\}^n \mapsto [0, 1]$ that maps the number of black draws to the unit interval is a possible statistical decision function.\\

We focus on the following class of decision functions $\delta \in \Gamma$, where each $\lambda$ indexes a particular decision function:

align*[align* omitted — 161 chars of source]

The empirical share of black balls in the sample $r / n$ provides the point estimate $\hat{\theta}$. The decision functions in $\Gamma$ specify the guess as a weighted average between the point estimate and the midpoint of the parameter space. The larger $\lambda$ is, the more weight is put on the point estimate. At the extremes, the guess is either the point estimate ($\lambda = 1$) itself or fixed at $0.5$ ($\lambda = 0$).\\

We begin by comparing the performance of the two decision functions with $\lambda = 1$ and $\lambda = 0.9$. We refer to the former as the as-if decision function (ADF), as it announces the point estimate as if it is the true parameter. For reasons that will later become clear, we identify $\lambda=0.9$ as the robust decision function (RDF). Following Wald.1950, we evaluate their relative performance by aggregating the vector of expected payoffs over the unit interval using the different decision-theoretic criteria. We set the number of draws $n$ to $50$.\\

Figure (ref) shows the sampling distribution of the number of black balls $R$ and the associated payoff of following the two decision functions for each possible draw. The true, but unknown, share in this example is $40\%$, i.e. $\theta_0 = 0.4$. The RDF outperforms the ADF for realizations of the point estimates smaller than its true value due to the shift towards $0.5$. At the same time, the ADF leads to a higher payoff at the center of the distribution.\\

figure[figure omitted — 177 chars of source]

\FloatBarrier

Figure (ref) shows the expected payoff for varying shares $\theta$ of black balls in the urn. On the left, we show the expected payoff at two selected points. While the ADF performs better than the RDF at $\theta = 0.1$, the opposite is true at $\theta = 0.4$. Thus, both decision functions are admissible, as neither outperforms the other for all possible true shares. On the right, we trace the expected payoff of both functions over the whole parameter space. Although the RDF outperforms the ADF for shares in the center of the parameter space, it performs worse at the boundaries. Overall, the performance of the RDF is more balanced across the whole parameter space, which motivates its name.

figure[figure omitted — 314 chars of source]

\FloatBarrier

Figure (ref) ranks the two functions according to different decision-theoretic criteria. Both decision functions have their lowest expected payoff at $\theta = 0.5$. As the RDF outperforms its ADF alternative at that point, the RDF is preferred based on the maximin and minimax regret criteria. The maximin and minimax regret criteria are identical in this setting, as the payoff at the true share is constant across the parameter space. Using the subjective Bayes criterion with a uniform prior, we select the ADF, as its better performance at the boundaries of the parameter space is enough to offset its worse performance in the center.

figure[figure omitted — 164 chars of source]

\FloatBarrier

Returning to the whole set of decision functions, we can construct the optimal statistical decision function in $\Gamma$ for the alternative criteria by varying $\lambda$ to maximize the relevant performance measure. For example, Figure (ref) shows the minimum and the uniformly weighted performance for varying $\lambda$.

figure[figure omitted — 470 chars of source]

\FloatBarrier

Neither of our two decision functions analyzed earlier turns out to be optimal, as $\lambda^*_{\text{Bayes}} \approx 0.96$ and $\lambda^*_{\text{Maximin}} \approx 0.87$. Overall, the performance measure is more sensitive to the choice of $\lambda$ under the maximin criterion than under subjective Bayes.\\

In summary, the urn example illustrates the performance comparison of alternative decision functions over the whole parameter space. It shows how to construct an optimal decision function within a class for alternative decision-theoretic criteria. Next, we move to the more involved setting of a sequential dynamic decision problem with ambiguous transitions that we analyze in our application.

\FloatBarrier

Data-driven robust Markov decision problem

We now outline the framework of an RMDP for the analysis of sequential decision-making in light of model ambiguity. From the perspective of statistical decision theory, any RMDP with a fixed level of robustness is a statistical decision function. Once a sample of transitions is available, we construct the ambiguity set of a given size and solve the RMDP for a robust decision rule.\\

We first present the general setup of an RMDP and discuss the construction of the ambiguity set. We then turn to the solution approach and describe our decision-theoretic analysis to determine the optimal level of robustness. Throughout, we address the new challenges of analyzing an RMPD as opposed to a standard MDP. In line with our application, we discuss an infinite horizon model in discrete time, stationary utility and transition probabilities, and discrete states and actions.\footnote{See Puterman.1994 for a textbook introduction to the standard MDP and Rust.1994 for a review of MDPs in economics and structural estimation.}\\

We focus our exposition on ambiguity in the transition dynamics of the Markov decision process. We do not address uncertainty about the parameters of the reward functions. Although our central insight to use statistical decision theory to determine the optimal level of robustness is also relevant for the parameters of the reward functions, we do not address uncertainty pertaining to these parameters, as each setting introduces its unique computational challenges Mannor.2019.

\FloatBarrier

Setting

We consider the following decision problem. At time $t = 0, 1, 2, \hdots$ a decision-maker observes the state of their environment $s_t \in \mathcal{S}$ and chooses an action $a_t$ from the set of admissible actions $\mathcal{A}$. The decision has two consequences. It creates an immediate utility $u(s_t, a_t)$, and the environment evolves to a new state $s_{t+1}$. The transition from $s_t$ to $s_{t+1}$ is affected by the action, and governed by a transition probability distribution $p(s_t, a_t)$.\\

Decision-makers take the future consequences of the current action into account. While a decision rule $d_t$ specifies the planned action for all possible states within period $t$, a policy $\pi =\{ d_0, d_1, d_2, \hdots \}$ is a collection of decision rules and specifies all planned actions for all time periods.\\

Figure (ref) depicts the timing of events in the decision problem. At the beginning of period $t$, a decision-maker learns about the utility of each alternative, chooses one according to the decision rule $d_t$, and receives its immediate utility. Then, the state evolves from $s_t$ to $s_{t+1}$, and the process repeats itself in $t + 1$.\\

figure[figure omitted — 2,585 chars of source]

\FloatBarrier

In a standard Markov decision process (MDP), a single transition probability distribution $p(s_t, a_t)$ is associated with each state-action pair. This distribution is assumed to be known, and thus the MDP incorporates risk only. In an RMDP, there is a whole set of distributions associated with each state-action pair collected in an ambiguity set $p(s_t, a_t) \in \mathcal{P}(s_t, a_t)$. For a particular RMDP, the ambiguity set is assumed to be known, and thus the RMDP incorporates risk for a given distribution and ambiguity about the true distribution.\\

In a standard MDP, the objective of a decision-maker in state $s_t$ at time $t$ is to choose the optimal policy $\pi^*$ from the set of all possible policies $\Pi$ that maximizes their expected total discounted utility $\tilde{v_t}^{\pi^*}(s_t)$ as formalized in Equation ((ref)):

align[align omitted — 176 chars of source]

The exponential discount factor $\delta$ parameterizes a taste for immediate over future utilities. The superscript of the expectation emphasizes that each policy induces a different probability distribution over sequences of possible futures. As long as transition probabilities used to construct the policy are in fact correct, the standard value function $\tilde{v_t}^{\pi^*}(s_t)$ measures the performance of the optimal policy.\\

In an RMDP, the goal is to implement an optimal policy that maximizes the expected total discounted utility under a worst-case scenario. Given the ambiguity about the transition dynamics, a policy induces a whole set of probabilities over sequences of possible future utilities $\mathcal{F}^\pi$, and the worst-case realization determines its ranking. The formal representation of the decision-maker's objective is Equation ((ref)):

align[align omitted — 241 chars of source]

We consider a setting where historical data provides information about the transition dynamics. In the data-driven standard MDP, the empirical probabilities $\hat{p}(s_t, a_t)$ serve as a plug-in for the truth, and the solution of the MDP provides an as-if decision rule. In a data-driven RMDP, the empirical probabilities are used to construct the ambiguity sets for the transitions, and the solution of the RMDP provides a robust decision rule.\\

We follow Ben-Tal.2013 and create the ambiguity sets using statistical hypothesis testing. We restrict attention to distributions we cannot reject with a certain level of confidence $\omega \in [0, 1]$ around the empirical probabilities and collect them in an estimated ambiguity set $\hat{\mathcal{P}}(s_t, a_t; \omega)$. Different values of $\omega$ result in different RMDPs, each with its own statistical decision function. Two special cases stand out. First, if $\omega = 0$, then a decision-maker treats the empirical probabilities as if they are correct. This case captures the notion of as-if decision-making. Second, for $\omega = 1$, a robust decision-maker considers the worst-case scenario over the whole probability simplex at each state-action pair when constructing the optimal policy.

\FloatBarrier

Solution

In a standard MDP, the objective is to maximize the expected total discounted utility as formalized in Equation ((ref)). This requires evaluating the performance of all policies based on all possible sequences of utilities and the probability that each occurs. Fortunately, the stationary Markovian structure of the problem implies that the future looks the same whether the decision-maker is in state $s$ at time $t$ or any other point in time. The only variable that determines the value to the decision-maker is the current state $s$. Thus, the optimal policy is stationary as well Blackwell.1965, and the same decision rule is used in every period. The value function is independent of time and of the solution to the following Bellman equation:

align[align omitted — 190 chars of source]

The as-if decision rule is recovered from Equation ((ref)) by finding the value $a \in \mathcal{A}$ that attains a maximum for each $s \in \mathcal{S}$.\\

Let $\mathbb{V}$ denote the set of all bounded real value functions on $\mathcal{S}$. Then, the Bellman operator $\tilde{\Lambda} : \mathbb{V} \rightarrow \mathbb{V}$ is defined as follows: For all $w\in\mathbb{V}$

align[align omitted — 219 chars of source]

Under mild conditions, $\tilde{\Lambda}$ is a contraction mapping and allows to compute the value function $\tilde{v}(\cdot)$ as its unique fixed point Denardo.1967.\\

For an RMDP, where transition probabilities are ambiguous, the contraction mapping property of the Bellman operator and the optimality of a stationary deterministic Markovian decision rule both require the assumption of rectangularity of $\mathcal{F}^\pi$ Iyengar.2005,Nilim.2005. As the realization of any particular distribution in a state-action pair does not affect future realizations, rectangularity is a form of an independence assumption. The uncertainty is uncoupled across states and actions. This approach rules out any kind of learning about future ambiguity from past experiences due to, for example, a common source of uncertainty across states. While restrictive, most applications rely on the rectangularity assumption, as general notions of coupled uncertainties are intractable Wiesemann.2013.\footnote{See Mannor.2016 and Goyal.2020 for recent attempts to introduce milder rectangularity conditions.}\\

We now develop the formal definition of rectangularity. Let $\mathcal{M}(\mathcal{S})$ denote the set of all probability distributions on $\mathcal{S}$. Then, the set of all conditional transition probability distributions associated with any decision rule $d$ is given by:

align*[align* omitted — 195 chars of source]

For every state $s \in \mathcal{S}$, the next state can be determined by any $p \in \hat{\mathcal{P}}(s, d(s); \omega)$.\\

A policy $\pi$ now induces a set of probability distributions $\mathcal{F}^\pi$ on the set of all possible histories $\mathcal{H}$. Any particular history $h = (s_0, a_0, s_1, a_1, \hdots)$ can be the result of many possible combinations of transition probabilities. Rectangularity imposes a structure on the combination possibilities.

AssumptionRectangularity The set $\mathcal{F}^\pi$ of probability distributions associated with a policy $\pi$ is given by \begin{align*} \mathcal{F}^\pi & = \bigg\{\mathbf{P} \mid \forall\, h\in \mathcal{H}:\, \mathbf{P}(h) =\prod^{\infty}_{t = 0} p(s_{t+1}|s_t, a_t), with p(s_t, a_t) \in \hat{\mathcal{P}}(s_t, d_t(s_t); \omega) for t = 0, 1, \hdots \bigg\} \\ &= \mathcal{F}^{d_0} \times \mathcal{F}^{d_1} \times \mathcal{F}^{d_2} \times \hdots = \prod^{\infty}_{t = 0} {F}^{d_t}, \end{align*} where the notation simply denotes that each element in $\mathcal{F}^\pi$ is a product of $p \in\mathcal{F}^{d_t}$, and vice versa Iyengar.2005.

Assumption (ref) formalizes the idea that ambiguity about the transition probability distribution is uncoupled across states and time. All elements of the ambiguity sets can be freely combined to generate a particular history.\\

The objective when facing ambiguity is to implement a policy $\pi^*$ that maximizes the expected total discounted utility under a worst-case scenario as presented in Equation ((ref)). Under the rectangularity assumption, the decision-maker faces the same uncertainty, whether he is in state $s$ at time $t$ or any other point in time. Thus, the value function is independent of time and solely depends on the current state $s$. It is the solution to the robust Bellman equation ((ref)), where the future value is evaluated using the worst-case element in the ambiguity set Iyengar.2005:

align[align omitted — 227 chars of source]

The robust decision rule is recovered from Equation ((ref)) by finding the value $a \in \mathcal{A}$ that attains a maximum for each $s \in \mathcal{S}$ under the worst-case scenario for all distributions in the ambiguity set.\\

The robust Bellman operator on $\mathbb{V}$ follows directly: For all $w\in\mathbb{V}$

align[align omitted — 242 chars of source]

Algorithm (ref) allows solving the RMDP by a robust version of the value iteration algorithm where $\kappa$ denotes a convergence threshold. The calculation of future values under the worst-case scenario is the key difference to the standard approach.

algorithm[algorithm omitted — 765 chars of source]

\FloatBarrier

Evaluation

The solution of an RMDP is tailored to the simultaneous worst-case realization of all distributions in all ambiguity sets. Although this conservative approach ensures a minimum performance over all distributions in the set, the performance of the robust decision rule in all other cases is disregarded. This indifference introduces a trade-off when determining the size of the ambiguity set Delage.2010. The larger the set, the more scenarios for which a minimum performance is ensured. However, the robust rule's general performance suffers. This trade-off is particularly pronounced when the actual structure of the decision problem exhibits coupled uncertainties that are ignored in the construction of the robust rule to ensure its computational tractability.\\

Statistical decision theory allows us to navigate the trade-off and determine the optimal level of robustness. Each RMDP is a different statistical decision function, and we consider the whole class of statistical decision functions each indexed by $\omega\in [0 , 1]$. Adapting our urn example from earlier to accommodate setting of a data-driven RMDP, the parameter space corresponds to the set of transition probability distributions $\mathcal{L(\mathcal{S}, \mathcal{A})} = \{p:\mathcal{S} \times \mathcal{A} \rightarrow \mathcal{M}(\mathcal{S})\}$. We observe data on the transition probabilities and measure the actual performance of a robust decision rule $\eta(\hat{p}; p_0, \omega)$ as the discounted sum of utilities, which depends on the estimate of the transition probabilities $\hat{p}$, the true underlying probabilities $p_0$, and the confidence level $\omega$ used to construct the robust decision function. The standard decision-theoretic criteria translate to this setting as follows:

align*[align* omitted — 730 chars of source]

Note that even for genuinely uncoupled uncertainties, the maximin criterion does not automatically select the most robust statistical decision function ($\omega = 1$). This particular decision function is based on the worst-case scenario over the full probability simplex at each state-action pair. In fact, the worst-case decision function might not be admissible in particular settings where it is weakly dominated by the as-if (or some other) decision function. Suppose, for example, the true distribution corresponds to the worst-case distributions. In this case, the distribution of sampled transitions is degenerate, as the worst-case scenario at each state-action pair is the certain transition to the state with the lowest future value Nilim.2005. Thus, the as-if and worst-case decision functions share the same performance. For all other true distributions, the as-if decision function may very well outperform the worst-case decision function if the sampled data is sufficiently informative.

\FloatBarrier

Bus replacement problem

We now study robust decision-making in the seminal bus replacement problem. First, we discuss the general setting and the details of the computational implementation. Second, we conduct an ex-post analysis of robust decision rules constructed for the observed sample of mileage transitions analyzed in Rust.1987. Third, considering the situation before any data is realized, we conduct an ex-ante analysis of robust decision functions with varying levels of robustness over the whole probability simplex, which allows us to determine the optimal level of robustness using statistical decision theory.

\FloatBarrier

Setting

The bus replacement model is set up as a regenerative optimal stopping problem Chow.1971. It is motivated by the sequential decision problem of a maintenance manager, Harold Zurcher, for a fleet of buses. He makes repeated decisions about their maintenance to maximize the expected total discounted utility under a worst-case scenario. Each month $t$, a bus arrives at the bus depot in state $s_t = (x_t, \epsilon_t)$ described by its mileage since the last engine replacement $x_t$ and other signs of wear and tear $\epsilon_t$. He faces the decision to either conduct a complete engine replacement $(a_t = 1)$ or perform basic maintenance work $(a_t = 0)$. The cost of maintenance $c(x_t)$ increases with the mileage state, while the cost of replacement $RC$ remains constant. In the case of an engine replacement, the mileage state is reset to zero. Note that we do not attempt to describe Harold Zurcher's decision-making process. Instead, we are interested in how a generic decision-maker should make decisions in this setting.\\

The immediate utility of each action is given by:

align*[align* omitted — 145 chars of source]

Decisions are made in light of uncertainty about next month's state variables captured by their conditional distribution $p(x_t, \epsilon_t, a_t)$.\\

Although in this framework, the utility and consequently the value function is finite in each state, they are not uniformly bounded. This property, however, is a crucial assumption for the results of Blackwell.1965 and Denardo.1967 on the contraction property of the Bellman operator and the stationarity of the optimal policy in the standard MDP setting. For the original as-if analysis, Rust.1988 circumvents this problem by imposing conditional independence between the observable and unobservable state variables, i.e. $p(x_{t+1}, \epsilon_{t+1}| x_t, \epsilon_t, a_t) = p(x_{t+1}| x_t, a_t)\thinspace q(\epsilon_{t+1}|x_{t+1})$, and assuming that the unobservables $\epsilon_t(a_t)$ are independent and identically distributed according to an extreme value distribution with mean zero and scale parameter one. These two assumptions, together with the additive separability between the observed and unobserved state variables in the immediate utilities, ensure that the expectation of the next period's value function is independent of the time. The regenerative structure of the process implies that the transition probabilities in case of replacement in any mileage state correspond to the probabilities of maintenance in the zero mileage state. Therefore, the expected value function is the unique fixed point of a contraction mapping on the reduced space of mileage states only. In addition, the conditional choice probabilities $P(a_t | x_t)$ have a closed-form solution McFadden.1973. We build on these results and extend them to our robust setting with ambiguous transition dynamics. The proof is available in Appendix (ref).\\

In the analysis of the original bus replacement problem, the distribution of the monthly mileage transitions are estimated in a first step and used as plug-in components for the subsequent analysis. We extend the original setup and explicitly account for the ambiguity in the estimation. Following the arguments on the regenerative structure of the process above, we incorporate ambiguity in the RMDP with ambiguity sets conditional on the mileage states $x$ only. We construct ambiguity sets $\hat{\mathcal{P}}(x; \omega)$ based on the Kullback-Leibler divergence $D_{KL}$ Kullback.1951 that are statistically meaningful, computationally tractable, and anchored in empirical estimates $\hat{p}(x)$.\\

Our ambiguity set takes the following form for each mileage state $x$:

align*[align* omitted — 221 chars of source]

where $J_x = \{j_1,\thinspace \dots,\thinspace j_{|J_x|}\}$ denotes the set of all states that have an estimated non-zero probability to be reached from $x$, $\mathring{\Delta}_{|J_{x}|} = \{p \in \mathbb{R}^{|J_{x}|}\,|\, p_i > 0 \text{ for all } i=1,\dots, |J_x| \text{ and } \sum_{i=1}^{|J_x|} p_i = 1\}$ is the interior of the $(|J_{x}| - 1)$ - dimensional probability simplex, and $\rho_{x}(\omega)$ captures the size of the set for the state $x$ with a given level of confidence $\omega$.\\

Iyengar.2002 and Ben-Tal.2013 provide the statistical foundation to calibrate $\rho_x(\omega)$ such that the true (but unknown) distribution $p_0$ is contained within the ambiguity set for a given level of confidence $\omega$. Let $\chi^2_{df}$ denote a chi-squared random variable with $df$ degrees of freedom, and let $F_{df}(\cdot)$ denote its cumulative distribution function with inverse $F^{-1}_{df}(\cdot)$. Then, the following approximate relationship holds as the number of observations $N_x$ for state $x$ tends to infinity Pardo.2005:

align*[align* omitted — 206 chars of source]

We can therefore calibrate the size of the ambiguity set based on the following relationship:

align[align omitted — 93 chars of source]

We use \citetalias{Rust.1987} original data to inform our computational experiments. His data consists of monthly odometer readings $x_t$ and engine replacement decisions $a_t$ for 162 buses. The fleet consists of eight groups that differ in their manufacturer and model. We focus on the fourth group of 37 buses with a total of 4,292 monthly observations. We discretize mileage into $78$ equally spaced bins of length $5,000$ and set the discount factor to $\delta=0.9999$.\\

Figure (ref) highlights the limited information about the true distribution of mileage utilization. It shows the number of observations available to estimate next month's utilization for different levels of accumulated mileage. While there are more than 1,150 observations on buses with less than 50,000 miles, there are only about 220 with more than 300,000.\\

figure[figure omitted — 180 chars of source]

We analyze a specific example of \citetalias{Rust.1987} bus replacement problem. We do not use his reported estimates of the maintenance and replacement costs. Given these estimates, decisions are mainly driven by the unobserved state variable $\epsilon_t$, and so ambiguity about the evolution of the observed state variable $x_t$ does not have a substantial effect on decisions. We ensure that a bus's accumulated mileage has a considerable impact on the timing of engine replacements by increasing the maintenance and replacement costs compared to their empirical estimates. Thus, we specify the following cost function $c(x_t) = 0.4\, x_t $ and set the replacement costs $RC$ to 50.\\

We solve the model using a modified version of the original nested fixed point algorithm (NFXP) Rust.1988, and we determine the worst-case transition probabilities in each successive approximation of the fixed point. Given the size of the ambiguity set, we can determine the worst-case probabilities as the solution to a one-dimensional convex optimization problem Iyengar.2005,Nilim.2005.\footnote{The core routines are implemented in our group's ruspy.2020 and robupy.2020 software packages and are publicly available.}

Ex-post analysis

We first study as-if and robust decision rules for \citetalias{Rust.1987} observed sample of mileage transitions. We present the estimated transition probabilities and the corresponding worst-case distributions. We then explore alternative decision rules based on several RMDPs, outline the resulting differences in maintenance decisions, and evaluate their relative performance under different scenarios.\\

Figure (ref) shows the point estimates $\hat{p}$ for the transition probabilities of monthly mileage usage. We pool all 4,292 observations to estimate this distribution by maximum likelihood, and thus the probability of the next period's mileage utilization is the same for each state $x_t$. We only observe increases of at most $J = 3$ grid points per month. For about 60% of the sample, monthly bus utilization is between 5,000 and 10,000 miles. Very high usage of more than 10,000 miles amounts to only 1.2%.\\

figure[figure omitted — 182 chars of source]

The confidence level $\omega$ and the available number of observations $N_x$ determine the size of the ambiguity set as outlined in Equation ((ref)). From now on, we mimic state-specific ambiguity sets by constructing them based on the average number of $55$ observations per state. Note that while the estimated distribution is the same for all mileage levels, its worst-case realization is not. However, there are only minor differences across mileage levels, so we focus our following discussion on a bus with an odometer reading of 75,000.\\

Figure (ref) shows the transition probabilities for different sizes of the ambiguity set. We vary the confidence level for the whole number of observations $(N_x = 55)$ on the left, while on the right, the level of confidence remains fixed $(\omega=0.95)$, and we cut the number of observations roughly in half. The larger the ambiguity set, the more probability is attached to higher mileage utilization, resulting in higher costs overall. For example, while the probability of mileage increases of 10,000 or more is an infrequent occurrence in the data, its probability increases first to 1.7%. It then doubles to 2.5% as we increase the confidence level. When only about half the data is available, this probability increases even further to 3.2%.\\

figure[figure omitted — 383 chars of source]

The decision-maker chooses whether to perform regular maintenance work on a bus or replace its complete engine each month. The assumed transition probabilities correspond to their worst-case transitions within the ambiguity set. As a result, any differences between the as-if and worst-case distributions translate into different maintenance decisions.\\

Figure (ref) shows the maintenance probabilities for different levels of accumulated mileage and alternative rules. Overall, the maintenance probability decreases with accumulated mileage, as maintenance becomes more costly than an engine replacement. Robust rules result in a higher probability of maintenance compared to the as-if decision rule. Under the worst-case transitions, a bus is more likely to experience higher usage during the period. As the cost of maintenance is determined by the mileage level at the beginning of the period, maintenance becomes more attractive. For example, again considering a bus with 75,000 miles, the as-if maintenance probability is 25%, while it is 33% $(\omega=0.50)$ and 43% $(\omega=0.95)$ following the robust rule.\\

figure[figure omitted — 176 chars of source]

\FloatBarrier

To gain further insights into the differences between the as-if and robust decisions, we simulate a fleet of 1,000 buses for 100,000 months under the alternative decision rules.\\

Figure (ref) shows the level of accumulated mileage over time for a single bus under different decision rules. It clarifies our simulation setup, where we apply different decision rules to the same bus. The realizations of observed transitions and unobserved signs of wear and tear remain the same. The bus accumulates more and more mileage until Harold Zurcher replaces the complete engine and the odometer is reset to zero. The first replacement happens after 20 months at 60,000 miles following the as-if decision rule, while it is delayed for another four months under the robust alternative $(\omega = 0.95)$. As its timing differs, the odometer readings will start to diverge after 20 months, even though monthly utilization remains the same.\\

figure[figure omitted — 198 chars of source]

\FloatBarrier

We now evaluate the as-if and robust decisions at the boundary of the ambiguity set. We measure the performance of the alternative decision rules based on their total discounted utility under different assumed and actual mileage transitions.\\

Figure (ref) shows the performance of the as-if decision rule over time when the worst-case distribution for a confidence level of 0.95 governs the actual transitions. It illustrates the sensitivity of the as-if decision rule to perturbations in the transition probabilities. The solid line corresponds to its expected long-run performance without misspecification of the decision problem, while the dashed line indicates its observed performance. After about 20,000 months, it accumulates the expected long-run average cost and performs about 14% worse overall.\\

figure[figure omitted — 195 chars of source]

\FloatBarrier

Figure (ref) shows the average difference in performance between the as-if and two robust decision rules with confidence levels of $0.50$ and $0.95$, respectively. The actual transitions follow the worst-case distribution with varying $\omega$. A positive value indicates that the robust decision rule outperforms the as-if decision rule. In the absence of any misspecification, the as-if decision rule must defeat any other decision rule. The same is true for the robust decision rule when the actual transitions are governed by the same $\omega$ used for their construction. Nevertheless, the as-if decision rule continues to outperform both robust decisions for moderate levels of $\omega$. For worst-case distributions with $\omega$ larger than 0.2, the first robust decision rule $(\omega=0.5)$ starts to beat the as-if decision rule. For the other robust decision rule $(\omega=0.95)$, the same is true for worst-case transitions of $\omega$ equal to 0.5.\\

figure[figure omitted — 396 chars of source]

\FloatBarrier

Ex-ante analysis

We now turn to the situation before any data are realized. We evaluate the ex-ante performance of as-if and robust decision functions over the whole probability simplex and determine the optimal level of robustness.\\

We operationalize our analysis as follows. In line with \citetalias{Rust.1987} assumption on the distribution of the mileage utilization, we specify a uniform grid with $0.1$ increments over the interior of the two-dimensional probability simplex $\mathring{\Delta}_3$. At each grid point, we draw 100 samples of 55 random mileage utilizations. For each sample, we solve several robust decision functions for a grid of $\omega = \{0.0, 0.1, \hdots, 1.0\}$ using the estimated transition probabilities. Note that the uncertainties are coupled across states, as the same underlying probability creates the sample of bus utilizations. Thus, the rectangularity assumption does not reflect the economic environment. However, we still impose it when constructing the robust decision functions to ensure tractability. We then simulate the implied decision rules' actual performance and compute their expected performance by averaging across the 100 runs for each grid point. Using this information, we measure the performance of the different decisions based on the maximin criterion, the minimax regret rule, and the subjective Bayes approach using a uniform prior.\\

In Figure (ref) we illustrate the differences in expected performance between a robust decision function $(\omega=0.1)$ and the as-if alternative over the probability simplex.

figure[figure omitted — 182 chars of source]

\FloatBarrier In the gray areas, the as-if decisions outperform the robust alternative based on their expected performance. The opposite is true for the black areas: robust decisions perform very well when the true probability of mileage increases of $5,000$ per month is high and when the true probability of increases amounting to $10,000$ is low. Otherwise, the as-if decisions outperform the robust alternative. Thus, no rule dominates the other, and it is essential to aggregate the performance over the whole probability simplex using decision theory before settling on a decision rule.\\

Figure (ref) ranks the as-if decisions against selected robust alternatives for the different performance criteria.

figure[figure omitted — 171 chars of source]

\FloatBarrier Based on a maximin criterion, decision functions rank higher when the confidence level $\omega$ used to construct them is greater. The decision function with $\omega = 0.3$ comes in first, while as-if decisions rank last. Thus, decision-makers can improve their worst-case outcomes by adopting a robust decision function. However, this comes at a cost, as indicated by the improved rankings for the as-if decision function as we move to different criteria. As-if decisions move to second place for minimax regret. The as-if decision rule comes in first when we aggregate performance across all states using a subjective Bayes approach with a uniform prior. Thus, our approach clarifies the trade-offs involved when choosing a particular decision function for decision-making.\\

We now determine the optimal size of the ambiguity set $\omega^*$ for each decision-theoretic criterium. Figure (ref) shows the minimum performance of the decision functions for varying levels of $\omega$ normalized between zero and one. Among all decision functions, robust decisions with $\omega=0.36$ have the highest minimum performance. They thus strike a balance between the conservatism of the worst-case approach and the protection against unfavorable transition probabilities. Based on the maximin criterion, the as-if decision function performs worst.

figure[figure omitted — 181 chars of source]

\FloatBarrier

The minimax regret criterion leads to a slightly reduced level of $\omega^*=0.1$. As-if decisions are optimal based on the subjective Bayes criterion with a uniform prior.

\FloatBarrier

Conclusion

Economists often estimate economic models on data and use the point estimates as a stand-in for the truth when studying the model's implications for optimal decision-making. This practice ignores model ambiguity, exposes the decision problem to misspecification, and ultimately leads to post-decision disappointment. We develop a framework to explore, evaluate, and optimize robust decision rules that explicitly account for the uncertainty in the estimation using statistical decision theory. We show how to operationalize our analysis by studying robust decisions in a stochastic dynamic investment model in which a decision-maker directly accounts for uncertainty in the model's transition dynamics.\\

As our core contribution, we combine ideas from data-driven robustness optimization Bertsimas.2018, robust Markov decision processes Ben-Tal.2009, and statistical decision theory Berger.2010 to optimize robustness in decision-making. This insight transfers directly to many other settings. For example, the COVID-19 pandemic provides a timely example of economists informing policy-making by using highly parameterized models in light of ubiquitous uncertainties Avery.2020. When analyzing these models, economists treat many of their parameters as if they are known. However, their actual values are uncertain, as they are often estimated based on external data sources. Using statistical decision theory, our research illustrates how to conduct robust policy-making and to evaluate its relative performance against policies that ignore uncertainty. Such an approach promotes a sound decision-making process, as it provides decision-makers with the tools to systematically navigate the uncertainties they face Berger.2021.

thebibliography\bibitem[Adda et al., 2017]{Adda.2017} Adda, J., Dustmann, C., and Stevens, K. (2017). \newblock The career costs of children. \newblock {\em Journal of Political Economy}, 125(2):293--337. \bibitem[Alismail et al., 2018]{Alismail.2018} Alismail, F., Xiong, P., and Singh, C. (2018). \newblock Optimal wind farm allocation in multi-area power systems using distributionally robust optimization approach. \newblock {\em IEEE Transactions on Power Systems}, 33(1):536--544. \bibitem[Andrews et al., 2017]{Andrews.2017} Andrews, I., Gentzkow, M., and Shapiro, J. M. (2017). \newblock Measuring the sensitivity of parameter estimates to estimation moments. \newblock {\em The Quarterly Journal of Economics}, 132(4):1553--1592. \bibitem[Andrews et al., 2020]{Andrews.2020} Andrews, I., Gentzkow, M., and Shapiro, J. M. (2020). \newblock On the informativeness of descriptive statistics for structural estimates. \newblock {\em Econometrica}, 88(6):2231--2258. \bibitem[Armstrong and Koles{\'a}r, 2021]{Armstrong.2021} Armstrong, T. B. and Koles{\'a}r, M. (2021). \newblock Sensitivity analysis using approximate moment condition models. \newblock {\em Quantitative Economics}, 12(1):77--108. \bibitem[Arrow, 1951]{Arrow.1951} Arrow, K. J. (1951). \newblock Alternative approaches to the theory of choice in risk-taking situations. \newblock {\em Econometrica}, 19(4):404--437. \bibitem[Avery et al., 2020]{Avery.2020} Avery, C., Bossert, W., Clark, A., Ellison, G., and Ellison, S. F. (2020). \newblock Policy implications of models of the spread of coronavirus: Perspectives and opportunities for economists. \newblock {\em NBER Working Paper}. \bibitem[Bagwell et al., 2021]{Bagwell.2021} Bagwell, K., Staiger, R. W., and Yurukoglu, A. (2021). \newblock Quantitative analysis of multiparty tariff negotiations. \newblock {\em Econometrica}, 89(4):1595--1631. \bibitem[Barnett et al., 2020]{Barnett.2020} Barnett, M., Brock, W., and Hansen, L. P. (2020). \newblock Pricing uncertainty induced by climate change. \newblock {\em The Review of Financial Studies}, 33(3):1024--1066. \bibitem[Ben-Tal et al., 2013]{Ben-Tal.2013} Ben-Tal, A., den Hertog, D., De Waegenaere, A., Melenberg, B., and Rennen, G. (2013). \newblock Robust solutions of optimization problems affected by uncertain probabilities. \newblock {\em Management Science}, 59(2):341--357. \bibitem[Ben-Tal et al., 2009]{Ben-Tal.2009} Ben-Tal, A., {El Ghaoui}, L., and Nemirowski, A. (2009). \newblock {\em Robust optimization}. \newblock Princeton University Press, Princeton, NJ. \bibitem[Berger, 2010]{Berger.2010} Berger, J. O. (2010). \newblock {\em Statistical decision theory and {B}ayesian analysis}. \newblock Springer, New York City, NY. \bibitem[Berger et al., 2021]{Berger.2021} Berger, L., Berger, N., Bosetti, V., Gilboa, I., Hansen, L. P., Jarvis, C., Marinacci, M., and Smith, R. D. (2021). \newblock Rational policymaking during a pandemic. \newblock {\em Proceedings of the National Academy of Sciences}, 118(4). \bibitem[Bertsimas et al., 2018]{Bertsimas.2018} Bertsimas, D., Gupta, V., and Kallus, N. (2018). \newblock Data-driven robust optimization. \newblock {\em Mathematical Programming}, 167(2):235--292. \bibitem[Bertsimas and Thiele, 2006]{Bertsimas.2006} Bertsimas, D. and Thiele, A. (2006). \newblock Robust and data-driven optimization: Modern decision making under uncertainty. \newblock In Johnson, M. P., Norman, B., and Secomandi, N., editors, {\em Models, Methods, and Applications for Innovative Decision Making}, pages 95--122. INFORMS. \bibitem[Blackwell, 1965]{Blackwell.1965} Blackwell, D. (1965). \newblock Discounted dynamic programming. \newblock {\em The Annals of Mathematical Statistics}, 36(1):226--235. \bibitem[Blundell and Shephard, 2012]{Blundell.2012} Blundell, R. and Shephard, A. (2012). \newblock {Employment, hours of work and the optimal taxation of low-income families}. \newblock {\em The Review of Economic Studies}, 79(2):481--510. \bibitem[Bonhomme and Weidner, 2020]{Bonhomme.2020} Bonhomme, S. and Weidner, M. (2020). \newblock Minimizing sensitivity to model misspecification. \newblock {\em arXiv preprint arXiv:1807.02161}. \bibitem[Brown et al., 2012]{Brown.2012} Brown, D. B., Giorgi, E. D., and Sim, M. (2012). \newblock Aspirational preferences and their representation by risk measures. \newblock {\em Management Science}, 58(11):2095--2113. \bibitem[Chernozhukov et al., 2020]{Chernozhukov.2020} Chernozhukov, V., Escanciano, J. C., Ichimura, H., Newey, W. K., and Robins, J. M. (2020). \newblock Locally robust semiparametric estimation. \newblock {\em arXiv preprint arXiv:1608.00033}. \bibitem[Chow et al., 1971]{Chow.1971} Chow, Y. S., Robbins, H., and Siegmund, D. (1971). \newblock {\em Great expectations: The theory of optimal stopping}. \newblock Houghton Mifflin, Boston, MA. \bibitem[Christensen and Connault, 2019]{Christensen.2019} Christensen, T. and Connault, B. (2019). \newblock Counterfactual sensitivity and robustness. \newblock {\em arXiv preprint arXiv:1904.00989}. \bibitem[Delage and Mannor, 2010]{Delage.2010} Delage, E. and Mannor, S. (2010). \newblock Percentile optimization for {M}arkov decision processes with parameter uncertainty. \newblock {\em Operations Research}, 58(1):203--213. \bibitem[Denardo, 1967]{Denardo.1967} Denardo, E. V. (1967). \newblock Contraction mappings in the theory underlying dynamic programming. \newblock {\em {SIAM} Review}, 9(2):165--177. \bibitem[Eaton et al., 2011]{Eaton.2011} Eaton, J., Kortum, S., and Kramarz, F. (2011). \newblock An anatomy of international trade: Evidence from french firms. \newblock {\em Econometrica}, 79(5):1453--1498. \bibitem[Eisenhauer et al., 2021]{Eisenhauer.2021} Eisenhauer, P., Gabler, J., and Janys, L. (2021). \newblock Structural models for policy-making: {C}oping with parametric uncertainty. \newblock {\em arXiv preprint arXiv:2103.01115}, submitted. \bibitem[French and Song, 2014]{French.2014} French, E. and Song, J. (2014). \newblock The effect of disability insurance receipt on labor supply. \newblock {\em American Economic Journal: Economic Policy}, 6(2):291--337. \bibitem[Gilboa, 2009]{Gilboa.2009} Gilboa, I. (2009). \newblock {\em Theory of decision under uncertainty}. \newblock Cambridge University Press, New York City, NY. \bibitem[Goh et al., 2018]{Goh.2018} Goh, J., Bayati, M., Zenios, S. A., Singh, S., and Moore, D. (2018). \newblock Data uncertainty in {M}arkov chains: Application to cost-effectiveness analyses of medical innovations. \newblock {\em Operations Research}, 66(3):697--715. \bibitem[Gotoh et al., 2021]{Gotoh.2021} Gotoh, J.-y., Kim, M. J., and Lim, A. E. B. (2021). \newblock Calibration of distributionally robust empirical optimization models. \newblock {\em Operations Research}. \bibitem[Goyal and Grand-Clement, 2020]{Goyal.2020} Goyal, V. and Grand-Clement, J. (2020). \newblock Robust {M}arkov decision process: {B}eyond rectangularity. \newblock {\em arXiv preprint arXiv:1811.00215}. \bibitem[Hansen and Sargent, 2016]{Hansen.2016} Hansen, L. P. and Sargent, T. J. (2016). \newblock {\em Robustness}. \newblock Princeton University Press, Princeton, NJ. \bibitem[He et al., 2019]{He.2019} He, S., Sim, M., and Zhang, M. (2019). \newblock Data-driven patient scheduling in emergency departments: A hybrid robust-stochastic approach. \newblock {\em Management Science}, 65(9):4123--4140. \bibitem[Hirano and Porter, 2009]{Hirano.2009} Hirano, K. and Porter, J. R. (2009). \newblock Asymptotics for statistical treatment rules. \newblock {\em Econometrica}, 77(5):1683--1701. \bibitem[Honor{\'e} et al., 2020]{Honore.2020} Honor{\'e}, B., J{\o}rgensen, T., and de Paula, {\'A}. (2020). \newblock The informativeness of estimation moments. \newblock {\em Journal of Applied Econometrics}, 35(7):797--813. \bibitem[Horta\c{c}su et al., 2019]{Hortacsu.2019} Horta\c{c}su, A., Luco, F., Puller, S. L., and Zhu, D. (2019). \newblock Does strategic ability affect efficiency? {E}vidence from electricity markets. \newblock {\em American Economic Review}, 109(12):4302--42. \bibitem[Igami, 2017]{Igami.2017} Igami, M. (2017). \newblock Estimating the innovator's dilemma: Structural analysis of creative destruction in the hard disk drive industry, 1981--1998. \newblock {\em Journal of Political Economy}, 125(3):798--847. \bibitem[Ilut, 2020]{Ilut.2020} Ilut, C. (2020). \newblock Paralyzed by fear: Rigid and discrete pricing under demand uncertainty. \newblock {\em Econometrica}, 88(5):1899--1938. \bibitem[Iskhakov et al., 2016]{Iskhakov.2016} Iskhakov, F., Lee, J., Rust, J., Schjerning, B., and Seo, K. (2016). \newblock Comment on “{C}onstrained optimization approaches to estimation of structural models". \newblock {\em Econometrica}, 84(1):365--370. \bibitem[Iyengar, 2002]{Iyengar.2002} Iyengar, G. N. (2002). \newblock Robust dynamic programming. \newblock {\em CORC Tech Report}. \bibitem[Iyengar, 2005]{Iyengar.2005} Iyengar, G. N. (2005). \newblock Robust dynamic programming. \newblock {\em Mathematics of Operations Research}, 30(2):257--280. \bibitem[Jin et al., 2020]{Jin.2020} Jin, X., Luo, D., and Zeng, X. (2020). \newblock Tail risk and robust portfolio decisions. \newblock {\em Management Science}, 67(5):3254--3275. \bibitem[J{\o}rgensen, 2021]{Jorgensen.2021} J{\o}rgensen, T. H. (2021). \newblock Sensitivity to calibrated paramters. \newblock {\em Review of Economics and Statistics}, forthcoming. \bibitem[Kaufman et al., 2017]{Kaufman.2017} Kaufman, D. L., Schaefer, A. J., and Roberts, M. S. (2017). \newblock Living-donor liver transplantation timing under ambiguous health state transition probabilities. \newblock {\em SSRN Working Paper}. \bibitem[Keith and Ahner, 2021]{Keith.2021} Keith, A. J. and Ahner, D. K. (2021). \newblock A survey of decision making and optimization under uncertainty. \newblock {\em Annals of Operations Research}, 300(2):319--353. \bibitem[Kitigawa, 2018]{Kitigawa.2018} Kitigawa (2018). \newblock Who should be treated? empirical welfare maximization methods for treatment choice. \newblock {\em Econometrics}, 86(2):591--616. \bibitem[Knight, 1921]{Knight.1921} Knight, F. H. (1921). \newblock {\em Risk, uncertainty and profit}. \newblock Houghton Mifflin Harcourt, Boston, MA. \bibitem[Kullback and Leibler, 1951]{Kullback.1951} Kullback, S. and Leibler, R. A. (1951). \newblock On information and sufficiency. \newblock {\em The Annals of Mathematical Statistics}, 22(1):79--86. \bibitem[Mannor et al., 2016]{Mannor.2016} Mannor, S., Mebel, O., and Xu, H. (2016). \newblock Robust {MDP}s with k-rectangular uncertainty. \newblock {\em Mathematics of Operations Research}, 41(4):1484--1509. \bibitem[Mannor et al., 2007]{Mannor.2007} Mannor, S., Simester, D., Sun, P., and Tsitsiklis, J. N. (2007). \newblock Bias and variance approximation in value function estimates. \newblock {\em Management Science}, 53(2):308--322. \bibitem[Mannor and Xu, 2019]{Mannor.2019} Mannor, S. and Xu, H. (2019). \newblock Data-driven methods for {M}arkov decision problems with parameter uncertainty. \newblock In Netessine, S., editor, {\em Operations Research & Management Science in the Age of Analytics}, pages 101--129. INFORMS. \bibitem[Manski, 2004]{Manski.2004} Manski, C. F. (2004). \newblock Statistical treatment rules for heterogeneous populations. \newblock {\em Econometrica}, 72(4):1221--1246. \bibitem[Manski, 2009]{Manski.2009} Manski, C. F. (2009). \newblock {The 2009 Lawrence R. Klein Lecture}: Diversified treatment choice under ambiguity. \newblock {\em International Economic Review}, 50(4):1013--1041. \bibitem[Manski, 2021]{Manski.2021} Manski, C. F. (2021). \newblock Econometrics for decision making: Building foundations sketched by {H}aavelmo and {W}ald. \newblock {\em Econometrica}, forthcoming. \bibitem[Marinacci, 2015]{Marinacci.2015} Marinacci, M. (2015). \newblock Model uncertainty. \newblock {\em Journal of the European Economic Association}, 13(6):1022--1100. \bibitem[McFadden, 1973]{McFadden.1973} McFadden, D. (1973). \newblock Conditional logit analysis of qualitative choice behavior. \newblock In Zarembka, P., editor, {\em Frontiers in {E}conometrics}, pages 105--142. Academic Press, New York City, NY. \bibitem[Meng et al., 2015]{Meng.2015} Meng, F., Qi, J., Zhang, M., Ang, J., Chu, S., and Sim, M. (2015). \newblock A robust optimization model for managing elective admission in a public hospital. \newblock {\em Operations Research}, 63(6):1452--1467. \bibitem[Niehans, 1948]{Niehans.1948} Niehans (1948). \newblock Zur {P}reisbildung bei ungewissen {E}rwartungen. \newblock {\em Swiss Journal of Economics and Statistics}, 84(5):433--456. \bibitem[Nilim and {El Ghaoui}, 2005]{Nilim.2005} Nilim, A. and {El Ghaoui}, L. (2005). \newblock Robust control of {M}arkov decision processes with uncertain transition matrices. \newblock {\em Operations Research}, 53(5):780--798. \bibitem[Pardo, 2005]{Pardo.2005} Pardo, L. (2005). \newblock {\em Statistical inference based on divergence measures}. \newblock Chapman & Hall, London, UK. \bibitem[Puterman, 1994]{Puterman.1994} Puterman, M. L. (1994). \newblock {\em {M}arkov decision processes: Discrete stochastic dynamic programming}. \newblock John Wiley & Sons, New York City, NY. \bibitem[Rahimian and Mehrotra, 2019]{Rahimian.2019} Rahimian, H. and Mehrotra, S. (2019). \newblock Distributionally robust optimization: A review. \newblock {\em arXiv preprint arXiv:1908.05659}. \bibitem[Reich, 2018]{Reich.2018} Reich, G. (2018). \newblock Divide and conquer: Recursive likelihood function integration for hidden {M}arkov models with continuous latent variables. \newblock {\em Operations Research}, 66(6):1457--1470. \bibitem[{robupy}, 2020]{robupy.2020} {robupy} (2020). \newblock A {P}ython package for robust optimization. \bibitem[{ruspy}, 2020]{ruspy.2020} {ruspy} (2020). \newblock An open-source package for the simulation and estimation of a prototypical infinite-horizon dynamic discrete choice model based on {R}ust (1987). \bibitem[Rust, 1987]{Rust.1987} Rust, J. (1987). \newblock Optimal replacement of {GMC} bus engines: An empirical model of {Harold Zurcher}. \newblock {\em Econometrica}, 55(5):999--1033. \bibitem[Rust, 1988]{Rust.1988} Rust, J. (1988). \newblock Maximum likelihood estimation of discrete control processes. \newblock {\em {SIAM} Journal on Control and Optimization}, 26(5):1006--1024. \bibitem[Rust, 1994]{Rust.1994} Rust, J. (1994). \newblock Structural estimation of {M}arkov decision processes. \newblock In Engle, R. and McFadden, D., editors, {\em Handbook of {E}conometrics}, pages 3081--3143. North-Holland Publishing Company, Amsterdam, Netherlands. \bibitem[Saghafian, 2018]{Saghafian.2018} Saghafian, S. (2018). \newblock Ambiguous partially observable {M}arkov decision processes: {S}tructural results and applications. \newblock {\em Journal of Economic Theory}, 178:1--35. \bibitem[Samuelson and Yang, 2017]{Samuelson.2017} Samuelson, S. and Yang, I. (2017). \newblock Data-driven distributionally robust control of energy storage to manage wind power fluctuations. \newblock In {\em 2017 IEEE Conference on Control Technology and Applications (CCTA)}, pages 199--204. \bibitem[Savage, 1954]{Savage.1954} Savage, L. J. (1954). \newblock {\em The foundations of statistics}. \newblock John Wiley & Sons, New York City, NY. \bibitem[Savitzky and Golay, 1964]{Savitzky.1964} Savitzky, A. and Golay, M. J. E. (1964). \newblock Smoothing and differentiation of data by simplified least squares procedures. \newblock {\em Analytical Chemistry}, 36(8):1627--1639. \bibitem[Smith and Winkler, 2006]{Smith.2006} Smith, J. E. and Winkler, R. L. (2006). \newblock The optimizer's curse: Skepticism and postdecision surprise in decision analysis. \newblock {\em Management Science}, 52(3):311--322. \bibitem[Stoye, 2009]{Stoye.2009} Stoye (2009). \newblock Minimax regret treatment choice with finite samples. \newblock {\em Journal of Econometrics}, 151(1):70--81. \bibitem[Stoye, 2012]{Stoye.2012b} Stoye (2012). \newblock Minimax regret treatment choice with covariates or with limited validity of experiments. \newblock {\em Journal of Econometrics}, 166(1):138--156. \bibitem[Su and Judd, 2012]{Su.2012a} Su, C.-L. and Judd, K. L. (2012). \newblock Constrained optimization approaches to estimation of structural models. \newblock {\em Econometrica}, 80(5):2213--2230. \bibitem[Tetenov, 2012]{Tetenov.2012} Tetenov (2012). \newblock Statistical treatment choice based on asymmetric minimax regret choice. \newblock {\em Journal of Econometrics}, 166(1):157--165. \bibitem[Wald, 1950]{Wald.1950} Wald, A. (1950). \newblock {\em Statistical decision functions}. \newblock John Wiley & Sons, New York City, NY. \bibitem[Wiesemann et al., 2013]{Wiesemann.2013} Wiesemann, W., Kuhn, D., and Rustem, B. (2013). \newblock Robust {M}arkov decision processes. \newblock {\em Mathematics of Operations Research}, 38(1):153--183. \bibitem[Zymler et al., 2013]{Zymler.2013} Zymler, S., Kuhn, D., and Rustem, B. (2013). \newblock Worst-case value at risk of nonlinear portfolios. \newblock {\em Management Science}, 59(1):172--188.