EconBase
← Back to paper

Dynamic Models with Robust Decision Makers: Identification and Estimation

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

106,425 characters · 0 sections · 98 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Dynamic Models with Robust Decision Makers: Identification and Estimation

\defaultbibliography{tp} \defaultbibliographystyle{chicago}

bibunit\begin{abstract} \singlespacing This paper studies identification and estimation of a class of dynamic models in which the decision maker (DM) is uncertain about the data-generating process. The DM surrounds a benchmark model that he or she fears is misspecified by a set of models. Decisions are evaluated under a worst-case model delivering the lowest utility among all models in this set. The DM's benchmark model and preference parameters are jointly underidentified. With the benchmark model held fixed, primitive conditions are established for identification of the DM's worst-case model and preference parameters. The key step in the identification analysis is to establish existence and uniqueness of the DM's continuation value function allowing for unbounded statespace and unbounded utilities. To do so, fixed-point results are derived for monotone, convex operators that act on a Banach space of thin-tailed functions arising naturally from the structure of the continuation value recursion. The fixed-point results are quite general; applications to models with learning and Rust-type dynamic discrete choice models are also discussed. For estimation, a perturbation result is derived which provides a necessary and sufficient condition for consistent estimation of continuation values and the worst-case model. The result also allows convergence rates of estimators to be characterized. An empirical application studies an endowment economy where the DM's benchmark model may be interpreted as an aggregate of experts' forecasting models. The application reveals time-variation in the way the DM pessimistically distorts benchmark probabilities. Consequences for asset pricing are explored and connections are drawn with the literature on macroeconomic uncertainty. Keywords: Robust control, ambiguity, model uncertainty, nonparametric identification, nonparametric estimation, entropy, change of measure. JEL codes: C14, C32, D81, E03 \end{abstract} \pagenumbering{arabic} \section{Introduction} A large and active literature explores the implications for individual decision making and policy design under model uncertainty or ambiguity, building on the decision-theoretic foundations of GS, HS2001-ack,HS2001, EpsteinSchneider, KMM2005,KMM2009, MMR, and Strzalecki. Various applications include monetary and fiscal policy design Giannoni,OnatskiStock,CCHS,Woodford,Karantounias, portfolio choice and asset allocation HST,BHS,EpsteinSchneiderAR,HS2017sets, dynamic contracting MiaoRivera, sovereign default PouzoPresno, climate policy Xepapadeas,BrockHansen, and understanding household and professional forecast survey data BBH,Szoke. This paper explores some issues regarding the econometrics of these models. In particular, we study identification and estimation of a class of dynamic models with a single decision maker (DM) who is uncertain about the data-generating process. We will deal mostly with environments in which the DM has multiplier or constraint preferences as in the “robustness” literature pioneered by Hansen and Sargent (see HS2008 and references therein), though extensions to some other classes of preferences will also be discussed. In this setting, the DM's decision problem may be summarized as follows. The DM has a benchmark model of the economy that he or she fears may be misspecified. The DM surrounds the benchmark model by a set consisting of all models whose discounted Kullback--Leibler discrepancy relative to the benchmark model does not exceed some threshold. Decisions are evaluated under a worst-case model that delivers lowest utility among all models within this set, as in the multiple prior framework of GS and EpsteinSchneider.\footnote{For a DM with multiplier preferences, a relative entropy penalty is instead appended to the DM's continuation value recursion. Nevertheless, the DM's behavior retains an ex-post interpretation that decisions are optimal under a worst-case model in a Kullback--Leibler neighborhood of the benchmark model.} The DM's fear of misspecification induces a wedge between the probability measure under which decisions are evaluated and the data-generating probability measure. In contrast, the two probability measures agree in conventional rational expectations models. This wedge must be accounted for when attempting to identify agents' preference parameters. Given the dynamic nature of the DM's problem, the worst-case model is that which lowers the DM's continuation value the most. The worst-case model and the DM's continuation value are pinned down jointly, by a particular nonlinear fixed point equation. This adds a further layer of complexity that must be dealt with when identifying model primitives and developing estimation and inference procedures. Robust decision rules and robust policies depend implicitly on the benchmark model. To date the literature has, with few exceptions, specified tightly-parameterized linear-Gaussian benchmark models. This is largely for the sake of analytic tractability, as it is one of the few instances where the DM's continuation value and worst-case model can be solved for in closed form, at least for certain specifications of the DM's period utility function. While analytically tractable, simple linear-Gaussian specifications often induce worst-case models that are time-invariant in the sense that they do not respond to fluctuations in state variables BHS,BS,BBH, and consequently do not deliver time-variation in prices of risk/uncertainty HS2017sets.\footnote{Sims also raised concerns as to whether the focus on simple linear models overlooks other, potentially more important aspects of model uncertainty. Moreover, the literature on nonlinear dynamic stochastic general equilibrium models has also emphasized the quantitative importance of differences between nonlinear models and log-linear approximations. See, e.g., FVRR and references therein.} More recently, increasing emphasis has been placed on more elaborate nonlinear benchmark specifications or richer preference structures incorporating preference shocks, learning, or multiple layers of uncertainty. With more complicated benchmark models and preferences, analytic tractability may be lost and issues of existence and uniqueness, model identification, and empirical implementation become more opaque. Prompted by these issues, this paper attempts to make progress on several questions, namely: Are there general conditions for existence and uniqueness of continuation values that do not rely on overly restrictive benchmark specifications? What features of model primitives might be identified in general nonlinear Markovian settings? What is required to estimate these models in such settings? We address these questions as follows. First, we study the identification of model primitives within a class of dynamic models featuring a single agent who solves an infinite-horizon robust decision problem. We allow for general nonlinear Markovian environments. The DM's benchmark model and preference parameters are jointly underidentified, even when the worst-case model is fully known. With the benchmark model held fixed, nonparametric identification of the worst-case model and local identification of the DM's preference parameters are established. The key regularity condition is a very mild condition on the distribution of utility growth, which can be easily verified. No further function-analytic conditions, such as compactness, are required. A key step in the identification analysis is to establish existence and uniqueness of the DM's continuation value function. The DM's preference for robustness induces a nonlinear adjustment to the continuation value recursion. As a consequence, the recursion is not a contraction mapping when the value function is allowed to be unbounded, which it is in almost all settings.\footnote{For instance, in linear-Gaussian settings the statespace is unbounded and the value function is affine in the state variable.} Existence and uniqueness of the value function is established by repurposing some tools from the modern statistics literature. The analysis is conducted within a Banach space of unbounded but “thin-tailed” functions that arises naturally from the structure of the recursion, specifically an exponential Orlicz class used in empirical process theory vdVW and modern high-dimensional probability theory Vershynin.\footnote{A special case among the class also has connections with information geometry PistoneSempi and exponential tilting Csiszar1995,KomunjerRagusa.} Monotonicity and convexity properties of the recursion are leveraged to establish existence. Establishing uniqueness requires ensuring that the conditional expectation operator associated with the DM's worst-case model does not move probability mass too far relative to the effect of discounting. Tail inequalities bounding the probabilities of large deviations of thin-tailed random variables are used for this step. Under restrictions on utilities, the value function recursion is isomorphic to that under Epstein--Zin--Weil (EZW) recursive utility and unit intertemporal elasticity of substitution (IES). The existence and uniqueness results therefore apply equally to such models. The existence and uniqueness results are leveraged to establish identification of the agent's worst-case model and preference parameters. As a byproduct, a general existence and uniqueness result is derived for fixed points of monotone, convex operators on classes of unbounded functions.\footnote{See BorovickaStachurski for related results for classes of bounded functions with an emphasis on models with EZW recursive preferences.} Its proof is constructive, using only a few basic results from the theory of integration. The result appears well suited to study existence and uniqueness of value functions in models with forward-looking agents more generally. To illustrate its usefulness, existence and uniqueness of value functions is established in two further applications. The first is models featuring a robust DM who learns about hidden states as in HS2007,HS2010, which nests models with EZW recursive utility and learning as well as other models of ambiguity studied by KMM2009 and JuMiao. The second application is dynamic discrete choice models Rust1987 allowing for unbounded utilities and continuous unbounded statespace. Second, we derive a set of perturbation results characterizing how the continuation value and worst-case model change as the benchmark model changes. The results provide a necessary and sufficient condition for the value function in the perturbed model to converge to the value function in the original model as the perturbation shrinks to zero. These results have several uses. Consider estimating the value function and worst-case model by first estimating the benchmark model (say, from time-series data on state variables or survey data) then solving the continuation value recursion under the estimated model. The perturbation results may be applied to establish consistency and convergence rates of estimators of the value function and worst-case model based on this “plug-in” procedure, treating the estimated model as a perturbation of the truth. The results also permit computation of approximate value functions in models where no closed-form solution exists by perturbing models where closed-form solutions do exist. As an example, it is shown how to compute approximate value functions in nonlinear environments with stochastic volatility by perturbing linear-Gaussian environments. The result may also be used to derive influence functions of plug-in estimators of various asset pricing functionals. Third, we consider an empirical application similar to BHS (see also HHL and BS). In contrast with earlier works, we specify the benchmark model as a covariate-dependent mixture of Gaussian vector autoregressions. This approach has several appealing features. In particular, it has a very natural and intuitive interpretation as a “mixture of experts” where each “expert” is summarized by a vector autoregression. The weights the DM assigns to each expert's forecast vary in a natural way with the state variables. Variation in the mixing weights generates nonlinearities in the conditional mean and conditional variance, which will be seen to generate important asset-pricing implications. In addition, this specification nests conventional linear-Gaussian models as a special case, making it well-suited to conduct a sensitivity analysis of departures from linearity and Gaussianity. The empirical findings are summarized briefly as follows. The time series of the realized change of measure between the benchmark and worst-case model is extracted and is seen to be volatile and counter-cyclical. The worst-case model pessimistically shifts mass towards regions of low consumption growth. Whereas the worst-case model in linear-Gaussian settings is time-invariant, here there is time-variation in the way the DM distorts his or her benchmark model to obtain the worst case. In particular, the DM's worst-case model features a much fatter left tail for consumption growth in “bad” economic states than in “good” states. Time-variation in the wedge between the benchmark and worst-case models generates time-variation in term structures of prices of risk/uncertainty which we explore. Further connections with the literature on macroeconomic uncertainty are also drawn. The remainder of the paper is as follows. Section (ref) describes the class of models under consideration. Section (ref) presents the identification results for continuation values, preference parameters, and the underidentification result. Section (ref) extends the existence and uniqueness results for value functions to models in which the DM is learning about hidden states. Section (ref) presents perturbation results and applies these to estimation. Finally, Section (ref) presents the empirical application. Appendix (ref) contains background material on Orlicz classes, Appendices (ref) and (ref) contain additional results for identification, and Appendix (ref) presents results for Rust-type dynamic discrete choice models. \section{Framework} This section describes the setup in a single-agent setting. Much of this section is a highly stylized summary of material in HS2008 to fix ideas and notation. Extensions to models with learning and other forms of ambiguity aversion are discussed in Section (ref). \subsection{Environment} Consider a discrete-time, infinite-horizon environment. Let $T$ denote the set of non-negative integers. At each date $t \in T$, the DM chooses a vector of controls $C_t \in \mc C_t$ (a constraint set). The source of risk is a time homogeneous, controlled Markov process $X = \{X_t : t \in T\}$ taking values in $\mc X \subseteq \mb R^d$. There is a conditionally deterministic state process $Z = \{Z_t : t \in T\}$ where $Z_t$ characterizes evolution of variables used to describe the constraint set (e.g. wealth or capital). The process $Z$ has law of motion $Z_{t+1} = z(C_t,Z_t,X_t,X_{t+1})$ known to the DM. \subsection{Preferences} First consider a DM with {\it multiplier preferences}, as introduced by HS2001 and axiomatized by Strzalecki. The DM's preference parameters are $(Q,U,\beta,\theta)$, where $U$ is the DM's period utility function, $\beta \in (0,1)$ is a time preference parameter, and $\theta>0$ a risk-sensitivity parameter. It is assumed throughout that $\theta$ is fixed, but the following analysis may be extended to accommodate environments in which $\theta$ is state dependent or, more generally, a stationary stochastic process as in BBH. The DM's benchmark model for evolution of $X$ is described by a Markov kernel $Q(\cdot|x,c)$ representing the conditional distribution of $X_{t+1}$ given $X_t = x$ and $C_t = c$. Let $\mb E^Q$ denote conditional expectation under the DM's benchmark model. Let $\mc F_t$ denote the DM's information set at date $t$ and let $\mc M_{t+1}$ denote the set of all $\mc F_{t+1}$-measurable random variables $m_{t+1}$ with $m_{t+1} \geq 0$ (almost surely) and $\mb E^Q[m_{t+1}|X_t,C_t] = 1$. Each $m_{t+1} \in \mc M_{t+1}$ is a Radon--Nikodym derivative that induces a (conditional) probability measure that is absolutely continuous with respect to the benchmark model. The DM's date-$t$ continuation value $V_t$ is defined by the recursion \begin{align} V_t & = \max_{C_t \in \mc C_t} \min_{m_{t+1} \in \mc M_{t+1}} \bigg( U(C_t,X_t) + \beta \mb E^Q \Big[ m_{t+1} \Big( V_{t+1} + \theta \log m_{t+1}\Big) \Big| X_t,C_t \Big] \bigg) \\ s.t. Z_{t+1} & = z(C_t,Z_t,X_t,X_{t+1}) \,. \notag \end{align} The term $\beta \theta \mb E^Q [ m_{t+1} \log m_{t+1} | X_t,C_t ]$ penalizes the Kullback--Leibler (KL) divergence between the alternate model induced by $m_{t+1}$ and the benchmark model $Q$. As $\theta$ increases, distortions away from $Q$ become increasingly costly. In the limit as $\theta \to \infty$, multiplier preferences approach expected utility preferences. Let $C_t^*$ denote the DM's optimal control at date $t$. The DM's worst-case model is induced by the change of measure \begin{align} m_{t+1}^* = \frac{ e^{-\theta^{-1}V_{t+1}}}{ \mb E^Q[e^{-\theta^{-1} V_{t+1}}|X_t,C_t^*]} \,, \end{align} (see, e.g., HS2008) which will be referred to as the worst-case belief distortion. The worst-case model assigns relatively more weight to events that reduce the DM's continuation value and relatively less weight to events that increase the DM's continuation value. Substituting the worst-case distortion into ((ref)) yields the recursion \begin{align} V_t = U(C_t^*,X_t) - \beta \theta \log \mb E^Q[e^{-\theta^{-1} V_{t+1}}|X_t,C_t^*]\,. \end{align} Closely related to multiplier preferences are constraint preferences. Let $Q$, $U$, and $\beta$ be as above. Also let $\mc C$ denote the constraint set for the sequence $C_0,C_1,\ldots$ and $\mc M$ be the set of all sequences $m_0,m_1,\ldots$ of Radon-Nikodym derivatives as defined above. Finally, let $M_{t+1} = M_t m_{t+1}$ for each $t \geq 0$ with $M_0 = 1$. The DM's date-$0$ problem is: \begin{align} V_0 & = \max_{\{C_t\}_{t \in T} \in \mc C} \min_{\{m_{t+1}\}_{t \in T} \in \mc M} \mb E \bigg[ \sum_{t=0}^\infty M_t \beta^t U(C_t,X_t) \bigg| X_0 ,C_0 \bigg] \\ \mbox{s.t. } & \phantom{==} \sum_{t=0}^\infty \beta^{t+1} \mb E \Big[ M_t \mb E[m_{t+1} \log m_{t+1} | \mc F_t] \Big| X_0 \Big] \leq \gamma \notag \,. \end{align} The constraint in ((ref)) makes clear the sense in which the DM is maximizing worst-case utility over a set of models: these are all models absolutely continuous with respect to $Q$ and whose discounted Kullback--Leibler discrepancy relative to $Q$ are no larger than $\gamma$. The DM's preference parameters consist of $(Q,U,\beta,\gamma)$. By introducing an additional control referred to as {continuation entropy}, HSTW show that constraint preferences may be studied in the same way as multiplier preferences with $\theta$ reinterpreted as a Lagrange multiplier on the model set in ((ref)). The continuation entropy at date $t$, denoted $\Gamma_t$, is defined recursively by \begin{equation*} \Gamma_t = \beta \mb E^Q \Big[ m_{t+1}^* (\Gamma_{t+1} + \log m_{t+1}^*) \Big| X_t \Big] \end{equation*} with $\Gamma_0 = \gamma$. Thus, $\Gamma_t$ is effectively the size of the neighborhood in the DM's date-$t$ problem. Continuation entropy provides a link between $\theta$ and $\gamma$. Although multiplier and constraint preferences are observationally equivalent, they induce different orderings over sequences $\{C_t\}_{t \in T}$. Strzalecki discusses the connection between multiplier and constraint preferences and other classes of preferences in decision theory. \subsection{Accommodating non-stationarity state variables} Growth in the conditionally deterministic state variables may be accommodated under the following mild condition. \@startsection{paragraph}{4}{\z@} {0pt \@plus1ex \@minus.2ex} {-1em} {\normalfont}{Condition S} \emph{(i) There exist $v : \mc X \to \mb R$ and $u : \mc X^2 \to \mb R$ such that \begin{align*} v(X_t) & = -\frac{1}{\theta} \left( V_t - \frac{1}{1-\beta} U(C_t^*,X_t) \right) \,, & u(X_t,X_{t+1}) & = U(C_{t+1}^*,X_{t+1}) - U(C_t^*,X_t) \,; \end{align*} (ii) $X$ is a (strictly) stationary and ergodic, first-order Markov process under $Q(\cdot|X_t,C_t^*)$. } Condition S is maintained throughout the paper. In stationary environments where the DM's optimal choice is a Markov policy $C_t^* = C^*(X_t)$ then Condition S is without loss of generality. In nonstationary environments, Condition S(i) is a homotheticity condition which allows the continuation value to be reformulated in terms of a scaled continuation value $v$ depending on the stationary process $X$ alone. Condition S(ii) is a stationarity condition. Section (ref) verifies Assumption S in a workhorse model featuring stochastic growth. In what follows, with some abuse of notation we write $Q(\cdot|X_t) = Q(\cdot|X_t,C_t^*)$. We also let $Q_0$ denote the stationary distribution of $X_t$ and $Q_0 \otimes Q$ denote the stationary distribution of $(X_t,X_{t+1})$. Under Condition S, it follows from equations ((ref)) and ((ref)) that $v$ solves the recursion \begin{equation} v(X_t) = \beta \log \mb E^Q \left[ \left. e^{ v(X_{t+1}) + \alpha u (X_t,X_{t+1}) } \right|X_t \right] \,, \end{equation} where \[ \alpha = -\frac{1}{\theta(1-\beta)}\,. \] Note that there is a one-to-one correspondence between $(\theta,\beta)$ and $(\alpha,\beta)$. The worst-case distortion may be expressed in terms of $v$ and $u$ as \begin{equation} m_{t+1}^* = m_v(X_t,X_{t+1}) = \frac{e^{ v(X_{t+1}) + \alpha u (X_t,X_{t+1} )}}{\mb E^Q[e^{ v(X_{t+1}) + \alpha u (X_t,X_{t+1} )}|X_t]} \,, \end{equation} and the continuation entropy is given by $\Gamma_t = \Gamma(X_t)$ where $\Gamma$ solves \begin{equation} \Gamma(X_t) = \beta \mb E^Q \Big[ m_{t+1}^* (\Gamma(X_{t+1}) + \log m_{t+1}^*) \Big| X_t \Big] \end{equation} with $\Gamma(X_0) = \gamma$. Equations ((ref)), ((ref)) and ((ref)) will be used extensively in what follows. \subsection{Running example} We use a simple example to illustrate the setup and foreshadow some identification issues that arise. Consider an exchange economy similar to that studied by HHL, BHS and BS. Let $U(C_t,X_t) = \log (c_t e^{\lambda'X_t})$ where $c_t$ is date-$t$ consumption and $\lambda \in \mb R^d$. Each period, the DM chooses how much to consume, $c_t$, and portfolio weights, $\pi_{t+1}$, hence $C_t = (c_t,\pi_{t+1})$. There is an exogenous Markov state process $X$. It is also assumed that aggregate consumption and dividends are both functions of $(X_t,X_{t+1})$, which is trivially the case in typical partial-equilibrium settings where date-$t$ consumption growth and dividends are themselves components of $X_t$. At date $t$, the DM solves \begin{align*} V_t & = \max_{C_t \in \mc C_t} \min_{m_{t+1} \in \mc M_{t+1}} \; \log (c_t e^{\lambda'X_t}) + \beta \mb E^Q \Big[m_{t+1} \Big( V_{t+1} + \theta \log m_{t+1} \Big) \Big| X_t,C_t \Big] \end{align*} subject to a budget constraint. In equilibrium, $c_t = c_t^*$ and $\pi_t = \pi^*$ where $\pi^*$ denotes the market-clearing vector of portfolio weights. The recursion in equation ((ref)) becomes \[ V_t = \log (c_t^*) + \lambda'X_t - \beta \theta \log \mb E^Q \Big[e^{-\theta^{-1} V_{t+1}} \Big| X_t,c_t^* \Big] \,. \] By homotheticity, $V_t - \frac{1}{1-\beta}(\log (c_t^*) + \lambda'X_t) = \zeta(X_t)$ for some $\zeta : \mc X \to \mb R$. Setting $v= -\theta^{-1} \zeta$ yields the recursion in equation ((ref)) with $u(X_t,X_{t+1}) = \log( c_{t+1}^*/ c_t^*) + \lambda'(X_{t+1}-X_t)$. The DM's first-order conditions deliver the Euler equation \begin{equation} \mb E^Q \left[ \left. m_{t+1}^* \beta \left( \frac{c_{t+1}^*}{c_t^*} \right)^{-1} R_{t+1} \right| X_t \right] = 1 \end{equation} where $R_{t+1}$ denotes the return on a traded asset from $t$ to $t+1$ and $m_{t+1}^*$ is from equation ((ref)). \@startsection{paragraph}{4}{\z@} {0pt \@plus1ex \@minus.2ex} {-1em} {\normalfont}{Linear-Gaussian (LG) example:} We use a parametric example from BHS to illustrate some ideas in a transparent way throughout the paper. Suppose the DM's benchmark model is \[ X_{t+1} = \mu + A X_t + \sigma \varepsilon_{t+1}\,, \] where the $\varepsilon_t$ are i.i.d. $N(0,I)$ and all eigenvalues of $A$ are inside the unit circle. Also let $ \log (c_{t+1}^*/c_t^*) = \lambda_0'X_t + \lambda_1' X_{t+1}$ for some fixed $\lambda_0,\lambda_1 \in \mb R^d$. A solution to ((ref)) is $v(x) = a + b x$ where \begin{align*} a & = \frac{\beta}{1-\beta} \Big( (\alpha \lambda_1 + b)'\mu + \frac{1}{2}(\alpha \lambda_1 + b)' \sigma \sigma'(\alpha \lambda_1 + b) \Big) \,, & b & = \alpha \beta (I - \beta A')^{-1}(\lambda_0 + A'\lambda_1) \,. \end{align*} This is the unique solution among affine functions. It remains to be seen whether solutions to recursion ((ref)) exist and are unique under departures from linearity and Gaussianity. The identification results in the next section provide affirmative answers to this questions. The worst-case belief distortion induced by $v$ is \[ m_{t+1}^* = e^{(\sigma'(\alpha \lambda_1 + b))'\varepsilon_{t+1} - \frac{1}{2}(\alpha \lambda_1 + b)'\sigma \sigma' (\alpha \lambda_1 + b)} \,. \] This belief distortion corresponds to shifting the mean of $\varepsilon_{t+1}$ from zero to $\sigma'(\alpha \lambda_1 + b)$. Thus, under the DM's worst-case model: \begin{equation*} X_{t+1} =\mu^* + A X_t + \sigma \varepsilon_{t+1} \end{equation*} where $\mu^* = \mu + \sigma \sigma'(\alpha \lambda_1 + b)$. The worst-case model corresponds to shifting the mean $\mu$ to $\mu^*$ irrespective of the current value of the state. There exists a continuum of $(\theta,\mu)$ that yield identical $\mu^*$. Thus, the DM's benchmark model and preference parameters are underidentified from data on the state $X_t$ and asset returns. As shown in the next section, joint underidentification of the benchmark model and preference parameters is generic. \section{Identification} This section presents three results about identification. First, primitive, directly verifiable conditions are derived for nonparametric identification of the DM's continuation value function, worst-case belief distortion, and continuation entropy given $(Q,U,\beta,\theta)$. The key step is to establish primitive conditions for existence and uniqueness of the DM's value function, which is of independent interest. Second, local identification conditions are derived for the preference parameters $(\beta,\theta)$ given $(Q,U)$. As a special case of the model is isomorphic to models with EZW recursive utility with unit IES, these results have direct implications for identification in EZW models also.\footnote{In those settings, agents are typically assumed to have rational expectations, in which case $Q$ can be identified with the data-generating probability measure.} Third, an underidentification result is stated which shows $(Q,\theta)$ are not identified even if ($U$,$\beta$) and the worst-case model are known. To locally identify preference parameters, it will be presumed that there exists an auxiliary vector of moment conditions holds under the worst-case model, namely: \begin{equation} \mb E^Q \Big[ m_{t+1}^* \beta \bs g(X_t,X_{t+1},Y_{t+1}) - \bs 1 \Big| X_t \Big] = \bs 0 \end{equation} where $Y_t$ is a vector of variables with support $\mb Y \subset \mb R^{d_y}$ such that the conditional distribution of $(X_{t+1},Y_{t+1})$ given $(X_t,Y_t)$ depends only on $X_t$, and the function $\bs g : \mc X^2 \times \mb Y \to \mb R^{d_g}$ is known. An example of this setting is the Euler equation ((ref)), where \begin{align} \bs g(X_t,X_{t+1},Y_{t+1}) = \left( \frac{c_{t+1}^*}{c_t^*} \right)^{-1} \bs R_{t+1} \end{align} where $\bs R_{t+1}$ is a vector of asset returns from date $t$ to $t+1$. The tuples $(Q,U,\beta,\theta)$ and $(Q',U',\beta',\theta')$ are said to be \emph{observationally equivalent} if the conditional moment restriction ((ref)) holds under both $(Q,U,\beta,\theta)$ and $(Q',U',\beta',\theta')$, where the worst-case belief distortion is constructed as in equations ((ref)) and ((ref)) under $(Q,U,\beta,\theta)$ and $(Q',U',\beta',\theta')$, respectively. Throughout this section, $U$ is assumed to be known by the econometrician. The parameter space for $Q$ is the set $\mc Q$ of all Markov transition kernels on $(\mc X,\mcr X)$ and the parameter space for $(\beta,\theta)$ is $B \times \Theta$ where $B = (0,1)$ and $\Theta = (0,\infty)$. \subsection{Existence and uniqueness of continuation values} The recursion for $v$ in equation ((ref)) may be written as the fixed-point equation $v = \mb T v$ with \[ \mb T f (x) = \beta \log \mb E^Q \Big[ e^{ f(X_{t+1}) + \alpha u(X_t,X_{t+1})} \Big| X_t=x \Big] \,. \] It should be understood that $\mb T$ and $v$ depend implicitly on $(\beta,\theta)$. This implicit dependence will be used to derive local identification conditions for $(\beta,\theta)$ in the next subsection. This subsection presents primitive, directly verifiable conditions on $(Q,U)$ under which the operator $\mb T$ has a unique fixed point within an appropriate class of functions for any $(\beta,\theta) \in B \times \Theta$. First, a word on two function classes that are not appropriate: (i) bounded functions and (ii) $L^p$ spaces. Most work to date has established existence and uniqueness within the class $B(\mc X)$ of bounded functions on $\mc X$ equipped with the sup norm (see the discussion after Proposition (ref)). Yet $B(\mc X)$ is an inappropriate class for workhorse parametric models where $v$ is unbounded, such as the LG model discussed above. A possible solution might be to truncate the support of $X$ at some arbitrarily large value. However, artificial restriction of the support of an unbounded state process to bounded sets can lead to uniqueness in the restricted problem even when the unrestricted problem does not have a unique solution (see Appendix (ref) for an example). Intuitively, this is because the nonlinear adjustment in the recursion means that all moments matter, and artificial truncation eventually has a material effect on sufficiently high moments. We therefore seek a class that accommodates unbounded functions. A natural class of unbounded functions is $L^p(Q_0)$, which consists of all functions with finite $p$th moment under $Q_0$. However, the operator $\mb T$ may not be defined on all of $L^p(Q_0)$ for any $1 \leq p < \infty$. For instance, in the LG example above with scalar $X_t$, for any $k \geq 2$ and any $1 \leq p < \infty$ the function $f(x) = x^{2k}$ belongs to $L^p(Q_0)$ but $\mb T f$ is not defined. Intuitively, the tails of $x^{2k}$ under $Q$ are too thick to be compatible with the nonlinear adjustment. To get around these issues, we embed the analysis within a Banach space of unbounded but “thin-tailed” functions. Let $\phi_r(x) = \exp(x^r) -1$ for $r \geq 1$. The \emph{Orlicz space} $L^{\phi_r}(Q_0)$ is (the equivalence class of) all measurable $f : \mc X \to \mb R$ for which \begin{equation*} \|f\|_{L^{\phi_r}(Q_0)} := \inf\left\{ c > 0 : \mb E^{Q_0}[\phi_r(|f(X_t)|/c)] \leq 1\right\} < \infty\,. \end{equation*} The \emph{Orlicz heart} $E^{\phi_r}(Q_0) \subset L^{\phi_r}(Q_0)$ consists of all $f \in L^{\phi_r}(Q_0)$ with $\mb E^{Q_0}[\phi_r(|f(X_t)|/c)] < \infty$ for each $c > 0$. To simplify notation we drop dependence of the spaces and norm on $Q_0$ and simply write $L^{\phi_r}$, $E^{\phi_r}$ and $\|\cdot\|_{\phi_r}$. Suppose $X_t$ (a scalar) is normally distributed under $Q_0$. The function $a_0 + a_1 x + a_2 x^2$ belongs to $L^{\phi_1}$ because $\mb E^{Q_0}[e^{(x/c)^2}] < \infty$ for all $c$ sufficiently large. As this expectation is infinite for $c$ sufficiently small, the function does not belong to $E^{\phi_1}$ if $a_2 \neq 0$. Similarly, the function $a_0 + a_1 x$ belongs to $L^{\phi_2}$ but not to $E^{\phi_2}$ if $a_1 \neq 0$, and belongs to $E^{\phi_r}$ for all $1 \leq r < 2$. Moreover, functions that grow no faster than $|x|^{2/r}$ for some $r > 1$ belong to $E^{\phi_s}$ for all $1 \leq s < r$. The spaces $E^{\phi_r}$ and $L^{\phi_r}$ are (separable and nonseparable, respectively) Banach spaces when equipped with the norm $\|\cdot\|_{\phi_r}$. Further properties of these spaces are described in Appendix (ref), for now we simply observe that $E^{\phi_r} \subseteq E^{\phi_s}$ and $L^{\phi_r} \subseteq L^{\phi_s}$ for each $1 \leq s \leq r$. Define the spaces $E_2^{\phi_r}$ and $L_2^{\phi_r}$ analogously to $E^{\phi_r}$ and $L^{\phi_r}$ for functions of $(X_t,X_{t+1})$ using the stationary distribution $Q_0 \otimes Q$ of $(X_t,X_{t+1})$. The only condition required for identification is that utility growth is thin-tailed. We verify this condition for some models at the end of this subsection. \@startsection{paragraph}{4}{\z@} {0pt \@plus1ex \@minus.2ex} {-1em} {\normalfont}{Assumption U} \emph{$u \in E^{\phi_r}_2$ for some $r > 1$.} Assumption U ensures $\mb T$ is a well-defined mapping from $E^{\phi_s}$ to $E^{\phi_s}$ for each $1 \leq s \leq r$. However, $\mb T$ may not be a contraction on $E^{\phi_s}$ (see Appendix (ref)). We therefore make use of certain monotonicity and convexity properties of $\mb T$ to establish existence and uniqueness of $v$. For the intuition, consider an increasing, convex function $T : \mb R \to \mb R$ (see Figure (ref)). The function $T$ can have zero, one, two, or a continuum of fixed points. If there is a point $\ol v$ such that $T(\ol v)$ lies on or below the 45 degree line and the sequence $T^n(\ol v)$ is bounded from below, then $T$ must have at least one fixed point. This is true of the blue, orange, and purple functions plotted in Figure (ref), but not the grey functions. On the other hand, if at every fixed point the function $T$ has a subgradient that is strictly less than 1 then $T$ must have at most one fixed point. The subgradient of the orange function exceeds 1 at its upper fixed point; similarly, the subgradients of the purple function are 1 along the continuum of fixed points on the 45 degree line. \begin{figure}[hptb] \vskip 8pt \begin{center} \begin{tikzpicture}[scale=0.8, every node/.style={scale=0.8}] \draw[step=1cm,lightgray,very thin] (-4.5,-1.5) grid (4.5,4.5); \draw[thick,<->] (-1,-1.5) -- (-1,4.5) node[anchor=west] {$T(v)$}; \draw[thick,<->] (-4.5,0) -- (4.5,0) node[anchor=west] {$v$}; \draw[thick,dashed] (-2.5,-1.5) -- (3.5,4.5) ; \draw[color=lightgray,domain=-1.5:4,thick,<->] plot ({\x},{\x }) ; \draw[color=lightgray,domain=-4:1.5,thick,<->] plot ({\x},{0.02*exp(\x+3) + 2.5}) ; \draw[color=purple,domain=-1.5:-0.5,thick,-] plot ({\x},{\x+1}) ; \draw[color=purple,domain=-3.5:-1.5,thick,<-] plot ({\x},{0.2*\x-.2}) ; \draw[color=purple,domain=-0.5:0.5,thick,->] plot ({\x},{3.2*\x+2.1}) ; \draw[color=orange,domain=-3:2,thick,<->] plot ({\x},{0.1*exp(\x+2) -1.25}) ; \draw[color=blue,domain=-4:3,thick,<->] plot ({\x},{0.0125*exp(\x+1) +0.1*\x +1}) ; \end{tikzpicture} \vskip 4pt \parbox{12cm}{\caption{ Intuition in one dimension. }} \end{center} \vskip -10pt \end{figure} Proposition (ref) in Appendix (ref) presents a reasonably general existence and uniqueness result which extends this reasoning from a one-dimensional setting to an infinite-dimensional setting. Here we give an heuristic description of the result and introduce relevant definitions. Given $f,g \in E^{\phi_s}$, write $f \leq g$ if $f(x) \leq g(x)$ holds $Q_0$-a.e.. Say that $\mb T $ is \emph{monotone} (or \emph{isotone}) if $f \leq g$ implies $\mb T f \leq \mb T g$ and {\it convex} (or \emph{order-convex}) if for each pair of functions $f$ and $g$ and each $\tau \in [0,1]$ we have $\mb T(\tau f + (1-\tau)g) \leq \tau \mb T f + (1-\tau) \mb T g$. \begin{lemma} Let Assumption U hold. Then $\mb T$ is a continuous, monotone and convex operator on $E^{\phi_s}$ for each $1 \leq s \leq r$. \end{lemma} The properties of $\mb T$ established in Lemma (ref) are used to establish existence of a fixed point $v \in E^{\phi_r}$. Uniqueness requires an appropriate notion of a subgradient of $\mb T$. Let $\mb E_v$ denote expectation under the conditional distribution induced by $m_v$ from equation ((ref)) and define \begin{align*} \mb D_v f(x) & = \beta \mb E_v[ f(X_{t+1}) | X_t = x] = \beta \mb E^Q[ m_v(X_t,X_{t+1}) f(X_{t+1}) | X_t = x] \,. \end{align*} HST showed that the operator $\mb T$ satisfies a subgradient inequality, which they used for pricing assets. In our notation, the subgradient inequality is: \begin{align} \mb T(v+f) - \mb Tv \geq \mb D_v f \,. \end{align} Given a linear operator $\mb K : E^{\phi_s} \to E^{\phi_s}$, let $\|\mb K\|_{E^{\phi_s}} = \sup\{ \| \mb K f\|_{\phi_s} : f \in E^{\phi_s}, \|f\|_{\phi_s} \leq 1\}$ denote its operator norm and $\rho(\mb K; E^{\phi_s}) = \lim_{n \to \infty} \|\mb K^n\|_{E^{\phi_s}}^{1/n}$ denote its spectral radius, where $\mb K^n$ denotes $\mb K$ applied $n$ times in succession. Define $\|\mb K\|_{L^{\phi_s}}$ and $\rho(\mb K; L^{\phi_s})$ analogously. \begin{lemma} Let Assumption U hold and fix any $v \in E^{\phi_{r'}}$ with $r' > 1$. Then: for all $1 \leq s < \infty$, $\mb D_v$ and $\mb E_v$ are continuous linear operators on $E^{\phi_s}$ and $L^{\phi_s}$ with $\rho(\mb D_v;E^{\phi_s}) \leq \rho(\mb D_v;L^{\phi_s}) < 1$. \end{lemma} The property $\rho(\mb D_v;E^{\phi_s}) < 1$ is analogous to the function $T$ having a subgradient less than 1 at its fixed points. This property, together with convexity, delivers uniqueness. It is always the case that $\rho(\beta \mb E^Q;L) = \beta < 1$ because $\mb E^{Q}$ is a weak contraction when $L$ is any $L^p$ or Orlicz class defined relative to $Q_0$.\footnote{The weak contraction property follows by Jensen's inequality, iterated expectations, and stationarity.} However, the stationary distribution under the worst-case model may be different from $Q_0$ in which case $\mb D_v$ is not, in general, a contraction (see Appendix (ref)). Nevertheless, functions in $E^{\phi_r}$ have sufficiently thin tails that, under repeated application of $\mb D_v$, probability mass only moves “so far” and the effect of the discounting by $\beta$ eventually dominates. The next theorem, which is the main result of this subsection, establishes nonparametric identification of $v$ given $(Q,U,\beta,\theta)$ within a class of “thin-tailed” functions. \begin{theorem} Let Assumption U hold. Then: $\mb T$ has a fixed point $v \in E^{\phi_r}$. Moreover, $v$ is the unique fixed point of $\mb T$ in $E^{\phi_s}$ for each $1 < s \leq r$. \end{theorem} \begin{remark} \normalfont The proof of Theorem (ref) also shows: (i) that $\ul v \leq v \leq \ol v$ with \begin{align*} \ul v(x) & = (\mb I - \beta \mb E^Q)^{-1} \beta \mb E^{Q} \big[ \alpha u(X_t,X_{t+1}) \big|X_t = x \big] \\ \ol v(x) & = (1-\beta) \sum_{i=0}^\infty \beta^{i+1} \log \mb E^{Q} \Big[ e^{\frac{\alpha}{1-\beta} u(X_{t+i},X_{t+i+1})} \Big|X_t = x \Big] \,, \end{align*} where $\mb I$ denotes the identity operator; and (ii) that fixed-point iteration on $\ol v$ will converge to $v$. \end{remark} It is worth noting an implication of the proof of Theorem (ref) for the class $E^{\phi_1}$ containing thicker-tailed functions not in $E^{\phi_s}$. Let $\mc V \subset E^{\phi_1}$ denote all fixed points of $\mb T : E^{\phi_1} \to E^{\phi_1}$. Note $\mc V$ always contains $v$ from Theorem (ref). Say $v$ is the \emph{smallest} fixed point of $\mb T : E^{\phi_1} \to E^{\phi_1}$ if $v' \geq v$ for each $v' \in \mc V$. Say $v' \in \mc V$ is \emph{stable} if $\rho(\mb D_{v'};E^{\phi_1}) < 1$ and \emph{unstable} if $\rho(\mb D_{v'};E^{\phi_1}) \geq 1$. Consider the orange function plotted in Figure (ref): its upper fixed point is unstable---iteration on a neighborhood of this fixed point may diverge---whereas its lower fixed point is stable. \begin{proposition} Let Assumption $U$ hold. Then: $v$ is both the smallest fixed point and the unique stable fixed point of $\mb T: E^{\phi_1} \to E^{\phi_1}$. \end{proposition} \begin{remark} \normalfont Existence of a fixed point $v \in E^{\phi_1}$ is guaranteed under the weaker condition $u \in E^{\phi_1}_2$. The stronger condition $u \in E^{\phi_r}_2$ for $r > 1$ (which implies $v \in E^{\phi_r}$) is used to establish stability of $v$ (which implies $v$ is both the smallest and unique stable fixed point of $\mb T$). \end{remark} \@startsection{paragraph}{4}{\z@} {0pt \@plus1ex \@minus.2ex} {-1em} {\normalfont}{Related results:} There exist several works establish existence and uniqueness of value functions using contraction or local contraction arguments (see, e.g., RTW,RZRP,RZRP_recursive,MDRV). However, $\mb T$ and $\mb D_v$ are generally neither contraction mappings nor local contraction mappings on $E^{\phi_s}$, as shown in Appendix (ref). There exists a recent related literature on existence and uniqueness of value functions using monotonicity and concavity/convexity of various operators (see MarinacciMontrucchio2010,Balbus,BorovickaStachurski,GuoHe,BloiseVailakis). Except for Balbus and GuoHe, these papers impose restrictions that rule out recursions of the form ((ref)). The results in these papers apply to classes of bounded functions and therefore require either that the state process $X$ has compact support and/or that utilities are bounded. Unfortunately, such restrictions are incompatible with conventional benchmark models, where $X$ is typically a Markov process with full support, and period utility functions, which are often of logarithmic or CRRA form. Artificially truncating the support of $X$ to be bounded in order to apply these results is not necessarily the right approach, as it may result in misleading conclusions about existence and uniqueness (see Appendix (ref)). Moreover, the above papers generally make use of fixed-point theorems relying on certain topological properties of the space $B(\mc X)$, such as “solidness” of positive cones. These properties are not shared by $L^p$ spaces with $p < \infty$ and Orlicz classes. HS2012 presented spectral conditions for existence of a fixed point in $L^1$ of a related recursion corresponding to EZW preferences allowing for unbounded state variables but did not study uniqueness. npsdfd established \emph{local} identification in the same EZW recursion allowing for unbounded state variables under spectral radius and Fr\'echet differentiability conditions but did not establish \emph{global} identification or existence. Theorem (ref) establishes both these properties, applies to a broader class of models, does not require a differentiability condition, and the spectral radius condition is verified directly. We close this subsection with a discussion of Assumption U for the framework described in Section (ref). \@startsection{paragraph}{4}{\z@} {0pt \@plus1ex \@minus.2ex} {-1em} {\normalfont}{LG environments:} Suppose $u(X_t,X_{t+1}) = \lambda_0'X_t + \lambda_1'X_{t+1}$ is normally distributed under $Q_0 \otimes Q$. Then $u \in L^{\phi_2}_2$ and so $u \in E^{\phi_r}_2$ for each $1 \leq r < 2$. The affine solution $v(x) = a + b'x$ is therefore the unique solution in $E^{\phi_s}$ for all $1 < s < 2$. \@startsection{paragraph}{4}{\z@} {0pt \@plus1ex \@minus.2ex} {-1em} {\normalfont}{Fat tails and rare disasters:} Consider a model featuring time-varying rare disasters from BS. Let $\log(c_{t+1}^*/c_t^*) = g_{t+1}$ where \[ g_{t+1} = \mu_g + w_{z,t+1} + \sigma w_{g,t+1} \,, \] with $w_{g,t+1} \sim N(0,1)$, $w_{z,t+1}|j_{t+1} \sim N(\mu_j j_{t+1}, \sigma_j^2 j_{t+1})$ where $\mu_j < 0$, $j_{t+1}|h_t$ is Poisson distributed with mean $h_t$ which follows an autoregressive gamma (ARG) process (see Appendix (ref) for details). Consumption growth is subject to occasional “disasters” when $j_t > 0$. The rate at which disasters arrive, $h_t$, is time-varying. Define the state as $X_t = (g_t,h_t)$ so that $u(X_t,X_{t+1}) = \lambda_0'X_t + \lambda_1'X_{t+1}$ with $\lambda_0 = (0,0)'$ and $\lambda_1 = (1,0)'$. By iterated expectations: \[ \mb E^{Q_0 \otimes Q}\left[ e^{c u(X_t,X_{t+1})}\right] = e^{c \mu_g + \frac{c^2 \sigma^2}{2}} \mb E^{Q_0} \left[ \exp \left( h_t \left( e^{c \mu_j + \frac{c^2 \sigma_j^2}{2}} - 1\right) \right) \right] \] which is finite only for values of $c$ close to zero because $h_t$ is Gamma distributed under $Q_0$ and the moment generating function of the Gamma distribution is defined only on a neighborhood of the origin. Therefore, $u \in L^{\phi_1}_2$ which violates Assumption U. Indeed, it is known that there may exist zero, one or two fixed points of the form $v = a + b'x$ under this specification. One could modify the above specification so that $w_{z,t+1}|j_{t+1} \sim N(\mu_j j_{t+1}^{1/\varsigma}, \sigma_j^2 )$ for some $\varsigma \in (1,2]$. Given the low frequency of jumps, this modification is likely to be difficult to distinguish empirically from the original specification. Under this modification, one may deduce that $u \in L^{\phi_\varsigma}_2$ and hence $u \in E^{\phi_r}_2$ for each $1 \leq r < \varsigma$, implying that there is a unique fixed point $v \in E^{\phi_s}$ for all $1 < s \leq r$. \subsection{Existence and uniqueness of continuation entropy} The continuation entropy recursion from equation ((ref)) may be expressed in operator notation as \begin{equation} (\mb I - \mb D_v) \Gamma = \chi_v \,, \end{equation} where $\chi_v(x) = \beta \mb E_v \big[ \log m_v(X_t,X_{t+1}) \big| X_t =x \big]$ is the discounted conditional entropy of $m_{t+1}^*$ (cf. equation ((ref))). Equation ((ref)) is a Fredholm equation of the second kind, which have been studied extensively in the applied mathematics literature and used in economics since at least Lucas1978 and TH. It is well known that $\Gamma :=(\mb I - \mb D_v)^{-1} \chi_v $ is the unique solution to ((ref)) in $E^{\phi_s}$ provided $(\mb I - \mb D_v)$ is continuously invertible on $E^{\phi_s}$ and $\chi_v \in E^{\phi_s}$. The spectral radius condition derived in Lemma (ref) is sufficient for invertibility, leading to the following result. \begin{theorem} Let Assumption U hold. Then: $\Gamma=(\mb I - \mb D_v)^{-1} \chi_v$ is the unique solution to ((ref)) in $E^{\phi_s}$ for each $1 \leq s \leq r$. \end{theorem} \subsection{Local identification of preference parameters} This section presents sufficient conditions for local identification of preference parameters $(\beta,\theta)$ given $(Q,U)$ based on the moment condition ((ref)). HST derived an observational equivalence proposition showing that $(\beta,\theta)$ are not separately identified from consumption and investment data alone in linear-quadratic-Gaussian environments. They also showed that data on prices of risky assets could be used to disentangle the two parameters. Intuitively, their positive result arises because varying $(\beta,\theta)$ generates variation in continuation values, and continuation values are reflected in prices of risky assets. The local identification results presented in this section may be viewed partly as a formalization of this intuition. Characterizing the precise source of variation in continuation values required for identification is a nontrivial task, however, as continuation values vary only implicitly as preference parameters vary. Though they did not study models with nonlinear fixed point constraints, our approach is similar in spirit to the general approach of CCLN for nonlinear semiparametric models. Local identification is linked to the rank of a particular matrix. By Theorem (ref) we know $v$ is \emph{globally} identified for given preference parameters. Therefore, here we derive local identification conditions for $(\beta,\theta)$ directly. As a consequence, the rank condition we require is weaker than that which would be required for local identification of $(\beta,\theta,v)$ jointly using the general framework for nonlinear models in CCLN. Our approach to establishing local identification can also be generalized to other models with recursive preferences. Throughout this subsection, let $(\beta_0,\theta_0)$ denote the true preference parameters. Say that $(\beta_0,\theta_0)$ is \emph{locally identified} given $(Q,U)$ if there exists a neighborhood $\mc N \subseteq B \times \Theta$ such that $(Q,U,\beta,\theta)$ and $(Q,U,\beta_0,\theta_0)$ are not observationally equivalent for any $(\beta,\theta) \in \mc N$ with $(\beta,\theta) \neq (\beta_0,\theta_0)$. There is a one-to-one mapping between $(\beta,\theta)$ and $(\alpha,\beta)$, so local identification of one guarantees local identification of the other. It is slightly cleaner to work with $(\alpha,\beta)$ than $(\beta,\theta)$ in what follows. Let $\alpha_0 = -\frac{1}{\theta_0(1-\beta_0)}$ denote the true value of $\alpha$. Let $v_{(\alpha,\beta)}$ denote the solution to the recursion ((ref)) for given $(\alpha,\beta)$. By ((ref)), the moment condition ((ref)) may be written as \begin{align*} \mb E^Q \bigg[ \underbrace{ \frac{e^{ v_{(\alpha_0,\beta_0)}(X_{t+1}) + \alpha_0 u (X_t,X_{t+1} )}}{e^{\beta_0^{-1}v_{(\alpha_0,\beta_0)}(X_t)}} }_{m_{t+1}^*} \beta_0 \bs g(X_t,X_{t+1},Y_{t+1}) - \bs 1 \bigg| X_t \bigg] & = \bs 0\,. \end{align*} We can view the conditional expectation on the left-hand side of the above display as a map from $(\alpha,\beta)$ into a $d_g$-vector of functions of $X_t$. Let \[ \rho(\alpha,\beta;X_t) = \mb E^Q \bigg[ \frac{e^{ v_{(\alpha,\beta)}(X_{t+1}) + \alpha u (X_t,X_{t+1} )}}{e^{\beta^{-1}v_{(\alpha,\beta)}(X_t)}} \beta \bs g(X_t,X_{t+1},Y_{t+1}) - \bs 1 \bigg| X_t \bigg] \,. \] To introduce the result, let $\bs g_{t+1} =\mb E^Q[\bs g(X_t,X_{t+1},Y_{t+1})|X_t,X_{t+1}]$ and $v_0 = v_{(\alpha_0,\beta_0)}$. Also let $\mb E_{v_0}^n$ denote iterated conditional expectation under the worst-case model at the true parameters. Thus, $\mb E_{v_0}^2 h(x) = \mb E^Q[m_{t+1}^* \mb E^Q[m_{t+2}^* h(X_{t+1},X_{t+2})|X_{t+1}] |X_t = x]$, and so on. The Fr\'echet derivatives of $\rho$ with respect to $\alpha$ and $\beta$ at $(\alpha_0,\beta_0)$ are \begin{align} \partial_\alpha \rho(\alpha_0,\beta_0;X_t) & = \mb E_{v_0} \left[ \left. (\beta_0 \bs g_{t+1} - \bs 1) \left( u(X_t,X_{t+1}) + \sum_{n=1}^\infty \beta^n \mb E_{v_0}^n u(X_{t+1}) \right)\right| X_t \right] \\ \partial_\beta \rho(\alpha_0,\beta_0;X_t) & = \frac{1}{\beta_0} \left( \mb E_{v_0} \left[ \left. (\beta_0 \bs g_{t+1} - \bs 1) \left( v_0(X_{t+1}) + \sum_{n=1}^\infty \beta^n \mb E_{v_0}^n v_0(X_{t+1}) \right) \right| X_t \right] - \bs 1 \right) \,. \end{align} Define: \[ \mf V = \mb E^{Q_0} \left[ \, \left( \begin{array}{c} \partial_\alpha \rho(\alpha_0,\beta_0;X_t)' \\ \partial_\beta \rho(\alpha_0,\beta_0;X_t)' \end{array} \right) \left( \begin{array}{c} \partial_\alpha \rho(\alpha_0,\beta_0;X_t)' \\ \partial_\beta \rho(\alpha_0,\beta_0;X_t)' \end{array} \right)'\, \right] \,. \] Let $\mc A = (-\infty,0) \times (0,1)$ denote the parameter space for $(\alpha,\beta)$. For the following result, we may view $\mb T$ as an operator from $\mc A \times E^{\phi_s}$ into $E^{\phi_s}$ for some $1 < s \leq r$. \begin{proposition} Let Assumption U hold, let $\mb T : \mc A \times E^{\phi_s} \to E^{\phi_s}$ be continuously Fr\'echet differentiable at $(\alpha_0,\beta_0,v_0)$, let each element of $\bs g_{t+1}$ have finite $2+\varepsilon$ moment under $Q_0 \otimes Q$ for some $\varepsilon > 0$, and let $\mf V$ be positive definite. Then: $(\beta_0,\theta_0)$ is locally identified. \end{proposition} The key condition for local identification is positive definiteness of $\mf V$. This condition essentially requires sufficient correlation of the residuals $(\beta_0 \bs g_{t+1}-\bs 1)$ with forward-looking expectations of $u$ and $v$ under the worst-case model. As the value function recursion is isomorphic to models with EZW recursive utility with unit IES, Proposition (ref) therefore provides sufficient condition for local identification of preference parameters in that setting also. Global identification conditions may be obtained under further structure on $\bs g$ though we defer this to future research. \subsection{Underidentification of the benchmark model and preference parameters} The LG example clearly illustrated joint underidentification of $Q$ and $\theta$: there is a continuum of $\theta$ and drift parameters $\mu$ that produce in the same worst-case model. This result is now generalized outside of LG environments. Although perhaps obvious, the result is informative in terms of pinpointing the cause of the underidentification. Specifically, for each $\theta > 0$ we construct an alternative model $Q_\theta$ by distorting $Q$ by an amount that is exactly offset when formulating the worst-case model under $Q_\theta$. Correspondingly, the distinct tuples $(Q_\theta,U,\beta_0,\theta)$ and $(Q,U,\beta_0,\theta_0)$ both induce the same worst-case model and are therefore observationally equivalent. \begin{proposition} Let Assumption U hold. Then: for each $\theta > 0$ there is a $Q_\theta \in \mc Q$ such that $(Q_\theta,U,\beta_0,\theta)$ and $(Q,U,\beta_0,\theta_0)$ are observationally equivalent. \end{proposition} Proposition (ref) holds under the conditions that are used to establish existence of continuation values and preference parameters. Thus, underidentification of benchmark models and preference parameters, even when the worst-case model is fully known, is generic. This result is reminiscent of other nonidentification results for Markov decision processes when agents' beliefs and preferences are allowed to vary (see, e.g., Rust1994, Section 3.5). \section{Learning} This section extends the previous existence and uniqueness results to a class of models where the DM learns about a hidden state, e.g. a regime, stochastic volatility, growth process, or time-varying parameter. This setting is relevant for the extension of multiplier preferences by HS2007,HS2010 to accommodate learning. This extension is also relevant for models with generalized recursive smooth ambiguity preferences of JuMiao, recursive smooth ambiguity preferences of KMM2009, and EZW recursive preferences with learning about hidden states as used, for example, by CLL. \subsection{Setting} Partition $X_t = (\varphi_t',\xi_t')$ where the DM observes only $\varphi_t$. Let $\mc O_t = \sigma(\varphi_t,\varphi_{t-1},\ldots,\varphi_0)$ denote the information set observable to the agent at date $t$. The DM's beliefs about $\xi_t$ are summarized by a posterior distribution $\Pi_t$ conditional on $\mc O_t$. As in the extension of multiplier preferences by HS2007,HS2010 to accommodate learning, the date-$t$ value function takes the form: \begin{equation} V_t = U(C_t^*,X_t) - \beta \theta \log \mb E^{\Pi_t}\! \left[ \left. \mb E^Q \left[ \left. e^{-\vartheta^{-1} V_{t+1}} \right| \mc O_t,\xi_t,C_t^* \right]^\frac{\vartheta}{\theta} \right| \mc O_t,C_t^* \right] \,, \end{equation} where parameters $\vartheta>0$ and $\theta>0$ encode concerns about misspecification of $Q$ and $\Pi_t$. When $U(C_t^*,X_t) = \log c_t$ then this recursion is isomorphic to that obtained under generalized recursive smooth ambiguity preferences of JuMiao with unit IES, where $\theta$ and $\vartheta$ are one-to-one transformations of the ambiguity aversion and risk aversion parameters, respectively. When $\vartheta = \theta$, recursion ((ref)) reduces to \[ V_t = U(C_t^*,X_t) - \beta \vartheta \log \mb E^{\Pi_t}\! \left[ \left. \mb E^Q \left[ \left. e^{-\vartheta^{-1} V_{t+1}} \right| \mc O_t,\xi_t,C_t^* \right] \right| \mc O_t,C_t^* \right] \,. \] With $U(C_t^*,X_t) = \log c_t$, this recursion corresponds to EZW recursive preferences with unit IES when learning about the hidden state. A final special case is obtained in the limit as $\vartheta \to \infty$ (thus, the agent is confident in $Q$ but has doubts about the hidden state), in which case: \begin{equation} V_t = U(C_t^*,X_t) - \beta \theta \log \mb E^{\Pi_t}\! \left[ \left. e^{-\theta^{-1} \mb E^Q \left[ \left. V_{t+1} \right| \mc O_t,\xi_t,C_t^* \right] } \right| \mc O_t,C_t^* \right] \,, \end{equation} as is obtained under recursive smooth ambiguity preferences of KMM2009. Several conditions are imposed to make the analysis tractable. First, the state is assumed to have a conventional hidden Markov structure, in which the conditional distribution factorizes as $Q(X_{t+1}|X_t,C_t^*) = Q_\varphi(\varphi_{t+1}|\xi_t,C_t^*)Q_\xi(\xi_{t+1}|\xi_t,C_t^*)$. This accommodates models with regime-switching studied by JuMiao as well as models with learning about a hidden growth term as in CLL and CMST. Our analysis readily extends to allow for realizations of $\varphi_t$ to influence future realizations of $\varphi$, but we maintain this simpler presentation for convenience. The most restrictive condition is a dimension reduction condition assuming $\Pi_t$ is summarized by a finite-dimensional sufficient statistic $\tilde \xi_t$. This is trivially true under Bayesian updating when $\xi_t$ is a hidden regime as in JuMiao or when the evolution of the full state $X_t$ under $Q$ is described by a Gaussian state-space model as in HS2007,HS2010, CLL, CMST, and several other works. In other settings, $\tilde \xi_t$ could be a sufficient statistic used to update beliefs in a boundedly-rational way. Under this condition, the effective state vector is $\tilde \xi_t$. Let $\tilde X_t = (\varphi_t',\tilde \xi_t')$ and let $\mc X_{\tilde X}$, $\mc X_{\tilde \xi}$, and $\mc X_\varphi$ denote the support of $\tilde X_t$, $\tilde \xi_t$, and $\varphi_t$. It is also assumed that learning is in a “steady state” under which the process $\{\tilde \xi_t: t \in T\}$ is stationary. Consider, for instance, LG environments in which learning about hidden states corresponds to updating beliefs via the Kalman filter. If the filter is not initialized in its steady-state then this process will typically be non-stationary. The stationary problem studied here can be viewed as a boundary problem once the filter has converged to its steady state. Solutions could be obtained by backwards iteration from the steady-state boundary solution.\footnote{A similar approach is taken by CDJL in models with an EZW agent who learns about parameters of the data-generating process.} Uniqueness of the boundary solution may be used to establish uniqueness of the backward iterates. The following assumption is maintained throughout this section (cf. Condition S in Section (ref)). \@startsection{paragraph}{4}{\z@} {0pt \@plus1ex \@minus.2ex} {-1em} {\normalfont}{Condition S-Learn} \emph{(i) $X$ is a stationary, first-order Markov process under $Q(\cdot|X_t,C_t^*)$ and the transition distribution factorizes as $Q(X_{t+1}|X_t,C_t^*) = Q_\varphi(\varphi_{t+1}|\xi_t,C_t^*)Q_\xi(\xi_{t+1}|\xi_t,C_t^*)$; \\ (ii) $\Pi_t(\xi_t) = \Pi_\xi(\xi_t|\varphi_t,\tilde \xi_t)$ where $\tilde \xi$ is updated according to a rule $\tilde \xi_{t+1} = \Xi(\tilde \xi_t,\varphi_{t+1})$; \\ (iii) $\{(\xi_t,\tilde X_t) : t \in T\}$ is strictly stationary; \\ (iv) There exist $v : \mc X_{\tilde \xi} \to \mb R$ and $u : \mc X_\varphi \to \mb R$ and such that \begin{align*} v(\tilde \xi_t) & = -\frac{1}{\theta} \left( V_t - \frac{1}{1-\beta} U(C_t^*,X_t) \right)\,, & u(\varphi_{t+1}) & = U(C_{t+1}^*,X_{t+1}) - U(C_t^*,X_t) \,. \end{align*} } \vskip -16pt Before proceeding, two examples of settings in which Condition S-Learn holds are given. For both examples, let $U(C_t^*,X_t) = \log (c_t^*)$ and let $\log(c_{t+1}^*/c_t^*)$ be a function of $\varphi_{t+1}$. \@startsection{paragraph}{4}{\z@} {0pt \@plus1ex \@minus.2ex} {-1em} {\normalfont}{Example: regime switching.} Suppose that $\xi_t$ denotes a hidden regime and evolves as a Markov chain on finite statespace $\{1,\ldots,N\}$ with transition matrix $\bs \Lambda$. Let $\Delta^{N-1}$ denote the simplex in $\mb R^N$. Let the conditional distribution of $\varphi_{t+1}$ given $\xi_t = \xi$ have density $q(\cdot|\xi)$. The posterior $\Pi_t$ is identified with a vector $\tilde \xi_t \in \Delta^{N-1}$ which is updated as: \[ \tilde \xi_{t+1} = \bs \Lambda \frac{\vec q(\varphi_{t+1}) \odot \tilde \xi_t}{\bs 1' (\vec q(\varphi_{t+1}) \odot \tilde \xi_t)} \,, \] where $\vec q(\varphi_{t+1})$ is the $N$-vector whose entries are $q(\varphi_{t+1}|\xi)$ for $\xi \in \{1,\ldots,N\}$, $\odot$ denotes element-wise product, and $\bs 1$ is a $N$-vector of ones Hamilton1994. \@startsection{paragraph}{4}{\z@} {0pt \@plus1ex \@minus.2ex} {-1em} {\normalfont}{Example: Gaussian state-space models.} Suppose $X$ evolves under $Q$ according to: \begin{align*} \varphi_{t+1} & = A \xi_t + u_{t+1} \,, & \xi_{t+1} & = B \xi_t + w_{t+1} \,, \end{align*} where $u_t$ and $w_t$ are i.i.d. $N(0,\Sigma_u)$ and $N(0,\Sigma_w)$, respectively, and where the maximum eigenvalue of $B$ is inside the unit circle. If $\xi_0\sim N(\tilde \mu_{0},\tilde \Sigma_{0})$ under $\Pi_0$ then $\xi_t \sim N(\tilde \mu_{t},\tilde \Sigma_{t})$ under $\Pi_t$. The matrix $\tilde \Sigma_{t}$ will converge to a fixed matrix $\bar \Sigma$ as $t \to \infty$. In this steady state, the sufficient statistic for $\Pi_t$ is $\tilde \xi_t = \tilde \mu_{t}$, which is updated as $\tilde \xi_{t+1} = B \tilde \xi_{t} + B \bar\Sigma A'(A\bar\Sigma A' + \Sigma_u)^{-1}(\varphi_{t+1} - A \tilde \xi_{t})$. \subsection{Existence and uniqueness of continuation values} The only existence and uniqueness result for value functions we are aware of in any of these setting is that of KMM2009, which applies to a more restrictive model (corresponding to $\vartheta = +\infty$), requires finite support of the state, and applies to the class of bounded functions. The results presented below relax these conditions. In view of Condition S-Learn, we again abuse notation slightly and drop dependence of conditional distributions on $C_t^*$. First consider the case with $\vartheta < \infty$. The recursion ((ref)) may be reformulated as the fixed-point equation $v(\tilde \xi) = \tilde{\mb T} v(\tilde \xi)$ where \[ \tilde{\mb T} f(\tilde \xi_t) = \beta \log \mb E^{\Pi_\xi}\! \left[ \left. \mb E^{Q_\varphi} \left[ \left. e^{\frac{\theta}{\vartheta} f(\Xi(\tilde \xi_t,\varphi_{t+1})) + \alpha u(\varphi_{t+1})} \right| \xi_t,\tilde \xi_t \right]^\frac{\vartheta}{\theta} \right| \tilde \xi_t \right] . \] The recursion ((ref)) in the limiting case with $\vartheta = +\infty$ may be reformulated as the fixed-point equation $v(\tilde \xi) = \tilde{\mb T} v(\tilde \xi)$ where \[ \tilde{\mb T} f(\tilde \xi_t) = \beta \log \mb E^{\Pi_\xi}\! \left[ \left. e^{ \mb E^{Q_\varphi} \left[ \left. f(\Xi(\tilde \xi_t,\varphi_{t+1})) + \alpha u(\varphi_{t+1}) \right| \xi_t,\tilde \xi_t \right] } \right| \tilde \xi_t \right] \,. \] The existence and unqiueness results presented below apply to either case, though the proofs are presented only for the more complicated case with $\vartheta < \infty$. The first condition required for identification of $v$ is an appropriate version of Assumption U. Let $\tilde Q_0$ denote the stationary distribution of $\tilde X_t$ and let $E^{\phi_r}_{\tilde X}$ denote the Orlicz heart consisting of all $f : \mc X_{\tilde X} \to \mb R$ for which $\mb E^{\tilde Q_0} [ e^{|u(\tilde X_{t+1})/c|^r}] < \infty$ for each $c > 0$. Similarly, let $E^{\phi_r}_\varphi \subset E^{\phi_r}_{\tilde X}$ and $E^{\phi_r}_{\tilde \xi} \subset E^{\phi_r}_{\tilde X}$ denote functions in $E^{\phi_r}_{\tilde X}$ depending only on $\varphi$ or only on $\tilde \xi$, respectively. \@startsection{paragraph}{4}{\z@} {0pt \@plus1ex \@minus.2ex} {-1em} {\normalfont}{Assumption U-Learn} $u \in E^{\phi_r}_\varphi$ for some $r > 1$. Assumption U-Learn depends only on the marginal distribution of the observed state and is therefore easy to verify. For example, suppose $U(C_{t+1}^*,X_{t+1}) = \log (c_t^*)$. JuMiao study an economy in which consumption and dividend growth is modeled as \begin{align*} \log(c_{t+1}^*/c_t^*) & = \kappa_{\xi_t} + u_{t+1} \,, & \log(d_{t+1}/d_t) & = \zeta \log(c_{t+1}/c_t) + g_d + w_{t+1} \,, \end{align*} where $u_t$ and $w_t$ are i.i.d. $N(0,\sigma_u^2)$ and $N(0,\sigma_w^2)$ and $\xi_t$ is a hidden regime. In this example, $\log(c_{t+1}^*/c_t^*) = \varphi_{t+1}$ and the stationary distribution of $u(\varphi_{t+1})$ is a finite mixture of Gaussians. Assumption U-Learn therefore holds for any $r < 2$. Similarly, Assumption U-Learn holds for any $r < 2$ in Gaussian state-space settings with $\log (c_{t+1}^*/c_t^*) = \lambda_1' \varphi_{t+1}$. The next theorem establishes nonparametric identification of $v$ given $(Q,U,\beta,\vartheta,\theta)$ within classes of “thin-tailed” functions. The result is derived by applying Proposition (ref) in Appendix (ref). The operator $\tilde{\mb T}$ is a continuous, monotone, convex operator on $E^{\phi_s}$ for each $1 \leq s \leq r$ (see Lemma (ref)) and satisfies a subgradient inequality similar to inequality ((ref)). Here, however, the subgradient is a discounted conditional expectation operator under a distorted posterior-predictive distribution. Analogous continuity and spectral radius conditions for the subgradient also hold (see Lemma (ref)). \begin{theorem} Let Assumption U-Learn hold. Then: $\tilde{\mb T}$ has a fixed point $v \in E^{\phi_r}_{\tilde \xi}$. Moreover, $v$ is the unique fixed point of $\tilde{\mb T}$ in $ E^{\phi_s}_{\tilde \xi}$ for each $1 < s \leq r$. \end{theorem} \begin{proposition} Let Assumption U-Learn hold. Then: $v$ is both the smallest fixed point and the unique stable fixed point of $\tilde{\mb T}: E^{\phi_1}_{\tilde \xi} \to E^{\phi_1}_{\tilde \xi}$. \end{proposition} It is possible to relax Assumption S-Learn to allow for $u$ to depend on $(\varphi_t,\varphi_{t+1})$. In this case, however, the effective state vector will be $\tilde X_t$ rather than $\tilde \xi_t$. The above results go through in this case also under an appropriate modification of Assumption U-Learn. Given Theorem (ref), one also may derive local identification results for $(\beta,\theta,\vartheta)$ using similar arguments to Proposition (ref). \section{Estimation} In taking the model to data, the econometrician must either choose a specific benchmark model or adopt a partial identification approach. The previous literature has done the former,\footnote{See, e.g., HST,HSW,AHS.} typically taking the benchmark model to be equal to the member of a parametric family that best approximates the data-generating process. This section develops estimation results under general conditions based on a plug-in estimator of the benchmark model. The results allow the first-stage estimate to be parametric or nonparametric. \subsection{Perturbing the benchmark model} Consider an alternate benchmark model $\hat Q \in \mc Q$. Let $\hat{\mb T}$ be defined by: \[ \hat{\mb T} f(x) = \beta \log \mb E^{\hat Q} \left[ \left. e^{ f(X_{t+1}) + \alpha u(X_t,X_{t+1})} \right| X_t=x \right] \,. \] Under some mild regularity conditions below, $\hat{\mb T}$ will be a well-defined operator and it will have a unique fixed point $\hat v \in E^{\phi_s}$ for each $1 < s \leq r$. This section derives conditions under which $\hat v$ converges to $v$ as $\hat Q$ converges to $Q$ in an appropriate sense. Let $\ll$ denote absolute continuity of measures. Say $\hat Q$ and $Q$ are \emph{everywhere mutually absolutely continuous} if $Q(\cdot|x) \ll \hat Q(\cdot|x) \ll Q(\cdot|x)$ for each $x$. We use the notation $Q \lll \hat Q \lll Q$ to denote everywhere mutual absolute continuity. Whenever this condition holds, write: \begin{align*} \hat \ell(x_0,x_1) & = \log( \hat Q(x_1|x_0) / Q(x_1|x_0) ) \,, \\ \hat \eta(x_0,x_1) & = \hat \ell(x_0,x_1) - \mb E^Q [\hat \ell(X_t,X_{t+1})|X_t = x_0] \,, \mbox{ and }\\ \kappa_{\hat \eta}(x_0) & = \log \mb E^Q[e^{\hat \eta(X_t,X_{t+1})}|X_t = x_0] \,. \end{align*} We will parameterize alternative models by viewing $\hat \ell$ or $\hat \eta$ as elements of Orlicz classes. To do so, let $N^{\phi_s}_2 = \{ f(x_0,x_1) + h(x_1) : f \in L^{\phi_1}_{2}, h \in E^{\phi_s}\}$ equipped with the $L^{\phi_1}_{2}$ norm. We extend $\mb T$ to have domain $N^{\phi_s}_2$ by defining $\mb T : N^{\phi_s}_2 \to L^{\phi_1}$ as: \[ \mb T f(x) = \beta \log \mb E^Q \left[ \left. e^{ f(X_t,X_{t+1}) + \alpha u(X_t,X_{t+1})} \right| X_t=x \right] \,. \] Note that this extension preserves the fixed points of $\mb T$. The operator $\hat{\mb T}$ may be related to the extension of $\mb T$ by noting that for any $f \in E^{\phi_s}$: \begin{align*} \hat{\mb T} f(x) & = \beta \log \mb E^{\hat Q} \left[ \left. e^{ f(X_{t+1}) + \alpha u(X_t,X_{t+1})} \right| X_t=x \right] = \mb T(\hat \ell + f)(x) = \mb T(\hat \eta + f)(x) - \beta \kappa_{\hat \eta} \,. \end{align*} To study how fixed points of $\hat{\mb T}$ relate to those of $\mb T$, we impose a mild regularity condition on $\hat Q$. If $X$ is stationary under $\hat Q$, let $\hat Q_0$ denote its stationary distribution and let $\hat \Delta$ and $\hat \Delta_2$ denote the Radon-Nikodym derivatives of $\hat Q_0$ and $\hat Q_0 \otimes \hat Q$ with respect to $Q_0$ and $Q_0 \otimes Q$. \@startsection{paragraph}{4}{\z@} {0pt \@plus1ex \@minus.2ex} {-1em} {\normalfont}{Assumption AM} Let $Q \lll \hat Q \lll Q$ and let either (a) or (b) of the following hold:\\ (a) $\hat \ell \in E^{\phi_r}_2$ \\ (b) $\hat \ell \in L^{\phi_1}_2$, $X$ is stationary under $\hat Q$ with $Q_0 \ll \hat Q_0 \ll Q_0$, and there is $p > 1$ such that $\mb E^{Q_0}[\hat \Delta(X_t)^p] < \infty$, $\mb E^{Q_0}[\hat \Delta(X_t)^{1-p}] < \infty$, and $\mb E^{Q_0 \otimes Q}[ \hat \Delta_2(X_t,X_{t+1})^p] < \infty$. Assumption AM(b) imposes a less restrictive tail condition on $\hat \ell$ than part (a) but carries the added requirement of stationarity. To understand this assumption, consider the LG setup from Section (ref). If $X_{t+1} = \hat \mu + A X_t + \sigma \varepsilon_{t+1}$ under $\hat Q$ (i.e. only the mean parameter is perturbed), then $\hat \ell \in E^{\phi_r}_2$ and so Assumption AM(a) holds for each $1 \leq r < 2$. If $X_{t+1} = \hat \mu + \hat A X_t + \hat \sigma \varepsilon_{t+1}$ under $\hat Q$, then $\hat \ell \in L^{\phi_1}_2$ and so Assumption AM(b) holds provided all eigenvalues of $\hat A$ are inside the unit circle. \begin{lemma} Let Assumptions U and AM hold. Then: $\hat{\mb T}$ has a fixed point $\hat v \in E^{\phi_r}$ and $\hat v$ is the unique fixed point of $\hat{\mb T}$ in $E^{\phi_s}$ for each $1 < s \leq r$. \end{lemma} One may also establish local Lipschitz and linearity results under a uniform version of Assumption AM. Let $M \geq \mb E^{Q_0 \otimes Q}[ \exp(|\frac{\alpha}{1-\beta} u(X_t,X_{t+1})|^r)]$ be a finite positive constant. \@startsection{paragraph}{4}{\z@} {0pt \@plus1ex \@minus.2ex} {-1em} {\normalfont}{Assumption AM2} Let $Q \lll \hat Q \lll Q$ and let either (a) or (b) of the following hold: \\ (a) $\hat \ell \in E^{\phi_r}_2$, $\| \hat \ell \|_{\phi_r} \leq M$, and $\mb E^{Q_0 \otimes Q}[ \exp(|\frac{1}{1-\beta} ( \hat \ell(X_t,X_{t+1}) +\alpha u(X_t,X_{t+1})|^r)] \leq M$ \\ (b) $\hat \ell \in L^{\phi_1}_2$, $X$ is stationary under $\hat Q$ with $Q_0 \ll \hat Q_0 \ll Q_0$, and there is $p > 1$ such that $\mb E^{Q_0}[\hat \Delta(X_t)^p]\leq M$, $\mb E^{Q_0}[\hat \Delta(X_t)^{1-p}] \leq M$, and $\mb E^{Q_0 \otimes Q}[ \hat \Delta_2(X_t,X_{t+1})^p] \leq M$. Let $C$ denote a positive constant depending only on $\alpha$, $\beta$, $\|u\|_{\phi_r}$, $r$, and $M$ under AM2(a) or on $\alpha$, $\beta$, $\|u\|_{\phi_r}$, $r$, $M$, $p$, and $\mb E^{Q_0 \otimes Q}[ \exp(q^2|\frac{\alpha}{1-\beta} u(X_t,X_{t+1})|^r)]$ under AM2(b) where $p^{-1} + q^{-1} = 1$. \begin{lemma} Let Assumption U and AM2 hold and let $\| \hat \eta\|_{\phi_1} \leq 1$. Then: \begin{equation*} \| \hat v - v \|_{\phi_1} \leq C \| \hat \eta\|_{\phi_1} \end{equation*} and \begin{equation*} C^{-1} \| \hat{\mb T} v - v\|_{\phi_1} \leq \| \hat v - v \|_{\phi_1} \leq C \| \hat{\mb T} v - v\|_{\phi_1} \,. \end{equation*} The inequalities also hold in $\|\cdot\|_{\phi_s}$ norm for every $1 \leq s \leq r$ under Assumption AM2(a). \end{lemma} The first inequality in Lemma (ref) shows $\hat v - v$ is locally Lipschitz in $\hat \eta$. It follows from the second inequality that the rate at which $\hat v$ converges to $v$ is equivalent to the rate at which $\hat{\mb T}v $ converges to $v$. Thus, it is \emph{necessary} that $\| \hat{\mb T} v - v\|_{\phi_1} \to 0$ in order that $\|\hat v - v\|_{\phi_1} \to 0$. To interpret the second inequality, note that \[ \hat{\mb T} v(x) - v(x) = \beta \left( \log \mb E_v \left[ \left. e^{\hat \eta(X_t,X_{t+1})} \right| X_t = x \right] - \log \mb E^Q \left[ \left. e^{\hat \eta(X_t,X_{t+1})} \right| X_t = x \right] \right) \,, \] i.e., the discounted difference between a certainty equivalent adjustment of $\hat \eta$ under the worst-case and benchmark models. For the following local linearization result, we view $\eta \mapsto \kappa_{\eta}$ as a map from $L^{\phi_1}_2$ to $L^{\phi_1}$ and index the subgradient by $h \in L^{\phi_1}_2$. The operator $\mb D_{h + v}$ is defined formally in Appendix (ref). \begin{lemma} Let Assumptions U and AM2 hold, let $\eta \mapsto \kappa_{\eta}$ be Fr\'echet differentiable at $\eta = 0$ and let $\|\mb D_{h+v} - \mb D_v\|_{L^{\phi_1}} \to 0$ as $\|h\|_{\phi_1} \to 0$. Then: \begin{equation*} \hat v - v = \sum_{n = 1}^\infty (\beta \mb E_v)^n \hat \eta + o(\|\hat \eta\|_{\phi_1}) \,. \end{equation*} \end{lemma} Lemma (ref) justifies the approximation $\hat v - v \approx \sum_{n = 1}^\infty (\beta \mb E_v)^n \hat \eta$ when $\|\hat \eta\|_{\phi_1}$ is small. This result shows that approximate value functions in models with rich dynamics may be approximated by perturbing simpler models with closed-form solutions. Appendix (ref) presents an example showing how to approximate continuation values in models featuring stochastic volatility by perturbing LG environments. The perturbation is in terms of the likelihood ratio relative to a model with a known solution, unlike usual perturbation methods that expand around a deterministic steady state. In that respect, it shares some similarities with the approach of KoganUppal used by HHLR and HHL to compute approximate continuation values by expanding a preference parameter about a value with a known solution. Here the expansion is in the (infinite-dimensional) score of the alternative model rather than a (scalar) preference parameter. Lemma (ref) may also be used to compute influence functions of plug-in estimators of asset pricing functionals. \subsection{Consistency and convergence rates for general estimators} We first consider plug-in estimators based on frequentist procedures then turn to Bayes procedures. Given a (parametric or nonparametric) first-stage estimator $\hat Q$ of $Q$, the continuation value recursion may be solved under $\hat Q$ to obtain a fixed point $\hat v$. Lemma (ref) guarantees existence and uniqueness of $\hat v$ provided $\hat Q$ satisfies Assumption AM. Given $\hat v$, the worst-case belief distortion may be estimated using: \[ m_{\hat v}(X_t,X_{t+1}) = e^{\hat v(X_{t+1}) + \alpha u(X_t,X_{t+1}) - \beta^{-1} \hat v(X_t)} \,. \] Let $\|f\|_p = \mb E^{Q_0 \otimes Q}[f(X_t,X_{t+1})^p]^{1/p}$ denote the $L^p(Q_0 \otimes Q)$ norm. Let $a_n$ be a positive sequence with $a_n \to 0$ as $n \to \infty$. \begin{proposition} Let Assumption U hold, let $\hat Q$ satisfy assumption AM2 wpa1, let $\|\hat \eta\|_{\phi_1} = o_p(1)$ and let $\| \hat{\mb T} v - v\|_{\phi_1} = O_p(a_n)$. Then: $\| \hat v - v\|_{\phi_1} = O_p(a_n)$, $\|m_{\hat v} - m_v\|_{p} = O_p(a_n)$ and $\|\frac{m_{\hat v}}{m_v}-1\|_{p} = O_p(a_n)$ for each $1 < p < \infty$. \end{proposition} For Bayes procedures, let $\Pi_n$ denote a posterior distribution for $Q$. In parametric models $\Pi_n$ can be a posterior over the parameters in $Q$, but we also allow for nonparametric settings in which $\Pi_n$ is a posterior over a nonparametric class of transition kernels. For each draw $\hat Q$ from $\Pi_n$ that satisfies Assumption AM, one can construct $\hat{\mb T} = \mb T(\hat Q)$ then compute its fixed point $\hat v = v(\hat Q)$, the belief distortion $m _{\hat v} = m_v(\hat Q)$, and so on, building up posterior distributions for these quantities across repeated draws. The next result presents conditions under which such a procedure is consistent and characterizes posterior contraction rates. \begin{proposition} Let Assumption U hold, let $\Pi_n(\mc A_n) = 1 + o_p(1)$ for a sequence of subsets $\mc A_n$ satisfying AM2 with $\sup_{\hat Q \in \mc A_n} \| \eta(\hat Q) \|_{\phi_1} = o(1)$ and $\sup_{\hat Q \in \mc A_n} \|({\mb T}(\hat Q)) v - v\|_{\phi_s} = O(a_n)$. Then: \begin{align*} \Pi_n ( \{ \hat Q : \| v(\hat Q) - v(Q)\|_{\phi_1} > C_n a_n \} ) & = o_p(1) \,, \\ \Pi_n ( \{ \hat Q : \| m_v(\hat Q) - m_v(Q)\|_p > C_n a_n \} ) & = o_p(1) \,, \mbox{ and }\\ \Pi_n ( \{ \hat Q : \| {\textstyle \frac{m_v(\hat Q)}{m_v(Q)}-1 } \|_p > C_n a_n \} ) & = o_p(1) \end{align*} for each $1 < p < \infty$ and each positive sequence $C_n \to \infty$. \end{proposition} \subsection{Mixtures of experts} The empirical approach we take in the next section is to treat the benchmark model $Q$ as a covariate-dependent mixture of Gaussian VARs. This model can be interpreted as a “mixture of experts” where each “expert” is represented by a Gaussian VAR(1) and the weight that the DM assigns to each expert's forecast is time-varying. Mixtures of experts have long been popular in statistics, machine learning, and computer science for solving prediction problems, including in various dynamic settings.\footnote{For early applications to time series see ZMA. For more recent applications to macroeconomic time series see VKG and KalliGriffin.} This model is attractive for our purposes for several reasons. First, the model is very flexible yet retains a clear interpretation which is not necessarily the case, say, with estimates of $Q$ based on other “flexible” estimation techniques such as kernels. Second, it is easy to compute transition densities and simulate from the model, facilitating easy computation of value functions and equilibrium prices. Third, the mixtures can approximate smooth conditional densities arbitrarily well as the number of mixing components increases (see, e.g., Norets2010mixture). Fourth, the procedure can be embedded in a state-space setting, which may be relevant for dealing with measurement error and/or mixed frequencies at which macroeconomic data are available. Finally, regularity conditions from Section (ref) guaranteeing existence of value functions and so on are easy to verify under transparent conditions. We treat the joint distribution for $(X_t,X_{t+1})$ as a $K$-component mixture of normals: \[ f(x_t,x_{t+1}) = \sum_{k=1}^K w_k \, \phi( (x_t',x_{t+1}')' ; \mu_k^{(2)} , \Omega_k^{(2)} ) \,, \] where $0 \leq w_k \leq 1$ with $\sum_{k=1}^K w_k = 1$, $\phi( x ; \mu, \Omega)$ denotes the normal probability density function with mean $\mu$ and covariance $\Omega$ (with dimensions conformable with $x$), and \begin{align*} \mu^{(2)}_k & = \left[ \begin{array}{c} \mu_k \\ \mu_k \end{array} \right] \,, & \Omega^{(2)}_k & = \left[ \begin{array}{cc} \Omega_k & \Omega_k A_k' \\ A_k \Omega_k & \Omega_k \end{array} \right] \,, \end{align*} where $A_k$ is a square matrix with all eigenvalues inside the unit circle and $\Omega_k$ is positive definite and symmetric. The process $X$ is strictly stationary and ergodic under this specification, with stationary density \[ f_0(x_t) = \sum_{k=1}^K w_k \, \phi( x_t ; \mu_k , \Omega_k ) \] and conditional density \[ f(x_{t+1}|x_t ) = \sum_{k=1}^K w_k(x_t) \, \phi( x_{t+1} ; (I - A_k) \mu_k + A_k x_t , \Sigma_k ) \] where \[ w_k(x_t) = \frac{w_k \, \phi( x_t ; \mu_k , \Omega_k ) }{\sum_{i=1}^K w_i \, \phi( x_t ; \mu_i , \Omega_i )} \,, \] and $\Sigma_k = \Omega_k - A_k^{\phantom \prime} \Omega_k A_k'$. The quantity $w_k(x_t)$ is the weight assigned to the $k$th forecasting model having observed $X_t = x_t$, the $k$th forecasting model itself being a Gaussian VAR(1) with mean $(I - A_k) \mu_k$, autoregressive coefficients $A_k$, and conditional variance $\Sigma_k$. State dependence of the weights generates time-variation in the conditional mean and conditional variance of $X$. Let $\mc Q_K$ denote the set of all such $K$-component mixtures. Also let $\bar {\mc Q}_K$ denote all $Q \in \mc Q_{K}$ whose $\mu_k$ are uniformly bounded and the smallest and largest eigenvalues of $\Omega_k$ and $\Omega_k^{(2)}$ are uniformly bounded away from $0$ and $+\infty$. Say that $Q_0$ has \emph{Gaussian-like tails} if it has (Lebesgue) density $q_0$ for which there exist $\ul c, \ol c, \ul s, \ol s \in (0,\infty)$ such that $\ul c \exp(-\frac{1}{2\ul s^2} \|x\|^2 ) \leq q_0(x) \leq \ol c \exp(-\frac{1}{2\ol s^2} \|x\|^2)$. Note that $Q_0$ does not necessarily have to be Gaussian to have Gaussian-like tails: it just must lie between some multiples of Gaussian distributions with possibly different covariance matrices. \begin{lemma} Let $Q$ have strictly positive conditional density $q(\cdot|x_t)$ on $\mb R^d$ for each $x_t$ and let the marginal $Q_0$ and joint $Q_0 \otimes Q$ distributions of $X_t$ and $(X_t,X_{t+1})$ have Gaussian-like tails. Then: any $\hat Q \in \mc Q_K$ satisfies Assumption AM(b). If, moreover, $u(X_t,X_{t+1}) = \lambda_0'X_t + \lambda_1'X_{t+1}$, then: Assumption U holds and Assumption AM2(b) holds for each $\hat Q \in \bar{\mc Q}_K$. \end{lemma} Consider an environment where the DM's benchmark model $Q$ is the closest approximation within the class $\mathcal Q_K$ to the true dynamics of $X$.\footnote{Here “closest” in the sense of minimizing average Kullback--Leibler divergence between the conditional densities under the data-generating process and $Q$.} As the conditions of the first part of Lemma (ref) hold, this guarantees existence and uniqueness of a fixed point $\hat v$ for each $\hat Q \in \mc Q_K$ under Assumption U. If the uniformity conditions in the second part of Lemma (ref) hold then we may apply the earlier consistency results. All that remains to check is whether the score terms vanish in the manner described by Propositions (ref) and (ref). \section{Empirical application} This section revisits an economy similar to that described in Section (ref) and studied by HHL, BHS, and BS, amongst others. We depart from the previous literature by modeling the DM's benchmark model as a mixture of experts as described in Section (ref). Our perspective here is to treat this application as a type of sensitivity analysis by examining how various equilibrium quantities differ under slightly more flexible, though still intuitive, nonlinear specifications for the benchmark model. As will be seen, introducing nonlinearities into state dynamics in this fashion generates interesting predictions about equilibrium prices and term structures relative to those obtained under LG specifications. This sections explores these differences and the channels through which they arise. \subsection{Setup} Preferences are as described in Section (ref) with $U(C_t,X_t) = \log (c_t)$. Similar to HHL, we use two state variables: aggregate consumption growth and the consumption-earnings ratio (both in logs). The data sourced from the NIPA tables, are at the quarterly frequency, and span 1947Q1 to 2018Q3. The two series are plotted in Figure (ref). The series are approximately uncorrelated and may be thought of as representing high- and low-frequency sources of risk. Two benchmark models for state dynamics are used. The first is a covariate-dependent mixtures of Gaussian vector autoregressions as described in Section (ref). We use $K = 4$ mixtures, though our results were reasonably insensitive to this choice. The second is a LG model where the state is treated as a first-order Gaussian vector autoregression. The first specification with $K>1$ allows time-variation in the conditional variance of $X$ whereas the LG model does not. Both models are estimated using Bayes procedures. We use the same priors on parameters common to both models. For the mixture specification, we use an adaptive sequential Monte Carlo algorithm HerbstSchorfheide to accommodate potential multi-modality of the posterior. \begin{figure}[p] \begin{center} \parbox{12cm}{\caption{ Time series of log consumption growth and log consumption-earnings ratio. Recession periods are indicated as shaded regions.}} \end{center} \end{figure} \begin{figure}[p] \begin{center} \parbox{12cm}{\caption{ Upper panel: Realized belief distortion $m_v(X_t,X_{t+1})$ for the mixture specification. Lower panel: Difference between the conditional means of future consumption growth under the benchmark and worst-case models for the mixture specification. Recession periods are indicated as shaded regions.}} \end{center} \end{figure} For each draw from the posterior, we calculate: (i) the stationary and transition distributions under the benchmark model, (ii) the value function $v$, from which we construct (iii) the worst-case belief distortion $m_v$, (iv) the stationary distribution under the worst-case model and the transition distribution under the worst-case model, (v) the continuation entropy function $\Gamma$, and (vi) term structures of the risk-free rate and excess returns on earnings strips.\footnote{We work with earnings data rather than dividend data to avoid potential seasonality in dividend series.} Computations for the mixture model are performed numerically using interpolation on a large grid.\footnote{As the state-space is compact when using a grid, Proposition (ref) guarantees existence and uniqueness of $v$ in the discretized problem. Our identification results remain relevant in this setting as they ensure that there is a unique solution for the actual un-discretized problem.} To focus on the role of varying the benchmark model, we calibrate the preference parameters to seemingly reasonable values rather than estimating them directly from data. A more thorough empirical investigation would estimate these parameters from data on asset returns. In particular, we fix the time preference parameter to $\beta = (0.98)^{1/4}$ and the risk-sensitivity parameter to $\theta = 7.367$. The implied return on a 30-year discount bond is around 3% per annum under this parameterization. Chernoff entropy and detection error probabilities can be used to interpret the scale of $\theta$; see AHS and HS2008. The posterior mean Chernoff entropy between $Q$ and the worst-case model is $0.0056$ for the mixture specification. The posterior mean half-life of detection-error probabilities is approximately 32 years. Thus, approximately an additional 32 years' worth of data is required in order for error probabilities of likelihood-ratio tests between $Q$ and the worst-case model to halve. The benchmark and worst-case models may therefore reasonably be viewed as statistically difficult to discriminate from one another given the length of data available. \subsection{The worst-case model: time-varying tails and pessimism} The upper panel of Figure (ref) plots time series of the realized worst-case belief distortion for the mixture specification. This series is constructed by taking the posterior mean of $m_v(X_t,X_{t+1})$ for each date $t$. The series is volatile and pronouncedly counter-cyclical, rising sharply during recessions. Comparing the time series for state variables in Figure (ref), the belief distortion also fluctuates at a higher frequency than both of the state variables. The lower panel of Figure (ref) plots the posterior mean difference between the conditional means of future consumption growth under the benchmark and worst-case models, in percent per year terms.\footnote{I.e., the posterior mean of $(\mb E^Q[\log (C_{t+1}/C_t)|X_t]-\mb E_v[\log (C_{t+1}/C_t)|X_t])\times 400$.} As can be seen, this series is time-varying and counter-cyclical, with the spread rising from below 0.3% outside of recession periods to around 0.5%--0.8% around recession periods. Thus, the worst-case model becomes relatively more pessimistic about consumption growth than the benchmark model during recession periods. The spread is also much more volatile around recession periods. In contrast, the spread is constant for the LG specification. The posterior mean difference is around 0.36% per annum for the LG model, which agrees with the average posterior mean spread for the mixture specification over the 284 quarters. Thus, the LG model matches the same average spread but misses an important dynamic component. The time-varying pessimism reported in Figure (ref) indicates that the wedge between the benchmark and worst-case models is time-varying. To explore this further and understand differences relative to a LG specification, Figures (ref) and (ref) display the conditional distribution for $X_{t+1}$ given $X_t$ under the benchmark and worst-case models in two states. The first is a “good” state (Figure (ref)) when consumption growth is one standard deviation higher than its mean (around 3.77%) and the consumption earnings ratio is one standard deviation lower than its mean. The second is a “bad” state (Figure (ref)) where consumption growth is one standard deviation lower than its mean (around -0.25%) and the consumption earnings ratio is one standard deviation high than its mean. Both figures show that the worst-case model assigns more mass to regions of low consumption growth relative to the benchmark model. In the good state, the benchmark and worst-case distributions look similar to those for the LG benchmark specification reported in Figure (ref). In the bad state, however, the conditional distribution in the benchmark model has a longer left tail for consumption growth and the worst-case model assigns relatively more mass far out in the left tail. This variation in the way the benchmark model is distorted to obtain the worst-case model generates the time-varying pessimism reported in Figure (ref). In contrast, for the LG benchmark specification, the worst-case model in the bad state (Figure (ref)) looks exactly as it does in the good state, modulo a change in location, with identical contours and marginals. This is entirely as expected: the worst-case model under the LG benchmark is also a Gaussian VAR(1) with a fixed location shift (cf. Section (ref)). The asymmetry in the way the left tails of consumption growth behave in the good versus bad states is reminiscent of the work on “investor fears” by BollerslevTodorov and BollerslevTodorovXu. Using S&P500 options data and model-free continuous-time nonparametric methods, these studies document important time-variation in the wedge between the objective and risk-neutral jump sizes and intensities, and asymmetries between the pricing of left- and right-tail risk, which are ascribed to fluctuations in investor fears. Of course, our frameworks and data sources are very different from these works. Nevertheless, in view of Figures (ref), (ref), and (ref), it is reasonable to expect that the time-variation in the way the benchmark model is distorted would lead to qualitatively similar pricing of tail events. \begin{figure}[p] \begin{center} \parbox{12cm}{\caption{ Mixture specification: Conditional distribution of $X_{t+1}$ given $X_t$ under the benchmark (red) and worst-case (blue) models in the “good” state. Data points are plotted in the center.}} \end{center} \end{figure} \begin{figure}[p] \begin{center} \parbox{12cm}{\caption{ Mixture specification: Conditional distribution of $X_{t+1}$ given $X_t$ under the benchmark (red) and worst-case (blue) models in the “bad” state.}} \end{center} \end{figure} \begin{figure}[p] \begin{center} \parbox{12cm}{\caption{ LG specification: Conditional distribution of $X_{t+1}$ given $X_t$ under the benchmark (red contours and marginals) and worst-case (blue contours and marginals) models in a “good” state.}} \end{center} \end{figure} \begin{figure}[p] \begin{center} \parbox{12cm}{\caption{ LG specification: Conditional distribution of $X_{t+1}$ given $X_t$ under the benchmark (red contours and marginals) and worst-case (blue contours and marginals) models in a “bad” state.}} \end{center} \end{figure} Time-variation in the benchmark and worst-case model in the mixture specification also leads to interesting properties of the implied stationary distribution, which is displayed in Figure (ref). Relative to the benchmark model, the stationary distribution under the worst-case model has a much fatter left tail for consumption growth---a long-run consequence of the distortion exhibited in Figure (ref)---and a slightly higher mean for the consumption-earnings ratio. \begin{figure}[p] \begin{center} \parbox{12cm}{\caption{ Mixture specification: Stationary distribution the benchmark (red contours and marginals) and worst-case (blue contours and marginals) models.}} \end{center} \end{figure} \begin{figure}[p] \begin{center} \parbox{12cm}{\caption{ Term structures of excess returns on earnings strips in the “good” state. Solid lines are posterior means, shaded bands are 90% pointwise credible sets.}} \end{center} \end{figure} \begin{figure}[p] \begin{center} \parbox{12cm}{\caption{ Term structures of excess returns on earnings strips in the “bad” state. Solid lines are posterior means, shaded bands are 90% pointwise credible sets.}} \end{center} \end{figure} \begin{figure}[p] \begin{center} \parbox{12cm}{\caption{ Term structures of excess returns on earnings strips in the “average” state. Solid lines are posterior means, shaded bands are 90% pointwise credible sets.}} \end{center} \end{figure} \subsection{Implications for asset prices} To explore the implications of the model for asset prices, we compute term structures of excess returns on earnings strips in different states.\footnote{I.e. $\log \mb E^Q[ E_{t+\tau}|X_t] - \log \mb E_v[ \beta^\tau (C_t/C_{t+\tau}) E_{t+\tau}|X_t] + \log \mb E_v[ \beta^\tau (C_t/C_{t+\tau}) |X_t]$ where $\tau$ is the horizon $E_t$ denotes earnings at date $t$. The term $\beta^\tau (C_t/C_{t+\tau})$ is the DM's stochastic discount factor for pricing claims to date $t+\tau$ payoffs at date $t$. The term $\log \mb E_v[ \beta^\tau (C_t/C_{t+\tau}) |X_t]$ corrects for the risk-free rate.} The posterior means in three states are plotted in Figure (ref) (good state), (ref) (bad state), and (ref) (an “average” state, where both state variables equal their mean). Each plot presents the posterior mean excess return in solid lines together with horizon-wise 90% credible sets as shaded regions. As can be seen, the term structures are time-varying, with a hump shape in the good state, an upwards-sloping shape in the bad state, and a downwards-sloping shape in the average state.\footnote{As the environment is ergodic, however, the long-end of the term structure remains fixed at around 0.75%. } This time-variation at the short end cannot be generated in LG benchmark specifications in this setting. HS2017sets provide a dynamic extension of max-min preferences in which agents consider both parametric and nonparametric families of models. Their extension of max-min preferences can generate state dependence in worst-case models and uncertainty prices even in LG environments. \subsection{Macroeconomic uncertainty} Finally, we compare three time series related to the model with other notions of macroeconomic uncertainty. The first series is the difference between the conditional mean of consumption growth under the benchmark and worst-case models, as in Figure (ref). The second is the continuation entropy a function of the realized state, i.e. $\Gamma(X_t)$. Both of these series are constant with a LG benchmark model but are time-varying for the mixture specification. The third series is the entropy of the experts' mixture weights, i.e. $-\sum_{k=1}^K w_k(X_t) \log w_k(X_t)$. Each of these series are distinct in nature: the first represents time-varying pessimism. The second represents the size (in terms of discounted relative entropy) of the set of models over which the agent is maximizing worst-case utility. The third series measures the dispersion in the forecast weights in the benchmark model. This third series may be interpreted as uncertainty among the mixture components, and is bounded between zero (where the weight is essentially one for one component and zero for all others) and $\log K$, when all components have equal weight. \begin{figure}[t] \begin{center} \parbox{12cm}{\caption{ Time series of the posterior means of the difference between the conditional mean of consumption growth under the worst-case and benchmark models, continuation entropy, and entropy of experts' weights in the benchmark model. Bloom2009 major stock-market volatility shock dates are indicated as shaded regions.}} \end{center} \end{figure} Figure (ref) plots three time series for the mixture specification alongside the (maximum) major stock-market volatility shock dates from Bloom2009. Each of the three series peaks around the Bloom2009 uncertainty dates in the late 1970s, early 80s and 90s, and 2008, but behave differently around the other dates. In particular, comparing with Figure (ref), fluctuations in the continuation entropy appear driven largely by fluctuations in the consumption-earnings ratio whereas the other series appear driven by both low- and high-frequency state variables. The Bloom2009 dates are essentially dates of stock market volatility shocks. The correlations of the three series with the CBOE S&P 100 Volatility Index (VXO) over the period 1986Q1 to 2018Q3 is 0.38 for the first two series (pessimism and continuation entropy) and 0.45 for the third (entropy of mixing weights). Another popular uncertainty measure are the JLN indices of macroeconomic uncertainty. The correlations of the indices of uncertainty of horizons 1, 3, and 12 months over the period 1960Q1 to 2018Q3 with our first uncertainty measure (pessimism) are all around 0.48, correlations with the second (continuation entropy) vary between 0.38 and 0.44, and correlations with our third measure (entropy of mixing weights) are all around 0.56. Correlations with the JLN indices of financial uncertainty display similar patterns but are weaker. \section{Conclusion} This paper studies identification and estimation of a class of dynamic models where the DM is endowed with multiplier or constraint preferences as in the “robustness” literature. The DM entertains a set of models surrounding a benchmark model that he or she fears may be misspecified. Decisions are evaluated under a worst-case model delivering lowest utility within this set. This paper derives primitive conditions for identification of the DM's worst-case model and preference parameters. The key step in the identification analysis is to establish existence and uniqueness of the DM's continuation value function allowing for unbounded statespace and unbounded utilities, both of which are important in applications. Extensions to models featuring other types of ambiguity aversion are discussed. For estimation, a perturbation result is derived which provides a necessary and sufficient condition for consistent estimation of continuation values and the worst-case model and allows convergence rates of estimators to be characterized. The result is also useful for computing approximate value functions in models for which no closed form solution exists by perturbing simpler models. An empirical application studies an endowment economy where the DM's benchmark model aggregates experts' forecasting models. Asset pricing consequences are discussed and some connections are drawn with the literature on macroeconomic uncertainty. Extensions of some results to models with learning have been sketched; we plan to pursue this in more detail going forwards. \singlespacing \putbib
bibunit