EconBase
← Back to paper

Scenario Synthesis and Macroeconomic Risk

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

85,959 characters · 27 sections · 42 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

\emergencystretch 3em

center[center omitted — 1,463 chars of source]

We introduce methodology to bridge scenario analysis and model-based risk forecasting, leveraging their respective strengths in policy settings. Our Bayesian framework addresses the fundamental challenge of reconciling judgmental narrative approaches with statistical forecasting. Analysis evaluates explicit measures of concordance of scenarios with a reference forecasting model, delivers Bayesian predictive synthesis of the scenarios to best match that reference, and addresses scenario set incompleteness. This underlies systematic evaluation and integration of risks from different scenarios, and quantifies relative support for scenarios modulo the defined reference forecasts. The framework offers advances in forecasting in policy institutions that supports clear and rigorous communication of evolving risks. We also discuss broader questions of integrating judgmental information with statistical model-based forecasts in the face of unexpected circumstances.

{\bf Keywords:} Macroeconomic Forecasting, Mixtures of Scenarios, Misclassification Rates, Entropic Tilting, Bayesian Predictive Synthesis, Judgmental Forecasting, Forecast Risk Assessment

\setstretch{1.0} \thispagestyle{empty} \if0\blind { \footnotetext{\\[3pt] {\bf Disclaimer}: The views expressed in this paper are those of the authors and do not necessarily reflect the views and policies of the Board of Governors, the Federal Reserve System, or the International Monetary Fund, its Management, or its Executive Directors.} \setcounter{footnote}{0} } \fi

Introduction

Macroeconomic policy institutions such as central banks rely heavily on forecasting methods. Monetary policymakers are regularly briefed on the economic outlook, alternative policy paths, and the balance of risks around the central forecast. Central bank staff rely on a combination of structural macroeconomic models, reduced-form empirical models, and judgmental approaches to prepare such monetary policy briefings. The central forecast is used as a basis for alternative policy path discussions, and the balance of risks is discussed more loosely based on scenario analysis.

The Bank of England pioneered communication of risk with fan charts in 1993; the Inflation Reports show central projections of inflation with charts that reflect uncertainty. Uncertainty intervals are derived from judgmental assessments of risk around the baseline BankofEngland1998. Since 1995, the U.S. Federal Reserve's Tealbook (TB) has presented scenarios as perturbations around baseline forecasts. Most major central banks now use some variant of these approaches. Fan charts and scenario analysis pose practical challenges as they require frequent updating and quantification of risks based on judgment. Hence central banks are relying more often on statistical methods to forecast macroeconomic risk. The density forecasting approach of \lq\lq Growth-at-Risk” (GaR: AdrianEtal2016,Adrianetal2019,PlagborgEtal2020,Adrianetal2022) is increasingly popular. The Tealbook has included GaR measures together with scenarios since 2017; other central banks have also implemented GaR approaches in addition to the more judgmental scenario methods FIGUERES2020109126, lenza2023density, GaR_BoE2019,GaR_BoE2024, Anesti2023, BdFGaR2022, BoIGaR2019.

Our focus here is on a formal statistical approach to integrating scenario-based balance of risk discussions with statistical forecasts. The methodology defines a synthesis of the baseline and scenarios that best match the statistical {\em reference} forecast distribution, the latter typically from GaR and/or a quantile regression model. The scenario synthesis assigns weights to each scenario, quantifying their relative concordance with the reference and so providing explication of why a certain set of scenarios is particularly relevant. The analysis also incorporates a synthetic \lq\lq backstop” scenario designed to address potential incompleteness of the defined scenario set. In practice, uncertainty measures are usually published only for the baseline; alternative scenarios are typically represented only in terms of point forecasts. We use extensions of the Bayesian decision analysis method of entropic tilting RobertsonET2005,TallmanWestET2022 to define full scenario forecast distributions as perturbations the baseline.

Analysis further addresses scenario information beyond a single point forecast, specifically to use scenario tail percentiles that reflect measures of scenario risk. This links to the desirability of scenario hypotheses that represent more radical perturbations of the baseline than has been typical (e.g., justiniano2008time, fernandez2011risk, AdrianBoyarchenko12, HeKrishnamurthy12a, BrunnermeierSannikov14, fernandez2023financial, fernandez2024search with structural models, and adrian2021multimodality, Caldara2021, carriero2024capturing with reduced-form models). That scenarios considered by policy institutions often represent only modest perturbations of the baseline is also partly addressed by our use of the synthetic backstop scenario. This can serve as a \lq\lq red flag" when the scenario set fails to account for risks-- especially tail risks-- supported under the statistical reference.

Relating baseline and scenario forecasts to the statistical reference exploits Bayesian predictive synthesis McAlinnWest2019,JohnsonWest2024 to motivate a discrete mixture (linear pool) of baseline and scenario distributions as a proxy for scenario-based forecasting. The \lq\lq match" of such a mixture with the statistical reference distribution uses a central measure of concordance between distributions, namely the {\em expected misclassification rate} (EMR). Identifying mixture weights to optimize EMR is then a formal Bayesian decision problem. Relative scenario probabilities based on this scenario:reference optimization guide evaluation and interpretation of the roles of scenarios. The analysis includes explicit statistical measures of {\em scenario set incompleteness} reflecting aspects of lack-of-concordance the scenario synthesis with the reference. This aids in the policy setting on the question of whether the baseline and chosen scenarios adequately reflect all the risks captured by the reference.

The case study draws on published versions of the TB. We use data from reports prepared for the December FOMC meetings in 2007 and 2018, giving predictions for 2008 and 2019, respectively. Following the TB, we focus on risks to real growth. Reanalysis incorporating other variables, such as inflation and unemployment at risk AdamsEtal2021,Kiley2022,Loria2024Inflation is straightforward but beyond our main scope here. Our detailed examples highlight the generation of scenario weights reflecting aspects of concordance with the reference, and also the questions of scenario set incompleteness. One example of that latter highlights the lack of a very negative, \lq\lq downside risk scenario" in both the 2007 and 2018 TB. This relates to the particular interest in our analysis when economic uncertainty is high so that defining an adequate baseline forecast is challenging. Then, listing and discussing a range of plausible scenarios, each with an assigned probability derived from the reference match, offers a richer perspective on informed decision-making under uncertainty.

Our analysis takes baseline and scenarios (as well as the reference) as given. In policy practice, of course, the back-and-forth between changes to statistical forecast distributions and the evolving narrative of scenarios provides a rich ground to rigorously examine shifts in the balance of risks. This was noted by Bernanke2023 and is germane to the TB, where scenario-based approaches to the balance of risk and statistical forecast distributions are discussed separately. As Federal Reserve Chair Jerome Powell noted during the Press Conference following the January 2025 FOMC meeting, {\em \color{RoyalBlue4} \lq\lq One of the things our staff does is they look at a range of possible outcomes. [...] There’ll be baseline, and then they’ll show six or seven alternative scenarios, including really good ones and not so good ones. And what those do is they spark [...] the policymakers to sort of think and understand about [...] the uncertainties that surround us.”} Our methodology provides formal cross-talk that can aid macroeconomic staff in policy institutions: it combines the communicative strength of narrative scenarios with the statistical rigor of predictive models, identifying the most relevant risks with easy-to-understand stories and quantifying the relevance of these stories.

Section (ref) discusses foundations and overviews methodology. Section (ref) addresses partial scenario information. Section (ref) introduces expected misclassification rates as distributional concordance metrics, with foundational insights. Section (ref) develops the embedding of scenario analysis in a fully Bayesian framework, with core theoretical summaries and aspects of computational implementation. Section (ref) summarizes key aspects of the detailed case study. Section (ref) links to broader questions of combining judgmental information with statistical model-based forecasts. The Appendix adds technical and methodological details. Summary comments define Section (ref).

Setting, Foundations and Perspective

Context and Goals

Interest lies in forecasting a vector outcome $\mathbf{y}$, such as a path of several macroeconomic indicators over multiple future time periods, based on the following ingredients.

itemize[itemsep=1pt,topsep=3pt] • A policy-based analysis produces a predictive density $p_0(\mathbf{y})$, referred to as the \bem{baseline density}. • Relative to the baseline, the policy analysis considers each of a set of \bem{alternative scenarios}; scenario $j$, labeled $\mathcal{S}_j$, generates a predictive density $p_j(\mathbf{y})$. These are regarded as hypothetical scenarios to be assessed relative to the baseline. • The baseline is a given forecast distribution in the policy setting, so not an hypothetical scenario; that understood, we use $\mathcal{S}_0$ and the index $j=0$ to designate the baseline. • Separately, a statistical model (e.g., the statistical GaR analysis) produces a full predictive density $p(\mathbf{y})$, referred to as the \bem{reference predictive density}.

The over-arching goal is to identify \lq\lq closeness" of each scenario to the reference $p(\mathbf{y}),$ and rank them relative to that assessment. The methodology we introduce addresses this, building on foundational statistical concepts and model developments now discussed.

Scenario Mixtures and Bayesian Predictive Synthesis

A Bayesian decision-maker in the policy setting can regard the set of scenario p.d.f.s $p_j(\mathbf{y})$ as \lq\lq information" to use in forming a policy-relevant overall forecast. This involves some form of \bem{pooling} of the predictions across the baseline and alternative scenarios. Here the foundational theory of Bayesian predictive synthesis (BPS)-- and the specific class of \lq\lq mixture BPS” models (McAlinnWest2019, section 2.2; JohnsonWest2024)-- applies. Under BPS, a valid Bayesian predictive analysis can be based on a \bem{scenario mixture}, i.e., a distribution with p.d.f. $f(\mathbf{y}|{\bm\alpha})$ that is a linear pool of the $p_j(\mathbf{y})$ with respect to probability weights $\alpha_j$ in a vector ${\bm\alpha},$ namely $f(\mathbf{y}|{\bm\alpha}) \propto \sum_j \alpha_j p_j(\mathbf{y}).$

A key theoretical aspect of mixture BPS is that it can address the broad question of \lq\lq scenario set (in-)completeness". That is, a setting in which the baseline $\mathcal{S}_0$ and all of the the alternative scenarios $\mathcal{S}_j$ considered are discordant with the reference $p(\mathbf{y}).$ This relates to the \lq\lq model set incompleteness” issue widely discussed in Bayesian econometrics. BPS theory addresses this by requiring an additional p.d.f. to extend the initial set and to use in the mixture. This has been exploited in BPS applications-- and in its generalization to decision-guided settings (BPDS: TallmanWest2023,ChernisTallmanKoopWest2024)-- by structuring the additional p.d.f. as a \bem{backstop} that can be expected to be supported by future data that is not so well-predicted by the initial model set. Examples in the above studies use an over-dispersed average of the initial mixture of model p.d.f., and this strategy can be adopted for scenario analysis. The specific construction of such a backstop scenario in the case study in Section (ref) provides an example of this modelling strategy.

Index the alternative scenarios from the policy setting by $j=1:J-1$ with the baseline $j=0$ and now with $j=J$ for the chosen backstop p.d.f. The latter is labeled $\mathcal{S}_J$ though it is a purely synthetic scenario chosen for the above purposes. Then the overall \bem{scenario mixture} p.d.f. is

equation[equation omitted — 113 chars of source]

Incomplete Specification of Scenario Forecast Distributions

Scenario p.d.f.s $p_j(\mathbf{y})$ are typically only partially specified. A common setting is that $\mathcal{S}_j$ defines point forecasts such as means or medians, with or without uncertainty measures such as a few other percentiles. The foundational concept is that the alternative scenarios represent economically relevant \lq\lq what-if?” perturbations of the baseline. Hence receiving such partial information on $\mathcal{S}_j$ indicates a modification of $p_0(\mathbf{y})$ to match that partial information. Our approach aims to identify $p_j(\mathbf{y})$ that is \lq\lq closest to" the baseline $p_0(\mathbf{y})$ subject to being consistent with that partial scenario information. The theoretical basis for methodology to do this, detailed in Section (ref), is that of entropic tilting (ET: TallmanWestET2022). Since its introduction by RobertsonET2005, ET--based methodology has seen increasing use in forecasting in econometrics, finance and related areas KrugerET2017,MetaxoglouPettenuzzoSmith2018,KoopMcIntyreMitchell2019,AntolinDiazPetrellaRubioRamirez2021,ClarkGanicsMertens2022,West2023constrainedforecasting,crump2021large. The current setting is different, though use here of ET is close in spirit and goals to its original use in imposing constraints on a given-- here the baseline-- forecast distribution.

Scenario-Reference Concordance

The goal of measuring concordance of scenarios with the statistical reference is now that of relating $f(\mathbf{y}|{\bm\alpha})$ in \eqn{scenariomixturef} to the reference p.d.f. $p(\mathbf{y})$. This is addressed by identifying the probability vector ${\bm\alpha}=(\alpha_0,\ldots,\alpha_J)'$ such that the scenario mixture is \lq\lq closest to" $p(\mathbf{y}).$ This requires specification of a utility function to characterize and quantify \lq\lq close" in comparing densities, and then the resulting methodology to evaluate ${\bm\alpha}$ and thus define both scenario-specific weights and the overall mixture synthesis. Section (ref) introduces a foundational metric for this-- based on a measure of concordance of $f(\mathbf{y}|{\bm\alpha})$ and $p(\mathbf{y})$ from traditional statistical classification. With some new and relevant theoretical results and motivating examples, this underlies its use in scenario synthesis.

Partial Scenario Information and Entropic Tilting

Partial Scenario Information

As noted in Section (ref) the common setting is that for each scenario only partial information relative to the fully specified baseline is provided. In many examples, the partial information can be represented as expectations of functions of $\mathbf{y}$, and this is the setting we adopt. Often, only the perturbed central tendency is reported. If taken as a mean, it would be a constraint on the expected value directly. If taken as the median, then it is formally defined as the expectation of an indicator function. Similar reasoning applies to other percentiles. Our analysis below addresses multiple scenario features simultaneously, such as a set of percentiles.

Suppose that $\mathcal{S}_j$ provides partial information on $p_j(\mathbf{y})$ in terms of $\mathbf{m}_j = \textrm{E}[\mathbf{s}_j(\mathbf{y})|\mathcal{S}_j]$ where $\mathbf{s}_j(\mathbf{y})$ is a $q_j-$vector of \bem{scenario scores}; call the given vector $\mathbf{m}_j$ the \bem{target score} for $\mathcal{S}_j$. In general, the definition of scores can be scenario-specific, but here we assume that $q_j=q$ and $\mathbf{s}_j(\mathbf{y})=\mathbf{s}(\mathbf{y})$ for all $j=\seq 1J$. A main case of interest has elements of $\mathbf{s}(\mathbf{y})$ as indicator functions in one or more of the univariate dimensions; then $\mathbf{m}_j$ is a given vector of percentiles of $p_j(\mathbf{y})$ in those dimensions. With $\mathcal{S}_j$ regarded as a perturbations of the baseline, methodology aims at identifying that $p_j(\mathbf{y})$ closest to the baseline $p_0(\mathbf{y})$ subject to being consistent with the forecast information $\mathbf{m}_j$. Entropic tilting (ET) results if we choose to define \lq\lq close to" in a K\"ullback-Leibler (KL) sense.

Entropic Tilting and Scenario-Baseline ET Weights

ET--based methodology, recently exploited in new ways in Bayesian predictive decision synthesis ChernisTallmanKoopWest2024, TallmanWest2023,TallmanWest2024, was originally used in imposing constraints on forecast distributions; that is the context here. In our setting, ET aims to identify $p_j(\mathbf{y})$ to minimize the KL divergence of the baseline $p_0(\mathbf{y})$ from $p_j(\mathbf{y})$ subject to $\mathbf{m}_j = \textrm{E}[\mathbf{s}_j(\mathbf{y})|\mathcal{S}_j] = \int_\mathbf{y} \mathbf{s}_j(\mathbf{y})p_j(\mathbf{y})d\mathbf{y}$. ET theory TallmanWestET2022 yields \beq{ETpj} p_j(\mathbf{y}) = k_j e^{{\bm\tau}_j'\mathbf{s}_j(\mathbf{y})} p_0(\mathbf{y})\quadwhere\quad k_j^{-1} = \int_\mathbf{y} e^{{\bm\tau}_j'\mathbf{s}_j(\mathbf{y})} p_0(\mathbf{y})d\mathbf{y},\end{equation} in which ${\bm\tau}_j$ is the (provably unique) \bem{tilting vector} such that the expectation constraint is satisfied. The implied identity $\mathbf{0} = \int_\mathbf{y} \{\mathbf{s}_j(\mathbf{y})-\mathbf{m}_j\} \textrm{exp}\{{\bm\tau}_j'\mathbf{s}_j(\mathbf{y})\} p_0(\mathbf{y})d\mathbf{y}$ is typically efficiently solved for ${\bm\tau}_j$ using simple Newton-Raphson.

In practice, it is typical that the baseline is represented in terms of a Monte Carlo (MC) sample, i.e., defined as a discrete distribution $\{ \mathbf{y}^i, w_0^i \}_{i=\seq 1n}$ with support points $\mathbf{y}^i$ having weight (probability) $w_0^i.$ This is particularly key in our setting as we will later use importance sampling to evaluate $p_0(\mathbf{y})$ relative to the statistical reference $p(\mathbf{y})$. Then expectations defining the ET tilting vectors ${\bm\tau}_j$ are trivially evaluated via simple Monte Carlo integration.

ET analysis can be regarded as using $p_0(\mathbf{y})$ as an importance sampling proposal with respect to a target p.d.f. $p_j(\mathbf{y}).$ This was recognized by RobertsonET2005 and provides useful numerical checks on consistency of the scenario-specific moment constraints with the baseline. On sample values $\mathbf{y}^i$, the implied normalized IS weights for MC integration in \eqn{ETpj} are $w_j^i \propto u_j^i w_0^i$ where $u_j^i\propto p_j(\mathbf{y}^i)/p_0(\mathbf{y}^i) = \textrm{exp}\{{\bm\tau}_j'\mathbf{s}_j(\mathbf{y}^i)\}$. The $u_j^i$ are called \bem{ET weights}. The standard expected sample size (ESS) can be evaluated on the $u_i$. ESS-- the reciprocal of the sum of squared $u_j^i$ over $i=\seq 1n$-- provides an overall assessment of concordance of the $\mathcal{S}_j$ constraints with the baseline TallmanWestET2022. This relates closely to the minimized KL divergence (e.g., GruberWest2016BA; GruberWest2017ECOSTA) but on an interpretable scale.

Predictive Concordance

Predictive Concordance and Misclassification Rates

Predictive concordance mooted in Section (ref) is presented here in a general setting comparing two density functions $p(\mathbf{y})$ and $f(\mathbf{y}).$ The scenario mixture setting then arises with $f(\mathbf{y})$ replaced by $f(\mathbf{y}|{\bm\alpha})$ of \eqn{scenariomixturef} for any given ${\bm\alpha}.$ Assume that $p(\mathbf{y})$ and $f(\mathbf{y})$ have the same support.

Suppose a random draw $\mathbf{y}$ is made from either $f(\mathbf{y})$ or $p(\mathbf{y})$ with equal probabilities. It is not disclosed which distribution generates the outcome $\mathbf{y}.$ Write $\mathcal{H}_p$ for the hypothesis that $\mathbf{y}\simp(\mathbf{y}),$ and $\mathcal{H}_f$ for the hypothesis that $\mathbf{y}\simf(\mathbf{y}).$ Since the choice is made with $\textrm{Pr}(\mathcal{H}_p)=\textrm{Pr}(\mathcal{H}_f)=0.5,$ the resulting posterior probabilities conditional on the observed $\mathbf{y}$ are $ \textrm{P}(\mathcal{H}_p|\mathbf{y}) = p(\mathbf{y})/\{p(\mathbf{y})+f(\mathbf{y})\}$ and $\textrm{Pr}(\mathcal{H}_f|\mathbf{y})=1-\textrm{P}(\mathcal{H}_p|\mathbf{y}).$

Now assume that $\mathbf{y}$ is actually a draw from $p(\mathbf{y})$, i.e., condition on $\mathcal{H}_p$. Before learning $\mathbf{y},$ the expected posterior probability on $\mathcal{H}_f$ is then \beq{Eppyf} \pi_{pf} \equiv E[P(\mathcal{H}_f|\mathbf{y}) |\mathcal{H}_p] = \int_\mathbf{y} P(\mathcal{H}_f|\mathbf{y}) p(\mathbf{y}) d\mathbf{y} = \int_\mathbf{y} \frac{\fyp(\mathbf{y})}{\{f(\mathbf{y})+p(\mathbf{y})\}}d\mathbf{y}. \end{equation} By symmetry, if $\mathbf{y}$ is actually from $\mathcal{H}_f,$ the expected posterior probability $\pi_{fp}= \textrm{E}[\textrm{P}(\mathcal{H}_p|\mathbf{y}) |\mathcal{H}_f]$ is obviously the same, $\pi_{fp} = \pi_{pf}$.

Predictive concordance of $f(\mathbf{y})$ with $p(\mathbf{y})$ is inherently measured by the \bem{expected misclassification rate (EMR)} $\pi_{pf}$. Higher values indicate that it is difficult to discriminate $f(\mathbf{y})$ from $p(\mathbf{y})$-- indicating that draws from $f(\mathbf{y})$ are more likely to be misclassified as coming from $p(\mathbf{y})$-- and vice-versa. This is a natural, interpretable metric to assess concordance-- or discordance-- of the two distributions.

In traditional classification in statistics and machine learning, the optimal Bayesian classifier judges $\mathbf{y}$ as coming from $f(\mathbf{y})$ with probability $\textrm{P}(\mathcal{H}_f|\mathbf{y}).$ Averaging across $\mathbf{y}\sim p(\cdot),$ and using standard terminology, $1-\pi_{pf}$ is then both the population {\em sensitivity} and (due to the comparison of just two distributions and the implied symmetry) the population {\em sensitivity} of the optimal Bayesian classifier. It follows that $1-\pi_{pf}$ is the traditional overall {\em accuracy} of the test comparing $f(\cdot)$ and $p(\cdot)$, and so EMR $\pi_{pf} =1-\textrm{\em accuracy}$ is the traditional {\em error rate}. Increasing EMR indicates decreased discrimination of $f(\cdot)$ from $p(\cdot)$. Judging $f(\cdot)$ to be \lq\lq close to" $p(\cdot)$ at higher values of $\pi_{pf}$ is thus theoretically fundamental and practically interpretable.

It is immediate that $\pi_{pf}\le 0.5$ with equality only when $f(\cdot)\equiv p(\cdot),$ defining the absolute scale for assessment of concordance. To prove this, note that $\pi_{pf} = \textrm{E}[r(\mathbf{y})/\{1+r(\mathbf{y})\} |\mathcal{H}_p]$ where $r(\mathbf{y})=f(\mathbf{y})/p(\mathbf{y})$ with $\textrm{E}[r(\mathbf{y}) |\mathcal{H}_p]=1.$ Now, $r/(1+r)$ is concave on $r>0$ so that $\pi_{pf} \le \textrm{E}[r(\mathbf{y})|\mathcal{H}_p]/\{1+\textrm{E}[r(\mathbf{y})|\mathcal{H}_p]\}=1/2.$ The upper bound is achieved when $f(\mathbf{y})\equiv p(\mathbf{y})$, i.e., $r(\mathbf{y})=1$ for all $\mathbf{y}.$

Now consider a decision setting where $f(\cdot)$ is to be chosen to be \lq\lq close to" $p(\mathbf{y}),$ and when $\mathbf{y}\sim p(\cdot)$. Choosing $f(\cdot)$ to maximize $\pi_{pf}$ subject to relevant constraints is the optimal decision with respect to the implied constrained version of utility function $\textrm{P}(\mathcal{H}_f|\mathbf{y})$. This defines the Bayesian foundation of use of EMR in the scenario synthesis development in Section (ref).

Relationships to K\"ullback-Leibler Divergence

Note that $\pi_{pf} = \textrm{E}[ 1/[1+\exp\{k(\mathbf{y})\}]|\mathcal{H}_p]$ where $k(\mathbf{y})=\log\{p(\mathbf{y})/f(\mathbf{y})\}$. Under $\mathcal{H}_p,$ the scalar random quantity $k(\mathbf{y})$ has expectation $\textrm{KL}(p\|f) \equiv \textrm{E}[k(\mathbf{y}) |\mathcal{H}_p] = \int_y \log\{p(\mathbf{y})/f(\mathbf{y})\}p(\mathbf{y}) d\mathbf{y}$, the K\"ullback-Leibler divergence {\em of} $f(\cdot)$ {\em from} $p(\cdot).$ Assuming this expectation is finite, the delta approximation yields $\pi_{pf}\approx 1/[1+\exp\{\textrm{KL}(p\|f)\}]$; thus choosing $f(\cdot)$ to maximize $\pi_{pf}$ is approximately the KL divergence minimizing solution. In many cases of practical relevance, this also provides a strict lower bound on $\pi_{pf},$ i.e., $\pi_{pf} \ge 1/[1+\exp\{\textrm{KL}(p\|f)\}];$ see Appendix (ref). Both the direct approximation and the lower bound are accurate in cases of higher concordance. Then, the symmetry of EMR in $f(\cdot)$ and $p(\cdot)$ implies that the same results hold with the two densities exchanged. With $\textrm{KL}(f\|p)$ the divergence of {\em of} $p(\cdot)$ {\em from} $f(\cdot),$ this immediately refines the lower bound to $\pi_{pf} \ge 1/[1+\exp(\kappa_{pf})]$ where $\kappa_{pf}= \min\{ \textrm{KL}(p\|f), \textrm{KL}(f\|p) \},$ with equality as the direct delta approximation. KL divergence always raise the question of directional definition. This does not arise in using $\pi_{pf}$ due to its symmetry, and this link to KL indicates the relevant \lq\lq symmetrization” of KL as $\kappa_{pf}.$ In cases of relatively good concordance, the two directional measures will also be close. Additional aspects of the relationship are discussed and exemplified in Appendix (ref).

EMR is fundamental for reasons discussed above; we have presented these connections to KL as it is a well-known measure. A major caveat is that it assumes KL measures are finite. There are important practical contexts where this is not so. An example has $f(\mathbf{y})$ Gaussian and $p(\mathbf{y})$ log T with any degrees of freedom; then $p(\mathbf{y})$ has no moments at all West2023constrainedforecasting and $\textrm{KL}(p\|f)$ is infinite. In contrast, $\pi_{pf}\in (0,0.5]$ always.

Scenario Synthesis

EMR and Optimizing Scenario Mixture Probabilities

The predictive concordance concept applies to the scenario mixture setting with $f(\mathbf{y})$ replaced by $f(\mathbf{y}|{\bm\alpha}) = \sum_{j=\seq 0J}\alpha_jp_j(\mathbf{y})$ at any chosen probability vector ${\bm\alpha}=(\alpha_0,\ldots,\alpha_J)'.$ Making dependence on ${\bm\alpha}$ explicit, \eqn{Eppyf} is now \beq{Eppyfalpha} \pi_{pf}({\bm\alpha}) = \int_\mathbf{y} \frac{\fyap(\mathbf{y})}{\{f(\mathbf{y}|{\bm\alpha})+p(\mathbf{y})\}}d\mathbf{y}. \end{equation} Values of ${\bm\alpha}$ yielding high values of $\pi_{pf}({\bm\alpha})$ define mixtures of baseline and scenarios \lq\lq close to” to the reference $p(\mathbf{y})$ in terms of probabilistic concordance. Suppose $\widehat{{\bm\alpha}}$ maximizes $\pi_{pf}({\bm\alpha})$ with maximum value $\widehat{\pi}_{pf}=\pi_{pf}(\widehat{{\bm\alpha}})$. Each element $\widehat\alpha_j$ of $\widehat{{\bm\alpha}}$ quantifies the extent to which $\mathcal{S}_j$ is concordant with the reference {\em\color{RoyalBlue4} relative to the other $\mathcal{S}_i$ for $i\ne j$}. The summary $\widehat{\pi}_{pf}$ is a concrete measure of the concordance of the set of predictions from the baseline and the scenarios combined. A low value of $\widehat{\pi}_{pf}$ indicates that none of the $\mathcal{S}_j$ nor their mixture are really concordant with the reference, relating to the scenario set incompleteness discussion of Section (ref). Thus $\widehat{\pi}_{pf}$ measures how \lq\lq discordant” the scenario set is with the reference statistical predictions. The weight $\widehat\alpha_J$ on the backstop provides additional information.

The framework addresses selection of ${\bm\alpha}$ as a decision problem that maximizes $\pi_{pf}({\bm\alpha})$ with regularization to penalize very small $\alpha_j.$ This is based on deeper foundational and theoretical development in the next subsection, and leads to choosing ${\bm\alpha}^*$ to maximize the objective function \beq{Logpostalpha} \lambda({\bm\alpha}) = \log\{\pi_{pf}({\bm\alpha})\} + \epsilon\sum_{j=\seq 0J} \log(\alpha_j), \quadsubject to \alpha_j>0 \ (j=\seq 0J) \quadand \sum_{j=\seq 0J}\alpha_j=1, \end{equation} where $\epsilon>0$ is a very small regularization parameter. As we now show, \eqn{Logpostalpha} is in fact the log of a formal posterior distribution so that the optimization seeks the posterior mode.

Bayesian Foundation

EMR is a Likelihood Function

Suppose that the economic reality $\mathbf{y}$ is generated from the reference $p(\mathbf{y})$ and consider an hypothetical/synthetic binary outcome $z$ generated from the Bernoulli distribution with success probability $\textrm{Pr}(z=1|\mathbf{y},{\bm\alpha}) = f(\mathbf{y}|{\bm\alpha})/\{p(\mathbf{y})+f(\mathbf{y}|{\bm\alpha})\}.$ Then $$p(z=1,\mathbf{y}|{\bm\alpha}) = \textrm{Pr}(z=1|\mathbf{y},{\bm\alpha}) p(\mathbf{y}|{\bm\alpha}) = \textrm{Pr}(z=1|\mathbf{y},{\bm\alpha})p(\mathbf{y}) = f(\mathbf{y}|{\bm\alpha})p(\mathbf{y})/\{p(\mathbf{y})+f(\mathbf{y}|{\bm\alpha})\}.$$ Now suppose you observe $z=1$ but not $\mathbf{y}$; EMR emerges via expectations over the \lq\lq missing data" $\mathbf{y},$ viz., $p(z=1|{\bm\alpha}) = \pi_{pf}({\bm\alpha}).$ Thus, $\pi_{pf}({\bm\alpha})$ is in fact a likelihood function for the {\em parameter} ${\bm\alpha}$ based on an hypothetical observation $z=1$ that classifies a random draw from $p(\mathbf{y})$ as coming from $f(\mathbf{y})$ under a 50:50 prior. The connection with the foundation of EMR in Section (ref) is immediate.

It follows that EMR-maximizer $\widehat{{\bm\alpha}}$ is a maximum likelihood estimate (MLE). Evaluating $\widehat{{\bm\alpha}}$ is probability simplex constrained convex optimization problem with a unique solution, the convexity and hence uniqueness being shown here in Appendix (ref). The solution will typically be a \bem{sparse mixture} of scenarios, with some zeros in $\widehat{{\bm\alpha}}.$ This follows from general results of optimization of convex functions over the probability simplex BoydVandenberghe2004. For some integer $k\in \{0:J\}$ a subset of $k$ of the $\widehat\alpha_j$ can be zero. There are cases when $k=0$ but $k>0$-- defining a sparse optimizing vector-- is more usual, especially with larger $J$ and diversity among the $p_j(\mathbf{y}).$ This relates to general features of optimization over the simplex; simplex constraints operate to shrink weights to the boundaries, effectively as $\ell_1$ shrinkage for sparsity brodie2009sparse. This underlies the notion of scalability of the analysis to larger numbers of scenarios.

However, sparsity in $\widehat{{\bm\alpha}}$ is unstable since it is not a genuine feature but is induced by the implicit prior $\ell_1$ penalty; its values are typically very sensitive to small changes in the input scenario and reference p.d.f.s. This pathology of sparsity inducing penalties was identified and documented in the context of forecasting by IllusionOfSparsity. In the current setting, take an example with two very similar scenario p.d.f.s; one of these scenarios will have a zero value in $\widehat{{\bm\alpha}},$ the other non-zero. Then, a very small change in either of the p.d.f.s-- or of the reference p.d.f.-- will flip the zero/non-zero pattern. At each of these extremes-- and for ranges of the $\alpha_j$ on these two scenarios bridging the extremes-- the resulting scenario mixture $f(\mathbf{y}|\widehat{{\bm\alpha}})$ will be almost unchanged. This sensitivity is undesirable; it is desirable to have similar probabilities on the two scenarios. The key point is that a uniform prior on ${\bm\alpha}$ favors overly sparse models when the likelihood function has modes at the simplex boundaries. This can be addressed by imposing additional constraints or, more foundationally, with a minimally informative \lq\lq regularizing" prior over ${\bm\alpha}$.

Priors and Penalties

The natural priors are Dirichlet, ${\bm\alpha}\sim \textrm{Dir}(\mathbf{a})$ having p.d.f. $p({\bm\alpha}) \propto \prod_{j=\seq 0J} \alpha_j^{a_j-1}$ over the simplex. Here $a_j>0$ for all $j$ and, with precision $a=\sum_{j=\seq 0J}a_j,$ the means are $a_j/a$ and prior joint mode has elements $\max\{0, (a_j-1)/(a-J-1)\}.$ A prior with each $a_j=1+\epsilon$ for a very small $\epsilon>0$ is \lq\lq minimally informative" subject to the joint prior mode being positive on each scenario. Modifications to $a_j=1+\epsilon_j$ to differentially favor scenarios {\em a priori} are obviously of interest, but for this paper the symmetric prior is adopted. For given $\epsilon,$ the prior joint mode and mean are then each $\mathbf{1}/(J+1),$ i.e., favoring a uniform set of scenario probabilities though with high uncertainty since $\epsilon$ is taken as very small. Under this prior ${\bm\alpha}\sim \textrm{Dir}(\mathbf{1}(1+\epsilon)),$ the log posterior is $\lambda({\bm\alpha})$ in \eqn{Logpostalpha}, up to an additive constant. The prior is zero at simplex boundaries, hence so is the resulting unimodal posterior. The posterior mode-- denoted by ${\bm\alpha}^*$-- maximizes EMR modified by the prior-based penalty that explicitly acts to move from the boundary zero MLE values in $\widehat{{\bm\alpha}}$ to small but non-zero values. This leads to more stable and robust results and addresses the issues discussed in the previous section.

Analysis requires choice of a (small) value of the regularizing hyper-parameter $\epsilon.$ Based on theory in Appendix (ref), the default recommendation is $\epsilon = c/(J+1),$ where $c=0.005.$ The value of $c$ can be modified somewhat up/down with minimal impact, while the scaling with number of scenarios is important in more heavily penalizing the MLE-based analysis in higher dimensions. Given $\epsilon>0,$ evaluation of the posterior mode ${\bm\alpha}^*$ to maximize \eqn{Logpostalpha} trivially modifies the probability simplex constrained convex optimization problem with a unique solution. See Appendix (ref).

It is also of interest to consider analyses with additional constraints on ${\bm\alpha}.$ A key example is to require $\alpha_0\ge \alpha_j$ for $j=\seq 1J,$ consistent with the view that the baseline is the \lq\lq modal" scenario. In general it is of interest to run comparative analyses with and without such constraints. Such a constraint simply modifies the Dirichlet prior by the indicator of the constraint; this does not impact the convexity of the optimization problem and is trivially implemented.

Monte Carlo Importance Sampling

Analysis {\em prima facie} relies on evaluating the p.d.f.s $p(\mathbf{y})$ and each $p_j(\mathbf{y}),$ and then performing the integration in \eqn{Eppyfalpha}. Analytic approximations to the integral may be explored. Specific approximations relate to measures of discriminatory information in classification using mixtures LinChanWest2015Biostatistics. In practice, however-- and as already noted in Section (ref)-- forecasts will typically be \lq\lq available” in terms of Monte Carlo (MC) samples, so direct evaluation of $\pi_{pf}$ by MC integration is a priority (and avoids concerns of assessing the quality of analytic approximations).

The analysis is implemented with the values of the p.d.f.s $p(\mathbf{y})$ and the $p_j(\mathbf{y})$ available only on a (large) sample of MC draws from the reference $p(\mathbf{y})$, a \bem{reference random sample}. The random sample $\mathbf{y}^i$, $(i=\seq 1n)$, is drawn from $p(\mathbf{y})$ and at the first step this defines an importance sample (IS) for the baseline $p_0(\mathbf{y})$ with normalized IS weights $w_0^i \propto p_0(\mathbf{y}^i)/p(\mathbf{y}^i).$ The discrete distribution $\{ \mathbf{y}^i, w_0^i \}_{i=\seq 1n}$ defines the MC approximation to the baseline for evaluation of expectations in the downstream analysis. A proviso is that $p(\mathbf{y})$ is a relevant importance sampling proposal; in particular, it should be heavier-tailed than $p_0(\mathbf{y})$. As in all applications of IS, monitoring efficiency measures such as the % effective sample size $\textrm{ESS}=n^{-1}100/\sum_{i=\seq1n} (w_0^i)^2$ provides guidance; initial analysis generating a relatively low ESS guides choice of a larger sample size. This IS analysis then underlies evaluation of scenario-specific ET parameters as in Section (ref), yielding ET weights $u_j^i\propto \textrm{exp}\{{\bm\tau}_j'\mathbf{s}_j(\mathbf{y}^i)\}$ on sample $\mathbf{y}^i$ defining $p_j(\cdot)$ relative to the baseline. Scenario-specific ESS measures using the $u_j^i$ weights are then relevant. In (rare) cases of a scenario that is really discordant with the baseline, a very low ESS indicates such. Refined but much more computationally demanding adaptive IS methods may be considered, but are outside our current scope. In any case, encountering such discordance would indicate that such a scenario might better be considered separately and its full distribution directly assessed.

This ET analysis leads to {\em compound weights} $w_j^i \propto u_j^iw_0^i$ relating $\mathcal{S}_j$ to the reference; these are called the \bem{ET$-$IS weights}. ESS measures can now also be evaluated on the $w_j^i$ to provide direct overall assessment of each $p_j(\cdot)$ relative to the reference $p(\cdot).$ Note that there can be cases where a scenario is more concordant with the reference than the baseline as some of our examples show. The reference sample and compound ET$-$IS weights are then ingredients in the direct evaluation of \eqn{Eppyfalpha} via MC integration.

Case Study

The case study draws from the Risk and Uncertainty analyses in the December \href{https://www.federalreserve.gov/monetarypolicy/files/FOMC20071211gbpt120071205.pdf}{2007} and \href{https://www.federalreserve.gov/monetarypolicy/files/FOMC20181219tealbooka20181207.pdf}{2018} Tealbooks Dec2007Tealbook,Dec2018Tealbook. The scenarios specify point forecasts for GDP growth, inflation, the unemployment rate, and other variables. The methodology applies to multiple variables and horizons, but this first application restricts attention to one-year ahead GDP growth, namely $\mathbf{y}=y$, now scalar. Analysis follows the processes discussed in the previous sections; Appendix (ref) gives a summary of the flow of analytic and computational details.

Reference Distribution

Among recent statistical approaches to risk assessment, Adrianetal2019 develop quantile regression models of conditional predictive distributions and show that financial markets provide useful risk information. This approach has influenced practice, being adopted for conditional one-year ahead forecasts of GDP growth, unemployment rate, and inflation, for example, by the Federal Reserve Board TBTVMR in the \lq\lq Time-Varying Macroeconomic Risk” exhibit in the Risk and Uncertainty section of Tealbook A, and the New York Federal Reserve Outlook@Risk in \href{https://www.newyorkfed.org/research/policy/outlook-at-risk}{Outlook-at-Risk}. Other central banks and international financial institutions have also adopted GaR approaches FIGUERES2020109126, lenza2023density, IMF2017, GaR_BoE2024, Anesti2023, BdFGaR2022, BoIGaR2019, BancoEspana2022, Bundesbank2023.

wraptable{r}{.5\textwidth}\caption{Reference percentiles} {.0\textwidth} \begin{tabular}{C{.02\textwidth}L{.07\textwidth}C{.09\textwidth}C{.09\textwidth}C{.01\textwidth}|C{.02\textwidth}L{.07\textwidth}C{.09\textwidth}} \hline \hline & & \multicolumn{2}{c}{NY Fed} &\multicolumn{2}{c} & & {Tealbook} \\ \hline & & \color{RoyalBlue4} 2007 & \color{RoyalBlue4} 2018 &&& & \color{RoyalBlue4} 2018 \\ \hline & \color{RoyalBlue4} P10: & $-$1.7\phantom{\ \ \,} & 0.0 &&& \color{RoyalBlue4} P5 & 0.7 \\ & \color{RoyalBlue4} P25: & 0.2 & 1.1 &&& \color{RoyalBlue4} P15 & 1.3 \\ & \color{RoyalBlue4} P50: & 1.8 & 2.1 &&& \color{RoyalBlue4} P50 & 2.5 \\ & \color{RoyalBlue4} P75: & 3.3 & 3.0 &&& \color{RoyalBlue4} P85 & 3.6 \\ & \color{RoyalBlue4} P90: & 4.8 & 4.0 &&& \color{RoyalBlue4} P95 & 4.3 \\ \hline \end{tabular} \begin{tabular}{p{.46\textwidth}}\scriptsize The NY Fed Outlook at Risks gives point forecasts for one-year ahead GDP growth. The Tealbook Time-Varying Macroeconomic Risk gives one-year ahead Tealbook forecast errors in the December 2018 Tealbook, which provides GDP growth point forecasts once the baseline forecast is added. \end{tabular}

The Federal Reserve Board started producing the Tealbook Time-Varying Macroeconomic Risk forecasts in 2017 but do not provide past values. The NY Fed started producing the Outlook-at-Risk forecasts only in 2023 but provides past values starting in 1989. As a result, our case study constructs and compares reference distributions from both the NY Fed and the Fed Board. In practice, both the NY Fed and the Tealbook give five predictive percentiles (Table (ref)). To construct the reference distribution, we fit skew-t distributions AzzaliniCapitanio2003 on these percentiles. Following Adrianetal2019, the four parameters of the skew-t minimize the squared distance between the given reference quantiles and those of the resulting skew-t. Table (ref) reports resulting parameters.

Baseline Distribution

Table (ref) shows baseline and alternative scenario projections. The Tealbook provides the point forecast and 70% intervals for the baseline. To construct the baseline distribution, we take the point forecast as the median and extremes of the 70% confidence interval as the 15th and 85th percentiles, respectively. We fit a skew-t with 50 degrees of freedom to these percentiles-- the choice of degrees of freedom allows some modest tail-weight beyond normal but effectively represents a \lq\lq close to normal" distribution.

table[table omitted — 2,346 chars of source]
figure[figure omitted — 754 chars of source]
wraptable{r}{.6\textwidth}\caption{Reference and baseline skew-t parameters} {0\textwidth} \begin{tabular}{C{.05\textwidth}L{.13\textwidth}L{.085\textwidth}R{.07\textwidth}R{.07\textwidth}R{.07\textwidth}R{.07\textwidth}R{.01\textwidth}}\hline\hline &Distribution & Type & \color{RoyalBlue4} lc & \color{RoyalBlue4} sc & \color{RoyalBlue4} sk & \color{RoyalBlue4} df &\\\hline \multirow{3}{*}{\rotatebox{90}{\color{RoyalBlue4} \ 2007}} &&&&&&&\\[-10pt] &Reference & NY Fed & 2.7 & 2.2 & $-$0.5 & 3.4 &\\[1pt] &Baseline & & 1.3 & 1.1 & 0.0 & 50.0 &\\[-10pt] &&&&&&&\\\hline \multirow{4}{*}{\rotatebox{90}{\color{RoyalBlue4} \ 2018}} &&&&&&&\\[-10pt] &Reference & NY Fed & 2.5 & 1.3 & $-$0.3 & 3.0 &\\[1pt] &Reference & Tealbook & 2.1 & 1.1 & 0.5 & 50.0 &\\[1pt] &Baseline & & 1.2 & 1.9 & 2.1 & 50.0 \\[-10pt] &&&&&&&\\\hline \end{tabular} \begin{tabular}{p{.555\textwidth}}\scriptsize The skew-t parameters are those for location (lc), scale (sc), skewness (sk) and degrees of freedom (df). \end{tabular}

Table (ref) shows parameters of the baseline and reference skew-t distributions; Figure (ref) shows the p.d.f.s. The baseline is much more precise than the NY Fed reference, with a lower scale and higher degrees of freedom. This raises questions for economic forecasting and policy design. In a world with a known \lq\lq true model" of the economy, the baseline and reference would be identical; as they differ in practice, interpreting the scenarios is the challenge. By comparison, the baseline is roughly as precise as the Tealbook reference, the latter having 50 degrees of freedom as a result of the optimization process in fitting the skew-t. This suggests that the Tealbook reference may underestimate risk. As Federal Reserve Chair Jerome Powell said during the Press Conference following the January 2025 FOMC meeting, we should not be surprised as {\em \color{RoyalBlue4} \lq\lq it is human nature, apparently, to underestimate [...] how fat the tails are.[...] We think of things in a normal distribution. And in the economy, it’s not a normal distribution.”}

Given reference and baseline distributions, our analysis proceeds based on Monte Carlo sampling from the reference. In developments below, the MC sample size is $10^6$ and resulting MC analysis summaries stable and robust across reanalyses with such a large sample size.

Scenarios

TB scenarios in Table (ref) provide only point forecasts. In the Dec. 2007 (2018) TB, we have 6 (4) alternative scenarios.\footnote{We exclude \lq\lq More room to grow" scenario in 2007, and both \lq\lq Supply constraints" and \lq\lq Lower oil prices" scenarios in 2018. These had GDP point forecasts identical to either the baseline or another scenario. If included in our synthesis, they would receive equal weight with either the baseline or one other scenario. } Figure (ref) shows scenarios on the baseline, indicating their concentration around the baseline median but with some indication of downside risk in the left tail.

Our first analysis treats the scenario point forecasts as medians\footnote{TB and other point forecasts might alternatively be treated as modes of scenario distributions. Our analysis can address that, based on new theoretical results (not reported here) showing why and how entropic tilting can be applied when point forecasts are modes. We use medians, however, based on fundamental concern for deeper representation of the probability distributions of scenarios, and embedding in more detailed analyses with multiple percentiles.} of the $p_j(y)$, and the ET construction maps the baseline to each scenario p.d.f. constrained to its specified median only; see the P50s in Table (ref). While the TB provides no measures of uncertainty around the scenario projections, such information could be useful and available in other applications. Section (ref) explores analyses with P15 and P85 constructed for each scenario.

As discussed in Section (ref), we augment the scenario set with a backstop located at the center of the scenarios while being relatively over-dispersed. Since the scenario information here is restricted to the median point forecasts, we first construct $p_j(y)$, $j=1,\ldots,J$, using ET as in Section (ref), then use the implied percentiles to define those of the backstop. Specifically, the backstop has $\text{P50}_{\text{B}}=\text{\small median}_{j=1,\ldots, J}\text{P50}_j$, $\text{P15}_{\text{B}}=\min_{j=1,\ldots, J}\text{P15}_j$, and $\text{P85}_{\text{B}}=\max_{j=1,\ldots, J}\text{P85}_j$, respectively. Numerical details are in Table (ref).

figure[figure omitted — 812 chars of source]

Figure (ref) shows the resulting $p_j(y)$ for four of the 2007 TB scenarios. Table (ref) shows corresponding ET--based ESS measures. When a scenario is close to the baseline (e.g., $\mathcal{S}_1$ and $\mathcal{S}_5$), the tilted distribution remains close and slightly asymmetric; otherwise, the tilted distribution can exhibit skewness and multimodality due to the mismatch between the scenario forecasts and baseline. The emergence of interesting shapes and multimodality in scenario distributions indicates the hypothesized state of the economy in the scenario is in regions poorly supported by the baseline. This is also related to the concept of modest policy intervention of LeeperZha, i.e. that we can analyze policy effects using the baseline as long as entertained policy interventions are small enough that economic agents would not change their behavior in response to the intervention. Related considerations are those of formally down-weighting extreme conditional assumptions in conditional forecasting AntolinDiazPetrellaRubioRamirez2021,ChernisTallmanKoopWest2024.

Scenario Synthesis based on Medians

Table (ref) shows optimal scenario mixture weights and other summaries from analyses. There is strong concordance between the mixture synthesis and the reference in the case of the 2018 TB (EMR=0.48), while the concordance is weaker in the 2007 Tealbook (EMR=0.43). Figure (ref) provides insights via comparison of p.d.f.s. For TB 2008, the mixture synthesis is light-tailed relative to the reference, as all scenarios lack probability on downside GDP ranges that the reference meaningfully supports. For the 2018 example, the scenario mixture supports positive GDP values-- partly as the baseline is already right skewed-- but assigns relatively limited support to negative GDP growth.

table[table omitted — 4,513 chars of source]
figure[figure omitted — 989 chars of source]

In the 2007 TB analysis, the baseline, $\mathcal{S}_4$ and backstop are roughly equally weighted, followed by $\mathcal{S}_2$. The two more extreme scenarios $\mathcal{S}_2$ and $\mathcal{S}_4$ get weight as they help to capture the spread of the reference, while the other scenarios get small weights-- they sit in the center of the reference and are very similar to the baseline, as shown by the ET ESS measures. In contrast, in the 2018 TB example the reference is more precise and the baseline is heavier weighted than other scenarios. Here the EMR of the baseline alone is quite high, so the alternative scenarios add small but limited value in contributing to approximating the reference.

A key feature of analysis is that $100{-}\textrm{ESS}$ for the synthesis is an absolute measure of scenario set incompleteness relative to the reference. In TB 2007, the ESS of $f(\mathbf{y}|{\bm\alpha}^*)$ is about 71-72%; we can say that the scenario set is about 28-29% incomplete. In contrast, in TB 2018 the ESS of $f(\mathbf{y}|{\bm\alpha}^*)$ is about 91%, which indicates that the TB scenarios much more adequately represent the risk and uncertainty in the economy as defined by the reference than in 2007. A hint of the inability of TB 2007 scenarios to properly capture the risks in the reference comes also from the substantial weight on the backstop; in contrast, in TB 2018, the backstop receives low weight. As a general point looking ahead, a backstop p.d.f. that is over-dispersed has general benefits, but in application other choices are possible and may be preferred. These examples highlight this, indicating that scenarios reflecting increased support in the upper and lower tails of the-- anticipated-- reference distribution are most relevant. Future applications might address this.

table[table omitted — 2,365 chars of source]

\FloatBarrier

Table (ref) summarizes 2018 analysis using the reference distribution based on percentiles from the Tealbook Time-Varying Macroeconomic Risk to compare with the NY Fed-based details above. Again the synthesis here uses only the scenario medians. In this case, the baseline is as precise as the reference and has similar tail-weight in terms of the skew-t degrees of freedom. The main point, however, is that the baseline dominates the scenarios, indicating that the hypothesized median shifts they represent add little to no value in predictive discrimination relative to the reference.

On modeling strategy, consider an example of \lq\lq normal" economics times as represented by 2018. In such settings, (i) the reference can be expected to be relatively light-tailed, (ii) scenarios can be expected to be modest in terms of varying from backstop, and (iii) the resulting scenario synthesis will be be close to the reference. There is limited scenario set incompleteness and the backstop will play a limited role. Contrast this with periods of higher uncertainty, such as in the 2007 context here where the reference distribution should have appropriately fatter tails. Then the synthesis of the baseline and the scenarios can substantially under-represent the reference unless the scenario set includes more extreme considerations. This mandates admitting extreme scenario considerations as a rational response to increased uncertainty in very uncertain economic times.

Scenario Synthesis based on P15, P50 and P85

As earlier noted, the methodology admits specification of multiple features of the scenarios so long as they can be represented as expectations under implicit scenario distributions. The case of multiple percentiles is practically key, and we visit this setting with summaries of further analysis in the TB 2007 context.

Figure (ref) and Table (ref) report the results when scenario p.d.f.s are tilted versions of the baseline that match the scenario-specific P15, P50, and P85. Here the scenario P15 and P85 are computed assuming distances of tail percentiles from median agree with the baseline: $\text{P15}_j=\text{P50}_j-(\text{P50}_0-\text{P15}_0)$ and $\text{P85}_j=\text{P50}_j+(\text{P85}_0-\text{P50}_0)$. The P15 and P85 for the scenarios in Table (ref) are similar to those in Table (ref) when the scenario is close to the baseline; they are naturally more discordant with increasing departure of the scenario percentiles from those of baseline (e.g., $\mathcal{S}_2$). Then, optimal weights and EMR result are quite similar to those in Table (ref). Other specifications of the P15 and P85 uncertainties-- specifications that are founded in scenario considerations-- may be quite different than the synthetic choices here, of course, and can be expected to lead to different results. A main point is that the methodology is open to-- and trivially applied to-- scenario specifications in terms of multiple percentiles, with negligible analytic and computational burden. Such specifications are increasingly common in application, and to be encouraged in policy research moving forward.

figure[figure omitted — 837 chars of source]
table[table omitted — 2,523 chars of source]

Distributional Forecasts and Judgment

At a general level, our focus is on use and reconciliation of information from statistical models and judgmental sources. Analysis is directional in that the specific goals are to assess judgmentally derived scenarios against a statistical reference distribution. In the broader context, the reverse is also of interest; that is, investigation of how a statistical forecast distribution may be \lq\lq tilted" in a direction deemed important from a judgmental point of view. The latter can often be proxy for information external to that underlying the statistical model.

A key context is that of unique, unexpected events and shifts in the structure of the economy that go well beyond existing model structure and assumptions. While structural economic models may provide a formal basis for longer-term adaptation, fully modeling the implications of regime shifts on the structure of the economy takes time. Short-term, judgmental adjustments can be most valuable for real-time decision making. Indeed, judgment plays a dominant role in decision making in other areas, such as among investment professionals in macroeconomic trading. In monetary policy settings, quantitative macroeconomic models provide a firmer basis, but undesirable policy recommendations from models are often attributed to persistent forecast errors. A \lq\lq good” policy maker would intervene to input judgment to address this.

Subjective information sources include surveys, market intelligence, and stress test outputs. Some comments on each are germane.

{\em i. Surveys.} Perhaps the key example is the US Survey of Professional Forecasters (SPF) a well-established source of community-wide forecast information. SPF now collates probabilistic forecasts on predefined bins for outcomes Croushore1993,DelNegro2023. Similar regular surveys are conducted by ECB, the Bank of England, and other institutions.

{\em ii. Market Intelligence.} Large and detailed information sets are commonly collected by central banks in order to inform policy makers on many dimensions of economic and financial market developments outside the scope of well-adopted structural macroeconomic models. Such models, aiming to reduce complexity and lead to openness and interpretation in economic terms, inevitably lack the ability to reflect the full complexity of structural changes or nonlinear dynamics that become practically relevant in more unusual circumstances. More complex economic and financial market intelligence-- in the form of summary external forecast information and judgment-- can then add real value, if recognized and appropriately integrated with the model-based forecasts.

{\em iii. Stress testing.} Initially developed by the IMF in the 1990s to assess financial system resilience, stress testing was later broadly adopted to assess macroeconomic risks from banking distress following the Great Financial Crisis Adrian2020. All major financial regulators now design and publish macroeconomic stress scenarios and assess financial stability relative to those scenarios. This typically includes scenarios of major downturns in macroeconomic aggregates such as real activity, inflation, and financial conditions. Scenario design emphasizes extreme economic and financial circumstances; the resulting outputs can provide a basis for judgmental modification of forecasts from the established reference econometric models that are not designed or customized to quickly and easily address such circumstances.

The overall question is that of intervention in the statistical model to incorporate such external information. The concept is long recognized and much methodology exists and is used in other areas of forecasting, such as commercial and financial applications (e.g., WestHarrision.JASA.1986; WestHarrison.JoF.1989; black1991asset; WestHarrison.YellowBook2ndEdn.1997; West2023constrainedforecasting). However, emphasizing and formalizing the question with respect to policy applications is highlighted and of renewed interest here.

Beyond contextual connections with the main theme of the current paper, key technical features of our scenario synthesis methodology relate directly to these complementary interests. From Section (ref) and generalizing the notation there, each of the constructed scenario p.d.f.s has the form $p_j(\mathbf{y}) \propto w_j(\mathbf{y}) p(\mathbf{y})$ based on the statistical reference p.d.f. $p(\mathbf{y})$ and scenario-specific weight-- or tilting-- function $w_j(\mathbf{y}).$ The latter is $w_j(\mathbf{y}) \propto w_0(\mathbf{y}) \exp\{ {\bm\tau}_j'\mathbf{s}_j(\mathbf{y})\} $ involving: (i) the ET term $\exp\{ {\bm\tau}_j'\mathbf{s}_j(\mathbf{y})\}$ used to define $p_j(\cdot)$ using the baseline and partial scenario information; and (ii) the baseline-reference IS weight function $w_0(\mathbf{y}) \propto p_0(\mathbf{y})/p(\mathbf{y}).$ The Monte Carlo methodology uses the discrete versions over the reference random samples $\mathbf{y}^i,$ i.e., $w_j^i = w_j(\mathbf{y}^i)$ for $j=\seq 0J.$

The normalized scenario p.d.f.s are then $p_j(\mathbf{y}) = c_j w_j(\mathbf{y}) p(\mathbf{y})$ where the $c_j$ are just normalizing constants. Thus, explicitly, the partial judgmental information that $\mathcal{S}_j$ encodes leads to a weighted modification of the statistical reference; on $\mathcal{S}_j$ alone, this can be regarded as the scenario-tilted version of the model-based forecast. This indicates that the overall question of conditioning a model-based forecast on what may be quite distinct forms of judgmental information is intimately addressed within our framework. It also follows that the BPS-justified scenario mixture p.d.f. is $f(\mathbf{y}|{\bm\alpha}) = w(\mathbf{y}|{\bm\alpha}) p(\mathbf{y})$ where $w(\mathbf{y}|{\bm\alpha}) = \sum_{j=\seq 0J} \alpha_j c_j w_j(\mathbf{y})$. This is true for any ${\bm\alpha},$ not just the EMR-optimized value central to our scenario synthesis goals. In other contexts, such as the above settings of modifying the reference $p(\mathbf{y})$ with judgment-based information summaries, this allows for context-specific specification of relative scenario weightings. Importantly, scenarios can address multiple aspects of the forecast distribution, including location shifts, scale and skewness perturbations-- both within any one scenario and with diversity across a scenario set. This may be particularly important to extension and evaluation of this approach in areas such as stress testing.

Summary Comments

The formal assessment and integration of partial scenario information with statistical forecast distributions is of interest in a range of policymaking settings. Our approach has been motivated by the monetary policy process, where policy decisions are firmly rooted in macroeconomic forecasts that involve not only the baseline forecast, but also alternative risk scenarios. The methodology offers a concrete and straightforward approach to evaluating baseline and judgmental scenario assessments-- with their intuitive and easily communicated bases-- against more formal statistical density forecasts of risk. Applications to the monetary policy process are highlighted, and offer a new frontier for practical yet rigorous policy making.

We believe the methodology will have broad appeal in other applications. For example, the IMF continuously monitors global financial stability, and publishes a formal biannual GaR-based global financial stability assessment in its Global Financial Stability Report. The statistical approach can be compared directly-- and in a quantitatively meaningful manner-- with the scenario-based risk assessment of the IMF's World Economic Outlook, published on the same schedule as the GFSR.

The approach can also be readily applied to other areas of institutional risk management and contexts such as portfolio choice applications. In risk management, as in financial institution supervisory stress testing, both scenario-based approaches and more statistical approaches are commonly deployed. Our framework and methodology offers a novel way to evaluate and integrate those two avenues in a concrete fashion. In portfolio allocation decision making, the role of priors is fundamental and features commonly in allocation decisions, yet the bridge between intuitive scenario based approaches and statistical modeling of return forecasts has received little attention. Again our framework offers steps ahead in this regard. Beyond these areas, there are also opportunities in commercial revenue and supply chain forecasting where ranges of forms of external/subjective information are often assessed in the context of formal models with a view to eventual decisions.

Forms of scenario information that may feed-into new applications are information sets with more than a few candidate percentiles of forecast distributions under any assumed scenario. If a scenario is specified in terms of a larger number of percentiles, then analysis begins to approximate that given a fully-specified scenario p.d.f. $p_j(\mathbf{y}).$ This is certainly of methodological interest, and may represent applied interests in settings where the scenarios are effectively replaced by forecast distributions from alternative/competing models. This latter setting is closer to the existing setting of BPS where multiple predictive distributions are considered by a Bayesian decision maker, and define analogue information to condition an initial reference forecast distribution. Some of the technical developments here are rather different-- and complementary-- to the general setting of BPS, but open up new questions for potential development and exploitation in forecast model comparison and synthesis. There are also questions of extension of the technical approach to address synthesis based on other forms of scenario information, e.g., point forecasts that are regarded as subjectively assessed modes or means rather than medians, and uncertainty in scenario information, e.g., percentiles provided with some notational $\pm$ uncertainties. The general ET framework in principle applies to such contexts, though details are to be developed for exploitation in any new applied context in which such scenario information sets arise.

We have noted that the methodology is applicable with $\mathbf{y}$ in several dimensions. Elements of $\mathbf{y}$ can include multiple economic indicators (real growth, inflation, unemployment, etc.,) as well as-- anchored at a current time period, such as the end of the current quarter-- multiple time periods ahead, such as the coming eight quarters. Scenario information to define (uncertain) constraints on forecasts of the state of the macroeconomy over multiple future time periods can than generate a range of scenarios. Technically and computationally, the methodology here extends immediately. We have experience with such extensions, and recognize questions that arise due to increasing dimension of $\mathbf{y}$. In technical essentials, the main questions there are not new, but have to do with scalability of importance sampling methodology with dimension, and of its close technical ally entropic tilting with increasing dimension of the underlying Bayesian decision-analytic utility (a.k.a. score) functions. These questions are addressed in all applications of these general approaches, and will need to be addressed in context-- in specific applied settings of scenario synthesis.

\if0\blind{

Acknowledgments

The authors thank colleagues at the Federal Reserve and the IMF-- including Gianni Amisano, Matthias Paustian, Sheheryar Malik, Jason Wu, and Pierre-Olivier Gourinchas-- for multiple discussions and feedback on the general area and a range of specific topics. We also thank Todd Clark of the Center for Financial Economics at Johns Hopkins University for his detailed, thoughtful comments on a first draft of our paper. The views expressed in this paper are those of the authors and do not necessarily reflect the views of the Federal Reserve System, the Federal Open Market Committee, or the International Monetary Fund, its Management, or its Executive Directors. }\fi

\FloatBarrier