EconBase
← Back to paper

Decision synthesis in monetary policy

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

66,492 characters · 31 sections · 34 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

RoyalBlue4 Decision Synthesis in Monetary Policy

\emergencystretch 3em

center[center omitted — 1,274 chars of source]

\thispagestyle{empty}

\setcounter{page}{1}

Introduction

Monetary policymakers are tasked with simple, but hard to achieve, objectives. A key example is the \lq\lq dual mandate" under which policymakers target future inflation rates and real activity using interest rates as the policy instrument. Decisions are made based on uncertain information from many sources. Here, the sources are models that generate predictive distributions for macroeconomic outcomes and policy instruments over multiple time periods. For a single model, an optimal policy path is evaluated via conditional forecasting and decision analysis. Applying standard methods, such as Bayesian model averaging (BMA), is one way to address model uncertainty-- routine decision analysis can then be applied to the weighted average of models. This traditional view, however, ignores the reality that models may each individually recommend very different optimal policy decisions. The question then arises of how to synthesize this information and, potentially, exploit it in the overall final decision process. This paper addresses this question.

The extensive Bayesian econometrics literature on model combination rarely highlights the fact that models are typically built for specific prediction and decision goals. Traditional BMA analysis weights models according to purely statistical model fit and only scores one-step ahead forecast outcomes. Extensions and alternatives have arisen to define model weightings based on aspects of past forecast performance with respect to specific forecast goals. MARTIN2023 survey Bayesian forecasting in economics and finance and review various forecast combination approaches, including some that are more explicitly concerned with goal-focused prediction mitchell-evaluating-2005, geweke-optimal-2011, conflitti-optimal-2015,Kapetanios2015,LoaizaMayaJOE2021,chernis-nowcasting-2022,AastveitEtAL2022, BernaciakGriffin2024. LavineLindonWest2021avs place many of the earlier approaches in a foundational Bayesian context. They justify model weights based on utilities in forecasting using historical model-specific \lq\lq scoring\rq\rq\ of past forecast outcomes. The underlying theoretical justifications come from Bayesian predictive synthesis (BPS) and the specific class of \lq\lq mixture BPS" models (McAlinnWest2018, section 2.2; JohnsonWest2022). However, while ultimate decision goals may be implicit in specific applications of model combination, they are rarely, if ever, taken into account in the analysis and resulting decision-making. This raises concerns. A model that has fit or forecast specific outcomes well in the past may be a good bet for use in resulting decision analysis, such as defining optimal decisions about values of policy instruments; however, there is no guarantee that this will be so.

Our view is that models that have recommended policy decisions that turned out to be \lq\lq good" should be more heavily weighted in looking ahead, just as past statistical predictive performance is generally positively weighted. The challenge is to operationalize the concept of \lq\lq good decision" performance. For example, a vector autoregression (VAR) model can be evaluated on forecast performance using a pseudo real-time forecasting exercise, but it is not clear how to evaluate such a model when it is used to advise policy decisions. We can, however, explore how the model would have advised on decisions in the past. Evaluations can then compare such analyses to decisions actually made by policymakers in the past (while recognizing that policymakers' past decisions were not necessarily \lq\lq good" and just outcomes of the rather amorphous reality of monetary policymaking).

Bayesian predictive decision synthesis (BPDS-- TallmanWest2023) has the potential to address these questions. An outgrowth of the theoretical BPS framework, BPDS explicitly addresses model scoring based on decision outcomes as well as predictive accuracy. In addition to reflecting historical outcomes of predictions and decisions, BPDS allows for differential model weighting based on expected decision outcomes. This is a decision parallel to the proven use of BPS models that incorporate outcome-dependent weights that modify BMA-like mixtures to differentially favour models in different parts of the future outcome space for pure forecasting. The latter concept was introduced by Kapetanios2015 whose empirically inspired developments recognized, for example, that one model may be better at predicting inflation when inflation is high and rising, while another model may be better when inflation is low and stable. BPS defines a theoretical Bayesian basis and broader methodological framework for this JohnsonWest2022. BPDS goes further by integrating both historical and expected decision outcomes; here we develop, extend, and exemplify BPDS in the central macroeconomic policy context.

BPDS applies the Bayesian mixture model approach of BPS using defined utility-- or \lq\lq score"-- functions that relate to explicit decision goals. This allows for multiple objectives (i.e., multi-attribute decision analysis). For example, a purely predictive vector score function can allow for multiple forecast horizons (e.g., to produce inflation near a target for each of the next eight quarters) and/or multiple outcome criteria (e.g., to separately reflect inflation targeting, interest rate smoothing, and stable growth patterns over coming quarters), among others. For policymakers juggling multiple objectives, this key feature is rather distinct from conventional approaches that adopt single, scalar criteria for model weighting. For example, a forecast combination approach might choose model weights based on the $h-$step ahead predictive likelihood for a single choice of $h$, with BMA simply focused on $h=1$. In contrast, BPDS can address multi-steps ahead in parallel, along with scoring of realized decisions that simultaneously target several macroeconomic outcomes.

The opportunities for exploring such practically relevant questions are highlighted in empirical studies. Our case study uses US macroeconomic data with multi-objective score functions to define BPDS model weights, demonstrating BPDS in macroeconomic forecasting and advisory decision-making. This is enabled by new methodological advances motivated by this applied context. First, in developing the scenario-based conditional forecasting analyses, we explicitly recognize the dual roles of policy decision variables (such as interest rates) as both outcomes and putative controls. This leads to an important theoretical extension of traditional conditional forecasting in which candidate decision scenarios are weighted by their predicted plausibility prior to assessing their likely implications in forecasting outcomes of other variables.\footnote{This is keeping in the spirit of LEEPER2003 who argue that the \lq\lq modest interventions"commonly used in reduced-form scenario analysis are unlikely to be subject to the Lucas Critique. We explicitly give higher weight to small interventions and down-weight large interventions.} This is important quite generally in conditional forecasting, as well as here in the resulting BPDS analysis. Second, we advance the BPDS methodology based on a new and practically critical focus on the relationships between multiple utilities defining the scoring of decision outcomes. Third, we discuss and promote the use of bounded utilities in model scoring, and customized forms of utility functions to guide model-specific and final BPDS-based optimal decisions.

BPDS Framework

We present and discuss the structure of BPDS at a particular point in time, ignoring the time dependency and relevance in the notation for clarity in communicating these essentials. Practical implementation in time series is of course sequential, with models at time $t$ depending on all relevant historical data and information.

Mixture BPDS and Decision Setting

At a given time point, let $\mathbf{y}$ denote the $q-$dimensional outcome variable of interest (e.g., inflation in each of the next $q$ quarters) and $\mathbf{x}$ the vector of control/decision variables (e.g., a target profile of central bank interest/base rates over the next $q$ quarters). Each of a set of $J$ models, ${\mathcal M}_j$, $j=\seq1J,$ predicts the outcome $\mathbf{y}$ via a predictive density $p_j(\mathbf{y}| \mathbf{x}, {\mathcal M}_j)$ conditional on any considered decision $\mathbf{x}.$ The policymaker responsible for ultimate decisions adopts a general BPDS approach with the overall conditional (on $\mathbf{x}$) predictive p.d.f.

equation[equation omitted — 183 chars of source]

with the following ingredients.

\em BPDS model probabilities

The {\em decision-dependent model probabilities} $\pi_j(\mathbf{x})$ can differentially weight models $j$ over the decision space $\mathbf{x}$. This incorporates any prior information relevant to model weighting based on past predictive model fit and decision outcomes, and now explicitly allows for adjustments based on a currently considered decision $\mathbf{x}$. Dependence of $\pi_j(\mathbf{x})$ on $\mathbf{x}$ is simply fundamental and critical in our policy setting.

\em BPDS calibration functions

The $\alpha_j(\mathbf{y}|\mathbf{x})$ are {\em calibration functions} that define outcome dependence of model weights over the outcome space of $\mathbf{y}$ for any chosen $\mathbf{x}$. This defines the opportunity to increase or decrease the model weights differentially over the outcome $\mathbf{y}$ space to address model-specific biases and preferences and address questions of model-specific calibration more generally. The BPDS mixture of \eqn{BPDSmixgeneral} has the equivalent form

equation[equation omitted — 156 chars of source]

where

equation[equation omitted — 258 chars of source]

with normalizing terms $k(\mathbf{x})$ and $a_j(\mathbf{x})$ explicitly dependent on $\mathbf{x}$. This form shows how the calibration functions $\alpha_j(\cdot|\cdot)$ modify the initial mixture pdfs $p_j(\cdot|\cdot) \to f_j(\cdot|\cdot)$ with corresponding changes of mixture weights $\pi_j(\mathbf{x})\to\tilde\pi_j(\mathbf{x}).$

Note that the choice of relevant calibration functions $\alpha_j(\mathbf{y}|\mathbf{x})$ will, in any given application, be partly dependent on characteristics of the model pdfs $p_j(\mathbf{y}|\mathbf{x},{\mathcal M}_j)$. In particular, the expectation of each $\alpha_j(\mathbf{y}|\mathbf{x})$ under $p_j(\mathbf{y}|\mathbf{x},{\mathcal M}_j)$ must be finite in order that \eqn{BPDSmixgeneral} defines a valid BPDS density $f(\mathbf{y}|\mathbf{x}).$ This supports the use of bounded scores, in general.

\em Baseline mixture component

The model index $j=0$ explicitly allows for a {\em baseline} model component ${\mathcal M}_0$ in the mixture p.d.f. $f(\cdot|\cdot)$ that can, among other things, address the ever-present issue of \lq\lq model set incompleteness" TallmanWest2023. ${\mathcal M}_0$ can be chosen to produce a pdf $f_0(\cdot|\cdot)$ that is over-dispersed relative to the mixture of the initial $J$ models, so supporting outcomes $\mathbf{y}$ that are unusual under the $J$ models. The baseline is then a suitable \lq\lq fall back" model for times when the other models are forecasting poorly.

\em Initial mixture

The special case with each $\alpha_j(\mathbf{y}|\mathbf{x})=1$ defines the {\em initial mixture} with no BPDS calibration. We use $p(\mathbf{y}|\mathbf{x})$ in notion; that is, $p(\mathbf{y}|\mathbf{x}) = \sum_{j=\seq 1J} \pi_j(\mathbf{x}) p_j(\mathbf{y}|\mathbf{x},{\mathcal M}_j)$.

Special cases fix ideas. First, if $\pi_j(\mathbf{x})=\pi_j$ with $\pi_0=0$ are model probabilities based on historical BMA analysis, and with $\alpha_j(\mathbf{y}|\mathbf{x})=1,$ then \eqn{BPDSmixgeneral} specializes to BMA. Thus, BMA analyses-- with or without this decision dependence in model-specific forecasts-- are very special cases of BPDS. Second, again with $\pi_j(\mathbf{x})=\pi_j$, $\pi_0=0$ and $\alpha_j(\mathbf{y}|\mathbf{x})=1,$ the decision-maker has the freedom to specify the initial mixture probabilities $\pi_j$ in other ways than with BMA. This includes using historical performance defined by scoring of past forecast outcomes, justifying various approaches to goal-focused model weighting LavineLindonWest2021avs,LoaizaMayaJOE2021 as special cases of BPDS. Third, mixture BPS McAlinnWest2018,JohnsonWest2022 is a special case in which models are combined with outcome-dependent weights. In these settings, $\pi_j(\mathbf{x})=\pi_j$ depends on past predictive performance, $p_j(\mathbf{y}|\mathbf{x}) = p_j(\mathbf{y})$ and $\alpha_j(\mathbf{y}|\mathbf{x})=\alpha_j(\mathbf{y})$ define outcome-dependent modifications of model probabilities, but there is no decision context so no $\mathbf{x}-$dependence. BPDS critically recognizes that the foundational BPS theory allows explicit incorporation of decision goals-- admitting the conditioning on $\mathbf{x}$ throughout all components of \eqn{BPDSmixgeneral}-- to extend the foregoing analyses.

With predictions of $(\mathbf{y}|\mathbf{x})$ based on \eqn{BPDSmixgeneralfform}, the Bayesian decision-maker acts to identify the optimal decision $\mathbf{x}$ based on a chosen utility function $U(\mathbf{y},\mathbf{x}).$ This involves numerical optimization to maximize the implied expected utility $\bar U(\mathbf{x}) = E_f[U(\mathbf{y},\mathbf{x})|\mathbf{x}]$ over the decision space of $\mathbf{x}.$ The notation $E_f[\cdot|\cdot]$ here explicitly represents expectation with respect to the BPDS distribution, and we use $E_p[\cdot|\cdot]$ to denote expectation under the initial mixture.

Decision-dependent Scores for Calibration Functions

The key step in integrating decision outcomes into relative model weights addresses the question of how each ${\mathcal M}_j$ would inform decisions if used alone. Given the predictive p.d.f. $p_j(\mathbf{y}|\mathbf{x},{\mathcal M}_j)$ and a chosen, potentially model-specific utility function $u_j(\mathbf{y},\mathbf{x}),$ acting based only on ${\mathcal M}_j$ leads to the optimal decision $\mathbf{x}_j$ that maximizes $E_{p_j}[ u_j(\mathbf{y},\mathbf{x})|\mathbf{x}]$ over $\mathbf{x}.$ The decision-maker has access to this set of model recommendations and is interested in model combination to preferentially weight \lq\lq good decision models" as well as models that generate good predictions. BPDS formalizes this with specified {\em score functions} $\mathbf{s}_j(\mathbf{y},\mathbf{x}_j),$ each being a $k-$vector of utilities that can be chosen to reflect both predictive and decision goals. The use of multi-dimensional scores addresses multiple goals simultaneously.

The Bayesian decision-theoretic development of TallmanWest2023 generates the resulting functional forms of the BPDS calibration functions as

equation[equation omitted — 156 chars of source]

where ${\bm\tau}(\mathbf{x})$ is a $k-$vector with elements differentially weighting the multiple utility dimensions of the score vector. The reasoning and theory behind this key result is as follows.

The initial mixture $p(\mathbf{y}|\mathbf{x})= \sum_{j=\seq 1J} \pi_j(\mathbf{x}) p_j(\mathbf{y}|\mathbf{x},{\mathcal M}_j)$ is the $\mathbf{y}-$margin of the joint distribution $p(\mathbf{y},{\mathcal M}_j|\mathbf{x}) = \pi_j(\mathbf{x}) p_j(\mathbf{y}|\mathbf{x},{\mathcal M}_j),$ $(j=\seq 1J)$. Under this initial distribution for any candidate decision $\mathbf{x}$, and with score vectors $\mathbf{s}_j(\mathbf{y},\mathbf{x}_j)$ defined and evaluated at model-specific optimal decisions $\mathbf{x}_j,$ the decision-maker has {\em initial expected score} $\mathbf{m}_p(\mathbf{x}) = \sum_{j=\seq0J} \pi_j(\mathbf{x})\mathbf{m}_{jp}(\mathbf{x})$ where $\mathbf{m}_{jp}(\mathbf{x}) = \int_{\mathbf{y}} \mathbf{s}_j(\mathbf{y},\mathbf{x}_j)p_j(\mathbf{y}|\mathbf{x},{\mathcal M}_j)d\mathbf{y}.$ Treating $\mathbf{m}_p(\mathbf{x})$ as a benchmark to improve on expectation, the BPDS theory enquires about distributions $f(\mathbf{y},{\mathcal M}_j|\mathbf{x})$ that yield expected scores $\mathbf{m}_f(\mathbf{x}) \ge \mathbf{m}_p(\mathbf{x}) + {\bm\epsilon}(\mathbf{x})$ for some non-negative $k-$vector (with at least one positive entry) ${\bm\epsilon}(\mathbf{x});$ this may be chosen to depend on $\mathbf{x},$ or may be a specified constant \lq\lq decision score improvement." Given $\mathbf{m}_f(\mathbf{x}),$ the BPDS theory identifies a unique $f(\cdot,\cdot|\mathbf{x})$ that minimizes the K\"ullback-Leibler (KL) divergence of $p(\cdot,\cdot|\mathbf{x})$ from $f(\cdot,\cdot|\mathbf{x})$ and has an expected score of exactly $\mathbf{m}_f(\mathbf{x}).$ The theory is that of {\em relaxed entropy tilting} TallmanWestET2022, TallmanWest2023,West2023constrainedforecasting and yields $f(\mathbf{y},{\mathcal M}_j|\mathbf{x}) \propto \pi_j(\mathbf{x}) \alpha_j(\mathbf{y},\mathbf{x}) p_j(\mathbf{y}|\mathbf{x},{\mathcal M}_j)$ with calibration function precisely as in \eqn{alphajBPDS}. The {\em tilting vector} ${\bm\tau}(x)$ is implicitly defined by the vector of $k$ {\em target score constraints} $E_f[ \mathbf{s}_j(\mathbf{y},\mathbf{x}_j) |\mathbf{x} ] = \mathbf{m}_f(\mathbf{x}).$

BPDS uses the initial mixture based on past performance, but then admits that small perturbations of model probabilities based on their expected performance may lead to improved scores. A stylized example has $J=2$ models that have forecast equally well to date. Traditional model averaging methods including BMA confer equal weights in the combination. If, however, the models have different expected scores $\mathbf{m}_{jp}(\mathbf{x})$, conferring slightly more weight on the model expected to lead to a higher score makes sense.

The entropic (or exponential) tilting theory is general. It is, of course, possible to tilt the initial joint distribution to most targets, so long as they are technically achievable under the initial distribution. However, an overly ambitious target score may result in empirically unreasonable results. Hence, we emphasize the importance of selecting $\mathbf{m}_f(\mathbf{x})$ that represents a \lq\lq small" improvement over the initial benchmark score $\mathbf{m}_p(\mathbf{x}).$ This is bolstered by the assumption that the initial model probabilities reflect the empirical plausibility of models, as well as any available information about historical predictive and decision performance. Further, as we demonstrate later in the case study, aspects of the computational methodology for model fitting in the sequential time series setting naturally inform on, and allow monitoring of, relevant choices of target expected scores.

BPDS Summary

This section has outlined the main ideas underlying BPDS and the key ingredients of the theory and resulting technical machinery. Specifications of score and utility functions, initial model probabilities, and target scores are all required for implementation and are, of course, application specific. The following section develops full details in the context of the macroeconomic decision-making application. In terms of computation, BPDS requires the use of posterior simulation methods (i.e., draws from conditional predictive densities from each model), as well as numerical optimization methods (i.e., to find the $\mathbf{x}_j$ and the overall optimal decision $\mathbf{x}$ from BPDS analysis).

BPDS for Optimal Monetary Policy Decisions

Following frs2019, we use quarterly macroeconomic and financial data from 1973:Q1 to 2022:Q2 from the FRED-QD database (Federal Reserve Bank of St. Louis). This includes GDP (log of real GDP), prices (log of GDP deflator), interest rate (the shadow rate,\footnote{This is the Federal Funds rate when the latter is positive, but can go negative when it is at the zero lower bound, taking into account unconventional monetary policy; see Wu-shadowrate.} which we treat as the policy rate), investment (ratio of real gross private domestic investment to GDP), real stock prices (log of the S&P500 deflated by GNP/GDP price index), and spread (between BAA bonds and the Fed funds rate). Models are run over multiple years. Each quarter they produce forecasts-- full predictive distributions in terms of Monte Carlo samples-- of outcomes of interest over the following $k=8$ quarters. This is conditional on candidate settings of the decision vector, which is taken as the trajectory of interest (shadow) rates over those quarters. Within-model decision analysis then delivers model-specific optimal decisions about these rates.

Models, Forecasts, and Model-specific Decisions

We consider $J=2$ models: ${\mathcal M}_1$ is a three-variable monetary policy VAR involving GDP, prices, and the interest rate; ${\mathcal M}_2$ is similar to the model of frs2019, a VAR with the same variables as ${\mathcal M}_1$ plus investment/GDP ratio, stock prices, and GZ credit spread. Following frs2019 we include five lags in the VARs, and each model is identified using the sign restrictions from Table 1 of that reference. In ${\mathcal M}_1$, these restrictions define supply, demand, and monetary policy shocks. In ${\mathcal M}_2$, investment and financial shocks are additionally identified. We condition on a given value of the policy rate and set monetary policy to be the driving shock. We do this by imposing restrictions on the set of structural shocks underlying the conditional forecasts. Structural shocks other than the monetary policy shock have zero means. We use the asymmetric conjugate prior of C2022QE, with the advantage that the marginal likelihoods for each can be easily calculated; prior hyperparameter choices are made to maximize the marginal likelihood as in this referenced paper. At each quarter, multi-step ahead predictions are based on simulations using the precision-based sampler of chan-cond-forecast. Details on the conditional forecast computations are summarized in Appendix B. In short, this generates $p_j(\mathbf{y}|\mathbf{x})$ with zero-mean constraints on all shocks apart from the monetary policy variable. We do not restrict the variance (i.e., \lq\lq soft" restrictions) such that we also have uncertainty around the path of $\mathbf{x}$. This can be thought of as conditional commitment-- we allow the possibility of $\mathbf{x}$ deviating from the proposed policy path with deviations informed by historical uncertainty around forecasts of the interest rate outcomes.

It is worth noting that using models such as VARs for policymaking is sometimes questioned due to the Lucas critique. However, papers such as LEEPER2003 argue in favor of VARs as long as the policy interventions being studied are modest. We ensure BPDS will favor such modest interventions through two mechanisms. First, as noted in Section (ref), we choose only modest improvements in the target score. Second, large policy interventions will be down-weighted through the initial probabilities reflecting the plausibility of an intervention; see details in Section (ref). Thus BDPS allows for policy analysis using VARs by emphasizing decision scenarios which represent \lq\lq modest interventions", circumventing the Lucas Critique. This is an important practical contribution to the VAR literature on policy and scenario analysis.

To find optimal decisions requires choice of utility function. The form adopted here is as follows. In the current quarter, $\mathbf{x} = (x_1,\ldots,x_k)'$ is the $k-$vector of interest rate values over the next $k$ quarters, and $\mathbf{y} = (y_1,\ldots,y_k)'$ is the corresponding $k-$vector of inflation rates. Whatever other variables are in ${\mathcal M}_j,$ interest focuses on the implied $p_j(\mathbf{y}|\mathbf{x},{\mathcal M}_j)$ required for BPDS \eqn{BPDSmixgeneral}. This conditional predictive is used in decision analysis with the same utility function for each model, namely $u_j(\mathbf{y},\mathbf{g},\mathbf{x}) = U(\mathbf{y},\mathbf{g},\mathbf{x})$ given by \beq{DSutility} U(\mathbf{y},\mathbf{g},\mathbf{x}) = - \sum_{h = \seq1k}\{\theta(y_h - y^{*})^2 + (1-\theta)(g_h - g^{*})^2 + (x_h - x_{h-1})^2 \} \end{equation} where $\mathbf{g}$ denotes GDP growth.

This is a conventional utility function that reflects the dual mandate of inflation rate targeting, moderate growth in real-activity, and interest rate smoothing over the next $k=8$ periods. The $\mathbf{y}$ terms relate an inflation targeting mandate of $y^{*}=2\%$ over the longer run. The $\mathbf{g}$ terms reflect GDP growth with a target of $g^{*}=2.5\%$, roughly the average growth rate of GDP from 1990 to 2024. The $\mathbf{x}$ terms encourage relatively constrained changes in quarter-to-quarter interest rates (a \lq\lq don't rock the boat" consideration, as large swings in interest rates can/will have otherwise undesirable effects on the macroeconomy). The model-specific optimal decision vector $\mathbf{x}_j$ then maximizes $E_{p_j}[ U(\mathbf{y},\mathbf{g},\mathbf{x})|\mathbf{x}]$ over $\mathbf{x}.$ Again, this analysis is repeated each quarter over time, producing rolling updates of the \lq\lq currently optimal" projections for interest rates over the coming eight quarters.

BPDS Model Specification

BPDS requires specification of a relevant class of baseline pdfs $p_0(\mathbf{y}|\mathbf{x})$, the model-specific vector score functions $\mathbf{s}_j(\mathbf{y},\mathbf{g},\mathbf{x}),$ the initial BPDS model probabilities $\pi_j(\mathbf{x})$ as functions of candidate decisions $\mathbf{x},$ and the target expected scores $\mathbf{m}_f(\mathbf{x})$ at any $\mathbf{x}$. These are discussed in turn. Beyond customizing BPDS to the specific application, we highlight methodological developments relevant to other applications. This includes, in particular: (a) the linkages of the $\pi_j(\mathbf{x})$ to $\mathbf{x}$ that are relevant more generally when $\mathbf{x}$ is both an outcome to be forecast as well as a putative decision variable; see Section (ref); (b) the relevance of dependence structure among the elements of the vector score under the initial distribution $p(\mathbf{y},{\mathcal M}_j)$; see Section (ref)). Notation considers the target variable $\mathbf{y}$ as including both inflation and GDP unless specifically denoted (as in the score/utility functions).

Baseline Distribution

Completing the main BPDS p.d.f. in \eqn{BPDSmixgeneral} requires the baseline $p_0(\mathbf{y}|\mathbf{x}).$ This is taken as a multivariate T distribution with 10 degrees of freedom, using the location from the initial mixture $p(\mathbf{y}|\mathbf{x})$, ignoring the baseline (i.e., with $\pi_0(\mathbf{x})=0)$ and the corresponding variance of that mixture inflated by 2. This defines a relevant, tractable ${\mathcal M}_0$ that can capture outcomes $\mathbf{y}$ that the set of VAR models are not predicting well for any $\mathbf{x}$ under consideration, and signal that to the decision-maker.

BPDS Score Functions

The dual mandate and interest rate smoothing from the model-specific decision analysis in Section (ref) are reflected in the BPDS score functions. We take

align[align omitted — 131 chars of source]

with elements

align[align omitted — 282 chars of source]

where $y^{*}= 2\%$ is the inflation target, $g^{*}= 2.5\%$ is the GDP growth target, and $z_{y}$, $z_{g}$ and $z_{x}$ are {\em score bandwidth} parameters. This defines a class of bounded score functions, always relevant in decision analysis and here ensuring that the entropically tilted BPDS p.d.f. of \eqn{fj} is always integrable. The score bandwidths are set so that a certain deviation $d_y = (y - y^*)^2$ has a score of ${\bm\varepsilon}$; given a choice of ${\bm\varepsilon}$, we set $z_y = d_y/\sqrt{-2\log({\bm\varepsilon})}$. Similar considerations apply to choosing $z_{x}$ and $z_{g}$. Our analyses use ${\bm\varepsilon} = 0.4$, $d_y = 2$, $d_g = 2$, and $d_x = 1$ to ensure the score function is dispersed enough to accommodate modest changes in the Federal Funds rate while being more lenient in deviations from the inflation target. Obvious modifications could incorporate horizon $h-$specific inflation targets and differentially weight the two exponential terms, but this form suffices for our main goals in this paper. Note also that, if inflation deviations from target and interest rate changes are \lq\lq small," then $s_{jh}(y_h, x_h) $ is approximately quadratic in $|y_h - y^{*}|$ and $|x_h - x_{h-1}|$ for all $h,$ perhaps a more familiar utility form.

Initial and Conditional Model Probabilities

For clarity in this section, we make explicit the dependency on time, so that the ingredients of the full BPDS predictive p.d.f. in \eqn{BPDSmixgeneral}-- with the exponential form of the calibration function of \eqn{alphajBPDS}-- are now indexed by current time $t$; that is, $$ f_t(\mathbf{y}_t|\mathbf{x}_t) \propto \sum_{j=\seq 0J} \pi_{tj}(\mathbf{x}_t) {\textrm e}^{{\bm\tau}_t(\mathbf{x}_t)'\mathbf{s}_{tj}(\mathbf{y}_t,\mathbf{x}_{tj})} p_{tj}(\mathbf{y}_t|\mathbf{x}_t,{\mathcal M}_j). $$ Bayesian model weighting based on historical predictive performance with respect to defined forecast goals LavineLindonWest2021avs is the starting point for specification of the $\pi_{tj}(\mathbf{x}_t).$ The general form adopted is \beq{BPDSprobs} \pi_{tj}(\mathbf{x}_t) \propto \pi_{tj} p_{tj}(\mathbf{x}_t|{\mathcal M}_j), \qquad j=\seq 0J, \end{equation} subject to summing to 1 over $j=\seq 0J$ and with ingredients as follows.

\em Initial model probabilities

Traditional analysis WestHarrison1997 defines the time $t$ initial model probabilities as Bayesian updates from those at $t-1$; that is, $\pi_{tj} \propto \pi_{t-1,j} p_{tj}(\mathbf{z}_{t-1,j}|{\mathcal M}_j)$ where the \lq\lq model marginal likelihood" term $p_{tj}(\mathbf{z}_{t-1,j}|{\mathcal M}_j)$ is the value of the one-step ahead predictive p.d.f. under ${\mathcal M}_j$ at the observed values of the last period outcomes $\mathbf{z}_{t-1,j}$ under that model. In our setting, this includes time $t-1$ outcomes of inflation $(y)$, interest rate $(x)$, and other economic indicators in ${\mathcal M}_j$. In general these can differ across models, but in consideration for the initial weights we restrict to variables common across models.

BPDS allows the decision-maker freedom to make alternative choices of the $\pi_{tj},$ and the goal and decision focus recommend modification of the standard BMA choice. BMA, after all, only reweights models based on one-step ahead predictive accuracy. Hence, we adopt two modifications based on recent literature consonant with the goal foci.

First, we use simple power discounting of historically accrued support across models, in which the time $t-1$ to time $t$ evolution is reflected in $\pi_{tj}\propto \pi_{t-1,j}^\gamma p_{tj}(\mathbf{z}_{t-1,j}|{\mathcal M}_j)$, where $\gamma$ is a discount factor in $(0,1]$, closer to 1 for most applications. This acts to discount historically accrued support for model $j$ at a per-time unit discount rate $\gamma$ prior to updating by the time $t-1$ information. Going back at least to JQSmith1979 and then, in a formal dynamic model uncertainty context, WestHarrison1989, power-discounting has been shown to be of value in empirical studies in implicitly allowing for time variation in the predictive relevance of different models Raftery10,Koop2013,ZhaoXieWest2016ASMBI. Our case study uses $\gamma = 0.95$.

Second, initial model probabilities are modified based on the recent relative performance of models with respect to the defined goals. This is the premise underlying the specific variants of BPS in the setting of adaptive variable selection, or BPS-AVS, in LavineLindonWest2021avs, and related developments in LoaizaMayaJOE2021, for example. This leads to the immediate BPDS extension of these prior approaches with $$\pi_{tj}\propto \pi_{t-1,j}^\gamma p_{tj}(\mathbf{z}_{t-1,j}|{\mathcal M}_j) \textrm{e}^{{\bm\tau}_{t-1}(\mathbf{x}_{t-1})' \mathbf{s}_{t-1,j}(\mathbf{y}_{t-1}, \mathbf{x}_{t-1,j})}.$$ Here the discounted past probabilities are updated with AVS-style weights using the realized BPDS calibration function with relative model scores based on the actual decision outcomes at the last time period. As a result, models are initially and naturally reweighted based on both predictive and decision outcome performance at the last time period.

The result at this stage is that models achieving \lq\lq good" recent trajectories of interest rate smoothness, as well as relatively accurate forecasting performance of realized inflation outcomes, will be rewarded with higher initial BPDS model probabilities looking forward. Of course, the specification can cut back to define special cases including BPS-AVS (with ${\bm\tau}_{t-1}(\mathbf{x}_{t-1})=\mathbf{0})$, and within that to traditional BMA (with $\gamma=1$) for comparisons.

Finally, the inclusion in BPDS of the baseline model and its forecast densities leads to a modification of these initial model probabilities to provide a non-zero value $\pi_{t0}$ for the baseline. We choose a fixed probability-- in our analysis $\pi_{t0} =0.1$ at each time point $t$-- and simply renormalize the $\pi_{tj}$ over $j=\seq 1J$ accordingly.

\em Informative conditioning on $\mathbf{x}_t$

As noted earlier, in our setting the future values of decision variables are also considered outcomes predicted under the models. The models each forecast the future evolution of interest rates as part of the complex, dynamic macroeconomic system, whereas for decisions we must condition on $\mathbf{x}_t$. This is reflected in the conditional (on $\mathbf{x}_t$) distributions $p_{tj}(\mathbf{y}_t|\mathbf{x}_t,{\mathcal M}_j)$ in BPDS where $\mathbf{x}_t$ is treated as known. The theoretical implication for the BPDS model probabilities is the term $p_{tj}(\mathbf{x}_t|{\mathcal M}_j) $ in \eqn{BPDSprobs}-- this is the value of the current marginal predictive p.d.f. of the vector $\mathbf{x}_t$ under ${\mathcal M}_j.$ Assuming the prior (to time $t$) probabilities $\pi_{tj}$ are specified, this form arises directly via Bayes' theorem. The act of conditioning on $\mathbf{x}_t$ is informative, and the implied update is, simply by Bayes' theorem, that in \eqn{BPDSprobs}. Critically, this implies that candidate decision values that are not well-supported under the joint distribution of a model are down-weighted. Conversely, at any candidate decision vector $\mathbf{x}_t,$ models that are more predictively supportive of the decision $\mathbf{x}_t$ will be relatively rewarded with higher values of resulting $\pi_{tj}(\mathbf{x}_t)$.

This is keeping with LEEPER2003 who argue that reduced form models are appropriate to use for policy analysis if the policy does not deviate too much from past policy changes. If the decision is very unusual agents may think they are in a new policy regime and adjust their behavior rendering a reduced form model less useful (i.e. the Lucas Critique). The technique above addresses these concerns by explicitly increasing the weight on models where the policy decision represents only a \lq\lq modest intervention'. Further, if the decision is unsupported by the model set then the baseline model, being dispersed with fat tails, will receive more weight. In this way BPDS up-weights models least susceptible to the Lucas Critique, and in periods when there is more uncertainty the weights increase on the baseline model.

In other applications of BPDS, the decision variables may be exogenous; that is, control variables that are to be chosen by the decision-maker but that are not forecast jointly with $\mathbf{y}_t$ in the set of models. In such cases, it will be common to assume that the external choice of $\mathbf{x}_t$ is not informative, and then \eqn{BPDSprobs} results in decision-independent BPDS probabilities $\pi_{tj}(\mathbf{x}_t) =\pi_{tj} $ based only on historical data and information.

BPDS Target Scores

The BPDS expected score $\mathbf{m}_f(\mathbf{x}) = E_f[\mathbf{s}(\mathbf{y}, \mathbf{x})]$ represents a target for improvement over the initial expected score $\mathbf{m}_p(\mathbf{x})=E_p[\mathbf{s}(\mathbf{y}, \mathbf{x})]$. In the multi-objective case, the resulting ${\bm\tau}(\mathbf{x})$ that defines $f(\mathbf{y}|\mathbf{x})$ to satisfy this target expectation is sensitive to both the relative scales and dependence of elements of $\mathbf{s}(\mathbf{y}, \mathbf{x})$ under the initial mixture $\mathbf{y} \sim p(\mathbf{y}|\mathbf{x})$ at any candidate decision $\mathbf{x}.$ As functions of $\mathbf{y}$, the elements of the random score vector $\mathbf{s}(\mathbf{y}, \mathbf{x})$ can be strongly correlated, leading to challenges in specifying relevant targets. This can also complicate the calculation of the implied BPDS tilting vector ${\bm\tau}(\mathbf{x})$ (i.e., the vector that is needed to satisfy $\mathbf{m}_f(\mathbf{x}) = E_f[\mathbf{s}(\mathbf{y}, \mathbf{x})]$ under the BPDS density of \eqn{BPDSmixgeneral}). We address this by explicitly recognizing score dependencies and defining an approach that explicitly incorporates dependence.

Some theoretical intuition is gained by considered cases of \lq\lq small perturbations" in which $\mathbf{m}_f(\mathbf{x}) - \mathbf{m}_p(\mathbf{x})$ has small elements. In this setting, entropic tilting theory in TallmanWestET2022 yields the second-order approximation ${\bm\tau}(\mathbf{x}) \approx \mathbf{V}_p(\mathbf{x})^{-1} (\mathbf{m}_f(\mathbf{x})-\mathbf{m}_p(\mathbf{x}))$ where $\mathbf{V}_p(\mathbf{x})$ is the variance matrix of $\mathbf{s}(\mathbf{y},\mathbf{x})$ under the initial mixture $p(\mathbf{y}|\mathbf{x}).$ This shows that the implied tilting vector will be very sensitive to the initial score scales and dependencies as reflected in $\mathbf{V}_p(\mathbf{x}),$ and suggests a prime focus on a {\em standardized score} scale; that is, define $\mathbf{C}_p(\mathbf{x})$ as the scaled eigenvector matrix such that $\mathbf{V}_p(\mathbf{x})=\mathbf{C}_p(\mathbf{x})\mathbf{C}_p(\mathbf{x})'$ and set the target score using $\mathbf{m}_f(\mathbf{x}) = \mathbf{m}_p(\mathbf{x}) + \mathbf{C}_p(\mathbf{x}){\bm\epsilon}(\mathbf{x})$ for a specified {\em standardized expected score vector} ${\bm\epsilon}(\mathbf{x})$. The usual convention is taken in which the eigenvector columns of $\mathbf{C}_p(\mathbf{x})$ are ordered according to decreasingly values of the corresponding eigenvalues, so that the first column is \lq\lq dominant," and so forth. This provides insights into how to practically define target scores related to the absolute standardized scale. As examples of the two extremes, taking ${\bm\epsilon}(\mathbf{x}) = \epsilon(\mathbf{x})\mathbf{1}$ for some scalar $\epsilon(\mathbf{x})$ represents targets deviating from the initial expected score in equal amounts of $\epsilon(\mathbf{x})$ along each of the standardized eigen dimensions. At the other extreme, and most relevant when there are strong score dependencies, taking ${\bm\epsilon}(\mathbf{x})= (\epsilon(\mathbf{x}),0,\ldots,0)'$ defines the resulting target $\mathbf{m}_f(\mathbf{x})$ based on the major, dominant eigen dimension alone. The latter is a starting point in general and is taken to define our BPDS case study that follows. In that setting, we choose $\epsilon(\mathbf{x})$ such that $\min \{(\mathbf{m}_f(\mathbf{x}) / \mathbf{m}_p(\mathbf{x}))\} = 0.75$ to define the maximum expected improved score in any dimension.\footnote{Additionally, due to the arbitrariness of the signs of eigenvectors, we apply a $\pm 1$ multiplier to the first column of $\mathbf{C}_p(\mathbf{x})$ so that the sum of the elements are positive, ensuring the target score improves upon $\mathbf{m}_p(\mathbf{x})$.} It is obviously straightforward to extend this methodology to define target scores impacted by higher eigen dimensions, though that is left for future applications.

BPDS Implementation and Optimal Decisions

The final step couples the decision-maker's utility with the BPDS predictive eqns. ((ref),(ref)) to define the optimal decisions from the model synthesis. The decision-maker can adopt any utility function, but an initial neutral analysis will be based on using the same form as usual in the model-specific decisions: the function $U(\mathbf{y},\mathbf{x})$ of \eqn{DSutility}. This is used here, computing $\mathbf{x}$ to maximize the implied expected utility function $\bar U(\mathbf{x}).$ We compare decisions recommended by BPDS to those from each of the models and to the traditional BMA-based analysis. BMA uses model weights proportional to the marginal likelihoods of the data that are common across models (including inflation, interest rate, and GDP). As earlier discussed in Section (ref), BMA arises as a special case of the BPDS analysis.

The computation of BPDS involves two key components. First, the overall optimization over $\mathbf{x}$ explores potential BPDS decisions and finds the optimizing vector $\mathbf{x}$. This requires an \lq\lq outer loop" numerical optimization to explore $\mathbf{x}$ space. This is done using the MATLAB implementation of particle swarm optimization. We use this due to occasional multi-modality in $\bar U(\mathbf{x})$ and because the algorithm supports parallelization which significantly speeds up optimization. Model specific optimizations are done using a trust region method, namely, Powell's Derivative Free Optimization Solvers (PDFO-- pdfo). Second, within each evaluation of a potential BPDS decision, it is necessary to compute the tilting vector ${\bm\tau}(\mathbf{x})$ given the constraint $E_f[\mathbf{s}(\mathbf{y},\mathbf{x})]=\mathbf{m}_f(\mathbf{x})$ for a target expected score $\mathbf{m}_f(\mathbf{x}).$ The theoretical basis of this is an implicit equation that is solved via standard, generic numerical optimization methods. Relevant details follow TallmanWest2023 and are summarized in our Appendix A.

BPDS forecast distributions are evaluated using importance sampling. At any given $\mathbf{x},$ the BPDS predictive distribution in \eqn{BPDSmixgeneral} is simulated by sampling from the $p_j(\mathbf{y}|\mathbf{x},{\mathcal M}_j)$ in proportions defined by the BPDS probabilities $\pi_j(\mathbf{x})$. Importance sampling weights are then proportional to the realized values of $\alpha_j(\mathbf{y}|\mathbf{x}).$ This provides for efficient computation as well as access to traditional methods and metrics-- such as the importance of sampling effective sample sizes (ESS-- GruberWest2016,GruberWest2017ECOSTA)-- to monitor and evaluate the quality of the resulting Monte Carlo approximations compared to resulting predictive expectations. Note that this evaluation can deliver such metrics to assess \lq\lq concordance" between the initial densities $p_j(\mathbf{y}|\mathbf{x},{\mathcal M}_j)$ and their corresponding BPDS-tilted versions $f_j(\mathbf{y}|\mathbf{x},{\mathcal M}_j)$ in \eqn{BPDSmixgeneralfform}, as well as that of the initial mixture $p(\mathbf{y}|\mathbf{x})$ and the resulting $f(\mathbf{y}|\mathbf{x}).$ More aggressive BPDS target scores will generally lead to lower concordance, and choices can be partly guided by such empirical evaluations.

Case Study

Overview

Analyses proceed sequentially on an expanding window of data beginning in 1992Q2. We compare decisions recommended by BPDS to those under BMA, discuss the individual models and how they are combined by BPDS and BMA, and highlight some operational BPDS details underlying insights into the resulting decision outcomes.

Optimal Decisions

Figure (ref) shows the actual policy rate each quarter along with the $1{-}8$-quarters-ahead policy recommendations that would have been made by BPDS and BMA. In using the shadow rate, the zero lower bound is not in effect and negative values for the policy rate are possible. Recommendations for negative values for the policy rate are not to be taken literally as advising cuts to a negative Federal Funds rate, but rather as a suggestion to undertake other forms of monetary easing that would be expected to proxy such cuts.

Since 2014, the optimal policy paths recommended by the two approaches are generally similar, though there are notable differences prior to that time. Some specific periods of interest are now highlighted.

\em 2014 to the present

During this period, BMA and BPDS provide similar recommendations that are often quite different from the actual policy rate. For almost all of these times, the policy recommendations are to cut interest rates, whereas (apart from 2019--2021) the actual policy rate increased. Some differences do arise between BMA and BPDS. For example, during the post-COVID inflation period, BPDS recommends a higher rate path. BPDS is closer to the decision actually made by the Federal Reserve, although according to BPDS, interest rates should decrease throughout 2023 and 2024.

\em The financial crisis and subsequent recession

It is during this period that the differences between BPDS and BMA are most acute. The actual policy rate fell slowly during this period. BPDS recommends rate cuts as well, initially at a more rapid rate than what actually occurred, but as of 2010, its recommendations are similar to the ones the policymakers actually made. In contrast, BMA recommends huge cuts to the policy rate right at the start of the financial crisis, but subsequently consistently argues for rate increases.

\em The first years of the 21st century

From 2003 through to the beginning of the financial crisis, the actual policy rate was gradually increasing. In this period, BMA consistently recommends rapid rate increases. In contrast, BPDS recommendations are generally similar to what actually transpired, apart from at the beginning of this period where the advice is to raise the policy rate more slowly than what actually occurred.

\em The 1990s

During this period, the pattern is more mixed. Optimal policy recommendations generated under each of BMA and BPDS often differ from actual decisions, with no consistent pattern; at times the recommended rates are higher than the actual policy rate, and at other times they are lower.

figure[figure omitted — 203 chars of source]

A general pattern, one that occurs throughout the sample period, is that BMA and BPDS typically recommend larger changes in policy rates than were actually implemented by policymakers. Part of this is presumably due to differences between the policymakers' utility function and those used in our analyses. Other possible explanations are that policymakers can affect expectations through their communications, which is a channel not captured in the model, or their models use a much steeper Phillips Curve. Also, we focus only on inflation up to two years ahead, without considering the possibility of an over- or under-shoot of inflation after eight quarters. In contrast, policymakers would generally aim for inflation to be sustainably at target over the longer-term.

Trajectories of BPDS and BMA Predictive Densities

Figures (ref)--(ref) shed light on these patterns. These images represent the time trajectories of predictive densities of inflation at each relevant horizon, using BPDS (Fig. (ref)) and BMA (Fig. (ref)), as well as their differences (Fig. (ref)). These indicate that the BPDS mixtures are less dispersed than the BMA mixture for much of the sample period; that is, BMA predictive distributions are relatively more heavy-tailed, especially at longer horizons. This is partly due to the BPDS score function emphasizing that the policymaker wants to avoid extreme inflation outcomes, and also accounts for why BMA often tells policymakers to make larger changes to the policy rate than BPDS, as discussed in Section (ref). The differences between BPDS and BMA become larger at longer forecast horizons. Medium- and long-term macroeconomic forecasting is difficult, which leads to standard methods such as BMA producing fairly dispersed predictive densities at longer horizons. BPDS, on the other hand, is reducing this effect, which dampens the BPDS optimal decisions and reduces predictive uncertainty relative to BMA. Then, differences between BPDS and BMA forecast densities are reduced after the financial crisis, which helps account for why their policy recommendations are similar in the last decade of the sample.

figure[figure omitted — 418 chars of source]
figure[figure omitted — 418 chars of source]
figure[figure omitted — 378 chars of source]

Model Probabilities

Figure (ref) shows trajectories of model probabilities under BPDS and BMA. At each $t$, these are the discounted AVS prior model probabilities $\pi_{tj}$, implied initial decision-dependent probabilities $\pi_{tj}(\mathbf{x}_t)$ evaluated at the BPDS-optimal decision $\mathbf{x}_t$, resulting BPDS probabilities $\tilde\pi_{tj}(\mathbf{x}_t)$ of \eqn{fj}, and standard BMA probabilities.

figure[figure omitted — 783 chars of source]

Under traditional BMA, the two model probabilities are appreciable until the financial crisis. After the start of the crisis, the less parsimonious ${\mathcal M}_2$, which includes additional financial variables, receives virtually all the weight. In contrast, BPDS weights vary more over time, allocating most of the weight to the parsimonious ${\mathcal M}_1$ for much of the period (i.e., 1997 through 2017), though ${\mathcal M}_2$ plays more of a role at both the beginning and end of the sample period. That BPDS generally favours the more parsimonious ${\mathcal M}_1$, with less dispersed forecast distributions, partially accounts for why BPDS often dampens extreme recommendations made when using BMA.

BPDS probabilities on the over-dispersed ${\mathcal M}_0$ are generally small, though with notable increases at the start of the COVID-19 pandemic. In such extreme times, when neither ${\mathcal M}_1$ nor ${\mathcal M}_2$ forecasts well, the increased probability on the fall-back ${\mathcal M}_0$-- though small-- provides an indicator of this.

The BPDS prior model probabilities $\pi_{tj}$ based on discounted AVS differ noticeably from BMA probabilities (except at the very start of the time period). A big impact then arises from the conditioning on information generated in the decision space to map these $\pi_{tj}$ to the decision-dependent weights $\pi_{tj}(\mathbf{x}_t)$ at the BPDS optimal decisions $\mathbf{x}_t$ at each time. Recall that this mapping theoretically properly takes into account the likelihood of future, as yet unobserved, interest rate outcomes; this relevant information is not accounted for in the prior weights $\pi_{tj}$ and is, of course, absent under BMA. The subsequent map from initial probabilities $\pi_{tj}(\mathbf{x}_t)$ to the BPDS weights $\tilde\pi_{tj}(\mathbf{x}_t)$ is wholly based on the impact of the entropic tilting towards \lq\lq more favourable" decisions in expectation. We see that the impact is rather small over time, and this is to be expected: the BPDS analysis uses \lq\lq small" perturbations of the initial mixture based on target expected scores that are only modest increases over those under the initial mixture. We expect to see slight tilting towards models that are expected to do well, but not large changes relative to the initial probabilities.

Additional Insights from BPDS Results

Time trajectories of the tilting vectors ${\bm\tau}_t(\mathbf{x}_t)$ evaluated at the optimal decisions $\mathbf{x}_t$ are shown in Figure (ref). The values generally tend to increase with horizon $h$, thus attaching more weight to longer forecasting horizons. This is partly to be expected due to the higher uncertainties at longer forecast horizons.

figure[figure omitted — 773 chars of source]

Figure (ref) plots trajectories of several effective sample size (ESS) measures arising from the importance sampling to simulate BPDS predictive distributions, as discussed in Section (ref). This provides a read-out of the extent of tilting the initial mixture $p(\mathbf{y}|\mathbf{x})$ to the BPDS mixture $f(\mathbf{y}|\mathbf{x})$, as well as that for tilting each of the individual model pdfs from $p_j(\mathbf{y}|\mathbf{x},{\mathcal M}_j)$ to $f_j(\mathbf{y}|\mathbf{x},{\mathcal M}_j)$ (again, time-indexed and updated throughout the time series). Until the COVID recession, the ESS of the initial mixture is stable between 90--95%, suggesting only a small amount of tilting, as desired. The COVID recession is a period of rapid change, as expected, as we see large changes in the initial weights and larger values of ${\bm\tau}$ required to achieve the desired target. The low value of ESS indicates that at that point, the target expected scores are unrealistic given the then-current state of the economy. However, the resulting decisions during this time period appear to be rather sensible. This means we do not need to be too concerned about the low ESS, which can, in any case, be redressed by simply increasing the overall Monte Carlo sample size accordingly. The ESS values of individual models are generally lower than that of the overall mixture and somewhat more volatile. One nice point is that, even when one of the models seems to suffer a low ESS, the BPDS mixture ESS is generally maintained at higher values. This indicates that BPDS is able to strike a balance in weighting expected versus historical performance of models on both predictive and decision outcomes.

Finally, Figure (ref) compares the realized trajectories of expected utilities under BPDS and BMA. Each uses the same utility function to define the final optimal policy path decision, so these are directly comparable, and the comparison is relevant in terms of the setting of forward, sequential decisions where a change to much lower values at any time point should signal concern to the decision-maker. BPDS is designed to target an expected utility higher than that of the initial mixture, but whether it achieves a higher expected utility than BMA-- which has different initial probabilities and lacks outcome-dependent weighting-- is a question for empirical study. In this example, as illustrated in the figure, BPDS utility does exceed that of BMA in virtually every period. After the financial crisis, the two are similar, consistent with the earlier finding that they typically produce similar decisions during this time period. However, before the financial crisis, there are several periods during which the BPDS expected utilities are substantially higher than those of BMA. These correspond to times where we see more differences between optimal policy path recommendations. Within these times there are some periods of greater concordance between BPDS and actual policy decisions, as well as more constrained (i.e., less extreme) recommended decisions under BPDS relative to BMA.

figure[figure omitted — 207 chars of source]

Summary Comments

BPDS is the formal, foundational Bayesian framework that extends traditional Bayesian model uncertainty analysis to address explicit use of model-specific decision outcomes as well as purely predictive performance in model comparison and combination. This paper has adapted the BPDS foundations to define implied methodology in formulating macro-economic decision-making when faced with multiple objectives and multiple outcomes of interest in the monetary policy setting.

Earlier applications of BPDS focused on financial forecasting and portfolio decisions TallmanWest2023,TallmanWest2024. In this setting, forecasting models do not (generally) depend on the decisions of interest, while utility functions may and often do depend on the models and their predictions. In contrast, the setting of monetary policy analysis is one in which the dependence of models and their forecasts on the decision variables (policy instruments) is simply fundamental. It is also a setting in which the decision variables are treated simultaneously as outcomes. The future paths of central bank interest rates, for example, are modelled as time series outcomes along with other economic and financial indicators in VAR models. This then leads to conditioning on decision variables to define predictions of other indicators, with consequent implications for relative model weights in the model uncertainty setting. This latter point is critical as it then leads to relatively up- or down-weighting a model based on how well-supported a particular candidate decision is under its predictions; to our knowledge, this is the first time this central question has been formally, statistically addressed. These central features of predictive decision-making in monetary policy contexts are addressed with extensions and customization of the existing theory of BPDS.

The BPDS perspective-- of integrating historical and expected decision outcomes with focused aspects of statistical predictive performance into relative model weightings-- is new to the policy arena. We argue for this perspective since policymakers are primarily interested in using sets of models for the eventual policy decisions. Pure forecasting exercises-- and evaluation and combinations of models for prediction {\em per se}-- are, of course, of parallel interest and importance. We emphasize that BPDS also involves addressing predictive performance on specific, defined outcomes of interest. But most importantly, by putting the spotlight on decision-making, we gain additional insights into policy-making that are not possible in exercises that focus solely on predictive performance.

In a recursive, real-time decision-making exercise, we find substantial differences at various periods of time between the policy recommendations of BPDS and the traditional Bayesian model averaging approach, though good concordance at other times. When recommended policy decisions differ between the approaches, in most cases the BPDS policy paths are more intuitively sensible and less extreme than under BMA, and more consistent with the actual decisions made by the policymakers at the time. The case study presented investigates and interprets aspects of BPDS in terms of differential model weights based on historical information alone, and then updated based on identified optimal decisions, with consequent insights into how the differences relative to standard BMA arise and are exploited. This case study is a first step towards broader development and evaluation of BPDS in a setting with larger numbers of econometric models. Parallel next steps can naturally be expected include broader evaluation in terms of multiple models, and developments for scenario forecasting. Such extensions, in collaboration with policymakers, will advance understanding the sensitivity of model-based recommendations relative to chosen potential economic scenarios.

\setstretch{1.1}

Acknowledgements

Research of Emily Tallman was partially supported by the US National Science Foundation through NSF Graduate Research Fellowship Program grant DGE 2139754. Any opinions, findings, and conclusions or recommendations expressed in this material are those of the authors and do not necessarily reflect the views of the National Science Foundation. Further, the views expressed in this paper are solely those of the authors and may differ from the official views of the Bank of Canada. No responsibility for the views expressed in this paper should be attributed to the Bank of Canada.

\setstretch{0.9}

\setstretch{1}