Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
87,393 characters · 24 sections · 75 citation commands
A New Perspective of the Meese-Rogoff Puzzle: Application of Sparse Dynamic Shrinkage
\affil[1]{ Department of Economics, University of Melbourne} \affil[2]{ Department of Econometrics and Business Statistics, Monash University}
\noindentKeywords: Bayesian Econometrics, Shrinkage Methods, Sparsity, Model Combination, Variable Selection, Exchange Rate Prediction
\noindentJEL Classification: C11, C14, C52, C53
Forecasting exchange rate is important for policy makers, investors and international traders. The seminal Meese–Rogoff Puzzle of meese1983empirical claims that no model can systematically beat the random walk, with subsequent literature predominantly focusing on the point forecast. Thirty years on, rossi2013exchange confirmed such claim with an exhaustive empirical work. On the other hand, rossi2013exchange also inspires that
This work motivates economic theoretic driven studies such as ferraro2015can,cheung2019exchange,engel2019uncovered,candian2023imperfect,neghab2024explaining amongst others. A more recent and comprehensive survey is conducted in fang202430, who thoroughly reviewed the literature in this area. They found no influential findings after rossi2013exchange, with the exception of ferraro2015can, even though many researchers have attempted to address the puzzle.
Challenges in exchange rate forecasting include: (1) dynamic instability, (2) limited sample size, and (3) model uncertainty. There are other concerns such as data frequency choice and the discussion of big and small economy, but we focus on the aforementioned three points in this paper. Our novel econometric approach addresses these three challenges and indeed provide a dominating predictive outcome in many out-of-sample metrics than the random walk. The foundation of our econometric models is built upon established economic theories, with historical research on the relationship between exchange rates and economic fundamentals respected in our framework.
From the econometric modelling perspective, many research addressed dynamic instability. Existing popular method such as time-varying parameter (TVP) models wolff1987time,canova1993modelling,mumtaz2013time,byrne2016exchange cater for gradual change of the parameters and regime-switching models engel1994can,nikolsko2012markov allow for sudden parameter change. These models are flexible and well-tested in time series forecasting. However, to our knowledge, no paper has seriously addressed the importance of model parsimony in the presence of limited sample size for exchange rate forecasting. A data-driven balance between flexibility and parsimony, in our view, could shed lights on the Meese–Rogoff prediction puzzle.
To answer challenges (1) and (2), we consider the dynamic shrinkage process (DSP) of kowal2019dynamic. The DSP is a process taking the form of a state space model, with the innovation taking a particular form such that the process could experience a decent duration of large negative values. See also hauzenberger2024dynamic, who applied DSP to their global shrinkage factor. When used to model the log volatility of a dynamic parameter, the DSP allows for the variation of the time-varying parameter to approach zero, reducing its temporal variation. This approach shrinks the parameters towards the value in the previous period rather than shrinking them to zero. The same idea can be found in dufays2021sparse, but their model is computationally intensive because it requires a joint movements of the parameters at adjacent time points. In contrast, the DSP works as a clip by collecting adjacent parameters into a single value through the innovation's volatility, which is a stark contrast to the standard TVP model that assumes constant volatility. Another closely related research is knaus2023dynamic, who proposed dynamic triple gamma prior to model innovations in the dynamic linear model framework.
In addition to shrinkage of volatility, huber2021inducing emphasized the empirical importance of sparsity. They pointed out that setting irrelevant parameters to zero may greatly improve prediction. Motivated by west2006bayesian, kalli2014time developed a Normal Gamma Autoregressive process (NGAR) that can mask segments of a time series of coefficients to be numeric zero. Similar ideas include the dynamic slab and spike process in rockova2021dynamic and the Markov mask of latent coefficients in uribe2020dynamic and bernardi2023dynamic.\footnote{ The idea of sparsity via masking can be traced back to nakajima2013bayesian.} lopes2022parsimony, instead, use a 4-component mixture prior to select between TVP and constant coefficient models, but they assume static mixtures. There is no dynamic framework to date that allow both sparsity and (potentially) multiple regimes of non-zero, but constant, coefficient in the TVP setting.
We bridge this gap by introducing the Markov Switching DSP (MSDSP), which allows for the time-varying coefficient to dynamically switch between the zero-state (achieving sparsity) and the DSP state (achieving shrinkage). To illustrate the flexibility of our proposed framework, we plot four synthetic parameter paths in Figure (ref). The first parameter $\beta_0$ is a constant $0$ which renders its corresponding covariate (or intercept) useless. The second parameter $\beta_1$ changes values twice, with all three values being none-zero. The third parameter $\beta_2$ is zero at the two ends of the illustrative period, and is only “activated” in the middle period. The last parameter $\beta_3$ experiences gradual change, becomes abruptly “inactivated” at zero, and then gradually move towards a fixed none-zero value. All these behaviors can be captured by our MSDSP framework.
In the exchange rate forecasting application, we follow the established benchmark models from rossi2013exchange and find that when we apply our proposed MSDSP method to the economic models, they substantially outperform the random walk across a range of metrics. In addition applying the MSDSP to enhance the individual models, we find that the MSDSP naturally lends itself to the Bayesian predictive synthesis (BPS). BPS is a novel and powerful model assembly method developed by mcalinn2019dynamic. It takes care of potential model bias and allows unrestricted model weights (even negative) to optimally combine many sub-models or experts, see mcalinn2020multivariate,tallman2024bayesian for recent applications. In the BPS context, the MSDSP can be used directly on model weights in model assembly. These weights can move as in a state space model with parsimony (DSP-state) or simply shut down (zero-state) for a while and then reactivate (or not). We apply the MSDSP assembly method and compare it to some existing approaches. We find that the MSDSP assembly method works the best among multiple well-known assembly methods. With this, we answer to the aforementioned challenge (3).
The remainder of the paper is organized as follows. Section 2 outlines the construction of the MSDSP framework for TVP models, along with the Bayesian inference proposed for the the framework. A simulation study to assess the advocacy of the framework is provided in Section 3. In Section 4, focus is given to a detailed application of the MSDSP framework on existing economic models to address the Meese-Rogoff puzzle, while Section 5 provides an application of the MSDSP in model assembly for exchange rate forecasting. Section 6 concludes.
The Markov Switching Dynamic Shrinkage Process (MSDSP) provides enhanced flexibility to the state space model that encapsulates the TVP specificaiton in many macroeconomic models. The state-space model, tracing back to kalman1960new, is a statistical framework for dynamic systems, with the most commonly used being the Gaussian Dynamic Linear Model (DLM) harrison1976bayesian, west2006bayesian, assuming linear relationship between the observables and the state, with normally distributed innovations. The TVP model for time series analysis belongs to this class, and can be specified as:
where the innovation $\epsilon_t$'s variance $\sigma^2_t$ and $\bm{\omega}_t$'s covariance matrix $\Omega_t$ are pre-determined or a constant. The variable of interest $y_t$ is a scalar, vector $\bm{x}_t$ is the $p\times 1$ observed data including the intercept. The time-varying parameter $\bm{\beta}_t$ is a $p\times 1$ vector.
Contemporary macroeconomic or financial applications emphasized the time varying volatility. For example, stochastic volatility (SV) has been advocated and evidenced by numerous empirical works primiceri2005time, justiniano2008time, nakajima2011time,chan2018bayesian,chan2024large. Here, we base our model on the DLM with SV specification of aguilar2000bayesian, who assumes (ref)-(ref), but with $\sigma_t^2=\exp (g_t)$, and
The log volatility $g_t$ follows an AR(1) process, with stationarity imposed by $|\phi_g| <1$. In addition, we assume a diagonal matrix for $\Omega_t$ in (ref) for simplicity.
Our goal is to extend the TVP modelling framework to account for both sparsity and shrinkage of the coefficient to a non-zero value. To this end, we extend the work of kowal2019dynamic, and introduce $p$ two-state Markov switching processes to allow for the time-varying coefficients to “switch off” to zero. The model, in full, takes the form of
for $i=1,..,p$, $t=1,...,T$ and $j,k\in \{0,1\}$. The $\circ$ in Equation (ref) is the Hadamard (pointwise) product, and $\bm{0}$ is a $p\times 1$ vector of zeros. The introduction of the $p\times 1$ variable, $\bm{s}_t=(s_{1t},...,s_{pt})$, allow for the elements of $\bm{\beta}_t$ to take value either the value of zero or the corresponding element of the “shadow” coefficient $\tilde{\bm{\beta}}_t$, based on the Markovian transition probability in (ref). We assume that Markovian switches operate independently across the $p$ dimension, such that $\bm{s}_i \perp\!\!\!\perp \bm{s}_j$ for $j\neq i$ and $\bm{s}_i = (s_{i1},...,s_{iT})$.
The element of the shadow coefficient $\tilde{\bm{\beta}}_t$, $\tilde{{\beta}}_{it}$ for $i=1,\dots,p$, follows the Dynamic Shrinkage Process (DSP) of kowal2019dynamic. The DSP allows for the latent coefficient itself to exhibit SV structure through the modelling of the volatility of each innovation term $\omega_{it}$ in (ref)-(ref). kowal2019dynamic proposed the use of the Z-distribution as a prior in (ref), with the feature of this particular distribution allowing for extreme negative values in the log volatility term $h_{it}$. The Z-distribution, introduced by Barndorff-Nielsen1982-yc, is a flexible class of normal mixture that allows for extreme skewness and fat-tails. In the DSP context, we employ the mixture constructed using the logistic transformation of the beta variable with parameters $\alpha_h$ and $\beta_h$, with zero location and unit scale for identification. In its extremes, \( \eta_{it} \to -\infty \) results in the log-variance \( h_{j,t} \) approaches $-\infty$. In such circumstances, the volatility of $\tilde{{\beta}}_{it}$ approaches zero, rendering a constant coefficient environment with $\tilde{{\beta}}_{it}=\tilde{{\beta}}_{i,t-1}$ .
Our proposed MSDSP specification offers additional parsimony and flexibility to the TVP model as follows. First, each element of $\bm{\beta}_t$ is governed by independent dynamic shrinkage and Markov switching processes, allowing for each covariate to behave independent through the sample period. The general framework also allows for the potential for the TVP to switch to zero for sparsity, to shrink to a constant value, or to revolve dynamically according to the state space structure. Our specification for sparsity and shrinkage is internally coherent, as also done in rockova2021dynamic, which is in contrast to the post-processing approaches of huber2021inducing,Hahn2015-an, among many.
The general MSDSP framework proposed here nests several existing models. The DSP proposed by kowal2019dynamic is its most obvious special case, which can be achieved when the transition probability \(P^i_{jj}=1\) for both \(j \in \{0,1\}\) and the two-state Markow switching process reduces to a one-state scenario. This coincide with the original TVP with DSP specification. In the case when \( \alpha_h = \beta_h = 1/2 \), and in the absence of dynamic structure, in the specification of the Z-distribution in (ref), the model corresponds to a horse-shoe prior also discussed in kowal2019dynamic. Of course, the original TVP model is also nested within our framework, with \( \bm{\beta}_{t} = \tilde{\bm{\beta}}_t \) and the latent time-varying parameter exhibit homoskedasticity. Lastly, if $\tilde{\bm{\beta}}_{it}$ is a constant, the model assumes a dynamic stochastic search variable selection (SSVS) method for the corresponding variable $\bm{x}_i$.
A closely related framework is that of bernardi2023dynamic, who employ one latent parameter process and one masking process, with roles similar to $\tilde{\beta}_{it}$ and $s_{it}$, respectively, in our framework. Their latent parameter process is a simple random walk (or some joint normal distribution) instead of DSP in this paper, so their latent model cannot afford short-term constant coefficient. The mask parameter is determined by an inverse logit function with input from another latent normal distribution, with the normality assumptions lending itself naturally the variational inference in their paper.
We use Bayesian computation to conduct inference on the MSDSP model. Using the Polya-gamma representation for the DSP polson2013bayesian and the mixture technique for the SV specification Omori2007-is, we seek to establish the augmented posterior:
where \(\bm{y}_{1:T}=(y_1,\dots, y_T)^\top\), \(\bm{x}_{1:T}=(\bm{x}_1,\dots,\bm{x}_T)\) with \(\bm{x}_t=\{x_{it}\}_{i=1}^p\) and the model unknowns are collected in \(\Psi\). The model unknowns include any auxiliary variables from the Polya-gamma resprentation and the SV mixture representation that is amendable to the Gibbs sampler in the Markov chain Monte Carlo (MCMC) sampling to estimate our posterior distribution. The elements of \(\Psi\), along with their respective description and dimensions are summarized in Table (ref)
Within our MCMC algorithm, we utilize the techniques to sample the components related to the DSP component outlined in kowal2019dynamic. Standard algorithms for Markov switching and SV inference are adopted. Algorithm (ref) outlines the steps for our MCMC inference, with full details of the algorithm and the prior distribution given in the Appendix. In all our inference, we construct \(G=4,000\) posterior samples after $30,000$ burn-in draws, with every 5th draw retained through thinning. Simulation consistency inference can be carried out from our $G$ this posterior samples. For instance, if the posterior expected value of $\bm{\beta}_t$, or simply $E(\bm{\beta}_t\mid \bm{y}_{1:T},\bm{x}_{1:T})$, can be estimated by the posterior sample mean $\frac{1}{G}\sum\limits_{g=1}^G \bm{\beta}_t^{(g)}$.
In the application of exchange rate forecasting, our primary focus is to construct the predictive distribution for future exchange rates. The model specified in Section (ref) specifies a contemporaneous regression model in (ref). Without an explicit dynamic for the covariate, $x_t$, constructing predictions for future time points is not feasible. In order to construct the $h-$step-ahead predictive distribution, we propose the direct $h-$step-ahead forecasting model by adjusting the measurement equation (ref) to
with $\epsilon_t \sim N\left(0,\exp(g_t)\right)$ and the remainder of the model components remain as defined in (ref)-(ref).
The out-of-sample $h$-step-ahead posterior predictive distribution, integrating out the uncertainty in the model unknowns, takes the form of
where $I_T=(\bm{y}_{1:T},\bm{x}_{1:T})$ denotes the information set up to time $T$. By construction, the predictive distribution, even for $h>1$, can be constructed using the covariate observed within the sample period. In addition, the implied TVP model structure, captured by \(p(y_{T+h} \mid x_T,\bm{\beta}_T,g_T)\), reflects the specific dynamic relationship between the exchange rate and the $h-$lag covariate without the need to further assume the evolution of such relationship between periods $T$ and $T+h$.
Note that when applying the model in (ref), the posterior distribution $p(\Psi \mid \bm{y}_{h+1:T}, \bm{x}_{1:T-h})$ needed in (ref) only include the inference of the latent quantities up to time point $T-h$. That is, we obtain $G$ posterior daws of $(\bm{s}_{T-h},\tilde{\bm{\beta}}_{T-h},\bm{h}_{T-h},g_{T-h})$. In order to produce the $h-$step-ahead prediction, we simulate forward using the dynamic specification in (ref)-(ref) to obtain $G$ draws of $\bm{\beta}_T$ and $g_T$, upon which $p(y_{T+h} \mid x_T,\bm{\beta}_T,g_T)$ conditions on. By simulation, we obtain $G$ draws of $y_{T+h}$, $\{y^{(g)}_{T+h}\}_{g=1}^G$, from its conditional predictive distribution. All predictive statistics in the application can be inferred from this framework. For example, the expected value $E(y_{T+h}\mid I_T)$ can be estimated by the sample mean $\frac{1}{G}\sum\limits_{g=1}^G y_{T+h}^{(g)}$. The predictive density $p(y_{T+h}\mid I_T)$ can be estimated by the sample mean of the Gaussian conditional densities implied by (ref) as $\frac{1}{G}\sum\limits_{g=1}^G p(y_{T+h}\mid x_T, \bm{\beta}^{(g)}_T, g_T^{(g)})$, evaluated over the plausible grid value of $y_{T+h}$. We evaluate the prediction using five metrics geared to evaluate the predictive distribution. These metrics are defined and discussed in Section (ref).
We carried out two simulation studies to illustrate the effectiveness of the MSDSP framework. In particular, we highlight the ability of the MSDSP in flexibly capturing multiple regimes, with the posterior inference of these dynamic parameters efficient for both sudden and gradual shifts in the TVP.
We consider a scenario with $5$ explanatory variables and an intercept, for $T = 300$ periods. The explanatory variables $\bm{x}_t=(1,x_{1t},...,x_{5t})$ are generated by following the structure from george1997approaches. First we generate six independent standard normal random variables, $z_0,...,z_5$. We generate five correlated covariate variables by $x_{1t} = 0.4z_0+z_1$, $x_{2t}=x_{1t}+1.8z_2$, $x_{3t}=0.4z_0+ z_3$, $x_{4t} = x_{3t} + 1.4z_4$, and $x_{5t} = x_{3t} + 1.4 z_5$, resulting in the covariate vector $\bm{x}_t$ that are independent across time $t$. The dependent variable $y_t$ is generated by $y_t = \bm{x}'_t\bm{\beta}_t + \epsilon_t$, where the error term $\epsilon_t$ is independently drawn from a standard normal distribution. The coefficients, $\bm{\beta}$, are chosen such that some of the parameter values abruptly changes between different regimes of constant values or zero. In particular, $\beta_0$ is set to zero, rendering a zero intercept. $\beta_2$ and $\beta_5$ are both zero across all time points, rendering zero impact from the covariates $x_{2t}$ and $x_{5t}$. $\beta_1$ has 3 regimes, with the parameter values being constant but non-zero in each regime. $\beta_3$ has 3 regimes, with the parameter taking non-zero constant values in the first two regimes before switching to a zero value in its last regime. $\beta_4$ also exhibit 3 regimes, but with the second of three regimes being the zero-state, while the other two being non-zero constants, implying that the impact $x_{4t}$ is switched on and off. The dynamic of the true parameter values are shown in the green dashed line in the top and middle panels of Figure (ref)
Figure (ref) depict the comparison of posterior inference of the TVP between the DSP and the MSDSP models. The top panel shows the posterior means (solid yellow line) and point-wise 95% density intervals (gray shade) inferred from the DSP model. The middle panel depicts the posterior means (solid yellow line) and point-wise 95% density intervals (gray shade) inferred from the MSDSP model. The bottom panel plots the true Markov switching state $\bm{s}_t$ (green dashed line) and the posterior probability of being in the non-zero DSP-state (yellow solid line). In inference of both the DSP and the MSDSP models, we use identical uninformative prior specification for the parameters related to the DSP component for comparability.
From the bottom panel in Figure (ref), the true state and the inferred state probability almost coincide, which indicates an almost perfect ability of the MSDSP model to switch from the zero state to the DSP state. Comparing the top and middle panels, we observe that when a coefficient's true value is in the zero-state, the MSDSP will quickly adapts to that state. This is particularly evident in the inference of $\beta_0, \beta_2$ and $\beta_5$, where the posterior uncertainty is much smaller for the MSDSP. Since the DSP is unable to shrink to to absolute zero-state, the the posterior inference over these true zero-states still exhibit a larger degree of uncertainty, even when the posterior means are hovering around zero. Similarly, the zero-state episodes for $\beta_3$ and $\beta_4$ are well captured with much tighter posterior intervals when the MSDSP is used. For non-zero coefficient values, the MSDSP and DSP performs similarly, with both models able to capture the shifts of the coefficients to a non-zero constant value effectively.
For this case, we simulate three explanatory variables, $\bm{x}_t=(1,x_{1t},...,x_{3t})$, from independent standard normal distributions with $T=300$. As in Case 2, $\bm{x}_t$'s are independent across time $t$. The coefficients, $\bm{\beta}$, are given in Figure (ref) and described in the Introduction. They are also shown as the green dashed lines in the top and middle panels in Figure (ref). Again, the dependent variable $y_t$ is generated by $y_t = \bm{x}'_t\bm{\beta}_t + \epsilon_t$, where the error term $\epsilon_t$ is independently drawn from a standard normal distribution.
This collection of coefficients encompasses all possible scenarios that the dynamic TVP, and we compare the DSP and MSDSP models in this setting. We highlight that the difference between Case 1 and Case 2 is that Case 2 exemplifies the gradual change in parameter in the scheme for $\beta_3$. As in Figure (ref) for Case 1, Figure (ref) depicts the posterior mean and point-wise 95% posterior interval of $\bm{\beta}_t$, with those produced by the DSP model depicted in the top panel, and those from the MSDSP depicted in the middle panel. The bottom panel depicts the posterior summary for the Markov switching variable $\bm{s}_t$. Consistent with what we observed in Case 1, when the coefficient enters the zero-state, the MSDSP model is able to capture this scenario with a much higher degree of certainty. When the parameters are non-zero, the MSDSP and DSP models perform similarly, including when the shift in the parameter is gradual, as in the case of $\beta_3$.
We employ the flexible MSDSP model to offer a novel econometric perspective on the Meese-Rogoff puzzle. In particular, we seek to establish if the model's ability to achieve sparsity and incorporate dynamic parameters that can potentially shrink toward a constant makes economic models competitive with random walk predictions. To set notation, let $e_t$ denote the logarithm of spot change rate between a country to the US dollar at time $t$, and we construct the response variable as $y_t = e_t - e_{t-1}$ as the growth rate of exchange rate at time $t$. We use the direct quotation method, hence a positive $y_t$ means that the domestic currency depreciates against the US dollar.
The Meese–Rogoff Puzzle meese1983empirical, meese1983out, meese1988was claimed that the random walk model is a competitive model that is difficult to beat for exchange rate prediction. Therefore, we view the random walk as the benchmark model and augment it with a SV component to reflect the contemporary development in the empirical macroeconomic modeling literature. Specifically,
If we restrict the $\bm{x}_t=1$ in the measurement equation, (ref), and assume a constant intercept, the MSDSP model reduces to the random walk model. We use both the random walk model with non-zero drift (RW-drift-SV) and the driftless random walk where $\mu=0$ (RW-SV). As pointed out by rossi2013exchange, the RW-SV is the toughest benchmark to beat. Since both RW-drift-SV and RW-SV models contains no external covariates, we apply the indirect method to recursively construct the $h$-period ahead forecast from these models.
We respect the existing literature, as summarized by rossi2013exchange, to include four different models driven by economic theory, along with three further predictive models based on commodity prices (oil, gold, and copper) inspired by ferraro2015can. All seven competing economic models can be written in the form of a linear model in (ref), but with varying set of covariates. We summarize the specification of the covariate dictated by each of the economic model in Table (ref).
The adoption of MSDSP (and its nested DSP) in the time-varying coefficient in these models allow for the coefficient to move over time, as well as allow for stochastic volatility in the predictive distribution of exchange rate changes. Both features in the econometric modelling of these economics theories have been tested in the contemporaneous macroeconomic literature. We employ both the MSDSP and the DSP specifications on all seven competing economic models listed in Table (ref). Our analysis provide a new perspective to the Meese-Rogoff puzzle, with the analysis allowing for additional flexibility in the economic models to allow them to adapt to changing economic environments.
The two random walk models, RW-SV and RW-drift-SV, along with the constant coefficient versions of the seven competing economic models, serve as benchmarking models in our analysis. We refer to the constant coefficient economic models simply as “linear” models. For the linear economic models, we consider expanding window as well as moving window of 100 months and 50 months to capture potential structural instability brutally. All models, including the RW-SV, RW-drift-SV and the linear benchmark models, are inferred in the Bayesian framework. Hence, we take account the parameter uncertainties and are able to provide a full profile of the predictive distribution, including densities and point forecasts.
We evaluate our predictive distributions over the forecast horizon $h=1,...,12$. Since we employ monthly data, these horizons corresponds to from 1 month to 1 year. All predictions are out-of-sample by using expanding as delineated in Section (ref), unless stated otherwise. For completeness, we also report the contemporaneous forecasts conducted in ferraro2015can, which is equivalent to predicting $y_{t+h}$ by assuming $x_{t+h}$ exists, in the appendix. \footnote{ This approach is a pseudo prediction since it ignores simultaneity. It is not a fair comparison to the random walk model, but can provide information for comparison between economic models by stripping off predictor's uncertainty.}
We consider five metrics to evaluate the predictive distributions of the competing models: log predictive likelihood (Bayes factor), continuous ranked probability score (CRPS), root mean squared forecast error (RMSFE), tail coverage rate, and model confidence set (MCS). All metrics are out-of-sample and calculated based on the Monte-Carlo estimate of (ref), as described in Section (ref). In what follows, we define $T_0$ as the initial sample period, and $T$ denotes the overall sample size, with the number of $h$-period ahead periods evaluated being $T-h-T_0+1$.
First, the predictive distribution is evaluated by the log predictive density ratio (LPDR), which evaluates the predictive performance of model $\mathcal{M}$ relative to the benchmark model $\mathcal{M}_B$:
Here, \[\log(p(Y_{T_0+h:T}\mid I_{T_0}, \mathcal{M}))=\sum_{t=T_0}^{T-h}\log(p(y_{t+h}\mid I_t, \mathcal{M}))\] is the cumulative log score, with \(\log(p(y_{t+h}\mid I_t, \mathcal{M}))\) denoting the log predictive density of the $h-$period ahead prediction constructed based on information set up to time $t$. The LPDR is promoted by Geweke2010-kb, and can be viewed as the log predictive Bayes factor of model \(\mathcal{M}\) relative to the benchmark \(\mathcal{M}_B\). A value that is larger than $3$ means that model $\mathcal{M}$ is strongly preferred to the benchmark model $\mathcal{M}_B$ (see Kass1995-ke). With the RW-SV being identified by rossi2013exchange as being the toughest model to beat, the RW-SV model serves as our $\mathcal{M}_B$ for all subsequent analyses.
The continuous ranked probability score (CRPS) assesses the full predictive distribution prioritizing the sharpness of the distribution and is thus less sensitive to tail behaviors gneiting2007strictly. It is defined, for model $\mathcal{M}$, as
where $F(r\mid I_t, \mathcal{M})=\int_{-\infty}^r p(\tilde{y}_{t+h}\mid I_t, \mathcal{M}) d\tilde{y}_{t+h}$ is the CDF for $y_{t+h}$ and $\bm{1}(.)$ is the indicator function which returns one when the condition in the parenthesis is true. We report the average CRPS
for model comparison. A smaller value means better performance.
For point forecast evaluation, the root mean squared forecast error (RMSFE) is one standard approach, and computed as follows:
where the conditional mean $E(y_{t+h}\mid I_t, \mathcal{M})$ is calculated by the sample average of the corresponding set of $\{y_{t+h}^{(g)}\}_{g=1}^G$.
To evaluate the tail behavior, we report the coverage rate (CR) at different tail quantiles. It is the ratio of the number of observations falling into a specified range to the total number of observations. For example, to assess the predictive 5% quantile, the coverage rate will be the number of observations that are below the predictive 5% quantile value to the total number of observations. Obviously, a model with a ratio closer to 5% would be favored in this example. In general, we define, for the lower-tailed $\alpha$% quantile prediction,
where $q^{\alpha,\mathcal{M}}_{t+h}$ is the predictive quantile from the predictive distribution, such that $Pr(y_{t+h}\leq q^{\alpha,\mathcal{M}}_{t+h}|I_t,\mathcal{M})=\alpha$. Analogously, it is straightforward to calculate the upper-tailed quantile coverage rate, with the “$\leq$” sign replaced with the “$\geq$” sign in all calculations above.
Lastly, we construct the Model Confidence Set (MCS, Hansen2011-hy, and executed in MSC_R) to select a list of `superior' models. The MCS controls for multiple comparisons to avoid inflated Type I error and is robust to model mis-specification, especially when evaluating out-of-sample predictive performance. We implement the MCS using three different out-of-sample loss functions for MCS: Squared Forecast Errors (SFE), negative Log Score, and Continuous Ranked Probability Score (CRPS). The MCS produces a set of “competitive” models within the given confidence set along with their respective rankings, with uncompetitive models discarded. We choose a 95% model confidence set in our application.
With the United Kingdom and the USA being key international financial market players, we present our application using the United Kingdom's British Pound (GBP/USD) exchange rate and discuss the results comprehensively. We also analyze the cases ofthe Canadian dollar (CAD/USD) and Japanese Yen (JPY/USD) exchange rates, but their results are provided in the appendix due to the volume of results from the variety of models investigated.
Following rossi2013exchange and Aristidou2022-ox, Aristidou2022-pg, we collected data from the following primary sources: FRED (Federal Reserve Economic Data), the IMF (International Monetary Fund), the OECD (Organisation for Economic Cooperation and Development), and the Bank of England. We employ the end-of-month nominal exchange rate for GBP/USD. As noted by rossi2013exchange, end-of-month data generally exhibits weaker predictive power, providing a lower bound for forecasting accuracy. The monthly total industrial production is used as a proxy for real GDP, while the real potential GDP is derived by using the one-sided HP-filter. The other times series include the Producer Price Index (PPI), monetary aggregates (M1), and commodity prices for crude oil, gold, and copper. All variables are standardized to have a mean of 0 and a variance of 1. Detailed data sources and series codes are provided in section A of the supplementary material.
Due to the availability of commodity prices and interest rates, the longest common sample period spans from January 1990 to June 2017, totaling T=330 months. The out-of-sample period is from Nov 2000 to June 2017, comprising 200 observations, leaving the initial sample period $T_0=130$ observations.
We first summarize that the economic models with the MSDSP flexibility perform the best in every out-of-sample evaluation metric. It consistently demonstrates superior forecasting performance compared to the random walks, producing higher LPDR, lower CRPS, lower RMSFE, more accurate coverage rate, as well as always attaining high ranks in the MCS assessment. At the same time, our results are also consistent with the literature, in that the random walk models are superior to the linear economic models (expanding or rolling window) in general. While there are a few exceptions to this general conclusion, but they are not systemic do not alter our overall conclusions that the added flexibility from the MSDSP prior indeed lead to better predictive performance of the economic models. Our observations also extend to the cases of Canadian dollar and Japanese Yen, with their detailed results provided in sections D and E of the supplementary material, respectively. In what follows, we discuss the GBP/USD results in more details.
Table (ref) reports the LPDR of the competing models, evaluated relative to the RW-SV benchmark, at selected horizons of $1,3, 6$ and $12$ months. In each row, the bolded statistic indicated the best predictive score, with positive numbers indicating that particular model outperforms the benchmark. It is evident that the MSDSP version of each economic model is always the best performing, with the MSDSP column always bolded. The results for other forecast horizons are provided in section C of the supplementary material and are consistent with those reported in Table (ref).
Importantly, the first three columns under “Linear” always produce negative LPDR, indicating that the basic economic models are always worse than the random walk, an observation that is consistent with the Meese–Rogoff Puzzle. Once the MSDSP is implemented, however, the results are overthrown. This is the decisive evidence that MSDSP can discover the nonlinearity and dynamic instability to improve exchange rate prediction. In addition, the predictive LPDR from the DSP column are all negative as well, highlighting the value added from the Markov switching process that we proposed, suggesting that the DSP itself is not parsimonious enough to produce a competitive density prediction relative to the RW-SV.
The above interpretation of results in Table (ref) is convincing, and more than sufficient to provide evidence for the predictive dominance of our MSDSP framework. Convincingly, the MSDSP outperforms the RW-SV at all horizon and with all economic models. Based on the magnitude of the LPDR itself, we can also infer the degree of the superiority of the predictive performance. For example, when $h=1$, the best performing model is the MSDSP-PPP, which is the time varying parameter PPP model with an MSDSP prior. The LPDR is larger than $3$ and hence it is strongly supported by data against the random walk model RW-SV. In fact, the MSDSP-PPP model strongly dominates the random walk at all horizons, providing a strong signal that MSDSP captures important data dynamic features. These observations are also confirmed by the analysis of CRPS metric, reported in Table (ref), with relative differences in this metric also confirming the superiority of the MSDSP models relative to the RW-SV, as well as its nested alternative the DSP specification.
From the angle of tail prediction, we report the 1-period ahead coverage rate for top and bottom $2.5\%$ and $10\%$, respectively, in Table (ref)\footnote{Due to page limitation, other horizon's results are available upon request.}. The MSDSP does not perform ideally in bottom $2.5\%$ and top $10\%$, but it is much better than the RW-SV in all other quantile regions. The MSDSP models also generally produces more accurate coverages than their linear counterparts. We observe that the RW-SV tends to provide tail predictions that produce under-coverage. For example, in the top panel (Top $2.5\%$) of Table (ref), the coverage rate from MSDSP is very consistent to $2.5\%$, but the random walk only covers $1.5\%$. In the third panel (Top $10\%$), the coverage rate from MSDSP is about $8\%$, but the random walk benchmark is only $4.5\%$. In summary, the MSDSP still performs better than the RW-SV.
The root mean squared forecast error (RMSFE) is the most common evaluation prediction metric for point forecasts, especially in the frequentist approach. The RW-SV model usually has its home advantage in this metric rossi2013exchange. We are excited to reveal that the MSDSP performs extremely well with point forecasts according to the RMSFE. Table (ref) shows the ratio of RMSFE of the competing models relative to that of the RW-SV benchmark model. A value less than $1$ means that the corresponding model produces smaller RMSFE statistic than the RW-SV benchmerk.
With only one exception, the MSDSP model produces smaller RMSFE statistics compared to the RW-SV model. The only exception is the linear economic model using copper prices, based on a 50-period rolling window scheme. Even in this case, the MSDSP method is still better than the RW-SV, and is only slightly worse than the linear counterpart. As in the case of density forecast, the MSDSP dominates the RW-SV benchmark, but the RW-SV is better than linear or DSP economic models in most cases. We also observe that the linear models based on commodity prices are also capable of performing better than the RW-SV model in point forecasting, although this cannot be generalized to density predictions.
We carried out pairwise Diebold-Mariano (DM) tests for all the models against the RW-SV. Unfortunately, we did not achieve many significant result despite the MSDSP column in Table (ref) indicating the direction for better performance. We also performed the Wilcoxon signed-rank test, which is a joint test based on the DM statistics from for all pairs of MSDSP models and the RW-SV benchmark ($7\times 12=84$). The joint test reports the p-value of $0.0000$, indicating that there is indeed a significant difference between the RMSFE of the MSDSP models compared to the RW-SV. We also performed the same Wilcoxon signed-rank test for each linear economic model with its MSDSP counterparts ($12$ statistics/horizons each). For all variants of the model, the test reports p-values less than $0.001$, with exception of the model that involve gold prices returning a p-value of $0.017$. This reflects a strong evidence that the superior performance of the MSDSP is supported by the data, even for point forecast.
Table (ref) reports the $95\%$ Model Confidence Sets for $h=1$ and $h=2$ months ahead forecasts based on three evaluation criteria: Squared Forecast Error (SFE), negative Log Score, and the Continuous Ranked Probability Score (CRPS). Each column report integer rank of each of the model included in the $95\%$ model confidence set, for the corresponding selection criterion and forecasting horizon. Symbol “-” indicates that the model is eliminated form the confidence set, which implies that the models are not ranked.
The MSDSP is always the number $1$ choice in the confidence sets. With only two exceptions, the MSDSP is always among top-10 choices. Considering there are 7 MSDSP economic models, this is a striking result. Amazingly, at horizon $h=1$, all three criteria admit that MSDSP method dominate any other models, since the seven MSDSP economic models ranks the top 7 in each of the model confidence sets. Although the RW-SV and RW-drift-SV models are not rejected and are retained in the confidence sets, their highest ranking is 8.
Once again, we observe that the economic models, along with the flexibility of the MSDSP, beats the RW-SV benchmark. Our MCS results also confirm prior discussions on the DSP model performance being inferior, and are often eliminated from the confidence set, highlighting the added value of Markov switching mechanism in exchange rate prediction. We also confirm the superiority of the RW-SV and RW-drift-SV compared to the linear economics models, with the benchmark models' rankings consistently higher than those of the linear models.
Since there has been weak consensus on the dominant model choice in exchange rate forecasting, a natural path is to combine the predictions via model assembly. Researchers applied the Bayesian Model Averaging (BMA) wright2008bayesian, della2009economic or dynamic Bayesian Model Averaging (DBMA) vcasta2024forecasting in exchange rate forecasting, combining many weak models to form a strong prediction. We do not intend to exhaustively explore the model combination method due the space limitation and our focus on MSDSP. We consider each prediction from a small model can be viewed as an expert's opinion, and propose a new model assembly method that naturally adapts to the MSDSP setting based on dynamic Bayesian predictive synthesis (DBPS) from mcalinn2019dynamic. The adaptation of the MSDSP allows for the predictive synthesis to dynamically adapt, with shrinkage of the combination parameters to a constant or a switch to no contribution in a single framework. We investigate if the use of MSDSP in this setting also leads to improved predictive performance in exchange rates.
We build our method on the dynamic Bayesian predictive synthesis from mcalinn2019dynamic. There is a burgeoning literature as johnson2017bayesian,McAlinn2020-ke,Aastveit2023-wv,Tallman2024-mu. The DBPS of mcalinn2019dynamic can be expressed as follows
where $v_t$ is defined via a standard beta-gamma random walk volatility model Prado2010-ec and $W_t$ follows a standard single discount factor specification West1999-ic. Each $\hat{y}_{t, j}$ is the prediction from model (or expert) $j$. The dynamic BPS considers distribution input, hence $\hat{y}_{t, j}$ is treated as random. The vector $\omega_t=(\omega_{t, 0},\omega_{t, 1},...,\omega_{t, L})'$ is the time-varying intercept and weights, with $L$ here denoting the number of models or experts. The intercept $\omega_{t, 0}$ captures the systemic time-varying bias. There is no restrictions to the model weights $\omega_t$. The absence of any restrictions provides a great deal of flexibility. For example, it allows a “bad” model to provide “good” results, as a negative weight is also permitted. It can also take advantage of a systemically upward biased model by giving it a discount, resulting the sum of model weights being less than unity.
The conditional posterior distribution of each model's prediction $\hat{y}_t=(\hat{y}_{t, 1}, ..,, \hat{y}_{t, L})'$ is updated as is given by:
where $ \alpha\left(y_t \mid \omega_t, \hat{y}_t\right)$ is the density implied by (ref), which is named as synthesis function. The predictors $\hat{y}_{t, j}$ are sampled from this conditional kernel. This sampling approach can be considered as tilting the originally standalone predictive distribution of model $j$ by the observations via the synthesis function and associated interactions with other models.
Conditional on a draw of $\hat{y}_t$, standard MCMC can be applied. To ensure reliability and consistency, both inference and prediction procedures are implemented by using the code provided on McAlinn's website. The hyper-parameters of the DBPS such as the discount factor are set the same as mcalinn2019dynamic. We have also performed robustness checks on these hyper-parameters and results are available upon request.
The idea of DBPS is very appealing to practitioners. But we do not want to increase additional computational cost on simulating $\hat{y}_t$. Hence, instead of inputting predictive distributions, we only input one prediction value of $\hat{y}_t$ such as the predictive mean and treat it as the data. This would free the cost of computation by make the assembly a two-stage task. In stage 1, we compute the prediction from each model. In stage 2, assemble these models by the synthesis function.
Our approach is closer to the pooling method such as static pooling Hall2007-ye,Geweke2011-ap and dynamic pooling Waggoner2012-vb,Del-Negro2016-gz. Some other variants include two-state Markov process pooling by Waggoner2012-vb and the infinite-state Markov process pooling by Jin2022-gj. There are two key differences. First, we use the MSDSP as the weight dynamics, so that sparsity is explicit. Second, we use the DBPS idea to incorporate the bias term in the synthesis function. Our method is also close to ensemble learning via stacking in machine learning WOLPERT1992-241, Breiman1996, Dietterich2000, Proscura2022, MUSLIM2023-200204.
The specifics are expressed as follows.
where $\omega_{t, 0}$ is the time-varying intercept and $\omega_{t, j}$, for $j=1,...,L$, is the time-varying weight for $j$-th model's prediction as in DBPS.\footnote{ We double use the notation $\omega$, which also exists in (ref), to represent the model weights. Since Equation (ref) is not explicit in this section, we hope readers can understand that these two $\omega$'s are different. } The difference in (ref) is that we treat $\hat{y}_{t,j}$ as data in MSDSP assembly. In this paper, we use out-of-sample predictive mean as $\hat{y}_{t,j}$ while other options such as the median is also feasible. One can also input multiple predictive values from one model to exploit various information content in our framework, because the computation is parallelizable and the sparsity is explicit. The volatility of $y_t$ is set as the standard SV in (ref). The dynamic of each $\omega_{t,i}$, for $i=0,1,...,L$, is set as an independent MSDSP process in (ref)-(ref).
For simplicity, we set $L=7$ base models as the linear economic models in Section (ref). For prediction at horizon $h$, simply replace $\hat{y}_{t, j}$ by $\hat{y}_{t+h, j}$ and $y_t$ by $y_{t+h}$, respectively. As long as the information set used to produce these predictions is up to time $t$, these are out-of-sample prediction. In addition to the MSDSP assembly, we also apply its nested alternative, the DSP, in model assembly, along with the standard BPS framework described in Section (ref), for comparison. Note that no MSDSP or DSP methods are used in the construction of our base models. The ideas applying “MSDSP-in-MSDSP”, that is ensembling MSDSP models with MSDSP weights are left for future research.
Because of the the two-stage procedure for model assembly, we cut the data at two time points: $T_0$ and $T_1$ with $T_0<T_1<T$. The initial training sample uses the information up to time $T_0$, which is the same as in Section (ref). The initial training sample serves to construct the predictions from the base models in this section. The period between $T_0$ and $T_1$ is used to train the synthesis function, and thus the structure of the model weights $\omega_t$ It cannot be inferred without input from the base models, so the two-stage method wins us some computational time, at the cost of a smaller testing sample. We choose $T_0$ and $T_1$ such that $T-T_1=100$ and $T_1-T_0=100$, and keep $T_0$ the same as in Section (ref). This idea of breaking the sample to three subset is analogous to the training, validation and testing sampling scheme in the machine learning literature, but we keep updating the data after each prediction and hence can exploit the information as much as possible.
Table (ref) reports the log predictive likelihood (LPL), Continuous Ranked Probability Score (CRPS), and Root Mean Squared Forecast Error (RMSFE) for horizons $h=1,...,12$ months. These values are based on the last $100$ samples, adjusted by the forecasting horizons. The methods include the MSDSP and DSP assembly, along with the standard DBPS of mcalinn2019dynamic. The RW-SV predictive metrics over the corresponding sample period are also reported for comparison. Note that all base models are the linear economic models with static parameters for both versions of model assembly and DBPS.
Amongst the model assembly method, the MSDSP performs better than the DSP assembly and the standard DPBS in all evaluation metrics at all horizon. Thus, we show here that the MSDSP assembly can easily dominate the state-of-the-art dynamic Bayesian predictive synthesis. Without the need to deploy heavy computation for the distributional input, the MSDSP assembly is scalable. Its Markov switching structure explicate sparsity to improve prediction, which is evidenced by the dominance of MSDSP assembly over the DSP assembly.
The MSDSP assembly is slightly worse than the random walk at short horizons but tends to improve at longer horizons. For example, the point forecast (RMSFE) from MSDSP is better than the random walk when $h>9$, which is empirically plausible (see also mark1995exchange, rapach2002testing, molodtsova2009out). A better model pool might improve the MSDSP assembly, but it is not our focus in this paper.
We propose the novel Markov Switching Dynamic Shrinkage Process method to address dynamic instability, model parsimony and sparsity, and use the proposed framework to address the renowned Meese–Rogoff puzzle, which states that the random walk is highly competitive in exchange rate forecasting. We consider adopting the MSDSP in two settings: directly incorporating the flexible structure into existing economic models; and applying the flexibility in model assembly of standard linear economic models.
In our application, we show that adopting the MSDSP directly to enhance the economic model structure leads to significant improvements in the predictive performance of the economic models relative to the benchmark random walk. The decisive evidence illustrate that the random walk model, even when adjusted to include stochastic volatility, is systematically inferior to the economic models improved with MSDSP priors for both point and density forecasts. This evidence, in itself, provide a new perspective on the Meese-Rogoff puzzle, highlighting the fact that the models resulting from economic theories possess predictive power when flexibility and sparsity is taken into account. Existing results literature have evaluated these models in its most stringent form via the constant coefficient linear model, under the assumption that the respective economic relationship is static, which, in the current economic environment, is unrealistic.
Furthermore, we provide a natural framework for model assembly by using the MSDSP toolkit. With a very primitive linear constant coefficient model pool, the MSDSP assembly consistently beat the state-of-the-art dynamic Bayesian predictive synthesis, while being scalable computationally. The MSDSP model assembly produces similar predictive performance as the benchmark random walk. Compared to the individual model improvement with MSDSP, there is still further improvements that can be investigated with future research, including the inclusion of more sophisticated base model and the incorporation of predictive distribution (rather than point) in the construction of predictive synthesis.