Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
85,371 characters · 19 sections · 39 citation commands
Forecasting Oil Prices Across the Distribution: A Quantile VAR Approach\@thefnmark\@footnotetextWe would like to thank Christiane Baumeister, Jamie L. Cross, Howard Bondell, Sylvia Kaufmann and participants at the 2025 European Seminar on Bayesian Econometrics (ESOBE) in Melbourne, and at the 2025 Milan Time Series Seminars (MiTSS) in Milan for their valuable comments. This paper is part of the research activities at the Centre for Applied Macroeconomics, Commodity Prices (CAMP) at the BI Norwegian Business School. The usual disclaimers apply.
\affil[1]{{ BI Norwegian Business School}} \affil[2]{{ Facultad de Administracion y Economia, Universidad Diego Portales}} \affil[3]{{ Adam Smith Business School, University of Glasgow}} \affil[4]{{ Centre for Applied Macroeconomics and Commodity Prices (CAMP)}}
\thispagestyle{empty} \onehalfspace \setcounter{page}{1}
Oil prices play a central role in shaping global economic activity, yet forecasting them remains exceptionally difficult. The comprehensive survey by AlquistKilianVigfusson2013 establishes that oil futures prices fail to systematically outperform a simple no-change forecast at horizons up to one year, while EllwangerSnudden2023 show that even sophisticated models struggle against a strengthened random walk benchmark\footnote{EllwangerSnudden2023 recommend using end-of-month daily prices rather than monthly averages to construct the random walk benchmark. Since real oil prices require deflation by the monthly CPI, we cannot implement this refinement. Nevertheless, monthly average real prices remain a relevant forecasting target for macroeconomic analysis and policy applications, where oil prices enter inflation forecasts and feed into planning decisions based on average costs over billing or budgeting cycles.}. Crisis periods expose complete forecast breakdown, cf.\ BaumeisterKilian2016: neither futures markets nor econometric models anticipated the 2014--2016 collapse, the April 2020 negative price event, or the 2022 Russia-Ukraine spike. Model combination approaches have improved average forecasts BaumeisterKilianLee2014,BaumeisterKilian2015,AagBjornlandEliassen2024, but offer limited insight into how predictive relationships differ across the distribution of oil price changes and can perform poorly during unusually large price movements. This limitation reflects a deeper challenge: oil price changes exhibit a heavy-tailed and asymmetric distribution, and standard econometric approaches, including Bayesian VARs with stochastic volatility CarrieroClarkMarcellino2016,BaumeisterKorobilisLee2022 or time-varying parameters, summarize uncertainty around a single conditional mean and therefore provide limited insight into the likelihood or nature of tail events.
Figure (ref) illustrates why mean-based approaches may miss economically relevant variation. Financial uncertainty is essentially orthogonal to the level of real Brent prices (left panel), yet systematically associated with time variation in return skewness (right panel). Periods of elevated uncertainty coincide with pronounced shifts in distributional asymmetry, even when the conditional mean exhibits little response.\footnote{This pattern is robust to alternative rolling windows, oil price measures, and transformations, and similar relationships obtain for other predictors such as geopolitical risk and financial spreads.} This suggests that macro-financial predictors may primarily load on the tails of the conditional distribution rather than on its center, motivating a forecasting framework that explicitly models heterogeneity across quantiles.
We address these challenges by developing a Quantile Bayesian Vector Autoregression (QBVAR) that models the entire conditional distribution of real oil price movements, allowing the dynamics and predictive content of economic fundamentals to vary across quantiles. Drawing on the Growth-at-Risk literature AdrianBoyarchenkoGiannone2019, which shows that financial conditions affect the conditional distribution of macroeconomic outcomes asymmetrically, we extend this insight to oil markets, where similar asymmetries may arise from the distinct mechanisms driving price collapses versus spikes. The distributional perspective enables us to identify which variables help predict downside and upside risks, evaluate forecast accuracy specifically in the tails, and assess performance during major episodes of oil market stress.
Our methodological contribution builds on two literatures. First, the Structural VAR tradition that dominates oil market analysis, tracing back to Hamilton1983 and developed through frameworks that treat oil prices as endogenous,\footnote{For earlier SVAR studies separating oil price and macroeconomic shocks, see BurbidgeHarrison1984, GisserGoodwin1986, BernankeGertlerWatson1997, and Bjornland2000.} including Kilian2009 and the Bayesian SVAR of BaumeisterHamilton2019. These methods have transformed our understanding of oil--macroeconomy linkages but focus on conditional means. Second, the Bayesian quantile regression literature, where YuMoyeed2001 established that maximizing the asymmetric Laplace likelihood yields valid quantile inference, and KozumiKobayashi2011 developed the Gibbs sampling algorithms that make Bayesian QVAR computationally feasible. Crucially, Srirametal2013 show that posterior inference concentrates around true quantile coefficients even when the asymmetric Laplace distribution is misspecified -- a “working likelihood” result that justifies our parametric approach when the true error distribution is unknown.
Our approach relates to, but is distinct from, recent work on oil price uncertainty and tail risks. AastveitCrossVanDijk2023 construct time-varying predictive densities from mixtures of conditional means and variances but do not model quantile-specific dynamics explicitly. Baumeisteretal2024 employ Bayesian Additive Regression Trees to generate nonlinear predictive densities and outperform linear quantile regressions, but derive tail forecasts from the predictive distribution rather than modeling quantile-specific dynamics directly. CarrieroClarkMarcellino2024 show that Bayesian VARs with stochastic volatility capture macroeconomic tail risks comparably to quantile regression, but symmetric distributional assumptions preclude predictor effects that vary across quantiles. By contrast, our Bayesian QVAR models the conditional distribution directly through quantile-specific coefficients, shocks, and factor structures within a multivariate framework. In addition, our factor-based specification avoids the variable-ordering dependence inherent in Cholesky-based QVAR implementations ChavleishviliManganelli2024, a concern shown by AriasRubioShin2023 to substantially affect VAR forecasting performance. The Bayesian framework further enables flexible regularization through hierarchical shrinkage priors, which is particularly valuable in the oil forecasting context where VARs with 12 or more lags are common.
Using monthly data from 1975 to 2025, we assess the out-of-sample performance of the QBVAR across horizons and quantiles, showing substantial gains relative to standard Bayesian VARs, Bayesian VARs with stochastic volatility, and the random walk.\footnote{In unreported analyses, we also considered a BVAR with factor stochastic volatility and leverage effects, following KastnerFruhwirthSchnatterLopes2017. This specification generally underperforms the BVAR-SV without leverage across most quantiles and horizons, so we retain the BVAR-SV as our primary stochastic volatility benchmark. Results can be obtained at request.} Three stylized facts emerge. First, the QBVAR improves median forecasts uniformly, achieving 2--5% gains over standard BVARs -- noteworthy given the difficulty of improving upon well-specified models for point prediction. Second, uncertainty and financial condition variables strongly predict downside risk, with left-tail improvements of 10--25% that intensify during crisis episodes, echoing the Growth-at-Risk finding that financial conditions affect the left tail more strongly than the right. Third, right-tail forecasting remains challenging; the QBVAR typically underperforms stochastic volatility models for upside risk, though forecast combinations recover this deficit. These results underscore the importance of modelling distributional dynamics when loss functions are asymmetric: variables with limited value for mean forecasts, such as crack spreads, financial uncertainty, or geopolitical risk, contain substantial information for the tails, with uncertainty and financial stress predicting downside risk through precautionary demand channels (e.g., Bloom2009,BaumeisterKilian2016) and demand-side measures providing some information about upside risk.
This paper makes three contributions. First, we develop a Bayesian Quantile VAR that models the entire conditional distribution of oil price movements through quantile-specific parameters and factor structures, offering an ordering-invariant alternative to existing implementations. Second, we provide new evidence that the determinants of oil price movements differ markedly across quantiles, revealing asymmetric mechanisms hidden in mean-based or density-forecasting models. Third, we show that the QBVAR delivers substantial tail-forecasting gains relative to both random-walk and state-of-the-art BVAR benchmarks, especially during major oil market disruptions. Together, these contributions demonstrate that modelling the conditional distribution, rather than only the mean or an implied density, yields meaningful improvements for risk assessment and policy-relevant forecasting.
The remainder of the paper is organized as follows. Section 2 outlines the Bayesian QVAR methodology. Section 3 describes the data, forecasting design, and out-of-sample results. Section 4 investigates forecast combination strategies and evaluates performance during major oil price episodes. Section 5 concludes.
Standard Bayesian VARs model the conditional mean of economic variables and characterize uncertainty through parametric distributional assumptions. While effective for point forecasting, this approach provides limited insight into tail risks or distributional asymmetries that may be crucial for policy and risk management. We develop a Quantile Vector Autoregression (QVAR) that directly models specific quantiles of the conditional distribution, allowing the dynamics and predictor effects to vary across different points of the distribution.
Our starting point is the Bayesian VAR with factor decomposition in the covariance matrix, as specified by Korobilis2022. For an $n \times 1$ vector of endogenous variables $\bm y_{t}$, this takes the form
where $\bm x_{t} = \left(\bm 1, \bm y_{t-1},..., \bm y_{t-p} \right)'$ contains an intercept and $p$ lags, $\bm \Phi$ is an $n \times (np+1)$ coefficient matrix, $\bm f_{t} \sim N\left( \bm 0, \bm I_r \right)$ is an $r \times 1$ vector of latent factors with $r \ll n$, $\bm \Lambda$ is an $n \times r$ matrix of factor loadings, and $\bm \Sigma = \text{diag}(\sigma_{1}^{2},...,\sigma_{n}^{2})$ is diagonal. This specification implies that cross-equation comovements are captured entirely by the common component $\bm \Lambda \bm f_{t}$, with the VAR covariance matrix given by $\text{cov} \left( \bm y_{t} \vert \bm x_{t} \right) = \bm \Lambda \bm \Lambda' + \bm \Sigma$.
We extend this framework to model specific quantiles of the conditional distribution. For quantile level $q \in (0,1)$, our Quantile BVAR (QBVAR) takes the form
where $\text{AL}(0,\sigma_{i,q},q)$ denotes the asymmetric Laplace distribution with location zero, scale $\sigma_{i,q}$, and asymmetry parameter $q$. Unlike the mean VAR in (ref), the QVAR features quantile-specific coefficients $\bm \Phi_{q}$, loadings $\bm \Lambda_{q}$, and factors $\bm f_{t,q}$, following the quantile factor approach of KorobilisSchroeder2024a. This allows the model to capture different factor structures at different points of the distribution---the factors driving median dynamics may differ substantially from those driving tail behavior. We normalize $\bm f_{t,q} \sim N \left( \bm 0, \bm I_r \right)$ for all $q$.
ChavleishviliManganelli2024 develop a frequentist QVAR that models quantile dynamics through a recursive triangular structure, relying on Cholesky decomposition to orthogonalize the system. While computationally convenient, this approach introduces variable ordering dependence. AriasRubioShin2023 demonstrate that such ordering choices can substantially affect forecasting performance in VARs with stochastic volatility, a concern that extends naturally to quantile settings where the conditional distribution varies across quantiles.
Our factor-based specification avoids this limitation. The common component $\bm \Lambda_{q} \bm f_{t,q}$ captures cross-equation comovements without requiring identification restrictions on the loadings, yielding forecasts that are invariant to variable ordering. This is particularly valuable for forecasting applications where the goal is distributional characterization rather than structural identification.
To facilitate posterior computation, we exploit the location-scale mixture representation of the asymmetric Laplace distribution KozumiKobayashi2011. The QVAR likelihood in (ref) can be equivalently written as
where $\bm z_{t,q} = (z_{1t,q}, \ldots, z_{nt,q})'$ with $z_{it,q} \stackrel{iid}{\sim} \text{Exp}(1)$, and the quantile-specific constants are $\theta(q) = \frac{1-2q}{q(1-q)}$ and $\tau^{2}(q) = \frac{2}{q(1-q)}$. The term $\theta(q) \bm z_{t,q}$ generates asymmetry: it equals zero at the median ($q=0.5$) but shifts the conditional mean for other quantiles, with the direction depending on whether $q$ is below or above 0.5.
This mixture representation renders the likelihood conditionally Gaussian in $(\bm \Phi_q, \bm \Lambda_q, \bm f_{t,q})$ given the auxiliary variables $\bm z_{t,q}$, enabling straightforward Gibbs sampling.
We employ hierarchical shrinkage priors that adapt to the sparsity structure of the data. For the VAR coefficients, we use a horseshoe prior CarvalhoPolsonScott2010:
where $\phi_{ij,q}$ denotes the $(i,j)$th element of $\bm \Phi_q$, $\psi_{ij,q}$ are local shrinkage parameters, $\kappa_q$ is the global shrinkage parameter, and $C^{+}(0,1)$ denotes the half-Cauchy distribution. This prior aggressively shrinks irrelevant coefficients toward zero while allowing important predictors to remain unshrunk.
For the factor loadings, we use independent normal priors: $\lambda_{ij,q} \sim N(0, \underline{V}_\lambda)$. The idiosyncratic scale parameters receive inverse-gamma priors: $\sigma_{i,q} \sim \text{IG}(\underline{a}_\sigma, \underline{b}_\sigma)$. We set $\underline{V}_\lambda = 1$, $\underline{a}_\sigma = 3$, and $\underline{b}_\sigma = 1$ throughout.
Posterior inference proceeds via Markov Chain Monte Carlo. Given the mixture representation (ref), all full conditional distributions are available in closed form. Let $k = np + 1$ denote the number of regressors per equation.
A distinctive feature of quantile regression is that the asymmetric Laplace distribution serves as a computational device for targeting specific quantiles rather than a statement about the true data generating process YuMoyeed2001, Srirametal2013. The posterior concentrates around the true quantile regression coefficients regardless of the actual error distribution, justifying the “working likelihood” interpretation.
For multi-step forecasting, we must decide how to generate predictions beyond the one-step-ahead horizon. Recall from equation (ref) that the ALD mixture representation introduces auxiliary variables $\bm z_{t,q}$ that enter both the conditional mean (through $\theta(q) \bm z_{t,q}$) and the conditional variance (through $\tau^2(q) \text{diag}(\bm z_{t,q})$). Iterating the QVAR forward using this full representation would require simulating future paths of $\bm z_{t+1,q}, \bm z_{t+2,q}, \ldots$, which creates two problems. First, the exponentially-distributed auxiliary variables can generate explosive forecast paths, particularly at extreme quantiles where $\theta(q)$ is large. Second, tracking these auxiliary variables through the forecasting recursion adds computational burden without clear accuracy gains---the quantile-specific information is already encoded in the estimated coefficients $\bm \Phi_q$ and $\bm \Lambda_q$.
We therefore adopt a standard VAR forecasting approach with Gaussian innovations after estimation. This is justified on two grounds. First, the ALD serves as a working likelihood for estimation: the goal is to obtain consistent estimates of the quantile-specific coefficients, not to characterize the true error distribution Srirametal2013. Once $\bm \Phi_q$ and $\bm \Lambda_q$ are estimated, forecasts can be generated using any reasonable innovation distribution. Second, for iterated multi-step forecasts, the accumulated shocks over multiple periods converge toward Gaussianity by the central limit theorem, making the normal approximation increasingly appropriate as the forecast horizon grows.
At each forecast origin $t$, we generate $h$-step-ahead quantile forecasts as follows. For each posterior draw $s = 1, \ldots, S$ of the parameters $\{\bm \Phi_q^{(s)}, \bm \Lambda_q^{(s)}, \bm \Sigma_q^{(s)}\}$:
This procedure integrates over parameter uncertainty while using the estimated quantile-specific dynamics to generate distributional forecasts. By estimating separate QVAR models at multiple quantile levels ($q \in \{0.1, 0.5, 0.9\}$), we trace out the conditional distribution of oil prices, capturing both central tendency and tail risks. Our objective is to characterize the shape of this distribution (its location, spread, and asymmetry) rather than to quantify uncertainty around individual quantile estimates.
We evaluate quantile forecasts using the quantile score (also known as the pinball loss or check function), which is the proper scoring rule for quantile forecasts GneitingRaftery2007. For quantile level $q$ and forecast error $u_t = y_{t+h} - \hat{Q}_t(q,h)$, the quantile score is
This loss function penalizes errors asymmetrically: at $q = 0.1$, overpredictions (predicting values too high) are penalized nine times more heavily than underpredictions, reflecting the greater concern about missing downside risks when forecasting the left tail. Lower quantile scores indicate more accurate forecasts.
Our empirical analysis employs monthly data spanning from January 1975 to February 2025, encompassing 602 observations and providing one of the longest time series available for oil price forecasting research. We construct two primary oil price series, both expressed in real terms using the U.S. Consumer Price Index: the Refiner Acquisition Cost (RAC) of imported crude oil from the Energy Information Administration's Monthly Energy Review, and Brent crude oil spot prices extended backward using RAC growth rates to ensure consistency with historical price movements. Both series undergo logarithmic transformation to achieve stationarity and facilitate interpretation of results in terms of percentage changes. The analysis below presents results for RAC forecasts; comparable results using Brent can be found in the online supplement.
The forecasting models incorporate a comprehensive set of predictor variables capturing different dimensions of oil market fundamentals and broader economic conditions. These include global economic activity measures, supply-side indicators, inventory and storage data, consumption patterns, geopolitical risk indices, financial market conditions, product market spreads, and uncertainty measures. All variables undergo appropriate transformations to ensure stationarity, with detailed descriptions, sources, and transformation codes provided in the Data Appendix. This rich set of predictors allows us to evaluate which economic fundamentals are most informative for quantile-specific oil price forecasts.
The baseline specification is a three-variable VAR system comprising oil prices, a measure of global oil market fundamentals (Global Liquid Fuels Consumption), and one additional predictor variable selected from the comprehensive set described in Table (ref). This design allows systematic evaluation of individual predictor contributions to distributional forecasts. Motivated by the distributional-risk perspective of the Growth-at-Risk literature AdrianBoyarchenkoGiannone2019, and drawing on recent oil market research for predictor selection Baumeisteretal2024, we focus on twelve representative predictors spanning distinct economic categories:
For the QBVAR, we consider lag lengths $p \in \{1, 2, 4, 12\}$. Our baseline results use $p = 4$, which balances flexibility with parameter parsimony. This is an important consideration for quantile regression, which requires sufficient observations to estimate quantile-specific coefficients reliably. The three-variable QBVAR(12) specification involves substantially more parameters, and while results remain reasonable, more parsimonious specifications have an edge on average. Results for alternative lag lengths are available in the Online Supplement. All Bayesian VAR models (QBVAR, BVAR, and BVAR-SV) employ a factor structure in the error covariance matrix with $r=1$ common factor throughout all exercises.\footnote{In exploratory analyses not reported here, we found that specifications with $r=2$ factors yield qualitatively similar results.}
We implement a recursive forecasting scheme where models are re-estimated at each forecast origin with an expanding estimation window. We generate forecasts at horizons $h = 1, 2, \ldots, 12$ months ahead for log growth rates of oil prices. To assess robustness and performance across different economic environments, we report results for two evaluation windows: 2008M1--2025M2 (205 monthly forecast origins, encompassing the financial crisis, commodity super-cycle collapse, and COVID-19 pandemic) and 2013M1--2025M2 (145 monthly forecast origins, focusing on the post-crisis period including the 2014--2016 oil price collapse). Results for additional evaluation windows are available in the Online Supplement.
We compare the QBVAR against three benchmark models representing the current state-of-the-art in oil price forecasting. First, BVAR(12), a three-variable Bayesian VAR that includes the same endogenous variables as the corresponding QBVAR specification (oil prices, Global Liquid Fuels Consumption, and the same additional predictor). To ensure comparability, both BVAR and BVAR-SV use the same factor structure in the error covariance matrix and the same horseshoe shrinkage prior on the VAR coefficients as the QBVAR. The only difference is that the BVAR targets the conditional mean with constant volatility, while the BVAR-SV incorporates stochastic volatility. Second, a no-change random walk (RW), the challenging univariate benchmark that has proven difficult to beat AlquistKilianVigfusson2013. Third, BVAR-SV(12), a three-variable specification (again with identical endogenous variables) that incorporates stochastic volatility and represents the state-of-the-art for density forecasting CarrieroClarkMarcellino2024. For BVAR and BVAR-SV benchmarks, we extract quantile forecasts from the posterior predictive distribution; for the QBVAR, we report the median across posterior draws from the quantile-specific model.
Performance is assessed using quantile score (QS) ratios, with values below 1.00 indicating QBVAR outperformance. We report results separately for QS10 (left tail), QS50 (median), and QS90 (right tail) to evaluate performance across the distribution. For all Bayesian models, posterior inference is based on Markov Chain Monte Carlo simulation. At each forecast origin, we run the Gibbs sampler for 3,000 iterations, discarding the first 1,000 as burn-in. From the remaining 2,000 post-burn-in draws, we retain every 5th draw to reduce autocorrelation, yielding 400 saved posterior draws per forecast step.
The Online Supplement provides comprehensive robustness analysis for all results presented in this paper. Specifically, it reports quantile score ratios for: (i) both oil price measures (U.S.\ Refiners' Acquisition Cost and Brent); (ii) alternative lag specifications for the QBVAR ($p \in \{1, 2, 4, 12\}$); (iii) four evaluation windows spanning different market regimes (2018M1--2025M2, 2013M1--2025M2, 2008M1--2025M2, and 2000M1--2025M2); and (iv) the complete set of twelve predictors, grouped by economic channel---uncertainty and financial conditions (JLN, JLNF12, EBP, GZ spread), global activity and supply (GECON, Kilian Index, OECD stocks, World Rig Count), and geopolitical risk (GPRH, GPRHT, GPRHA, Crack spread). These results are reported against all three benchmark models: BVAR(12), the random walk, and BVAR-SV(12). The supplement also includes forecast combination results examining optimal weighting between QBVAR and BVAR-SV forecasts.
Table (ref) reports out-of-sample Quantile Score (QS) ratios from our QBVAR(4) model relative to the BVAR(12) benchmark, using four representative predictors that capture distinct economic channels: macroeconomic uncertainty (JLN), geopolitical risk (GPRHT), financial conditions (GZ spread), and global real activity (GECON).\footnote{The Online Supplement reports results for all twelve predictors, four different estimation samples, and QBVAR specifications with lag orders $p\in \{1,2,4,12\}$; the conclusions are qualitatively unchanged.} To confirm robustness across time periods, we report results for two distinct evaluation windows.
Two key features of these results merit discussion. First, we observe consistent improvements in the median forecast (QS50) across all predictors, with ratios below unity in nearly every case. Gains are modest but remarkably stable: the GECON Index delivers 4--5% improvements at medium horizons, while other predictors yield 2--5% gains that persist across both evaluation periods.
Second, the results reveal striking asymmetries in tail forecasting that vary systematically by predictor type. The uncertainty measure JLN delivers substantial left-tail improvements---up to 15% at horizon 6---while showing no gains (and occasional deterioration) in the right tail. Conversely, the geopolitical risk index GPRHT excels at right-tail forecasting, with improvements of 14--16% at horizons 4--6, but performs poorly in the left tail. The GZ spread and GECON exhibit more balanced performance at medium horizons, with gains concentrated around the median and one tail, though left-tail performance weakens at longer horizons, particularly for GECON. These asymmetries suggest that different economic fundamentals are informative for different parts of the oil price distribution: uncertainty measures help predict downside risk, while geopolitical indicators capture upside risk.
Beating a simple random walk (RW) remains one of the most difficult challenges in forecasting commodity returns. Table (ref) presents Quantile Score ratios for the QBVAR(4) relative to this benchmark, using four predictors that illustrate distinct forecasting channels: macroeconomic uncertainty (JLN), global economic conditions (GECON), financial stress (GZ spread), and oil market fundamentals (OECD stocks).\footnote{The Online Supplement reports results for all twelve predictors, four different estimation samples, and QBVAR specifications with lag orders $p\in \{1,2,4,12\}$; the conclusions are qualitatively unchanged.}
Three findings stand out. First, the QBVAR achieves meaningful improvements in median forecasts (QS50) at short horizons, with gains of 6--9% for $h \leq 2$, but these advantages dissipate at longer horizons where ratios converge to unity. This pattern of significant short-run predictability that fades over time is consistent with the well-documented difficulty of beating random walks in commodity markets.
Second, the left tail (QS10) reveals the QBVAR's main value proposition. The uncertainty measure JLN delivers remarkable improvements across all horizons and both evaluation periods, with gains ranging from 12% to 24%. Financial conditions captured by the GZ spread show similarly consistent left-tail gains of 9--18%. These results suggest that uncertainty and credit stress provide substantial advance warning of oil price declines---precisely the downside risk that matters most for risk management.
Third, right-tail forecasting (QS90) proves more challenging but shows horizon-dependent gains. At horizons 2--5, real activity (GECON) and inventory measures (OECD stocks) deliver 3--8% improvements, though these evaporate at longer horizons. The contrast with left-tail results is striking: while uncertainty predicts downside risk persistently, upside risk remains harder to forecast beyond the near term.
Taken together, these results highlight a fundamental asymmetry in oil price predictability. Measures of uncertainty and financial stress provide advance signals of elevated downside risk, even though adverse price realizations themselves tend to occur abruptly. By contrast, upside risks appear to be driven by demand- or inventory-related developments that are harder to anticipate beyond short horizons. From a forecasting perspective, these findings underscore the value of explicitly modeling the conditional distribution: while mean forecasts offer limited gains over a random walk, quantile-based approaches deliver economically meaningful improvements precisely where risk management and policy concerns are most acute.
Stochastic volatility is known to enhance forecast accuracy, particularly for tail predictions. Table (ref) presents Quantile Score ratios for the QBVAR(4) relative to a BVAR(12) with stochastic volatility (BVAR-SV), using four predictors that span uncertainty (JLN), financial conditions (GZ spread, EBP), and global activity (GECON).\footnote{The Online Supplement reports results for all twelve predictors, four different estimation samples, and QBVAR specifications with lag orders $p\in \{1,2,4,12\}$; the conclusions are qualitatively unchanged.} Since our quantile-based approach inherently captures heteroskedasticity through quantile-specific coefficients, we incorporate stochastic volatility only in the benchmark.
Three findings emerge. First, the QBVAR achieves consistent, if modest, improvements in median forecasts (QS50), with gains of 1--4% that persist across horizons and evaluation periods. The consistency is notable: for JLN and GZ spread, every QS50 entry falls below unity.
Second, left-tail forecasting (QS10) reveals the QBVAR's comparative advantage. The uncertainty measure JLN delivers gains of 3--24%, with improvements in 15 of 16 entries. Financial conditions indicators perform similarly: GZ spread shows gains of up to 21%, and EBP up to 14%. These results demonstrate that quantile-specific modeling captures downside risk dynamics that even flexible stochastic volatility specifications miss.
Third, the QBVAR offers no improvement for right-tail forecasting (QS90). Every entry exceeds unity, often substantially (1.10--1.50). This asymmetry has a clear interpretation: oil price collapses build through observable financial stress and uncertainty that the QBVAR exploits, while price spikes arrive as less predictable supply disruptions that stochastic volatility handles better through its flexible variance dynamics.
For practitioners, these results suggest that quantile-based and volatility-based approaches are complementary rather than competing: stochastic volatility excels at capturing unanticipated right-tail price movements, while the QBVAR is particularly effective at identifying persistent downside risks relevant for risk management.
The results in Section 3 reveal a striking pattern: the QBVAR excels at left-tail forecasting but struggles with the right tail, particularly against the BVAR-SV benchmark. This asymmetry suggests that the two approaches capture fundamentally different aspects of oil price dynamics---the QBVAR exploits observable predictors of downside risk, while stochastic volatility models better accommodate the sudden variance expansions characteristic of price spikes. A natural question arises: can we combine these complementary strengths to achieve improvements across the entire distribution?
This section investigates forecast combination strategies that blend QBVAR and BVAR-SV predictions, then evaluates whether these gains hold during the major oil price events that matter most for risk management.
We investigate whether convex combinations of forecasts from QBVAR(4) and BVAR-SV(12), one of the most competitive benchmarks, can enhance predictive performance.\footnote{The online supplement considers different lag-length specifications and sample sizes.} To this end, we evaluate three distinct combination strategies.
First, we employ a fixed-weight scheme, defined as:
where $\tau \in \{0.1, 0.5, 0.9\}$ denotes the quantile, $\hat{q}^{Comb}_t(\tau,h)$ is the combined quantile forecast, $\hat{q}^{QBV}_t(\tau,h)$ represents the QBVAR forecast, $\hat{q}^{BVSV}_t(\tau,h)$ denotes the BVAR-SV(12) forecast, and $h$ is the forecasting horizon. Here, $\lambda$ is the fixed weight assigned to the QBVAR forecast, unchanged across each forecasting step, predictor, and horizon.
Second, we adopt a dynamic weighting strategy based on historical performance. At each forecasting step, for each horizon and predictor, we compute the weight assigned to QBVAR, $\lambda_t(\tau,h)$, using the most recent $S$ rolling observations:
where $\lambda_t(\tau,h)$ is the weight for QBVAR at time $t$, quantile $\tau$, and horizon $h$. The right-hand side represents the ratio of quantile scores over the past $S$ observations, assigning greater weight to the model with lower (better) quantile scores, thus prioritizing forecasts with stronger recent performance. For this exercise we set $S=50$ months.
Finally, we implement a strategy that optimizes weights at each forecasting step based on recent performance. Using the latest $S$ observations, we determine the optimal weight for each period, applied to the forecast combination. The weight $\lambda_t(\tau,h)$ varies by quantile $\tau$ and horizon $h$, computed as:
where $\lambda^*_t(\tau,h)$ minimizes the average quantile score over the prior $S$ observations, optimizing the forecast combination for each step. For this exercise, we set $S=75$ months.\footnote{Results using different values for $S$ for both methods are available upon request.}
To understand when each model contributes most, Figures (ref) and (ref) display Quantile Score ratios for the combined forecast relative to BVAR-SV across a grid of combination weights $\lambda \in [0,1]$, where $\lambda = 1$ corresponds to pure QBVAR and $\lambda = 0$ to pure BVAR-SV. The combined forecast follows equation ((ref)) using JLNF12 as the predictor.
For right-tail forecasts (QS90), Figure (ref) shows that intermediate weights perform best. At horizons $h > 4$, the optimal combination lies near $\lambda = 0.5$, suggesting QBVAR contains useful right-tail information that is revealed when combined with BVAR-SV. Notably, $\lambda = 0.4$ improves performance at all horizons, while $\lambda = 0.5$ does so at all but one.
For left-tail forecasts (QS10), Figure (ref) tells a different story. At 7 of 12 horizons, the optimal weight is $\lambda = 1.0$, meaning QBVAR fully encompasses BVAR-SV for downside risk prediction. Where the optimum falls below unity, it consistently exceeds 0.5, reinforcing the QBVAR's dominance in left-tail forecasting.
These patterns motivate our combination strategies: conservative weights ($\lambda \approx 0.4$--$0.5$) should improve right-tail forecasts while preserving most left-tail gains. Importantly, the weight profiles confirm that QBVAR is the primary source of improvements in downside risk forecasting, whereas BVAR-SV contributes mainly to right-tail performance at longer horizons. More broadly, the relative usefulness of QBVAR and BVAR-SV is quantile- and horizon-dependent, reinforcing the case for forecast combinations.
The preceding results highlight a clear complementarity: QBVAR delivers strong left-tail performance, while BVAR-SV contributes mainly to right-tail accuracy at longer horizons. Forecast combinations therefore provide a natural way to exploit these strengths simultaneously. We first examine whether simple averaging rules, applied uniformly across horizons, predictors, and time periods, can outperform BVAR-SV across quantiles. Table (ref) reports Quantile Score ratios for fixed weights $\lambda \in \{0.4, 0.5, 0.6\}$ using two representative predictors: JLNF12 (which showed strong QBVAR performance) and JLN (which showed mixed right-tail results in Section 3).\footnote{Results for additional predictors, sample sizes, and lag-length specifications can be found in the online supplement with similar patterns.}
The results confirm that naive averaging succeeds where pure QBVAR failed. With JLNF12, the conservative weight $\lambda = 0.4$ achieves improvements in all 36 entries across quantiles and horizons. On average, this strategy improves left-tail forecasts by 9%, median forecasts by 1%, and, crucially, right-tail forecasts by 4%. Even more aggressive weights ($\lambda = 0.5, 0.6$) deliver gains in nearly all configurations, with only 1--2 exceptions.
The JLN predictor reveals the limits of fixed weights. While left-tail improvements remain substantial (12--16% on average), right-tail performance becomes mixed at $\lambda \geq 0.5$, with deterioration at longer horizons. This heterogeneity motivates adaptive weighting schemes.
Taken together, the combination results point to a clear asymmetry. QBVAR is the dominant source of forecast improvements for downside risk, while stochastic volatility contributes mainly to right-tail performance at longer horizons. Simple fixed-weight combinations preserve most of the substantial left-tail gains delivered by QBVAR while improving right-tail forecasts, without requiring real-time estimation of time-varying weights. These findings highlight the value of forecast combinations that exploit the complementary strengths of quantile-based and volatility-based approaches.
While fixed-weight combinations already deliver robust improvements, allowing combination weights to adapt to recent forecasting performance may yield additional gains. Hence, rather than fixing weights ex ante, we allow $\lambda_t(\tau,h)$ to vary at each forecasting step based on recent performance, following equation ((ref)).
Table (ref) reports results for four representative predictors. Results indicate that performance-based weighting delivers consistent gains across predictors and horizons. With JLNF12, improvements span all quantiles: 9% for QS10, 1% for QS50, and 4% for QS90 on average. The GPRHA predictor achieves similar success with gains across 35 of 36 entries. Even predictors that showed mixed pure-QBVAR performance (JLN, GZ spread) now deliver left-tail improvements of 9--14% alongside modest median gains, though right-tail results remain mixed for GZ spread at longer horizons.
These gains are consistent with performance-based schemes endogenously assigning greater weight to QBVAR during periods of heightened downside risk, when tail dynamics dominate forecast accuracy. Relative to fixed-weight combinations, the benefits of adaptation are therefore concentrated in the left tail, while median and right-tail improvements remain more modest.
Overall, allowing combination weights to adapt to recent forecasting performance yields incremental improvements over fixed-weight schemes, while preserving, and in some cases strengthening, the dominant role of QBVAR in forecasting downside risk.
While performance-based weighting adapts to recent relative accuracy, it still constrains weights to follow a fixed updating rule. We therefore consider a fully flexible approach that directly optimizes the combination weight at each forecasting step. Specifically, this strategy minimizes the quantile score over recent periods, as defined in equation ((ref)). Table (ref) reports results for four predictors spanning different economic channels.
The results highlight the benefits and limits of fully optimized combination weights. This strategy maximizes left-tail gains: JLN achieves average improvements of 21%, while GZ spread and JLNF12 deliver 19% and 18%, respectively. Median forecasts show modest but consistent gains of 0--2%. Right-tail performance is more nuanced: JLNF12 improves in 9 of 12 horizons (2% average gain), while predictors such as JLN exhibit deterioration at longer horizons, suggesting that aggressive optimization can overfit right-tail dynamics. The GECON index achieves more balanced improvements, with average gains of 8% for QS10, 1% for the median, and no gains for QS90.
A critical question is whether the QBVAR delivers reliable performance during episodes of large oil price movements, since accurate tail forecasting is particularly valuable in such environments. To examine this issue, we evaluate the left-tail forecasting performance, measured by QS10, of the QBVAR(4) and competing benchmarks during three major oil price declines: the Global Financial Crisis from 2008 to 2009, the oil price collapse of 2014 to 2015, and the COVID-19 outbreak in 2020. We also assess right-tail forecasting performance, measured by QS90, during three major oil price surges: the Gulf War period from 1990 to 1991, the commodity super-cycle of 2007 to 2008, and the post-COVID surge associated with the Russian invasion of Ukraine from 2021 to 2022.
Figure (ref) illustrates the magnitude of oil price movements during these episodes. In particular, the figure shows that major oil price events are associated with large and rapid movements, often concentrated in the tails of the distribution. This motivates evaluating forecast performance conditional on such episodes, where tail accuracy is most economically relevant.
Table (ref) summarizes the key economic and geopolitical mechanisms behind each episode, distinguishing between negative and positive oil price shocks. This classification guides the interpretation of tail-specific forecast performance.
\paragraph{Price Decline Episodes.}
Table (ref) reports QS10 ratios relative to the RW benchmark during price declines, comparing QBVAR against BVAR(12) and BVAR-SV(12). Three patterns emerge. First, RW consistently underperforms. Most entries fall below unity across all models, confirming that structured approaches add value during crises. Second, the QBVAR frequently achieves the largest gains, with improvements reaching 70% during the 2014--16 collapse (JLNF12 at $h=6$). Third, performance varies by event: all models excel during COVID-19 and the 2014--16 collapse, but results for the 2008 Financial Crisis are more mixed, with BVAR-SV often competitive at short horizons.
\paragraph{Price Surge Episodes.}
Table (ref) reports QS90 ratios during price surges. The patterns differ markedly from declines. First, RW is more competitive---many entries exceed unity, particularly during the 2021--22 crisis. Second, the QBVAR nonetheless achieves substantial gains in specific configurations: during the 2007--08 commodity boom, improvements reach 30% at medium horizons for several predictors. Third, the Gulf War episode shows the strongest QBVAR performance, with gains of 10--50% at horizons 9--12 across most predictors. The 2021--22 crisis proves most challenging, with QBVAR underperforming at short horizons but recovering at $h \geq 6$.
Taken together, the event-based analysis confirms and sharpens the full-sample findings. During major oil price declines, the QBVAR consistently delivers the largest improvements in left-tail forecasting accuracy, particularly at medium and longer horizons, highlighting the importance of quantile-specific dynamics when financial stress and uncertainty intensify. By contrast, performance during price surges is more heterogeneous, with QBVAR gains concentrated in specific episodes such as the Gulf War and the commodity super-cycle, and weaker results during the post-COVID period. These results underscore a fundamental asymmetry in oil price dynamics: downside risks tend to build gradually through observable macro-financial channels, while upside risks are more episodic and driven by hard-to-predict supply disruptions.
This paper develops a Bayesian Quantile Vector Autoregression that models the entire conditional distribution of oil prices through quantile-specific coefficients and factor structures. Our out-of-sample evaluation spanning 1975--2025 yields three main findings. First, the QBVAR improves median forecasts by 2--5% relative to standard Bayesian VARs, demonstrating that quantile-specific dynamics benefit even point prediction. Second, uncertainty and financial condition variables strongly predict downside risk, with left-tail improvements of 10--25% that are particularly pronounced during crisis episodes. Third, right-tail forecasting remains challenging: the QBVAR typically underperforms stochastic volatility models for the 90th percentile, suggesting oil price collapses build through observable financial stress while spikes arrive as less predictable supply disruptions.
These findings have practical implications for risk assessment. Traditional models focusing on expected price movements may systematically underestimate downside risks during periods of rising financial stress or macroeconomic uncertainty. Our results show that financial conditions and uncertainty measures substantially improve left-tail predictions, providing earlier warning of potential price collapses. This is valuable for investors managing portfolio risk, firms hedging commodity exposure, and policymakers concerned with macroeconomic stability. The persistent difficulty in forecasting right-tail events suggests that upside risks may require hybrid approaches combining quantile-specific coefficients with stochastic volatility, or alternative predictors designed to capture supply disruption risks.