Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
78,829 characters · 20 sections · 97 citation commands
Let the Tree Decide: FABART A Non-Parametric Factor Model
\noindentKeywords: Bayesian FAVAR, Asymmetry, Regression Tree models, Oil supply News, Non-parametric techniques \noindentJEL Codes: C11, C32, C38, Q43\\
\onehalfspacing \setcounter{footnote}{0}
Recent episodes of heightened volatility—driven by geopolitical uncertainty and the COVID-19 pandemic—have underscored the importance of capturing nonlinearities in economic time series models. In such periods, standard linear methods often fail to deliver reliable forecasts or accurate assessments of macroeconomic dynamics. Principal Component Analysis (PCA) and traditional linear factor models, widely used for dimensionality reduction in large datasets, impose restrictive linearity assumptions on the relationship between observed variables and latent factors. These limitations can weaken the predictive performance of factor-augmented frameworks, particularly in the presence of threshold effects, regime shifts, or extreme realizations.
To address these shortcomings, a growing literature has developed nonlinear dimension reduction techniques that are better equipped to uncover complex structures and enhance predictive accuracy. Early contributions by BaiNg2008 introduced quadratic principal components and targeted predictors, showing that even simple forms of nonlinearity can yield substantial forecasting gains. More recently, PelgerXiong2022 and hauzenberger2023real highlight that allowing for state-dependent or nonlinear mappings between observables and latent factors can significantly improve model performance, particularly during recessionary episodes and periods of economic stress.
Building on these insights, this article proposes a novel Factor-Augmented Vector Autoregressive (FAVAR) model that integrates Bayesian Additive Regression Trees (BART) into the factor extraction stage—the FABART model. By introducing a nonparametric mapping between observed variables and latent factors, FABART captures complex, nonlinear relationships that linear methods may miss, while maintaining the interpretability of factor-augmented frameworks.
A simulation experiment that compares the model performance with a linear FAVAR model and Monte Carlo simulations demonstrate that FABART accurately recovers the relationship between latent factors and observables in the presence of nonlinearities without imposing spurious nonlinearities when the true data-generating process is linear. This robustness highlights its adaptability to different underlying structures and its ability to avoid overfitting.
The empirical application first assesses the forecasting performance of the FABART model relative to established linear benchmarks, with particular attention to periods of heightened volatility, such as the COVID-19 pandemic. Full-sample results show that FABART delivers forecast accuracy that is comparable to linear alternatives for real oil prices, but substantially improves forecasts for industrial production, particularly during periods of economic stress. For financial variables, forecast performance is more mixed: FABART delivers better density forecasts for the excess bond premium during periods of heightened volatility, but performs comparably to the linear BVAR in the case of the S&P 500 index. Overall, the results suggest that forecasting performance deteriorated for both linear benchmarks—the BVAR and the linear FAVAR model—during and after the COVID-19 shock, whereas FABART remained broadly stable. These findings are consistent with recent evidence highlighting the advantages of flexible nonlinear models in capturing instability during periods of hightened turbulence (e.g., huber2020nowcasting, hauzenberger2023real, baumeister2024risky), and underscore the relative strength of FABART in delivering more stable and reliable forecasts under these conditions.
In addition, this article investigates the transmission of oil supply news shocks by exploiting the flexible structure of FABART combined with Generalized Impulse Response Functions (GIRFs), which allow for a nonparametric assessment of nonlinearities. The results reveal pronounced sign asymmetries at the aggregate level: positive oil supply shocks induce stronger and more persistent contractions in real activity and inflation than the expansions triggered by negative shocks, consistent with the evidence documented by hamilton2003oil and balke2002oil. A similar pattern emerges for employment at the U.S. state level, with negative shocks generating weaker improvements in employment compared to the sharper contractions observed after positive shocks.\footnote{Moreover, the cross-sectional analysis highlights substantial heterogeneity in the magnitude of labor market responses: states with larger manufacturing sectors experience deeper employment contractions, while mining-intensive states are relatively less affected. These findings align with recent studies emphasizing the role of sectoral composition in amplifying the effects of macroeconomic shocks across regions (e.g., mumtaz2018state).}
The remainder of this article is structured as follows: Section (ref) introduces the empirical model, estimation approach, and identification strategy. Section (ref) presents the Monte Carlo results. Section (ref) discusses the empirical findings. Section (ref) concludes.
The factor-augmented vector autoregressive (FAVAR) model, introduced by Bernanke2005, offers a reduced-dimensional vector autoregression (VAR) framework that integrates information from a large set of variables, addressing omitted-variable concerns by extracting latent factors from high-dimensional datasets. This method allows latent factors to co-evolve with observed variables. Extending this approach, I incorporate Bayesian additive regression trees (BART) to flexibly sample latent factors within a nonparametric framework, capturing complex, nonlinear dependencies while preserving interpretability. In line with other structural factor model frameworks, the Factor Bayesian Additive Tree (FABART) model leverages rank reduction techniques to expand the information set, facilitating the identification of structural disturbances without compromising degrees of freedom (korobilis2013assessing).
In the following, the empirical framework is introduced through the factor-augmented VAR model by Bernanke2005 and extended to sample the factor loadings via Bayesian additive regression additive trees.
In the absence of non-parametric factor loadings the linear FAVAR set-up is determined by the following measurement equation
where $X_{i,t}$ with i=$1,...,N$ denotes a large panel of N observable variables and $Y_t = \{Z_t,F^1_t, \dots, F^j_t\}$ is a matrix that contains $J=1,...,j$ latent factors that capture the shared movement patterns of the underlying series at each point in time. $Z_t$ is treated as an observed factor i.e. with factor loading equal to unity. In the application outlined in this paper the observed variable $Z_t$ is real oil prices.\footnote{In prominent frameworks that identify monetary policy shocks using FAVAR models (e.g., Bernanke2005; baumeister2013changes), the Federal Funds rate is typically treated as an observed factor. This assumption reflects its informational availability to both the econometrician and the monetary authority.} The relationship between the large granular dataset $X_{i,t}$ and the latent factors in $F_t$ is established by the factor loadings $\Gamma_i$ and $\epsilon_{i,t}$ are the i.i.d. errors with $E(\epsilon_{i,t}'\epsilon_{i,t})=\Sigma$, where $\Sigma$ is a diagonal matrix.
The transition equation ((ref)) with $M$ endogenous variables contained in matrix $Y_t$ of size $MxT$ and lag length $L=12$ is given by
The objective of the FABART method is to examine the non-linearities in the estimation of the unobserved factors, specifically in the measurement equation ((ref)). In line with dolado2020quantile, hauzenberger2023real, and korobilis2024monitoring, I aim to explore the non-parametric relationship between the factor and its loadings. Introducing such a flexible structure in the measurement equation enables mapping the high-dimensional dataset $X_{i,t}$ without imposing a restrictive linear relationship between the factors and the observables. This flexibility is essential for effectively reducing the dimensionality of large data panels while capturing complex dependencies in the data.
Several approaches have been proposed to introduce non-linearities in the relationship between latent factors and observables in $X_{i,t}$. Some methods achieve this by allowing for time-varying factor loadings (del2008dynamic), while others introduce time variation in the factor dynamics (e.g., mumtaz2012evolving; korobilis2013assessing). My approach builds on this literature by explicitly modeling the non-parametric link between the factor and its loadings, offering a more flexible and data-driven framework for capturing structural changes in the underlying data relationships. Specifically, I postulate that $X_{t}=(X_{1}',..X_{N}')$ evolves according to a general multivariate model of the form
where $F(Y_t)= (f_1(Y_t),... f_N(Y_t)')$ is a N-dimensional vector of non-linear functions estimated non parametrically via Bayesian additive regression trees (BART).
Each $f_i(Y_t)$ is approximated as
with $g_{is}(Y_t| \tau_{is},\mu_{iS})$ being a single regression function, $f_i$ denoting the sum of the regression trees and $\tau_{is}$ are the tree structures related to the $i$th element in $X_{t}$. $\mu_{iS}$ are tree specific terminal nodes and $S$ measures the total number of trees.
As illustrated in huber2020nowcasting, the space of each explanatory variable is recursively partitioned by every regression tree using binary splitting rules. Each rule involves selecting one explanatory variable and a threshold \( c \), determining whether an observation lies above or below that threshold. Based on these rules, observations are allocated to terminal nodes, which capture the fitted values conditional on the associated partition. These splitting rules can be written as \( \{ Y \in A_r \} \) if \( Y \) lies within the partition set \( A_r \) for \( r = 1, \dots, b \), and \( \{ Y \notin A_r \} \) otherwise.
Let \( \mathbf{Y}_{\cdot m} \) denote the \( m \)-th column of \( Y \), containing the latent factors. A binary split is defined as:
For each of the \( N \) variables in \( X\), the fitted values from a regression tree are obtained via the following step function:
where \( \mathbbm{1} \{ \cdot \} \) denotes the indicator function, which takes the value 1 if \( Y \) falls within the partition region \( A_r \) defined by the sequence of splitting rules encoded in the tree structure. Each component of the regression tree is treated as an unknown parameter and estimated from the data. The parameter set includes: the terminal node values \( \mu \), the number of terminal nodes \( b \), and all elements defining the tree structure—namely, the splitting variable \( \mathbf{Y}_{\cdot m} \) and the threshold \( c \) that governs the binary split.
The prior distributions introduced by chipman2010bart are fundamental in mitigating the risk of overfitting. The joint prior for the parameters governing the \(S \) trees at horizon \( h \) is decomposed as:
where the conditional prior follows:
The structure of the tree \( T_s^{(h)} \) depends on the likelihood that a given node at depth \( d = 0,1,2, \dots \) is an internal node rather than a terminal one. This probability is determined as $\alpha(1 + d)^{-\beta}$, with \( \alpha \in (0,1) \) and \( \beta > 0 \). Higher values of \( \beta \) and lower values of \( \alpha \) strengthen the prior belief that the tree structure remains simple (i.e., shorter). As recommended by chipman2010bart, I set \( \alpha = 0.95 \) and \( \beta = 2 \). The threshold parameter \( c \) follows a uniform distribution over the observed range of the variables. By default, the selection of splitting variables is assumed to be uniform across regressors.
To define \( p(\mu_s^{(h)} | T_s^{(h)}) \), chipman2010bart propose rescaling the dependent variable such that its values fall within the interval \([-0.5, 0.5]\). Consequently, the function \( m_h(x_t, z_t, w_{t+h}^{(h)}) \) is expected to remain within this range. The prior distribution of \( p(\mu_s^{(h)} | T_s^{(h)}) \) follows a normal distribution \( N(0,S) \), where the variance \( S \) is defined as $ S = \frac{1}{2\kappa(0.5)^2}$, as recommended by chipman2010bart \( \kappa = 2 \) . Given this prior, the probability that the conditional mean of the dependent variable lies between \([-0.5,0.5]\) is approximately 95%.
For the variance term \( \sigma_{j} \), I adopt a conjugate inverse \( \chi^2 \) prior. The hyperparameters of this prior are specified using an estimated variance \( \hat{\sigma}_{j} \) derived from a linear regression. As noted in mumtaz2022impulse, when the model exhibits non-linearity, this estimate is typically biased upwards. In line with common practice, the hyperparameters are chosen to satisfy \( P(\sigma_j < \hat{\sigma}_j) = 0.75 \) and $\nu_j = T/2$ . The number of trees \( S \) is fixed and set to 250, e.g., see chipman2010bart or huber2020nowcasting).
The Bayesian "sum-of-trees" framework developed by chipman2010bart enables posterior sampling through an iterative Bayesian backfitting MCMC algorithm. This approach regularizes each tree using a prior, ensuring that individual trees act as weak learners.\footnote{Although conceptually similar, BART differs technically from the gradient boosting method of friedman2001greedy by incorporating a prior to constrain the trees and applying Bayesian backfitting across a fixed number of trees.}
The sampling of each tree structure \( T_{js} \), with index \( j = 1, \dots, N \) denoting the different equations of the model, follows the Bayesian backfitting algorithm introduced by chipman2010bart. The posterior distribution of the trees is computed by integrating out the terminal node parameters \( \mu_{js} \), and is given by:
Here, the partial residual vector \( R_{js} \) is defined as:
The likelihood term \( p(R_{js} \mid \tau_{js}, \sigma_j) \) facilitates efficient estimation by marginalizing over \( \mu_{js} \).
Tree structures are sampled using the Metropolis-Hastings (MH) algorithm, as outlined in chipman1998bayesian. A candidate tree \( \tau_{js}^* \) is proposed from a distribution \( q(\tau_{js}^i, \tau_{js}^*) \) based on the current tree state \( \tau_{js}^i \). The acceptance probability for \( \tau_{js}^{i+1} \) is given by:
The proposal distribution \( q(\tau_{js}^i, \tau_{js}^*) \) incorporates four types of moves:
By integrating out \( \mu_{js} \), this approach maintains computational efficiency while ensuring a stable parameter space for posterior inference.
To fully implement the MCMC algorithm the latent states conditional distribution of the BART parameters need to be drawn. This is akin to the strategy employed by huber2022inference, who extend the Bayesian Additive Vector Autoregression Trees (BAVART) model by integrating latent state sampling to accommodate mixed-frequency data. Specifically, leveraging insights from the literature on effect size estimation in black-box models (see crawford2019variable), a linear approximation to $F(Y_t)$ is constructed. This approximation enables the efficient estimation of latent states using conditionally Gaussian methods.
This linear approximation was initially introduced in the context of Gaussian kernel (GK) and Gaussian process (GP) regressions. However, its applicability extends far beyond these methods and is relevant to a broad spectrum of machine learning and econometric techniques, see ish2019interpreting. The primary requirement is the estimation of a nonlinear function $F = (F(Y_1), \dots, F(Y_T))$ over $T$ observations.
To illustrate the concept of linear approximation, consider a standard linear regression model where the effect size (or regression coefficient) quantifies the extent to which $Y$, containing both latent factors and observables, projects onto the granular dataset $X_{i}$.
This projection is defined as: \( \hat{A_i} = \text{Proj}(Y, X_{i}) = Y^{\dagger} X_{i}\), where $Y^\dagger$, a $K \times T$ matrix, represents the Moore-Penrose inverse. If $Y$ is of full rank, the projection simplifies to $(Y'Y)^{-1}Y'X_{i}$, making the effect size equivalent to the least squares estimate, which captures the marginal impact of explanatory variables on dependent variables.
Building on crawford2018bayesian and crawford2019variable, this approach is extended by projecting $X_i$ onto the matrix $F$, yielding the following effect size estimate:
\( \tilde{A_i} = \text{Proj}(F, X_i) = F^\dagger X_i\).
The core intuition behind this approximation is that regressing $X_i$ onto $F$ provides insight into how much of its variance is accounted for by $F$. This offers a clear interpretation of the relationship between each factor $F_i, (i = 1, \dots, K)$ and the corresponding $X_i$. Without additional regularization, the projection naturally implies that $F\tilde{A_i} \approx X_i$.
$\tilde{A_i}$ allows to produce a linear approximation to the non-parametric multivariate model and obtain the elements of $\Gamma$ Factor loadings matrix for each $N$-variable from the measurement equation ((ref)), such that
In matrix notation this would correspond to
where $\Lambda$ and $\Psi$ are proxied by $\tilde{A_i}$, representing the elements of the factor loading matrix $\Gamma$, which has dimensions $(N+1) \times (J+1)$. The structure of this matrix allows certain variables to exhibit a contemporaneous relationship with the observable variable $Z_t$, as indicated by the non-zero elements of $\Psi$.
Based on this approximation to a linear model with Gaussian shocks the latent factors are drawn based on the carter1994gibbs algorithm as in mumtaz2010evolving.
The algorithm for sampling from the full conditional posterior distribution is detailed below. The MCMC algorithm sequentially draws samples from the full conditional posterior distributions described in the preceding subsections. Given a set of initial values, the algorithm follows these steps in each iteration\footnote{Steps 2., 3. and 6. are implemented using the R package dbarts.}:
The algorithm generates a total of 30,000 draws, with the initial 15,000 discarded as burn-in.
Given the nonlinear nature of the model, impulse responses are obtained using the Generalized Impulse Response Function (GIRF) framework proposed by koop1996impulse. The GIRF at horizon $k$ is defined as:
where $\Psi_t$ denotes the set of model parameters, and $e_j$ represents the structural shock. The long-run mean of the data, denoted as $\overline{Y}$, is given by:
By using the long-run mean instead of the historical values of the series, the influence of specific data point selections should be reduced. Since the model is nonlinear, the impulse responses are computed via Monte Carlo simulation, accounting for history-dependent effects and potential asymmetries in shock transmission.
The covariance structure of the reduced-form residuals, denoted as $\boldsymbol{\Sigma}$, is specified as:
where $\mathbf{A}$ represents a lower triangular matrix and $\mathbf{Q}$ is an orthogonal matrix satisfying $\mathbf{Q}'\mathbf{Q} = \mathbf{I}$. The structural shocks, $\boldsymbol{\varepsilon}_t$, are then obtained as:
where $\mathbf{u}_t$ represents the reduced-form errors. Following kanzig2021macroeconomic, the identification of an oil supply news shock is established by assuming it corresponds to the first element of $\boldsymbol{\varepsilon}_t$, i.e., $\varepsilon_{1t}$. This is achieved through the use of an instrumental variable, $m_t$, satisfying:
with the orthogonality condition $E(m_t \varepsilon_{jt}) = 0$ for all $j \neq 1$.
In the empirical application, the instrumental variable $m_t$ follows kanzig2021macroeconomic and is constructed based on variations in oil futures prices around OPEC announcements.
The covariance structure of the reduced-form residuals, denoted as $\boldsymbol{\Sigma}$, is specified as:
where $\mathbf{A}$ represents a lower triangular matrix and $\mathbf{Q}$ is an orthogonal matrix satisfying $\mathbf{Q}'\mathbf{Q} = \mathbf{I}$. The structural shocks, $\boldsymbol{\varepsilon}_t$, are then obtained as:
where $\mathbf{u}_t$ represents the reduced-form errors. Following kanzig2021macroeconomic, the identification of an oil supply news shock is established by assuming it corresponds to the first element of $\boldsymbol{\varepsilon}_t$, i.e., $\varepsilon_{1t}$. This is achieved through the use of an instrumental variable, $m_t$, satisfying:
with the orthogonality condition $E(m_t \varepsilon_{jt}) = 0$ for all $j \neq 1$. The strength of the instrument is summarized by the reliability ratio:
This is carried out using a frequentist approach. Importantly, equation (ref) is not introduced as an additional equation in the model. Instead, for each draw of the transition equation parameters, the first column of $\mathbf{A}$ is computed by projecting $m_t$ on the reduced-form residuals using the standard frequentist method, as in miranda2023identification.
In the empirical application, the instrumental variable $m_t$ follows kanzig2021macroeconomic and is constructed based on variations in oil futures prices around OPEC announcements.
The primary objective of this section is to assess the FABART model’s ability to capture nonlinear dynamics in economic data. To this end, I develop a simulation experiment that generates data reflecting nonlinear features commonly observed in macroeconomic time series. The central question is whether FABART can successfully recover these nonlinear relationships without imposing restrictive parametric assumptions.
Model performance is evaluated within a factor model framework by comparing outcomes across three data-generating processes (DGPs): one linear and two nonlinear.
The factor evolves according to an autoregressive process of order three:
where \( e_t \) is normally distributed with variance \(\sigma^2 \). The coefficient vector \( \beta \) is parameterized as \( \beta = (0.6, -0.3, 0.2) \). The setup follows a standard factor model in which the $N$=20 observed variables \( X_t \) is driven by $J$=1 latent factor \( F^{J}_t \). The baseline (linear) specification is given by:
where \( X_t \) is a \( 20 \times T \) vector of observed variables, \( B \) is a \(20 \times 1 \) factor loading matrix, and \( V_t \) follows a normal distribution with mean zero and diagonal covariance matrix \( R \) (\( 20 \times 20 \)). The elements of \(B\) are drawn independently from a uniform distribution over the interval \([-0.9, 0.9]\). The diagonal elements of \(R\), representing the idiosyncratic variances of each series, are independently sampled from a uniform distribution.
To introduce nonlinearities, Equation (ref) is extended in two ways:
To evaluate model performance in recovering the latent factor, a recursive one-step-ahead forecasting exercise is conducted across all synthetic datasets. The factor process $F_t$ is held constant across data-generating processes to isolate differences in model behavior, allowing for a direct comparison between FABART, a linear FAVAR, and a Random Walk (RW) benchmark. Table (ref) reports the Root Mean Squared Errors (RMSEs) for each method, normalized relative to the RW benchmark, which is scaled to 1. Across all specifications, the FABART model delivers the lowest forecast errors.
Figure (ref) displays the one-step-ahead forecast paths for the latent factor $F_t$ and a representative observable $X_t$, comparing predictions from the FABART and linear FAVAR models. Visually, the FABART forecasts align more closely with the true series, particularly under nonlinear data-generating processes, highlighting the model’s flexibility in capturing complex dynamics.
A Monte Carlo experiment based on 100 replications of each data-generating process (DGP)—in which the latent factor is held fixed while the observable variables \(X_t\) vary across replications—provides a deeper assessment of the properties of the FABART model. The results of this exercise are presented in the Annex. Figure (ref) displays the distribution of estimation errors for the latent factor \(F_t\), obtained using the FABART methodology. The first column illustrates the posterior estimates of \(F_t\) across replications, while the second column shows the corresponding estimation errors, defined as the difference between each posterior draw and the true factor. The figure highlights that, on average, errors are centered around zero across all DGPs and time periods. Estimation is most accurate in the linear case, with slightly larger errors observed under the nonlinear DGPs. This is reflected in high average correlations between the estimated and true factors: 0.982 for the linear DGP, 0.968 for the quadratic nonlinear DGP, and 0.966 for the hyperbolic tangent case. These findings confirm that the proposed methodology accurately recovers the latent factor across varying data-generating environments.
The analysis in this section applies the FABART methodology to evaluate its forecasting performance across a range of macro-financial variables—including oil prices, industrial production, and financial market indicators—and to examine the transmission of oil price shocks to the U.S. economy. The results are presented in two parts.
The first part assesses the model’s forecasting accuracy relative to established benchmarks, with particular attention to periods of heightened volatility, such as the COVID-19 pandemic, to evaluate its ability to capture complex macroeconomic dynamics.
The second part investigates nonlinearities in the transmission of oil price shocks, providing empirical evidence on how the sign of such shocks affects key macroeconomic outcomes in the U.S. economy on aggregate and on the differences at the state level dimension.
The baseline specification exploits the potential of rank-reduction techniques by incorporating a set of 187 macroeconomic and financial time series. This approach enables an in-depth analysis of oil price dynamics, leveraging variables emphasized in recent contributions to the literature—such as baumeister2015forecasting, kanzig2021macroeconomic, and baumeister2024risky. It also expands the information set relevant to key determinants of oil price dynamics (Table (ref)) by including the full FRED-MD monthly database of macroeconomic indicators proposed by mccracken2016fred. A detailed list of the FRED-MD variables and transformation codes is provided in Table (ref) located in the Apppendix. To facilitate a state-level analysis of the employment response to oil price shocks, the dataset further includes total nonfarm employment for all U.S. states.
This broad information set facilitates a richer understanding of the structural and cyclical drivers of oil prices by capturing a wide range of economic and financial conditions. At the same time, the use of dimension reduction techniques ensures that the resulting framework remains computationally tractable, even in high-dimensional settings, while still allowing for the analysis of the dynamics and interactions of the broader set of included macroeconomic and financial series.
A key variable of interest in this study is the real price of WTI crude oil, a widely used global benchmark for crude and refined petroleum products. To obtain a consistent measure of real oil price dynamics, Brent prices are deflated using the U.S. Consumer Price Index (CPI).
The selection of predictor variables reflects their theoretical and empirical relevance in capturing the fundamental drivers of oil prices, including supply conditions, demand pressures, and broader macroeconomic influences. This setup is motivated by the extensive literature on oil price modeling, which emphasizes the interplay between structural and cyclical components.
For instance, baumeister2015forecasting highlight the importance of including oil production, inventory levels, and indicators of economic activity to adequately capture both supply- and demand-side shocks. Their findings underscore the role of global economic momentum and inventory dynamics in shaping oil price fluctuations, particularly in capturing nonlinear and episodic responses to macroeconomic and geopolitical developments.
On the supply side, global oil production and petroleum inventories serve as key indicators of physical market fundamentals, reflecting adjustments in output and stockpiles that buffer short-run imbalances. On the demand side, global industrial production and composite indicators of economic activity capture fluctuations in energy consumption over the business cycle. Together, these variables provide a comprehensive basis for modeling the dynamics of real oil prices across regimes and time horizons.
Table (ref) provides an overview of the relevant oil market drivers considered in the analysis.
In selecting the dataset for this analysis, I follow baumeister2015forecasting, kanzig2021macroeconomic, miescu2024cardiff, and baumeister2024risky. The partial dependence analysis by baumeister2024risky highlights key nonlinear relationships influencing oil price predictions. While moderate changes in predictors like global fuel consumption and petroleum inventories have limited effects, extreme variations—such as sharp contractions in the GECON indicator, steep declines in rig counts, or significant increases in the gasoline-Brent spread—induce disproportionate and asymmetric impacts on oil prices.
All variables are measured at a monthly frequency and are transformed into growth rates by taking the first difference of their natural logarithms. Exceptions to this transformation include the GECON indicator, which is inherently stationary due to its construction. The transformation of real oil prices differs across empirical applications: for the forecasting exercise, the series is included in levels, consistent with the approach in baumeister2024risky, while the structural analysis employs the log difference of real oil prices, in line with the specification used by miescu2024cardiff.
To benchmark the performance of the FABART model, a linear BVAR is estimated using a selected subset of macroeconomic and oil market variables. These include the real WTI oil price, world oil production, world oil inventories, global industrial production, U.S. industrial production, and the U.S. Consumer Price Index (CPI). A detailed description of these variables is provided in Table (ref) in the Annex. In addition, a linear FAVAR model is estimated using the same variable set as the FABART model, allowing for a direct comparison between the nonlinear and linear factor-augmented approaches.
To investigate cross-state heterogeneity in the transmission of oil price shocks, the analysis incorporates state-level nonfarm employment data. These series, sourced from the U.S. Bureau of Labor Statistics and accessed via the Federal Reserve Bank of St. Louis FRED database, provide a consistent measure of labor market conditions across all U.S. states. Each series is first seasonally adjusted using standard procedures based on the X-13ARIMA-SEATS methodology, and then transformed by computing the logarithm of the first difference to ensure stationarity and focus on cyclical dynamics.
In addition, to assess how differences in economic structure influence the transmission of oil price shocks, the analysis follows mumtaz2018state and draws on data describing the industrial composition of each state. Specifically, the dataset includes annual state-level GDP by industry from 1963 to 2013, averaged over time to mitigate short-run volatility and emphasize long-run structural characteristics. These data are provided by the U.S. Bureau of Economic Analysis (BEA), with industry classifications based on the Standard Industrial Classification (SIC) system through 1996 and the North American Industry Classification System (NAICS) from 1997 onward.
This subsection evaluates the forecasting performance of the FABART model relative to two linear benchmarks: a Bayesian Vector Autoregression (BVAR) estimated on a selected subset of macroeconomic and oil market variables, and a Factor-Augmented VAR (FAVAR) that uses the same variables and factor dimension as the FABART model. Forecast accuracy is assessed using both root mean squared errors (RMSEs) and predictive densities estimated via kernel methods.
Evaluation of predictive densities via kernel methods allows for the capture of non-normal features such as asymmetry and fat tails, which tend to become more pronounced at forecast horizons beyond one month (see geweke2010comparing). This evaluation criterion is selected because, while the posterior predictive densities generated by linear models—such as Bayesian Vector Autoregressions (BVARs) and Factor-Augmented VARs (FAVARs)—are Gaussian by construction, the FABART model accommodates richer distributional characteristics. This distinction becomes especially relevant in periods of heightened uncertainty, where there is evidence of deviations from normality in economic time series (see alessandri2017financial).
Formally, let $M = M_0$ denote the benchmark model, the FABART model introduced in this study. Two alternative models are considered for comparison, denoted as $M_1$ and $M_2$: (1) a fixed-coefficient Bayesian VAR, which assumes time-invariant relationships among variables; and (2) a linear FAVAR model, which incorporates latent factor structures while maintaining Gaussian density assumptions. For both the linear FAVAR and the FABART, seven factors are considered ($J = 7$), which explains approximately 47% of the total variance in the dataset. This choice is consistent with the findings of stock2005implications, who estimate seven dynamic factors using a comparable macroeconomic dataset, and mccracken2016fred, who report similar explanatory power in their FRED-MD-based factor model. For all three specifications, the sample spans 1974:M01–2024:06 with a lag length of $L = 12$. The final observation used for estimation is 2023:06, with forecast evaluation conducted over the period 2023:07–2024:06. As is standard for US data, the overall prior tightness is set to $\iota = 0.1$. As in banbura2010large, the tightness of the sum of coefficients prior is set as $\lambda = 10\iota$, and a flat prior is imposed on the constant term (see, e.g., alessandri2017financial).
For each model $M \in \{M_0, M_1, M_2\}$, the log predictive likelihood is computed as:
where $\log p(Z_{t+h} | Z_t)$ denotes the predictive likelihood of a variable of interest, $h$ is the forecast horizon, and $t = S+1, \dots, T-h$ defines the forecast evaluation window, with $S < T$.
I evaluate forecast accuracy across one-, three-, and twelve-month-ahead horizons using root mean squared errors (RMSE) and predictive log scores (LS). Table (ref) reports the results for the full sample.
Forecast accuracy varies significantly across models, variables, and forecast horizons. For real oil prices, the three models deliver broadly similar predictions. In contrast, the performance gap is more pronounced for industrial production: the FABART model consistently outperforms the linear benchmarks, achieving lower RMSEs and better log scores than both the BVAR and FAVAR across all horizons. In terms of log scores, the deterioration in forecast densities is most acute for the FAVAR at the three-month horizon, reflecting a substantial worsening in density forecasts. In contrast, FABART maintains more contained log scores, indicating a more stable forecast performance.
Turning to financial variables, the forecast performance becomes more mixed. For the excess bond premium (EBP), the BVAR achieves the lowest RMSEs at one and three months, but its log scores do not consistently outperform those of the FABART model. FAVAR performs weakest in both metrics. For the S&P 500 index, the linear BVAR displays the strongest forecasting performance at the 1-month horizon. While results are more comparable to those of the FABART model at longer horizons, the BVAR achieves slightly better log scores throughout, particularly in the short and medium term. The FAVAR model performs notably worse across both metrics, especially in terms of density forecasting accuracy.
Taken together, the full-sample results underscore the relative strength of the FABART model, particularly in terms of forecast accuracy and density fit for macroeconomic variables. Nevertheless, the BVAR model retains some advantage in short-horizon forecasts for financial series.
To assess how the FABART model performed relative to linear benchmark models over time, Figures (ref)--(ref) display the RMSE and cumulative log scores for 1-month-ahead forecasts of selected macroeconomic and financial variables over the 2014:01–2024:06 period. The top panels show the RMSEs, while the bottom panels report the cumulative log scores. Following standard practice, I compute the log score as the logarithm of the predictive density evaluated at the realized value. For comparability across models and series, I report the cumulative sum of these log scores in absolute value.\footnote{Log scores are based on the logarithm of the predictive density evaluated at the realized outcome. Reporting their cumulative sum in absolute value facilitates interpretation: lower values indicate that the model consistently assigned higher probability to outcomes that actually occurred, reflecting better density forecast performance. See geweke2010comparing or alessandri2017financial.} In this convention, lower values signal superior predictive performance, as they correspond to higher likelihood assigned to the realized outcomes. Shaded areas correspond to the observed series.
The first column of Figure (ref) presents the results for real oil prices. Forecast performance is broadly comparable across the three models—both in terms of RMSE and log score accumulation—though the FABART model (red line) performs slightly better, as indicated by a lower cumulative log score.
The second column displays the results for industrial production. Across all models, RMSEs increase sharply during the COVID-19 period and its immediate aftermath. This deterioration in forecast performance coincides with the extreme year-on-year fluctuations visible in the shaded series. As also reported in Table (ref), average RMSE performance is broadly similar across models. However, FABART exhibits notably greater stability in density forecasts, as evidenced by a smoother and less volatile accumulation of log scores. This suggests that FABART delivers more robust predictive distributions, particularly during periods of heightened volatility.
Figure (ref) displays the forecast evaluation across time for two selected financial variables. The first column presents the results for the excess bond premium. The linear BVAR displays smaller RMSEs, but it also exhibits a notable decline in density forecast performance following the COVID period, consistent with the increased volatility in the series. By contrast, the FABART model maintains a more stable log score trajectory, underscoring its robustness in capturing predictive densities during periods of financial stress.
The second column shows the results for the S&P 500 index. Forecast accuracy, measured by the RMSE, deteriorates sharply in early 2020 across all models. The cumulative log scores indicate that the FAVAR model performs notably worse, while the linear BVAR and FABART models exhibit similar and more stable density forecast performance.
This relative advantage of FABART aligns with recent literature emphasizing the limitations of linear models during extreme events. As emphasized by huber2020nowcasting, the COVID-19 shock generated data far outside historical ranges, rendering traditional mixed-frequency VARs unreliable and motivating the adoption of nonlinear approaches. In this spirit, hauzenberger2023real show that nonlinear dimension reduction techniques—such as autoencoders and kernel-based principal components—markedly improve forecast accuracy for inflation during turbulent periods. To examine how these modeling differences play out during episodes of heightened uncertainty, Figure (ref) zooms in on the critical months of Spring 2020, when both industrial production and oil markets experienced unprecedented disruptions. Industrial production saw extreme outliers, including the largest monthly change in the post-1974 period, while oil markets were hit by the combined shock of the COVID outbreak and the collapse of OPEC+ negotiations. As noted by baumeister2024risky, traditional models failed to anticipate the severity of the oil price collapse. The figure provides a granular view of how each model’s forecast densities responded in quasi real time.
Figure (ref) (LHS) illustrates the one-month-ahead predictive densities for real oil prices across the different models during the onset of the COVID-19 shock. All models gradually incorporated the steep decline in oil prices, with observed values (green dots) consistently falling within the forecast bands. Over time, the posterior densities of all specifications shifted into negative territory, reflecting the sharp deterioration in oil market conditions. Among the three models, FABART exhibited a faster adjustment to the downturn, particularly between February and March 2020. Its nonlinear structure allowed for greater responsiveness to the incoming data, leading to forecast distributions that re-centered more accurately around the observed outcomes. In contrast, the linear BVAR and FAVAR models adapted more slowly, resulting in delayed recognition of the oil price collapse.
The one-month-ahead predictive densities for industrial production are presented in Figure (ref) (RHS). The comparative advantage of the FABART model becomes particularly evident during the early phase of the COVID-19 shock. In the forecast made in February 2020—targeting March outcomes—FABART already shows signs of a downward adjustment. This adaptability becomes even more pronounced in the March forecast (for April 2020), where the linear models entirely fail to anticipate the sharp contraction in industrial production. In contrast, FABART captures the deterioration in real activity more accurately, with its predictive density re-centering closer to the observed outcome.
A substantial literature has established that oil price increases have more pronounced macroeconomic effects than decreases, also referred to as sign asymmetry. hamilton2003oil shows that output declines following oil price hikes are not mirrored by equivalent expansions after price declines. His results suggest that increases in oil prices—particularly those that reverse earlier declines—contain significant predictive content for GDP, whereas decreases do not. This asymmetry has been linked to mechanisms such as uncertainty, adjustment costs, and irreversible investment. Moreover, the analysis in balke2002oil indicates that neither monetary policy nor standard price dynamics fully explain the asymmetric responses of output to oil shocks. Their findings instead highlight the role of real-side frictions—such as sectoral reallocation and capital adjustment—as key drivers of the observed nonlinearity.
To assess whether such asymmetries are present within the FABART framework, this section investigates potential nonlinearities in the transmission of oil price shocks to both the U.S. and global economy, with particular attention to the sign of the shock.\footnote{The analysis uses pre-pandemic aggregate and state-level data, as detailed in Section (ref). The monthly sample spans from 1974:01 to 2016:12, aligning with recent studies on oil supply news shocks and their macroeconomic effects (e.g., kanzig2021macroeconomic,miescu2024cardiff,caravello2024disentangling,forni2023asymmetric,mumtaz2024distribution). As noted by kanzig2021macroeconomic, the external instrument is available only from 1984:04 onward, therefore the estimation of the $\mathbf{A}$ matrix is limited to that subsample.} The nonlinear nature of the model allows to capture the possibility that positive and negative shocks propagate through distinct channels and with differing intensity.
In line with kanzig2021macroeconomic, the size of the shock is calibrated to generate a 10% increase in the real price of oil in response to an oil supply news shock. Figure (ref) displays the median responses to this shock. Overall, the results are broadly consistent with the findings of kanzig2021macroeconomic in terms of both the sign and magnitude of the responses. In particular, oil supply shocks lead to higher inflation and a contraction in real economic activity, consistent with standard transmission mechanisms, as also shown in mumtaz2024distribution, who employ a FAVAR model with an external instrument to study oil price shocks. Moreover, the relevance of oil shocks for short-term inflation dynamics is supported by jacquinot2009assessment, who, using a DSGE framework calibrated to the euro area, find that oil price changes are a key driver of short-run inflation fluctuations. Consistent with the findings of kanzig2021macroeconomic, the shock also leads to a significant increase in global oil inventories on impact, followed by sluggish growth. This pattern reflects forward-looking behavior by market participants anticipating future scarcity or higher oil prices.
Figure (ref) displays the impulse responses to positive and negative oil price shocks, with the latter mirrored for comparability.\footnote{For clarity, the Appendix includes a separate figure showing the original responses to negative oil price shocks.} The results reveal clear sign asymmetries: positive shocks induce a persistent increase in inflation, which is notably stronger than the disinflationary response to negative shocks. However, their effects on real activity unfold more gradually—U.S. and global industrial production decline significantly only after approximately 20 months. By contrast, negative oil price shocks trigger an immediate and significant decline in industrial production, with effects lasting up to one year. This pattern challenges the conventional view that oil price reductions are expansionary. As emphasized by hamilton2003oil, such nonlinearities may stem from asymmetries in adjustment costs, heightened uncertainty, and the irreversibility of investment in energy-intensive durable goods. While oil price increases can disrupt production plans and erode consumer confidence, declines often fail to stimulate activity symmetrically due to frictions in reallocating resources and expectations that lower prices may not persist.
In terms of other key oil market drivers considered in the analysis, a decline in oil prices leads to a modest fall in OECD oil production, significant up to the 10-month horizon, after which the response becomes centered around zero and remains insignificant until the end of the forecast horizon. In contrast, following a positive shock, oil production initially shows no significant response, but a contraction materializes after about 20 months.
Recent literature offers complementary and contrasting interpretations of such asymmetries. forni2023asymmetric, for instance, emphasize the role of uncertainty in shaping the macroeconomic effects of oil supply news shocks. In their framework, uncertainty rises in response to both positive and negative shocks, but its macroeconomic consequences are asymmetric: it amplifies the contractionary impact of oil price increases while dampening the expansionary effects of oil price declines. This helps to explain the muted real activity responses to negative shocks in their findings. By contrast, the results presented here point to significant real effects following both positive and negative shocks. As this paper does not impose an explicit uncertainty mechanism, alternative transmission channels—such as demand reallocation, precautionary behavior, or financial frictions—may play a more dominant role in shaping the dynamics of industrial production, particularly in response to negative price shocks.
In line with this perspective, the theoretical framework developed by miescu2024cardiff embeds uncertainty and labor market frictions into a nonlinear DSGE model with forward-looking agents and endogenous job separation. Rather than distinguishing between positive and negative shocks, their model focuses on shock size and shows that larger oil supply news shocks lead to disproportionately stronger effects on real activity, employment risk, and financial variables such as risk premia. In particular, larger shocks increase the likelihood of job loss and precautionary savings, which in turn amplify economic downturns. While the framework successfully captures inflation persistence and uncertainty-driven downturns, it does not account for the sign-dependent contractionary effects of negative shocks observed in this paper. This contrast reinforces the idea that the macroeconomic impact of oil shocks depends not only on their size or direction, but also on how risk and uncertainty are propagated through different economic frictions.
A further comparison with caravello2024disentangling underscores the importance of distinguishing between sign and size nonlinearities. Their methodological contribution shows that local projection models augmented with even or odd functions of shocks can disentangle these two forms of asymmetry. Applying their framework to oil supply news shocks, they find evidence of size nonlinearities but not sign-dependent responses. In contrast, the results presented here suggest a clear sign asymmetry, both in the timing and magnitude of inflation and output responses. This divergence suggests that within the FABART framework, future work should further examine how shock size, uncertainty, and sign interact to drive macroeconomic asymmetries in the response to oil price movements—particularly in light of evidence by arce2025economic showing that heightened oil supply uncertainty, even in the absence of actual supply cuts, can generate inflationary pressures and sector-specific output declines that resemble the effects typically associated with negative oil supply shocks.
While the preceding analysis has focused on the U.S. aggregate effects of oil price shocks on a range of macro-financial variables, the following analysis shifts to the U.S. federal state level and examines employment responses. The model underlying the state-level analysis is identical to the one previously presented, as the number of variables included in the observation equation remains unchanged.
To examine whether positive and negative shocks induce asymmetric employment contractions across U.S. states, Figure (ref) plots the two-year cumulative employment responses. The response to positive shocks is sign-adjusted and placed on the x-axis, while the response to negative shocks is plotted on the y-axis. The red dashed 45-degree line represents the benchmark of symmetry: states lying on this line would experience equal employment losses from both types of shocks. The majority of states fall below the line, suggesting that employment contractions tend to be larger following positive oil price shocks than following negative ones. This asymmetry is consistent with the findings of hamilton2003oil and balke2002oil, who emphasize the role of real-side frictions—such as adjustment costs, investment irreversibility, and uncertainty—in amplifying the effects of oil price increases relative to declines.
To formally assess sign asymmetry, I compare the two-year cumulative responses to positive and negative shocks across the posterior distribution. The results, presented in Appendix Figure (ref), indicate that the posterior mass lies predominantly to the left of zero, suggesting that positive shocks induce significantly larger employment contractions than negative shocks. This Bayesian evidence is corroborated by a one-sample t-test on the cross-sectional distribution of state-level median cumulative responses, reported in Appendix Table (ref), which rejects the null hypothesis of a zero mean difference at the 1% significance level.
As an additional exercise, I assess whether structural characteristics help shape the magnitude of employment responses to oil price shocks across states. Specifically, I examine whether variation in industrial composition—particularly the relative importance of manufacturing and mining sectors—can account for part of the observed cross-sectional heterogeneity. The results, presented in the Appendix (Table (ref)) , are based on a cross-sectional regression that relates the cumulative employment responses after 2 years to sectoral shares and suggest that sectoral specialization may condition the intensity of oil price shock transmission: states with larger manufacturing sectors tend to experience deeper employment contractions, while those with greater mining exposure appear more insulated. These findings are consistent with mumtaz2018state, who document a similar role for industrial structure in shaping state-level income responses to aggregate uncertainty shocks.
Taken together, the results point to two key patterns in state-level employment responses to oil price shocks: (i) sign asymmetries, with positive shocks triggering larger contractions than negative ones, and (ii) cross-sectional heterogeneity in the magnitude of these responses, with industrial composition playing an important role. Nevertheless, the mechanisms driving these asymmetries merit further investigation. Future research could examine how labor market frictions, regional demand conditions, or distributional channels contribute to these heterogeneous effects. Recent evidence indicates that oil shocks can have unequal consequences across income groups, suggesting an additional transmission channel via regional differences in household exposure (e.g., berisha2021income, mumtaz2024distribution). Building on the discussion in the previous subsection, further work on the interaction between shock size, sign, and uncertainty in the transmission of oil shocks could shed light on how these dimensions jointly shape regional labor market dynamics (e.g., arce2025economic,forni2023asymmetric).
This article develops the FABART model, a novel framework that integrates Bayesian Additive Regression Trees (BART) into a Factor-Augmented Vector Autoregressive (FAVAR) structure to forecast U.S. macro-financial variables and analyze asymmetries in the transmission of oil supply news shocks. By combining flexible nonparametric methods with latent factor structures, FABART effectively captures complex, nonlinear relationships between observables and latent factors, particularly during periods of economic and geopolitical instability.
A simulation experiment comparing FABART to linear alternatives and a Monte Carlo simulation demonstrate that the model accurately recovers the relationship between latent factors and observables in the presence of nonlinearities, while remaining consistent when the true data-generating process is linear. This robustness highlights FABART’s adaptability to different underlying structures and its ability to avoid overfitting.
The empirical application evaluates the forecasting performance of FABART relative to standard linear benchmarks across a range of U.S. macro-financial variables. Results show that FABART substantially improves forecast accuracy for industrial production, while delivering comparable performance for real oil prices and mixed results for financial variables. Importantly, FABART maintains stable forecasting performance during periods of heightened volatility, such as the COVID-19 pandemic, when the accuracy of linear models tends to deteriorate.
The article also uncovers pronounced sign asymmetries in the transmission of oil supply news shocks: positive shocks induce stronger and more persistent contractions in real activity and inflation than the expansions triggered by negative shocks. These patterns are consistent with mechanisms such as adjustment costs, investment irreversibility, and heightened uncertainty that amplify the macroeconomic effects of positive oil shocks. Extending the analysis to the U.S. state level, the results reveal similar sign asymmetries in employment responses, with positive shocks generating sharper employment contractions than the improvements following negative shocks.
The FABART framework captures key asymmetries and nonlinear propagation channels in the transmission of oil supply news shocks. However, understanding the mechanisms driving the emergence of asymmetric responses remains an important avenue for future research. Further work could explore how adjustment costs, investment irreversibility, and heightened uncertainty amplify the contractionary effects of positive oil price shocks relative to the muted expansions following negative shocks. At the state level, a more granular analysis of sectoral specialization, regional demand dynamics, and distributional channels may help explain the observed cross-sectional heterogeneity in employment responses. Finally, extending the framework to examine the interaction between uncertainty, shock size, and sign dependence would provide additional insights into the nonlinear transmission of oil shocks across both aggregate and regional outcomes, building on the recent advances discussed in this article.