Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
53,009 characters · 14 sections · 45 citation commands
Forecasting With Factor-Augmented Quantile Autoregressions: A Model Averaging Approach
\setcounter{page}{1}
Model averaging is an alternative to model selection which has seen an increase in popularity over the recent years. In a model selection approach, the researcher attempts to find a single best model for a given purpose, while ignoring all the information in other models. Under the model averaging approach, however, the researcher combines information from all competing models by obtaining a weighted average of the competing models' estimators. This can be seen as an insurance against selecting a very poor model, as it allows the researcher to diversify and account for model uncertainty, thus enabling an improvement in out-of-sample performance. This is because, as argued by inter alia hendry_pooling_2004, wallis_combining_2005, even a simple combination of forecasts with equal weights can never produce a forecast that is worse than the worst individual forecast. Frequentist model averaging (FMA), despite having a shorter history in economics than Bayesian Model Averaging (BMA), has started to receive a lot of attention in the past decade or so, as, in contrast with the BMA, it requires no priors and the corresponding estimators are totally determined by the data moral-benito_model_2015 .\\
At the same time, the recent availability of large datasets has generated interest in models with many possible predictors. Factor-augmented regressions, in particular, have been proven to forecast a particular series relatively well compared to the original predictors, with the benefit of significant dimension reduction stock_forecasting_2002. However, factors are usually determined and ordered by their importance in driving the co-variability of many predictors, which may not necessarily be consistent with their forecast power over a particular series of interest. Model specification is therefore necessary in order to determine which factors should be included in the forecast regression, in addition to specifying the number of lags of all independent variables. \\
In this paper, we consider different model averaging methodologies for the combination of different quantile autoregressive (QAR) models, in an attempt to best forecast the quantiles of inflation and growth. kapetanios_forecast_2008 provide compelling reasons for using model averaging for the purpose of forecasting, in order to address the problem of model uncertainty. Combinations of forecasts often outperform individual forecasts, as models may be incomplete in different ways, thus averaging can offset different biases present in each model granger_can_1996, granger_spurious_1974. Nevertheless, this prior empirical work on forecast combinations of inflation focuses on point estimates for the conditional mean of inflation kapetanios_forecast_2008 and such point forecasts ignore the risks and uncertainty around the central forecast. In the case of growth, for example, such forecasts may paint an overly optimistic picture of the state of the economy adrian_vulnerable_2019. Distribution or quantile forecasts, on the other hand, provide a more complete picture of the conditional dependence structure of the variables examined and allow us to forecast higher moments. Predicting several conditional quantiles that can characterise the entire distribution of future growth and inflation can be integral in order to properly assess growth vulnerability and inflation stability adrian_vulnerable_2019, manzan_are_2013. \\
The intersection of latent factors with quantile regression models is fairly recent. ando_quantile_2011 have considered a quantile regression model with factor-augmented predictors, whose effect is allowed to vary across the different quantiles. More recently, ando_quantile_2020 introduced a new procedure for analysing the quantile co-movement of a large number of time series based on a large scale panel data model with factor structures. In their study, the latent factors are allowed to vary across the different quantiles of the variables from which they are extracted and, as such, their model is a quantile factor model. Similarly, chen_quantile_2019 estimate scale-shifting factors and quantile dependent loadings, thus factors may shift characteristics (moments or quantiles) of the distribution of the set of directly observable measures, other than its mean, and factor loadings are allowed to vary across the distributional characteristics of each variable. phella_consistent_2020 contributed to this relevant literature by using mean shifting factors as a method of dimension reduction and use these latent factors as additional regressors in quantile autoregressive models. Results showed that the distribution of CPI inflation is best modelled parametrically with a quantile autoregression of order 1, while the distribution of GDP growth is best modelled as a factor-augmented quantile autoregression and such latent factors impact certain quantiles differently. Therefore, our pool of candidate models in the forecasting exercise in this paper will include multiple lag QARs, factor augmented QARs or QARs augmented with targeted macroeconomic variables.\\
Nevertheless, in order for the model averaging approach to outperform the model selection, weights assigned to each model need to be correctly chosen and how such weights are chosen can have a significant impact on forecast performance. buckland_model_1997 and burnham_model_2002 construct model averaging weights based on the values of the Akaike information criterion(AIC) and the Bayesian Information Criterion (BIC). In a seminal article, hansen_least_2007 proposed that weights in least squares model averaging should be chosen over a discrete set by minimizing a Mallow's criterion, as such an estimator is asymptotically optimal in terms of optimising the mean squared error. Meanwhile, hansen_jackknife_2012 propose a Jackknife model averaging (JMA) approach for least squares regression where weights are selected by minimising a leave-one-out cross-validation criterion function, which was extended for models with dependent data by zhang_model_2013. More recently, cheng_forecasting_2015 demonstrated that both the Mallows and leave-h-out cross-validationn criteria remain valid in factor-augmented regression forecasts, as the factor estimation error is negligible, without any restrictions on the relation between N and T.\\
lu_jackknife_2015 extended the JMA of hansen_least_2007 to the quantile regression framework and also proposed a Mallows-type information criterion for QR model averaging, the Quantile Regression Information Criterion (QRIC), which has a computational advantage over the Jackknife approach. We therefore employ this multitude of weighting methodologies over the competing models. We utilise the AIC and the BIC, which use exponential weighting and are often referred to as \enquote{smoothed} averaging buckland_model_1997, the Quantile Regression Information Criterion (QRIC) and the Jackknife weighting method, as those are outlined in lu_jackknife_2015. We also choose to include in our forecast performance comparison the quantile autoregressive model of order 1, QAR(1), as a naive benchmark model similar to the one proposed in pasaogullari_simple_2010 and pasaogullari_simple_2010, as well as a full model which includes all the possible predictors under consideration.\\
We employ these methodologies to determine which model averaging weighting choice is the best for forecast performance, in terms of coverage and final prediction error when predicting one-quarter-ahead GDP growth and CPI inflation for the United Kingdom.\footnote{Along with the coverage rates, we also consider an interval score in order to obtain information regarding the magnitude of violations when they take place.} We specifically employ these methodologies within our estimation sample only, in order to assign fixed weights to the competing models and then use these corresponding weights to obtain the averaged model and produce forecasts, which are in turn evaluated by the aforementioned performance measures. Our results demonstrate that, on average, the equal weights or the weights obtained by the AIC and BIC perform equally well, but are somehow outperformed by the QRIC and the Jackknife approach on the majority of the quantiles of interest of GDP growth, both in terms of coverage and final prediction error. Furthermore, the QRIC and Jackknife model averages ourperform the full model where all available information has been used. On the other hand, the Quantile Autoregression of order 1 of CPI inflation, a model that is often chosen as the best predictor in forecasts of the average value, outperforms all model averaged methodologies. \\
The remainder of the paper is organised as follows. In Section 2, we outline the framework, present the range of models we consider and describe the model averaging methodologies. Section 3 examines the forecast evaluations with respect to the two forecast performance measures. Concluding remarks are given in Section 4 and information regarding mnemonics and the competing models are referred to an appendix.\\
We begin by outlining the factor model used in the sequel. Let
\\ where $X_t$ is an $N \times 1$ vector of observable variables characterising the economy, $\Lambda_t$ is an $N \times k$ matrix of factor loadings, $F_t$ is a $k \times 1$ vector of the $k$ latent common factors and $e_t$ is an $N \times 1$ vector of idiosyncratic disturbances. The errors are allowed to be both serially and (weakly) cross sectionally correlated.\\
As in stock_forecasting_2002, factors are extracted via the principle components approach and the estimated factors and estimated factor loadings are defined as:
The resulting principal components estimator of $F$ is then $\hat{F}=\frac{X'\hat{\Lambda}}{N}$, where $\hat{\Lambda}$ is set equal to the eigenvectors of $X'X$ corresponding to its $k$ largest eigenvalues. In the remainder of this paper, the number of factors $k$ would remain fixed and can be estimated using the information criteria outlined in bai_determining_2002, who take into account the sample size both in the cross-section and time-series dimensions.\\
Ideally, we would like to include all the macroeconomic variables in $X_t$ as additional regressors in the quantile autoregressive model, however, due to the curse of dimensionality, we wish to reduce the dimension of $X_t$ with the use of factors as a way of summarising all the available information. Let therefore $\lbrace y_{t}, V_{t}, F_{t} \rbrace_{t=1}^T$ be a random sample, where $y_t$ is a scalar dependent variable (in this work the CPI inflation rate or the GDP growth rate), $V_t=\lbrace (V_{1,t}, V_{2,t},...) \rbrace$ is a set of targeted macroeconomic variables of countably finite dimension and $F_t=\lbrace (F_{1,t}, F_{2,t},...) \rbrace$ is a set of latent factors of countably finite dimension that need to be estimated a priori from a large panel dataset.\footnote{The vector $F_{t}$ is unobservable, therefore, in practice, we replace the infeasible vector $F_{t-1}$ with the feasible vector $\hat{F}'_{t} $, where $\hat{F}_{t}=(\hat{F}_{1,t},...,\hat{F}_{k,t}) \in \Re^k, k\in \aleph$, is the vector of estimated factors from the panel data $X_{i,t}$. However, given that we are not interested in coefficient inference, it is sufficient that factor estimation error has already been proven to be negligible.} \\
Without loss of generality, in the remainder of the section we assume that $Y_{1,t}=1$. Under the assumption that the conditional distribution of $y_{t+1}$ given $Z_{t}$, where $Z_{t}=(Y_{t}, V_{t}, F_{t})$, is continuous, we can then define the $\tau^{th}$ conditional quantile of $y_{t+1}$ given $Z_{t}$ as the measurable function $q_{\tau}$ satisfying the conditional restriction
\\ We consider a sequence of approximating models $m=1,2,...,M$, where the $m^{th}$ model uses $r_m$ regressors belonging to $Z_{t}$ and $M$ may go to infinity with the sample size. We write the $m^{th}$ approximating model as the Quantile Autoregression (QAR) of order p,
\\ where $ \theta_{(m)}\equiv(\theta_{1(m)},...,\theta_{r_m(m)})'$, $Z_{t(m)}=(Z_{t1(m)},...,Z_{tr_m(m)} )'$, $Z_{tj(m)}$ for $j=1,...,r_m$, are variables in $Z_t$ that appear as regressors in the $m^{th}$ model and $\theta_{j(m)}$ are the corresponding coefficients. The $\tau^{th}$ Quantile Autoregression Estimator (QARE) of $\theta_{(m)}$, proposed by koenker_quantile_2006, is defined as
\\ where $\rho_{\tau}(e)=e(\tau-\mathbbm{1}(e<0))$ is the \enquote{tick} loss function. Let $\hat{\epsilon}_{t+1(m)}(\tau)\equiv y_{t+1}-\hat{\theta}'_{(m)}(\tau)Z_{t(m)}$ be the quantile residual and let $w(\tau)\equiv(w_1(\tau),...,w_m(\tau))'$ be the weight vector in the unit simplex of $\Re^M$ and $\mathcal{W}= \lbrace w \in [0,1]^M: \sum_{m=1}^M w_m(\tau)=1 \rbrace$. Therefore, for $t=1,...,T$, the model averaging estimator for the $\tau^{th}$ quantile is given by\\
\\ It is evident in equation ((ref)) that the weight vector $w_m(\tau)$ differs across different quantiles. This implies that a heavier importance could be placed on different competing models depending on the quantile under consideration. For notational simplicity however, we drop this dependence hereinafter. Furthermore, the weight vector is independent of time. kascha_combining_2010 have previously found that time-varying weights for combining inflation density forecasts provide no advantage over fixed weights. Furthermore, in context, the computational cost of recursive weights with an extensive pool of candidate quantile models can become particularly high. As a result, unless there is strong evidence of a structural change that would imply different suitable models at each period, there is no significant gain in employing time-varying weights.\\
Both the AIC and BIC are often referred to as \enquote{smoothed} averaging buckland_model_1997, which use exponential weights of the form $\frac{exp(\frac{-Inf_m}{2})}{\sum_{j=1}^M exp(-Inf_j)}$, where $Inf_m$ is an information criterion for the $m^{th}$ model. The aforementioned criteria assess each model fit, while at the same time penalizing for the number of estimated parameters, albeit the BIC penalises model complexity more heavily. Both the AIC and BIC criteria are easy to compute regardless of the number of competing models $M$ we are considering and they have been proved to outperform the simplest model combination, i.e. equal weights. In the quantile regression context, following machado_robust_1993, for the $m^{th}$ model, the AIC and BIC are respectively defined as,
\\ where $r_m$ is the dimension of the independent vector $Z_{t(m)}$.\\
Thr AIC and BIC weights for model $m$ are thus respectively defined as,
The QRIC is a Mallows-type information criterion for QR model averaging, whose criterion function has been outlined in lu_jackknife_2015. Mallows' type criteria tend to compare the predictive ability of subset models to that of a full model, but they still balance the trade off between obtaining a \enquote{good} model that contains as few variables as possible. Letting therefore $\hat{\epsilon}_{t+1(m)}(\tau)\equiv y_{t+1}-I_{t(m)}'\hat{\theta}_{(m)}(\tau)$ and $\hat{\epsilon}_{t+1}(w)\equiv \sum_{m=1}^M w_m \hat{\epsilon}_{t+1 (m)}(\tau)$, the QRIC can be defined as,
where $Q_T(w)=\frac{1}{T}\sum_{t=0}^T\rho_{\tau}(\hat{\epsilon}_{t+1}(w))$ indicates the average in-sample QR prediction error. $F$ and $f$ denote the CDF and PDF, respectively, of $\epsilon_{t+1}(\tau)$ and $\sum_{m=1}^M w_m r_m$ signifies the number of effective parameters in the combined estimator. Therefore, in order to choose the weight vector $w$, by the QRIC, we must estimate the sparsity function of $\epsilon_{t+1}(\tau)$, $s(\tau)=\frac{\tau(1-\tau)}{f(F^{-1}(\tau))}$. Following koenker_quantile_2005 this can be estimated by,
\\ where $\tilde{F}^{-1}_T$ is an estimate of the quantile function $F^{-1}$ of $\epsilon_{t+1}(\tau)$ based on the quantile residuals obtained from the largest approximating model, $h_T=T^{-\frac{1}{5}} \lbrace 4.5 \phi^4(\Phi^{-1}(\tau))/[2\Phi^{-1}(\tau)^2+1]^2 \rbrace^{\frac{1}{5}}$ and $\phi$ and $\Phi$ are the standard normal PDF and CDF, respectively. The resulting empirical QRIC weighting vector is then defined as,\footnote{There is no closed form solution for equation ((ref)) but the optimal weight vector can be found by linear programming as in typical quantile regressions.}
\\ It is worth noting that due to the presence of the sparsity function, the QRIC may not perform as well on extreme quantiles where the sparsity is low. However, the same holds true for the quantile regression estimator, which is why our analysis excludes the most extreme tails of the distribution and only focuses on values where $\tau \in [0.1,0.9]$.\\
The Jackknife selection of the weighting vector $w$ is the most distinct, as rather than imposing the weighting on the fitted value of the dependent variable it, in practice, weighs the estimated quantile regression coefficients and uses the average coefficient estimator in the forecasting exercise. The Jackknife weight vector is optimal in terms of minimising the final prediction error (FPE), one of our forecast performance measures, in the sense of akaike_statistical_1970. For all the competing models under consideration $m=1,...,M$, let $\hat{\theta}_{t(m)}$ denote the jackknife estimator of $\hat{\theta}_{(m)}$ in model $m$ with the $t^{th}$ observation excluded from the estimation. The jackknife choice of the weight vector $\hat{w}^{Jknife} =(\hat{w}^{Jknife}_1,...,\hat{w}^{Jknife}_M)$ is obtained by choosing $w \in \mathcal{W}$ to minimise the leave-one-out cross-validationn criterion function,
\\ The resulting Jackknife weighting vector is therefore defined as,
\\ Is worth noting that, though the leave-one-out cross validaion criterion function is convex in $w$ and can be minimised by running the quantile regression of $y_{t+1}$ on $Z'_{t(m)}\hat{\theta}_{(m)}$, it cannot guarantee that the resulting solution lies in $\mathcal{W}$. However, one can express the constrained minimisation problem in ((ref)) as a linear programming problem (see e.g. lu_jackknife_2015, p. 43).\\
The quantile autoregression framework allows us to obtain forecasts for the conditional quantiles of the dependent variable, however, once we obtain the realisation for the dependent variable, it only provides us with the actual value of the dependent variable, but not its quantile at that period. For example, in our empirical context, we can forecast the $\tau^{th}$ quantile of the inflation rate $H$-periods ahead, however, fast forward $H$ periods later what we obtain is the value of the inflation rate, $y_{t+H}$, and not the value of the $\tau^{th}$ quantile, $Q_{y_{t+H}}(\tau)$. This implies that evaluating the forecast performance of the aforementioned weighting methodologies is not a trivial issue, but it also allows for a certain degree of flexibility. We therefore evaluate the suggested methodologies using two different forecast performance measures.\\
Though we may not be able to obtain an observation for the realisation of the $\tau^{th}$ quantile, in practice quantile forecasts correspond to one-sided interval forecasts of the dependent variable of interest. Therefore, the unconditional coverage testing framework of christoffersen_evaluating_1998 poses a suitable measure for evaluating the \enquote{accuracy} of the quantile forecast. Correct coverage tests in practice aim at testing whether a sequence of conditional quantile forecasts satisfies certain optimality conditions. Here, we are not interested in testing in absolute terms if the quantile forecasts are correct, but which method for choosing the weight vector might be better, thus we are interested into which method provides as with the most correct coverage rate. \\
Assume we obtain a sequence of out-of-sample forecasts,$\lbrace \hat{q}_{t+h|t}(\tau) \rbrace_{t=0}^P$, of the conditional quantile $Q_{y_{t+H}}(\tau)$ of the time series $y_{t+H}$. This quantile forecast is an upper limit for an interval forecast for the time series $y$, for time $t+H$, made at time $t$, for the coverage probability $\tau$.\footnote{If the $\tau^{th}$ quantile forecast, $q_{t+H|t}(\tau)$ is accurate, then in practice the number of times that the realisation of $y_{t+H}$ falls below the forecast value should on average equal the nominal quantile level for which we are forecasting.} However, as we move across the quantiles and towards the upper tail of the distribution, it is more likely that most of the observations will fall below the estimated quantile, which can result in better coverage results by construction, rather than due to better forecast performance. Therefore, for each quantile under evaluation, we choose to obtain a two-sided interval forecast, $\lbrace (L^{\tau}_{t+H|t}(p), U^{\tau}_{t+H|t}(p)) \rbrace_{t=0}^P$, where $L^{\tau}_{t+H|t}(p)$ and $U^{\tau}_{t+H|t}(p)$ are the lower and upper limits of the ex-ante interval forecast for quantile $\tau$, for time $t+H$, made at time t, for the nominal coverage probability, p. We can then define an indicator variable, $\mathcal{I}^{\tau}_{t+H}$ for time $t+H$, where
\\
We can therefore measure the efficiency of this forecast interval by measuring the amount of times that the indicator variable takes a value of 1 as a proportion of the total number of prediction periods. For example, in the one step ahead forecast, the coverage rate is defined as
\\ where $P$ s the number of periods from the sample set aside for forecast evaluation. The quantile forecast is dependent on the weight vector chosen by each method, though this dependence has been dropped for simplicity, therefore, the closer the empirical coverage rate to the nominal coverage probability, p, the more \enquote{accurate} the method in terms of approximating the true quantile of the distribution.\\
Although the unconditional coverage measure is a suitable evaluation method for quantile forecasts, it can only provides us with an indication of whether the realised observation falls within the forecast interval or not. In case the realised observation falls outside the interval forecast, implying we have a violation, we have no information on how far away it is. It is true that in the spirit of quantile regression the only relevant issue is only whether or not the observation falls below an estimated quantile and no attention is given on how far an observation is from the quantile. Nevertheless, in this empirical context where the prediction period is rather small and thus empirical coverage rates might not provide a clear image on which methodology performs better, we also wish to examine the behaviour of observations that fall outside the interval forecast. We therefore complement the unconditional coverage measure with the interval score of gneiting_strictly_2007. This scoring rule has an intuitive appeal in that the forecaster incurs a penalty, the size of which depends on the relevant confidence level, if the observation misses the interval, but is also rewarded for narrow prediction intervals. In our empirical context, the interval score can be defined for a nominal coverage probability $p$ as,
\\ where the dependencies of the lower and upper bound, $L$ and $U$, on $\tau$, $t$, $t+h$ and $p$ has been dropped for simplicity. It is evident that this scoring rule imposes the same penalty for violations that occur below the lower bound and above the upper bound. It is worth mentioning that one could examine an asymmetric form of penalisation that would take into consideration the spectrum of the distribution available below the lower bound or above the upper bound, which is dependent on the quantile of interest considered, since this could impact the possible magnitude of a violation. However, in this context, this is not expected to influence the results in a significant way and has not been considered.\\
We therefore obtain the one-step-ahead average interval score as,
\\ where $P$ s the number of forecast evaluation periods and the dependence of the Interval Score on the weight vector $w$ is due to the dependence of lower and upper bounds of our confidence intervals on $w$, which has been dropped for simplicity. In this case a lower interval score is desirable, since it implies that even if the realised observation falls outside the forecast interval, the magnitude of the violation is not large.\\
Different loss functions $\mathcal{L}$ arguably correspond to different optimal forecasts. For example, letting $\hat{\epsilon}_{t+H}\equiv y_{t+H}-\hat{q}_{t+H|t}$ be the forecast error, if a quadratic loss function is used such that $\mathcal{L}(\hat{\epsilon}_{t+H})=\hat{\epsilon}_{t+H}^2$, then the optimal forecast for $y_{t+H}$ would be conditional mean. Similarly, if the absolute value loss function is used such that $\mathcal{L}(\hat{\epsilon}_{t+H})=|\hat{\epsilon}_{t+H}|$, then the optimal forecast for $y_{t+H}$ would correspond to the conditional median. Based on this idea, giacomini_evaluation_2005 argue that, in the case conditional quantiles, the corresponding loss function would be the asymmetric linear loss function, $\mathcal{L}(\hat{\epsilon}_{t+H})\equiv \rho_{\tau}(\hat{\epsilon}_{t+H})\equiv \hat{\epsilon}_{t+H}(\tau-\mathbbm{1}(\hat{\epsilon}_{t+H}<0))$ and therefore the optimal forecast for $y_{t+H}$ is its conditional $\tau^{th}$ quantile. Therefore, in the model averaging framework, an alternative forecast performance measure for the one-step-ahead forecast is the Final Prediction Error (FPE), which, as a function of the weight vector, is defined as,
\\ where $\hat{w}_m$ is chosen by one of the suggested methodologies. In this case, the parameter $\tau$ describes the degree of asymmetry in the loss function, where a value less than one-half indicates that over-predicting results to a greater loss for the forecaster than under-predicting by the same magnitude.\footnote{When the parameter $\tau$ equals one-half, over-predicting and under-predicting generates the same loss, thus we converge to the absolute value loss function where the optimal forecast is the conditional median.} Therefore, the smaller the FPE, the better the weight vector method in terms of the out-of-sample quantile prediction error. \\
Inflation is one of the most important variables due to its dominant role in many macroeconomic models (levin_is_2002, angeloni_new_2006). Quarterly \enquote{Inflation Reports} have become the new norm for the majority of central banks and even though the costs and benefits of transparency are still widely debated, it is broadly agreed that a central bank should be concerned with inflation forecasting. Nevertheless, as argued by, inter alia, faust_forecasting_2013, henry_is_2004, a point forecast of inflation without some measure of associated uncertainty is arguably of little value. Policymakers forecasting inflation need to not only consider the most likely outcome for inflation, but all possible paths that inflation can take, which involves examining the dynamics and higher moments of inflation. Similarly, over the recent years policy-makers have shifted focus towards downside risk for GDP growth rather than point forecasts for the conditional mean of growth. This is due to the fact that such point forecasts ignore the risks surrounding these central forecasts and thus may paint an overly optimistic picture of the state of the economy adrian_vulnerable_2019. A density forecast therefore gives a more complete characterisation of future growth and inflation prospects and forecasting future conditional quantiles is a computationally easy method to obtain it, while allowing different quantiles to exhibit different sensitivity to predictors.\\
In this empirical study we examine, under a model averaging approach, which methodology for choosing the appropriate weight vector is better suited for forecasting the one-quarter-ahead annual GDP growth rate and CPI inflation rate. Several of our competing models will involve latent factors as a way to summarise a large amount of information from different macroeconomic variables. We therefore consider for the estimation of factors series containing data on inflation, real activity and indicators of money and key asset prices for the United Kingdom. We will undertake the analysis using a sample which includes quarterly data from 174 macroeconomic variables spanning from the second quarter of 1991 to the second quarter of 2018, with a total of $T=109$ observations.\footnote{This dataset is a subset of the dataset used by ellis_what_2014 to create a time-varying factor augmented VAR model for the UK monetary transmission mechanism and details regarding the variable used can be found in the Appendix.} All the data has been stationarised prior to use. Given that latent factors have been proven important in modelling and predicting our variables of interest we will extract two latent factors from the macroeconomic dataset. The latent factors are extracted using a recursive estimation scheme, thus in each period within the prediction sample all available past information is utilised, but the results hold under alternative estimation schemes as well. Furthermore, given the close relationship between growth and inflation, we shall include these variables as individual regressors. Therefore, there would overall be 5 possible regressors for each dependent variable, as shown in Tables (ref) $\&$ (ref).\\
For each dependent variable of interest we have constructed 57 non-nested candidate models with all the possible combinations between the five regressors and an intercept, where the smallest models has at least two regressors. Our smallest and largest model can therefore be characterised by the following regressors, $\lbrace 1, r_1 \rbrace$ and $\lbrace 1,r_1, r_2, r_3, r_4, r_5 \rbrace$, respectively. We split the sample into an estimation sample of size $T_1=\frac{T}{2}$, used to determine the appropriate weight vector for model averaging and an evaluation sample of size $T_2=T-T_1$, used for forecast performance evaluation. We then choose the relevant weight vector for the 57 competing models under the four methodologies presented and then using the averaged model construct one-period-ahead forecasts for 9 different quantiles where $\tau \in [0.1,0.9]$. We also construct forecasts with the naive quantile autoregressive model of order 1, a full model which includes all the aforementioned regressors, as well as an average model which assigns an equal weight to all 57 competing models. \\
As it was outlined in the framework section, the 57 competing models get assigned a corresponding weight for model averaging by each of the four methodologies, which are then used to obtain a forecast average of the one-step-ahead GDP growth and CPI inflation rate. Tables (ref) -(ref) in the Appendix demonstrate the allocated weights by each methodology across the 57 competing models, for each of the two dependent variables. It is evident that the AIC and BIC allocate similar weights, which are roughly equal across all the models under consideration, so their out-of-sample performance is expected to be similar. On the other hand, the QRIC allocates significantly high weights to specific models, assigning a zero weight to the majority of the competing models. It is also evident that this criterion favours low-dimensional models, in contrast with the Jackknife method, which favours higher dimension models. Furthermore, Tables (ref) -(ref) in the Appendix show the empirical coverage rates and final prediction error for all the methodologies under consideration. From those results, as expected given the allocated weights, the out-of-sample performance of a simple average model of equal weights is comparable to the performance of an average model where weights are assigned using the AIC or BIC, for both growth and inflation. For brevity purposes therefore, in the remainder of the paper we shall be comparing the forecasting performance of the naive model, the full model and the average models with an equal weighting and with weighting assigned by the QRIC and Jackknife criterion. \\
Figure (ref) shows the empirical coverage rate by each of the remaining competing methods. The top panel has GDP growth as the dependent variable and the bottom panel is for CPI inflation. We evaluate $9$ equidistributed points for $\tau \in [0.1, 0.9]$, where for each quantile of interest we obtain an interval forecast with a nominal coverage probability of $10\%$.\footnote{For example, if the quantile of interest is the conditional median, we obtain an interval forecast by estimating the $45^{th}$ and $55^{th}$ conditional quantile} The shaded region demonstrates the confidence interval for the null hypothesis that the empirical coverage rate is equal to nominal coverage probability of $10\%$.\footnote{With larger nominal coverage probabilities (e.g. $20\%$), the absolute performance of the methodologies, in particular with respect to coverage, improves due to the larger forecast intervals, but the comparative performance between methodologies remains identical.}\\
Focusing on GDP, in terms of coverage, although it is not clear whether a single methodology outperforms all its competitors, several conclusions can be made. First, we can see that all the methodologies tend to provide more conservative interval forecasts at the lower tails of the distribution, but such intervals become more liberal as we move towards the upper tail. Furthermore, on several occasions we see that the null hypothesis of nominal coverage of $10\%$ is not satisfied at certain points of the distribution by several methodologies, indicating that perhaps additional regressors should be considered. Nevertheless, overall, the QRIC weighting methodology seems to have a rather competitive performance, with empirical coverage rates close to the nominal probability for a substantial range of the distribution. More interestingly, the fact that the full model does not outperform several of the model averaging methodologies proves that there are gains to be made by considering an array of competing models, even if such models do not utilise the full information. This might be, as argued by granger_can_1996, due to presence of different estimation biases in the competing models that may cancel each other out and, as a result, provide us with a more accurate forecast than the full model.\\
With respect to CPI inflation, there exists a much clearer picture. Firstly, all the methodologies provide liberal forecast intervals for the majority of the distribution, with the exception of the most extreme tails. A simple average is often outperformed by the QRIC and Jackknife methodologies, but overall it is evident that the naive QAR(1) model seems to be outperforming all the competitors, by achieving coverage rates that are the closest to $10\%$ for the majority of the quantiles of interest.\\
However, taking into consideration the interval score can provide more clarity to our conclusions.\footnote{It is worth noting once again that one could consider an asymmetric penalty, in order to penalise violations according to which quantile of the distribution once is interested and the maximum possible violation that can incur at that point.} As it is evident in Figure (ref), the interval score for the QRIC and Jackknife approaches is significantly lower when compared to the other methodologies, particularly in the case of GDP growth. This result indicates that even though the QRIC and Jackknife approaches may have a significant number of violations (i.e. the realised observation falls outside the interval forecast) and thus not as accurate coverage rates, the magnitude of the violation is not large and/or the prediction intervals provided by these methodologies are narrower. This implies that for several of the prediction periods, the realised observation is not far from the lower or upper bound of the forecast interval. In contrast, the simple average model may have a similar number of violations, but when a violation takes place the forecast error is significantly larger. Furthermore, similarly to what we have identified before, the full model seems to be outperformed by certain model averaging methodologies.\\
When forecasting CPI inflation, a similar picture is painted as the one for GDP growth, in that a simple average is associated with larger mean forecast errors across the whole distribution. Similar with its coverage performance, the QAR(1) seems to be outperforming all the methodologies under consideration, implying that the naive model might be the best for predicting the quantiles of CPI inflation. Another take away from this measure of CPI inflation is the fact that the interval score and thus the forecast error across all the methodologies seems to be smaller in the upper tail of the distribution. This could be explained by a higher variation in the right tail of the CPI inflation distribution, as this has been found in phella_consistent_2020, which enables for better forecasting regardless of the allocation of weights across the different competing models. Conversely, downturns are more difficult to predict, probably due to the zero lower bound that is present, thus lower quantiles are associated with larger violations. \\
Looking in conjunction at the two figures, we can conclude that, in the case of GDP growth, the QRIC and, to a certain extent, the Jackknife model averages can produce reliable forecasts for a large spectrum of the distribution, though such forecasts tend to be rather liberal. Notably however, for CPI inflation, the QAR(1) seems to be performing the best with coverage rates close to $10\%$ across most of the evaluated quantiles and with violation magnitudes lower than the ones produced by the QRIC and Jackknife. This is in line with the notion present in the relevant literature that past inflation is the best predictor for future inflation. On the other hand, the fact that the QRIC and Jackknife model averages for GDP growth outperform the full model, demonstrates the importance of considering the aggregation of multiple competing models. Furthermore, the fact that for GDP growth the methodologies that forecast better are the ones which attribute significant weights to models with latent factors (see Table (ref) in Appendix), is an indication that such latent factors are relevant for out-of-sample estimation of conditional quantiles of growth. This is a complementary result to that present in the in-sample estimation literature, where latent factors, summarising a larger information set, were found to carry relevant information for in sample estimation of GDP growth, but not CPI inflation.\\
The evidence when considering the final prediction error are much clearer than that for coverage. It is clear in both panels of Figure (ref) that the final prediction error of the QRIC and Jackknife methodologies seem to be the best for most of the quantiles considered, albeit the full model is competitive in the extreme right tail. This implies that a model average with distinctive weighting serves as a better predictor for $y_{t+1}$. The good performance of the QRIC methodology is not surprising in this case, as the quantile regression criterion was constructed in order to minimise this type of prediction error. Once again, we see that the naive QAR(1) model of CPI inflation outperforms all the competing methodologies, implying that conditional quantiles estimated solely with past inflation information can act as the best predictor for the future value of inflation. This is a striking difference with mean regressions of the UK CPI inflation, where forecast performance relative to the AR benchmark model is improved when forecasts are combined kapetanios_forecast_2008, though this work utilises information from multiple data sources and data types in the forecast combination, including subjective data. \\
The distinctive differences and similarities between the two performance measures in all three figures highlights certain empirical facts. Firstly, it is evident that despite the computational ease, assigning weights according to the Akaike and Bayesian Information Criteria does not improve forecasting performance over a model averaging approach where each competing model is assigned an equal weight. Most importantly, the distinctive performance of the different model averaging methodologies demonstrates the importance of choosing the assigned weights correctly. Secondly, the differences between the forecasting performance of the naive QAR(1) models in the case of GDP growth versus CPI inflation, demonstrates fundamental differences regarding which variables are the best predictors in each individual case, despite the close relationship between inflation and growth.\\
Similarly with the results of in-sample estimation found in phella_consistent_2020, past inflation is the best predictor for the future path of inflation, while GDP growth can be predicted best if additional latent factors and models which include such factors are taken into consideration. The former has been long discussed in the case of mean regressions of inflation, however these results prove that this remains true in the case of conditional quantiles as a way to obtain a trace of the inflation distribution and therefore a measure of the risk and uncertainty surrounding future inflation. Lastly, the good performance in forecasting GDP growth of the QRIC and Jackknife approach, which assign higher weights in models with latent factors, is also in line with the in-sample literature, where these latent factors had been shown to have a distinctive impact across different parts of the growth distribution. These results once again highlight the necessity of high dimensional macroeconomic data when extracting latent factors and at the same time demonstrates the importance of such latent factors in quantile regression models for different variables of interest.\\
We have constructed forecasts of the conditional quantiles for the GDP growth rate and CPI Inflation rate of the United Kingdom, using a model averaging approach where the weight vector has been chosen by different criteria: the AIC, BIC, QRIC and Jackknife. The literature has long supported that forecast combinations of average inflation can improve out-of-sample performance, however this literature ignores the risks and uncertainties around central forecasts. Similarly for growth, although Bayesian model averages have long been used, frequentist model averaging has not gained significant attention. This work addresses both issues, by dealing with the conditional quantiles of growth and inflation and thus their whole distributions.\\
We find that, in terms of coverage the distinctive weighting across competing models imposed by the QRIC and Jackknife can outperform an equal weighting, as well as the full model, for forecasting GDP growth, when an interval score is simultaneously considered. On the other hand, the benchmark QAR(1) model of CPI inflation outperforms all model averaging methodologies. In terms of final prediction error, a similar picture exists. Overall, the results demonstrate the importance of high dimensional macroeconomic data and the inclusion of latent factors, as a way to summarise such data, when forecasting GDP growth but confirm that past inflation is the bet predictor for future inflation. Previously, latent factors were shown to be relevant for in-sample estimations of the growth distribution. Although good in-sample performance does not guarantee good out-of-sample performance, this work demonstrates that in the case of growth, latent factors are also relevant and important in out-of-sample forecasts of the conditional quantiles.