Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
56,904 characters · 11 sections · 41 citation commands
Consistent Specification Test of the Quantile Autoregression
\setcounter{page}{1}
Quantile regression (QR) enables the analysis of a continuous range of conditional quantile functions, providing a more complete picture of the conditional dependence structure of the variables examined, rather than a single measure of conditional location. As the awareness of the importance of data heterogeneity increases, quantile regression has become even more relevant koenker_quantile_2017. At the same time, the recent availability of large datasets has generated interest in models with many possible predictors. Inference methods overcoming the curse of dimensionality have become increasingly popular and factors in particular, have been proven useful in overcoming the limited information bias. This arises from the fact that the information set of decision makers is sufficiently larger than the information set captured by conventional empirical models. \\
The intersection of latent factors with quantile regression models is fairly recent. ando_quantile_2011 have considered a quantile regression model with factor-augmented predictors, whose effect is allowed to vary across the different quantiles. Their study models the quantile structure in a cross-sectional context. More recently, ando_quantile_2020 introduced a new procedure for analysing the quantile co-movement of a large number of time series based on a large scale panel data model with factor structures. In their study the latent factors are allowed to vary across the different quantiles of the variables from which they are extracted and as such their model is a quantile factor model. Similarly, chen_quantile_2019 estimate scale-shifting factors and quantile dependent loadings, thus factors may shift characteristics (moments or quantiles) of the distribution of the set of directly observable measures, other than its mean, and factor loadings are allowed to vary across the distributional characteristics of each variable. We add in this relevant literature by using mean shifting factors, that is factors that are only allowed to alter the location of the observable measures, as a method of dimension reduction and use these latent factors as additional regressors in quantile autoregressive models. \\
Meanwhile, in the empirical literature, the majority of the work has examined whether univariate regression models should be augmented with factors, in estimates and forecasts of the conditional mean (stock_forecasting_2002, bernanke_measuring_2005). However, point estimates and forecasts for the conditional mean of macroeconomic variables, like growth and inflation, ignore the risks around this central estimates adrian_vulnerable_2019. Furthermore, over the recent years, the tendency of policy-makers to focus more on the downside risk of GDP growth demonstrates the need to analyse the conditional quantiles of GDP growth and examine how they should be modelled individually. Similarly, as supported by levin_is_2002 and angeloni_new_2006, the inflation rate is one of the most important variables, due to its dominant role in many macroeconomic models and as argued by, inter alia, henry_is_2004, the dynamic behaviour of the inflation rate has a number of economic implications. In order to correct for this omission, we trace the conditional distributions of GDP growth and CPI inflation in the United Kingdom by estimating several conditional quantiles. \\
In this paper therefore, our contribution to the existing literature is twofold. Firstly, we allow for the interaction of mean shifting factors, summarising a larger information set, with quantile autoregressive models and furthermore, as a result, obtain a trace of the conditional distribution of a variable of interest. The latter enables us to assess the upside and downside risks present for the dependent variable. We focus on the GDP growth and CPI inflation in the United Kingdom and find that quantile autoregressive models for GDP growth should include latent factors, as a way to summarise multiple macroeconomic variables. Such factors have non-uniform effects on different tails of the growth distribution. However, these latent factors do not carry relevant information for modelling the inflation rate distribution. We also find evidence of a decrease in the asymmetry of UK GDP growth following the recent financial crisis and significant time variation in downside risk. Meanwhile, in the case of inflation upside risk has varied significantly over the years despite a rather stable central tendency. \\
Nevertheless, for any post estimation inference to be valid, the correct specification of the empirical model used needs to be assessed. Therefore, we also propose a test for the joint hypothesis of correct dynamic specification and no omitted latent factor for the quantile autoregression (QAR), as that is outlined in koenker_quantile_2006. The test we suggest is related to the conditional moment tests of bierens_consistent_1982, bierens_consistent_1990, as well as the specification test of parametric quantile models of escanciano_specification_2010\footnote{Though similar, the work by escanciano_specification_2010 cannot account for the factor estimation error present in our context due to the use of unobservable factors that need to be estimated a priori.}. In practice, we propose two tests: the first one can be used to determine if the quantile autoregression is correctly specified, conditional on all available past information of the dependent variable and latent factors. In cases where the null hypothesis fails to reject when this test is applied, we have evidence that the QAR is correctly specified. If however, the null hypothesis is rejected, we perform a second test in order to determine whether the misspecification arises because latent factors are an omitted variable or because the model is dynamically misspecified. The second test therefore involves testing the null hypothesis of the Factor-Augmented QAR (FA-QAR) being correctly specified. Valid asymptotic critical values are obtained via a bootstrap procedure based on resampling functions of the entire history of available information as in corradi_information_2009.\\
The remainder of the paper is organised as follows. In Section 2, we outline the framework and describe the testing procedure in order to examine whether a Quantile Autoregression should be augmented with latent factors. We define the test statistics and study the asymptotic properties of the suggested statistics. Section 3 demonstrates some finite sample performance results based on a limited Monte Carlo simulation. In Section 4, we examine the distributions of UK growth and inflation, based on the use of a large-scale macroeconomic dataset and discuss the empirical findings. Concluding remarks are given in Section 5 and information regarding proofs is referred to an appendix.\\
We begin by outlining the factor model used in the sequel. Let
where $X_t$ is an $N \times 1$ vector of observable variables characterising the economy, $\Lambda_t$ is an $N \times k$ matrix of factor loadings, $F_t$ is a $k \times 1$ vector of the k latent common factors and $e_t$ is an $N \times 1$ vector of idiosyncratic disturbances. The errors are allowed to be both serially and (weakly) cross sectionally correlated.\\
As in stock_forecasting_2002, factors are extracted via the principle components approach and the estimated factors and estimated factor loadings are defined as:
The resulting principal components estimator of $F$ is then $\hat{F}=\frac{X'\hat{\Lambda}}{N}$, where $\hat{\Lambda}$ is set equal to the eigenvectors of $X'X$ corresponding to its $k$ largest eigenvalues. In the remainder of this paper the number of factors $k$ would remain fixed and can be estimated using the information criteria outlined in bai_determining_2002 who take into account the sample size both in the cross-section and time-series dimensions.\\
Suppose we observe a real-valued dependent variable $y_t$ and a high-dimensional information vector $X_t$. Due to the curse of dimensionality we wish to reduce the dimension of $X_t$ with the use of factors, as a way to summarise all the available information. The desirable information vector is therefore $I_{t-1}=(Y'_{t-1}, F'_{t-1}) \in \Re^d$, $d=(p+1)+k$, where $F_{t-1}=(F_{1,t-1},...,F_{k,t-1}) \in \Re^k, k\in \aleph$, is the vector of latent factors and $Y_{t-1}=(1, y_{t-1},...,y_{t-p}) \in \Re^{p+1}$, where $A'$ denotes the transpose of $A$. In practice, the vector $F_{t-1}$ is unobservable, therefore, we replace the infeasible information set $I_{t-1}$ with the feasible information vector $\hat{I}_{t-1}=(Y'_{t-1}, \hat{F}'_{t-1}) \in \Re^d$, $d=(p+1)+k$, where $\hat{F}_{t-1}=(\hat{F}_{1,t-1},...,\hat{F}_{k,t-1}) \in \Re^k, k\in \aleph$, is the vector of estimated factors from the panel data $X_{i,t-1}$. In the remainder of this paper we assume that the time series process $\lbrace(Y_t, F_{t-1}')':t=0,\pm 1,\pm 2,...\rbrace$, defined on the probability space $(\Omega, \mathcal{A})$ is strictly stationary and ergodic.\\
Under the assumption that the conditional distribution of $Y_t$ given $I_{t-1}$ is continuous, we can then define the $\tau^{th}$ conditional quantile of $Y_t$ given $I_{t-1}$ as the measurable function $q_{\tau}$ satisfying the conditional restriction
\\ We use the Quantile Autoregression (QAR) model of order p as that is set out in koenker_quantile_2006 so that the variable of interest can then be characterised by the following equation
where ${u_t}$ is a sequence of i.i.d. standard uniform random variable, while $\theta_i(u_t)$ are unknown functions $[0,1] \rightarrow \Re$. Provided that the right hand side of ((ref)) is monotone increasing in $u_t$, the $\tau^{th}$ conditional quantile function of $y_t$ takes the following form,
\\ Our objective is to test whether factors are an omitted variable in a quantile autoregressive model for $y_t$, while at the same time controlling for dynamic misspecification. In the sequel, we are testing the following hypothesis:
against its respective alternative
where $\mathcal{B}$ is a family of uniformly bounded functions from $\mathcal{T}$ to $\Theta$.\\
The null hypothesis states that if the specification is correct, the probability that the observed value of $y_t$ falls below the estimated quantile should, on average, equal the nominal quantile level of interest ($\tau$), with probability one. Hypothesis $H_{0,1}$ is the joint hypothesis that the QAR model is not dynamically misspecified and factors are not an omitted variable. If $H_{0,1}$ did not include such latent factors, the null hypothesis would simplify to the null hypothesis of the correct specification of parametric dynamic quantile models, tested in escanciano_specification_2010, where the conditioning information vector includes only directly observable variables.\\
Testing for the null hypotheses $H_{0,1}$ is a challenging problem, since it involves an infinite number of conditional moments indexed by $\tau \in \mathcal{T}$, where $\mathcal{T}$ is a compact set comprising of the range of quantiles of interest ($\mathcal{T}\subset(0,1)$). Therefore, following bierens_consistent_1990 we can characterise $H_{0,1}$ by the infinite number of unconditional moment restrictions:
where $\Phi$ is an arbitrary Borel Measurable bounded one-to-one mapping from $\Re^d$ to $\Re^d$ and $\xi \in \Re^d$ is a vector of weights. Conditioning on $I_{t-1}$ is equivalent to conditioning on the bounded vector $\Phi(I_{t-1})$, for $I_{t-1}$ and $\Phi(I_{t-1})$ generate the same Borel field. Furthermore, we wish for the weight attached to past observations to decrease over time, thus we define the weighting vector $\xi$ as in de_jong_bierens_1996. Let therefore, $exp(\xi' \Phi(Z_t))=exp(\sum_{j=1}^{t-1} \xi'_j \Phi(Z_{t-j}))$ and $\Xi=\lbrace \xi_j: a_j\leq \xi_j \leq \gamma_j, j=1,2; \vert a_j \vert, \vert \gamma_j \vert \leq \Gamma j^{-\kappa}, \kappa \geq 2\rbrace$.\\
Given therefore a sample $\{(y_t, \hat{I}'_{t-1})': 1 \leq t \leq T\}$ and an estimated parameter value $\hat{\theta}(\tau)$ we consider the following quantile empirical process:
for the Quantile Autoregression Estimator (QARE), proposed by a koenker_quantile_2006), defined as
where $\rho_{\tau}(u)=u(\tau-\mathbbm{1}(u<0))$ is the \enquote{tick} loss function.\\
The null hypothesis holds when the process $S_{1,T}(\xi,\hat{\theta}_T(\tau))$ is close to zero for almost all $(\xi',\tau)' \in \Re^d \times \mathcal{T}$, and thus the test statistic is based on a distance from a standardised sample analogue of $E\{[\mathbbm{1}(y_t-Y_{t-1}\hat{\theta}_T(\tau)\leq 0)-\tau]exp(\xi'\Phi(\hat{I}_{t-1}))\}$ to zero. Some popular norms we could consider are the following:
where $\Phi_1$ and $\Phi_2$ are some integrating measures on $\mathcal{Y}$ and $\mathcal{T}$ respectively and $\mathcal{Y}$ is a generic compact subset of $\Re^d $ containing the origin.\\
The test we propose therefore rejects $H_0$ for \enquote{large} values of such functionals. If $H_{0,1}$ is not rejected then one can conclude that the QAR model is correctly specified with no omitted variables and thus can proceed with inference. If however, $H_{0,1}$ is rejected we still need to ascertain the source of the rejection. A logical next step would therefore be to augment the quantile autoregression model with the feasible estimated factors. We therefore proceed to test the following null hypothesis:
against its negation
In this case, the null hypothesis $H_{0,2}$ states that the Factor-augmented QAR (FA-QAR) is correctly specified. It follows that the test statistic of interest for $H_{0,2}$ will be based on the quantile empirical process $S_{2,T}(\xi,\hat{\theta}_T(\tau)):= T^{-\frac{1}{2}}\sum_{t=1}^{T} [\mathbbm{1}(y_t-Y_{t-1}\hat{\theta}_{1,T}(\tau)-\hat{F}_{t-1}\hat{\theta}_{2,T}(\tau)\leq 0)-\tau]exp(\xi'\Phi(\hat{I}_{t-1}))$. In this instance the factor estimation error not only appears in the conditioning information vector but also influences the indicator function element of the statistic. If the null hypothesis fails to reject then the FAQAR is correctly specified, hence factors were originally an omitted variable. If on the other hand, $H_{0,2}$ is also rejected then we have evidence that the linear specification of the model may not be appropriate and a non-parametric approach may be more suitable, however this is beyond the scope of this paper.\\
Let $\Vert A \Vert =[tr(A'A)^{\frac{1}{2}}]$ denote the norm of matrix A. Throughout, we let $F_t$ be the $k \times 1$ vector of true factors and $\lambda_i$ be the true loadings, with $F$ and $\Lambda$ being the corresponding matrices. The relevant assumptions for the latent factors are those used in bai_inferential_2003 to derive the limiting distributions of the estimated factors, factor loadings and common components and are provided explicitly in the Appendix as Assumptions A-F. We rely on those same assumptions, to demonstrate that factor estimation error does not influence our test statistic.\\
To derive the asymptotic results of the quantile empirical process, we need to further consider the following assumptions. Let, for each $t \in \mathbb{Z}$, $\mathcal{F}_t=\sigma (I_t', I_{t-1}',\ldots)$ be the $\sigma$-field generated by the information set obtained up to time $t$. Define also the family of conditional distributions $F_b(y)\mathrel{\mathop:}= P(Y_t \leq y \vert I_{t-1}=b)$ . Let $f_b$ be the density function of the cumulative distribution function (cdf) $F_b$. In particular, $f_{I_{t-1}}(y)$ denotes the density of $Y_t$ given $I_{t-1}$, evaluated at $y$. Also, the family $\mathcal{B}$, in which the parameter $\theta$ takes values, is endowed with the sup norm\footnote{i.e. $\Vert \theta\Vert_{\mathcal{B}}=\sup_{\tau \in \mathcal{T}} \vert \theta(\tau) \vert$.}. Lastly, given that similar assumptions are needed both when testing the QAR model under $H_{0,1}$ and FA-QAR model under $H_{0,2}$, let $m_j(I_{t-1},\theta(\tau))$ be the conditional quantile function under consideration and $H_{0,j}$ the null hypothesis to be tested. Thus, $m_1(I_{t-1},\theta(\tau))=Y_{t-1} \theta(\tau)$ is the conditional quantile function in the QAR case and $m_2(I_{t-1},\theta(\tau))=Y_{t-1}\theta_1(\tau)+F_{t-1}\theta_2(\tau)$ in the FA-QAR case. \\
Assumption G: Time series Model Checks
\quad \quad 1. $\lbrace(Y_t, F_t')':t=0, \pm 1, \pm2,\ldots\rbrace$ is a strictly stationary and ergodic process. Under $H_{0,j}$, $\lbrace \mathbbm{1}(y_t-m_j(I_{t-1},\theta(\tau))\leq 0)-\tau,\mathcal{F}_t\rbrace$ is a martingale difference sequence for all $\tau \in \mathcal{T}$.\\
\quad \quad 2. $m_j(I_{t-1},\theta(\tau))$ is non-decreasing in $\tau$ a.s.\\
\quad \quad 3. The family of distributions functions $\lbrace F_b, b \in \Re^d \rbrace$ has Lebesque measures $\lbrace f_b, b \in \Re^d \rbrace$ that are uniformly bounded away from zero for the quantiles of interest.\\
Assumption H: Class of functions\\ For each general $\theta_1 \in \mathcal{B}$,\\
\quad \quad 1. There exists a vector of functions $g_{t-1}: \Theta\rightarrow \Re^q$ such that $g_{t-1}(\theta_1(\tau))$ is $\mathcal{F}_{t-1}$-measurable for each $t \in \mathcal{Z}$, and satisfies for all $k<\infty$,
\\
\quad \quad 2. For all sufficiently small $\delta>0$,
\\
\quad \quad 3. Uniformly in $\tau \in \mathcal{T}$, $E \vert g_{t-1}(\theta_1(\tau)) \vert^2 <\infty$, and uniformly in $(\xi, \tau) \in \mathcal{Y} \times \mathcal{T}$,
\\
Assumption I: Compactness of the parameter space\\ The parametric space $\Theta$ is compact in $\Re^p$. The true parameter $\theta(\tau)$ belongs to the interior of $\Theta$ for each $\tau \in \mathcal{T}$ and $\theta \in \mathcal{B}$. The class $\mathcal{B}$ satisfies\\
\quad \quad \quad \quad \quad \quad \quad \quad $\int_0^{\infty} \big( log(N_{[\cdot]}\delta^2, \mathcal{B}, \Vert \cdot \Vert_{\mathcal{B}}) \big)^{\frac{1}{2}}d\delta < \infty$.\\
Assumption J: Estimator Consistency\\ The estimator $\hat{\theta}_T$ satisfies that $P(\hat{\theta}_T \in \mathcal{B})\to 1$ as $T \to \infty$, and the following asymptotic expansion under $H_0$, \\
\quad \quad \quad \quad $Q_T(\tau)=\sqrt{T}(\hat{\theta}_T(\tau)-\theta(\tau))$\\
\quad \quad \quad \quad \quad \quad \quad $=\frac{M^{-1}}{q(\tau)}T^{-\frac{1}{2}} \sum_{t=1}^T (\mathbbm{1}(y_t-m_j(I_{t-1},\theta(\tau))\leq 0)-\tau)I_{t-1}+o_p(1)$, \quad uniformly in $\tau \in \mathcal{T}$,\\
where $M=E(\textbf{I'I})$ is a positive definite matrix, and $q(\tau)=f_{\epsilon}(F_{\epsilon}^{-1}(\tau))$ is the reciprocal of the sparsity function. Furthermore, the process $Q_T(\tau)$ converges weakly to a zero mean Gaussian process $Q(\tau)$ with covariance function
\\ where $ L(\theta(\tau_1), \theta(\tau_2)) =\lim_{T\rightarrow \infty} T^{-1} \sum_{t=1}^T \sum_{s=1}^T E \Big[ (\mathbbm{1}(y_t-m_j(I_{t-1},\theta(\tau_1))\leq 0)-\tau_1)I_{t-1} \times (\mathbbm{1}(y_s-m_j(I_{s-1},\theta(\tau_2))\leq 0)-\tau_2)I_{s-1} \Big]$.\\
Assumption G1 is standard in time series model checks and is always true in the present context where the information set contains all relevant past history, while G3 is necessary for the asymptotic tightness of the process $S_{j,T}(\xi, \theta_T)$. Assumption H is satisfied for the Linear Quantile Autoregression model under consideration. Sufficient conditions for Assumption I of monotone classes of functions applying to the QAR model can be found in Theorem 2.7.5 in vaart_weak_2000. Meanwhile, the asymptotic normality of the quantile regression process has been established in the literature under a variety of conditions, see for example Theorem 1 in gutenbrunner_regression_1992. \\
In this subsection we establish the limit distribution of the quantile-marked empirical process $S_{j,T}(\xi,\hat{\theta}_T(\tau))$ under the null hypothesis $H_{0,j}$.\\
We begin by showing that the factor estimation error in the feasible information vector is negligible and therefore the proposed statistics converge to the equivalent statistics with the infeasible information vector that includes the true latent factors and a term which converges to zero.\\
Lemma 1: Let Assumptions A-F (shown in Appendix) hold. Then:\\
\\
It is immediate to see that, in the first instance when dealing with $S_{1,T}(\xi,\hat{\theta}_T(\tau))$ factor estimation error is only present in the conditioning set and thus only appears in the weighting exponential function, while in the test statistic $S_{2,T}(\xi,\hat{\theta}_T(\tau))$, the factor estimation error is present in both the exponential function and in the indicator function element. Therefore the proof of the statement for $H_{0,1}$ follows from the proof of statement for $H_{0,2}$ show in the Appendix.
We then proceed to state the limiting distribution of the test statistic recognising that in addition to factor estimation error, which has been found to be negligible, there is also parameter estimation error present. Define the function $G(\xi, \theta(\tau))= E[g_{t-1}(\theta(\tau))f_{I_{t-1}}(m_j(I_{t-1},\theta(\tau))exp(\xi' \Phi(I_{t-1}))]$, $\xi \in \Upsilon$, $\tau \in \mathcal{T}$. Also note that under a suitable central limit theorem $S_{j,T}(\xi,\theta_T(\tau))$ converges to a zero mean Gaussian process $S_{j,\infty}(\xi,\theta(\tau))$ with covariance function given by $Cov_{\infty}(\nu_1, \nu_2)=(\min \lbrace \tau_1,\tau_2 \rbrace-\tau_1\tau_2) E \big[exp((\xi_1-\xi_2)'\Phi(I_0)) \big]$.\\
Theorem 1: Let Assumptions A-J hold. Then, under the null hypothesis $H_{0,j}$, $S_{j,T}(\xi,\hat{\theta}_T(\tau))\xrightarrow{d} \sup\limits_{\xi \in \Upsilon,\tau \in \mathcal{T}} \vert S_{j}(\xi,\theta(\tau)) \vert$, where $S_{j}(\xi,\theta(\tau)) $ is a zero mean Gaussian process with covariance function
where, $\nu_1=(\xi_1',\tau_1)'$, $\nu_2=(\xi_2',\tau_2)'$.\\
Details regarding the convergence of the statistic can be found in the Appendix, while the consistency properties of tests based on continuous functionals has been shown in escanciano_specification_2010.\\
Under the null hypothesis, the quantile error $S_{j,T}(\xi,\theta(\tau))$ converges to a Gaussian process with zero mean and a given covariance structure. However, when the estimated parameter $\hat{\theta}_T$ is used in $S_T(\xi,\hat{\theta}_T)$ the parameter estimation error affects its asymptotic properties. Given that the asymptotic null distribution of $S_{j,T}(\xi,\hat{\theta}_T)$ will be dependent on the data generating process and thus is not nuisance parameter free, critical values for the test statistics cannot be tabulated for general cases.\\
Bootstrap methods have been proposed in the literature for quantile regression models (see, e.g., \citet*{hahn_bootstrapping_1995, parzen_resampling_1994}). With respect to quantile regression models with time series, gregory_smooth_2018 established a smooth tapered block bootstrap procedure. Their work studied the properties of the block bootstrap with smoothing of both data observations via kernel smoothing techniques and data blocks by tapering, which demonstrated the validity of the block bootstrap in dynamic quantile models. At the same time their work extended the validity of the moving block bootstrap fitzenberger_moving_1998 to quantile regression under weaker conditions than previously considered.\\
In our context, under both the null hypotheses, $[\mathbbm{1}(y_t-m_j(I_{t-1},\theta(\tau) \leq 0)-\tau]exp(\xi'\Phi(\hat{I}_{t-1})$ is a martingale difference sequence, therefore resampling blocks of length one, as in the iid case preserves the first order validity of the block bootstrap corradi_information_2009. In order to achieve higher order refinements, the block bootstrap with an increasing block size would have been necessary, given however that our statistics depend on the nuisance parameters, $\xi$, that are not identified under the null we cannot obtain such refinements. Therefore, in order to preserve the temporal ordering, we proceed to jointly resample ($y_t, Y_{t-1}, \hat{F}_{t-1}, exp(\sum_{j=1}^{t-1} \xi'_j \Phi(I_{t-j}))$) by drawing $T-1$ independent draws. For each bootstrap replication we use the same set of resampled values across $\xi \in \Xi$. \\
The bootstrap analogues of $S_{1,T}(\xi,\hat{\theta}_T(\tau))$ and $S_{2,T}(\xi,\hat{\theta}_T(\tau))$, say $S^*_{1,T}(\xi,\hat{\theta}^*_T(\tau))$ and $S^*_{2,T}(\xi,\hat{\theta}^*_T(\tau))$ respectively, are then defined to be
\\
Consequently, one could then obtain the corresponding bootstrap functionals
For any bootstrap replication we compute the bootstrap functional $CvM^*_{j,T}$($KS^*_{j,T}$). Performing then B bootstrap replications, with B large, we compute the quantiles of the empirical distribution of the B bootstrap statistics. The null hypothesis $H_{0,j}$ is rejected if $CvM_{j,T}$($KS_{j,T}$) based on the original sample is greater than the $(1-\alpha)^{th}$ percentile of the corresponding bootstrap distribution, where $\alpha$ is the level of significance. This is because $CvM_{j,T}$($KS_{j,T}$) has the same limiting distributions as its corresponding bootstrapped statistics, which ensures an asymptotic size equal to $\alpha$, while under the alternative $CvM_{j,T}$($KS_{j,T}$) diverges to infinity while the corresponding bootstrap statistics maintain its well defined distribution, ensuring asymptotic power.
We consider two simulation cases to asses the finite sample performance of the proposed test statistics. The data generating process follows the following autoregressive process:\\
where $\upsilon_t$ are independent Uniform $(0,1) $random variables . Motivated by our subsequent empirical application, the lag is taken as $p=1$ and the varying coefficients are defined as follows:\\
where the intercept coefficient $\rho_0(\upsilon)$ is from a normal distribution while the slope coefficients are constant. \\
As it is evident under Case 1 the null hypothesis $H_{0,1} $ is satisfied while under Case 2 the null hypothesis $H_{0,1}$ is violated but $H_{0,2}$ is satisfied. We consider multiple sample sizes, including $T=100$, $T=300$, $T=500$ and $T=1000$ , carry out $1000$ Monte Carlo Simulations, and for each of them perform $300$ bootstrap replications. In all the replications the nominal probability of rejecting a correct null hypothesis is 0.05. The conclusions with other nominal values are similar.\\
Following the relevant literature, we choose the exponential function instead of the indicator function as it has been confirmed that exponential-based tests have higher performance than indicator-based tests. Furthermore, following the earlier relevant consistent test literature (see, e.g. bierens_consistent_1982, bierens_consistent_1990), we choose $\Phi$ as the $arctan$ function. In the experiment we consider a grid of $\mathcal{T}$ in $m=17$ equidistributed points from $\omega=0.1$ to $1- \omega=0.9$. Denote by $\mathcal{T}_m=\lbrace \tau_q \rbrace_{q=1}^m$ the point in the grid, with $w=\tau_1<\cdots<\tau_m=1-\omega$. Also, as already mentioned we wish to weigh more heavily the more recent lags. Let $\Psi_{\tau}(\epsilon_t)=\mathbbm{1}(y_t-\rho_0(\tau)-\rho_1(\tau)y_{t-1}\leq 0)-\tau$ and let $exp(\xi'_Z \Phi(Z_t))=exp(\sum_{j=1}^d \xi'_j \Phi(Z_{t-j}))$, where $d=d(t)=\min \lbrace t-1,c \rbrace$ with $c<\infty$ and $\Xi={\xi_j: a_j\leq \xi_j \leq \gamma_j, j=1,2; \vert a_j \vert, \vert \gamma_j \vert \leq \Gamma j^{-\kappa}, \kappa \geq 2}$. It is immediate to see therefore that the weight attached to past observations decreases over time. We define $\Gamma$ over a grid, $\Gamma \in [0,3]$ and evaluate $g=30$ equidistributed points along the grid. Therefore the $CvM$ and $KS$ statistics are computed as:\\
\\
\\ The theory allows for $m \rightarrow \infty$ as $T \rightarrow \infty$ and the $\lbrace \tau_q \rbrace_{q=1}^m$ generated independently from a distribution on $\mathcal{T}$, but for simplicity in the computations we assume $m$ fixed and $\lbrace \tau_q \rbrace_{q=1}^m$ deterministic throughout the remainder of this paper.\\
In Table 1 we report the rejection frequencies of the test based on the two continuous functionals, the Kolmogorov-Smirnov and the Cramer-von-Mises, for the model under the two data generating processes. The empirical size, though smaller than the nominal $5\%$ is satisfactory for the $CvM$ functional even with a low number of observations, in contrast with the $KS$ functional which remains severely undersized. With regards to the empirical power, we observe that the $CvM$ has the highest rejection frequencies, however both functionals capture the misspecification under the DGP of Case 2 relatively well, when the null hypothesis is that of a standard quantile autoregressive model with no additional factors (i.e. $H_{0,1}$). This is in line with previous empirical results which have shown that the $CvM$ functional outperforms the $KS$ in terms of power escanciano_specification_2010. Furthermore, these rejection rates increase with both the sample size and the strength of the alternative hypothesis imposed (i.e. with higher values of $\beta_1$).\\
The limited simulation study taken suggests that even with relatively small sample sizes (which might be a common scenario in empirical applications) the test exhibits fairly good size accuracy and power performance, particularly when using the CvM functional. As a result, in the subsequent empirical application we will be basing our analysis only on the Cramer-von-Mises functional\footnote{It is worth nothing, that one could also employ the Kupiec functional as well.}. \\
GDP growth and inflation are two of the most important variables due to their dominant role in many macroeconomic models (levin_is_2002,angeloni_new_2006). In the empirical literature, the majority of the work has examined how univariate regression can be improved by augmenting such models with factors as a way to summarise large amounts of information in estimates and forecasts of the conditional mean of these variables (stock_forecasting_2002, bernanke_measuring_2005). Nevertheless, such point forecasts ignore the risks around the central forecast. For example, in the case of growth a central forecast may paint an overly optimistic picture of the state of the economy, which is why the policymakers' focus on downside risk has increased in recent years adrian_vulnerable_2019. Similarly, as argued by, inter alia, henry_is_2004, the dynamic behaviour of inflation in particular has a number of economic implications. \\
\enquote{Fan charts} have been extensively used by the Bank of England in its Inflation Reports, to describe its best provision of future inflation to the general public since 1997, however, more recently, a number of inflation-targeting central banks have started to publish both GDP growth and inflation distributions. Examining the dynamics and higher moments of inflation and growth is of extreme importance and the ability to test whether these conditional quantile models should be augmented with latent factors representing a larger information set (e.g. financial conditions) is necessary.
In this section, in an illustrative empirical analysis, we aim to examine whether the quantiles of the GDP growth rate and the CPI inflation rate are best modelled as univariate regressions or factor augmented models. We consider for the estimation of factors, data series containing data on inflation, real activity and indicators of money and key asset prices for the United Kingdom. We will employ the specification test using a sample which includes quarterly data from 177 macroeconomic variables, including the annual inflation and GDP growth rates, spanning from the second quarter of 1991 to the second quarter of 2018, with a total of $T=109$ observations. According to the criteria outlined in bai_determining_2002 the number of common latent factors present in the data is one and for both of the dependent variables in consideration, the autoregressive model of order 1 has been deemed appropriate by the Schwarz criterion outlined in machado_robust_1993.\\
We entertain for the GDP growth rate and the CPI inflation rate, both a Linear QAR model of order 1, as in ((ref)), and a factor augmented QAR of order 1, with parameters estimated by quantile autoregression. In order to test $H_{0,1}: \quad E[\mathbbm{1}(y_t-\theta_0(\tau)-y_{t-1}\theta_1(\tau)) \leq 0)-\tau|I_{t-1}]=0$, a.s. versus $H_{A,1}: \quad Pr\{E[\mathbbm{1}(y_t-\theta_0(\tau)-Y_{t-1}\theta_1(\tau)) \leq 0)-\tau|I_{t-1}]=0\}$, we construct the $S_{1,T}$ statistics as defined in ((ref)), setting $I_{t-1}=[Y_{t-j},F_{1,t-j}]'$. Similarly in order to test $H_{0,2}: \quad E[\mathbbm{1}(y_t-\theta_0(\tau)-Y_{t-1}\theta_1(\tau)-F_{t-1}\theta_2(\tau)) -\leq 0)-\tau|I_{t-1}]=0$, a.s. versus $H_{A,1}: \quad Pr\{E[\mathbbm{1}(y_t-\theta_0(\tau)-y_{t-1}\theta_1(\tau)-f_{t-1}\theta_2(\tau)) \leq 0)-\tau|I_{t-1}]=0\}$, we construct the $S_{2,T}$ statistic. We use the exponential function, and set $\Phi$ as the inverse tangent function, while setting $\xi_j=\Gamma(j+1)^{-2}$, where $\Gamma$ is defined over a fine grid, $\Gamma=$ ( $\Gamma_1 \atop \Gamma_2$)$\in [0,3] \times [0,3]$. Similarly with the simulation setup we take $m$ equidistributed points $\lbrace \tau_g \rbrace_{g=1}^m$ from $\tau_1=0.1$ to $\tau_m=0.9$, but perform the test for multiple choices of $m$. In Table (ref) we report the p-values for the $CvM$ statistic obtained with the iid bootstrap with the number of bootstrap replications set to 500.\\
From the results in Table (ref), we can conclude that the $QAR(1)$ model of GDP growth is misspecified at $1\%$ nominal level, conditional on an information set that includes growth lags and the estimated factor, for all the number of quantiles estimated . We see that the omnibus test based on the $CvM_{1,T}$ strongly rejects this model and the rejection appears stronger when the grid over $\mathcal{T}$ becomes finer, a result that is consistent with the fact that the power of the test improves as $m \rightarrow \infty$. Our expectations that the factor augmented model should be an omitted variable in the QAR model is verified with the corresponding p-values, where we see that the misspecification is eliminated as we fail to reject the null hypothesis $H_{0,2}$ (the FA-QAR columns), which indicates that the factor-augmented QAR is correctly specified. On the other hand, we see that the $QAR(1)$ model of CPI Inflation is correctly specified and in fact the estimated latent factor is not an omitted variable and does not carry relevant information. This result is consistent with the literature supporting that past inflation is the most important determinant and thus best predictor for future inflation. \\
Based on these results we clearly see that estimated latent factors may often carry information that is relevant for the estimation of variables of interest and having a test which enables us to test if they are an omitted variable, is of critical importance in empirical analysis. We further complement the testing procedure by estimating multiple quantiles of growth and inflation, employing the Factor Augmented Quantile Autoregression and the Quantile Autoregression respectively, in an attempt to trace out their distributions. In addition we examine the impact of the independent variables of our regression on the dependent variable.\\
Figure (ref) shows the scatter plots of one-quarter-ahead CPI inflation against the current inflation rate and one-quarter-ahead GDP growth against the current growth rate and the current latent factor, along with the corresponding univariate regressions for the tenth, fiftieth and ninetieth quantile and the least squares regression line. The slopes for the CPI inflation lag appear relatively stable across the quantiles and so do the slopes of the GDP growth lag. Interestingly, the slopes of the latent factor differ significantly across the quantiles, implying that its impact is distinctively different across the spectrum of the growth distribution. Furthermore we see that for all the variables concerned the mean and median regression appear identical which one might consider as an indication of symmetry in the distribution of the dependent variables.
Indeed, Figure (ref) reveals that the autoregressive coefficient of GDP growth changes slightly across the evaluated quantiles and appears to become weaker across the upper tail of the distribution. This is not the case however for the latent factor coefficient. The current latent factor seems to be statistically not significant in the lower tails of the distribution, but its impact increases on the middle and towards upper tail of the distribution, as the coefficient becomes more negative. This implies that the latent factor is not uniformly informative for predicting tail outcomes. For lower quantiles of the GDP growth the lag coefficient is the sole determinant, however, as we start to consider the upper tail of the distribution, the latent factor seems to provide significant information. In the case of inflation, the autoregressive coefficient is statistically different from zero across all the estimated quantiles. The coefficient has a bigger impact on higher quantiles and in the extreme upper tail approaches a unit root process (see Figure (ref)).\\
Furthermore, in figure (ref) we have traced out the tenth, fiftieth and ninetieth conditional quantiles across the sample time span. The figure demonstrates three important empirical facts. On the one hand, if we focus on the periods before and around the 2008 economic crisis we can see an asymmetry in the distribution of GDP growth as the difference between the ninetieth quantile and the median is significantly smaller than the difference between the median and tenth quantile. This asymmetry however seems to dissipate following the recent financial crisis up until the end of the sample. Furthermore, the distribution of GDP growth has remained fairly stable across the sample, with the exception of the 2008 recession where GDP had experienced a significant decline. The distribution of CPI Inflation on the other hand has experienced smaller fluctuations throughout the sample and demonstrates a more symmetric behaviour. It is noticeable however, that the inflation distribution had narrowed significantly around 2015 when the interest rate was close to the zero lower bound.\\
Based on these indications, we finally proceed to smooth the estimated quantile distributions each quarter by interpolating between the estimated quantiles using the skewed $t-distribution$ by azzalini_distributions_2003, characterised by four parameters that pin down the location, $\mu$, scale,$\sigma$, fatness, $\nu$ and shape, $\alpha$. We fit the skewed $t-distribution$ in order to smooth the quantile function and recover a probability density function:
\\ where, $t(\cdot)$ and $T(\cdot)$ respectively denote the PDF and CDF of the student $t-$distribution. For each quarter we choose the four parameters $\lbrace \mu, \sigma,\nu,\alpha \rbrace$ of the skewed $t-distribution$ $f$ to minimise the squared distance between our estimated quantile function $m_j(I_{t-1}, \hat{\theta}(\tau))$ and the quantile function of the skewed $t-distribution$ $F^{-1}(\tau;\mu_t, \sigma_t,\nu_t,\alpha_t)$ to match the $5,25,75,95$ percent quantiles.This approach, first presented in adrian_vulnerable_2019, is computationally less burdensome while also making fewer parametric assumptions compared to alternative ways of estimating conditional predictive distributions hamilton_new_1989, smith_asymmetric_2016 and it allows us to obtain an estimated conditional distribution of the GDP growth and CPI inflation, plotted in Figures (ref) $\&$ (ref) . \\
As it is evident in Figure (ref) the entire distribution and not just the central tendency of the GDP growth has evolved over time. For example the 2008 recession was associated with a rather symmetric distribution, albeit one with a significantly lower mean, while the antecedent expansion was associated with left skewed distributions. Furthermore the right tail of the distribution appears to be more stable than the median and lower tails which exhibit a stronger time variation. The distinctive behaviour in the tails of the conditional distribution indicates that when it comes to growth downside risk varies much more strongly than upside risk over time, a fact that has also been found to be true for the US by adrian_vulnerable_2019. The distribution of inflation on the other hand, with the exception of some outliers at the beginning of the sample, has remained fairly more stable over time. The fluctuations in the central tendency have been significantly smaller than the fluctuations in the tails of the distribution as well as the kyrtosis. Furthermore, we see that the left tail of the distribution has remained fairly stable, a fact associated with the zero lower bound of nominal interest rates, while the median and the right tail in particular exhibit stronger time series variation. We see therefore, that upside risk for inflation varies much more strongly and was a key ingredient during the 2008 financial crisis and for the following years, where as we see the right tail of the distribution had significantly shifted to the right. This in fact supports the thought that, though inflation had on average remain unaltered during the crisis, the inflation distribution was still affected by the volatility in the macro-economy in a multitude of ways, in terms of location, scale, fatness and shape. The significant variation in the distributions of both variables of interest demonstrate the importance of examining higher moments and dynamics rather than simple point forecasts which ignore the risks surrounding that central forecast.\\
Econometric modelling often requires the specification of conditional quantile models for a range of quantiles of the conditional distribution. However, such models may often rely on unobservable variables that require estimation prior to modelling. For the evaluation therefore of quantile autoregressive models, we propose a test for the joint hypothesis of correct dynamic specification and no omitted latent variable with \enquote{valid} asymptotic critical values obtained via a bootstrap procedure based on the entire history of available information. We have demonstrated in a simulation study the consistency of the proposed test and its high power. An illustrative empirical implementation of the test suggests that the modelling of the conditional quantiles of UK GDP growth can be improved with the inclusion of estimated latent factors characterising the economy, while the CPI inflation rate does not require such additional information. Furthermore an empirical analysis of the conditional dependence structure of both variables demonstrate that downside risk for growth exhibits higher fluctuations, while in the case of inflation upside risk becomes more relevant.\\