Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
94,370 characters · 17 sections · 83 citation commands
Pair copula constructions of point-optimal sign-based tests for predictive linear and nonlinear regressions
Keywords: Predictive regressions, D-vine, truncation, sequential estimation, persistent regressors, sign test, exact inference, adaptive method, power envelope.
JEL Codes: C12, C13, C15, C22, C52
{\hskip 1.5em} Predictive regressions are frequently encountered in the economics and finance literature. Within this framework, the regressors are often highly persistent (and potentially nonstationary) with errors that exhibit contemporaneous correlation with the disturbances in the predictive regression. This leads to endogeneity, as a result of which the least-squared estimator of the regression parameters is biased. In such settings, t-type tests possess a non-standard distribution and inference using asymptotic critical values is no longer valid [see mankiw1986we and stambaugh1999predictive among others]. The econometric analysis of predictive regressions has been addressed extensively and numerous papers have suggested solutions to overcome this problem. These include reducing the bias in finite samples [see nelson1993predictable and stambaugh1999predictive among others] or self generated instrumental variables (IVs) that eliminate endogeneity [see magdalinos2009limit, phillips2013predictive among others]. However, many of these studies impose strict assumptions on the degree of persistency of the regressors [see phillips2014confidence for an overview]. In this paper, we propose pair copula constructed point-optimal sign-based tests (PCC-POS-based tests hereafter) in the context of linear and nonlinear predictive regressions. The proposed tests are robust in the presence of persistent/endogenous regressors and errors, heterogeneous and persistent volatility, and disturbances that exhibit serial (nonlinear) dependence. Moreover, they are exact, consider the entire dependence structure of the signs and may be inverted to build confidence regions for the parameters of the regression function.
Sign-based tests, such as those proposed by dufour1995exact, campbell1997exact, luger2003exact, and dufour2010exact are randomized tests with a randomized distribution under the null hypothesis of unpredictability [see pratt2012concepts for a review of randomized tests]. Hence, under mild assumptions these procedures are distribution-free and do not suffer from the issues encountered by t-type statistics in finite samples. These class of tests are valid in the presence of non-normal distributions and heteroskedasiticty of unknown form [see boldin1997sign and taamouti2015finite for a review of sign-based tests]. Furthermore, dufour2010exact show that the heteroskedasticity and autocorrelation corrected tests developed by white1980heteroskedasticity (more commonly referred to as \textquotedblleft HAC\textquotedblrightprocedures) are plagued with low power when the errors follow GARCH-type structures or there is a break in variance. To address these issues, dufour2010exact propose point-optimal sign-based (POS-based) inference to test whether the conditional median of a response variable is zero against a linear regression alternative, where these procedures are further extended to nonlinear models.
In an earlier paper, we proposed an extension of the POS-based tests within a predictive regression framework. However, in order to obtain feasible test statistics, we imposed a Markovian assumption on the sign process. This paper relaxes the Markovian assumption on the signs, by decomposing the joint distribution of the signs using models proposed by panagiotelis2012pair for multivariate discrete data based on pair copula constructions. The latter allow us to build feasible test statistics that are robust against heavy-tailed and asymmetric distributions, provided that the errors have zero median conditional on their own past and the explanatory variables. The PCC-POS-based tests are shown to be robust against non-standard distributions and heteroskedasticity of unknown form, and have the highest power among certain parametric and nonparametric tests that are intended to be robust against heteroskedasticity. Moreover, as in dufour2010exact, they can be inverted to produce a confidence region for the vector (sub-vector) of the parameters.
Although, the literature surrounding sign-based and sign-ranked inference is vast, the focus of the POS-based tests constructed by dufour2010exact is to maximize power at a nominated point in the alternative parameter space, such that the power of the POS-based test is close to that of the power envelope - i.e. maximum attainable power for a given testing problem [see king1987towards]. Similarly, the PCC-POS-based tests are Neyman-Pearson type tests based on the signs, and as in dufour2010exact a practical problem concerns finding an alternative at which the power of the PCC-POS-based tests is close to that of the power envelope. By conducting simulations exercises, dufour2010exact find that the power of the POS-based tests is shifted close to the power envelope, when 10% of the sample is used to estimate the alternative and the remaining portion to calculate the test-statistic. Our simulations using a variety of split-sample PCC-POS-based tests confirm these findings.
Due to the nonlinear nature of the signs, there is inherent uncertainty regarding the structure of sign dependence. Therefore, it is important to consider the entire dependence structure of the signs. One approach for computing the joint distribution of the signs $s(y_1),\cdots,s(y_T)$, where $s(y_i)=\mathbbm{1}_{\mathbb{R}^+ \cup\{0\} }\{y_i\}$, entails taking advantage of copula functions [see sklar1959fonctions], which express the joint distribution of the signs in terms of i) the marginal distributions of the individual signs; and ii) the copula models capturing the dependence of the $T$ signs. As the signs are discrete, the likelihood function of the POS-based tests under the alternative hypothesis can then be calculated using rectangle probabilities and in turn estimated using copulas with closed analytical form. However, this approach would not yield feasible test statistics, as the number of multivariate copulas that need to be evaluated increase at an exponential rate with growing sample sizes $T$. As a result of this curse of dimensionality, the literature concerning calculating probability mass functions (p.m.f hereafter) using discrete data is limited to low-dimensional data and copulas that are fast to calculate [see nikoloulopoulos2008multivariate,nikoloulopoulos2009finite and li2010two].
To propose feasible POS-based test statistics within a predictive regression framework, we use a discrete analogue of the vine PCCs introduced by panagiotelis2012pair. The likelihood function of the signs under the alternative hypothesis can then be decomposed as a vine PCC under certain set of conditions that are later outlined in the paper. The most important advantage of this method is that for a sample of size $T$, only $2T(T-1)$ bivariate copula evaluations are required, as opposed to $2^T$ multivariate copula evaluations using rectangle probabilities. Another advantage of the vine PCC methodology is that model selection techniques can be used to identify the conditional independence in the process of signs in order to create more parsimonious PCC models.
The rest of the paper is organized as follows: in Section (ref), we motivate the use of the discrete analogue of the vine PCC for building POS-based tests. In Section (ref), we outline the conditions under which vine PCCs can be implemented and we also discuss the choice of the PCC model. We then propose PCC-POS-based test statistics for linear and nonlinear models. Secion (ref), discusses the estimation approach implemented for the vine PCC models. In Section (ref), we derive the power envelope and highlight the choice of the alternative hypothesis for computing the PCC-POS-based test statistic. In Section (ref), we discuss the problem of finding a confidence set for a vector (subvector) of parameters using the projection techniques. In Section (ref), we assess the performance of the proposed tests in terms of size and power. Finally, in Section (ref), we conclude the findings of the paper.
{\hskip 1.5em}Consider a stochastic process $Z=\{Z_t=(y_t,\bm{x}_{t-1}'):\Omega\rightarrow\mathbb{R}^{(k+1)},t=1,2,\cdots\}$ defined on a probability space $(\Omega,\mathcal{F},P)$. Within a framework of a predictive regression $y_{t}$ can linearly be explained by a vector variable $\bm{x}_{t-1}$
where $y_t$ is the dependent variable and $\bm{x}_{t-1}$ is an $(k+1)\times 1$ vector of stochastic explanatory variables, say $\bm{x}_{t-1}=[1,x_{1,t-1},\cdots,x_{k,t-1}]'$, $\bm{\beta} \in \mathbb{R}^{(k+1)}$ is an unknown vector of parameters with $\bm{\beta}=[\beta_0,\beta_1,\cdots,\beta_k]'$ and \[ \varepsilon_t\mid X\sim F_t(.\mid X) \] where $F_{t}(.\mid X)$ is an unknown conditional distribution function and X=$[\bm{x}_0',\cdots,\bm{x}_{T-1}']'$ is an $T\times (k+1)$ information matrix.
We follow coudin2009finite by considering the median as an alternative measure of central tendency, which differs from the Martingale Difference Sequence (MDS hereafter) assumption. In the latter, it is generally assumed that for an adapted stochastic sequence $\{Z_t,\mathcal{F}_{t}, t=1,2,\cdots\}$, where $\mathcal{F}_{t}$ is a $\sigma$-field in $\Omega$, $\mathcal{F}_s\subseteq\mathcal{F}_t$ for $s< t$, $\mathbb{E}\{\varepsilon_t\mid\mathcal{F}_{t-1}\}=0$, $\forall t\geq 1$. Thus the alternative implies imposing a median-based analogue of the MDS on the error process - namely we suppose that $\varepsilon_t$ is a strict conditional mediangale
with \[ \bm{\varepsilon}_{0}=\{\emptyset\},\quad\bm{\varepsilon}_{t-1}=\{\varepsilon_1,\cdots,\varepsilon_{t-1}\},\quad\text{for}\quad t\geq2. \]
Note ((ref)) entails that $\varepsilon _{t}\mid X$ has no mass at zero for all $t$, which is only true if $\varepsilon_{t}\mid X$\ is a continuous variable. Model ((ref)) in conjunction with assumption ((ref)) allows the error terms to possess asymmetric, heteroskedastic and serially dependent distributions, so long as the conditional medians are zero. Assumption ((ref)) allows for many dependent schemes, such as those of the form $\varepsilon_1=\sigma_1(x_1,\cdots,x_{t-2})\epsilon_1$, $\varepsilon_t=\sigma_1(x_1,\cdots,x_{t-2},\varepsilon_1,\cdots,\varepsilon_{t-1})\epsilon_t$ ,$t=2,\cdots,n$, where $\epsilon_1,\cdots,\epsilon_n$ are independent with a zero median. In time-series context this includes models such as ARCH, GARCH or stochastic volatility with non-Gaussian errors. Furthermore, in the mediangale framework the disturbances need not be second order stationary.
Suppose, we wish to test the null hypothesis of unpredictability
against the alternative
where $\bm{0}$ is a $(k+1)\times 1$ zero vector. Define the vector of signs as follows
where for $t=1,\cdots,T$
We consider Neyman-Pearson type test based on the signs. Thus, to build POS-based tests for testing the null hypothesis ((ref)) against the alternative ((ref)), we first define the likelihood function of the sample in terms of signs $s(y_1),\cdots,s(y_T)$ conditional on $X$
with
and \[ P[s(y_1)=s_1\mid\text{\b{S}}_{0}=\text{\b{s}}_{0},X]=P[s(y_1)=s_1\mid X], \] where each $s_t$ for $1\leq t\leq T$ takes two possible values of 0 and 1. Under model ((ref)) and assumption ((ref)), the variables $s(\varepsilon_1),\cdots,s(\varepsilon_T)$ and in turn $s(y_1),\cdots,s(y_T)$ are i.i.d conditional on $X$, according to the distribution \[ P[s(\varepsilon_1)=1\mid X]=P[s(\varepsilon_1)=0\mid X]=\frac{1}{2},\quad t=1,\cdots,T. \] This results holds true iff for any combination of $t=1,\cdots,T$ there is a permutation $\pi: i\rightarrow j$ such that the mediangale assumption holds for $j$. Then the signs $s(\varepsilon_1),\cdots,s(\varepsilon_T)$ are i.i.d [see Theorem (ref) of coudin2009finite]. Therefore, under the null hypothesis of unpredictability we have
Consequently, under the null hypothesis of orthogonality, the log-likelihood function conditional on $X$ is given by \[ L_0(U(T),\bm{0},X)=\prod\limits_{t=1}^{T}P[s(y_t)=s_t\mid X]=\left(\frac{1}{2}\right)^T. \] On the other hand, under the alternative we have \[ L_1(U(T),\bm{\beta}_1,X)=\prod\limits_{t=1}^{T}P[s(y_t)=s_t\mid \text{\b{S}}_{t-1}=\text{\b{s}}_{t-1}, X], \] where now for $t=1,\cdots,T$ \[ y_t=\bm{\beta}_1'\bm{x}_{t-1}+\varepsilon_t. \]
In an earlier paper, we considered optimal sign-based tests (in the Neyman-Pearson sense), which maximize power under the constraint $P[\text{Reject } H_0\mid H_0]\leq \alpha$, where $\alpha$ is an arbitrary significance level [see lehmann2006testing]. Let $H_0$ and $H_1$ be defined by ((ref)) and ((ref)) respectively. Then under the assumptions ((ref)) and ((ref)), the log-likelihood ratio
is most powerful for testing $H_0$ against $H_1$ among level $\alpha$ tests based on the signs $(s(y_1),\cdots,s(y_T))'$, where $c$ is the smallest constant such that \[ P[SL_T(\bm{\beta}_1)>c\mid H_0]\leq \alpha, \] and $\alpha$ is an arbitrary significance level.
For POS-based tests within a predictive regression framework, the test statistic requires the calculation of $P[y_t\geq0\mid\text{\b{S}}_{t-1}=\text{\b{s}}_{t-1},X]$ and $P[y_t<0\mid\text{\b{S}}_{t-1}=\text{\b{s}}_{t-1},X]$. The latter is not easy to compute, as it involves the distribution of the joint process of signs $s(y_1),\cdots,s(y_T)$, conditional on $X$, which is unknown. Therefore, to obtain feasible test statistics, we may impose a Markovian assumption on the sign process. However, it may be important to capture the dependence structure of the entire process.
In this paper, we consider the entire dependence structure of the vector of signs by taking advantage of copulas. The Theorem of sklar1959fonctions states that there exists a copula $C$ such that
where $F$ is a conditional joint cumulative distribution function (CDF hereafter) of the signs vector $\text{\b{S}}=(s(y_1),\cdots,s(y_T))'$ with conditional marginal distribution functions $F_j$ for $j=1,2,\cdots,T$. Copula $C(.)$ is unique for continuous variables, but for discrete variables, it is unique only on the set \[ \text{Range}(F_1)\times\cdots\times\text{Range}(F_T), \] which is the Cartesian product of the ranges of the conditional marginal distribution functions. To illustrate an example of non-uniqueness in the discrete case, let us consider a sample of two discrete binary variables, say $s(y_1)$ and $s(y_2)$, with corresponding conditional marginal distribution functions $F_1$ and $F_2$. We know that $F_j\sim\text{Bernoulli}(p_j)$ for $j=1,2$, such that
Thus, $\text{Range}(F_1)=\{0,1-p_1,1\}$ and $\text{Range}(F_2)=\{0,1-p_2,1\}$, with the copula only being unique for $C(1-p_1,1-p_2)$, noting that $C(0,1-p_j)=0$ and $C(1,1-p_j)=1-p_j$ for $j=1,2$. However, this non-uniqueness does not preclude the use of parametric copulas for modelling discrete data [see. joe1997multivariate, song2009joint].
By considering this bivariate example, the conditional p.m.f can be expressed in terms of rectangle probabilities, \begingroup \allowdisplaybreaks
\endgroup and in turn in terms of copulas as follows \begingroup \allowdisplaybreaks
\endgroup which implies that the $T$-variate conditional likelihood function ((ref)) can be expressed in terms of $2^T$ finite differences \begingroup \allowdisplaybreaks
\endgroup
Evidently, the calculation of the conditional likelihood function ((ref)) using this approach would require $2^T$ multi\-variate copula evaluations, which is not computationally feasible. However, by employing the vine PCC introduced later in the paper, we will show that this number can be reduced to only $2T(T-1)$ bivariate copula evaluations. The latter method provides us with flexibility, since any multivariate discrete distribution can be decomposed as a vine PCC under a set of conditions that are discussed in the following Section.
{\hskip 1.5em}In this Section, we derive POS-based tests in the context of linear and nonlinear regression models based on vine PCC decomposition. Following a structure similar to dufour2010exact, we first consider the problem of testing whether the conditional median of a vector of observations is zero against a linear regression alternative. We further consider the conditions under which the conditional likelihood function under the alternative hypothesis can be decomposed as a vine PCC, and as such, choose an appropriate vine model. These results are later generalized to test whether the coefficients of a possibly nonlinear median regression function have a given value against an alternative nonlinear median regression.
{\hskip 1.5em}Consider the problem of testing the null hypothesis of unpredictability ((ref)) against the alternative ((ref)), using the test statistic ((ref)) and given the assumptions ((ref)) and ((ref)). As it was shown in Section (ref), under the alternative hypothesis the conditional likelihood function can be expressed as
Let $s(y_j)$ be a scalar element of $\text{\b{S}}_{t-1}$, with $\text{\b{S}}_{t-1}^{\backslash j}=\text{\b{S}}_{t-1}\backslash s(y_j)$ such that \[ \text{\b{S}}_{t-1}^{\backslash j}=\left\{s(y_1),s(y_2),\cdots,s(y_{j-1}),s(y_{j+1}),\cdots,s(y_{t-1})\right\} \] and $s(y_t)\notin \text{\b{S}}_{t-1}$. By choosing a single element of $\text{\b{S}}_{t-1}$, say $s(y_j)$, we would have \begingroup \allowdisplaybreaks
\endgroup where the bivariate conditional probability in ((ref)) can be expressed in terms of copulas as follows \begingroup \allowdisplaybreaks
\endgroup Further, let $\text{\b{S}}_{t-1}^{\backslash i,j}=\text{\b{S}}_{t-1}^{\backslash j}\backslash s(y_ i)$, such that $s(y_ i)$ is a scalar element of $\text{\b{S}}_{t-1}^{\backslash j}$. Then the arguments $F_{s(y_t)\mid\text{\b{S}}_{t-1}^{\backslash j}}$ and $F_{s(y_j)\mid\text{\b{S}}_{t-1}^{\backslash j}}$ in copula expression ((ref)) can be presented by the general form \begingroup \allowdisplaybreaks
\endgroup Thus, decomposition ((ref)), and in turn ((ref)) can be applied recursively to the elements of the conditional likelihood function ((ref)), such that it is expressed in terms of bivariate copulas. Let $\text{\b{S}}_{t-1}=\{s(y_1),\cdots,s(y_{t-1})\}$ be the variables that $s(y_t)$ for $t=2,\cdots,T$ is conditioned on. We follow joe2014dependence, by letting $\underline{\sigma}_{t-1}=\{\sigma(1,t),\cdots,\sigma(t-1,t)\}$ be a permutation of $\text{\b{S}}_{t-1}$, such that $s(y_t)$ is paired sequentially first with $\sigma(1,t)$, then $\sigma(2,t)$ and finally $\sigma(t-1,t)$, where in the $r^{\text{th}}$ step $(2\leq r\leq t-1)$, $\sigma(r,t)$ is paired to $t$ conditional on $\sigma(1,t),\cdots,\sigma(r-1,t)$. For $n\leq3$ (i.e. $t=2,3$) there are only three possible permutations with $\underline{\sigma}_1=\{s(y_1)\}$ for $t=2$, and $\underline{\sigma}_2=\{s(y_1),s(y_2)\}$, as well as $\underline{\sigma}_2=(s(y_2),s(y_1))$ for $t=3$ respectively. Therefore, under assumptions ((ref)) and ((ref) ), and with $T\leq 3$, let $H_{0}$ and $H_{1}$ be defined by ((ref)) - ((ref)), then the Neyman-Pearson type test-statistic based on the signs $(s(y_1),\cdots,s(y_T))'$ can be expressed as \begingroup \allowdisplaybreaks
\endgroup for $t=2,\cdots,T$, where \begingroup \allowdisplaybreaks
\endgroup and such that \begingroup \allowdisplaybreaks
\endgroup
For $T>3$, the permutations $\underline{\sigma}_{t-1}$ are dependent on the choice of the permutations at stages $3,\cdots,t-1$. Therefore, an issue that requires considerable attention is whether there exists a decomposition such as the one considered in the earlier example for $T>3$. Furthermore, the conditional likelihood function expressed in terms of bivariate copulas by recursively using ((ref)) and ((ref)), assumes that a single copula is specified for each conditional bivariate distribution $F_{s(y_t),s(y_j)\mid \text{\b{S}}_{t-1}^{\backslash j}}$ in decomposition ((ref)) over all possible values of $\text{\b{S}}_{t-1}^{\backslash j}$. This implies that the copula is unique for the Cartesian product of the ranges of conditional CDFs $F_{s(y_t)\mid \text{\b{S}}_{t-1}^{\backslash j}}$ and $F_{s(y_j)\mid \text{\b{S}}_{t-1}^{\backslash j}}$. Therefore, the decomposition must be such that each conditional bivariate distribution in said vine has a constant conditional copula [see panagiotelis2012pair]. For a constant conditional copula to exist, the following conditions outlined by panagiotelis2012pair must be satisfied.
To satisfy the above conditions, the vine decomposition must be such that the strength of the dependence of the conditional bivariate distribution does not vary much across different values of the conditioning set [see panagiotelis2012pair]. As we are dealing with time-series data, the D-vine decomposition yields a constant dependence structure over different values of $\text{\b{S}}_{t-1}^{\backslash j}$, and is thus, the most appropriate and intuitive choice for the decomposition of the conditional likelihood function ((ref)).
The D-vine PCC (figure (ref)) is constructed by $T-1$ trees, say $D=\{\mathcal{T}_1,\cdots,\mathcal{T}_{T-1}\}$, comprised of the edges $\xi(D)=\xi(\mathcal{T}_1)\cup\cdots\cup \xi(\mathcal{T}_{T-1})$, where $\xi(\mathcal{T}_l)$ refers to the edges of the tree $\mathcal{T}_l$. In the first tree $T_1$, the marginals $F(s_1\mid X),F(s_2\mid X),\cdots,F(s_T\mid X)$, are arranged as nodes in consecutive order, say $N(\mathcal{T}_1):=\{1,2,\cdots,T-1,T\}$, where the nodes are of degree two, meaning that no more than two edges is connected to each node. The corresponding edges join the adjacent nodes, such that $\xi(\mathcal{T}_1):=\{12,23,\cdots,(T-1,T)\}$. Next, the edges of the first tree $\xi(\mathcal{T}_1)$ become the nodes of $\mathcal{T}_2$, a process which is completed in a recursive manner, such that $N(\mathcal{T}_{l+1})=\xi(\mathcal{T}_l)$, with the edges of each tree joining the adjacent nodes, and with the mutual elements between the nodes becoming the conditioning set.
To express the conditional likelihood function ((ref)) as a D-vine, we begin by calculating the conditional marginals $F_1,\cdots,F_T$, where $F_t\sim Bernoulli(p_t)$, for $t=1,\cdots,T$, with conditional CDFs that are expressed as ((ref)). Therefore, under assumptions ((ref)) and ((ref)), we have
In the case of continuous variables, say $\{s^*(y_t)\in \mathbb{R},t=1,\cdots,T\}$, the construction of the D-vine involves an iterative copula evaluation process for the trees $\mathcal{T}_1,\cdots,\mathcal{T}_{T-1}$. This leads to $T(T-1)/2$ copula evaluations, which correspond to one copula evaluation for each edge [see Appendix]. On the other hand, for discrete variables, the conditional p.m.fs are expressed as in ((ref)), which requires the evaluation of the following four copulas \begingroup \allowdisplaybreaks
\endgroup where $F_{t\mid \mathbf{\backslash j}}^+=P[s(y_t)\leq s_t \mid\text{\b{S}}_{t-1}^{\backslash j}=\text{\b{s}}_{t-1}^{\backslash j},X]$ and $F_{t\mid \mathbf{\backslash j}}^-=P[s(y_t)\leq s_t-1 \mid\text{\b{S}}_{t-1}^{\backslash j}=\text{\b{s}}_{t-1}^{\backslash j},X]$. Henceforth, $4\times T(T-1)/2$ bivariate copulas need to be evaluated in the case of discrete data.
Let us express the conditional joint p.m.f of the signs as follows
where following the results by stoeber2013simplified, if the D-vine is expressed as a vine-array $A=(\sigma_{lt})_{1\leq l\leq t\leq T}$, where $l=2,\cdots,T-1$ is the row with tree $\mathcal{T}_l$, and column $t$ has the permutation $\underline{\sigma}_{t-1}=(\sigma(1,t),\cdots,\sigma(t-1,t))$ of the previously added variables, such that
then
where following joe2014dependence, the copula densities in expression ((ref)) are calculated by
which leads to the following proposition.
Under the null hypothesis, the signs $s(y_1),\cdots,s(y_T)$ are i.i.d. according to Bernoulli $Bi(1,0.5)$, with the distribution of $SL_T(\bm{\beta}_1)$ only depending on the weights $a_t(\bm{\beta}_1)$, without the presence of any nuisance parameters. Assumption ((ref)) implies that tests based on $SL_{T}(\bm{\beta}_{1})$, such as the test given by ((ref)), are distribution-free and robust against heteroskedasticity of unknown form. On the other hand, under the alternative hypothesis, the power function of the test depends on the form of the distribution of $\varepsilon_t$. A special case is where $\varepsilon_1,\cdots,\varepsilon_T$ are independently distributed according to $N(0,1)$, which leads to the optimal test statistic assuming the following form \[ SL_{T}(\bm{\beta} _{1})=\sum\limits_{t=2}^{T}\sum\limits_{l=t-1}^{2}\ln c_{\sigma_{lt}t,\mid \sigma_{1t},\cdots,\sigma_{t-1,t}}+\sum\limits_{t=2}^{T}\ln c_{\sigma_{1t}t}+\sum\limits_{t=1}^{T}s(y_t)a_t(\bm{\beta}_1)>c_1(\bm{\beta}_1), \] where \[ a_t(\bm{\beta}_1)=\ln\left\{\frac{\Phi(\bm{\beta}' \bm{x}_{t-1})}{1-\Phi(\bm{\beta}' \bm{x}_{t-1})}\right\}, \] where $\Phi(.)$ is the standard normal distribution function. The distribution of $SL_{T}(\bm{\beta}_{1})$ can be simulated under the null hypothesis with sufficient number of replications, and the critical values can be obtained to any degree of precision.
{\hskip 1.5em}We now consider the nonlinear predictive regression model
where $\bm{x}_{t-1}$ is a $(k+1)\times 1$ vector of stochastic explanatory variables, such that $\bm{x}_{t-1}=[1,x_{1,t-1},\cdots,x_{k,t-1}]'$, $f(\,\cdot \,)$ is a scalar function, $\bm{\beta} \in \mathbb{R}^{(k+1)}$ is an unknown vector of parameters and
where as before $F_{t}(.\mid X)$ is a distribution function and $X=[\bm{x}_0',\cdots,\bm{x}_{T-1}']$ is an $T\times (k+1)$ matrix. Suppose that the error process $\{\varepsilon_t,t=1,2,\cdots\}$ is a strict conditional mediangale, such that
with \[ \bm{\varepsilon}_{0}=\{\emptyset\},\quad\bm{\varepsilon}_{t-1}=\{\varepsilon_1,\cdots,\varepsilon_{t-1}\},\quad\text{for}\quad t\geq2 \] and where ((ref)) entails that $\varepsilon_t\mid X$ has no mass at zero, i.e. $P[\varepsilon_t=0\mid X]$=0 for all $t$. We do not require that the parameter vector $\bm{\beta}$ be identified.
We consider the problem of testing the null hypothesis
against the alternative hypothesis
We construct a test statistic for testing $H(\beta_0)$ against $H(\beta_1)$ in a similar manner to Section (ref), by transforming model ((ref)) to \[ \tilde{y}_t=g(\bm{x}_{t-1},\bm{\beta},\bm{\beta}_0)+\varepsilon_t,\quad t=1,\cdots,T \] where $\tilde{y}_t=y_t-f(\bm{x}_{t-1},\bm{\beta}_0)$ and $g(\bm{x}_{t-1},\bm{\beta},\bm{\beta}_0)=f(\bm{x}_{t-1},\bm{\beta})-f(\bm{x}_{t-1},\bm{\beta}_0)$. Notice that testing $H(\bm{\beta}_0)$ against $H(\bm{\beta}_1)$ is equivalent to testing \[ \bar{H}_0:g(\bm{x}_{t-1}, \bm{\beta},\bm{\beta}_0)=\bm{0},\quad\text{for}\quad t=1,\cdots,T \] where $\bm{0}$ is a $(k+1)\times 1$ zero vector, against the alternative \[ \bar{H}_A: g(\bm{x}_{t-1},\bm{\beta},\bm{\beta}_0)=f(\bm{x}_{t-1},\bm{\beta}_1)-f(x_{t-1},\bm{\beta}_0),\quad\text{for}\quad t=1,\cdots,T. \] For $\tilde{U}(T)=(s(\tilde{y}_1),\cdots,s(\tilde{y}_T))'$, where for $1\leq t\leq T$
As before, the conditional joint p.m.f of the process of signs is expressed as
Furthermore, the D-vine-array $\tilde{A}=(\tilde{\sigma}_{lt})_{1\leq l\leq t\leq T}$, is such that $l=2,\cdots,T-1$ is the row with tree $\mathcal{T}_l$, and column $t$ has the permutation $\tilde{\underline{\sigma}}_{t-1}=(\tilde{\sigma}_{1t},\cdots,\tilde{\sigma}_{t-1,t})$ of the previously added variables. Then
which leads to the following corollary.
Consider a linear function $f(\bm{x}_{t-1},\bm{\beta})=\bm{\beta}'\bm{x}_{t-1}$, and assume that under the alternative hypothesis $\varepsilon_t$ for $t=1,\cdots,T$ follows a standard normal distribution (i.e. $\varepsilon_t\sim N(0,1))$. Then the statistic for testing $H(\bm{\beta}_0)$ against the alternative $H(\bm{\beta}_1)$ is given by \begingroup \allowdisplaybreaks
\endgroup where \[ \tilde{a}_t(\bm{\beta}_0\mid\bm{\beta}_1)=\ln\left\{\frac{\Phi((\bm{\beta}_1-\bm{\beta}_0)'\bm{x}_{t-1})}{1-\Phi((\bm{\beta}_1-\bm{\beta}_0)'\bm{x}_{t-1})}\right\}, \] such that $\Phi(.)$ is the standard normal distribution function. As in Section (ref), the distribution of $SN_{T}(\bm{\beta}_0\mid\bm{\beta} _{1})$ can be simulated under the null hypothesis with sufficient number of replications and the relevant critical values can be obtained to any degree of precision.
{\hskip 1.5em}In this Section, we first consider the issue of estimating the bivariate copulas in the D-vine decomposition and suggest a sequential estimation strategy for the parameters of the copulas. We then turn our attention to the problem of selecting a class of parametric bivariate copulas. The choice of the latter has an important implication on introducing dependence to the vector of signs.
{\hskip 1.5em}The calculation of the test statistics in Section (ref) requires four bivariate copula evaluations at $T(T-1)/2$ distinct points, leading to a total of $2T(T-1)$ copula evaluations. The estimation of the D-vine is often facilitated with the maximum likelihood (MLE hereafter). However, since the latter requires optimization with respect to at least $2T(T-1)$ copula parameters, sequential estimation procedures are favored for faster computation times, with the caveat that the increased speed comes at the cost of efficiency. Furthermore, the sequential estimates may be provided as starting points for the simultaneous numerical optimization using MLE [see czado2012maximum, haff2012comparison and dissmann2013selecting among others]. We assume that the copulas are specified parametrically, given by an appropriate parameter (vector). More specifically, let $\pmb{\theta}_{l}=(\theta_{1,k}',\cdots,\theta_{T-l,k}')'$ be the set of all the parameters to be estimated for tree $\mathcal{T}_l$, $l=1,\cdots,T-1$ of the D-vine, with $k=l-1$ conditioning variables. Therefore, $\pmb{\theta}=(\pmb{\theta}_1',\cdots,\pmb{\theta}_{T-1}')'$ is the entire set of the parameters that need to be estimated for the D-vine decomposition. To estimate the parameter vector $\pmb{\theta}$, we follow a sequential estimation strategy proposed by czado2012maximum, whereby first, the parameters of the unconditional bivariate copulas are estimated. These parameters are then utilized as means of estimating the parameters of bivariate copulas with a single conditioning variable. The latter are then used to estimate the pair-copulas with two conditioning variables, and so on. This bivariate copula estimation approach is continued sequentially until all parameters are estimated.
In the first step, the marginals are obtained by computing the conditional Bernoulli CDFs ((ref)) using an arbitrary distribution, such as the standard normal distribution considered in Section (ref). The second step of the process involves estimating the parameters of the unconditional copula, by fixing the marginals with their aforementioned estimates and maximizing the bivariate likelihood corresponding to each copula in each tree $\mathcal{T}_l$ to obtain $\pmb{\hat{\theta}}_l=(\hat{\theta}_{1,k}',\cdots,\hat{\theta}_{T-l,k}')$ for $l=1,\cdots,T-1$ and $k=l-1$ . As all the variables are discrete, the log-likelihood function, say, for the unconditional copula $C_{t,t+1}$ for $t=1,\cdots,T-1$ for the signs $\left(s(y_{i,t}),s(y_{i,t+1})\right)$, $i=1,\cdots,n-1$ is expressed as \[ L(\theta_{t,0})=\sum\limits_{i=1}^{T-1}\log\left\{\sum\limits_{\{a_1,a_2\}\in \{-,+\}^{2}}(-1)^{a_{j}}C_{t,t+1}\left(F_t(s_{i,t}^{a_1}\mid X;\mathbf{\hat{\bm{\beta}}_1}),F_{t+1}(s_{i,t+1}^{a_2}\mid X;\mathbf{\hat{\bm{\beta}}_1});\theta_{t,0}\right)\right\}. \] The estimate of the copula parameter, $\hat{\theta}_{t,0}$ for $t=1,\cdots,T-1$, is then obtained as follows \[ \hat{\theta}_{t,0}=\arg \max_{\theta_{t,0}} L(\mathbf{\theta}_{t,0}), \] which under regularity conditions solves \[ \frac{\partial L(\mathbf{\theta}_{t,0})}{\partial \theta_{t,0}}=0. \]
Let us illustrate this process with an example: once the marginals are obtained, the next step involves estimating the parameters $\theta_{t,0}$ for $t=1,\cdots,T-1$ of the unconditional copulas. Next, we are interested in estimating $\theta_{t,1}$ for $t=1,\cdots,T-2$. Define \[ \hat{u}_{t\mid t+1}=F_{t\mid t+1}\left(s_t\mid s_{t+1},X;\hat{\theta}_{t,0} \right), \] and \[ \hat{v}_{t+2\mid t+1}=F_{t+2\mid t+1}\left(s_{t+2}\mid s_{t+1},X;\hat{\theta}_{t+1,0} \right), \] for $t=1,\cdots,T-2$. The data $\hat{u}_{t\mid t+1}$ and $\hat{v}_{t+2\mid t+1}$ is then used to estimate the parameters $\theta_{t,1}$ for $t=1,\cdots,T-2$, denoted by $\hat{\theta}_{t,1}$. This procedure is repeated sequentially until all parameters are estimated. haff2010simplified show that under regularity conditions, the sequential estimates are asymptotically normal; however, as noted earlier their asymptotic covariance is \textquotedblleft intractable\textquotedblrightand the faster computation time comes at the cost of efficiency. Therefore, the sequential estimates can be utilized as the starting values of the high-dimensional MLE.
Another approach for estimating one-parameter pair-copulas in the sequential estimation procedure for copula families with a known relationship to Kendall's $\tau$ consists of inverting the empirical Kendall's $\tau$ based on, say, $\hat{u}_{t}$ and $\hat{u}_{t+1}$ for $t=1,\cdots,T-1$ for the edges of the first tree. However, we provide a caveat that the Kendall's $\tau$ of discrete data does not correspond to the Kendall's $\tau$ of the bivariate copulas [see denuit2005constraints]. denuit2005constraints show that by continuous extension of the discrete variables with a perturbation with values in $[0,1]$, the continuous features of Kendall's $\tau$ are adaptable to discrete data. In other words \[ \tau(s^*(y_{t}),s^*(y_{t+1}))=4\int\int_{[0,1]^2}C_{t,t+1}^*\left(\hat{u}_{t},\hat{u}_{t+1}\right)dC_{t,t+1}^*\left(\hat{u}_{t},\hat{u}_{t+1}\right)-1 \] for $t=1,\cdots,T-1$, such that $\hat{u}_{t}=F\left(s^*_t\mid X;\hat{\bm{\beta}}_1\right)$, \[ s^*(y_t)=s(y_t)+U-1 \] where $U$ is a continuous random variable in $[0,1]$. A natural choice for $U$ is the uniform distribution.
{\hskip 1.5em}Many different classes of parametric bivariate copulas have extensively been studied and reviewed by joe2014dependence. These include the Archimedean, elliptical, extreme value or max-id copula families that can specify the dependence structure of the vector of signs. As the dependence is introduced by the copula family, the type and the degree of the dependence between the signs depends on the choice of the copula. The literature surrounding the goodness-of-fit of copulas is extensive and has been analyzed by genest2006goodness, genest2009goodness, and berg2009copula, among many others. genest2009goodness categorize goodness-of-fit tests into three broad categories: procedures for testing particular dependence structures such as Gaussian or Clayton family; procedures that may be used for any classes of copulas, but require a strategic choice for their implementation; and finally, the so-called \textquotedblleft blanket tests\textquotedblright that apply to all classes of copulas and require no strategic choice for their use. A simple procedure proposed by joe1997multivariate involves specifying the Akaike information critetion (AIC) to different copulas and using it as a copula selection criterion, which is particularly attractive as it allows for the automation of the copula selection process [see. czado2012maximum]. The AIC specified to the copulas of, say, the first tree of the D-vine, can be expressed as follows \[ AIC=-2\sum\limits_{i=1}^{t}\log c_{t,t+1}(\hat{u}_{i,t},\hat{u}_{i,t+1};\hat{\theta}_{1,k})+2l \] for $t=1,.\cdots,T-1$ and $k=1,\cdots,T-1$, and where $l$ is the number of parameters $\theta_{1,k}$. panagiotelis2012pair suggest that while dependence structures such as tail dependence are weak in discrete data, the choice of the copulas could still have a significant effect on the joint pmf of the signs. They considered Gaussian, Clayton and Gumbel copulas in constructing the D-vines for Bernoulli margins by keeping the marginal probabilities and dependence constant, and have found that in the case where the probabilities of zero marginals and joint probabilities in the data is high, preference goes to the use of the Gumbel copula over the other two alternatives.
Within the context of our work, the mediangale assumption ((ref)) implies that the signs $s(y_1),\cdots,s(y_T)$ conditional on $X$, and in turn $\hat{u}_1,\cdots,\hat{u}_T$ exhibit serial nonlinear dependence. Thus, the issue with specifying a copula-based model for the signs is that the distribution of the signs must imply an identity correlation matrix, where independence is only sufficient for uncorrelatedness. The literature generally deals with this issue by capturing the serial nonlinear dependence by imposing the identity correlation matrix on a multivariate Student's $t$ distribution. Henceforth, we consider the \textquotedblleft jointly symmetric\textquotedblrightcopulas proposed by oh2016high, where the latter can be constructed with any given (possibly asymmetric) copula family. In addition, when they are combined with symmetric marginals, they ensure an identity correlation matrix. A \textquotedblleft jointly symmetric\textquotedblrightcopula is defined as follows
The general idea is that the average of mirror image rotations of a possibly asymmetric copula along each axis generates a jointly symmetric copula [see oh2016high]. For instance, the marginals can be assumed to possess standard normal distributions, while the nonlinear dependency is modeled using jointly symmetric copulas.
Therefore, using the jointly symmetric copula family, the bivariate copulas can be evaluated by considering
\[ C^{JS}\left(\phi_t(s_t\mid X;\hat{\bm{\beta_1}}),\phi_{t+1}(s_{t+1}\mid X;\hat{\bm{\beta_1}})\right)=\frac{1}{4}\sum\limits_{k_1=0}^{2}\sum\limits_{k_2=0}^{2}\left(-1\right)^R C(\phi_t(s_t\mid X;\hat{\bm{\beta_1}}),\phi_{t+1}(s_{t+1}\mid X;\hat{\bm{\beta_1}})) \] \[ where\quad R=\sum\limits_{i=1}^n\bm{1}\{k_i=2\},\quadand\quad\tilde{u}_i=
,\quad \forall i\in\{t,t+1\}. \] for $t=1,\cdots,T-1$.
{\hskip 1.5em}Following joe2014dependence, we refer to a D-vine as a $p$-truncated D-vine, if the copulas in the trees $\mathcal{T}_{p+1},\cdots,\mathcal{T}_{T}$ are $C^{\perp}$, where by definition \[ C^{\perp}\left(u_1,\cdots,u_n\right)=\prod\limits_{t=1}^{T}u_t,\quad\text{with}\quad (U_1,\cdots,U_T)\sim U(0,1), \] implying $U_1\perp U_2 \perp \cdots \perp U_T$. The POS-based tests constructed using the Markov assumption of order one can be regarded as a special case of the PCC-POS based tests, whereby the former can be constructed by a $1$-truncated D-vine, which only depends on $C_{12},C_{23},\cdots,C_{T-1,T}$, or rather $C_{12},C_{\sigma_{13}3},\cdots,,C_{\sigma_{1T}T}$ using a vine array representation, given the following array
Similarly, a $2$-truncated D-vine, depends on the copulas $C_{12},C_{23},\cdots,C_{T-1,T}$ and $C_{13\mid 2},C_{24\mid 3},\cdots,C_{T-2,T\mid T-1}$ or $C_{\sigma_{1t}}t$ for $t=2,\cdots,T$ and $C_{\sigma_{2t}t\mid\sigma_{1t}}$ for $t=3,\cdots,T$ using the vine array representation. Therefore, for a $p$-truncated D-vine, ((ref)) is modified to \begingroup \allowdisplaybreaks
\endgroup for $t-1\geq p$.
In this Section, we follow dufour2010exact by first showing the analytical derivation of the power envelope function of the PCC-POS-based tests. We then suggest using simulations as means of approximating the said function, by showing the difficulty of inverting the latter to find the optimal alternative. Thereafter, we propose an adaptive approach based on the split-sample technique to choose an alternative which has a power function close to that of the power envelope.
{\hskip 1.5em}Point-optimal tests trace out the power envelope (i.e. the maximum attainable power) for any given testing problem [see king1987towards]. However, in practice the alternative hypothesis $\bm{\beta}_1$ is unknown and a problem consists of finding an approximation for it, such that the power function is maximized and is close to that of the power envelope. Following dufour2010exact and dufour2001finite, we propose an adaptive approach based on the split-sample technique to choose an alternative $\bm{\beta}_1$ that yields the greatest power function and makes size control easier [see dufour2001finite and dufour2010exact for an overview]. We follow dufour2010exact by presenting the analytical derivation of the power envelope of the PCC-POS tests for predictive regressions, which can be purposed as a benchmark for comparing the power functions of the PCC-POS tests for different sample splits.
We have shown in Section (ref) that the PCC-POS tests are a function of $\bm{\beta}_1$. In other words, \[ SN_{T}(\bm{\beta}_0\mid\bm{\beta}_{1})=\sum\limits_{t=2}^{T}\sum\limits_{l=t-1}^{2}\ln c_{\tilde{\sigma}_{lt}t\mid \tilde{\sigma}_{1t},\cdots,\tilde{\sigma}_{t-1,t}}+\sum\limits_{t=2}^{T}\ln c_{\tilde{\sigma}_{1t}t}+\sum\limits_{t=1}^{T}\ln\left\{\frac{1-p_t[\bm{x}_{t-1},\bm{\beta}_0,\bm{\beta}_1\mid X]}{p_t[\bm{x}_{t-1},\bm{\beta}_0,\bm{\beta}_1\mid X]}\right\}s(y_t-f(\bm{x}_{t-1},\bm{\beta}_0)). \] which in turn implies that its power function, say $\Pi(\bm{\beta}_0,\bm{\beta}_1)$, is also a function of $\bm{\beta}_1$ \[ \Pi(\bm{\beta}_0,\bm{\beta}_1)=P[SN_{T}(\bm{\beta}_0\mid\bm{\beta} _{1})>c_1(\bm{\beta}_0,\bm{\beta}_1)\mid H(\bm{\beta}_1)] \] where $c_1(\bm{\beta}_0,\bm{\beta}_1)$ is the smallest constant that satisfies $P[SN_T(\bm{\beta}_0\mid\bm{\beta}_1)>c_1(\bm{\beta}_0,\bm{\beta}_1)\mid H(\bm{\beta}_0)]\leq \alpha$, and where $\alpha$ is an arbitrary significance level. Theorem (ref) provides the theoretical results for the power function of the PCC-POS tests.
Under the assumption that the signs follow an RMT-process, $\rho_t(u)$ can be estimated using the results from Theorem 2 of heinrich1982factorization. Given that point-optimal tests are optimal at a specific point in the alternative parameter space, the power envelope of the PCC-POS tests, say $\bar{\Pi}(\bm{\beta}_1)$, is obtained for values of $\bm{\beta}$, such that $\{\bm{\beta}: \bm{\beta}=\bm{\beta}_1, \forall \bm{\beta}_1\in \mathbb{R}^{(k+1)}\}$. Finding values of $\bm{\beta}_1$ for a PCC-POS test at level $\alpha$, with a power function that is close to the power envelope can be achieved by inverting the power envelope function. However, in a much simpler case of POS tests for i.n.i.d data, dufour2010exact show that the inversion of the power function is not a straightforward task and obtaining an exact solution is not feasible. Therefore, simulations are used as means of approximating the power envelope function and finding the optimal alternative for the PCC-POS test.
{\hskip 1.5em}As noted earlier, the power function of the PCC-POS test statistic depends on the alternative $\bm{\beta}_1$, which in practice is unknown and needs to be approximated. To make size control easier and to choose an approximation to $\bm{\beta}_1$ such that the power function of the test statistic is close to that of the power envelope, we follow dufour2010exact by proposing an adaptive approach based on the split-sample technique for choosing the alternative. For an extensive review of adaptive statistical methods, we refer the reader to o2004applied. Furthermore, the application of the split-sample technique in parametric settings can be studied by consulting dufour2003point and dufour2008finite.
The split-sample technique involves splitting a sample of size $T$ into two independent subsamples, say $T_1$ and $T_2$, such that $T=T_1+T_2$. The first subsample is then used to estimate the alternative $\bm{\beta}_1$, while the other is purposed for computing the PCC-POS test statistic. Assuming that $f(\bm{x}_{t-1},\bm{\beta})=\bm{x}_{t-1}'\bm{\beta}$, the alternative $\bm{\beta}_1$ can be estimated using OLS \[ \hat{\bm{\beta}}_{(1)}=(X_{(1)}'X_{(1)})^{-1}X_{(1)}'y_{(1)}. \] We provide a caveat that the OLS estimator is sensitive to extreme outliers, which motivates the use of robust estimators [see. maronna2019robust for a review of robust estimators]. Using $\hat{\bm{\beta}}_{(1)}$ and the observations in the second independent subsample, we compute the test-statistic as follows \begingroup \allowdisplaybreaks
\endgroup where for $t=T_1+2,\cdots,T$ and D-vine-array $\tilde{A}=(\tilde{\sigma}_{lt})_{1\leq l\leq t\leq n}$, $l=T_1+1,\cdots,T-1$ is the row with tree $\mathcal{T}_l$, and column $t$ has the permutation $\tilde{\underline{\sigma}}_{t-1}=(\tilde{\sigma}_{(T_1+1)t},\cdots,\tilde{\sigma}_{t-1,t})$ of the previously added variables, $p_t[\bm{x}_{t-1},\bm{\beta}_0,\bm{\beta}_{(1)}\mid X]=P_t[\varepsilon_t\leq \bm{x}_{t-1}'(\bm{\beta}_0-\bm{\beta}_1)\mid X]$, and $\tilde{\text{\b{S}}}_{t-1}=s(y_{t-1}-\bm{x}_{t-2}'\bm{\beta}_0),\cdots,s(y_{T_1+2}-\bm{x}_{T_1+1}'\bm{\beta}_0)$.
The choices for the subsamples $T_1$ and $T_2$ can be arbitrary. However, our simulations show that the proportion of the observations retained for estimating the alternative and in turn for computing the PCC-POS test statistic has an impact on the power of the test. We find that the power function of the split-sample PCC-POS test (SS-PCC-POS test hereafter) is closest to that of the power envelope, when a relatively small number of observations is retained for estimating the alternative, with the rest used for computing the test statistic - findings that are in line with dufour2010exact. Specifically, by considering all the DGPs in our simulations study, we find that the subsamples $T_1$ and $T_2$ must in turn contain roughly $10\%$ and $90\%$ of the observations in the entire sample respectively.
{\hskip 1.5em}In this Section, we lay out the theoretical framework for building confidence regions for a vector (sub-vector) of the unknown parameters $\bm{\beta}$, say $C_{\bm{\beta}}(\bm{\alpha})$, at a given significance level $\alpha $, using the proposed PCC-POS tests. Consider model ((ref)) such that $f(\bm{x}_{t-1},\bm{\beta})=\bm{x}_{t-1}'\bm{\beta}$. Suppose we wish to test the null hypothesis ((ref)) against the alternative hypothesis ((ref)). Formally, this implies finding all the values of $\bm{\beta}_{0}\in \mathbb{R}^{k}$ such that
where the critical value is given by the smallest constant $c(\bm{\beta} _{0},\bm{\beta}_{1})$ such that
Thus, the confidence region of the vector of parameters $\bm{\beta}$ can be defined as follows:
Given the confidence region $C_{\bm{\beta} }(\alpha )$, confidence intervals for the components and sub-vectors of vector $\bm{\beta}$ can be derived using the projection techniques [see dufour2010exact and coudin2009finite]. Confidence sets in the form of transformations $T$ of $\bm{\beta}\in \mathbb{R}^{m}$, for $m\leq (k+1)$, say $T(C_{\bm{\beta}}(\alpha ))$, can easily be found using these techniques. Since, for any set $C_{\bm{\beta} }(\alpha )$
we have
where
\newline From ((ref)) and ((ref)) , the set $T(C_{\bm{\beta} }(\alpha ))$ is a conservative confidence set for $T(\bm{\beta} )$ with level $1-\alpha $. If $ T(\bm{\beta} )$ is a scalar, then we have
{\hskip 1.5em}In this Section, we assess the performance of the proposed 10% SS-PCC-POS tests (in terms of size control and power) by comparing it to other tests that are intended to be robust against non-standard distributions and heteroskedasticity of unknown form. We consider DGPs under different distributional assumptions and heteroskedasticities. For each DGP, we further consider different correlation coefficients between the errors of the predictive regression and the disturbances of the regressors. In the first subsection, the DGPs are formally introduced, and in the following subsection, the performance of the proposed 10% SS-PCC-POS tests are compared to that of the other tests considered in our study.
{\hskip 1.5em}We assess the performance of the proposed 10% SS-PCC-POS tests in terms of size and power, by considering various DGPs with different symmetric and asymmetric distributions and forms of heteroskedasticity. The DGPs under consideration are supposed to mimic different scenarios that are often encountered in practical settings. The performance of the 10% SS-PCC-POS tests is compared to that of a few other tests, by considering the following linear predictive regression model
where $\beta$ is an unknown parameter and
where $\theta=0.9$ and
for $\rho=0,0.1,0.5,0.9$ and $\varepsilon_t$ and $w_t$ are assumed to be independent. The initial value of $x$ is given by: $x_0=\frac{w_0}{\sqrt{1-\theta^2}}$ and $w_t$ are generated from $N(0,1)$. The errors $\varepsilon_t$ are i.n.i.d and are categorized by two groups in our simulation study. In the first group, we consider DGPs where the errors $\varepsilon_t$ possess symmetric and asymmetric distributions:
The second group of DGPs represents different forms of heteroskedasticity:
We consider the problem of testing the null hypothesis $H_0: \beta=0$. Our Monte Carlo simulations compare the size and power of the 10%-PCC-POS test to those of t-test, t-test based on white1980heteroskedasticity variance-correction (WT-test hereafter), and the sign-based test proposed by dufour1995exact. Due to computational constraints, we perform only $M_1=999$ simulations to evaluate the probability distribution of the 10% SS-PCC-POS test statistic and $M_2=1,000$ iterations for approximating the power functions of the proposed PCC-POS test and other tests. In all simulations, we consider a sample size of $T=50$. As the sign-based statistic of dufour1995exact has a discrete distribution, it is not possible to obtain test with a precise size of $5\%$; therefore, the size of the test is $5.95\%$ for $T=50$.
{\hskip 1.5em}The results of the Monte Carlo study corresponding to DGPs described in Section (ref) are presented in figures (ref)-(ref). These figures compare the performance of the 10% SS-PCC-POS test in terms of size and power, to those of the t-test, t-test based on White's (1980) variance-correction, as well as the sign-based procedure proposed by dufour2010exact. The results are described in detail below.
First, figure (ref) considers the case where the error terms $\varepsilon_t$ are normally distributed. At first glance, we note that all tests control size. Evidently, our test is outperformed by the t-test, as well as the t-test based on White's (1980) variance-correction. The former is expected, since for normally distributed error terms, the t-test is the most powerful test. However, the 10% SS-PCC-POS test outperforms the sign-based procedure proposed by dufour1995exact [CD (1995) hereafter]. Furthermore, changing the correlation coefficient $\rho$ does not seem to lead to visually significant differences in the performance of the tests.
Second, figure (ref) presents the results of the performance of the aforementioned tests, when the errors $\varepsilon_t$ follows Cauchy distribution. It is evident that the 10% SS-PCC-POS test outperforms all other tests. Moreover, the t-test and WT-test do not possess much power for low correlation coefficient (0 and 0.1) values, $\rho$. However, as the correlation between $u_t$ and $w_t$ increases, the gap between the power functions narrows significantly.
Third, in figures (ref) and (ref), we have considered the cases where the errors in turn follow $t(2)$ and mixture distributions. In the former case of $t(2)$ distributed errors, the 10% SS-PCC-POS test outperforms the rest; however, for almost all correlation coefficients $\rho$, the gap between the power functions is rather small, albeit it is narrowest for $\rho=0.9$. In the case of errors with mixture distribution, our 10% SS-PCC-POS test is still the most powerful test. On other hand, it is evident that the t-test and WT-test do not possess much power for small values (0 and 0.1) of correlation coefficient $\rho$. However, the power functions increase and converge to those of the other tests, as the correlation increases.
An interesting observation is the stark contrast between the power of the 10% SS-PCC-POS test and the t-test, when the errors follow the Cauchy, $t(2)$ and normal distributions respectively. The Cauchy and $t(2)$ distributions possess heavy tails, in the presence of which the standard error of the regression coefficients is inflated, which in turn leads to low power when the mean is used as a measure of central tendency. For instance, the Cauchy distribution has the heaviest tails among the considered DGPs, as a result of which the t-test and WT-test have very low power. By noting that a Student's $t$ distribution with $\nu$ degrees of freedom converges to the Cauchy distribution for $\nu=1$ and to the normal distribution as $\nu\rightarrow \infty$, it would be interesting to see at which degree of freedom the 10% SS-PCC-POS test is outperformed by the t-test and WT-tests. Figures (ref)-(ref) suggest that for different values of $\rho$ in ((ref)) the t-test and WT test outperform the 10% SS-PCC-POS test for $\nu=4$. Interestingly, figure (ref) shows that the tails of the $t(2)$ distribution are substantially heavier than that of the $t(4)$, which may explain the transition.
Finally, in figures (ref) and (ref), the errors are normally distributed with different forms of heteroskedasticity. In the first case [see figure (ref)], there is a break in variance, in the presence of which our test outperforms the other tests. Furthermore, the t-test and WT-test do not possess any power for low correlation (0 and 0.1) values of $\rho$. However, with increasing values of the correlation coefficient the power curves of all test appear to converge. In the other case [see figure (ref)], the variance follows a GARCH(1,1) model with a jump in variance. In this case, our test is only outperformed by the CD (1995) test, which has the greatest power. Nevertheless, the 10% SS-PCC-POS test still outperforms the t-test and WT-test.
\FloatBarrier
{\hskip 1.5em}In this paper, we extend the exact point-optimal sign-based procedures proposed by dufour2010exact to a predictive regression framework. We showed that by implementing the procedures for pair copula constructions of discrete data, we can derive exact and distribution-free sign-based statistics for dependent data in the context of linear and nonlinear predictive regressions, without imposing any potentially restrictive assumptions. The proposed tests are valid, distribution-free and robust against heteroskedasticity of unknown form. Furthermore, they may be inverted to produce a confidence region for the vector (sub-vector) of parameters of the regression model.
We further suggest a sequential estimation strategy for the D-vine PCC and discuss the choice of the copula family. As the proposed sign statistics depend on the alternative hypothesis, another problem consists of finding an alternative that controls size and maximizes the power. In line with dufour2010exact, we find that when 10% of sample is used to estimate the alternative and the rest to compute the test-statistic, our procedures have the optimal power and are closest to the power envelope.
Finally, we present a Monte Carlo study to assess the performance of the proposed tests in terms of size control and power, by comparing them to some other tests that are intended to be robust against heteroskedasticity. We consider a variety of different DGPs and we show that the 10% split-sample point-optimal sign-test based on pair copula constructions is superior to the t-test, dufour1995exact sign-based test, and the t-test based on white1980heteroskedasticity variance correction in most cases.