EconBase
← Back to paper

New robust inference for predictive regressions

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

63,577 characters · 11 sections · 128 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

New robust inference for predictive regressions

abstractWe propose a robust inference method for predictive regression models under heterogeneously persistent volatility as well as endogeneity, persistence, or heavy-tailedness of regressors. This approach relies on two methodologies, nonlinear instrumental variable estimation and volatility correction, which are used to deal with the aforementioned characteristics of regressors and volatility, respectively. Our method is simple to implement and is applicable both in the case of continuous and discrete time models. According to our simulation study, the proposed method performs well compared with widely used alternative inference procedures in terms of its finite sample properties in various dependence and persistence settings observed in real-world financial and economic markets.

\noindentKeywords: predictive regressions, robust inference, near nonstationarity, heavy tails, nonstationary volatility, endogeneity.

\noindentJEL Codes: C12, C22

Introduction

Many papers in the literature have focused on econometric analysis of predictive regressions for stock returns (see Phillips2015, Phillips2015, for an up-to-date review). Predictive regression data is known to have several problematic characteristics, especially in statistical inference of stock return predictability. First, it is widely believed that the popular regressors, such as dividend-price and earnings-price ratios, used in the predictive regressions have near unit roots and their innovations are correlated with stock returns in the long run. The characteristics of the regressors, which are persistence and endogeneity jointly cause standard hypothesis tests to become substantially biased (see stambaugh1999predictive). Second, there is some evidence that supporting volatility of stock returns is stochastic and highly persistent (see, e.g., jacquier2004bayesian and hansen2014estimating). cavaliere2004testing shows that persistent stochastic volatility may cause substantial size distortions on standard tests developed mostly under the assumption that a volatility process is stationary with a constant unconditional mean, such as stationary GARCH-type models. Lastly, there are several other characteristics of predictive regression data that include heavy-tailedness of regressors as well as jumps, structural breaks and regime switching in volatility. These characteristics may also yield jointly or individually a significant distortion of standard hypothesis tests for predictive regressions.

In this paper, we propose a new method for robust inference on parameters of predictive regression models under the aforementioned characteristics of predictive regression data. Our approach relies on a simple nonlinear instrumental variable (IV) estimation and a nonparametric volatility correction. The nonlinear IV estimator in our approach is an IV estimator with the instrument being the sign transformation of the regressor. This particular IV estimator was first proposed by cauchy1836lxxviii, and is called the Cauchy estimator. As is known in the literature, the use of the instrument can effectively eliminate the problems caused by the persistent endogeneity, heavy-tailedness and other problematic characteristics of the regressors (see SoShin, CJP2016 and kim-meddahi-2019). On the other hand, volatility correction is used to deal with the problems caused by the presence of heterogeneity and persistence in stock return's volatility. As for the volatility correction, we consider a standard kernel-based nonparametric estimator of volatility.

Many authors have studied the issue of persistent endogeneity of regressors in predictive regressions. Among many of them, CampbellYogo2006, ChenDeo2009 and PhillipsMagdalinos2009 have proposed tests of return predictability, which are aimed at dealing with persistent and endogenous regressors. Though their tests perform well under the presence of persistent endogeneity of regressors, they are not expected to deal with other problematic characteristics of predictive regression data effectively. Our simulation study shows they have serious size distortions under the null of no predictability when volatility is persistent or incorporates structural breaks or regime switching. In contrast, the robustness of our approach is quite evident. Our approach always yields almost exact sizes in a variety of designs considered in our simulation study. Moreover, the robustness of our approach is obtained with no significant loss of power. The discriminatory powers of our test are comparable to the tests by CampbellYogo2006 and ChenDeo2009, which are optimal for the basic Gaussian model.

Our work is closely related to CJP2016, who propose an inference approach for predictive regressions. Similar to our method, their approach relies on the Cauchy estimator to eliminate the problems caused by the problematic characteristics of the regressors. They also use a nonparametric volatility correction. However, their approach for volatility correction is quite different from ours, and its applicability is limited to a predictive regression equipped with appropriate high frequency data since their method and theory are developed in a continuous time framework. More precisely, their volatility correction, called the time change, requires uniformly consistent estimation of a quadratic variation of a stock price for which high frequency observations of the stock price are necessary. Consequently, their approach requires the assumption that the sampling interval decreases to zero, and applications of their method on relatively low frequency data, i.e. monthly or quarterly data, are largely restricted. However, predictive regressions are often estimated using monthly or quarterly data. In contrast, our method can be applied to a discrete time model and a discrete sample collected from an underlying continuous time model as in CJP2016. Our simulation study shows that both our method and the method by CJP2016 perform well and have good size and power performances under continuous time settings considered in this paper. However, unlike our method, the CJP2016 method is not applicable under discrete time settings. Therefore, we may say that our method is more flexible and widely applicable since it can be applied to both high and low frequency data.

The rest of the paper is organized as follows. Section 2 introduces the predictive regression models, persistent volatility, and the Cauchy estimator. Section 3 proposes the robust inference method and presents its asymptotic properties. Section 4 generalizes the baseline predictive regression models, which have one persistent volatility factor, to have a two-factor volatility, where one factor is persistent and the other is transient such as a stationary GARCH process. Section 5 provides numerical results on finite sample properties of the proposed robust inference approach. Section 6 makes some concluding remarks.

The online supplementary appendix provides a discussion of the Cauchy estimator and general nonlinear IV estimators with the relevant asymptotic results that, in particular, point to the importance and usefulness of the Cauchy estimator (Appendix A); useful auxiliary results (Appendix B); the proofs of the main results in the paper (Appendix C); and some additional simulation results on finite sample performance of inference approaches dealt with (Appendix D).

Predictive Regressions

Research Problems and Models

Throughout the paper, we consider $(\mathcal{F}_t)$-adapted processes defined on a filtered probability space $(\Omega,\mathcal{F}, (\mathcal{F}_t)_{t\geq0}, P)$ equipped with an increasing filtration $(\mathcal{F}_t)$ of sub-$\sigma$-fields of $\mathcal{F}$. We consider a test for no predictability of the process $(y_t)$ (e.g., the time series of excess stock returns) based on some covariate process $(x_t)$ (e.g., the time series of price-to-dividend ratios). We consider the linear predictive regression model

align[align omitted — 84 chars of source]

where $(u_t)$ is a martingale difference sequence (MDS) with respect to $(\mathcal{F}_t)$. In particular, $(u_t)$ is conditionally heteroskedastistic. Following the usual specification for a volatility model, we assume that

align*[align* omitted — 38 chars of source]

where $(v_t)$ is a volatility process and $(\varepsilon_t)$ is an MDS with respect to $(\mathcal{F}_t)$.

ass(a) $(v_t)$ is $(\mathcal{F}_{t-1})$-adapted and is defined on $[\underline{v},\bar{v}]$ for some $0<\underline{v}<\bar{v}<\infty$, (b) $E(\varepsilon_t^2|\mathcal{F}_{t-1})=1$, and (c) $\sup_{t\geq1}E(\varepsilon_t^4|\mathcal{F}_{t-1})<\infty$.

The conditions (a)-(b) in Assumption (ref) are not stringent, and are required for the identification of the conditional variance of $u_t$. In particular, the conditional variance of $u_t$ given $\mathcal{F}_{t-1}$ is well identified and we have $E(u_t^2 | \mathcal{F}_{t-1}) = v_t^2$. Our test relies on uniform convergence results for a nonparametric estimator of the volatility process $(v_t)$. The condition (c) is used to obtain a uniform convergence rate of the nonparametric estimator of the volatility process $(v_t)$. Note that Assumption (ref) implies $\sup_{t\geq1}E(u_t^4)<\infty$, and hence, it rules out a predictive regression model having a heavy-tailed regression error $(u_t)$.

As for a nontrivial example, we let $v_t = f(z_{t-1})$ and $\varepsilon_t \sim iid \mathbb{N}(0,1)$, where $f$ is a positive function and $z_t$ is an $\mathcal{F}_t$-adapted process. Then $(v_t,\varepsilon_t)$ satisfies Assumption 2.1. If we assume that $f$ is bounded above, then $u_t = v_t \varepsilon_t$ satisfies $\sup_{t\leq 1} E(|u_t|^4 | \mathcal{F}_{t-1})<\infty$ a.s. for any $\mathcal{F}_t$-adapted process $(z_t)$ since $\varepsilon_t \sim iid\mathbb{N}(0,1)$. Moreover, $u_t$ is not uniformly bounded, i.e., there does not exist $M$ such that $|u_t|\leq M < \infty$ with probability one even if $f$ is bounded above, since the standard normal random variable $\varepsilon_t$ is not uniformly bounded. Further examples of martingales with bounded conditional moments of MDS summands are provided by more general martingale transforms and randomly stopped sums of independent r.v.'s (see Remark 3.3 in VPRI).

The hypothesis of no predictability of $(y_t)$ corresponds to the hypothesis $\beta=0$ in predictive model (ref). It is well-known that the standard OLS-based $t$-test is not robust with respect to a wide range of statistical problems in predictive regression data. For instance, the standard OLS estimator of $\beta$ is not asymptotically Gaussian under $H_0: \beta=0$ if $(x_t)$ is endogenous and (nearly) nonstationary (see elliott1994inference, elliott1994inference, PhillipsNear, PhillipsNear, PM, PM) or is stationary with infinite second moment (e.g., GrangerOrr, GrangerOrr, EKM, EKM, ibragimov2015heavy, ibragimov2015heavy, and references therein), even when there is no heteroskedasticity and $v_t = \sigma$ is constant.\footnote{The endogeneity of the covariate $x_t$ refers to the existence of nonzero long run covariance between innovations of $u_t$ and $x_t.$}

In the case of predictive regressions for stock returns, the returns process $(y_t)$ is widely believed to have time-varying stochastic volatility (see CJP2016 and references therein). Moreover, the volatility process is typically very persistent. For example, many authors have found that the autoregressive parameter for the dynamics of the volatility process is close to one under some appropriate functional transformations. In particular, jacquier2004bayesian and hansen2014estimating, provide convincing evidence that the logarithm of the volatility process follows a near unit root process for a wide range of equity and foreign exchange rate time series. It is well known that the presence of persistent volatility may cause the distribution of the standard $t$-statistic to be far from standard normal, yielding a substantial distortion in testing relying on standard normal critical values (see, e.g., chung2007nonstationary, CJP2016 and KimPark1).

The Cauchy Estimator

Our inference method is based on the Cauchy estimator. To effectively explain the main idea, we consider model (ref) with no intercept term, i.e., $\alpha=0$, and introduce the Cauchy estimator $\check{\beta}$ for $\beta$, which is given by

align*[align* omitted — 105 chars of source]

where $sign(\cdot)$ is the sign function defined as $sign(x)=1$ for $x\geq0$, and $sign(x)=-1$ for $x<0$. Thus, $\check{\beta}$ is an instrumental variable (IV) estimator with the instrument $sign(x_{t-1}).$ This particular IV estimator was first proposed by cauchy1836lxxviii. See, among others, SoShin, PPC, CJP2016 and kim-meddahi-2019 for econometric applications of the Cauchy estimator.

Under Assumption (ref) (b), not only $\varepsilon_t$, but also $sign(x_{t-1})\varepsilon_t$, hereafter denoted by $\xi_t$, is an MDS with respect to the filtration $(\mathcal{F}_t)$ with $E(\xi_t^2|\mathcal{F}_{t-1})=1$. Let us define a continuous time partial sum process $(W_T(r), 0\leq r\leq 1)$ by

equation[equation omitted — 77 chars of source]

The stochastic process $(W_T(r))$ takes values in $\mathbf{D}_{\mathbb{R}}[0,1]$, where $\mathbf{D}_E[0,1]$ denotes the space of c\`{a}dl\`{a}g functions from $[0,1]$ to $E\subset \mathbb{R}^d$ for some positive integer $d.$ Under Assumption (ref) (b)-(c), the partial sum process $(W_T(r))$ follows the usual functional central limit theorem (CLT) for martingales, see, e.g., Theorem 18.2 of billingsley1986convergence, that is, \[ W_T \to_d W \] in $\mathbf{D}_\mathbb{R}[0,1]$, where $W$ is a standard Brownian motion. The convergence $W_T \to_d W$ is to be interpreted as the weak convergence in the probability measures on $\mathbf{D}_\mathbb{R}[0,1]$. In our context, it is more convenient, and so is assumed, to endow $\mathbf{D}_E[0,1]$ with the uniform topology rather than the usual Skorohod topology (see billingsley1986convergence).

The use of the Cauchy estimator in our inference method is motivated by the above functional CLT for $(W_T(r))$. To convey the main idea, assume that the volatility process $(v_t)$ is observable. Recall that the numerator of $\check{\beta}$ is $\sum_{t=1}^T sign(x_{t-1})y_t$, and it becomes $\sum_{t=1}^T v_t \xi_t$ under $\beta=0$, where $\xi_t = sign(x_{t-1})\varepsilon_t$. One then may construct a robust test for the null hypothesis $H_0: \beta=0$ against the alternative $H_1: \beta\neq 0$ using the following statistic

align[align omitted — 105 chars of source]

In particular, for $\beta=0$, \[ \tau(v) = \frac{1}{T^{1/2}} \sum_{t=1}^T \xi_t = W_T(1) \to_d W(1) = \mathbb{N}(0,1). \] In practice, however, the volatility process $(v_t)$ is not observable, and hence, the above inference procedure using $\tau(v)$ is not feasible. In Section 3, a feasible version of the Cauchy based inference method above will be fully addressed under our construction of the persistent volatility introduced in Section 2.3.

Persistent Volatility

This subsection presents a time-varying and persistent volatility model, which is a well-known stylized fact for many financial returns. We define a stochastic process $\sigma_T$ on $\mathbf{D}_{\mathbb{R}^+}[0,1]$ as $\sigma_T(r) = v_{[Tr]}$. We assume that $\sigma_T$ has a limiting process $\sigma$ defined over $0\leq r\leq 1$ such that $(W_T, \sigma_T)$ converges to $(W, \sigma)$ jointly, where $(W_T)$ is defined as in (ref). Specifically, we consider the following assumption.

assThere exists a positive process $\sigma$ on $\mathbf{D}_{\mathbb{R}^+}[0,1]$ such that \[ (W_T, \sigma_T) \to_d (W, \sigma) \] in $\mathbf{D}_{\mathbb{R} \times \mathbb{R}^+}[0,1]$, where $W$ is a standard Brownian motion with respect to the filtration to which $W$ and $\sigma$ are adapted.

The above assumptions hold for wide classes of models, such as models with nonstationary volatility, regime switching, and structural breaks in volatility. It also holds for the processes with $v_t = \sigma(t/T),$ where $\sigma(s)$ is a deterministic function on $[0, 1],$ considered by cavaliere-taylor-2007, xu-phillips-2008 and harvey-leybourne-zu, among others.\footnote{Assumption 2.2 is a simplified version of the condition $v_{[Tr]}/a_T\to_d \sigma_r$ considered by Assumption 2 of cavaliere-taylor-2009. We rule out the explosive volatility settings with $a_T\to\infty$, and consider the stable volatility processes with $a_T=1$ for simplicity. The results in the paper can be obtained under the explosive volatility assumption with $a_T\to\infty$ at the cost of a more involved analysis.} The assumptions also hold for processes with nonstationary volatilities considered by hansen1995regression and chung2007nonstationary, who assume that $v_t^2$ is a smooth positive transformation of a (near) unit root process, i.e., $v_t^2 = \sigma^2(T^{-1/2}z_{t-1})$ for a unit root process $z_t$. One should note that Assumption (ref) is more general than the volatility models considered in the aforementioned literature and, in particular, allows the volatility to be stochastically discontinuous, which are desirable properties for modelling financial volatility having structural breaks or regime switching.

Assumptions (ref) and (ref) rule out some cases of globally homoskedastic processes, such as stationary GARCH processes. In Section 4, we generalize the model to have a two-factor volatility, one for a nonstationary long run component and the other one for a stationary short run component, and show the validity of our robust method introduced in Section 3 for the generalized model with the two-factor volatility.

Under our construction of the persistent volatility, the asymptotic behavior of the Cauchy estimator can be obtained immediately. The asymptotics of the Cauchy estimator $\check{\beta}$ are mainly determined by $\sum_{t=1}^T v_t\xi_t$ since $\check{\beta} = \beta + \sum_{t=1}^T v_t\xi_t/\sum_{t=1}^T |x_{t-1}|$. Note that $T^{-1/2}\sum_{t=1}^{[Tr]} v_t\xi_t = \int_0^r \sigma_T(s)dW_T(s)$ for $r\in[0,1]$, and the weak convergence of the stochastic integral $\int \sigma_T(r)dW_T(r)$ is well documented in the literature (see, e.g., Theorem 2.1 of hansen-1992 and Theorem 4.6 of KP), and we have $(\int \sigma_T(r) dW_T(r)) \to_d (\int \sigma(r) dW(r))$.

lemUnder Assumptions (ref) and (ref), \[ \left(\sum_{t=1}^T |x_{t-1}|/\sqrt{T} \right)\left(\check{\beta}-\beta\right) \to_d \int_0^1 \sigma(r)dW(r). \]

Two main implications of Lemma (ref) are (i) the limit distribution of the Cauchy estimator is generally non-Gaussian and (ii) the rate of convergence of the Cauchy estimator is nonstandard and unknown. These asymptotic properties of the Cauchy estimator subsequently imply that the usual $t$-test relying on the standard normal table becomes an invalid testing procedure for the null hypothesis of $\beta = 0$. The limit $\int_0^1 \sigma(r)dW(r)$ is Gaussian if and only if the limiting volatility process $\sigma$ is independent of $W$. In this case, $\int \sigma(r)dW(r)$ has a mixed normal distribution and $\int_0^s \sigma(r)dW(r) =_d \mathbb{MN}(0,\int_0^s \sigma^2(r)dr)$. If the independence condition is violated, then $\int \sigma(r)dW(r)$ becomes a non-Gaussian martingale in general.

Clearly, $\check{\beta}$ requires an extremely mild condition for consistency, that is $\sum_{t=1}^T |x_{t-1}|/\sqrt{T}\to_p\infty$. For example, if there exists a sequence $p_T$ of positive numbers such that \[ \left(p_T^{-1}\sum_{t=1}^T|x_{t-1}|\right)^{-1} = O_p(1), \] then $\check{\beta} - \beta = O_p(T^{1/2}/p_T)$ by Lemma (ref). For a wide class of time series, the consistency condition $T^{1/2}/p_T\to0$ is satisfied since $p_T\geq T$ unless $x_t\approx 0$ for most $t=1,\cdots,T$. Though it is not necessary in our subsequent theory, one may explicitly obtain the sequence $p_T$ for some time series satisfying required regularity conditions.

example(a) For weakly stationary processes $(x_t)$ with $E|x_t|<\infty$, $p_T=T$; (b) For stationary $\alpha$-stable $(x_t)$ with $0<\alpha<1$, $p_T=T^{1/\alpha}\ell(T)$ for some slowly varying function $\ell$ (see EKM, EKM, PSolo, PSolo, and references therein); (c) For the case of unit root and near unit root time series $(x_t)$, $p_T=T^{3/2}$ (see PhillipsUR, PhillipsUR, PhillipsNear, PM, PM, IP, IP, and references therein); (d) For fractionally integrated $I(d)$ processes $(x_t)$ with $1/2<d<3/2,$ $p_T=T^{d+1/2}\ell(T)$ for some slowly varying function $\ell$ (see Baillie, Baillie; Lemma 3.4 in PhillipsFrac1, PhillipsFrac1; PhillipsFrac2, PhillipsFrac2, wangetal and chanwang and references therein);

New Robust Inference Approach

Now we introduce our test for no predictability in the regression ((ref)). The test is motivated by $\tau(v)$ in (ref). Since $(v_t)$ is not observable, we replace $v_t$ by its consistent estimator $\hat{\sigma}((t-1)/T)$, and we consider the test statistic $\tau(\hat{\sigma})$ defined as

align[align omitted — 123 chars of source]

where

align[align omitted — 200 chars of source]

where $\hat{u}_t$ the OLS residuals given as $\hat{u}_t = y_t - \hat{\beta} x_{t-1}$ with the OLS estimator $\hat{\beta}$. Here $K_h(t)=K(t/h)$ with a kernel function $K$ and bandwidth $h$.

The validity of our approach requires that $\hat{\sigma}(r)$ is close enough to $\sigma_T(r)=v_{[Tr]}$ for most $r\in[0,1]$. We first establish a uniform convergence result

align[align omitted — 109 chars of source]

for some $\mathcal{C}_h\subset [0,1]$. Invoking the convergence in Assumption (ref), $\sigma_T\to_d \sigma$ is interpreted as the weak convergence in the probability measures on $\mathbf{D}_{\mathbb{R}^+}[0,1]$ endowed with the uniform topology. By virtue of the so-called Skorohod representation theorem (e.g., Pollard (1984), pp. 71-72), it is indeed possible to construct $\sigma_T$ and $\sigma$ on a common probability space, up to the distributional equivalence, so that $\sigma_T\to_{a.s.}\sigma$ uniformly on $[0,1]$. For our development of the uniform convergence results (ref), we assume that $\sigma_T$ is defined up to the distributional equivalence such that $\sigma_T\to_{a.s.}\sigma$ uniformly on $[0,1]$. This assumption is not restrictive since we are interested in the convergence of $\hat{\sigma}^2$ to $\sigma_T^2$ rather than $\sigma^2$.

For the nonparametric estimator $\hat{\sigma}$, we assume the kernel function $K$ satisfies the following assumption.

ass(a) a nonnegative kernel $K$ has a compact support $[0,1]$ with $\int_0^1 K(r)dr=1$, (b) $|K(r) - K(s)|\leq \bar{K}|r-s|$ for all $r,s\in\mathbb{R}$, and $\sup_r K(r)<\bar{K}$ for some $0<\bar{K}<\infty$,

The condition (b) in Assumption (ref) is standard in the investigation of uniform consistency. In Assumption (ref) (a), we assume a nonstandard assumption that $K$ is a one-sided kernel, which is unnecessary in developing the uniform consistency (ref). When we establish $\tau(\hat{\sigma})\to_d \mathbb{N}(0,1)$ under $\beta=0$, however, it is important to make $\hat{\sigma}(t/T)$ measurable with respect to $\mathcal{F}_{t+1}$ so that we can apply a martingale CLT to $\tau(\hat{\sigma})$.\footnote{For the same reason, hansen1995regression considered a one-sided kernel.} For a more precise explanation, we write

equation[equation omitted — 143 chars of source]

where

align*[align* omitted — 469 chars of source]

In the decomposition (ref), $\hat{\sigma}_1^2(r) - \sigma_T^2(r)$ is a bias term since $E(u_t^2 | \mathcal{F}_{t-1}) = v_t^2 = \sigma_T^2(t/T)$, whereas $\hat{\sigma}_2^2(r)$ is a variance term involving a martingale. On the other hand, $\hat{\sigma}_3^2$ and $\hat{\sigma}_4^2$ are error components induced by using $\hat{u}_t$, instead of $u_t$, in the kernel estimation of $\sigma_T^2$. Under Assumption (ref) (a), $\hat\sigma_1^2(t/T)$ is $\mathcal{F}_{t-1}$-adapted, whereas $\hat\sigma_2^2(t/T)$ is $\mathcal{F}_t$-adapted. Consequently, $\tilde{\sigma}^2(t/T)$, where $\tilde{\sigma}^2 = \hat\sigma_1^2 + \hat\sigma_2^2$, is $\mathcal{F}_t$-adapted, from which one may show that $\tau(\tilde{\sigma})\to_d \mathbb{N}(0,1)$, where $\tau(\tilde{\sigma})$ is defined as $\tau(\hat{\sigma})$ but with $\hat{\sigma}$ replaced $\tilde{\sigma}$, by a martingale CLT as long as $|\tilde{\sigma}^2(r)-\sigma_T^2(r)|=o_p(1)$ for most $r\in[0,1]$. However, $\hat{\sigma}_3^2(t/T)$ and $\hat{\sigma}_4^2(t/T)$ are not $\mathcal{F}_t$-measurable since $\hat{\beta}$ is not $\mathcal{F}_t$-measurable for any $t<T$. Consequently, we cannot directly apply a martingale CLT to show $\tau(\hat{\sigma})\to_d \mathbb{N}(0,1)$. Alternatively, for these two terms, it is shown that they have negligible effects in the test statistic $\tau(\hat{\sigma})$, and we have $\tau(\hat{\sigma}) = \tau(\tilde{\sigma})(1+o_p(1))$ as long as $\hat{\beta}\to_p\beta$ sufficiently quickly. For the asymptotic negligibilities of $\hat{\sigma}_3^2$ and $\hat{\sigma}_4^2$, we assume

assFor any deterministic sequence $(c_t)_{t=1}^T$ such that $0\leq c_t\leq 1$ for all $t$, $\sum_{t=1}^T c_t x_{t-1}u_t = O_p\left(T^p \left(\sum_{t=1}^T x_{t-1}^2\right)^{1/2}\right)$ for some $p\in[0,1/8)$.

Assumption (ref) is very general and many time series models satisfy the condition. In particular, it holds with $p=0$ if $(x_t)$ is either (near) unit root or stationary with finite variance. Moreover, if $(x_t)$ is stationary with unbounded variance, then the condition holds with $p=0$ under some additional conditions on $(x_t)$ and $(u_t)$ (see, e.g., Samorodnitsky2007).

lemIf Assumptions (ref) and (ref) hold, then $|\hat{\beta} - \beta| = O_p\left(T^p \left(\sum_{t=1}^T x_{t-1}^2\right)^{-1/2}\right)$.

We will show below that the rate of convergence of $\hat{\beta}$ in Lemma (ref) is enough to obtain the required uniform convergences of $\hat{\sigma}_3^2$ and $\hat{\sigma}_4^2$ as well as their asymptotic negligibility in the test relying on the statistic (ref).

On the other hand, the convergence $|\hat{\sigma}_1^2(r)- \sigma_T^2(r)|\to_p0$ requires that $\sigma_T^2$ be left-continuous at $r$ due, in particular, to the fact that $K$ is a one-sided kernel having support $[0,1]$. However, $\sigma_T$ may have countably many jumps since $\sigma_T\in\mathbf{D}[0,1]$. In particular, at a discontinuity point $r$ with $\sigma(r)\neq \sigma(r-)$, we have $|\hat{\sigma}_1^2(r)- \sigma_T^2(r-)|\to_p0$ instead of $|\hat{\sigma}_1^2(r)- \sigma_T^2(r)|\to_p0$. Therefore, the set $\mathcal{C}_h$ in (ref) should effectively exclude a set of discontinuity points as well as its neighborhoods so that the uniform convergence result holds. Under our convention of $\sigma_T\to_{a.s.}\sigma$ uniformly on $[0,1]$, we only need to consider $\sigma$'s discontinuity points, and we define

align[align omitted — 161 chars of source]

Clearly, $\mathcal{C}_h$ is a set of left-continuity points, and we establish the uniform convergence result (ref) over $\mathcal{C}_h$.

A martingale exponential inequality can be used to show the asymptotic negligibility of the variance component $\hat{\sigma}_2^2(r)$ uniformly in $r$ (see, e.g., victor1999general and bercu2008exponential). In this paper, we use the two-sided exponential inequality in bercu2008exponential under which we can relax the moment condition for $(\varepsilon_t)$ at the cost of an assumption on the stochastic order of the extremal process of $(\varepsilon_t)$.

assFor some $q\in[0,1/8)$, $\max_{1\le t\le T}|\varepsilon_t| = O_p(T^q)$.

Assumption (ref) is not stringent, and a wide class of time series models for $\varepsilon_t$ satisfies the condition.\footnote{Instead of Assumption (ref), one may obtain the subsequent results by assuming an additional moment condition, i.e., $E|\varepsilon_t|^{4r}<\infty$ for some $r\geq1$. For a relevant approach, the reader is referred to, e.g., Theorem 2.1 of wang2014uniform.} For instance, if $\varepsilon_t$ is a Gaussian process with $cov(\varepsilon_1, \varepsilon_T)\log T\to0$, then $\max_{0\leq t\leq T}|\varepsilon_t| = O_p(\sqrt{\log T})$ and the condition (a) holds for any $q>0$ (see, e.g., Theorem 2.5.2 of leadbetter1988extremal).

assAs $h\to0$ and $T\to\infty$, (a) $h T^{1/2-2p}\to\infty$ where $p\in[0,1/8)$ is defined as in Assumption (ref), and (b) $h T^{1-4q}\to\infty$ and $hT^{2q}\to0$, where $q\in[0,1/8)$ is defined as in Assumption (ref).

Assumption (ref) provides the connections among the stochastic orders in Assumptions (ref)-(ref) and the bandwidth $h$. If we let $h=c T^{-\alpha}$ for $c,\alpha>0$ as in the typical situation, then Assumption (ref) holds for $2q< \alpha <\min\{1/2 -2p, 1-4q\}$. Note that such $\alpha$ always exists for $p,q\in[0,1/8)$. In particular, if $p=q=0$, then $h=c T^{-\alpha}$ satisfies Assumption (ref) for $0<\alpha<1/2$. We note that the condition (a) is used to guarantee $\hat{\sigma}_3^2$ and $\hat{\sigma}_4^2$ being asymptotically negligible in our inference method. In condition (b), $h T^{1-4q}\to\infty$ is needed for the uniform convergence of $\hat{\sigma}_2^2$, whereas $hT^{2q}\to0$ is used to effectively handle discontinuity points of $\sigma^2$ at which $\hat{\sigma}^2$ becomes inconsistent.

propLet Assumptions (ref)-(ref) and (ref)-(ref) hold. As $h\to0$ and $T\to\infty$, we have \begin{align*} &(a) \quad \sup_{r\in \mathcal{C}_h} |\hat{\sigma}_1^2(r)-\sigma_T^2(r)|=o_p(1), &&(b) \quad \sup_{h\leq r\leq 1} |\hat{\sigma}_2^2(r)|=O_p\left(T^{2q} \left(\log (hT)/ (hT)\right)^{1/2}\right),\\ &(c) \quad \sup_{h\leq r\leq 1} |\hat{\sigma}_3^2(r)|=O_p\left(T^{2p} /(hT)\right), &&(d) \quad \sup_{h\leq r\leq 1} |\hat{\sigma}_4^2(r)|= O_p\left(T^{2p} /(hT)\right), \end{align*} and the uniform convergence result (ref) holds.

Under Assumption (ref), we have \[ T^{2p} /(hT) = o\left( T^{2q} \left(\log (hT)/ (hT)\right)^{1/2}\right), \] from which we can see that the error components $\hat{\sigma}_3^2$ and $\hat{\sigma}_4^2$ have smaller orders than $\hat{\sigma}_2^2$. Indeed, it is shown in the proof of Theorem (ref) that $\hat{\sigma}_3^2$ and $\hat{\sigma}_4^2$ have negligible effects in the test statistic $\tau(\hat{\sigma})$, and we have $\tau(\hat{\sigma})=\tau(\tilde{\sigma})(1+o_p(1))$, where $\tilde{\sigma}^2 = \hat{\sigma}_1^2 + \hat{\sigma}_2^2$. However, the orders of $\hat{\sigma}_1^2$ and $\hat{\sigma}_2^2$ are not sufficiently small to show directly that $\tau(\tilde{\sigma}) = \tau(\sigma_T)(1+o_p(1))$ even though $\tilde{\sigma}^2$ converges uniformly to $\sigma_T^2$. That is because the convergence rate of $\hat{\sigma}_2^2(r)\to_p0$ is not fast enough for the direct approximation $\tau(\tilde{\sigma}) = \tau(\sigma_T)(1+o_p(1))$, and the convergence rate of $|\hat{\sigma}_1^2(r)-\sigma_T^2(r)|\to0$ depends on the degrees of left-continuity of $\sigma_T^2$ which are unknown in general.

Alternatively, we use the weak convergence of the stochastic integral $\int_0^1 (\sigma_T(r)/\tilde{\sigma}(r))dW_T(r)$, as in Lemma (ref), jointly with the facts that $\tilde{\sigma}(t/T)$ is $\mathcal{F}_t$-adapted and $\sup_{r\in \mathcal{C}_h}|\tilde{\sigma}^2(r)-\sigma_T^2(r)|=o_p(1)$. Here, in particular, we require $\mathcal{C}_h\to_{a.s.}[0,1]$ which holds when $\sigma$ has finitely many jumps almost surely.

ass$\sigma$ has finitely many jumps almost surely.\footnote{In other words, $\sigma$ is of finite activity in the sense that the probability measure of any set $\{\omega: r\mapsto \sigma(r,\omega) \text{ has finitely many jumps in }r\in[0,1]\}$ is one.}
thmLet Assumptions (ref)-(ref) and (ref)-(ref) hold. As $h\to0$ and $T\to\infty$, we have the following. (a) Under $\beta=0$, $\tau(\hat{\sigma})\to_d \mathbb{N}(0,1)$. (b) If $\beta\neq0$ and $\sum_{t=1}^{T-1}|x_t|/\sqrt{T}\to_p \infty$, then \[ |\tau(\hat{\sigma})|\geq \frac{|\beta|}{\underline{v}} \frac{1}{\sqrt{T}}\sum_{t=1}^{T-1}|x_t| + O_p(1)\to_p \infty, \] and hence, $P\left[|\tau(\hat{\sigma})|>c\right]\to1$ for any positive constant $c$.
remarkThe asymptotic power result in Theorem (ref) (b) implies that the testing power is mainly determined by the asymptotic behavior of $\sum_{t=1}^{T-1}|x_t|$. As an illustration, assume that there exists a diverging sequence $p_T$ such that $p_T/T^{1/2}\to\infty$ and \begin{align} \frac{1}{p_T}\sum_{t=1}^T |x_{t-1}|\to_d P \end{align} for a random variable $P>0$. Clearly, a wide class of time series satisfies (ref) (see, e.g., Example (ref)). It then follows immediately from Theorem (ref) (b) that $|\tau(\hat{\sigma})|\to_p\infty$ and the speed of divergence is no slower than $p_T/\sqrt{T}$. Heuristically, we may consider the power of the proposed test by considering local alternatives in which $\beta\neq0$, but $\beta\to0$ at an appropriate rate. For our purpose, let $(P_T(r), 0\leq r\leq 1)$ be a continuous time process defined as \[ P_T(r) = \frac{1}{p_T}\sum_{t=1}^{[Tr]} \frac{|x_{t-1}|}{\sigma_T(t/T)} \] for a diverging sequence $p_T$ such that $p_T/T^{1/2}\to\infty$. We then assume instead of (ref) that \begin{align} P_T \to_d P \end{align} for a stochastic process $(P(r),0\leq r\leq 1)$ having a positive support. If the convergence (ref) holds jointly with the convergence in Assumption (ref), then we may develop the power of the proposed test by considering the following alternative hypothesis \begin{align} \beta = \bar{\beta} \times \frac{\sqrt{T}}{p_T} \end{align} for a constant $\bar{\beta}\in\mathbb{R}\setminus \{0\}$. Clearly, the hypothesis (ref) can be interpreted as a local alternative since $\beta\neq0$ and $\beta\to0$. Our construction of the local alternative (ref) is useful to develop the asymptotic power result in a unified framework, especially, when the covariate $(x_t)$ is general, but satisfies (ref). Under the local alternative (ref), one can easily deduce from the proof of Theorem (ref) with (ref) that \[ \tau(\hat{\beta})\to_d \bar{\beta}P(1) + \mathbb{N}(0,1). \] Clearly, $\tau(\hat{\sigma})$ under (ref) is not Gaussian asymptotically unless $P(1)$ is either constant or Gaussian. \begin{example} Let $\sigma_T(r)=\sigma$ for all $r\in[0,1]$. (a) $(x_t)$ be a stationary process such that $T^{-1}\sum_{t=1}^T |x_t|\to_p E|x_1|<\infty$. Under (ref) with $p_T=T$, $\tau(\hat{\sigma})\to_d \mathbb{N}\big((\bar{\beta}/\sigma)E|x_1|,\,\, 1\big)$. (b) $(x_t)$ be a unit root (or near unit root) process such that $T^{-3/2}\sum_{t=1}^T |x_t|\to_d \int_0^1 |X(r)|dr$, where $(X(r))$ is the limiting Brownian motion (or Ornstein-Uhlenbeck process) of $(x_t)$ such that $X_T\to_d X$ for $X_T(r) = T^{-1/2}x_{[Tr]}$.\footnote{The reader is referred to PhillipsNear and park2003strong for more discussion about the near unit root process and its limiting behaviors. Here, in particular, the Ornstein-Uhlenbeck process $X$ follows \[ dX(r) = c X(r) dr + dV(r), \quad X(0)=0, \] where $V$ is Brownian motion.} Under (ref) with $p_T=T^{3/2}$, $\tau(\hat{\sigma})\to_d (\bar{\beta}/\sigma)\int_0^1 |X(r)|dr + \mathbb{N}(0, 1)$. \end{example} When $\sigma_T(r)=\sigma$ for all $r\in[0,1]$ and $(x_t)$ is stationary with $E(x_t^2)<\infty$, the usual $t$-test procedure is a valid hypothesis testing procedure for the model (ref). In this case, the asymptotic power property of the usual $t$-test is also well known under the local alternative hypothesis (ref) with $p_T = T$, and is given by \[ \text{$t$-statistic}\to_d \mathbb{N}\big((\bar{\beta}/\sigma) (E(x_t^2))^{1/2},\,\, 1\big). \] The ratio of the asymptotic biases of the $t$-test and our test, obtained in Example (ref) (a), is given by $(E(x_t^2))^{1/2} / E|x_t|$. Importantly, the ratio is always greater than one as long as $E(x_t^2)<\infty$ due to Jensen's inequality. This implies that the usual $t$-test is more powerful than our test under the ideal assumptions, even though the statistics in both tests diverge at the same rate $T^{1/2}$ under a fixed alternative hypothesis. However, when one or more of the ideal assumptions are violated, our test remains valid, whereas the usual $t$-test becomes invalid. This is another example of the traditional issue of trade-off between efficiency and robustness.
remarkA number of works in statistics and econometrics have focused on robust inference using sign tests applied to different models, including time series regressions (see, among others, DH, CD, SoShin1, delaPena, IB, kim-meddahi-2019, and references therein). For instance, CD propose sign tests for testing independence of a zero median time series $Y_t$ with $P(Y_t=0)=0,$ e.g., a time series with continuous distributions symmetric about zero, of past values of $Y_t$ and another time series $X_t.$ The tests in CD are based on the observation that, under the above independence/orthogonality hypothesis, for any $T\ge 1,$ the sign statistic like $S_0=0.5(\sum_{t=1}^T sign(Y_tX_{t-1})+T)$ and its more general analogues follow a Binomial distribution with parameters $T$ and 0.5: $S_0\sim Bi(T, 0.5)$ (the results in IB imply that sign tests for general zero median or symmetric processes $Y_t$ can be based on similar statistics with randomization over zero values of $Y_t$). Efron, Edelman, Pinelis, DH, and delaPena consider related testing procedures based on bounds for tail probabilities of $t$-statistics of a parameter of interest (e.g., a location parameter or a regression/autoregression coefficient) under symmetry assumptions implied by (sharp) bounds on tail probabilities of weighted sums of i.i.d. symmetric Bernoulli r.v.'s. Naturally, in the time series regression context, the above sign-based inference approaches are more robust to moment assumptions and heavy tails than the inference procedures based on the Gaussian asymptotics for the full-sample OLS and Cauchy estimators. Typically, the sign-based tests can be used without any moment conditions on the time series considered, e.g., under infinite variances. However, they usually require symmetry or zero median assumptions on the processes. Such assumptions are often too restrictive in empirical applications, including the analysis of financial markets due to the stylized fact of gain-loss asymmetry in financial returns (see, among others, cont and references therein). Further, sign-based tests are less efficient than those on the Gaussian asymptotics for the OLS estimator under the validity of the latter tests.
remarkOur method can be applied to a discrete time model and a discrete sample collected from an underlying continuous time model as in CJP2016. The main difference between our approach to CJP2016 is that we do not require the assumption $\delta\to0$, where $\delta$ is the sampling interval of the discrete samples. Clearly, the method of CJP2016 is applicable to high frequency data. Therefore, we may say that our method is more flexible since it can be applied to both high and low frequency data. The price we have to pay for the flexibility is the persistent volatility assumption $\sigma_T \to_d \sigma$ in Assumption (ref). Persistent volatility is a well-known stylized fact of financial time series and, in our view, is best considered within the model formulation. Our method is also comparable to the IVX approach proposed by PhillipsMagdalinos2009. The IVX approach is based on a self-generated instrument obtained by differencing the predictor $x_t$ and using an autoregressive filter to construct the instrument. As is shown in PhillipsMagdalinos2009, the IVX approach is robust to a (near) unit root or mildly explosive predictor. The Cauchy estimator approaches to inference, including ours and CJP2016, are restricted to a single regressor.\footnote{In general, a test relying on a single regressor exhibits size distortion when some relevant regressors are omitted. To overcome the issue induced by a single regressor in our approach, one may extend our approach to a multivariate setting based on the recent paper by shephard2020 in which a multivariate extension of the Cauchy estimator is proposed. An alternative extension is to use the parsimonious system approach (see Ghysels-Hill-Motegi-2020 and Xu-Guo-2022) which is based on a set of misspecified regression models with only one group of regressors, allowing a single regressor for each regression. We leave these extensions for future research.} Unlike the Cauchy based inferences, the IVX approach is applicable to predictive regressions with multiple regressors. We also note that the IVX approach allows for conditional heteroskedasticity. However, to our knowledge, it is not known whether the IVX approach is valid when the volatility is persistent or the predictor is heavy-tailed with infinite second moments, and a continuous time extension of the IVX approach is not available in the literature. Therefore, our method and the IVX may be regarded as complementing each other. hansen1995regression provides a nonparametric GLS method for regression models with nonstationary volatility using the estimator $\hat{\sigma}$ to correct the heteroskedasticity. One should note that the assumptions on the limiting volatility $\sigma$ are more general than those in hansen1995regression and other work in the literature on the topic. In particular, the assumptions in hansen1995regression do not allow for structural changes or regime switching in the volatility process as the limiting volatility is assumed to have continuous sample paths almost surely. In contrast, the limiting volatility is allowed to have an arbitrary number of jumps in this paper, and hence, structural changes or regime switching are allowed. Moreover, we further extend our model to have a two-factor volatility in Section 4.

An Extension to Two-Factor Volatility Models

In this section, we generalize the model (ref) to have a two-factor volatility in the regression error $(u_t)$. More specifically, we assume that $(\varepsilon_t)$ is conditionally heteroskedastic, rather than conditional homoscedastic as is assumed in Assumption (ref) (b).

ass(a) $E(\varepsilon_t^2|\mathcal{F}_{t-1}) = w_t^2$ and $E(w_t^2)=1$, (b) $\max_{t\geq 1}E(|w_t|^{2\eta_1})<\infty$ for some $\eta_1>2$, (c) $(w_t)$ is $\alpha$-mixing such that the mixing coefficient $\alpha$ satisfies $\alpha(k)\leq Ak^{-\eta_2}$ for some $A<\infty$ and $\eta_2 > (2\eta_1 + 2)/(\eta_1 - 2)$, and (d) $hT^{\eta_3}\to\infty$ for some $\eta_3> (\eta_2(1-2/\eta_1) - 2/\eta_1 -2)/(\eta_2 +2)$.

Under Assumptions (ref) (a) and (ref), the regression error $(u_t)$ in the model (ref) can be written as $u_t = v_t w_t e_t$, where $(e_t)$ is an MDS with respect to $(\mathcal{F}_t)$ such that $E(e_t^2 | \mathcal{F}_{t-1})=1$. Clearly, $(u_t)$ has two volatility factors, $(v_t)$ and $(w_t)$, where $(v_t)$ is the long run component by Assumption (ref) and $(w_t)$ is the short run component by Assumption (ref) (c). Moreover, under Assumption (ref) and the condition $E(w_t^2)=1$ in Assumption (ref), we can identify and estimate the persistent volatility component $v_t$ by the nonparametric estimator (ref). In particular, under Assumption (ref), we may show that \[ \sup_{h\leq r\leq 1} \left|\frac{1}{hT}\sum_{t=1}^T (w_t^2 - 1) K_h(r-t/T)\right| = O_p\left((\log T/ (hT))^{1/2}\right) \] using an exponential inequality for a strongly mixing process (see, e.g., liebscher1996strong, vogt2012nonparametric) and kristensen2009uniform). The above uniform convergence result for the mixing process is sufficient to develop the required uniform convergence of the volatility estimator as well as the validity of the inference procedure proposed in Section 3. We also note that the conditions for $h$ and $T$ in Assumption (ref) (d) and Assumption (ref) hold simultaneously for any $p,q\in[0,1/8)$ as long as Assumption (ref) (d) holds for some $\eta_3>1/4$, which is not stringent. For instance, if $(w_t)$ is a stationary GARCH(1,1) process and $\beta$-mixing with exponential decay, which hold under some mild conditions (see, e.g., carrasco2002mixing and francq2006mixing), then both Assumption (ref) (d) and Assumption (ref) hold for any $p,q\in[0,1/8)$ since Assumption (ref) (d) holds for any $\eta_3 >1$.

For our purpose, we again consider the decomposition (ref) of the nonparametric estimator (ref), and write $\hat{\sigma}^2 = \sum_{k=1}^4 \hat{\sigma}_k^2$. We then can obtain the uniform convergence rate of each component $\hat{\sigma}_k^2$ for $k=1,2,3,4$ as in Proposition (ref), and establish the validity of the inference method relying on the test statistic (ref) as in Theorem (ref).

corLet Assumptions (ref) (a) and (c), (ref), (ref)-(ref) and (ref) hold. As $h\to0$ and $T\to\infty$, Proposition (ref) and Theorem (ref) remain valid.

Monte-Carlo Simulations

This section provides the numerical results on finite sample performance of the proposed robust test based on $\tau(\hat{\sigma})$. We present the comparisons of the finite sample properties of the test with the test proposed by CJP2016 (denoted as Cauchy RT; RT for random time) and also two other tests considered in CJP2016: the Bonferroni $Q$-test of CampbellYogo2006 (denoted as BQ) and the restricted likelihood ratio test of ChenDeo2009 (denoted as RLRT).

We consider two different settings for simulation models: continuous time and discrete time DGPs. As for the continuous time DGPs, we follow the simulation designs of CJP2016. The data is generated in the continuous time setting using the following DGP:

eqnarray[eqnarray omitted — 214 chars of source]

where $W_{1t}$ and $W_{2t}$ are Brownian motions with $E(W_{1t}W_{2t})=-0.98t$. We set the constant term in the predictive regression to be zero and use recursive de-meaning. We assume that the continuous time models are observed at $\delta$-intervals over $T$ years with $\delta=1/252$, which corresponds to daily observations of size $252 T$.

The volatility process considered in the numerical results is assumed to follow one of the following models:

itemize• Model CNST. Constant volatility: $ \sigma_t^2=\sigma_0^2$, $\sigma_0=1$. • Model SB. Structural break in volatility: $\sigma_0+(\sigma_1-\sigma_0) 1\{t/T\geq 4/5\}$ with $\sigma_0=1$ and $\sigma_1=4$. • Model GBM. Geometric Brownian motion: $d\sigma_t^2=\frac{1}{2}\frac{\bar{\omega}^2}{T}\sigma_t^2dt+\frac{\bar{\omega}^2}{\sqrt{T}}\sigma_t^2dZ_t$, where $Z_t$ is a Brownian motion with $E(W_{1t} Z_t) = -0.4 t$, and $\bar{\omega} = 9$. • Model RS. Regime switching: $\sigma_t=\sigma_0(1-s_t)+\sigma_1s_t$, where $s_t$ is a homogeneous Markov process indicating the current state of the world which is independent of both $Y_t$ and $X_t$ with the state space $\{0,1\}$ and the transition matrix \[P_t=\begin{pmatrix} 0.8 & 0.2\\ 0.8 & 0.2 \end{pmatrix} + \begin{pmatrix} 0.2 & -0.2\\ -0.8 & 0.8 \end{pmatrix}\exp\left(-\frac{\bar{\lambda}}{T}t\right), \] where $\bar{\lambda}=60$, $\sigma_0=1$ and $\sigma_1=4$. The process $s_t$ is initialized by its invariant distribution.

We set the number of years $T\in\{5,20,50\}$ (which corresponds to 60, 240 and 600 monthly data) and consider the values $\bar{\kappa}\in\{0,5,10\}$ for the persistence parameter $\bar{\kappa}$ of $X_t$ in ((ref)).

As indicated before, in contrast to the Cauchy RT test in CJP2016, our test is applicable, not only in the continuous time models, but also in the discrete time framework. We consider the following discrete time models in the analysis of the finite sample performance of the tests:

eqnarray[eqnarray omitted — 175 chars of source]

for $t=2, \dots,T$, where $T\in\{60, 240, 600\}$ (the same number of monthly observations as in continuous time simulations) and the same values of $\bar{\beta}$ and $\bar{\kappa}$. Here the innovations $(\varepsilon_t, \eta_t)$ are assumed to be multivariate normal with the correlation coefficient $-$0.98.

For the volatility processes in the discrete time setting, we consider three specifications: Model CNST and Model SB as in the continuous time setup, and GARCH volatility dynamics with

align*[align* omitted — 176 chars of source]

In the numerical analysis, we consider the ARCH(1) processes with $\theta=0,$ $\alpha=0.5773$ (stationary with infinite fourth moment); $\theta=0,$ $\alpha=0.7325$ (stationary with infinite third moment); IGARCH(1,1) models with $\alpha=0.9$, $\theta=0.1$ and $\alpha=0.1$, $\theta=0.9$ (nonstationary). Note that the ARCH(1) processes in our simulations violate the moment conditions in Assumption (ref). As shown in our simulation results below, our approach has reliable size and power properties even though the required moment conditions are violated.\footnote{See, among others, MS, DM, IPS, and references therein for the results on moment properties of GARCH processes and their importance in robust econometric inference.}

Finite Sample Size Properties

In this section, we analyze finite sample size properties of the no predictability tests by setting $\bar{\beta}=0$ in the regression models (ref) and (ref). The numerical results on the finite sample size properties are presented in Tables (ref)-(ref).

Table (ref) provides the finite sample size results for models CNST, SB, GBM and RS in the continuous time setting. The finite sample size values for the OLS, BQ, RLRT and Cauchy RT tests are exactly the same as those reported in CJP2016. These numerical results show that the size of the OLS, BQ and RLRT tests is highly distorted for most of the time-varying volatility models considered. In contrast, the rejection probabilities of the proposed test are very close to their nominal levels, such as the Cauchy RT test, regardless of the values of $\bar{\kappa}$ and $T$, and the volatility models we consider in our simulations. For the $5\%$ test, rejection probabilities stay between $4\%$ to $8\%$ without any exception.

As mentioned before, the Cauchy RT test is inapplicable in the discrete time settings. Table (ref) provides the numerical results on finite sample size properties of all the tests except Cauchy RT under the discrete time settings. The quantitative and qualitative comparisons of the size properties of the tests are similar to the continuous time case. In summary, the finite sample size properties reported in Tables (ref)-(ref) show that the proposed test has a reliable size performance and is widely applicable for both discrete and continuous time settings.

Finite Sample Power Properties

Figures (ref)-(ref) present the results on finite sample power properties of the tests considered.\footnote{We report the power properties for Models SB, GBM and RS (continuous time) as well as Models SB and GARCH (discrete time) in Section 5.2. The power properties for the other models, Model CNST (continuous time) as well as Models CNST and ARCH (discrete time), are presented in the Supplementary Online Appendix.} In our simulations, we consider the DGPs in (ref) in continuous time and (ref) in discrete time with $\bar{\beta}$ ranging from 0 to 20. All the power curves presented in the figures are size-adjusted. Taking into account the results of finite sample size performance of the tests and their comparisons, we mainly focus on two tests: the Cauchy RT and our test, in the analysis of finite sample power properties. For comparison, the analysis also provides the numerical results on the finite sample power of the OLS, BQ and RLRT tests.

In Figure (ref) for the case of the structural break in volatility, one observes that the Cauchy RT test appears to be superior to other testing approaches (except in the cases with $\bar{\kappa}=0$ and large $\bar{\beta}$). At the same time, the proposed test $\tau(\hat{\sigma})$ also appears to have good finite sample power properties especially in the case of highly persistent predictors.

Figure (ref) provides the numerical results on finite sample properties of the tests in the geometric Brownian motion case. For the case of a unit root regressor, the power properties of the proposed test based on $\tau(\hat{\sigma})$ appear to outperform those of the Cauchy RT test which in turn outperforms other tests considered. However, the finite sample power performance of the Cauchy RT test improves in the case of near unit root regressor with $\bar{\kappa}=5$ and $\bar{\kappa}=20$.

The power curves for the regime switching case presented in Figure (ref) demonstrate that the test based on the proposed test $\tau(\hat{\sigma})$ has better power properties than other tests in the case $\bar{\kappa}=0.$ For the case of the near unit root persistence in the regressor, the power properties of the Cauchy RT test appear to be better than those of the test based on $\tau(\hat{\sigma})$ for relatively small sample sizes (small values of $T$). However, as the sample size increases, the test based on $\tau(\hat{\sigma})$ becomes more powerful than the Cauchy RT test (see, e.g., Figure (ref) for the case $\bar{\kappa}=5$ and $T=50$). For large deviations from a unit root regressor ($\bar{\kappa}=20$), the Cauchy RT is more powerful than other tests, but the power curves appear to be very similar.

Figures (ref)-(ref) present the numerical results on power properties under discrete time settings for all the tests considered except Cauchy RT which is inapplicable in discrete time settings. Results in the figures are provided for the cases of the structural break in volatility (Figure (ref)); the GARCH cases with $(\alpha,\theta)=(0.9,0.1)$ (Figure (ref)) and $(\alpha,\theta)=(0.1,0.9)$ (Figure (ref)). For all the cases, the conclusions on power properties of the tests and their comparisons are virtually the same as in the continuous time framework.

Overall, the numerical results on finite sample properties of the tests indicate good performance of the test based on $\tau(\hat{\sigma})$ in comparison to the Cauchy RT. Again, the latter test is inapplicable in the discrete time settings. Their relative finite sample performances vary across different models. Which test should be used in practice depends on the availability of high frequency data as well as the size-power trade-off for a specific model. Therefore, the test proposed in this paper and the Cauchy RT complement rather than substitute one another.

{0.3ex} \newcolumntype{C}{>{\arraybackslash}X}

Conclusion

Endogenously persistent regressors have been extensively analyzed in the predictive regression literature. A widely believed characteristic of stock returns is heteroskedastic and persistent volatility, which is often ignored in the predictive regression literature except for CJP2016. These two characteristics cause standard hypothesis tests to become substantially biased and often over-reject the null of no predictability. The main contribution of this paper is to provide an inference method that is designed to be robust to these problematic characteristics of predictive regression data. The proposed method relies on the Cauchy estimator and a kernel-based nonparametric correction of volatility. Its theoretical validity is provided by analyzing the asymptotic size and power properties. Moreover, it is shown through a simulation study that the proposed method has a reliable finite sample performance compared to the most advanced existing inference methods.

Our inference method is comparable to the method proposed by CJP2016. Similar to our method, their approach relies on the Cauchy estimator and a nonparametric volatility correction. However, their approach to the volatility correction is quite different from ours, and its applicability is limited to a predictive regression equipped with appropriate high frequency data. In terms of finite sample properties, our method and the method by CJP2016 perform well and have good size and power performances under continuous time settings. However, unlike our method, CJP2016 method is not applicable under discrete time settings. In contrast, our method can be applied to a discrete time model as well as a discrete sample collected from an underlying continuous time model. Therefore, our method is more flexible and widely applicable since it can be applied to both high and low frequency data.

A further approach to robust inference in predictive regressions under heterogeneous and persistent volatility as well as endogenous, persistent or heavy-tailed regressors is provided by the simple to implement robust $t$-statistic inference approach (see IbragimovMuller2010) based on asymptotically normal group Cauchy estimators of a regression parameter of interest. This approach will be explored in a companion paper now in preparation.