Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
57,026 characters · 19 sections · 154 citation commands
Robust Cauchy-Based Methods for Predictive Regressions
\noindentKeywords: predictive regressions, robust inference, near nonstationarity, heterogeneity, heavy tails, persistent volatility, endogeneity.
JEL Codes: C12, C22
Predictive regressions play a central role in empirical finance, providing a framework for assessing whether financial or macroeconomic variables can forecast future returns. Prominent applications include the forecasting of equity and aggregate returns (see, among others, CampbellYogo2006, CampbellYogo2006; GoyalWelch, GoyalWelch; campbell, campbell; Hirshleifer, Hirshleifer; KJ, KJ; Rapach, Rapach; MR, MR; Goyal2024, Goyal2024, and references therein) and tests of market efficiency (e.g., Fama1, Fama1, Fama, Fama2; the review in Martin, Martin, and references therein). A large body of work has examined the econometric properties of predictive regressions (see Phillips2015, Phillips2015 for a review), highlighting several statistical challenges that complicate inference on return predictability.
A key difficulty arises from the persistence and endogeneity of commonly used predictors. Variables such as the dividend--price and earnings--price ratios typically exhibit near–unit–root behavior, and their innovations are correlated with future returns over long horizons. This combination induces substantial biases in conventional hypothesis tests (see, e.g., stambaugh1999predictive, stambaugh1999predictive; KimPark1, KimPark1). In addition, stock return volatility is stochastic and highly persistent (jacquier2004bayesian, jacquier2004bayesian; hansen2014estimating, hansen2014estimating), and cavaliere2004testing shows that such volatility can lead to severe size distortions in tests that rely on stationarity. Predictive regression data also often exhibit heavy tails, jumps, structural breaks, and regime switching, further undermining standard inference methods.
A large literature has proposed procedures to address persistent endogeneity in predictive regressions. Notably, CampbellYogo2006, ChenDeo2009, PhillipsMagdalinos2009, and KMS2015, among others, develop inference methods that perform well in the presence of persistent and endogenous regressors. However, these approaches typically rely on assumptions such as homoskedasticity or weak dependence in volatility and may not remain valid under empirically relevant features such as persistent or nonstationary volatility, structural breaks, or regime switching. For example, ibragimov2023new show that standard tests can suffer from severe size distortions in the presence of persistent volatility.
To address these limitations, CJP2016 propose an inference method (the Cauchy RT) based on the Cauchy estimator and a time-change transformation in a continuous-time framework.\footnote{See also BuKimWang for an alternative method robust to endogenously persistent or heavy-tailed regressors and persistent volatility in continuous time.} ibragimov2023new introduce another approach (the Cauchy VC), also based on the Cauchy estimator but with a nonparametric volatility correction, which applies to both continuous- and discrete-time models.
This paper proposes two practical tests that serve as robust alternatives to these methods. The proposed procedures are robust to heterogeneous and persistent volatility, as well as to endogenous, persistent, and heavy-tailed regressors. Both methods employ Cauchy estimation, as in CJP2016 and ibragimov2023new, to address endogeneity, persistence, and heavy tails. The two approaches differ in their treatment of heterogeneous volatility: the first extends the $t$-statistic–based group inference of IbragimovMuller2010 to asymptotically normal Cauchy estimators, while the second is a hybrid procedure that combines Cauchy and OLS estimation by using the Cauchy estimator for the coefficient and OLS residuals for variance estimation.
The proposed methods are simple to implement and avoid the technical complexities associated with existing approaches, such as the time-change transformation in CJP2016 or the nonparametric volatility correction in ibragimov2023new. Although they rely on an asymptotic exogeneity condition on volatility, simulation results show that they perform well in finite samples, even under mild violations of this condition. Moreover, the proposed methods apply to both continuous- and discrete-time models. Overall, our procedures complement existing Cauchy-based methods and provide a practical alternative for robust inference under a broad range of empirically relevant environments.
In addition to the continuous-time framework, we evaluate the performance of our methods in discrete-time predictive regressions and compare them with the IVX approach of KMS2015. The IVX method is specifically designed to address persistence and endogeneity and performs well under homoskedastic or weakly dependent volatility. Our results highlight that the two approaches are complementary. While IVX delivers strong performance under standard conditions, it may suffer from size distortions in the presence of nonstationary or persistent volatility. By contrast, our methods are designed to remain robust to such features, including heavy tails and time-varying volatility. This comparison underscores the advantages of each approach across empirically relevant settings.
The remainder of the paper is organized as follows. Section 2 introduces the predictive regression model and the Cauchy estimator. Section 3 develops the proposed inference procedures and establishes their theoretical properties. Section 4 extends the analysis to models with multiple predictors and intercepts. Sections 5 and 6 present simulation results and an empirical application. Section 7 concludes. All proofs are provided in the Appendix, and additional theoretical analysis, as well as supplementary simulation and empirical results, are reported in the accompanying Online Appendix.
Throughout the paper, we consider $(\mathcal{F}_t)$-adapted processes defined on a filtered probability space $(\Omega,\mathcal{F}, (\mathcal{F}_t), P)$ equipped with an increasing filtration $(\mathcal{F}_t)$ of sub-$\sigma$-fields of $\mathcal{F}$. Our objective is to test the (un)predictability of the process $(y_t)$ (e.g., the time series of excess stock returns) based on a covariate process $(x_t)$ (e.g., the time series of price–dividend ratios). As usual, we consider the linear predictive regression model
Following the standard specification for a volatility model, we assume that
where $(v_t)$ is a volatility process and $(\varepsilon_t)$ is a martingale difference sequence (MDS) with respect to $(\mathcal{F}_t)$. We impose the following regularity conditions on $(v_t, \varepsilon_t)$.
Conditions (a) and (b) are standard and ensure that the conditional variance of $u_t$ is identified: $E(u_t^2|\mathcal{F}_{t-1})=v_t^2$. Condition (c) is a conditional Lindeberg condition, which holds, for example, if $\sup_{1\le t\le T} E(|\varepsilon_t|^{2+\delta}|\mathcal{F}_{t-1})$ is bounded for some $\delta>0$ with probability one. See ibragimov2023new and references therein for further discussion and examples of processes satisfying Assumption (ref).
The hypothesis of unpredictability of $(y_t)$ corresponds to $H_0:\beta=0$ in regression (ref). It is well-known that standard OLS $t$-statistic inference is not robust to many empirically relevant features of financial data. For instance, the OLS estimator of $\beta$ is not asymptotically Gaussian under $H_0$ if $(x_t)$ is endogenous and (nearly) nonstationary (see elliott1994inference, elliott1994inference; PhillipsNear, PhillipsNear; GP, GP; PM, PM; KMS2015, KMS2015), or even if $(x_t)$ is stationary but has infinite variance (e.g., GrangerOrr, GrangerOrr; EKM, EKM; ibragimov2015heavy, ibragimov2015heavy). This non-Gaussianity persists even when the errors are homoskedastic with $v_t^2=\sigma^2$ for all $t$.\footnote{As usual, the endogeneity of $(x_{t-1})$ refers to the existence of nonzero long-run covariance between the innovations of $(y_t)$ and $(x_{t-1})$.} Furthermore, stock return data exhibit time-varying and stochastically persistent volatility, which causes the distribution of the OLS $t$-statistic to deviate from standard normality, leading to size distortions in conventional tests (see CJP2016, CJP2016; ibragimov2023new, ibragimov2023new).
Both inference methods proposed in this paper build upon the following Cauchy estimator of $\beta$ (assuming no intercept, i.e., $\alpha=0$):
where $\text{sign}(\cdot)$ denotes the sign function, $\text{sign}(x)=1$ for $x\ge0$ and $\text{sign}(x)=-1$ for $x<0$. The estimator $\check{\beta}$ can be interpreted as an instrumental variable (IV) estimator using $\text{sign}(x_{t-1})$ as an instrument for $x_{t-1}$ (see, e.g., SoShin, SoShin; BREITUNG2015, BREITUNG2015; kim-meddahi-2019, kim-meddahi-2019; shephard2020, shephard2020).
Under Assumption (ref), $\text{sign}(x_{t-1})\varepsilon_t$ (denoted by $\xi_t$) is a unit-variance MDS with respect to $(\mathcal{F}_t)$. Define the continuous-time partial sum process $(W^T(r), 0\le r\le1)$ by \[ W^T(r) = T^{-1/2} \sum_{t=1}^{[Tr]} \xi_t, \] which takes values in $\mathbf{D}_{\mathbb{R}}[0,1]$, the space of c\`adl\`ag functions on $[0,1]$ with values in $\mathbb{R}^d$. By the functional central limit theorem for martingales (Theorem 18.2 of billingsley1986convergence, billingsley1986convergence), we have $W^T \Rightarrow W$ in $\mathbf{D}_{\mathbb{R}}[0,1]$, where $W$ is a standard Brownian motion.
For the volatility process $(v_t)$, define $\sigma^T(r)=v_{[Tr]}$ on $\mathbf{D}_{\mathbb{R}^+}[0,1]$. Then the Cauchy estimator can be expressed in terms of $\sigma^T$ and $W^T$ as
Following ibragimov2023new, we assume that the volatility process $\sigma^T$ is persistent in the sense that it admits a limiting process $\sigma$ defined on $[0,1]$ such that $(W^T,\sigma^T)\Rightarrow(W,\sigma)$ jointly.
Assumption (ref) encompasses a wide class of models, including those with nonstationary volatility, regime switching, or structural breaks.\footnote{Assumptions (ref) and (ref) exclude some globally homoskedastic processes, such as stationary GARCH models. However, the hybrid testing procedure proposed later remains valid under $T^{-1}\sum_{t=1}^T v_t^2\to_p\omega^2>0$, which includes conditionally heteroskedastic but globally homoskedastic processes, such as stationary GARCH models (see also Section 4 of ibragimov2023new, ibragimov2023new).} It also covers cases with deterministic volatility $v_t=\sigma(t/T)$, as in cavaliere-taylor-2007,cavaliere-taylor-2008, xu-phillips-2008, and harvey-leybourne-zu, among others.\footnote{Assumption (ref) is a simplified version of the condition $v_{[Tr]}/a_T\Rightarrow\sigma_r$ in Assumption 2 of cavaliere-taylor-2009. We focus on stochastically bounded volatilities with $a_T=1$, excluding explosive volatility settings ($a_T\to\infty$) for simplicity.} It further includes nonstationary volatility processes such as those in hansen1995regression and chung2007nonstationary, where $v_t^2$ is a smooth positive transformation of a (near) unit root process. Overall, Assumptions (ref) and (ref) are general enough to allow for stochastic and discontinuous volatility—features commonly observed in financial returns.
Under Assumptions (ref) and (ref), the properly normalized Cauchy estimator satisfies \[ \left(\sum_{t=1}^T |x_{t-1}|/\sqrt{T}\right)\!(\check{\beta}-\beta) \Rightarrow \int_0^1 \sigma(r)\,dW(r), \] by standard results on the convergence of stochastic integrals (see hansen-1992, hansen-1992; kurtz1991weak, kurtz1991weak; ibragimov2023new, ibragimov2023new). The limit $\int_0^1 \sigma(r)\,dW(r)$ is in general a non-Gaussian martingale, becoming Gaussian only if $W$ and $\sigma$ are independent. In that case, $\int_0^1 \sigma(r)\,dW(r)$ is a scale mixture of normals with variance $\int_0^1 \sigma^2(r)\,dr$, denoted \[ \int_0^1 \sigma(r)\,dW(r)\;=_d\; \mathbb{MN}\!\left(0, \int_0^1 \sigma^2(r)\,dr\right). \] We formalize the independence assumption as follows.
Assumption (ref) requires the volatility process $\sigma^T$ to be asymptotically independent of the martingale $W^T$, but does not preclude finite-sample dependence. For example, consider \[ \sigma^T\!\left(t/T\right) = T^{-\delta} f(x_{t-1}, \varepsilon_t) + \sigma_0^T\!\left(t/T\right), \quad \delta>0, \] where $f:\mathbb{R}^2\to\mathbb{R}^+$ is bounded and $\sigma_0^T$ is independent of $W^T$ with $(W^T,\sigma_0^T)\Rightarrow(W,\sigma)$, where $W$ and $\sigma$ are independent. For any $\delta>0$, the volatility process $\sigma^T$ in this example satisfies Assumption (ref), even though $\sigma^T$ and $W^T$ may be dependent for any fixed $T>0$.
In the following sections, we develop inference methods based on the Cauchy estimator. Section 3 focuses on predictive regressions with a single predictor and no intercept, while Section 4 extends the analysis to models with multiple predictors and an intercept.
The first approach relies on $t$-statistic-based inference using group estimates of $\beta$, as proposed by IbragimovMuller2010 (see also IM1, IM1; Section 3.3 of ibragimov2015heavy, ibragimov2015heavy). The method is based on normalized Cauchy estimators—specifically, the numerator of the Cauchy estimator divided by $\sqrt{T}$ in the full-sample case:
Following the $t$-statistic approach, we partition the sample into a fixed number $q\ge2$ of approximately equal groups of consecutive observations. The observation $(y_t,x_{t-1})$ at time $t$ belongs to the $j$th group $\mathcal{G}_j$ if \[ t\in\mathcal{G}_j = \{s:(j-1)[T/q]<s\le j[T/q]\}, \quad j=1,\dots,q. \] We compute the normalized Cauchy statistic in (ref) within each group:
The $t$-statistic based on the $q$ group statistics $\{\check{\gamma}_j\}_{j=1}^q$ is given by
where \[ \bar{\gamma} = q^{-1}\sum_{j=1}^q \check{\gamma}_j, \qquad s_{\gamma}^2 = (q-1)^{-1}\sum_{j=1}^q (\check{\gamma}_j-\bar{\gamma})^2. \] Under the null hypothesis $H_0:\beta=0$, the test rejects $H_0$ in favor of $H_A:\beta\neq0$ if $|t_q(\check{\gamma})|>cv_q(\alpha)$, where $cv_q(\alpha)$ denotes the two-sided $t$-critical value at level $\alpha$, i.e.\ $P(|T_{q-1}|>cv_q(\alpha))=\alpha$ for $T_{q-1}\sim t_{q-1}$ (one-sided tests are analogous).
To study the asymptotic behavior of $\{\check{\gamma}_j\}_{j=1}^q$, we decompose \[ \check{\gamma}_j = \zeta_j + \psi_j, \] where \[ \zeta_j = \beta\sqrt{\frac{q}{T}}\sum_{t\in\mathcal{G}_j}|x_{t-1}|, \qquad \psi_j = \sqrt{\frac{q}{T}}\sum_{t\in\mathcal{G}_j}\text{sign}(x_{t-1})u_t. \] Under Assumption (ref), $\{\psi_j\}_{j=1}^q$ forms a sequence of martingale differences uncorrelated across groups, yielding the following asymptotic characterization.
The statistics $\{\check{\gamma}_j\}_{j=1}^q$ do not satisfy the standard condition in IbragimovMuller2010, which requires estimators $\{\tilde{\beta}_j\}_{j=1}^q$ such that \[ \{m_T(\tilde{\beta}_j-\beta)\}_{j=1}^q \to_d \{V_j Z_j\}_{j=1}^q, \] for some $m_T\to\infty$, $Z_j\stackrel{iid}{\sim}\mathbb{N}(0,1)$, and $\{V_j\}$ independent of $\{Z_j\}$. By contrast, Lemma (ref) shows that $\{\check{\gamma}_j\}_{j=1}^q$ lack such a diverging normalization. Consequently, as shown in Proposition (ref), the $t$-statistic approach yields correct asymptotic size but is consistent only for a restricted class of covariates, excluding (near) unit-root processes. This inconsistency arises precisely because the asymptotics of $\{\check{\gamma}_j\}_{j=1}^q$ do not involve a diverging sequence (see proofs of Proposition (ref) and Corollary (ref)).
Nevertheless, with additional regularity conditions, if $\{c_T^{-1}\sum_{t\in\mathcal{G}_j}|x_{t-1}|\}_{j=1}^q\to_d\{D_j\}_{j=1}^q$ for positive random variables $\{D_j\}$ and a sequence $c_T/\sqrt{T}\to\infty$, then the Cauchy estimator $\check{\beta}_j$ computed within each group satisfies \[ \{m_T(\check{\beta}_j-\beta)\}_{j=1}^q \to_d \{P_j\}_{j=1}^q, \] for $m_T=c_T\sqrt{q/T}$. In general, however, $\{P_j\}_{j=1}^q$ are non-Gaussian, especially when $(x_t)$ is (near) unit root and endogenous. Applying the $t$-statistic approach to $\{\check{\beta}_j\}_{j=1}^q$ thus yields consistency for broader classes of covariates but may incur size distortions due to non-Gaussianity.
Proposition (ref) shows that the $t$-statistic approach is conservative under $H_0$ and consistent under $H_A$ when $(x_t)$ is stationary with a finite first moment. It is thus valid and robust to persistent heteroskedasticity and endogenously heavy-tailed covariates. However, it becomes inconsistent for highly persistent covariates, such as (near) unit-root processes. To illustrate, consider the generalized local-to-unity framework of dou2021generalized, where $X^T(r)=x_{[Tr]}$ for $r\in[0,1]$ and
with $X$ a stationary continuous-time Gaussian ARMA process.\footnote{See dou2021generalized for a detailed discussion.}
When $(x_t)$ is highly persistent, $t_q(\check{\gamma})$ converges to $\mathcal{D}_q$ rather than diverging, with lower bound $(q-1)^{-1/2}$. Simulations in Section 5 confirm that rejection probabilities remain high even when $t_q(\check{\gamma})$ is asymptotically bounded. For $q=2$,
The ratio form in (ref) implies large realizations of $\mathcal{D}_2$ in finite samples, producing high rejection rates even under inconsistency. Figure (ref) plots the simulated density of $\mathcal{D}_2$ when $X$ is Brownian motion.\footnote{Based on 100{,}000 simulated draws.} The minimum simulated value is 1.15, and $\mathbb{P}(|\mathcal{D}_2|>cv_2(0.05))=0.15$ with $cv_2(0.05)=4.303$.
We now propose a simple hybrid test that remains consistent for a broad class of covariates. Under Assumptions (ref)–(ref), \[ \frac{\sum_{t=1}^T |x_{t-1}|}{\sqrt{T}}(\check{\beta}-\beta) = \frac{1}{\sqrt{T}}\sum_{t=1}^T \text{sign}(x_{t-1})u_t \to_d \int_0^1\sigma(r)\,dW(r)=\omega Z, \] where $Z\sim\mathbb{N}(0,1)$ and $\omega^2=\int_0^1\sigma^2(r)\,dr$.
A key feature of the Cauchy estimator $\check{\beta}$ is that its properly normalized limit distribution is invariant to the data-generating process of $(x_t)$. By contrast, the OLS estimator’s variance depends on both $(u_t)$ and $(x_t)$, complicating variance estimation even under homoskedasticity. For $\check{\beta}$, the asymptotic variance depends solely on $u_t$, requiring only heteroskedasticity-robust adjustments.\footnote{See shephard2020, Section 4.3, for related discussion.}
We define the hybrid test statistic as \[ \tau(\check{\beta}) = \frac{\check{\gamma}}{\hat{\omega}}, \] where $\check{\gamma} = (\sum_{t=1}^T |x_{t-1}|/\sqrt{T})\check{\beta}$ as in (ref), and \[ \hat{\omega}^2 = \frac{1}{T}\sum_{t=1}^T \hat{u}_t^2, \qquad \hat{u}_t = y_t - \hat{\beta}x_{t-1}. \] Here, $\hat{\omega}^2$ estimates $\omega^2=\int_0^1\sigma^2(r)\,dr$ using OLS residuals. As noted by shephard2020, the Cauchy-based variance estimator performs poorly because the Cauchy estimator converges more slowly and less efficiently than OLS when $(x_t)$ is heavy-tailed or nearly integrated. Hence, we use OLS residuals to improve efficiency.\footnote{A related approach is employed by KMS2015 in the IVX framework of PM.}
We assume:
Assumption (ref) is very general and holds in many time-series settings. It is weaker than Assumption 3.2 of ibragimov2023new, which requires $\sum_{t=1}^T x_{t-1}u_t = O_p\!\left(\sqrt{T^p\sum_{t=1}^T x_{t-1}^2}\right)$ for $p\in[0,1/16)$. As shown in ibragimov2023new, this holds with $p=0$ when $(x_t)$ is either (near) unit root or stationary with finite variance; it also applies to certain stationary heavy-tailed processes (see, e.g., Samorodnitsky2007, Samorodnitsky2007).
Under Assumption (ref), \[ |\hat{\beta}-\beta| = o_p\!\left(\sqrt{T\big/\sum_{t=1}^T x_{t-1}^2}\right), \quad\text{and hence}\quad \hat{\omega}^2\to_p\omega^2. \] The asymptotic properties of $\tau(\check{\beta})$ follow.
The conclusions of Proposition (ref) remain valid under weaker conditions. For instance, if Assumptions (ref) and (ref) hold and \[ \frac{1}{T}\sum_{t=1}^T v_t^2\to_p\omega^2>0, \qquad \frac{1}{\sqrt{T}}\sum_{t=1}^T \text{sign}(x_{t-1})u_t\to_d\omega Z, \] where $Z\sim\mathbb{N}(0,1)$ is independent of $\omega^2$, then $\tau(\check{\beta})$ retains its asymptotic validity. These conditions include stationary volatility with $E[v_t^2]=\omega^2$. Hence, Assumptions (ref) and (ref) can be interpreted as primitive sufficient conditions accommodating persistent volatility in predictive regression data.
Remark. Proposition (ref)(a) also holds if $\tau(\check{\beta})$ uses $\bar{\omega}^2=T^{-1}\sum_{t=1}^T y_t^2$ instead of $\hat{\omega}^2$, since $\beta=0$ under $H_0$. Moreover, the corresponding test remains consistent when $(x_t)$ is stationary with finite variance or follows a generalized local-to-unity process dou2021generalized. However, it can be inconsistent for heavy-tailed $(x_t)$. For instance, if $(x_t)$ is i.i.d.\ $\alpha$-stable with $\alpha\in(0,2)$ and independent of $(u_t)$, then \[ \bar{\omega}^2 = \beta^2 \left(\frac{1}{T}\sum_{t=1}^T x_{t-1}^2\right)(1+o_p(1)), \quad \tau(\check{\beta}) = \frac{\sum_{t=1}^T |x_{t-1}|}{\sqrt{\sum_{t=1}^T x_{t-1}^2}}(1+o_p(1)) = O_p(1), \] by the generalized central limit theorem (see feller1971, feller1971; logan1973, logan1973; Davis1983, Davis1983; DavisResnick1986, DavisResnick1986). Thus, the use of $\hat{\omega}^2$ (or another consistent estimator under both $H_0$ and $H_A$) is crucial for the consistency of the hybrid test.
This section extends the inference methods developed in Section 3 to models with multiple predictors and to regressions including an intercept. Our goal is not to design efficient procedures but to provide simple and robust inference methods that rely on minimal assumptions on the predictors and volatility processes.
Consider a predictive regression with $K$ predictors $x_t = [x_{1,t}, \ldots, x_{K,t}]'$:
The objective is to test the joint predictability of the covariates, that is, \[ H_0: \beta_{1,K} = \cdots = \beta_{K,K} = 0. \]
We construct a testing procedure for $H_0$ based on the univariate inference methods in Section 3. Specifically, we estimate $K$ univariate predictive regressions \[ y_t = \beta_k x_{k,t-1} + u_{k,t}, \qquad k = 1, \ldots, K, \] and test each null hypothesis \[ H_0^{(k)}: \beta_k = 0, \qquad k = 1, \ldots, K. \] Clearly, $H_0$ implies $H_0^{(k)}$ for all $k$. The converse also holds under mild regularity conditions, as shown below.
Lemma (ref) justifies the use of multiple hypothesis testing based on univariate Cauchy estimators.\footnote{See Harvey2015 for an application of the multiple-testing framework in predictive regressions, and KMS2015 for joint-predictability tests in the IVX framework. Note that the IVX approach may lose validity under heavy-tailed predictors or continuous-time data, whereas our method remains robust in such settings.} In conjunction with the hybrid test introduced in Section 3.2, we compute the statistic $\tau(\check{\beta}_k)$ for each parameter $\beta_k$, where $\check{\beta}_k$ denotes the corresponding Cauchy estimator. Let $p_k$ denote its $p$-value. The joint null hypothesis $H_0$ is rejected at level $\alpha$ if $\min_k p_k \le \alpha / K$, following the Bonferroni correction.
This approach directly extends the univariate robust inference procedure to a multivariate setting and requires only mild conditions for the equivalence between $H_0$ and $\{H_0^{(k)}\}_{k=1}^K$. The Bonferroni correction imposes no assumptions on the joint distribution of the test statistics, which motivates its use here (see Holm1979, Holm1979; BenjaminiHochberg1995, BenjaminiHochberg1995; Shaffer1995, Shaffer1995).
We also note that under some additional regularity conditions, the joint hypothesis can be tested directly using a Wald-type statistic: \[ W = \left( \sum_{t=1}^T z_{t-1} y_t \right)' \left( \hat{\omega}^2 \sum_{t=1}^T z_{t-1} z_{t-1}' \right)^{-1} \left( \sum_{t=1}^T z_{t-1} y_t \right). \] In particular, under $H_0$, \[ \left(\hat{\omega}^2 \sum_{t=1}^T z_{t-1} z_{t-1}' \right)^{-1/2} \left( \sum_{t=1}^T z_{t-1} y_t \right) \to_d \mathbb{N}(0, I_K), \] and hence $W \to_d \chi_K^2$.\footnote{As mentioned earlier, $z_t$ can be interpreted as an instrument. Therefore, one may use an alternative instrument, as in shephard2020, and construct a Wald-type test accordingly.} We provide a detailed analysis of the Wald-type test in the Online Appendix, Section A.1.
While both the Wald test and the Bonferroni-adjusted marginal tests evaluate the same null hypothesis, they are not equivalent procedures except under special cases such as asymptotic orthogonality of the regressors. The Wald test exploits the full covariance structure and yields an elliptical rejection region, whereas the Bonferroni procedure yields a rectangular region and controls family-wise error. Consequently, the procedures may differ in finite-sample and asymptotic power when regressors are correlated.\footnote{Again, our goal is not to design efficient procedures but to provide simple and robust methods that rely on minimal assumptions. We leave a systematic comparison between the Bonferroni-type multiple-testing procedure and the Wald-type joint test for future research.}
The analysis in Section 3 assumes that the intercept $\alpha=0$ in (ref). When $\alpha \neq 0$, it must be properly accounted for. A natural starting point is the demeaned model
where $\bar{z}_s = s^{-1}\sum_{t=1}^s z_t$ for $z_t\in\{y_t, x_{t-1}, u_t\}$. However, $(u_t - \bar{u}_T)$ is not a martingale difference sequence (MDS) with respect to $(\mathcal{F}_t)$, invalidating the martingale CLT used in Sections 2 and 3. Specifically, the Cauchy estimator becomes \[ \check{\beta} - \beta = \Biggl( \sum_{t=1}^T |x_{t-1} - \bar{x}_T| \Biggr)^{-1} \sum_{t=1}^T \text{sign}(x_{t-1} - \bar{x}_T)(u_t - \bar{u}_T), \] which is problematic because: (i) $u_t - \bar{u}_T$ is not an MDS, and (ii) $\text{sign}(x_{t-1} - \bar{x}_T)$ is not $\mathcal{F}_{t-1}$-measurable. Thus, the theory in Section 3 is not directly applicable.\footnote{Recursive demeaning using $\bar{y}_t$ instead of $\bar{y}_T$ does not resolve this issue since $u_t - \bar{u}_t$ is not an MDS either.}
To restore the MDS property, we instead difference the model: \[ y_t - y_{t-1} = \beta (x_{t-1} - x_{t-2}) + (u_t - u_{t-1}), \] and estimate this first-differenced (FD) model on alternating subsets of observations.\footnote{Differencing $y_t$ to eliminate the nonzero mean is a widely used technique in the analysis of dynamic panel data with individual heterogeneity. See, among others, AndersonHsiao1981; AndersonHsiao1982; ArellanoBond1991.} We focus on the even-indexed observations and define the modified Cauchy estimator: \[ \check{\beta}_e = (D_T^e)^{-1} \sum_{t=2}^{T/2} \text{sign}(x_{2t-2})(y_{2t} - y_{2t-1}), \quad D_T^e = \sum_{t=2}^{T/2} \text{sign}(x_{2t-2})(x_{2t-1} - x_{2t-2}). \]
This estimator has two key properties. First, for even-indexed data, the regression error $u_t^e = u_{2t} - u_{2t-1}$ forms an MDS with respect to $\mathcal{F}_t^e := \mathcal{F}_{2t}$ for $t=1,\ldots,T/2$.\footnote{For odd-indexed data, $u_t^o = u_{2t+1} - u_{2t}$ forms an MDS with respect to $\mathcal{F}_t^o := \mathcal{F}_{2t+1}$, yielding an analogous estimator $\check{\beta}_o$.} Second, $\check{\beta}_e$ can again be viewed as an IV estimator, but it uses $\text{sign}(x_{t-2})$, which is $\mathcal{F}_{t-2}$-measurable, as the instrument.\footnote{More generally, one may use $\text{sign}\!\left(\sum_{l\le2}w_l x_{t-l}\right)$ for deterministic weights $\{w_l\}$, provided $ E[\text{sign}(\sum_{l\le2}w_l x_{t-l})(x_{t-1}-x_{t-2})]\neq0$.} Hence, $\text{sign}(x_{2t-2})(u_{2t}-u_{2t-1})$ is an MDS with respect to $(\mathcal{F}_t^e)$.
The inference procedures of Section 3 remain valid for $\check{\beta}_e$. In particular, the hybrid test in Section 3.2 can be implemented as
where $\check{\gamma}_e = D_T^e \check{\beta}_e / \sqrt{T/2}$, and \[ \hat{\omega}^2 = \frac{1}{T}\sum_{t=1}^T \hat{u}_t^2, \qquad \hat{u}_t = (y_t - \bar{y}_T) - \hat{\beta}(x_{t-1}-\bar{x}_T), \] with $\hat{\beta}$ the OLS estimator from the demeaned model (ref). Note that $\hat{\omega}^2$ is based on the full sample, whereas $\check{\beta}_e$ uses only even-indexed data. The asymptotic validity of this hybrid procedure is established next.
Although the odd-indexed estimator $\check{\beta}_o$ has analogous properties, $\check{\beta}_e$ and $\check{\beta}_o$ are typically dependent, with the dependence structure determined by the DGP of $(x_t)$. Hence, unless additional assumptions are imposed, we restrict attention to a single subset of observations—either with even or odd indices.\footnote{Using only half of the data is not uncommon in predictive regressions. See, for example, ZhuCaiPeng2014 and Liu2019, who employ long-lag differencing to eliminate intercepts. In addition, Dufour uses a split-sample approach to address inference problems under a Markovian structure.}
Consistency of the hybrid test with an intercept requires \[ \frac{1}{\sqrt{T/2}}\sum_{t=1}^{T/2}\operatorname{sign}(x_{2t-2})(x_{2t-1}-x_{2t-2})\to_p\infty. \] This holds for most stationary processes $(x_t)$ if \[ E[\operatorname{sign}(x_{t-1})x_t] \neq E[|x_{t-1}|]. \] The condition may fail for certain unit-root processes. For instance, for a random walk $x_t=x_{t-1}+\varepsilon_t^x$, it does not hold. More generally, in the local-to-unity model of Phillips_Magdalinos_2007book,
with $\varepsilon_t^x$ satisfying Assumption LP therein, the consistency condition becomes
The first term diverges if and only if $\delta < 1$ (see Lemma 3.2 of Phillips_Magdalinos_2007book, Phillips_Magdalinos_2007book). Since $\varepsilon_t^x$ and $x_{t-1}$ may be dependent, it is generally the case that $ E[\operatorname{sign}(x_{2t-2})\varepsilon_{2t-1}^x] \neq 0$, implying that the second term also diverges. Consequently, inconsistency arises only when $\delta = 1$ and $ E[\operatorname{sign}(x_{2t-2})\varepsilon_{2t-1}^x] = 0$. In all other cases (i.e., $0 \le \delta < 1$ or when the expectation is nonzero), the test remains consistent.
Moreover, the robustness to a nonzero intercept is not without cost. Specifically, the test based on $\tau(\check{\beta}_e)$ is less powerful than that based on $\tau(\check{\beta})$ when the intercept is zero. A detailed power comparison is provided in Section A.2 of the Online Appendix.
This section investigates the finite-sample performance of the proposed methods. Two sets of simulation experiments are conducted. The first set is based on a continuous-time model and compares our robust $t$-statistic–based tests, $t_q(\check{\gamma})$ for $q \in \{8,12,16\}$, and the hybrid test $\tau(\check{\beta})$ with the Cauchy RT test of CJP2016 and the Cauchy VC test of ibragimov2023new. In this setting, the intercept is suppressed in the test statistics. The second set is based on a discrete-time predictive regression model and compares our procedures, $\tau(\check{\beta}_e)$ and $\tau(\check{\beta}_o)$, with the IVX test of KMS2015. In this setting, the statistics are computed with the intercept included in the estimation.\footnote{The number of replications is 10,000 for each setting.}
Following CJP2016 and ibragimov2023new, we consider the continuous-time predictive regression model
where $V_t$ and $W_t$ are Brownian motions with $ E(V_t W_t) = -0.98t$. The constant term in the predictive regression is set to zero, and recursive demeaning is applied as CJP2016. The model is observed at interval $\Delta = 1/252$, corresponding to daily observations, so that a sample of length $T$ years contains $252T$ observations.
The volatility process $\sigma_t$ follows one of the following specifications:
Following CJP2016, we set $T \in \{20, 50, 100\}$ (corresponding to 240, 600, and 1200 monthly observations) and $\bar{\kappa} \in \{0, 5, 10\}$ for the persistence parameter in (ref). We consider a two-sided test of $H_0: \beta = 0$ against $H_A: \beta \neq 0$.
We first assess the empirical size of each test under $H_0: \beta = 0$. The results for the four volatility models (CNST, SB, RS, and GBM) are reported in Table (ref). Overall, both the $t$-statistic–based tests and the hybrid method exhibit satisfactory size performance, closely matching the nominal levels and performing comparably to the Cauchy RT and Cauchy VC tests. Among the $t$-based procedures, moderate partition numbers ($q = 12$ or $16$) provide the most stable results, whereas smaller $q$ values tend to be mildly undersized. In the GBM case, where volatility is endogenously persistent, the $t$-statistic–based tests become slightly conservative but remain competitive with the Cauchy RT and VC methods.
Next, we analyze the finite-sample power properties of the tests. We consider $\beta \in \{0.004 k, k=1,\cdots,5\}$ under the same volatility specifications. The results are summarized in Tables (ref)–(ref). The proposed tests exhibit power comparable to that of the Cauchy RT and Cauchy VC procedures. For small samples ($T = 20$), the Cauchy RT and VC tests occasionally show higher power, but the difference diminishes as $T$ increases. In certain settings, our methods even outperform the existing approaches. For instance, $t_{16}(\check{\gamma})$ dominates under $\beta = 0.02$, $\bar{\kappa} = 0.5$, and regime switching volatility (Table (ref)), whereas the hybrid test $\tau(\check{\beta})$ performs best under $\beta = 0.004$, $\bar{\kappa} = 20$, $T = 20$, and regime switching volatility.
In summary, all four robust inference procedures—Cauchy RT, Cauchy VC, $t_q(\check{\gamma})$, and $\tau(\check{\beta})$—deliver accurate size control and strong discriminatory power under endogenously persistent regressors and persistent volatility. While the Cauchy RT requires high-frequency data and a time transformation, and the Cauchy VC involves nonparametric volatility filtering with a tuning parameter, our proposed $t$-statistic and hybrid methods are much simpler to implement and require neither. Hence, these approaches are best viewed as complementary: the Cauchy RT and Cauchy VC are preferable in high-frequency environments, whereas our procedures provide robust and easily implementable alternatives in more general settings. It is also worth emphasizing that the proposed methods, like the Cauchy VC, are applicable to both continuous- and discrete-time models, whereas the Cauchy RT method is restricted to the continuous-time framework.
We now examine the finite-sample performance of the proposed tests in a discrete-time setting with an intercept, comparing them to the IVX test of KMS2015. The DGP is specified as
for $t = 2, \ldots, 12T$, where $T \in \{20, 50, 100\}$ corresponds to 20, 50, and 100 years of monthly data (i.e., 240, 600, and 1200 monthly observations, respectively). The constant term in the predictive regression is set to zero. We set $\beta \in \{0.5k : k = 0, 1, \ldots, 5\}$ and $\bar{\kappa} \in \{0, 50, 100\}$, and consider a one-sided test of $H_0\!: \beta = 0$ against $H_A\!: \beta > 0$.\footnote{The IVX test of KMS2015 performs well in two-sided testing for a broad class of models with constant volatility. However, as shown in DEMETRESCU2023105271, the IVX method exhibits severe size distortions in one-sided tests when regressors are highly persistent and endogenous. For this reason, we focus on the one-sided case to highlight the performance of our methods in this setting. Simulations for two-sided and joint hypothesis tests are also reported in Section B of the Online Appendix. Notably, while the IVX approach demonstrates reliable size control under constant volatility, it exhibits significant size distortions in the presence of non-constant volatility. In contrast, our methods remain robust to such heteroskedasticity.}
The innovation process $\eta_t$ follows an MA($q$) process:
where $(\varepsilon_t, v_t)$ are jointly normal with correlation $-0.98$. In our simulations, we consider an MA(1) process with $C_0 = C_1 = 1/\sqrt{2}$.\footnote{Additional simulation results with an MA(3) process are reported in the Online Appendix, Section B.1. Moreover, as discussed in Section 4.2, the intercept-robust version of our test, $\tau(\check{\beta}_e)$, becomes inconsistent when $x_t$ follows the local-to-unity process (ref) with $\delta = 1$ and $\eta_t$ is serially independent. For this reason, we do not consider i.i.d.\ $\eta_t$ in our simulations. See the Online Appendix, Section A.2, for a detailed theoretical power analysis of the test statistic $\tau(\check{\beta}_e)$.} The volatility process $\sigma_t$ follows the same specifications as in the continuous-time simulations, except that the GBM model is excluded.
We implement the hybrid tests based on the even and odd observations, denoted by $\tau(\check{\beta}_e)$ and $\tau(\check{\beta}_o)$, respectively (see (ref)). For comparison, we also include the IVX test of KMS2015, in which the intercept is included in the estimation.
The results, summarized in Tables (ref)–(ref), indicate that the proposed tests exhibit excellent size control under the null hypothesis across all DGPs, whereas the IVX test is substantially oversized, particularly when volatility is nonstationary or exhibits structural breaks. Furthermore, the hybrid approaches demonstrate nontrivial power, even though they are constructed using only half of the observations.
Overall, these findings corroborate the theoretical robustness of our methods. They remain valid under heavy-tailed, endogenous, and persistent regressors, as well as under heteroskedastic and persistent volatility. In contrast, the IVX test exhibits severe size distortions in one-sided tests across all cases, including the constant volatility setting, as documented by DEMETRESCU2023105271.
We also conduct two-sided tests and report the results in the Online Appendix (Section B.2). In this case, the IVX method exhibits good size performance under constant volatility but suffers from severe size distortions in settings with time-varying volatility, whereas our methods maintain accurate size control across all specifications.
Taken together, these results suggest that the proposed robust procedures provide a practical and reliable alternative to existing inference methods for predictive regressions in both continuous- and discrete-time frameworks.
To illustrate the empirical performance of the proposed tests relative to the Cauchy RT and Cauchy VC tests, we reexamine the dataset used by CJP2016 to test the predictability of stock returns using the dividend--price (D/P) and earnings--price (E/P) ratios as predictors.\footnote{Further empirical results, including intercept-robust estimations and Wald-type tests compared to IVX KMS2015, are detailed in the Online Appendix C.} For stock returns, we employ the NYSE/AMEX value-weighted index and the S&P 500 index obtained from the Center for Research in Security Prices (CRSP). The dividend--price ratio is defined as the annual dividend divided by the current total market value. Further details on data construction are provided in Section 6.1 of CJP2016.
Following CJP2016, we estimate two types of predictive regressions: one based on all returns and another based only on returns generated from the diffusive component of stock prices, obtained by first testing for jumps and removing observations corresponding to detected jumps. In all cases, we apply one-sided tests.
The results are reported in Table (ref). As shown in Panels C and D, none of the tests reject the null hypothesis of unpredictability for the S&P 500 data when the E/P ratio is used as a predictor. By contrast, when the D/P ratio serves as a predictor, the proposed tests---$t_q(\check{\gamma})$ with $q=12,16$ and $\tau(\check{\beta})$---reject the null of unpredictability for several cases: CRSP (yearly without jump removal; quarterly with jump removal) and S&P 500 (quarterly and yearly without jump removal; yearly with jump removal). In contrast, the Cauchy RT test fails to reject the null in all cases, while the Cauchy VC test yields qualitatively similar conclusions to our proposed tests, except that it additionally rejects the null for CRSP (monthly with jump removal) and S&P 500 (monthly with jump removal; quarterly with jump removal).
Consistent with our simulation evidence, the Cauchy RT test demonstrates strong finite-sample power but requires high-frequency data due to its reliance on a continuous-time approximation.\footnote{For the Cauchy RT test in our simulations, we estimate the discretized time-changed regression using $n$ lower-frequency observations, with $n = 12T$ (approximately monthly), generated by the random time-sampling scheme described in Section 5 of CJP2016.} The mixed empirical results---where the Cauchy RT test fails to reject the null while both the proposed methods and the Cauchy VC test do reject---may reflect the limited accuracy of the continuous-time approximation when applied to monthly, quarterly, or yearly data. Evaluating the robustness of the continuous-time approximation underlying the Cauchy RT test remains an interesting topic for future research.
This paper introduces two robust inference methods for predictive regressions, addressing key econometric challenges commonly encountered in empirical finance, such as endogenously persistent or heavy-tailed regressors and persistent volatility in errors. Building on the Cauchy estimation framework, we develop two simple yet theoretically rigorous procedures: a $t$-statistic–based approach and a hybrid method. Both methods are computationally straightforward and applicable to continuous- and discrete-time models alike.
Simulation evidence demonstrates that the proposed tests perform well in finite samples, maintaining correct size and competitive power under a wide range of data-generating processes, including those characterized by stochastic volatility, structural breaks, and regime switching. Although our procedures require the assumption of asymptotically exogenous volatility, they exhibit excellent robustness and complement existing methods, including the IVX method of KMS2015, the Cauchy RT test of CJP2016, and the Cauchy VC test of ibragimov2023new.
In an empirical application to stock return predictability, we use the dividend--price and earnings--price ratios as predictors for excess returns on major U.S.\ equity indices. The results indicate that the dividend--price ratio possesses predictive power, while the earnings--price ratio does not significantly forecast returns. Overall, the proposed inference procedures offer a practical, theoretically sound, and implementable alternative to existing methods for robust inference in predictive regressions.
\setcounter{section}{0} \setcounter{subsection}{0} \setcounter{figure}{0} \setcounter{table}{0} \renewcommand\arabic{section}{Appendix \Alph{section}} \renewcommandS.\arabic{figure}{\Alph{section}\arabic{figure}} \renewcommandS.\arabic{table}{\Alph{section}\arabic{table}}
\setcounter{section}{0} \setcounter{equation}{1}