Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
120,255 characters · 15 sections · 0 citation commands
LOCALLY TRIMMED LEAST SQUARES: CONVENTIONAL INFERENCE IN POSSIBLY NONSTATIONARY MODELS
It is well known that under nonstationarity regression estimators do not have conventional limit distributions in general. As a consequence, the inferential procedures developed for stationary data are not applicable under nonstationarity. A number of early studies in the area of nonstationary econometrics (e.g. Phillips and Hansen, 1990; Johansen, 1995; Phillips, 1995; Robinson and Hualde 2003) develop inferential procedures suitable for nonstationary models, however these methods are not valid in general under stationarity. In fact, it is well known that methods such as FMLS (c.f. Phillips, 1995) may exhibit severe size distortions even under local deviations from the (fractional) unit root paradigm. This duality in inference, has made empirical work in time series econometrics elusive. Practitioners typically need to make preliminary (some times ad hoc) assumptions about the persistence level in the data or apply some sort of pre-testing -and therefore expose inference to problems associated to pre-testing- before proceeding to estimation and inference. A number of studies has attempted to address this issue using conservative confidence intervals (for a review see Mikusheva (2007), Phillips (2014) and the references therein). The more recent work of Magdalinos and Phillips (2009; MP hereafter) (see also Kostakis, Magdalinos and Stamatogiannis (2015) for refinements and additional results) follows a completely different direction. MP propose an IV estimator (IVX) that that has mixed Gaussian limit distribution at the expense of an arbitrary reduction in the convergence rate, relative to that of the OLS estimator.
In this paper we follow an approach similar to the pioneering work of MP. To fix ideas consider the simple model
where $x_{k}$ is a nearly integrated (NI) and $x_{k}$ predetermined with respect to some martingale difference error term ($u_{k}$).\ MP construct the so called IVX instrument by applying the following linear filtering to the OLS instrument ($x_{k}$)
for some $c_{z}<0$ and $0<b<1$. This linear filtering transforms $x_{k}$ into a mildly integrated process (e.g. see Giraitis and Phillips, 2006; Phillips and Magdalinos, 2007) that is less persistent than a NI array (e.g. $x_{k}$). By choosing $b$ arbitrary close to unity, the\ reduction in the signal of the instrument results into an arbitrary small reduction in the convergence rate of the IVX estimator, relative to that of the OLS, and this is sufficient for a martingale CLT to operate, rendering IVX based inference conventional. The choice of $b$ is important to inference with smaller $b$\ resulting in better size control at the expense of asymptotic power. Note that as $b\uparrow1$, $Z_{kn}$\ approximates the NI process $x_{k}$ and the IVX estimator resembles the behaviour of the OLS estimator. The recent work of Yang, Long, Peng and Cai (2019) generalises the IVX method to regression models with serially correlated regression (parametric AR) errors, whilst Demetrescu, Georgiev, Rodrigues and Taylor (2020) apply a modified version of the IVX estimator to test for episodic predictability in stock returns.
We consider an alternative method for reducing the signal of the OLS instrument. Let $K$ be an integrable kernel function and set \[ Z_{kn}=K\left[c_{n}\left(k/n-\tau\right)\right]x_{k}\text{,} \] where $c_{n}$ is a positive deterministic sequence such that$\ c_{n}^{-1}+c_{n}n^{-1}\rightarrow0$ and $0<\tau<1$. For simplicity set $\tau=1/2$ and $K(0)=1$. In this case the kernel function extracts information from the OLS instrument for observations near the middle of the sample. In particular, $Z_{kn}\approx x_{k}$ i.e. when $k\approx n/2$, and $Z_{kn}\approx0$\ when $k$ is far from $n/2$. In other words certain chronological trimming applies around the “chronological point $\tau$”. By allowing the $c_{n}$\ sequence to diverge at an arbitrary slow rate, the resultant IV (LTLS) estimator attains an arbitrary slower convergence rate relative to the OLS estimator. In principle, it is possible to extract information around multiple chronological points $0<\tau_{1}<...<\tau_{l_{n}}<1$ where $l_{n}$\,is either fixed or $l_{n}\rightarrow\infty$ such that $l_{n}=o(c_{n})$. In this case the relevant instrument is
As long as the LTLS estimator converges at slower rate, than the OLS estimator, limit theory is mixed Gaussian for nonstationary regressor covariates and Gaussian for stationary. In particular, the reduction in the signal of the OLS instrument allows a martingale CLT (c.f. Wang, 2014) to operate even if $x_{k}$ is nonstationary. Notice that if $c_{n}$ is too small or if too many chronological points ($l$) are employed, then $Z_{kn}$ approximates the OLS instrument and as a consequence LTLS based inference resembles OLS based inference. This can be easily seen, if a vanishing sequence $c_{n}$ is employed. Note that for $c_{n}\rightarrow0$, $Z_{kn}\approx l_{n}K(0)x_{k}$.
Our theoretical framework allows for a wide range of stationary and nonstationary linear processes as well NI arrays. In particular, $x_{k}$ can be a stationary or a nonstationary fractional process. Consider the LTLS estimator of $\beta$ in ((ref)) that utilises the instrument of ((ref)) i.e. $\hat{\beta}=\sum_{k=1}^{n}Z_{kn}y_{k}/\sum_{k=1}^{n}Z_{kn}x_{k}$. Let $t\in\lbrack0,1]$ and suppose that $x_{k}$ is a nonstationary process such that for some $d_{n}\rightarrow\infty$, $d_{n}^{-1}x_{\lfloor nt\rfloor}\Rightarrow X_{t}$ in $D[0,1]$\ where $X_{t}$ is a continuous process. For instance $X_{t}$ can be a fractional BM or a fractional Ornstein-Uhlenbeck process (see Remark (ref)\ below) depending on some memory or near-to-unity nuisance parameter.\ Then we have \[ d_{n}\sqrt{\frac{nl_{n}}{c_{n}}}\left(\hat{\beta}-\beta\right)\rightarrow_{d}\mathbf {MN}\left(0,E\left(u_{t}^{2}\right)\frac{\int_{ \mathbb{R} }K^{2}(x)dx}{\left(\int_{ \mathbb{R} }K(x)dx\right)^{2}\int_{0}^{1}X_{t}^{2}dt}\right). \] Because $c_{n}\rightarrow\infty$ and $l_{n}=o(c_{n})$ the convergence rate of the LTLS is slower than that of the OLS estimator ($d_{n}\sqrt{n}$). Further, note that nuisance parameters affect the limit distribution only via the mixing variate $\left[\int_{0}^{1}X_{t}^{2}dt\right]^{-1}$ and as a consequence the studentised LTLS estimator has standard normal limit distribution. Interestingly, the limit variance shown above is the same, up to a constant, to that of the FMLS estimator for the case where $x_{k}\sim I(1)$.
We mention that the constant that features in the limit variance of the LTLS estimator above, can be made arbitrarily small by an appropriate choice of the kernel function. For example suppose that $K(x)=\left(2\pi\varsigma^{2}\right)^{-1/2}\exp\left(-\frac{x^{2}}{2\varsigma^{2}}\right)$. Then \[ E\left(u_{t}^{2}\right)\int_{ \mathbb{R} }K^{2}(x)dx\Big/\left(\int_{ \mathbb{R} }K(x)dx\right)^{2}=\frac{E\left(u_{t}^{2}\right)}{2\sqrt{\pi\varsigma^{2}}}\rightarrow0 \] as $\varsigma^{2}\rightarrow\infty$.\footnote{
} Nevertheless, choosing a large value of the kernel variance parameter has the same effect as choosing a small value for $c_{n}$. Therefore as $\varsigma^{2}\rightarrow\infty$, the LTLS estimator approximates the OLS estimator.
It should be further noted that for nonstationary fractional covariates (i.e. $I(d)$, $d>1/2$), methods like FMLS (e.g. Phillips, 1995) or the spectral GLS of Robinson and Hualde (2003) (see also Hualde and Robinson, 2010) are asymptotically equivalent the Gaussian pseudo maximum likelihood and therefore asymptotically efficient (c.f. Phillips, 1991). The key feature of these methods is to induce asymptotically mixed Gaussian estimators by certain modification in the dependent variable that involves (fractionally) differencing the covariates. In the context of ((ref)) such differencing takes the form $(I-L)^{\hat{d}}x_{k}$, where $L$ is the lag operator and$\ \hat{d}$\ is a preliminary estimator for the memory parameter of $x_{k}$. Nevertheless, if there is a local deviation (order $O(n^{-1})$) from the (fractional) unit root model, the aforementioned methods yield mixed Gaussian limit theory only if the following quasi fractional differencing is applied \[ \left(I-\left(c/n\right)L\right)^{\hat{d}}x_{k}, \] where $c$ is a local to unity parameter. A non trivial value for the local to unity parameter however renders the aforementioned methods infeasible because of the lack of identifiability of $c$. It is well known that if $c\neq0$, inference based on methods like FMLS are prone to severe size distortions even if there is moderate correlation between the regressor and the regression error.
The remaining of this work is organised as follows. Section 2 provides basic limit theory for locally trimmed functionals of stationary and nonstationary processes. This limit theory is utilised in Section 3 for exploring the limit properties of the LTLS estimation and inference. Section 4 provides a simulation study and Section 5 an empirical application on the predictability of stock returns.
Throughout this paper we make use of the following notation. For two deterministic sequences $a_{n}$ and $b_{n}$, $a_{n}\sim b_{n}$ denotes $\lim_{n\rightarrow\infty}a_{n}/b_{n}=1$. $1\left\{ A\right\} $ is the indicator function on set $A$. We may write the integral $\int_{ \mathbb{R} }f(x)dx$ as $\int f$. $\Rightarrow$ denotes weak convergence in the space $D[0,1]$. For a vector $x$, $\left\Vert x\right\Vert $\ is its inner product norm and $x^{\prime}$ its transpose. By $[x]$ we denote the integer part of a positive number $x$. Finally, diag$\{a_{1},...,a_{p}\}$ denotes a $p\times p$ diagonal matrix with elements $\{a_{1},...,a_{p}\}$ on the main diagonal, $\to_d$ denotes the convergence in distribution and $Y:=\mathbf{MN}(\mathbf{0}, \sum)$ denotes a Gaussian variate (mixing normal) with characteristic function $f(t)=E e^{it' Y}=Ee^{-t'\sum t/2}$.
In this section we develop basic limit theory for locally trimmed (LT) sample functionals of stationary and nonstationary processes. Our basic limit theory is utilised in Section 3 for the asymptotic analysis of the LTLS estimator.\ Let $\left\{ x_{k}\right\} _{1\leq k\leq n}$ be a scalar time series process and $\{X_{nk}\}_{1\leq k\leq n,n\geq1}$ be some scalar random array. Further, let $K $ be an integrable kernel function and $\ g(.)=\left[g_{1}(.),...,g_{p}(.)\right]^{\prime}$, where, for each $i=1,...,p$, $g_{i} $ is a measurable function. For $l\in \mathbb{N} $ and$\ 0<\tau_{1}<...<\tau_{l}<1$, set
where $c_{n}$ is a sequence of positive constants, $l$ either fixed or $l\rightarrow\infty$ as $n\rightarrow\infty$, and $u_{k}$ together with an appropriate filtration $\left\{ \mathcal{F}_{k}\right\} $ forms a martingale difference sequence (such that $X_{nk}$, $x_{k}$ are $\mathcal{F}_{k-1}$-measurable). The limit theory of the LTLS estimator relies on the asymptotics of $\left\{ S_{jn,l},M_{jn,l}\right\} _{j=1}^{2}$. Limit theory for the functionals $\left\{ S_{1n,l},M_{1n,l}\right\} $ is relevant for stationary regressors whilst $\left\{ S_{2n,l},M_{2n,l}\right\} $ for nonstationary. In fact, it is assumed that $X_{nk}$ satisfies some FCLT.\ The term $S_{2n,l}$ resembles certain functionals considered by Phillips, Li and Gao (2017) who study the estimation of cointegrated models with smooth time varying parameters (TVP). The aforementioned work considers terms of the form
where $X_{nk}$\ is an $I(1)$ process\ normalised by $\sqrt{n}$.\ As explained below, under our assumptions $X_{nk}$\ can be an appropriately normalised $I(d)$, $d>1/2$, process or a NI array (possibly driven by fractional errors). Therefore the limit results provided in this section are also relevant to the estimation of TVP models for the case where the covariate is a general nonstationary process satisfying some FCLT (see Assumption A3 below).\
To facilitate basic limit results, we make use of the following conditions.
We remark that the innovation process $\left\{ \mathbf{ \eta }_{k},\mathcal{F}_{k}\right\}_{k\ge 1} $ used in {\bf A1} is standard in literature so that both $M_{1n, l}$ and $M_{2n, l}$ have a martingale structure. The uniform integrability conditions (a) and (b) are weak in comparison with the high moments used in previous works. See, for instance, Wang (2014) and Wang and Phillips (2009a, b). Since $\Sigma$ is required to be a positive definite matrix, condition (c) excludes the process $u_k$ to be ARCH and GARCH models. The condition (c) is required for technical reasons, which seems to be difficult to reduce at the moment.
Stationary process given in {\bf A2} is extensively used in empirical applications where examples include short and long memory (fractional) processes. Typical examples on nonstationary processes satisfying {\bf A3} have the form:
where $\rho =1+c /n$ with $ c\in\mathbb{R} $ and $\sum_{i=0}^{\infty }\phi _{i}^{2}<\infty.$ For the latter specification, ((ref)) holds with $X_t$ being a fractional Ornstein-Uhlenbeck process. See, for instance, Buchmann and Chan (2007), Wang and Phillips (2009a, b) and Wang (2015).
As for {\bf A4}, the restriction on compact support for $K(x)$ can be relaxed if we have more conditions on $l_n$. Indeed, in the following main results, {\bf A4} can be replaced by the following:
We now introduce the limit theory for LT sample functionals. Since there are essential difference between $M_{1n, l}$ and $M_{2n, l}$, the main results will be presented based on stationary and nonstationary processes, separately.
Theorem (ref) provides limit theory for rescaled functionals of nonstationary processes (i.e. $d_{n}^{-1}x_{k}$ as given A3). For the purposes of regression analysis, limit theory for non rescaled processes (i.e. $x_{k}$) is more relevant. Following Park and Phillips (1999, 2001), we assume that the function $g(.)=[g_{1}(.),...,g_{p}(.)]^{\prime}$ is asymptotically homogeneous, i.e. for large $\lambda$ \[ g_{i}(\lambda x)\approx\pi_{i}(\lambda)H_{i}(x),\text{ }i=1,...,p \] where $\pi_{i}$ (positive real valued function)\ is the “asymptotic order”\ of $g_{i}$\ and$\ H_{i}$ is the “asymptotic homogeneous function”\ of $g_{i}$ that is assumed continuous. Several specifications of interest satisfy these conditions e.g. polynomial functions, logarithmic, indicator functions and distribution type of functions e.g. see Park and Phillips (2001) for more details. Set $\pi\left(.\right):=$diag$\{\pi_{1}\left(.\right),...,\pi_{p}\left(.\right)\}$ and $H(.)=[H_{1}(.),...,H_{p}(.)]^{\prime}$. The following result is the counterpart of Theorem (ref) for additive transformations of non rescaled sequences.
The limit theory presented in the previous section is subsequently utilised for deriving the properties of the LTLS estimator and a related t-statistic. We consider nonlinear models of the form
where $f$ is a known regression function $(\mu,\beta)$ unknown parameters and the covariate $x_{k}$ can be nonstationary process or a stationary one amenable to the limit theory of Theorem (ref) or Theorem (ref) respectively. Further, $x_{k}$ is predetermined with respect to the error $u_{t}$ in the sense $x_{k}$ is $\mathcal{F}_{k-1}$-measurable and $\left\{ u_{k},\mathcal{F}_{k}\right\} $ is a martingale difference (c.f. Assumptions {\bf A1-A3}). Similar nonlinear models with a predetermined covariate have been considered for example by Park and Phillips (1999, 2001) and Chan and Wang (2015), in a parametric set up, and by Wang and Phillips (2009a,b, 2011, 2012) in a nonparametric set-up.\footnote{Here we consider nonlinear models in $x_{k}$ only. Our results can be generalised to models that are both nonlinear in $x_{k}$ and the parameters along the lines of Chan and Wang (2015) for instance.}
Let $K$ be a kernel function satisfying {\bf A4}(a) or {\bf A4$^*$}(a). Let $\tau_j=j/(l_n+1), j=1, ..., l_n$, $c_{n}$ and $l_{n}$ be deterministic sequences satisfying {\bf A4}(b) and (c) or {\bf A4$^*$}(b) and (c). We also allow for $l_n$ to be a fixed constant. Set
Our aim is to estimate the unknown parameter $\beta$ in ((ref)) by using the following instrument for $f(x_{k})$
As remarked in Section 1, due to the integrability of $K$, a trimming effect applies around the chronological point(s) (cp(s) hereafter) $\tau_{j}$ which in turn reduces the signal of the OLS instrument $f(x_{k})$. The reduction is more pronounced when the distance between $k/n$ and$\ \tau_{j}$ is large, and/or the sequence $c_{n}$ diverges fast. Clearly, for $K_{kn}=1$ we get the OLS estimator as a special case. The reduction in the instrument signal enables an extended martingale given by Wang (2014) to operate. As a result the estimator under consideration has a mixed Gaussian limit distribution, making pivotal inference possible.
A trimming method is also crucial for demeaning $\left\{ y_{k}\right\} $ i.e. taking into account the unknown intercept $\mu$. Let $K_{kn}^{\ast}$, $k=1,...,n$\ be additive functionals of certain integrable kernel function.\ For any sequence $\left\{ a_{k}\right\} _{k=1}^{n}$\ let
\ We will consider two possibilities for $K_{kn}^{\ast}$. Either
where $K^{\ast}$ satisfies {\bf A4}(a), $\tau_j=j/(l_n+1), j=1,2,..., l_n,$ are given above and $\ 0<\tau^{\ast}<1$. The first term in ((ref)) involves a trimmed sample mean around an array of several cps, whilst the second is a trimmed sample mean based on a single fixed cp.\ Define the LTLS estimator as \[ \hat{\beta}:=\frac{\sum_{k=1}^{n}Z_{kn}\overline{y}_{k}}{\sum_{k=1}^{n}Z_{kn}\overline{f}_{k}}. \] The employment of a “trimmed”\ sample mean is crucial for obtaining mixed Gaussian limit theory. Notice that \[ \hat{\beta}=\beta+\frac{1}{\sum_{k=1}^{n}Z_{kn}\overline{f}_{k}}\left\{ \sum_{k=1}^{n}f_{k}K_{kn}u_{k}-\frac{\left(\sum_{k=1}^{n}f_{k}K_{kn}\right)\sum_{k=1}^{n}K_{kn}^{\ast}u_{k}}{\sum_{k=1}^{n}K_{kn}^{\ast}}\right\} . \] For nonstationary $x_{k}$ the two martingale terms shown above converge jointly to a bivariate mixed Gaussian limit. In particular, \[ \left[\sqrt{\frac{c_{n}}{n}}\sum_{k=1}^{n}f\left(d_{n}^{-1}x_{k}\right)K_{kn}u_{k},\ \sqrt{\frac{c_{n}}{n}}\sum_{k=1}^{n}K_{kn}^{\ast}u_{k}\right]\rightarrow_{d} \mathbf{MN}\left(\mathbf{0},V\right), \] for some random matrix $V$. Note that if instead the standard demeaning was employed (i.e. $K^{\ast}=1$), then \[ \left[\sqrt{\frac{c_{n}}{n}}\sum_{k=1}^{n}f\left(d_{n}^{-1}x_{k}\right)K_{kn}u_{k},\ \frac{1}{\sqrt{n}}\sum_{k=1}^{n}u_{k}\right]\nrightarrow_{d}\mathbf{MN}\left(\mathbf{0},V\right), \] for some random matrix $V$, despite the fact that each of the components on the l.h.s. above converges weakly to some (mixed) Gaussian limit.
To investigate the limit properties of the LTLS estimator $\hat{\beta}$ in detail, set \[ \lambda_{n}:=\frac{nl_{n}}{c_{n}}\text{\qquad\ and \qquad}\lambda_{n}^{\ast}:=\frac{nl_{n}^{\ast}}{c_{n}}, \] where \[ l_{n}^{\ast}:=\left\{
\right.. \] The sequences $\lambda_{n}$, $\lambda_{n}^{\ast}$ give the order of the terms\footnote{Note that by standard arguments (Euler summation) \[ \sum_{k=1}^{n}K_{kn}^{\ast}\sim\frac{nl_{n}^{\ast}}{c_{n}}\int K^{\ast}. \] } $\sum_{k=1}^{n}K_{kn}$ and $\sum_{k=1}^{n}K_{kn}^{\ast}$ which in turn\ determine the convergence rate of the LTLS estimator. Further set
We have the following main results for the asymptotics of the LTLS estimator $\hat{\beta}$. Theorem (ref) is for stationary regressor. Limit theory in nonstationary case is given in Theorem (ref).
To end this section, we consider the following $t$-statistic for the hypothesis $H_{0}:\beta=\beta_{0}$ (for some $\beta_{0}\in\mathbb{R}$)
where \[ \mathcal{A}_{n}:=\left[1,\text{ }-\frac{\sum_{k=1}^{n}f_kK_{kn}}{\sum_{k=1}^{n}K_{kn}^{\ast}}\right],~~\text{ }\mathcal{C}_{n}:=\sum_{k=1}^{n}Z_{kn}\overline{f}_{k}, \] \[ \mathcal{V}_{n}:=\left[
\right], \] and $\tilde{\sigma}^{2}:=n^{-1}\sum_{k=1}^{n}\tilde{u}_{k}^{2}$, where $\tilde{u}_{k}$ are residuals from OLS estimation of ((ref)). The limit properties of $\hat{T}$ under the null hypothesis are demonstrated by Theorem (ref) below.
We next investigate the final sample performance of the t-test based on the LTLS estimator. In particular we test the hypothesis $H_{0}:\beta=0$ against $H_{1}:\beta\neq0$ at 5% significance level. The vector $\left[\xi_{k},u_{k}\right]$ process is generated by \[ \left[
\right]\sim i.d.N\left(\mathbf{0},\left[
\right]\right) \] Further, for $k=1,...,n$ the process$\ \left\{ y_{k}\right\} $ is generated by \[ y_{k+1}=\beta x_{k}+u_{k+1} \] where $\left\{ x_{k}\right\} $ is either a NI array of the form
with $c\leq0$ and $x_{0}=0$\ or a type II fractional process (e.g. see Robinson and Hualde, 2003) of the form
Let $\varphi_{\varsigma^{2}}(x)$ be the density of a $N\left(0,\varsigma^{2}\right)$ variate. Next, set $\tilde{\sigma}_{u}^{2}=n^{-1}\sum_{k=1}^{n}\tilde{u}_{k}^{2}$, $\tilde{\sigma}_{\xi}^{2}=n^{-1}\sum_{k=1}^{n}\tilde{\xi}_{k}^{2}$,$\ \tilde{\delta}=\frac{n^{-1}\sum_{k=1}^{n}\tilde{u}_{k}\tilde{\xi}_{k}}{\sqrt{\tilde{\sigma}_{u}^{2}\tilde{\sigma}_{\xi}^{2}}}$, where $\tilde{u}_{t}$ and $\tilde{\xi}_{k}$ are OLS residuals from the regressions \[ y_{k+1}=\tilde{\mu}+\tilde{\beta}x_{k}+\tilde{u}_{k+1}\qquad\text{and}\qquad x_{k}=\tilde{\mu}_{x}+\tilde{\rho}x_{k-1}+\tilde{\xi}_{k} \] respectively.\ Finally, $\left\{ \tau_{j}\right\} _{j=1}^{l_{n}}$ are equispaced points on $(0,1)$.
We consider 3 set-ups for kernel functionals and cps.
In S1 and S2 multiple cps are used for both $K_{kn}$\ and $K_{kn}^{\ast}$\ whilst in S3 $K_{kn}^{\ast}$\ involves a single cp. Contrary to S1, in S2 a data driven approach is followed for the determination of the number of cps ($l_{n}$). As remarked in Section 1, a small $c_{n}$ and/or large number of cps\ results in a LTLS estimator approximately equal to the OLS estimator. The OLS estimator in general has a good power properties but is severely oversized when endogeneity is strong (i.e. when $\left\vert \delta\right\vert $\ is close to one). In S2 a large number of cps is utilised when endogeneity is weak whilst for $l_{n}$ drops as $\left\vert \delta\right\vert $ approaches one. A similar data-driven approach is utilised in S3. In this case $c_{n}$ is very small (vanishing) for $\delta$ close to zero, whilst $c_{n}$ is large (diverging) for $\left\vert \delta\right\vert $ close to one. Further, in S3 the choice of the kernel variance is also data driven. Preliminary simulations have shown that superior performance is attained when $\varsigma^{2}=0.1$ for $\delta\approx0$ and $\varsigma^{2}=1$ for $\left\vert \delta\right\vert \approx1$. Therefore, $\hat{\varsigma}^{2}=\tilde{\sigma}_{u}^{2}\left(0.1+0.9\left\vert \tilde{\delta}\right\vert \right)$ provides an interpolation between these values based on the actual data.
For S1 and S2 we use the test statistic of ((ref)). For S3 we use $\mathcal{A}_{n}^{\ast}:=\left[1,\text{ }-\frac{\sum_{k=1}^{n}f(x_{k})K_{kn}^{\ast}}{\sum_{k=1}^{n}K_{kn}^{\ast}}\right]$ instead of $\mathcal{A}_{n}$ in ((ref)). Note that given the configuration of S3, in the nonstationary case, $\frac{\sum_{k=1}^{n}f(x_{k})K_{kn}^{\ast}}{\sum_{k=1}^{n}K_{kn}^{\ast}}=O_{p}\left(\pi_{f}(d_{n})\right)$ whilst $\frac{\sum_{k=1}^{n}f(x_{k})K_{kn}}{\sum_{k=1}^{n}K_{kn}^{\ast}}$ that appears in $\mathcal{A}_{n}\ $is $O_{p}\left(\pi_{f}(d_{n})\log n\right)$. Therefore, the employment of $\mathcal{A}_{n}^{\ast}$ results in giving less weight in the term that correspondents to the studentisation of the intercept correction. Note that in infinite samples the utilisation of $\mathcal{A}_{n}^{\ast}$ does not result in a consistent estimator for the limit variance of $\hat{\beta}-\beta$. Nevertheless, our simulation results reveal that in finite samples a superior performance in attained when $\mathcal{A}_{n}^{\ast}$ is employed.
Table 1 shows the size properties of the LTLS based t-tests, for the case the regressor is a NI array generated by ((ref)). The number of replication paths is set 10,000 throughout. We also consider the IVX based test (see eq. (20) in Kostakis et al, 2015) and the OLS based t-test. We allow for several values of the correlation parameter ($\delta=\{-.95,-.5,0,.5,.95\}$) the near to unity parameter ($c=\{0,-5,-10,-20,-50\}$) and sample size ($n=\{250,500,750,1000\}$). We use the notation T1, T2 and T3 to denote the LTLS t-statistics that correspond to set-ups S1, S2 and S3 respectively. In general, all LTLS based test exhibit good size control. Under S1 and S2 the tests are moderately oversized for small samples sizes when $c=0$ and correlation $\left\vert \delta\right\vert =.95$. Figure 1 and Figure 2 show the empirical power ($n=250$) of the LTLS and IVX tests for $c=0$ and $c=-20$ respectively. It can be seen from these figures that T3 attains better performance than the other LTLS based tests under consideration (i.e. T1 and T2). In particular, the performance of the the LTLS t-test under S3 is almost identical to that of the IVX based test. This is somewhat surprising given that under S3 the studentisation used does not lead to a consistent estimator for the limit variance of the LTLS estimator. As noted above, under S3 the term that provides studentisation to the intercept correction is of slightly smaller order of magnitude (i.e. $\log n$) than the corresponding term in $\mathcal{A}_{n}$. The simulation study provided suggests that this misbalancing leads to some finite sample improvement. Hosseinkouchack and Demetrescu (2019) provide finite sample improvements to the the IVX method. These authors show that the IVX t-statistic distribution is skewed relative to the $N(0,1)$ in finite samples when endogeneity is strong. It is reasonable to expect that a similar phenomenon holds for the LTLS distribution in finite samples. It seems that the utilisation $\mathcal{A}_{n}^{\ast}$ provides a rebalancing to the test statistic that corrects for deviations from the standard normal distribution. A rigorous analysis for the performance of the T3 in finite samples, would require developing higher order limit theory. A development in this direction is challenging from a technical point of view and will be left for future work.
We next consider the case where the regressor is a non stationary fractional process (i.e. ((ref))). The finite sample size performance of T3 and the LS based test procedure are shown in Table 2.\footnote{Preliminary simulation results show that the performance of T1 and T2 in the fractional case is comparable to that in the NI case.} It can be seen from Table 2 that the T3 test provides good size control for a wide range levels in persistence and endogeneity. On the other hand LS based test may exhibit serious oversizing. In particular, for $\delta=-.95$ empirical size ranges from three times ($d=0.75$) to six times ($d=1.2$) the nominal one. Finally, Figure 3 shows the finite power of T3 for $n=250$, $d=\{0.8,1,1.2\}$ and $\delta=\{0,-.5,-.95\}$. As expected, better power performance is attained for more persistent regressors.
\thispagestyle{empty}
A large literature in empirical finance is devoted to the investigation of the hypothesis that stock returns can be predicted with publicly available information. For a review of existing work see for example Welch and Goyal (2008) and for more recent developments Kostakis, Magdalinos and Stamatogiannis (2015). Typically empirical work in this area involves inferential procedures, for the hypothesis $H_{0}:\beta=0$, in the context of predictive regressions of the form
where $r_{k}$ are stock returns relating to some stock index,$\ x_{k}$ some predictive variable and $u_{t}$ a martingale difference regression error. Usually some financial ratio (e.g. dividend yield, earnings to price ratio, book to market ratio) or some macroeconomic variable (e.g. inflation) is considered as a possible predictor for future returns. Phillips (2015) provides a review for the econometric methodology employed in the predictive regressions literature. Most studies (e.g. Welch and Goyal, 2008) are utilising methods that are only valid for stationary $x_{k}$ despite the fact that there is strong evidence that that in certain datasets various financial and macroeconomic variables are consistent with nonstationary processes (e.g. see Kostakis et al, 2015; Table 4). To the best of our knowledge, Campbell and Yogo (2006)\ is the first work that explicitly provides an attempt to address the possibility that the regressor is nonstatationary. In particular, Campbell and Yogo (2006) develop a testing procedure for the case the predictor is a NI array based on conservative confidence intervals. Kostakis et al (2015) consider a modified version of the Magdalinos and Phillips (2009) IVX, that involves a finite sample correction relating to intercept estimation, to examine the return predictability hypothesis. The IVX estimator yields conventional inference for the case where $x_{k}$ is a NI or mildly integrated array (e.g. Phillips and Magdalinos, 2007) or a stationary linear process. IVX instruments are also employed in the recent work of Demetrescu et al (2020) who propose inferential procedures for detecting episodic predictability in stock returns. The IVX method has been also employed in the recent work of Yang, Long, Peng and Cai (2020) who investigate predictability in the U.S. housing index return.
An important issue that has been largely overlooked in most studies in this area, is that stock returns series typically exhibit very weak persistence relative to most popular predictors. In particular, in many datasets short-term returns appear to be close to $I(d)$ processes with $d\approx0$, whilst several predictors appear to be $I(d)$ with $d>1/2$ i.e. nonstationary processes. Regressing a stationary processes on a possibly nonstationary leads to misbalancing. As emphasised by Phillips (2015), misbalancing may result to asymptotically vanishing estimators. For instance if $r_{k}\sim I(d)$ with $d<1/2$ (stationary long memory) and $x_{k}\sim I(d)$ with $d>1/2$, then then OLS estimator for $\beta$ in ((ref)) is $\tilde{\beta}\rightarrow_{P}0$.
Only a few studies in this area attempt to address the issue of misbalancing.\ Marmer (2007) points out that a nonlinear relationship between returns and predictive variables is a plausible justification for this discrepancy in persistence. It is known for instance that integrable and bounded transformations of persistent processes may exhibit very weak signal (e.g. Park and Phillips, 1999, 2001; Park, 2003). Therefore, suppose that $r_{k+1}=f\left(x_{k}\right)+u_{k+1}$ where $f$ is either integrable and compactly supported or the indicator function $1\left\{ .<0\right\} $. The predictor $x_{k}$ in this case has only “spatial episodic”\ impact on returns when the predictive variable visits the support of $f$ (integrable case) or when it assumes negative values (indicator case).\ For DGPs of this kind it is difficult distinguishing $r_{k}$ from the martingale difference error $u_{k}$, despite the fact $r_{k}$ is a function of a persistent process\ (see for example Kasparis, Andreou and Phillips (2015), Figure 6; or Phillips (2015), Figure 2). Marmer (2007) develops a RESET type of functional form test for detecting possibly nonlinear components (e.g. integrable) of some predictor in the stock return series. A similar approach is also followed by Kasparis (2010) and Kasparis et al (2015), who utilise test statistics that involve integrable transformations of the predictor. The presence of integrable transformations in the test statistics results in conventional inference but can also detect weak signal nonlinear components affecting the returns series (for more details see p. 473-474 in Kasparis et al, 2015). Bollerslev, Osterrieder, Sizova and\ Tauchen (2013) follow a different approach for addressing the issue of misbalancing. These authors consider vix and realised volatility as possible predictors of stock returns. Using preliminary estimations they find that the aforementioned predictors exhibit long memory with memory parameters $d\approx0.4$, whilst stock returns appear to have a memory parameter $d\approx0$. In view of this, Bollerslev et al (2013) consider prefiltered predictors of the form $\left(I-L\right)^{\hat{d}}x_{k}$ where $x_{k}$\ is some volatility variable. Notice that the fractionally differenced process is approximately $d\approx0$. Finally, Demetrescu et al (2020) develop inferential procedures capable of detecting episodic predictability is stock returns for the case where the predictors that are either $I(0)$ or NI. In particular, they consider a potentially nonlinear relationship between returns and the predictive variables of the form $r_{k+1}=f_{n}\left(x_{k},k/n\right)+u_{k+1}$, where $f_{n}\left(x_{k},k/n\right)=\mu+k_{n}\beta\left(k/n\right)x_{k}$, $\beta\left(.\right)$ is a TVP depending on the rescaled time trend $k/n$, and $k_{n}$ an appropriate sequence. This formulations allows for “time episodic”\ impact of the predictor to the returns variable. Demetrescu et al (2020) achieve conventional inference by either utilising IVX instruments or the so called type II instruments of Breitung and Demetrescu (2015).\footnote{The method of Demetrescu et al (2020) can be used in conjunction with various instruments including LTLS. Such a development would require additional theoretical work and is left for future research.}
In this work we address the issue of misbalancing by consider predictability over longer horizons. In particular, we employ LTLS based inference in predictive regressions of the form
where $m\geq1$. The specification of ((ref)) has been considered by other studies that investigate return predictability over long horizons (see for example Bandi and Perron, 2008; Hjalmarsson, 2011). The data are taken from the updated 2018 Welch and Goyal dataset\footnote{The data are download from Amit Goyal's webpage: http://www.hec.unil.ch/agoyal/
}. The returns variable is constructed from the SP500 index ($I_{k}$) as follows $r_{k+m}=\ln(I_{k+m})-\ln(I_{k})$. We are using monthly and quarterly observations. Therefore, for monthly data, $r_{k+m}$ should be understood as $m$ months ahead returns, and for quarterly observations as $m$ quarters ahead. By construction returns are log-price differences. Therefore, the persistence of the returns series tends to increase as the horizon increases. Table 3 provides memory estimates for the return series over different horizons and frequencies. In particular, we use the local Whittle estimator (LW; e.g. see Robinson, 1995) and the exact local Whittle (ELW) of Shimotsu and Phillips (2005). The bandwidth employed is of the form $n^{b}$. Shimotsu and Phillips (2005) consider $b=0.65$ for the bandwidth exponent. Here we also consider $b=0.55$ and $b=0.75$. Moreover, we report memory estimates for the earnings to price ratio (EP). The particular series appears to be less persistent than dividend yield and book to marker ratio that are commonly used in empirical work. For this reason we will concentrate on EP whose memory characteristics are closer to those of the returns series. It can be seen from Table 3 that the EP appears to be nonstationary at both frequencies and for all bandwidth choices with minimal memory estimate $0.76$. Further, the memory characteristics of the returns series appear to resemble those of the EP variable over longer horizons i.e. $m=24$ for monthly data and $m=12$ for quarterly, when $b=0.65, 0.75$.
Figure 3 reports values for the LTLS $\hat{T}$-statistics for the hypothesis $H_{0}:\beta=0$ vs $H_{1}:\beta\neq0$ (c.f. equation ((ref))). These values are plotted against the predictability horizon parameter $m$. We consider three configurations for kernels, cps and bandwidth sequences consistent with the set-ups S1, S2 and S3 given in the previous section. In particular, for S1 and S2 we choose $K(x)=\varphi_{0.1\tilde{\sigma}_{u}^{2}}(x)^{1/2}$, $K(x)^{\ast}=\varphi_{\tilde{\sigma}_{u}^{2}}(x)^{1/2}$. It can be seen from Figure 3 that there is evidence for predictability only for longer horizons under S1 and S3. For monthly data, the null hypothesis is rejected at a 5%\ level under S1 and S3 for for $m$ greater than 6 and 5 respectively. For quarterly data the null is rejected under S1 and S3 for $m$ greater than 12 and 10 respectively. These findings are consistent with those of Bandi and Perron (2008) how find strong predictability (by volatility predictors) over longer horizons.
Throughout the section, we assume that $C, C_0, C_1, C_2,...$ are positive constants that may take a different value in each appearance and let $K_{kn}:=\sum_{j=1}^{l_{n}}K\left[c_{n}(k/n-\tau_{j})\right]$ as in ((ref)).
We start with two preliminary lemmas, which provide significant extension to Lemma 4.1 of Hu, Phillips and Wang (2019) and include ((ref)) and ((ref)) as a corollary. The proofs of these two lemmas will be given in Sections 6.7 and 6.8, respectively.
Let $\{X_{n, k}\}_{k\geq 1,n\geq 1}$, where $X_{n, k}=(X_{nk,1},...,X_{nk,p})$, be a vector random array. When there is no confusion, we also use the notation $X_{nk}=X_{n,k}.$ Let $\{v_{k}\}_{k\ge 1}$ be a sequence of random variables, and $G(q)=G(q_{1},...,q_{p})$ and $K(x)$ be Borel functions on $ \mathbb{R} ^{p}$ and $ \mathbb{R} $, respectively. For $0<\tau _{1}<\tau _{2}<...<\tau _{l}<1$, set
where $\{c_{n}\}_{n\ge 1}$ is a sequence of positive constants. Our first result investigates the asymptotics of $S_{n,l}$.
If we are only interested in the boundedness of $S_{n, l}$, condition (b) can be reduced as seen in the following result.
We start with the limit result for $S_{1n,l_{n}}$, i.e., ((ref)). For $\mathbb{\alpha}\in\mathbb{R}^{p}$, let $v_k=\alpha'g(x_k)$. Since $\{x_k\}_{k\ge 1}$ is an ergodic stationary sequence with $E\|g(x_k)\|^{2+\delta}<\infty$ for some $\delta>0$, it is readily seen that $\{v_k\}_{k\ge 1}$ is stationary and ergodic, and condition (b) of Lemma (ref) holds with $A_0=Ev_1$ (see, for instance, Kallenberg (2002, Chapter 10)). ((ref)) follows from Lemma (ref) with $G(x)\equiv 1$.
We next consider $M_{1n,l_{n}}$, i.e., ((ref)). Set $Q_{k,n}:=\sqrt{\frac{c_{n}}{nl_n}}\mathbb{\alpha}^{\prime}g(x_{k})K_{kn}$ where $\mathbb{\alpha}\in \mathbb{R} ^{p}$. Note that
by using Lemmas (ref) and (ref) with $G(x)\equiv 1$, $v_k=\left[\mathbb{\alpha}^{\prime}g(x_{k})\right]^{2}$ and $A_0=E\left[\mathbb{\alpha}^{\prime}g(x_{k})\right]^{2}$. It follows from Hall and Heyde (1980, Theorem 3.2) or Wang (2014, Theorem 2.1) that, equation\ ((ref)) will follow, if we prove
Note that for any $A>0$,
Similar arguments used in ((ref)) show that the first term as $n\rightarrow\infty$ first and then as $A\rightarrow\infty$
By Lemma (ref) with $G(x)\equiv 1$ and $v_k=1$, as $n\rightarrow\infty$, the second term \[ II_{2n}(A)\leq \|\alpha\|^4A^{4}\left(\frac{c_{n}}{nl_n}\right)^{2}\sum_{k=1}^{n}K_{kn}^{4}=o_{P}(1). \] Combining these facts together, we establish ((ref)). The proof of Theorem (ref) is complete. $\Box$
The result for $S_{2n,l_{n}}$, i.e., ((ref)) follows from Lemma (ref) with $v_k\equiv 1$.
We next consider $M_{2n,l_{n}}$, i.e., ((ref)). Set $Q_{k,n}:=\sqrt{\frac{c_{n}}{nl_n}}\mathbb{\alpha}^{\prime}g(X_{nk})K_{kn}$ where $\mathbb{\alpha}\in \mathbb{R} ^{p}$. Noting that $\int_0^1 g(X_{n, [nt]})dt$ is a continuous functional of $X_{n,\left[nt\right]}$, the limit result of ((ref)), jointly with ((ref)), will follow if we prove that
on $D_{ \mathbb{R} ^{2}}[0,1]$. First note that, by using Lemmas (ref) and (ref) with $v_k\equiv 1$,
indicating that
It follows from Theorem 2.1 of Wang (2014), the limit result of\ ((ref)) will follow, if we prove
and
In fact, by recalling the fact that $||g||^4$ is still continuous, it follows from Lemma (ref) with $v_k=1$ again that \[ \Big[\max_{1\leq k\leq n}\left\vert Q_{k,n}\right\vert \Big]^{4}\leq\sum_{k=1}^{n}Q_{k,n}^{4}\leq \|\alpha\|^4\Big(\frac{c_{n}}{nl_n}\Big)^{2}\sum_{k=1}^{n}\left\Vert g(X_{nk})\right\Vert ^{4}K_{kn}^4=o_{P}(1), \] yielding ((ref)). Similarly, by recalling $l_n/c_n\to 0$, we have
which shows ((ref)). The proof of Theorem (ref) is complete. $\Box$
To show Theorem (ref), we only prove ((ref)) since ((ref)) is a direct consequence of ((ref)) and Theorem (ref).
Notice that, by the condition (b), we may write
where $\ R(\lambda,x)=[R_{1}(\lambda,x),...,R_{p}(\lambda,x)]^{\prime}$ and
Now ((ref)) follows from Theorem (ref) with $g(x)=H(x)$ if we prove
for any $\alpha=(\alpha_1,...,\alpha_p)'\in \mathbb{R} ^{p}$.
We only prove ((ref)) with $i=2$ since the proof of $|\alpha'\Delta_{1n}|=o_P(1)$ is similar except simpler. Recall $K_{kn}:=\sum_{j=1}^{l_{n}}K\left[c_{n}(k/n-\tau_{j})\right]$ and set, for $A>0$, \[ \widetilde{R}_{n,l_{n}}(A)=\sqrt{\frac{c_{n}}{nl_n}}\sum_{k=1}^{n}\alpha^{\prime}\pi\left(d_{n}\right)^{-1}R(d_{n},X_{nk})I\left\{ \left\vert X_{nk}\right\vert \leq A\right\}K_{kn}u_{k}. \] Note that as $n\rightarrow\infty$ first and then $A\rightarrow\infty$
For any $\epsilon>0$\ and $A>0$, we have \[ P\left(\left\vert\alpha'\Delta_{2n}\right\vert \geq\epsilon\right)\leq P\left(\alpha'\Delta_n\neq\widetilde{R}_{n,l_{n}}(A)\right)+\epsilon^{-2}E\left[\widetilde{R}_{n,l_{n}}(A)\right]^{2}. \] Now $|\alpha'\Delta_{2 n}| =o_P(1)$ follows from ((ref)) and the fact that as $n\rightarrow\infty$ for any $A>0$
where $\epsilon_{n}=\max_{1\le i\le p} |[\pi_i(d_n)]^{-1}a_i(d_n)|\rightarrow 0$ and we have used ((ref)) of Lemma (ref) (with $G(x)\equiv 1$ and $v_k\equiv 1$). The proof of Theorem (ref) is now complete.
{\it Proofs of ((ref)) and ((ref))} are essentially the same as that of ((ref)). We only provide a outline for ((ref)). For any $\alpha, \beta\in \mathbb{R}$, let
where $K_{kn}^{\ast}:=\sum_{j=1}^{l_{n}}K^{\ast}\left[c_{n}\left(k/n-\tau_{j}\right)\right]$. As in the proof of ((ref)), we have
Note that, by using ((ref)) and Lemmas (ref) and (ref),
indicating
Similarly, we may prove that ((ref)) and ((ref)) hold with $Q_{kn}$ being replaced by $\widetilde Q_{k,n}$. As a consequence, ((ref)) follows from Wang (2014) as in the proof of Theorem (ref). $\Box$
We only prove Theorem (ref) since the proof of Theorem (ref) is similar except simpler. Let
Recall ((ref)) and $Z_{kn}=f(x_k)K_{kn}$ and note that $\frac {c_n}{nl_n^*}{\sum_{k=1}^{n}K_{kn}^{\ast}}=\int K^*+o(1)$. It is readily seen from ((ref)) of Theorem (ref) and Theorem (ref) that
where
Similarly, we have
where
Since both $C_n$ and $A_n$ are continuous functionals of $X_{n, [nt]}$, a simple application of ((ref)) and ((ref)) yields that
as required. The proof of Theorem (ref) is complete. $\Box$
We only prove Theorem (ref) under conditions of Theorem (ref) since the proof under conditions of Theorem (ref) is similar. In addition to $A_{2n}, B_{1n}$, $B_{2n}$, $A_n$ and $B_n$ in the proof of Theorem (ref), we define
As in the proof of ((ref)), by letting $D_{n}=$diag$\left\{ \pi(d_{n})\sqrt{\lambda_{n}},\sqrt{\lambda_{n}^{\ast}}\right\} $, we have
Since $\tilde{\sigma}^{2}=\sigma_u^{2}+o_{P}(1)$ under given assumptions, by using the similar arguments as in the proofs of ((ref)) and ((ref)), it follows from ((ref)) that
as required. $\Box$
We only prove ((ref)), as the proof of ((ref)) is similar except more simpler. We start with the proof of ((ref)) by assuming that there exists an $A>0$ such that $K(x)=0$ if $|x|\ge A$ and $K(x)$ is Lipschitz continuous on $ \mathbb{R} $. This restriction will be removed later.
Without loss of generality, suppose $A=1$. Set $\delta_{1n, j}=[n(\tau_j-1/c_n)]\vee 1$, $\delta_{2n, j}=[n(\tau_j+1/c_n)]\vee 1$ and $\delta_{n, j}=[n\tau_j]$. Recall $\tau_j=j/(l_n+1)$. Since
by letting $R_{1n, j}=\frac {c_n}n \sum_{k=\delta_{1n, j}}^{\delta_{2n,j}}\ v_k\ K\big[c_n(k/n-\tau_j)\big]$ and
we have
Since $ \frac{1}{l_{n}}\sum_{j=1}^{l_{n}}\, G\big( X_{n, \delta_{n,j}}\big) =\int_0^1 G(X_{n, [nt]})dt+o_P(1)\to_d \int_0^1G(X_t)dt$, it suffices to show that
To prove ((ref)), we start with some preliminaries. Recalling $X_{n,[nt]}\Rightarrow X_t$ on $D_{ \mathbb{R} ^{p}}[0,1]$ and the limit process $X(t)$ is path continuous, we have $X_{n,[nt]}\Rightarrow X_t$ on $D_{ \mathbb{R} ^{p}}[0,1]$ in the sense of uniform topology. See, for instance, Section 18 of Billingsley (1968). This fact implies that
and by the tightness of $\{X_{n,[nt]}\}_{ 0\le t\le 1}$, for any $\varepsilon>0$ and $\delta>0$, there is some $\tilde{\delta}=\tilde{\delta}(\varepsilon,\delta)>0$ such that
holds for all sufficiently large $n$. In terms of ((ref)), for any $\delta>0$, we have
We are now ready to prove ((ref)), starting with $j=1$.
For any $N>0$, we let $G_N(x)=G(x)\xi_N(x)$ with
and
Note that, as $n\to\infty$ first and then $N\to\infty$,
and
where $C_N:=\sup_{x} |G_N(x)|<\infty$ is a constant depending only on $N$, due to the continuity of $G(x)$. Result ((ref)) with $j=1$ will follow if we prove
as $n\to\infty$. Indeed, by virtue of ((ref)) and ((ref)), we have $E|\widetilde R_{1n}|\to 0$ and then $\widetilde R_{1n}=o_P(1)$ for each $N\ge 1$. This, together with ((ref)), yields $ R_{1n}=o_P(1)$.
Since, as $n\to\infty$,
to prove ((ref)), it suffices to show that $\max_{1\le j\le l_n} E|A_n(\tau_j)| \to 0$, where
Let $\gamma=\gamma_n$ be integers such that $ \gamma \to \infty$ and $\gamma\,c_n/n\to 0$, $T_{1n, j}=[\delta_{1n, j}/\gamma]$ and $T_{2n, j}=[\delta_{2n, j}/\gamma]$. Noting ((ref)), we may write
Recall $\sup_{k\ge 1}E|v_k|<\infty$ by condition (b), it is readily from the the Lipschitz condition on $K(x)$ that
uniformly in $1\le j\le l_n$. Similarly, by using condition (b), we have
where
and we have used the fact that $ \max_{1\le j\le l_n} \Big|A_{4n} (\tau_j) -\int K \Big| \to 0. $ Combining all these facts, we prove ((ref)), and complete the proof of $ R_{1n}=o_P(1)$.
We next show $R_{2n}=o_P(1)$. Let $ \widetilde R_{2n}=\frac{1}{l_{n}}\sum_{j=1}^{l_{n}}\, \widetilde R_{2n, j}$, where
In terms of ((ref)), we have
as $n\to\infty$ first and then $N\to\infty$. Result $R_{2n}=o_P(1)$ will follow if we prove $ \widetilde R_{2n}=o_P(1),$ for each fixed $N\ge 1$.
Recall that $G_N(x)$ is continuous with compact support. For any $\epsilon>0$, there exists a $\delta_\epsilon>0$ so that $|G_N(x)-G_N(y)|\le \epsilon$ whenever $||x-y||\le \delta_\epsilon$. Write
By virtue of the facts above and ((ref)), it is readily seen that
where $C_N$ is a constant depending only on $N$. Now, for any $\eta_1>0$ and $ \eta_2>0$, let $\epsilon=\eta_1\eta_2$ and $n_0$ be large enough so that, for all $n\ge n_0$ [recall ((ref))],
It is readily seen that, for all $n\ge n_0$,
where $\bar \Omega_{\delta_\epsilon}$ denotes the complementary set of $ \Omega_{\delta_\epsilon}$ and $C_{N}$ is a constant depending only on $N$. This yields $ \widetilde R_{2n}=o_P(1),$ for each fixed $N\ge 1$, and completes the proof of $R_{2n}=o_P(1)$ .
We finally remove the restriction on $K$ and then conclude the proof of Lemma (ref). If $K$ has compact support, then there exists $A_1>0$ such that $K(x)=0$ holds for all $|x|\ge A_1$. If $K$ is eventually monotonic, then for any $\epsilon>0$, we can also choose a constant $A_1:=A_1(\epsilon)>0$ such that $K(x)$ is monotonic on $(-\infty, -A_1)$ and $(A_1,\infty)$ and $\int_{|x|>A_1} K(x)dx<\epsilon$ (in order to simplify the notations, here we use the same notation $A_1$ to denote the constant).
Since $K\ge 0$ with $\int K<\infty$, for any $\epsilon>0$, there exists an $A:=A_\epsilon\ge A_1+1$ such that
where $K_{\epsilon, A} (x)=0$ if $|x|\ge A$ and $K_{\epsilon, A}(x)$ is Lipschitz continuous on ${ \mathbb{R} }$. Let $\widetilde K(x)=K(x)-K_{\epsilon, A} (x) $ and
It suffices to show that, as $n\to\infty$ first and then $\epsilon\to 0$,
The proof of ((ref)) is similar to that of ((ref)). Indeed, by letting
we have
as $n\to\infty$ first and then $N\to\infty$. Hence it suffices to show that, for each fixed $N\ge 1$, $ S_{n, \epsilon, N}=o_P(1)$ as $n\to\infty$ first and then $\epsilon \to 0$. Note that
and, if $K(x)$ is monotonic on $(-\infty, -A)$ and $(A,\infty)$ then for sufficiently large $n$, uniformly for $1\le j\le l_n$,
Hence, in terms of the uniformed boundedness of $G_N(x)$, we have
as $n\to\infty$ first and then $\epsilon\to0$. Hence $ S_{n, \epsilon, N}=o_P(1)$ as $n\to\infty$ first and then $\epsilon \to 0$. The proof of ((ref)) is completed. $\Box$
We first prove ((ref)). Using similar arguments as in the proof of ((ref)) or ((ref)), it suffices to show that, as $n\to\infty,$
Take $\eta_{n,i,j}=\frac 12 n(\tau_i+\tau_j).$ Note that $c_n (k/n-\tau_i)\ge c_n(j-i)/(2l_n)$ if $k\ge \eta_{n,i,j}$ and $|c_n (k/n-\tau_j)|\ge c_n(j-i)/(2l_n)$ if $k\le \eta_{n,i,j}$. It follows from $K(x)\le C/(1+|x|)$ that
as required.
The proof of ((ref)) is similar to that of ((ref)) and hence the details are omitted. Result ((ref)) follows easily from ((ref)) and ((ref)). As for ((ref)), it follows from the similar arguments as in the proof of ((ref)) and the fact: as $n\to\infty$,
due to ((ref)) and ((ref)). $\Box$
Using similar arguments as in the proof of ((ref)) or ((ref)), it suffices to show that \[ \widetilde{I}_{n}:=\frac{c_{n}}{n\sqrt{l_{n}}}\sum_{k=1}^{n}\sum_{j=1}^{l_{n}}K\left[c_{n}\left(k/n-\tau_{j}\right)\right]K^{\ast} \left[c_{n}\left(k/n-\tau^{\ast}\right)\right]\rightarrow 0. \] We first assume that $l_n\rightarrow\infty$. For any $n\in \mathbb{N}$, there exists $i_n\in \mathbb{N}$ so that $i_n-1<\tau^{\ast}\le i_n$. Therefore, for any $j\ne i_n+1, i_n, i_n-1$, we have $|\tau_{j}-\tau^{\ast}|\ge c_n(|j-i_n|-1)/(2l_n)$. This implies that $|k/n-\tau_j|\ge c_n(|j-i_n|-1)/(4l_n)$ or $|k/n-\tau^{\ast}|\ge c_n(|j-i_n|-1)/(4l_n)$. Recall that $K(x)\leq C/(1+|x|)$ and $K^*(x)\leq C/(1+|x|)$, we have
Therefore, by noting that $l_n=o(c_n)$ and $l_n\rightarrow \infty$,
Next, we assume that $l_n=l$ and $\tau^*, \tau_j, j=1,\dots, l$ are fixed constants. If $\tau^{\ast}\neq \tau_j$, then
This implies $\widetilde{I}_n\rightarrow 0$ for $\tau^*\not \in \{\tau_j, j=1,\dots, l\}$. $\Box$