EconBase
← Back to paper

Adaptive inference for a semiparametric generalized autoregressive conditional heteroskedasticity model

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

86,175 characters · 22 sections · 88 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Adaptive inference for a semiparametric generalized autoregressive conditional heteroskedasticity model

frontmatter\begin{aug} , \and \thankstext{t1}{Correspondence to: Department of Statistics & Actuarial Science, University of Hong Kong, Pokfulam Road, Hong Kong. E-mail address: [email removed]} \address{ Center for Statistical Science\\ Department of Industry Engineering\\ Tsinghua University\\ Beijing 100084, China\\ \printead{e1}\\ \phantom{E-mail:\ }\printead*{e2} } \address {Department of Statistics\\ \quad & Actuarial Science\\ University of Hong Kong\\ Hong Kong\\ \printead{e3} } \end{aug} \begin{abstract} This paper considers a semiparametric generalized autoregressive conditional heteroskedasticity (S-GARCH) model. For this model, we first estimate the time-varying long run component for unconditional variance by the kernel estimator, and then estimate the non-time-varying parameters in GARCH-type short run component by the quasi maximum likelihood estimator (QMLE). We show that the QMLE is asymptotically normal with the parametric convergence rate. Next, we construct a Lagrange multiplier test for linear parameter constraint and a portmanteau test for model checking, and obtain their asymptotic null distributions. Our entire statistical inference procedure works for the non-stationary data with two important features: first, our QMLE and two tests are adaptive to the unknown form of the long run component; second, our QMLE and two tests share the same efficiency and testing power as those in variance targeting method when the S-GARCH model is stationary. {\bf JEL Classification}: C12, C14, C58. \end{abstract} \begin{keyword} \kwd{Adaptive inference; Lagrange multiplier test; Portmanteau test; QMLE; Semiparametric BEKK model; Semiparametric GARCH model} \end{keyword}

Introduction

Since the seminal work of Engle:1982 and Bollerslev:1986, the generalized autoregressive conditional heteroskedasticity (GARCH) model is perhaps the most influential one to capture and forecast the volatility of economic and financial return data. However, the GARCH model is often used under the stationarity assumption. Due to business cycle, technological progress, preference change and policy switch, the underlying structure of data may change over time (see Hansen:2001). Hence, a non-stationary GARCH model with time-varying parameters seems more appropriate to fit the return data in applications; see, for example, MS:2004, SG:2005, ER:2008, FSR:2008, PR:2014, Truquet:2017 and the references therein.

In this paper, we consider a semiparametric GARCH (S-GARCH) model of order $(p, q)$

flaligny_t=&\sqrt{\tau_t}u_t with \tau_t=\tau(t/T),\\ u_t=&\sqrt{g_t}\eta_t and g_t=\omega_0+\sum_{i=1}^q\alpha_{i0} u_{t-i}^2+\sum_{j=1}^{p}\beta_{j0} g_{t-j},

for $t=1,..., T$, where $\tau(x)$ is a positive smoothing deterministic function with unknown form on the interval $[0,1]$, $u_t$ is a covariance stationary GARCH$(p, q)$ process with $\omega_0>0$, $\alpha_{i0}\geq0$ and $\beta_{j0}\geq0$, and $\{\eta_t\}$ is a sequence of independent and identically distributed (i.i.d) random variables with $E\eta_t^2=1$. The specification that $\tau_t$ is a function of ratio $t/T$ rather than time $t$ is initiated by Robinson:1989, and since then, it has become a common scaling scheme in the time series literature; see, for example, DR:2006, CT:2007, XP:2008, ZW:2009, ZhangW:2012, ZS:2013, and Zhu:2019 to name just a few. In ((ref))--((ref)), the smooth long run component $\tau_t$ is to depict time-varying parameters in volatility, and the GARCH-type short run component $u_t$ is to capture the temporal dependence.

By using different specified forms of $\tau(x)$, the S-GARCH model nests many often used models, including, for example, the standard GARCH model in Bollerslev:1986, the spline-GARCH model in ER:2008, and the smooth-transition GARCH model in AT:2013. The statistical inference for these models has been well studied. However, when the specification of $\tau(x)$ is unspecified, the statistical inference for the S-GARCH model has been less attempted. For $p=q=1$, HL:2010 considered the estimation for the S-GARCH model. For $p=0$ (i.e., $\beta_{j0}\equiv0$), PR:2014 constructed a score test to check the nullity of all $\alpha_{i0}$, and Truquet:2017 later proposed a projection-based estimation and a related Wald test to detect the nullity of some of $\alpha_{i0}$. For the general S-GARCH model, the statistical inference methodologies, including estimation, testing and model checking, are not available in the literature.

In this paper, we provide an entire inference procedure for the S-GARCH model to fill this gap. First, we give a two-step estimation for the model: the function $\tau(x)$ is estimated by the kernel estimator at step one, and the unknown parameter vector in the parametric process $u_t$ is estimated by the quasi maximum likelihood estimator (QMLE) at step two. Although the nonparametric estimator at step one has a slower convergence rate, we show that the QMLE at step two is asymptotically normal with a parametric convergence rate. Moreover, we construct a new Lagrange multiplier (LM) test for detecting the linear parameter constraint, and propose a new portmanteau test for model checking. The asymptotic null distributions of the LM and portmanteau tests are established. Since our entire inference methodologies allow for unspecified form of $\tau(x)$ and higher order $(p, q)$, they alleviate the potential risk of model-misspecification, leading to a broad application scope to handle the non-stationary data. Finally, we extend the two-step estimation to a multivariate semiparametric BEKK (S-BEKK) model, and establish the asymptotic normality of the corresponding QMLE.

Our two-step estimation was previously adopted by HL:2010 to study the multivariate S-BEKK(1, 1) model. For the univariate S-GARCH model, we find a much simpler expression for the asymptotic variance of the QMLE, making the related inference methodologies easy-to-implement. Meanwhile, we find that the asymptotic variance of the QMLE is adaptive to the unknown form of $\tau(x)$. Consequently, the efficiency of the QMLE and the power of its related LM and portmanteau tests are invariant regardless of the form of $\tau(x)$. However, we can show that the QMLE of the multivariate S-BEKK model no longer enjoys such an adaptiveness feature as in the univariate S-GARCH model. Our two-step estimation also shares the similar idea as the variance targeting (VT) estimation in FHZ:2011, which is only applicable for the stationary S-GARCH model (i.e., $\tau(x)\equiv \tau_0$). The difference is that our first step estimator of $\tau(x)$ is non-parametric, while the first step estimator of $\tau_0$ in the VT method is parametric. It turns out that our method requires more involved proof techniques. Interestingly, when the S-GARCH model is stationary, our QMLE is asymptotically as efficient as the QMLE in the second step estimation of the VT method, although the first step estimator of our method has a slower convergence rate than that of the VT method. On the contrary, when the S-GARCH is non-stationary, our QMLE is still valid with the same efficiency as the stationary case due to its adaptiveness feature, while the QMLE in the VT method is not applicable any more.

The remainder of the paper is organized as follows. Section 2 presents the two-step estimation procedure and establishes its related asymptotics. Section 3 gives a LM test for the linear parameter constraint. Section 4 introduces a portmanteau test and obtains its limiting null distribution. Section 5 makes a comparison with other estimation methods. Section 6 extends the two-step estimation into the multivariate S-BEKK model. Simulation results are reported in Section 7, and applications are given in Section 8. Concluding remarks are offered in Section 9. Proofs of all theorems are relegated to the Appendix.

Two-step estimation

Let $\theta=(\alpha_{1},...,\alpha_{q},\beta_{1},...,\beta_{p})'\in\Theta$ be the parameter vector in model ((ref)), and $\theta_0=(\alpha_{10},...,\alpha_{q0},\beta_{10},...,\beta_{p0})'\in\Theta$ be its true value, where $\Theta\subset \mathbb{R}_+^{p+q}$ is the parameter space, and $\mathbb{R}_+=(0,\infty)$. This section gives a two-step estimation procedure for the S-GARCH model in ((ref))--((ref)). Our procedure first estimates the nonparametric function $\tau(x)$ in ((ref)), and then estimates the parameter vector $\theta_0$ in ((ref)).

Estimation of $\tau(x)$

This subsection provides a (Nadaraya-Watson) kernel estimator of $\tau(x)$. To this end, we first need an assumption for the identification of $\tau_t$.

ass$(\mathrm{i})$ $\sum_{i=1}^{q}\alpha_i+\sum_{j=1}^{p}\beta_j<1$; $(\mathrm{ii})$ $\omega=1-\sum_{i=1}^{q}\alpha_i-\sum_{j=1}^{p}\beta_j$.

Assumption (ref)(i) is equivalent to the covariance stationarity of model ((ref)), and Assumption (ref)(ii) is to ensure $Eu_t^2=1$. Under Assumption (ref), we have $$ y_t^2=\tau(t/T)+\tau(t/T)(u_t^2-1):=\tau(t/T)+v_t, $$ where $v_t:=\tau(t/T)(u_t^2-1)$ is a zero-mean process. In other words, $y_t^2$ can be rewritten as a standard non-parametric regression problem with a time-varying mean. Following Hafner and Linton (2010), it is reasonable to estimate $\tau(x)$ by

flalign*\widetilde{\tau}(x)=\frac{\sum_{s=1}^{T}K_h\big(x-\frac{s}{T}\big)y_s^2}{\sum_{s=1}^{T}K_h\big(x-\frac{s}{T}\big)},

where $K_h(\cdot)=K(\cdot/h)/h$ with $K(\cdot)$ being a kernel function and $h$ being a bandwidth. Since $(1/T)\sum_{s=1}^{T}K_h(x-s/T)=1+O(1/(Th))$ under mild conditions, it is more convenient to estimate $\tau(x)$ by

flalign\widehat{\tau}(x)=\frac{1}{T}{\sum_{s=1}^{T}K_h\Big(x-\frac{s}{T}\Big)y_s^2}.

To obtain the asymptotic distribution of $\widehat{\tau}(x)$, the following assumptions are needed.

ass$(\mathrm{i})$ $\tau:[0,1]\to \mathbb{R}_{+}$ is twice continuously differentiable; $(\mathrm{ii})$ $0<\underline{\tau}\leq \inf_{x\in[0,1]}\tau(x)\leq \sup_{x\in[0,1]}\tau(x)\leq \overline{\tau}$, where $\underline{\tau}$ and $\overline{\tau}$ are two positive constants.
ass$(\mathrm{i})$ $K:[-1,1]\to \mathbb{R}_+$ is symmetric about zero, bounded and Lipschitz continuous with $\int_{-1}^{1}K(x)dx=1$ and $C_r=\int_{-1}^{1}x^rK(x)dx$; $(\mathrm{ii})$ $h\to0$ and $Th\to\infty$ as $T\to \infty$.
ass$Eu_{t}^4<\infty$.

Assumption (ref)(i) imposes a smoothness condition on $\tau(x)$, and similar conditions have been used in DR:2006, HL:2010, and CH:2016. Assumption (ref)(ii) is in line with the condition that the intercept term in the standard GARCH model has positive lower and upper bounds. Assumption (ref)(i) holds for many often used kernels, and the bounded support condition on $K(x)$ is just to simplify analysis. Assumption (ref)(ii) requires that $h$ converges to zero at a slower rate than $T^{-1}$, and later a more restrictive $h$ is needed for the asymptotics of the estimator of $\theta_0$. Assumption (ref) is stronger than Assumption (ref)(i), and it is used to ensure that the asymptotic variance of $\widehat{\tau}(x)$ is well defined.

Let $z_t=u_t^2-1$. The asymptotic normality of $\widehat{\tau}(x)$ is given below.

thmSuppose Assumptions (ref)--(ref) hold. Then, for any $x \in (0,1)$, $$\sqrt{Th}\big(\widehat{\tau}(x)-\tau(x)-h^2b(x)\big)\to_{\mathcal{L}}N(0,V(x))\mbox{ as }T\to\infty,$$ where `$\to_{\mathcal{L}}$' stands for the convergence in distribution, $$b(x)=\frac{C_2}{2}\frac{\partial^2 \tau(x)}{\partial x^2}\mbox{ and }V(x)=\tau^2(x)\Big\{\int_{-1}^{1} K^2(x)dx\Big\}\sum_{j=-\infty}^{\infty}E(z_tz_{t-j}).$$

Based on $\widehat{\tau}(x)$ in ((ref)), we estimate $\tau_t$ by $\widehat{\tau}_t=\widehat{\tau}(t/T)$. In practice, $\widehat{\tau}_t$ may have the boundary problem. To circumvent this problem, we follow CH:2016 to adopt the reflection method proposed by HW:1991. That is, we generate pseudo data $y_t=y_{-t}$ for $-[Th]\leq t\leq -1$ and $y_t=y_{2T-t}$ for $T+1\leq t\leq T+[Th]$, and then modify $\widehat{\tau}_t$ as

flalign\widehat{\tau}_t=\frac{1}{T}\sum_{s=t-[Th]}^{t+[Th]}K_h\Big(\frac{t-s}{T}\Big)y_s^2.

Intuitively, the reflection method makes the boundary points behave similarly as the interior ones. Similar to CH:2016, it can be seen that the reflection method gives a bias term of order $O(h^2)$, and hence it does not affect the asymptotics of the estimator of $\theta_0$. Although $\widehat{\tau}_t$ in ((ref)) is used for numerical calculations, our proofs below will be based on $\widehat{\tau}_t=\widehat{\tau}(t/T)$ in ((ref)) to ease the presentation.

Estimation of $\theta_0$

This subsection considers the QMLE of $\theta_0$. Based on Assumption (ref)(ii), we write the parametric $g_t$ in ((ref)) as

equation[equation omitted — 171 chars of source]

By assuming that $\eta_{t}\sim N(0, 1)$, the log-likelihood function (multiplied by -2 and ignoring constants) of $\{y_t\}$ is

equation[equation omitted — 151 chars of source]

Unfortunately, $L_T(\theta)$ is infeasible for computation, since $\{u_t\}$ are unobservable. Thus, we have to replace $\{u_t\}$ by $\{\widehat{u}_{t}\}$, and consider the following feasible log-likelihood function

equation[equation omitted — 205 chars of source]

where $\widehat{u}_{t}=y_{t}/\sqrt{\widehat{\tau}_{t}}$, and $\widehat{g}_t(\theta)$ is computed recursively by

equation[equation omitted — 199 chars of source]

with given constant initial values $$\widehat{u}_0=u_0,...,\, \widehat{u}_{1-q}=u_{q-1},\, \widehat{g}_0(\theta)=g_0,...,\, \widehat{g}_{1-p}(\theta)=g_{1-p}.$$ Based on $\widehat{L}_T(\theta)$ in ((ref)), our QMLE of $\theta_0$ is defined as $$ \widehat{\theta}_T=\arg\min_{\theta\in\Theta}\widehat{L}_T(\theta). $$

To establish the asymptotics of $\widehat{\theta}_{T}$, denote $\mathcal{A}_{\theta}(z)=\sum_{i=1}^{q}\alpha_iz^i$ and $\mathcal{B}_{\theta}(z)=1-\sum_{i=1}^{p}\beta_iz^i$ with the convention $\mathcal{A}_{\theta}(z)=0$ if $q=0$ and $\mathcal{B}_{\theta}(z)=1$ if $p=0$. The following additional assumptions are imposed.

ass$(\mathrm{i})$ $\Theta$ is compact; $(\mathrm{ii})$ if $p>0$, the polynomials $\mathcal{A}_{\theta_0}(z)$ and $\mathcal{B}_{\theta_0}(z)$ have no common roots, $\mathcal{A}_{\theta_0}(1)\neq 0$, and $\alpha_{q0}+\beta_{p0}\neq 0$; $(\mathrm{iii})$ $\theta_0$ is an interior point of $\Theta$.
ass$E|u_t|^{4(1+\delta_0)}<\infty$ for some $\delta_0>0$.
ass$(\mathrm{i})$ $\eta_t$ has a continuous and almost surely positive density on $\mathbb{R}$ with $E\eta_t^2=1$; $(\mathrm{ii})$ $E|\eta_t|^{4+4/\delta_0+\delta_1}<\infty$ for some $\delta_1>0$, where $\delta_0>0$ is defined as in Assumption (ref).
ass$h=c_hT^{-\lambda_h}$ for some $1/4<\lambda_h<1/2$ and $0<c_h<\infty$.

We offer some remarks on the aforementioned assumptions. Assumption (ref) is regular, and it has been used in HK:2003 and FZ:2004 to study the QMLE for the stationary GARCH model. Assumption (ref) is stronger than Assumption (ref), which is needed for the variance target estimator in FHZ:2011 but not for the QMLE in FZ:2004. Assumption (ref)(i) gives the identification condition for $\theta_0$ based on the QMLE, and ensures that the GARCH process $u_t$ is $\beta$-mixing (see CC:2002). Assumption (ref)(ii) is stronger than the condition $E\eta_t^4<\infty$, which is necessary to derive the asymptotic normality of the QMLE for the stationary GARCH model (see HY:2003). We resort to the stronger conditions of $u_t$ and $\eta_t$ in Assumptions (ref) and (ref)(ii) due to the existence of $\tau(x)$ in the S-GARCH model. Note that if $\eta_t$ has a light tail (for example, $\eta_t\sim N(0, 1)$), Assumption (ref)(ii) holds for a small value of $\delta_0$, and $u_t$ (or the data $y_t$) in Assumption (ref) is thus allowed to be heavy-tailed. Assumption (ref) requires a more restrictive condition on the bandwidth $h$ than Assumption (ref)(ii), and similar conditions have been adopted by HL:2010, PR:2014, and Truquet:2017. The reason is because an undersmoothing $h$ is needed to make the estimation bias from $\widehat{\tau}_t$ negligible so that the $\sqrt{T}$-convergence of $\widehat{\theta}_{T}$ holds.

Denote $\kappa=E\eta_{t}^{4}$, $g_t=g_{t}(\theta_0)$, $\psi_t=\psi_t(\theta_0)$ with $\psi_t(\theta)=\{\partial g_t(\theta)/\partial \theta\}/g_t(\theta)$, and

flalignJ_{1}&=E(\psi_t\psi_t'),\,\,\, J_{2}=E(g_t^2)E\big(\psi_t/g_t\big)E\big(\psi_t'/g_t\big).

Now, we are ready to give the asymptotics of $\widehat{\theta}_T$ in the following theorem.

thmSuppose Assumptions (ref)--(ref), (ref)(i)--(ii), and (ref)--(ref) hold. Then, $\mathrm{(i)}$ $\widehat{\theta}_{T}\to_{p} \theta_0$ as $T\to\infty$; $\mathrm{(ii)}$ furthermore, if Assumption (ref)(iii) holds and Assumption (ref)(ii) is replaced by Assumption (ref), $$ \sqrt{T}(\widehat{\theta}_T-\theta_0)\to_{\mathcal{L}} N(0,\Sigma)\mbox{ as }T\to\infty, $$ where $\Sigma=(\kappa-1)J_1^{-1}(J_1+J_2)J_{1}^{-1}$, and $J_1$ and $J_2$ are defined in ((ref)).
remWe can simply estimate $\Sigma$ by its sample version $\widehat{\Sigma}_{T}$, where \begin{flalign} \widehat{\Sigma}_{T}=(\widehat{\kappa}_{T}-1)\widehat{J}_{1T}^{-1}(\widehat{J}_{1T}+\widehat{J}_{2T})\widehat{J}_{1T}^{-1} \end{flalign} with \begin{flalign} \begin{split} \widehat{\kappa}_{T}=\frac{1}{T}\sum_{t=1}^{T}\widehat{\eta}_t^4,\,\,\widehat{J}_{1T}=\frac{1}{T}\sum_{t=1}^{T}\widehat{\psi}_t\widehat{\psi}_t' and \,\, \widehat{J}_{2T}=\Big(\frac{1}{T}\sum_{t=1}^{T}\widehat{g}_t^2\Big)\Big(\frac{1}{T}\sum_{t=1}^{T}\frac{\widehat{\psi}_t}{\widehat{g}_t}\Big) \Big(\frac{1}{T}\sum_{t=1}^{T}\frac{\widehat{\psi}_t'}{\widehat{g}_t}\Big). \end{split} \end{flalign} Here, $\widehat{\eta}_{t}=\widehat{\eta}_{t}(\widehat{\theta}_{T})$ with $\widehat{\eta}_{t}(\theta)=\widehat{u}_{t}/\sqrt{\widehat{g}_{t}(\theta)}$, $\widehat{\psi}_{t}=\widehat{\psi}_{t}(\widehat{\theta}_{T})$ with $\widehat{\psi}_{t}(\theta)=\{\partial \widehat{g}_t(\theta)/\partial \theta\}/\widehat{g}_t(\theta)$, and $\widehat{g}_{t}=\widehat{g}_{t}(\widehat{\theta}_{T})$. Under the conditions of Theorem (ref), we have $\widehat{\Sigma}_{T}\to_{p}\Sigma$ as $T\to\infty$.

Interestingly, the preceding theorem shows that the asymptotic variance of $\widehat{\theta}_T$ is independent of $\tau(x)$. Following the viewpoint of Robinson:1987, it means that $\widehat{\theta}_T$ is adaptive to the unknown form of $\tau(x)$. This adaptiveness feature ensures that the efficiency of $\widehat{\theta}_T$ and the power of its related tests are unchanged regardless of the form of $\tau(x)$.

The LM test

Since Engle:1982 and Bollerslev:1986, testing for the nullity of the parameters in the GARCH model is important in applications. This problem can be further generalized to consider the following linear constraint hypothesis

equation[equation omitted — 56 chars of source]

where $R$ is a given $d\times (p+q)$ matrix of rank $d$, and $r$ is a given $d\times 1$ constant vector. In this section, we construct a Lagrange multiplier (LM) test statistic $LM_T$ for $\mathbb{H}_0$, where $$LM_T=\frac{1}{T}\frac{\partial\widehat{L}_T(\widehat{\theta}_{T|0})}{\partial\theta'} \widehat{J}_{1T|0}^{-1}R'\big(R\widehat{\Sigma}_{T|0}R'\big)^{-1}R\widehat{J}_{1T|0}^{-1} \frac{\partial\widehat{L}_T(\widehat{\theta}_{T|0})}{\partial\theta}.$$ Here, $\widehat{\theta}_{T|0}$ is the constrained QMLE of $\theta_0$ under $\mathbb{H}_0$, and $\widehat{J}_{1T|0}$ and $\widehat{\Sigma}_{T|0}$ are defined in the same way as $\widehat{J}_{1T}$ and $\widehat{\Sigma}_{T}$, respectively, with $\widehat{\theta}_{T}$ replaced by $\widehat{\theta}_{T|0}$. The following theorem gives the limiting null distribution of $LM_T$.

thmSuppose the conditions in Theorem (ref)(i) hold, with Assumption (ref)(ii) replaced by Assumption (ref). Then, under $\mathbb{H}_0$, $$ LM_T\to_{\mathcal{L}}\chi^2_d\mbox{ as }T\to\infty, $$ where $\chi^2_{d}$ is the chi-square distribution with the degrees of freedom $d$.

Based on Theorem (ref), we can set the rejection region of $LM_{T}$ at level $\alpha$ as $\{LM_T>\chi_d^2(\alpha)\},$ where $\chi_d^2(\alpha)$ is the $\alpha$-upper percentile of $\chi_d^2$.

As $\widehat{\theta}_{T}$, our $LM_T$ has the adaptiveness feature, and it has a much broader application scope than the existing LM tests. Specifically, the LM test in Bollerslev:1986 is only applicable for the stationary GARCH model, but our $LM_T$ has the superior ability to tackle the non-stationary S-GARCH model. For the case of $p=0$, the score test in PR:2014 can detect the null hypothesis that all $\alpha_{i0}$ are zeros, and the Wald test in Truquet:2017 can check the null hypothesis that some of $\alpha_{i0}$ are zeros. However, it seems non-trivial to extend these two tests for the general null hypotheses in (3.1), although the score test in PR:2014 can be extended to detect the null hypothesis that all $\alpha_{i0}$ and $\beta_{j0}$ are zeros. Besides $LM_T$, the Wald and likelihood ratio (LR) tests could also be constructed for $\mathbb{H}_0$. When some of $\alpha_{i0}$ or $\beta_{j0}$ are allowed to be zeros under $\mathbb{H}_0$, the Wald and LR tests render non-standard limiting null distributions (see FZ:2009 for general discussions), which have to be simulated by the bootstrap method. In contrast, $LM_T$ always has the standard chi-square limiting null distribution, even when all of the null coefficients are not pinned down in $\mathbb{H}_0$\footnote{Following the arguments in FZ:2007, our QMLE $\widehat{\theta}_{T}$ can not be asymptotically normal if $\theta_0$ lies on the boundary of $\Theta$ (i.e., some of $\alpha_{i0}$ or $\beta_{j0}$ are zeros). Since the Wald (including $t$) and LR tests depend on $\widehat{\theta}_{T}$, they can not have the standard chi-square limiting null distribution any more if $\theta_0$ lies on the boundary of $\Theta$ under $\mathbb{H}_0$. Unlike Wald and LR tests, the limiting distribution of our LM test $LM_{T}$ depends on the one of $(RJ_1^{-1}R')^{-1}RJ_1^{-1}\frac{1}{\sqrt{T}}\frac{\partial\widehat{L}_T(\widehat{\theta}_{T|0})}{\partial\theta}$, which is always asymptotically normal no matter whether $\theta_0$ lies on the boundary of $\Theta$ or not. Hence, it turns out that $LM_T$ always has the standard chi-square limiting null distribution. For more discussions on this context, we refer to Pedersen:2017 and JLZ:2020.}. For practical convenience, we thus only focus on the LM test in this paper, and the consideration of Wald and LR tests is left for future study.

Portmanteau test

Since LB:1978, the portmanteau test and its variants have been a common tool for checking the model adequacy in time series analysis. For the stationary GARCH model, LM:1994 proposed a portmanteau test for model checking. However, their test is invalid for the non-stationary S-GARCH model. In this section, we follow the idea of LM:1994 to construct a new portmanteau test to check the adequacy of S-GARCH model, and our test seems to be the first formal try in the context of semiparametric time series analysis.

Let $\widehat{\eta}_{t}$ be the model residual defined as in ((ref)). The idea of our portmanteau test is based on the fact that $\{\eta_{t}^2\}$ is a sequence of uncorrelated random variables under ((ref))--((ref)). Hence, if the S-GARCH model is correctly specified, it is expected that the sample autocorrelation function of $\{\widehat{\eta}_{t}^2\}$ at lag $k$, denoted by $\widehat{\rho}_{T,k}$, is close to zero, where $$\widehat{\rho}_{T,k}=\frac{\sum_{t=k+1}^{T}\big(\widehat{\eta}^2_{t} -\overline{\widehat{\eta}^2}\big)\big(\widehat{\eta}^2_{t-k}-\overline{\widehat{\eta}^2}\big)}{\sum_{t=1}^{T}\big(\widehat{\eta}^2_{t} -\overline{\widehat{\eta}^2}\big)^2}$$ with $\overline{\widehat{\eta}^2}$ being the sample mean of $\{\widehat{\eta}_{t}^2\}$. Let $\widehat{\rho}_T=(\widehat{\rho}_{T,1},...,\widehat{\rho}_{T,\ell})'$ for an integer $\ell\geq1$, and

flalign\Sigma_{P1}&=(I_{\ell}, -H,-DJ_1^{-1})\in \mathbb{R}^{\ell\times(\ell+1+p+q)},\\ \Sigma_{P2}&= \left(\begin{matrix} (\kappa-1)I_{\ell}& F & D-FE\big(\frac{\psi_t'}{g_t}\big)\\ *&Eg_t^2&-Eg_t^2E\big(\frac{\psi_t'}{g_t}\big)\\ *&*&J_1+J_2 \end{matrix}\right)\in \mathbb{R}^{(\ell+1+p+q)\times(\ell+1+p+q)}

be a symmetric matrix, where $D=(D_1',..., D_\ell')'$ with $D_k=E\{(\eta_{t-k}^2-1)\psi_t'\}$, $H=(H_1,..., H_\ell)'$ with $H_k=E\{g_t^{-1}(\eta_{t-k}^2-1)\}$, and $F=(F_1,...,F_{\ell})'$ with $F_k=E\{g_t(\eta_{t-k}^2-1)\}$. To facilitate our portmanteau test, we need the limiting distribution of $\widehat{\rho}_T$ below.

thmSuppose the conditions in Theorem (ref)(ii) hold. Then, if the S-GARCH model in ((ref))--((ref)) is correctly specified, $$ \sqrt{T}\widehat{\rho}_T\to_{\mathcal{L}} N(0,\Sigma_{P})\mbox{ as }T\to\infty, $$ where $\Sigma_{P}=(\kappa-1)^{-1}\Sigma_{P1}\Sigma_{P2}\Sigma_{P1}'$, and $\Sigma_{P1}$ and $\Sigma_{P2}$ are defined in ((ref))--((ref)).

As in Remark (ref), $\Sigma_{P}$ can be consistently estimated by its sample version $\widehat{\Sigma}_{P}$. Based on $\widehat{\Sigma}_{P}$, our portmanteau test statistic is defined as $$Q_{T}(\ell)=T\widehat{\rho}_T'\widehat{\Sigma}_P^{-1}\widehat{\rho}_T.$$ If the S-GARCH model is correctly specified, $Q_{T}(\ell)\to_{\mathcal{L}}\chi^2_{\ell}$ as $T\to\infty$ by Theorem (ref). So, if the value of $Q_{T}(\ell)$ is larger than $\chi_\ell^2(\alpha)$, the fitted S-GARCH model is inadequate at level $\alpha$. Otherwise, it is adequate at level $\alpha$. In practice, the choice of lag $\ell$ depends on the frequency of the series, and one often chooses $\ell$ to be $O(\log(T))$, delivering 6, 9 or 12 for a moderate $T$. We shall hightlight that $Q_{T}(\ell)$ also has the adaptiveness feature as $LM_T$, and it is essential to detect the adequacy of the short run GARCH component $u_t$ but not the long run component $\tau_t$, since the form of $\tau_t$ is unspecified in the S-GARCH model.\footnote{To detect whether $\tau_t$ is a constant over time (i.e., $y_t$ follows a standard GARCH model), one can use the strict stationarity test in Hong:2017 to check the variance stationarity of $y_t$.}

Comparisons with other estimation methods

This section compares our two-step estimation method with the three-step estimation method in HL:2010 and the variance targeting (VT) estimation method in FHZ:2011.

Comparison with three-step estimation method

Our two-step estimation method is the same as the first two estimation steps in HL:2010, where they gave the following asymptotic normality result for the S-GARCH($1, 1$) model $$\sqrt{T}(\widehat{\theta}_T-\theta_0)\to_{\mathcal{L}} N(0,\Sigma_{\dag})\mbox{ as }T\to\infty, $$ where $\Sigma_{\dag}=J_1^{-1}[(\kappa-1)J_1+J_3+J_4+J_4']J_1^{-1}$ with $J_3=(M-E\psi_t)(M-E\psi_t)'Z_1$, $J_4=Z_2(M-E\psi_t)'$,

flalign*M=\sum_{j=0}^{\infty}\alpha_{10}\beta_{10}^{j}E\Big(\frac{u^2_{t-j-1}\psi_t}{g_t}\Big),\,\, Z_{1}=\sum_{j=-\infty}^{\infty}E(z_tz_{t-j})\,\, and \,\,Z_{2}=\sum_{j=0}^{\infty}E\big\{z_{t}(\eta_{t-j}^2-1)\psi_{t-j}\big\}.

Indeed, we can show that $\Sigma_{\dag}$ and $\Sigma$ are equivalent. Since $\Sigma_{\dag}$ involves three infinite summations $M$, $Z_1$ and $Z_2$, a consistent estimator for $\Sigma_{\dag}$ then involves laborious tuning and smoothing. On the contrary, our $\Sigma$ has a much simpler expression, and it can be directly estimated as shown in Remark (ref).

In HL:2010, they further proposed an updated estimator at step three, and claimed this updated estimator can achieve the semiparametric efficiency bound when $\eta_t\sim N(0, 1)$. Following their idea, we can also update our estimator $\widehat{\theta}_{T}$ to $\widecheck{\theta}_{T}$ at step three. Specifically, we first update the nonparametric part estimator $\widehat{\tau}_t$ to $\widecheck{\tau}_{t}=\widecheck{\tau}(t/T)$, where

flalign*\widecheck{\tau}(x)=\widehat{\tau}(x)-\Big[\frac{1}{T}\sum_{t=1}^{T}K_h\Big(x-\frac{t}{T}\Big)\frac{\partial^2 \widehat{l}_t(\widehat{\tau},\widehat{\theta}_T)}{\partial\tau^2}\Big]^{-1}\Big[\frac{1}{T}\sum_{t=1}^{T}K_h\Big(x-\frac{t}{T}\Big)\frac{\partial \widehat{l}_t(\widehat{\tau},\widehat{\theta}_T)}{\partial\tau}\Big]

with $$ \widehat{l}_t({\tau},\widehat{\theta}_T)=\log\widehat{g}_t(\widehat{\theta}_T)+\log(\tau)+\frac{y_t^2}{{\tau}\widehat{g}_t(\widehat{\theta}_T)}. $$ Then, based on $\widecheck{u}_t^2=y_t^2/\widecheck{\tau}_t$ and some given initial values, we calculate $$ \widecheck{g}_t(\theta)=1-\sum_{i=1}^{q}\alpha_i-\sum_{j=1}^{p}\beta_j+\sum_{i=1}^{q}\alpha_i\widecheck{u}_{t-i}^2 +\sum_{j=1}^{p}\beta_j\widecheck{g}_{t-j}(\theta), $$ and update the parametric part estimator $\widehat{\theta}_{T}$ to $\widecheck{\theta}_{T}$ as follows: $$ \widecheck{\theta}_T=\widehat{\theta}_T- \Big[\frac{\partial^2\widecheck{L}_T^*(\widehat{\theta}_T)}{\partial\theta\partial\theta'}\Big]^{-1} \frac{\partial\widecheck{L}_T^*(\widehat{\theta}_T)}{\partial\theta}, $$ where

flalign*\frac{\partial\widecheck{L}_T^*(\widehat{\theta}_T)}{\partial\theta}=&\sum_{t=1}^{T}\Big[\widecheck{g}_t(\widehat{\theta}_T)^{-1}\frac{\partial\widecheck{g}_t(\widehat{\theta}_T)}{\partial\theta}-\widecheck{G}_T(\widehat{\theta}_T)\Big](1-\widecheck{\eta}_t(\widehat{\theta}_T)^2),\\ \frac{\partial^2\widecheck{L}_T^*(\widehat{\theta}_T)}{\partial\theta\partial\theta'}=&\sum_{t=1}^{T}\Big[\widecheck{g}_t(\widehat{\theta}_T)^{-2}\frac{\partial\widecheck{g}_t(\widehat{\theta}_T)}{\partial\theta}\frac{\partial\widecheck{g}_t(\widehat{\theta}_T)}{\partial\theta'}-\widecheck{G}_T(\widehat{\theta}_T)\widecheck{G}(\widehat{\theta}_T)'\Big]

with $\widecheck{G}_T(\theta)=\frac{1}{T}\sum_{t=1}^{T}\widecheck{g}_t(\theta)^{-1}\frac{\partial\widecheck{g}_t(\theta)}{\partial\theta}$. Below, we give the limiting distribution of $\widecheck{\theta}_T$.

thmSuppose the conditions in Theorem (ref)(ii) hold. Then, $$ \sqrt{T}(\widecheck{\theta}_T-\theta_0)\to_{\mathcal{L}} N(0,\Sigma^*)\mbox{ as }T\to\infty, $$ where $\Sigma^{*}=(\kappa-1)J_1^{*-1}(J_1^*+J_2^*)J_{1}^{*-1}$ with \begin{flalign*} J_1^*=&E\{(\psi_t-E\psi_t)(\psi_t-E\psi_t)'\},\\ J_2^*=&\frac{\omega_0^2}{\gamma_0^2}\Big[Eg_t^{-1}E\psi_t-Eg_t^{-1}\psi_t\Big]\Big[Eg_t^{-1}E\psi_t-Eg_t^{-1}\psi_t\Big]', \end{flalign*} and $\omega_0=1-\sum_{i=1}^{q}\alpha_{i0}-\sum_{j=1}^{p}\beta_{j0}$ and $\gamma_0=1-\sum_{j=1}^{p}\beta_{j0}$.

The preceding theorem shows that $\widecheck{\theta}_T$ can not achieve the semiparametric efficiency bound as $J_2^*$ is positive definite. Hence, it seems unnecessary to consider the third estimation step in HL:2010. Note that the above updating procedure was also given by BKRW:1993, in which they showed the updated estimator can achieve the semiparametric efficiency bound when the data are independent. However, when the data are dependent, their conclusion may not be true as demonstrated by Theorem (ref). The failure of $\widecheck{\theta}_T$ in our case possibly results from the violation of the following condition

flalign\frac{1}{\sqrt{T}}\Big\{\frac{\partial\widecheck{L}_T^*(\widehat{\theta}_T)}{\partial\theta}- \frac{\partial{L}_T^*(\widehat{\theta}_T)}{\partial\theta}\Big\}=o_p(1),

where $\frac{\partial L_T^*(\theta)}{\partial\theta}$ is defined in the same way as $\frac{\partial\widecheck{L}_T^*(\theta)}{\partial\theta}$ with $\widecheck{u}_{t}$ and $\widecheck{g}_{t}({\theta})$ replaced by $u_t$ and ${g}_{t}({\theta})$, respectively. In BKRW:1993, a condition similar to ((ref)) was proved for the independent data. However, their technical treatment does not work in our time series setting. This is because in the updating procedure at step three, the process $\widecheck{g}_t$ utilizes the information before and after time period $t$, so that they are not independent of $\{u_s^2\}_{s\neq t}$.

Comparison with VT estimation method

Our two-step estimation method also has a linkage to the VT estimation method in FHZ:2011, and this aspect has not been explored before. The VT method is designed for the following covariance stationary GARCH($p, q$) model

flalign\begin{split} &y_t=\sqrt{h_t}\eta_t\\ &with h_t=\tau_0\Big(1-\sum_{i=1}^{q}\alpha_{i0}-\sum_{j=1}^{p}\beta_{j0}\Big)+\sum_{i=1}^q\alpha_{i0} y_{t-i}^2+\sum_{j=1}^{p}\beta_{j0} h_{t-j}, \end{split}

where $\tau_0$ is a positive parameter, and $\alpha_{i0}$, $\beta_{j0}$ and $\eta_t$ are defined as before. Indeed, model ((ref)) is just our stationary S-GARCH model, and it is also an alternative reparametrization version of the conventional covariance stationary GARCH model. Since $Ey_t^2=\tau_0$ under model ((ref)), the VT method first estimates $\tau_0$ by $\overline{\tau}_{T}$, and then estimates $\theta_0$ by the QMLE $\overline{\theta}_{T}$, where

flalign\begin{split} \overline{\tau}_{T}=\frac{1}{T}\sum_{t=1}^{T} y_{t}^2 and \overline{\theta}_T=\arg\min_{\theta\in\Theta}\overline{L}_T(\theta) with \overline{L}_T(\theta)=\sum_{t=1}^{T}\frac{\overline{u}_t^2}{\overline{g}_t(\theta)}+\log\overline{g}_t(\theta). \end{split}

Here, $\overline{u}_{t}=y_{t}/\sqrt{\overline{\tau}_T}$, and $\overline{g}_{t}(\theta)$ is defined in the same way as $\widehat{g}_t(\theta)$ in ((ref)) with $\widehat{u}_{t}$ replaced by $\overline{u}_{t}$. Clearly, the difference of two methods is that our method estimates the unknown function $\tau(x)$ nonparametrically, while the VT method estimates the unknown constant parameter $\tau_0$ by the sample mean of $y_{t}^2$. It turns out that two methods require different technical treatments and give different application scopes. From a statistical point of view, the proof techniques for VT method rely on the facts that the objective function $\overline{L}_T(\theta)$ is differential around $\tau_0$ and the first step estimator $\overline{\tau}_T$ is $\sqrt{T}$-consistent. However, neither of these facts holds for our method, and we thus need develop new proof techniques based on more restrictive conditions for $u_t$ and $\eta_t$. From a practical point of view, our method works for the either stationary or non-stationary S-GARCH model, while the VT method does only for the stationary S-GARCH model. Hence, our method has a much broader application scope than the VT method.

By revisiting Theorem 2.1 in FHZ:2011, we further find that the asymptotic variance of $\overline{\theta}_T$ is the same as the one of $\widehat{\theta}_{T}$ in Theorem (ref). That is, our QMLE $\widehat{\theta}_{T}$ and the QMLE $\overline{\theta}_T$ in the VT method have the same asymptotic efficiency, although our first step estimator has a slower convergence rate $\sqrt{Th}$ than the parametric convergence rate $\sqrt{T}$. This novel feature has not been discovered in the literature, and it makes our two-step method more attractive than the VT method, since our QMLE does not suffer any efficiency loss for the stationary S-GARCH model, and at the same time, our QMLE can still work with the same efficiency (due to the adaptiveness feature) for the non-stationary S-GARCH model. As expected, similar features also hold for our tests $LM_{T}$ and $Q_{T}(\ell)$, and these findings will be further illustrated by simulation studies.

Extension to multivariate S-BEKK model

In this section, we extend the two-step estimation for the S-GARCH model to the multivariate semiparametric BEKK (S-BEKK) model. Let $\{\mathbf{y}_t\}_{t=1}^{T}$ be a sequence of random vectors with dimension $N\geq 1$. Assume $\mathbf{y}_t$ satisfies the following S-BEKK model

flalign\mathbf{y}_t=&\mathbb{\pmb{\tau}}_{t}^{1/2}\mathbf{u}_t with \mathbb{\pmb{\tau}}_t=\mathbb{\pmb{\tau}}(t/T),\\ \mathbf{u}_t=&\mathbf{g}_t^{1/2}\pmb{\eta}_t and \mathbf{g}_t=W_0+\sum_{i=1}^{q}A_{i0}\mathbf{u}_{t-i}\mathbf{u}_{t-i}'A_{i0}'+\sum_{j=1}^{p}B_{j0}\mathbf{g}_{t-j}B_{j0}',

for $t=1,..., T$, where $\mathbb{\pmb{\tau}}(x)\in\mathbb{R}^{N\times N}$ is a positively definite, smoothing and deterministic matrix with unknown form on the interval $[0,1]$, $\mathbf{u}_t$ is a covariance stationary BEKK$(p, q)$ process parameterized by $N\times N$ matrices $A_{i0}$, $i=1,...,q$, $B_{j0}$, $j=1,...,p$ and $W_0:=I_N-\sum_{i=1}^{q}A_{i0}A_{i0}'-\sum_{j=1}^{p}B_{j0}B_{j0}'$, and $\{\pmb{\eta}_t\}$ is a sequence of i.i.d random vectors satisfying $E\pmb{\eta}_t\pmb{\eta}_t'=I_N$. Clearly, our S-BEKK model reduces to the standard BEKK model in EK:1995 when $\mathbb{\pmb{\tau}}(x)$ is a constant matrix, and it includes the first-order S-BEKK model in HL:2010 as a special case.

Let $\mathrm{tr}(A)$ and $\mathrm{det}(A)$ be the trace and determinant of a matrix $A$, respectively, $\mathrm{vec}(A)$ be the vectorization of a matrix $A$ by stacking its columns, $A\otimes B$ be the Kronecker product between two matrices $A$ and $B$, and $A^{\otimes 2}=A\otimes A$. Denote $\pmb{\theta}=(\mathrm{vec}(A_1)',...,\mathrm{vec}(A_q)',\mathrm{vec}(B_1)',$ $...,\mathrm{vec}(B_p)')'\in\pmb{\Theta}$ be the unknown parameter of $\mathbf{u}_t$, and $\pmb{\Theta}\subset \mathbb{R}^{\mathrm{dim}(\mathbb{\pmb{\theta}})} $ be the parameter space, where $\mathrm{dim}(\mathbb{\pmb{\theta}})$ stands for the dimension of $\pmb{\theta}$. Similar to the S-GARCH model, we consider the two-step estimation for the S-BEKK model. At step one, we estimate $\mathbb{\pmb{\tau}}_{t}$ by $\widehat{\mathbb{\pmb{\tau}}}_{t}=\widehat{\mathbb{\pmb{\tau}}}(t/T)$, where

flalign*\widehat{\mathbb{\pmb{\tau}}}(x)=\frac{1}{T}{\sum_{s=1}^{T}K_h\Big(x-\frac{s}{T}\Big)\mathbf{y}_s\mathbf{y}_s'}.

At step two, we consider the QMLE of $\pmb{\theta}_0$ given by $ \widehat{\pmb{\theta}}_T=\arg\min_{\pmb{\theta}\in\pmb{\Theta}}\widehat{\mathbf{L}}_T(\pmb{\theta}), $ where $$ \widehat{\mathbf{L}}_T(\pmb{\theta})=\sum_{t=1}^{T}\widehat{\mathbf{l}}_t(\pmb{\theta})\quad\mbox{with}\quad \widehat{\mathbf{l}}_t(\pmb{\theta})= \mathrm{tr}\big(\widehat{\mathbf{g}}_t(\pmb{\theta})^{-1}\widehat{\mathbf{u}}_t\widehat{\mathbf{u}}_t'\big) +\log\det\big(\widehat{\mathbf{g}}_t(\pmb{\theta})\big). $$ Here, $\widehat{\mathbf{g}}_t(\pmb{\theta})$ is calculated recursively by $$ \widehat{\mathbf{g}}_t(\pmb{\theta})=I_N-\sum_{i=1}^{q}A_{i}A_{i}'-\sum_{j=1}^{p}B_{j}B_{j}' +\sum_{i=1}^{q}A_{i}\widehat{\mathbf{u}}_{t-i}\widehat{\mathbf{u}}_{t-i}'A_{i}'+\sum_{j=1}^{p}B_{j}\widehat{\mathbf{g}}_{t-j}(\pmb{\theta})B_{j}' $$ with $\widehat{\mathbf{u}}_t=\widehat{\mathbb{\pmb{\tau}}}_t^{-1/2}\mathbf{y}_t$, $t=1,...,T$, and some given constant initial values $\widehat{\mathbf{u}}_0=\mathbf{u}_0,..., \widehat{\mathbf{u}}_{1-q}=\mathbf{u}_{1-q}$, $\widehat{\mathbf{g}}_0(\pmb{\theta})=\mathbf{g}_0,..., \widehat{\mathbf{g}}_{1-p}(\pmb{\theta})=\mathbf{g}_{1-p}$.

To give the asymptotic distribution of $\widehat{\pmb{\theta}}_T$, we need the following notations. Let $\mathcal{A}_i=A_i^{\otimes 2}$ for $i=1,...,q$, $\mathcal{B}_j=B_j^{\otimes 2}$ for $j=1,..., p$, and $\mathbb{B}^k(1:N^2,1:N^2)$ be the upper-left $N^2\times N^2$ submatrix of $\mathbb{B}^k$, where $$ \mathbb{B}=\left(

matrix[matrix omitted — 155 chars of source]

\right). $$ Furthermore, let $\pmb{\xi}_t=\mathrm{vec}(\pmb{\eta}_t\pmb{\eta}_t'-I_N)$, $\mathbf{\Upsilon}(x)=[\mathbb{\pmb{\tau}}(x)^{-1/4}\otimes \mathbb{\pmb{\tau}}(x)^{1/4}]$, $ \mathbf{\Omega}_0=I_{N^2}-\sum_{i=1}^{q}\mathcal{A}_{i0}-\sum_{j=1}^{p}\mathcal{B}_{j0}$, $\mathbf{\Gamma}_0=I_{N^2}-\sum_{j=1}^{p}\mathcal{B}_{j0}$, $\mathbf{Q}_t=(\mathbf{Q}_{t,1}',...,\mathbf{Q}_{t,\mathrm{dim}(\mathbb{\pmb{\theta}})}')'$, $\mathbf{N}=(\mathbf{N}_1',...,\mathbf{N}_{\mathrm{dim}(\mathbb{\pmb{\theta}})}')'$ and $\mathbf{M}=(\mathbf{M}_1',...,\mathbf{M}_{\mathrm{dim}(\mathbb{\pmb{\theta}})}')'$, where

flalign*\mathbf{Q}_{t,m}=&\mathrm{vec}\Big(\frac{\partial\mathbf{g}_t}{\partial\mathbb{\pmb{\theta}}_m}\Big)'(\mathbf{g}_t^{-1/2})^{\otimes2},\quad\quad\quad \mathbf{N}_m=E\Big[\mathrm{vec}\Big(\frac{\partial\mathbf{g}_t}{\partial\mathbb{\pmb{\theta}}_m}\Big)'(\mathbf{g}_t^{-1}\otimes I_N)\Big],\\ \mathbf{M}_m=&E\Big[\mathrm{vec}\Big(\frac{\partial {\mathbf{g}}_t}{\partial\mathbb{\pmb{\theta}}_m}\Big)'(\mathbf{g}_t^{-1})^{\otimes2}\mathbf{T}_t\Big],\quad\mathbf{T}_t=\sum_{k=0}^{\infty}\mathbb{B}_0^k(1:N^2,1:N^2)\Big(\sum_{i=1}^{q}\mathcal{A}_{i0} \Big)[I_N\otimes\mathbf{u}_{t-k-i}\mathbf{u}_{t-k-i}'].

The following theorem establishes the asymptotic normality of $\widehat{\pmb{\theta}}_T$.

thmSuppose Assumptions C.1--C.7 in JLZ:2019 hold. Then, $$ \sqrt{T}(\widehat{\mathbb{\pmb{\theta}}}_T-\mathbb{\pmb{\theta}}_0)\rightarrow_{\mathcal{L}} N(0,\mathbf{\Sigma})\mbox{ as }T\to\infty, $$ where $\mathbf{\Sigma}=\big(E[\mathbf{Q}_t\mathbf{Q}_t']\big)^{-1}\big[\mathbf{J}_1+\mathbf{J}_2+\mathbf{J}_3+\mathbf{J}_3'\big]\big(E[\mathbf{Q}_t\mathbf{Q}_t']\big)^{-1}$ with \begin{flalign*} \mathbf{J}_1=&E\Big[\mathbf{Q}_t\mathrm{Var}(\pmb{\xi}_t)\mathbf{Q}_t'\Big],\\ \mathbf{J}_2=&[\mathbf{M}-\mathbf{N}]\Big\{\int_{0}^{1}\mathbf{\Upsilon}(x)\mathbf{\Omega}_0^{-1}\mathbf{\Gamma}_0 E\Big[(\mathbf{g}_t^{1/2})^{\otimes2}\mathrm{Var}(\pmb{\xi}_t)(\mathbf{g}_t^{1/2})^{\otimes2}\Big]\mathbf{\Gamma}_0'\mathbf{\Omega}_0^{'-1} \mathbf{\Upsilon}(x)dx\Big\}[\mathbf{M}-\mathbf{N}]',\\ \mathbf{J}_3=&E\Big[\mathbf{Q}_t\mathrm{Var}(\pmb{\xi}_t)(\mathbf{g}_t^{1/2})^{\otimes2}\Big] \Big\{\mathbf{\Gamma}_0'\mathbf{\Omega}_0^{'-1}\int_{0}^{1}\mathbf{\Upsilon}(x)dx\Big\}[\mathbf{M}-\mathbf{N}]'. \end{flalign*}

When $p=q=1$, it can be shown that our asymptotic variance-covariance matrix $\mathbf{\Sigma}$ is equivalent to the one obtained in HL:2010, but with a relatively simpler expression. Moreover, Theorem (ref) indicates that the effect of nonparametric part $\mathbb{\pmb{\tau}}_t$ on $\mathbf{\Sigma}$ is reflected by the term $\mathbf{\Upsilon}(x)$ existing in $\mathbf{J}_2$ and $\mathbf{J}_3$. When $N=1$, we have $\mathbf{\Upsilon}(x)\equiv1$, and hence $\widehat{\mathbb{\pmb{\theta}}}_T$ has the adaptiveness feature as demonstrated before. When $N>1$, the form of $\mathbb{\pmb{\tau}}(x)$ has an impact on $\mathbf{\Upsilon}(x)$, except for some special cases (e.g., $\mathbb{\pmb{\tau}}(x)=\tau(x) I_N$ with $\tau(x)>0$). Therefore, $\widehat{\mathbb{\pmb{\theta}}}_T$ does not have the adaptiveness feature in the multivariate case.

Simulations

This section gives the simulation studies for the QMLE $\widehat{\theta}_{T}$ and the tests $LM_{T}$ and $Q_{T}(\ell)$. To facilitate it, we first show how to choose the bandwidth $h$.

Choice of bandwidth

The practical implementation of our entire methodologies needs to choose the bandwidth $h$. The methods in terms of mean squared error criterion (see, e.g., HL:2010) usually yield a bandwidth of order $T^{-1/5}$, which does not satisfy Assumption (ref). In what follows, we give a two-step cross-validation (CV) procedure to choose $h$ such that Assumption (ref) is satisfied.

alg{(CV bandwidth selection procedure)} \begin{enumerate} • Set a pilot bandwidth $h_{0}=T^{-\lambda_0}$ with $\lambda_0\in(1/4,1/2)$, and then obtain the pilot estimates $\widehat{\tau}_{t,0}$ and $\widehat{u}_{t,0}$. Choose a pilot GARCH (or ARCH) model for the process $u_t$, and based on $\{\widehat{u}_{t,0}\}_{t=1}^{T}$, estimate this pilot model by the QMLE to get the pilot estimates $\{\widehat{g}_{t,0}\}_{t=1}^{T}$. • With $\{\widehat{g}_{t,0}\}_{t=1}^{T}$, define a CV criterion as $$CV(h)=\sum_{t=1}^{T}\Big\{\frac{y_t^2}{{\widehat{\tau}_{-t}(h)\widehat{g}_{t,0}}}-1\Big\}^2,$$ where $\widehat{\tau}_{-t}(h)$ is a leave-one-out estimate of $\tau_t$ with respect to the bandwidth $h$, based on all observations except for $y_t$. Select our bandwidth as $h_{cv}=\arg\min_{h\in\mathcal{H}}CV(h)$, where $\mathcal{H}=[c_{\min}T^{-\lambda_0},c_{\max}T^{-\lambda_0}]$ with two positive constants $c_{\min}$ and $c_{\max}$. \end{enumerate}

Let $\widehat{\mathrm{Var}}(y_t)$ be the sample variance of $\{y_{t}\}_{t=1}^{T}$. To compute $h_{cv}$ in Algorithm (ref), we suggest to choose $\lambda_0=2/7$, $c_{\min}=0.5\widehat{\mathrm{Var}}(y_t)^{\lambda_0}$ and $c_{\max}=3\widehat{\mathrm{Var}}(y_t)^{\lambda_0}$, which will be used and demonstrated with good performance in our simulation studies below. For the pilot model in Algorithm (ref), it could be taken based on either some prior information or the Bayesian information criterion (BIC).

Simulations for the estimation

In this subsection, we examine the finite-sample performance of the QMLE $\widehat{\theta}_{T}$. We generate 1000 replications of sample size $T=2000$ and $4000$ from the following two data generating processes (DGPs)

flalign*&DGP 1 : \,\,The S-ARCH(2) model with \alpha_{10}=\alpha_{20}=0.3; \\ &DGP 2 : \,\,The S-GARCH(1, 1) model with \alpha_{10}=0.1 and \beta_{10}=0.8,

where the function $\tau(x)$ is designed as follows

flalign[No change]&\quad \tau(x)=1;\\ [Linear change]&\quad \tau(x)=1+2x;\\ [Cyclical change]&\quad \tau(x)=1+\sin(4\pi x)/2,

and the error $\eta_t$ follows $N(0,1)$, $st_{10}$, and $st_{5}$. Here, $st_{\nu}$ is the standardized Student-$t$ distribution with unit variance.

For each replication, we compute $\widehat{\theta}_{T}$ by using the Epanechnikov kernel $K(x)=\frac{3}{4}(1-x^2)\mathbf{1}(|x|\leq 1)$ and choosing the bandwidth $h=h_{cv}$ according to Algorithm (ref) with the (G)ARCH model in DGP as the pilot model. Table (ref) reports the sample bias, sample empirical standard deviation (ESD) and average asymptotic standard deviation (ASD) of $\widehat{\theta}_T$ based on 1000 replications for each DGP, where the ASD is calculated as in Remark (ref). From Table (ref), we find that (i) the biases of $\widehat{\theta}_{T}$ are small in each case; (ii) regardless of the specification of $\tau(x)$ and the distribution of $\eta_t$, the values of ESD and ASD are close to each other, especially for large $T$; (iii) when the value of $T$ increases, the value of ESD decreases; (iv) $\widehat{\theta}_{T}$ becomes less efficient with a larger value of ESD as the thickness of $\eta_t$ becomes heavier; (v) the value of ESD is almost invariant with respect to the specification of $\tau(x)$, meaning that $\widehat{\theta}_{T}$ is adaptive as expected. Under the same settings as in Table (ref), we also examine the finite-sample performance of the standard QMLE in Bollerslev:1986, and find that when $\tau(x)\sim (\ref{nochange})$, the standard QMLE is more efficient than $\widehat{\theta}_{T}$; but when $\tau(x)\sim (\ref{linearchange})$ or $(\ref{cyclicalchange})$, the standard QMLE suffers from larger bias and discrepancy between ESD and ASD. For saving the space, these results for the standard QMLE are not reported here. Overall, our QMLE $\widehat{\theta}_{T}$ has a satisfactory performance in all considered cases, and the standard QMLE should not be used for the non-stationary S-GARCH model.

table[table omitted — 6,859 chars of source]

Simulations for the testing

In this subsection, we examine the finite-sample performance of $LM_{T}$ and $Q_{T}(\ell)$. We generate 1000 replications of sample size $T=2000$ and $4000$ from the following two DGPs

flalign*&DGP 3 : \,\,The S-GARCH(1, 2) model with \alpha_{10}=\beta_{10}=0.3 and \alpha_{20}=0.03k;\\ &DGP 4 : \,\,The S-GARCH(2, 1) model with \alpha_{10}=\beta_{10}=0.3 and \beta_{20}=0.03k,

where $k=0,1,..., 10$, $\tau(x)$ is designed as in DGPs 1--2, and $\eta_t\sim N(0, 1)$. For each DGP, the model with respect to $k=0$ is taken as its null model. That is, the S-GARCH($1, 1$) model is the null model for both DGP 3 and DGP 4.

Next, we fit each replication by its related null model, and then apply $LM_{T}$ to detect the null hypothesis of $k=0$ as well as $Q_{T}(\ell)$ to check whether this fitted null model is adequate. Based on 1000 replications, the empirical power of $LM_{T}$ and $Q_{T}(\ell)$ is plotted in Fig\,(ref) and Fig\,(ref) for DGP 3 and DGP 4, respectively, where we take the level $\alpha=5\%$ and the lag $\ell=6, 9$, and $12$, and the sizes of both tests correspond to the results for $k=0$.

figure[figure omitted — 431 chars of source]
figure[figure omitted — 195 chars of source]

From Figs\,(ref)--(ref), we can find that (i) all tests have precise sizes; (ii) the power of all tests becomes large as the value of $T$ or $k$ increases; (iii) $LM_{T}$ is more powerful than all $Q_{\ell}$, and $Q_6$ is generally more powerful than $Q_9$ and $Q_{12}$; (iv) all tests are more powerful to detect the mis-specification of ARCH part in DGP 3 than the mis-specification of GARCH part in DGP 4; (v) all tests are adaptive, since their power is unaffected by the form of $\tau(x)$. In summary, all tests have a good performance especially for large $T$.

Comparison with three-step estimation method

In this subsection, we compare the finite-sample performance of $\widehat{\theta}_T$ and the three-step estimator $\widecheck{\theta}_{T}$ by investigating their bias difference and efficiency ratio (componentwisely) defined respectively as $$ d(\gamma)=(|\mbox{the Bias of }\widehat{\gamma}_T|-|\mbox{the Bias of }\widecheck{\gamma}_T|)\times 100 \mbox{ and } R(\gamma)=\frac{\mbox{the ESD of }\widehat{\gamma}_{T}}{\mbox{the ESD of }\widecheck{\gamma}_{T}}, $$ where $\gamma$ denotes any entry of $\theta_0$, and the Bias and ESD of each estimator are computed based on 1000 replications. We calculate the values of $d(\gamma)$ and $R(\gamma)$ under the same simulation settings as in Subsection (ref), and only report the results for the case of $\tau(x)\sim$ ((ref)) in Table (ref) due to the adaptiveness of $\widehat{\theta}_T$ and $\widecheck{\theta}_{T}$. From Table (ref), we can find that (i) both estimators have a comparable bias performance; (ii) when $\eta_t\sim N(0, 1)$, $\widecheck{\theta}_T$ is more (or less) efficient than $\widehat{\theta}_T$ in DGP 1 (or DGP 2), indicating that $\widecheck{\theta}_T$ does not achieve the semiparametric efficiency bound as indicated in Theorem (ref); (iii) when $\eta_t$ has a heavier distribution (e.g., $\eta_t \sim st_{5}$), $\widehat{\theta}_T$ exhibits more efficiency advantage over $\widecheck{\theta}_{T}$. In summary, our simulation results suggest that it is unnecessary to further update $\widehat{\theta}_T$ to $\widecheck{\theta}_T$.

table[table omitted — 2,281 chars of source]

Comparison with the VT method

In this subsection, we compare the finite-sample performance of $\widehat{\theta}_{T}$, $LM_T$ and $Q_T(\ell)$ with those of $\overline{\theta}_{T}$, $LM_{T}^{vt}$ and $Q_T^{vt}(\ell)$, respectively, where $\overline{\theta}_{T}$ defined in ((ref)) is the QMLE from the VT method, and $LM_{T}^{vt}$ and $Q_T^{vt}(\ell)$ are defined in the same way as $LM_T$ and $Q_T(\ell)$ with $\widehat{\theta}_{T}$ replaced by $\overline{\theta}_{T}$. Note that when the S-GARCH($p, q$) model is stationary, $\overline{\theta}_{T}$ is asymptotically normal, and $LM_{T}^{vt}$ and $Q_T^{vt}(\ell)$ have the same limiting null distributions as those of $LM_T$ and $Q_T(\ell)$.

First, we compare the efficiency of $\widehat{\theta}_{T}$ and $\overline{\theta}_{T}$ by looking at the following ratio $$R_{qmle}(\gamma)=\frac{\mbox{the ESD of }\widehat{\gamma}_{T}}{\mbox{the ESD of }\overline{\gamma}_{T}},$$ where the ESD of each estimator is computed based on 1000 replications. Table (ref) reports the values of $R_{qmle}(\gamma)$ when the DGP is a stationary S-ARCH(2) (or S-GARCH($1, 1$)) model with $\tau(x)\sim (\ref{nochange})$, $\eta_t\sim N(0, 1)$, and three different choices of $\theta_0$. From this table, we find that as expected, all the values of $R_{qmle}(\gamma)$ are close to 1, indicating that $\widehat{\theta}_{T}$ and $\overline{\theta}_{T}$ have the same asymptotic efficiency when the S-GARCH model is stationary.

table[table omitted — 2,696 chars of source]

Second, we compare the power of $LM_{T}$ and $LM_{T}^{vt}$ and that of $Q_{T}(\ell)$ and $Q_{T}^{vt}(\ell)$ by looking at the following two ratios $$R_{lm}=\frac{\mbox{the power of }LM_{T}}{\mbox{the power of }LM_{T}^{vt}}\quad\mbox{ and }\quad R_{q}(\ell)=\frac{\mbox{the power of }Q_{T}(\ell)}{\mbox{the power of }Q_{T}^{vt}(\ell)},$$ where the power of each test is computed based on 1000 replications. Table (ref) reports the values of $R_{lm}$ and $R_{q}(\ell)$ (for $\ell=6, 9$, and $12$), when the data are generated from a stationary S-GARCH($1, 2$) model in DGP 3 with $\tau(x)\sim (\ref{nochange})$. The results for DGP 4 are quite similar and hence omitted to save space. From Table (ref), we can see that (i) the values of $R_{lm}$ are close to 1 in all examined cases; (ii) when the value of $T$ or $k$ is small, the values of $R_{q}(\ell)$ are slightly less than one, meaning that $Q_T^{vt}(\ell)$ could be more powerful than $Q_T(\ell)$; (ii) when the value of $T$ or $k$ becomes large, the power advantage of $Q_T^{vt}(\ell)$ disappears as the values of $R_{q}(\ell)$ are close to 1. These findings demonstrate that when the S-GARCH model is stationary, our two tests have the same power performance as their counterparts from the VT method especially for large $T$. We also highlight that when the S-GARCH model is non-stationary, our unreported results show that $LM^{vt}_{T}$ and $Q_T^{vt}(\ell)$ can cause a severe over-sized problem, and hence they can not be used in this case.

table[table omitted — 1,594 chars of source]

Applications

In this section, we re-study the US dollar to Indian rupee (USD/INR) exchange rate series and FTSE-index series in Truquet:2017, with respect to in-sample fitting and out-of-sample prediction.

USD/INR exchange rate

This subsection considers the USD/INR exchange rate series from December 19th, 2005 to February 18th, 2015. The log returns (in percentage) of this series having $T=2301$ observations in total are denoted by $\{y_t\}$, and they are plotted in the upper panel of Fig\,(ref). We apply the non-parametric strict stationarity test in Hong:2017 (with the same settings as in their simulation) to $\{y_t\}$ and find this test statistic has a p-value close to zero, indicating a strong evidence against the strict stationarity. Thus, using a non-stationary model to fit this series seems appropriate. In Truquet:2017, this return series is fitted by a semiparametric ARCH(1) model with a time-varying intercept and a constant lag-1 ARCH parameter. Motivated by this, we use an ARCH(1) model as the pilot model in Algorithm (ref) to choose the bandwidth $h=0.0358$, and then calculate the series $\{\widehat{u}_{t}\}$. Based on $\{\widehat{u}_{t}\}$, the BIC selects $p=q=1$ for the S-GARCH model, and hence we fit this return series by the S-GARCH($1, 1$) model with $\widehat{\alpha}_{1T}=0.0762_{(0.0231)}$, $\widehat{\beta}_{1T}=0.8443_{(0.0475)}$, and $\widehat{\tau}_t$ being plotted in the middle panel of Fig\,(ref), where the values in parentheses are the asymptotic standard errors, and the bandwidth $h=0.0833$ is re-chosen by using a GARCH($1, 1$) model as the pilot model in Algorithm (ref). For this fitted S-GARCH($1, 1$) model, the p-values of the portmanteau tests $Q_{T}(6)$, $Q_{T}(9)$, and $Q_{T}(12)$ are 0.6472, 0.7530, and 0.8268, respectively, implying that our fitted short run GARCH($1, 1$) component is adequate. In view of the plot of $\{\widehat{\tau}_t\}$ in Fig\,(ref), we can find that the long run component $\tau_t$ has relatively larger values around years 2009 and 2014. Moreover, we also plot the estimated volatilities based on either S-GARCH or GARCH model in the bottom panel of Fig\,(ref), from which we can see that compared with the S-GARCH model, the GARCH model tends to underestimate the volatilities during 2008-2009 and 2013-2014, and overestimate the volatilities during other periods.

figure[figure omitted — 652 chars of source]

FTSE-index

This subsection considers the FTSE-index series from January 4th, 2005 to March 4th, 2015. We study the log returns of this index series with $T=2568$ observations in total, which is denoted by $\{y_t\}$ and plotted in the upper panel of Fig\,(ref). As the previous example, we use the non-parametric strict stationarity test in Hong:2017 to $\{y_t\}$, and find a strong evidence (with the p-value close to zero) against the strict stationarity. Since Truquet:2017 suggested a semiparametric ARCH(5) model with a time-varying intercept and constant ARCH parameters to fit this return series, we take an ARCH(5) model as the pilot model in Algorithm (ref), and then select the bandwidth $h=0.0865$ as a result. Based on this choice of $h$, we compute $\{\widehat{u}_{t}\}$ and select $p=q=1$ according to the BIC. Hence, we fit this return series by the S-GARCH($1, 1$) model with $\widehat{\alpha}_{1T}=0.1098_{(0.0165)}$, $\widehat{\beta}_{1T}=0.8433_{(0.0233)}$, and $\widehat{\tau}_t$ being plotted in the middle panel of Fig\,(ref), where the bandwidth $h=0.0941$ is re-chosen by using a S-GARCH($1, 1$) model as the pilot model in Algorithm (ref). Further, the portmanteau tests $Q_{T}(6)$, $Q_{T}(9)$, and $Q_{T}(12)$ (with p-values equal to $0.5326$, $0.5335$, and $0.2800$, respectively) suggest that this fitted short run GARCH($1, 1$) component is adequate. From the middle panel of Fig\,(ref), we find that the long run component $\tau_t$ for the FTSE return series only has a clear peak around 2009. This may imply that the stock market index series has a different long run structure with the exchange rate series. Moreover, we also plot the estimated volatilities based on either S-GARCH or GARCH model in the bottom panel of Fig\,(ref), from which we can see that the estimated volatilities from two models are quite close except around years 2008-2009, during which the GARCH model tends to underestimate the volatilities.

figure[figure omitted — 299 chars of source]

Forecasting comparisons

This subsection makes a forecasting comparison among S-GARCH($1, 1$) model, S-ARCH($q$) model, GARCH($1, 1$) model in Bollerslev:1986, and LS-ARCH($q$) model (i.e., the locally stationary ARCH($q$) model) in FSR:2008 for the USD/INR and FTSE return series. Note that the S-ARCH($q$) model can locally approximate the semiparametric ARCH($q$) model in Truquet:2017, where $q=1$ (or 5) is suggested for the USD/INR (or FTSE) return series. Hence, we follow Truquet:2017 to select $q$ for the S-ARCH($q$) and LS-ARCH($q$) models.

Next, we compare all four models in terms of the averaged QLIKE loss function in Patton:2011. Specifically, we use the in-sample data $\{y_t\}_{t=1}^{T_0}$ to make a $t_0$-step ahead forecast $\widehat{y}_{{T_0}+t_0|T_0}^2$ for the out-of-sample data point $y_{T_0+t_0}^2$, and then compute the averaged QLIKE by $$\mbox{QLIKE}(t_0)=\frac{1}{T-t_0-1499}\sum_{T_0=1500}^{T-t_0}\log\widehat{y}_{{T_0}+t_0|T_0}^2+\frac{y_{T_0+t_0}^2}{\widehat{y}_{{T_0}+t_0|T_0}^2}.$$ The model with the smaller value of $\mbox{QLIKE}(t_0)$ has the better $t_0$-step ahead forecasting performance.

Moreover, we introduce how each model computes $\widehat{y}_{{T_0}+t_0|T_0}^2$. For the S-GARCH($1, 1$) model, we fit the model via the two-step estimation based on the in-sample data $\{y_t\}_{t=1}^{T_0}$, where the bandwidth $h$ is chosen by Algorithm (ref) with a pilot GARCH(1, 1) model. With the kernel estimate $\widehat{\tau}_{T_0}$ and QMLE $\widehat{\theta}_{T_0}$, we then obtain $\widehat{y}_{T_0+t_0|T_0}^2=\widehat{\tau}_{T_0}g_{T_0+t_0|T_0}(\widehat{\theta}_{T_0})$, where $g_{T_0+t_0|T_0}(\widehat{\theta}_{T_0})$ computed as for volatility prediction in the GARCH($1, 1$) model is the $t_0$-step ahead prediction of $g_{T_0+t_0}$. A similar way is used for the S-ARCH($q$) model to compute $\widehat{y}_{{T_0}+t_0|T_0}^2$. For the GARCH($1, 1$) model, we fit the model via the VT estimation based on the in-sample data $\{y_t\}_{t=1}^{T_0}$, and then compute $\widehat{y}_{{T_0}+t_0|T_0}^2$ in the conventional way. For the LS-ARCH($q$) model, we follow the method in FSR:2008 to compute $\widehat{y}_{{T_0}+t_0|T_0}^2$. That is, we treat the last $\widetilde{T}$ in-sample data points as if they came from a stationary ARCH($q$) process, and then estimate the parameters based on these $\widetilde{T}$ data points and compute $\widehat{y}_{{T_0}+t_0|T_0}^2$ as for the stationary ARCH($q$) model. Here, the tuning parameter $\widetilde{T}$ is chosen by minimizing the QLIKE, i.e.,

equation*[equation* omitted — 162 chars of source]

where $\mathcal{T}=\{50, 100, ..., 500\}$, and $\widehat{y}_{t+1|t}^2(T)$ computed as for the stationary ARCH($q$) model is the prediction of $y_{t+1}^2$ based on the data $\{y_{i}\}_{i=t-T+1}^{t}$.

Table (ref) reports the values of QLIKE($t_0$) for all four models, where the prediction horizon $t_0$ is taken as $1, 5, 10$, and $22$, corresponding to daily, weekly, biweekly, and monthly predictions, respectively. The DM test in DM1995 is implemented to compare the forecasting accuracy between the model with smallest value of QLIKE and other three models. From this table, we find that for both series, the S-GARCH model has the smallest value of QLIKE for $t_0=1$ and 5, while the GARCH model has the smallest value of QLIKE for $t_0=10$ and 22. In terms of DM test, we find that the model with smallest value of QLIKE does not exhibit significantly forecasting accuracy than its three competitors for USD/INR series, while it has significantly forecasting accuracy than two ARCH-type competitors for FTSE series. These findings are consistent with those in FSR:2008 and Truquet:2017 that (S-)GARCH models could deliver better forecasts than LS-ARCH and S-ARCH models. For the S-GARCH model, we simply just use the latest in-sample long-run component estimator $\widehat{\tau}_{T_0}$ to predict the out-of-sample long-run component $\tau_{T_0+t_0}$. So far, we do not know how to find an “optimal” way under certain criterion to predict $\tau_{T_0+t_0}$, and this dilemma seems to exist in most of nonparametric methods. We believe that with a better prediction of $\tau_{T_0+t_0}$, our S-GARCH model could deliver better prediction performances especially at longer prediction horizons.

table[table omitted — 1,537 chars of source]

Concluding remarks

This paper provides a complete statistical inference procedure for the S-GARCH model. Our methodologies including the estimation and testing focus on the QMLE of non-time-varying parameters in GARCH-type short run component. Since this QMLE is based on the estimate of the long run component, we develop new proof techniques to derive its asymptotic normality, and find that its asymptotic variance is adaptive to the long run component with unknown form. By comparing the results with those in HL:2010, we find a much simpler asymptotic variance expression for the QMLE, bringing the convenience of use to practitioners. By comparing with the QMLE from the VT method in FHZ:2011, we find that our QMLE not only enjoys a broader application scope to deal with the non-stationary S-GARCH model, but also avoids any efficiency loss when the S-GARCH model is stationary. All of these interesting features have not been unveiled before in the literature, and they make our QMLE and its related Lagrange multiplier and portmanteau tests more appealing in practice.

Finally, we suggest some future research topics. First, it is interesting to extend our study to the robust estimation context. This could give us more efficient estimators and more powerful tests for dealing with heavy-tailed data. Second, a similar semiparametric framework as ((ref)) can be posed into many variants of the standard GARCH model (e.g., the asymmetric power-GARCH model in PWT:2008 and the asymmetric log-GARCH model in FWZ:2013), and our methodologies could be applied to these new resulting semiparametric models. Third, another possible work is to relax the smooth condition of the long run component to allow for abrupt changes. This seems challenging and may require more non-trivial technical treatments.

Acknowledgments

The authors greatly appreciate the helpful comments and suggestions of two anonymous referees, Associate Editor, and Co-Editor. Jiang acknowledges that his work was partly carried out during the visit in University of Hong Kong and University of Illinois at Urbana-Champaign, and his work is supported by China Scholarship Council (No. 201906210093). Li's work is supported by the NSFC (Nos. 11771239 and 71973077) and the Tsinghua University Initiative Scientific Research Program (No. 2019Z07L01009). Zhu's work is supported by Hong Kong GRF grant (Nos. 17306818 and 17305619), NSFC (Nos. 11690014 and 11731015), Seed Fund for Basic Research (No. 201811159049), and Fundamental Research Funds for the Central University (19JNYH08).