EconBase
← Back to paper

A new GARCH model with a deterministic time-varying intercept

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

68,677 characters · 12 sections · 82 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

A new GARCH model with a deterministic time-varying intercept

abstractIt is common for long financial time series to exhibit gradual change in the unconditional volatility. We propose a new model that captures this type of nonstationarity in a parsimonious way. The model augments the volatility equation of a standard GARCH model by a deterministic time-varying intercept. It captures structural change that slowly affects the amplitude of a time series while keeping the short-run dynamics constant. We parameterize the intercept as a linear combination of logistic transition functions. We show that the model can be derived from a multiplicative decomposition of volatility and preserves the financial motivation of variance decomposition. We use the theory of locally stationary processes to show that the quasi maximum likelihood estimator (QMLE) of the parameters of the model is consistent and asymptotically normally distributed. We examine the quality of the asymptotic approximation in a small simulation study. An empirical application to Oracle Corporation stock returns demonstrates the usefulness of the model. We find that the persistence implied by the GARCH parameter estimates is reduced by including a time-varying intercept in the volatility equation. \\ \\

\thispagestyle{empty}

Introduction

\setcounter{page}{1} It is common for long financial time series to exhibit gradual change in the unconditional volatility. Failure to model such change can lead to an overstatement of the persistence of volatility (mikoschstarica2004a, mikoschstarica2004b, staricagranger2005). In this paper, we propose a new model that captures this type of nonstationarity in a parsimonious way. The most popular volatility models are the autoregressive conditional heteroskedasticity (ARCH) model by Engle82 and its generalization, the generalized ARCH (GARCH) model by Bollerslev86 and Taylor86. A major reason for their popularity is that they are easy to fit and likelihood-based inference can be done. This paper extends the GARCH model to accommodate nonstationarity in a way that preserves the parametric likelihood framework. The model, called the additive time-varying (ATV-)GARCH model, augments the volatility equation by a deterministic time-varying intercept. The model combines the flexibility of the time-varying GARCH (tvGARCH) class of models (see ChenH16) with the financial motivation of multiplicative decomposition of volatility (see Engle&Rangel08).

In the nonparametric literature several papers consider time-varying GARCH models. DahlhausSR06 generalized the ARCH model to a nonstationary ARCH model with time-varying coefficients. The time-varying ARCH (tvARCH) process can be locally approximated by a stationary ARCH process. Such processes are referred to as locally stationary. The locally stationary tvARCH process was further extended to tvGARCH; see rao2006, Rohan13, ChenH16, kristensen2019local and KarmakarRichterWu. Truquet17 introduced a semiparametric ARCH model where some, but not all, parameters are time-varying. In statistical tests for detecting non-time-varying parameters in models fitted on real data, he finds that the null hypothesis tends to be rejected for the constant but not the ARCH parameters.

We show that the entirely parametric ATV-GARCH model admits a representation as a time-varying GARCH process. Many asymptotic results are available for time-varying GARCH processes. dahlhaus2019towards, henceforth DRW, developed a general theory for nonlinear locally stationary processes. In this class of models the parameters are Lipschitz continuous parameter curves. DRW proved a global law of large numbers and a global central limit theorem based on the stationary approximation. We apply these results to the ATV-GARCH model.

Time-varying GARCH models have not been much used to model financial data. Instead, multiplicative decomposition of volatility into a stationary GARCH component and a financially motivated slowly moving component has received a lot of attention in the literature. The first papers examining the multiplicative decomposition were Feng04 and BellegemS04, followed by Engle&Rangel08, fengmcneil2008, BrownleesG10, MishraSU10, MazurPipien2012, AmadoT13 and lintonkoo2015. For surveys, see VBellegem12 and Terasvirta12. For a more recent model with a multiplicative structure, see ES2018. We show that a particular choice of the multiplicative decomposition yields a model that can be asymptotically expressed as an ATV-GARCH model. The ATV-GARCH model preserves the financial motivation of variance decomposition while the framework of tvGARCH processes can be utilized to derive asymptotic results. The ATV specification can be viewed as a parsimoniously parameterized alternative to the very general tvGARCH by rao2006.

In the ATV-GARCH model, volatility is mean-reverting towards a time-varying mean, but with constant persistence. This way, the model captures structural change that slowly affects the amplitude of a time series while keeping the short-run dynamics constant over time. The model is particularly well suited for situations in which the volatility of an asset or index is smoothly increasing or decreasing over time. The idea is that short-run fluctuations in volatility are stationary, but long-run structural changes make the assumption of stationarity inappropriate. This view is complementary to the interpretation by Diebold86 that a neglected level shift in the intercept of a GARCH model can overstate persistence and erroneously leads one to an integrated GARCH by EngleB86. LamoureuxL90 investigate Diebold’s conjecture in finite samples. hillebrand2005 proves that the sum of the estimated autoregressive parameters of the conditional variance converges to one if a time series contains parameter changes in the conditional volatility. See also mikoschstarica2004a, who argue that assuming a constant unconditional variance can lead to spurious Integrated GARCH (IGARCH) models. A similar comment applies to the fractionally integrated GARCH (FIGARCH) process introduced by baillie1996. The FIGARCH model features a parameter that measures the degree of fractional integration of the volatility process. Neglecting structural breaks can cause overestimation of the parameter and in consequence an overestimation of persistence (see hillebrand2005 and the references therein). It is therefore possible that allowing for a model that incorporates changes in the intercept can make the evidence of long memory disappear. By making the intercept a smooth function of time, it is possible to capture level shifts that occur gradually, rather than abruptly as in volatility models with structural breaks; see e.g. chu1995 and smith2008.

The ATV-GARCH model is globally nonstationary but locally stationary. The parameters of the model are estimated globally by QMLE. The primary goal of this paper is to establish consistency and asymptotic normality of the QMLE of the parameters of the ATV-GARCH model based on the stationary approximation. More precisely, we make use of the theory of locally stationary processes in DRW. BHK2003, henceforth BHK, and FrancqZakoian2004 proved consistency and asymptotic normality of the QMLE of the parameters of strictly stationary GARCH processes under mild conditions. For earlier developments, see leehansen1994 and lumsdaine1996. We use the theory in DRW to extend the results to a nonstationary setting. To the best of our knowledge, this paper is the first one to give rigorous proofs of consistency and asymptotic normality in a parametric GARCH model with a deterministic time-varying intercept.

Our model is related to the GARCH-X model of HanPark2012 and HanK14, the FIGARCH model with a time-varying intercept of baillie2009 and the smoothly time-varying parameter GARCH model of ChenGL14, among others. In the GARCH-X model there is an additive component in the variance equation, but instead of being smooth and deterministic it is a positive-valued function of an exogenous stochastic variable. It may be viewed as a special case of the functional coefficient GARCH model by MedeirosW09. baillie2009 proposed a FIGARCH model with a time-varying intercept given by a trigonometric function. However, the authors do not provide results on the consistency and asymptotic normality of the QMLE of the parameters. ChenGL14 proposed a model in which all parameters are time-varying in a way similar to the ATV-GARCH model, but estimation and inference are Bayesian. ambrovzevivciute2008tvgarch study the statistical properties of the tvGARCH$(1,1)$ model with logistic coefficients, but do not consider inference.

The plan of the paper is as follows. In Section 2 we introduce the ATV-GARCH model and show that it can be locally approximated by a stationary GARCH process. Maximum likelihood estimation of the model is discussed in Section 3. We prove consistency and asymptotic normality of the QMLE of the parameters based on the stationary approximation. In Section 4 we provide a simulation study of the QMLE of the parameters. In Section 5 we demonstrate the usefulness of the model by an empirical application to Oracle Corporation stock returns. Section 6 concludes. Proofs can be found in the Appendix.

Throughout the paper, we use $\left\lVertX\right\rVert_p=(\mathbb{E}|X|^p)^{1/p}$ to denote the norm of a random variable $X$, and when applied to a vector or matrix, we use $|\cdot|$ to denote the maximum norm. We use $C_1,C_2,\ldots$ as generic constants, not necessarily the same across contexts.

The model

In this section, we define the ATV-GARCH model and show that a particular choice of multiplicative decomposition yields a model that is asymptotically equivalent to an ATV-GARCH model. We show that the ATV-GARCH model belongs to the class of locally stationary time-varying GARCH processes. We discuss some technical results related to the moment assumptions required in the proofs of consistency and asymptotic normality of the QMLE of the parameters.

The additive time-varying GARCH model

The ATV-GARCH model is defined by augmenting the volatility equation of the GARCH model of Bollerslev (1986) and Taylor86 by a deterministic time-varying intercept:

equation[equation omitted — 88 chars of source]

where $\varepsilon_{t}$ is IID$(0,1)$, while the volatility equation is given by

equation[equation omitted — 136 chars of source]

In applications, $X_{t,T}$ will typically be a log-return of a financial asset or an index. To keep the notation simple, the volatility equation ((ref)) is of order one. The time-varying intercept is defined as $\alpha_0(t/T;\boldsymbol{\theta}):= \alpha_0 + g(t/T, \boldsymbol{\theta})$, where $g(t/T, \boldsymbol{\theta})$ is a Lipschitz continuous function in the parameters $\boldsymbol{\theta}$; the details will follow. The double subscript $(t,T)$ is used to emphasize that we are working with triangular arrays in rescaled time.

Interestingly, the ATV-GARCH model can be derived from a multiplicative decomposition of volatility. Such a decomposition typically contains two parts: a slowly changing component that captures some feature of the data-generation process, and a transient component that generates the daily volatility. A general multiplicative decomposition can be stated by considering ((ref)) with

align[align omitted — 158 chars of source]

where $g_{t,T}$ is some function capturing the slowly changing component. The function $g_{t,T}$ has been motivated by financial arguments. This is the case in e.g. ES2018, where the authors posit that the slowly changing component stems from a flexible leverage multiplier. An alternative specification is to make $g_{t,T}$ a function of rescaled time $t/T$, $g_{t,T}=g(t/T)$, as in AmadoT13. Here, we take the latter route and further require $g_{t,T}$ to be Lipschitz continuous. To derive the ATV-GARCH model, rewrite ((ref)) as follows: $$ \sigma^2_{t,T} =\alpha_0g_{t,T} + \alpha_1\frac{X^2_{t-1,T}g_{t,T}}{g_{t-1,T}} + \beta_1\frac{\sigma^2_{t-1,T}g_{t,T}}{g_{t-1,T}}.$$ Note that by Lipschitz continuity, we have $$\left|g(t/T)-g((t-1)/T)\right| \leq C\left|t/T-(t-1)/T\right|=C/T,$$ for some positive constant $C,$ so as $T\to\infty,$ $g_{t,T}/g_{t-1,T}\to 1.$ Asymptotically, the conditional variance in the multiplicative decomposition is equal to

align*[align* omitted — 97 chars of source]

Now, by writing $$g_{t,T} = 1 +\frac{1}{\alpha_0}g\left(t/T; \boldsymbol{\theta}\right)$$ for some parametric function $g(\cdot),$ we obtain the specification ((ref)). See Truquet17 for a similar development in the ARCH context. Consequently, financial insights originally used to motivate the multiplicative model ((ref)) apply equally to the ATV-GARCH model. AmadoT13 also briefly discuss an additive decomposition of volatility, but do not consider it further.

A computational advantage of the ATV-GARCH model is that the parameters can be estimated by QML in a single step. AmadoT13 estimated the parameters of the multiplicative decomposition ((ref)) by an iterative procedure called estimation by parts, which is detailed in SongFK05. Further, the asymptotic theory of QML estimation of the ATV-GARCH model relies on assumptions that are straightforward to verify and comparable to the ones in strictly stationary GARCH processes.

The observations are now assumed to come from an ATV-GARCH$(p,q)$ process with volatility equation

equation[equation omitted — 179 chars of source]

where $\alpha_0>0$, $\alpha_1, \alpha_2, \ldots, \alpha_p\geq0$, $\beta_1, \beta_2, \ldots, \beta_q\geq0$ and $g(t/T;\boldsymbol{\theta}_1 )>0$. We follow AmadoT13 and parameterize $g(\cdot)$ as a linear combination of logistic transition functions

equation[equation omitted — 147 chars of source]

where

equation[equation omitted — 157 chars of source]

is the general logistic function with $\gamma>0,\ c_{1}<\ldots<c_{K}.$ When $L>1$, the parameters satisfy $\gamma_{l}>0,l=1,\ldots,L,\ c_{11}<c_{12}<\ldots<c_{1K}<c_{21}<\ldots<c_{LK}$. We use the partition $\boldsymbol{\theta} = (\boldsymbol{\theta}^{\intercal}_1, \boldsymbol{\theta}^{\intercal}_2)^{\intercal}$, where $\boldsymbol{\theta}_1 = (\alpha_{01}, \ldots, \alpha_{0L}, \gamma_1, \ldots, \gamma_L, \boldsymbol{c}_{1}^{\intercal}, \ldots, \boldsymbol{c}_{L}^{\intercal})^{\intercal}$ and $\boldsymbol{\theta}_2 = (\alpha_0,\alpha_1, \ldots, \alpha_p, \beta_1, \ldots, \beta_q)^{\intercal}$. The ATV-GARCH model is unidentified if in ((ref)) at least one $\alpha_{0l}=0$, $l=1, \ldots, L$. A Lagrange multiplier (LM) test of the null hypothesis $\alpha_{01}=...=\alpha_{0L}=0$, based on approximating the alternative by a Taylor expansion around the null hypothesis as in LST88b, is presented in ABT2023b. This circumvents the identification problem that invalidates the standard asymptotic inference; see e.g. Davies77 who first pointed out this problem and discussed a solution to it.

When $L=1$ in ((ref)) and $K=1$ in ((ref)), the change in the unconditional variance is monotonic. Nonmonotonic change can be achieved by choosing $K>1$ or by setting $L>1$ in ((ref)). Typical transition functions are illustrated in Figure (ref). Panels (a) and (b) show how the logistic transition function behaves as functions of the shape parameter $\gamma$ and location parameter $c$. Panel (c) depicts a double increase in volatility and panel (d) illustrates a case where volatility first increases and then decreases. As shown by hornik1989, continuous functions on compact subsets of $\mathbb{R}$ can be uniformly approximated as closely as desired by linear combinations of logistic functions. This indicates the flexibility of the function $g(t/T;\boldsymbol{\theta}_1 )$.

figure[figure omitted — 1,085 chars of source]

The GARCH equation ((ref)) can easily be made asymmetric. The simplest way to do so is to add an indicator variable for negative returns as in the GJR-GARCH model by GlostenJR93. For notational simplicity we retain the form ((ref)).

Local stationarity

Since the intercept in ((ref)) is deterministically time-varying, the process $X_{t,T}$ is nonstationary. Standard asymptotic results for stationary and ergodic processes do not apply to the ATV-GARCH model. However, the rescaling device $t/T$ enables a meaningful asymptotic theory for processes that can be locally approximated by stationary processes.

This rescaling device is used by dahlhaus1997, dahlhaus2000 and DahlhausSR06 in the definition of local stationarity. In the following we work with $X_{t,T}^{2}$ instead of $X_{t,T}$ (cf. DahlhausSR06). By the triangle inequality, decompose the difference between $X_{t,T}^{2}$ and the stationary approximation $\widetilde{X}_t^{2}(u)$ at $u \in [0,1]$ as

equation[equation omitted — 201 chars of source]

It is seen from ((ref)) that if $t/T$ is close to $u$, then $X_{t,T}^{2}$ and $ \widetilde{X}_{t}^{2}(u)$ should be close and that the degree of the approximation depends on the rescaling factor $T$ and the deviation $|t/T-u|$. The process $\{X^2_{t,T}\}$ is said to be locally stationary if (Dahlhaus and Subba Rao 2006)

equation[equation omitted — 135 chars of source]

where $u\in[0,1]$ and $\widetilde{X}^2_t(u)$ is the stationary approximation at $u$.

rao2006 considers the following class of time-varying GARCH (tvGARCH) processes:

align[align omitted — 292 chars of source]

where the parameters $\alpha _{0}(t/T)$, $\alpha _{i}(t/T)$, $i=1,\ldots ,p$ , and $\beta _{j}(t/T)$, $j=1\ldots ,q$, are smooth functions of time. The stationary approximation $\widetilde{X}_t(u)$ of the tvGARCH process at $u$ is given by

align*[align* omitted — 248 chars of source]

Theorem 2.1 of rao2006 states conditions under which a nonstationary, nonlinear process with time-dependent parameters can be locally approximated by a stationary process. Applied to the tvGARCH process, the conditions are as follows (cf. Subba Rao 2006, Section 5.1).

itemize• The parameter curves $\{\alpha_0(\cdot)\}$, $\{\alpha_i(\cdot)\}$ and $\{\beta_j(\cdot)\}$ are Lipschitz continuous. • \begin{equation*} \sup_{u}\left\{ \sum_{i=1}^{p}\alpha _{i}\left( u\right) +\sum_{j=1}^{q}\beta _{j}\left( u\right) \right\} <1-\eta \end{equation*} for some $\eta >0$. • $\mathbb{E}(\varepsilon_{t}^{2}) = 1$.

We now return to the ATV-GARCH$(p,q)$-version of the model in ((ref)) and ((ref)) and show that it is locally stationary. The ATV-GARCH model is a parsimonious parameterization of the more general tvGARCH process ((ref)) and ((ref)), with $$\alpha _{0}(t/T; \boldsymbol{\theta}):=\alpha_{0}+g(t/T; \boldsymbol{\theta}_1),$$ where $g(t/T; \boldsymbol{\theta}_1)$ is defined in ((ref)) and ((ref)), $ \alpha _{i}(t/T)=\alpha _{i}$, $i=1,\ldots ,p$, and $\beta _{j}(t/T)=\beta _{j}$, $j=1,\ldots ,q$. The intercept and unconditional variance are time-varying, but the persistence is constant. The stationary approximation is obtained by fixing the intercept at the value the function $g(\cdot; \boldsymbol{\theta}_1)$ takes at $u$. The ATV-GARCH process is locally stationary and we have that

equation[equation omitted — 210 chars of source]

We state the result as a proposition.

\newtheorem{prop}{Proposition}

propAssume that the parameter space $\varTheta$ is compact, $\sum_{i=1}^{p}\alpha _{i} +\sum_{j=1}^{q}\beta _{j}<1$ and $\mathbb{E}(\varepsilon _{t}^{2})=1$. Then the ATV-GARCH model is locally stationary.
proofConditions (ii) and (iii) are satisfied by assumption. For (i), see Appendix A.3.
rmkThe condition on the GARCH coefficients is a necessary condition for weak stationarity of the GARCH$(p,q)$\ process. For estimation by QML, it is a stronger assumption than needed for strictly stationary GARCH processes, where a weaker condition on the coefficients can be obtained in terms of the top Lyapunov exponent; see BHK and FrancqZakoian2004.

Moment assumptions and the stationary approximation

In order to prove consistency and asymptotic normality of the QMLE of the parameters of the ATV-GARCH model, we require assumptions on the moment structure of both the process $\{X_{t,T}\}$ and the errors $\{\varepsilon_t\}$. The GARCH$(p,q)$ process has a finite second moment under the condition $\sum_{i=1}^{p}\alpha _{i} +\sum_{j=1}^{q}\beta _{j}<1$ coupled with the existence of a second moment of the error term. To prove asymptotic normality of the QMLE in strictly stationary GARCH processes, it suffices to assume a finite fourth moment of the errors; see BHK and FrancqZakoian2004. For the ATV-GARCH model, we also have to assume that the process possesses a finite fourth moment. To explain why, we require some additional results from rao2006. Following Subba Rao (2006), the tvGARCH$(p,q)$ process $\{X_{t,T}\}$ admits the state space representation (assume without loss of generality that $p,q\geq 2$)

equation[equation omitted — 154 chars of source]

with

equation*[equation* omitted — 172 chars of source]
equation*[equation* omitted — 117 chars of source]

and

equation*[equation* omitted — 345 chars of source]

a $(p+q-1)\times (p+q-1)$ matrix where $\boldsymbol{\tau }_{t}(u)=(\beta _{1}(u)+\alpha _{1}(u)\epsilon_{t-1}^{2},\beta _{2}(u),\ldots ,\beta _{q-1}(u))$, $ \boldsymbol{\alpha }(u)=(\alpha _{2}(u),\ldots ,\alpha _{p-1}(u))$ and $\mathbf{z}_{t-1}^{2}=(\varepsilon _{t-1}^{2},0,\ldots ,0)\in \mathbb{R}^{q-1}$. The stationary process $\{\widetilde{\mathcal{X}}_t(u)\}$ at $u$ is given by

equation[equation omitted — 131 chars of source]

We now show that the ATV-GARCH process admits the representation ((ref)), and that the results in Subba Rao (2006) apply to it. The vectors $\mathbf{b}_{t}(u)$ are defined as in ((ref)) with $\alpha_0 + g(u)$ in place of $\alpha_0(u)$. The matrices $\mathbf{A}_t(u)$ are matrices with constant parameters. By Lipschitz continuity of $g(u)$, it follows from Subba Rao (2006, Theorem 2.1) that

equation[equation omitted — 156 chars of source]

where $W_t$ and $V_{t,T}$ are stochastic processes. For asymptotic results for locally stationary processes, conditions of the type

equation[equation omitted — 161 chars of source]

are needed (see Appendices A.1 and A.2). For consistency, we need ((ref)) to be satisfied with $n=1$ and for asymptotic normality with $n=2$. It is obvious that ((ref)) requires the existence of moments of $W_t$ and $V_{t,T}$, to which we turn next.

Define the stationary sequence $$ \widetilde{\mathbf{b}}_{t}=\left( \sup_{u\in[0,1]}\alpha_0(u), 0, \ldots, 0 \\ \right)^\mathsf{T} \in \mathbb{R}^{p+q-1}. $$ The quantity $\widetilde{\mathbf{b}}_{t}$ is deterministic and bounded. The matrices included in Assumption 2.1 of Subba Rao that bound $\mathbf{A}_t(u)$ can be taken to be $\mathbf{A}_t(u)$ in ((ref)) with constant parameters, denoted here $\mathbf{A}_t$. Define the matrix $(\mathbf{A})_n = \left(\mathbb{E}\left|A_{t, ij}\right|^n\right)^{1/n}$ and let $\lambda_{\text{spec}}(\mathbf{A})$ denote the largest absolute eigenvalue of $\mathbf{A}$. By Proposition 2.1 of Subba Rao (2006), if

(i) $\left\lVert\widetilde{\mathbf{b}_t}\right\rVert_n^n<\infty$

(ii) $\lambda_{\text{spec}}((\mathbf{A})_n)<1-\delta$ for some $\delta>0$, then

equation[equation omitted — 66 chars of source]

and

equation[equation omitted — 70 chars of source]

By boundedness of the intercept, (i) is fulfilled. For asymptotic normality, we have to assume that (ii) holds with $n=2$. Because the matrices $\mathbf{A}_{t}$ are matrices with constant parameters, (ii) with $n=2$ is equivalent to the necessary and sufficient condition for the existence of a fourth moment of the GARCH process in he1999, and LingMcaleer. To see this, consider the ATV-GARCH$(1,1)$ model with one transition function. The state space representation is given by ((ref)) with

align*[align* omitted — 260 chars of source]

We have

align*[align* omitted — 603 chars of source]

Since the trace of a matrix is equal to the sum of its eigenvalues, the condition $\lambda_{\text{spec}}((\mathbf{A})_{2})<1-\delta $ translates into

equation[equation omitted — 163 chars of source]

which, less the term $\delta$, is the fourth moment condition in he1999. Figure 3 depicts the region of existence of the fourth moment for the ATV-GARCH$(1,1)$ model with Gaussian errors: $\mathbb{E}(\varepsilon_{t}^{2})=1$ and $\mathbb{E}(\varepsilon_{t}^{4})=3$ in ((ref)).

figure[figure omitted — 226 chars of source]

Maximum likelihood estimation

In this section we consider estimation of the ATV-GARCH model by QML. Estimation is based on a representation of $\sigma^2_{t,T}$ in terms of past observations $X_{t-i,T}$ and the time-varying intercept $g_{t-i+1,T}$, $1 \leq i < \infty$. We use results from DRW for processes that can be locally approximated by stationary processes to derive the asymptotic properties of the QMLE.

Representation for the conditional variance

BHK derived a representation $h_{t,T}(\boldsymbol{\theta})$ for the conditional variance of the GARCH$(p,q)$ process in terms of past observations $X_{t-i,T}$. Our aim is to obtain a similar representation for the ATV-GARCH process in terms of $X_{t-i,T}$ and the time-varying intercept $g_{t-i+1,T}(\boldsymbol{\theta}_1)$, $1 \leq i < \infty$. The BHK representation was also used by ChenH16 in their non-parametric tvGARCH model. First, we define a parameter space similar to the one in BHK. Let $\boldsymbol{\theta} = (\boldsymbol{\theta}_1^{\intercal},\boldsymbol{\theta}_2^{\intercal})^{\intercal}$ be the partitioning of the parameters in Section 2.1, where $\boldsymbol{\theta}_1 \subset \mathbb{R}^{\text{dim}(\boldsymbol{\theta}_1)}$ contains the parameters in the parameterisation of the time-varying intercept and $\boldsymbol{\theta}_2 \subset \mathbb{R}^{p+q+1}$ the remaining GARCH parameters. Let $0<\underline{\vartheta}<\overline{\vartheta}$, $0<\rho_0<1$ and $q\underline{\vartheta}<\rho_0$. Define

equation[equation omitted — 122 chars of source]

and

align[align omitted — 346 chars of source]

The sequence $h_{t,T}(\boldsymbol{\theta})$ is computed from

align[align omitted — 227 chars of source]

where

align[align omitted — 83 chars of source]

and $g_{t,T}(\boldsymbol{\theta}_1)$ is a short-hand notation for $g(t/T;\boldsymbol{\theta}_1)$.

The formulas for the coefficients $c_i(\boldsymbol{\theta})$ in ((ref)) are given in BHK. By Lemma 3.1 of BHK, formula (3.4), the coefficients $c_{i}(\boldsymbol{\theta})$ satisfy

align[align omitted — 108 chars of source]

for some constant $C_2$. The coefficients $d_{i}(\boldsymbol{\theta})$ obey a similar formula and (see Appendix A.3)

align[align omitted — 108 chars of source]

BHK proved that under the assumption $\mathbb{E}\ln\sigma^2_0<\infty$, the representation ((ref)) with $d_{i}(\boldsymbol{\theta})=0$, $1 \leq i < \infty$, yields $\sigma^2_{t,T}$ almost surely. In the ATV-GARCH model, the result follows similarly by considering quantities that bound $\sigma^2_{t,T}.$ Assume for notational simplicity the ATV-GARCH$(1,1)$ model. Then from ((ref)),

align[align omitted — 147 chars of source]

Define the nonstationary sequence $\Phi_{t,T} = \alpha_0 + g_{t,T}(\boldsymbol{\theta}_1) + \alpha_1X^2_{t-1,T}$. Iterating ((ref)), we obtain

align[align omitted — 150 chars of source]

Lemma 2.2 of BHK yields that the left-hand side of ((ref)) with $g_{t,T}(\boldsymbol{\theta}_1) = 0$ in ((ref)) converges almost surely as $i\to\infty$. Define $\Phi_{t}^* = \alpha_0 + \sup_{[0,1]}g_{t,T}(\boldsymbol{\theta}_1) + \alpha_1X^{*2}_{t-1}$, and similarly $\sigma^{*2}_{t}$. By weak stationarity of the bounding process, $\mathbb{E} \sigma^{*2}_0<\infty$, and therefore also $\mathbb{E}\ln \sigma^{*2}_0<\infty$. Lemma 2.2 of BHK now also yields almost sure convergence of the bounding process $\Phi_t^*$, which in turn implies that the original process also converges. Next we use the following zero-one law, which is a consequence of the first Borel-Cantelli lemma: if $\{Z_t\}$ is a sequence and $\sum_{t=1}^{\infty} P(|Z_t| > \delta ) < \infty$ for all $\delta>0$, then $Z_T \to 0$ a.s., as $T\to\infty$; see gut2005. Applying this theorem to the right-hand side of ((ref)) yields

align*[align* omitted — 98 chars of source]

so

equation*[equation* omitted — 59 chars of source]

almost surely as $i\to\infty$ since $\beta_1^i$ decays exponentially and $\sigma^2_{t,T}$ is bounded by $\sigma^{*2}_t,$ which by weak stationarity is bounded in expectation. This proves the representation ((ref)) for the ATV-GARCH$(1,1)$ model. Let $\boldsymbol{\theta}_0$ denote the "true" parameter vector and note that $\sigma^2_{t,T}=h_{t,T}(\boldsymbol{\theta}_0)$. The representation for the ATV-GARCH$(p,q)$ model follows along similar lines. This scheme generates the nonstationary sequence $h_{t,T}(\boldsymbol{\theta})$ without the need to initialize it with arbitrary starting values but, as noted by FrancqZakoian2004, the computational cost of the procedure is of order $O(T^2)$ instead of $O(T)$.

Likelihood functions

We now turn to QML estimation of the ATV-GARCH model ((ref)) and ((ref))--((ref)). The Gaussian log-likelihood function of $(X_{1,T} \ldots, X_{T,T})$, is given by

equation[equation omitted — 114 chars of source]

where

equation*[equation* omitted — 152 chars of source]

is the log-likelihood for observation $t$. The QMLE is defined as

equation[equation omitted — 185 chars of source]

Denote the score for observation $t$ by $\boldsymbol{s}_{t,T}(\boldsymbol{\theta})= \partial l_{t,T}(\boldsymbol{\theta})/\partial \boldsymbol{\theta}$ and the Hessian by $\boldsymbol{H}_{t,T}(\boldsymbol{\theta})= \partial^2 l_{t,T}/(\boldsymbol{\theta})\partial \boldsymbol{\theta}\partial\boldsymbol{\theta}^{\intercal}.$ We have

align[align omitted — 437 chars of source]

and

align[align omitted — 875 chars of source]

In practice, we only observe $(X_{1,T}, \ldots, X_{T,T})$, so the log-likelihood ((ref)) cannot be computed. Hence, as in BHK, we replace $l_{t,T}(\boldsymbol{\theta})$ with

equation*[equation* omitted — 171 chars of source]

where

equation[equation omitted — 230 chars of source]

The truncated log-likelihood function is given by

equation[equation omitted — 112 chars of source]

Analogously to ((ref)), the truncated QMLE is defined as

equation[equation omitted — 176 chars of source]

and, analogously to ((ref)) and ((ref)) the expressions for the score and Hessian for observation $t$ are $\bar{\boldsymbol{s}}_{t,T}(\boldsymbol{\theta})$ and $\bar{\boldsymbol{H}}_{t,T}(\boldsymbol{\theta})$.

Our proofs of consistency and asymptotic normality of the QMLE are based on the local approximation of the log-likelihood function $l_{t,T}(\boldsymbol{\theta})$. Thus we also need the log-likelihood function of the stationary process $\widetilde{X}_t(u)$, defined by

equation[equation omitted — 124 chars of source]

where

equation*[equation* omitted — 194 chars of source]

Furthermore,

equation*[equation* omitted — 228 chars of source]

for $u\in[0,1]$, where $g_{t,T}(\boldsymbol{\theta}_1)$ is approximated by $g_{t}(u, \boldsymbol{\theta}_1)$. Similarly to ((ref)) and ((ref)), the expressions for the score and Hessian for observation $t$ are $\widetilde{\boldsymbol{s}}_t(u,\boldsymbol{\theta})$ and $\widetilde{\boldsymbol{H}}_t(u,\boldsymbol{\theta})$.

Main results

We are now ready to present the main results on consistency and asymptotic normality of the QMLE of the parameters of the ATV-GARCH model.

To show consistency of the QMLE, we make the following assumptions.

itemize• The random variables $\varepsilon _{t}$ are IID with $\mathbb{E}(\varepsilon_0) = 0$ and $\mathbb{E}(\varepsilon_0^2) = 1$. $\mathbb{E}\left|\varepsilon _{0}^{2}\right|^{1+d}<\infty$, for some $d>0$. The random variable $\varepsilon^2_0$ is non-degenerate and $\underset{t\to0}{\lim}\ t^{-\mu}\mathbb{P}\{\varepsilon^2_0\leq t\}=0$, for some $\mu>0.$ • The parameter space $\Theta$ is compact and the vector $\boldsymbol{\theta}_{0}\in \mbox{int}(\Theta).$ • The equation $$\boldsymbol{\lambda}^{\intercal}\frac{\partial \alpha_0(u,\boldsymbol{\theta}_0)}{\partial \boldsymbol{\theta}} =\boldsymbol{0}$$ for a vector of constants $\boldsymbol{\lambda}$ implies $\boldsymbol{\lambda} = \boldsymbol{0}.$ • The polynomials $\mathcal{A}(z)=\alpha_1z+\alpha_2z^2+\ldots+\alpha_pz^p$ and $\mathcal{B}(z)=1 - \beta_1z-\beta_2z^2-\ldots-\beta_qz^q$ are coprime on the set of polynomials with real coefficients. • $\sum_{i=1}^{p}\alpha _{i}+\sum_{j=1}^{q}\beta_j<1.$
rmkThe assumption that the random variables $\varepsilon _{t}$ are IID$(0,1)$ was made in Proposition 1 to show that the ATV-GARCH model can be locally approximated by a stationary process. The remaining assumptions in (A1) are inherited from BHK. The assumption (A2) is a standard assumption used for proving consistency and asymptotic normality. As BHK noted, (A2) precludes zero coefficients for $\boldsymbol{\alpha}=(\alpha_1, \ldots, \alpha_p)^\intercal$ and $\boldsymbol{\beta}=(\beta_1, \ldots \beta_q)^\intercal$. (A3) is an identification condition on the time-varying intercept. In the appendix, we show that the identification condition holds for an intercept containing one logistic transition function with one $c$ parameter. It is likely to hold more generally, but to avoid difficulties in proving it, we leave it as an assumption. (A4) is an identification condition. (A5) together with the first part of (A1) is a sufficient condition for weak stationarity of the standard GARCH process.

We are now ready to present the following theorem.

thmUnder (A1)--(A5), \begin{equation*} \widehat{\boldsymbol{\theta}}_T\overset{P}{\to}{\boldsymbol{\theta}_0}, as T\to\infty. \end{equation*}
proofSee Appendix A.4.
rmkUnder (A1)–(A5), the ATV-GARCH process is locally stationary. In the proof of the theorem, we show that the sequence of log-likelihood functions $l_{t,T}(\boldsymbol{\theta})$ can be locally approximated by the stationary process $\widetilde{l}_t(u,\boldsymbol{\theta})$. Using the global law of large numbers (Theorem 2.7(i) of DRW), our proof also requires $\sup_{u\in[0,1]}\left\lVert\widetilde{l}_t(u,\boldsymbol{\theta})\right\rVert_1 <\infty$.
rmkA main difference between the ATV-GARCH model and the nonparametric maximum likelihood estimation of parameter curves in DRW is that in a nonparametric framework, smoothness conditions on the log-likelihood function are imposed, whereas in the parametric ATV-GARCH model we need to show that Lipschitz- or H\"{o}lder-type conditions hold for the log-likelihood function and its derivatives (the score and the Hessian). In the proof we show that a condition of the type ((ref)) with $n=1$ holds for the log-likelihood function.
rmkTheorem 2.7 (i) of DRW gives convergence in $L^{1}$. Therefore, we are not able to show strong (almost sure) consistency of the QMLE of the parameters.
rmkIn practice, we have to use the truncated estimator $\bar{\boldsymbol{\theta}}_{T}$ instead of the estimator $\widehat{\boldsymbol{\theta}}_{T}$. To prove consistency of $\bar{\boldsymbol{\theta }}_{T}$ from consistency of $\widehat{\boldsymbol{\theta }}_{T}$, it is sufficient to show that the truncated log-likelihood function $\bar{L}_{T}(\boldsymbol{\theta })$ converges uniformly to the log-likelihood function $L_{T}(\boldsymbol{\theta})$ (cf. BHK).

To show asymptotic normality of the QMLE, we make the following additional assumptions.

itemize$\mathbb{E}|\varepsilon_t|^{4+d}<\infty$, for some $d>0$. • $\lambda_{\text{spec}}((\mathbf{A})_2)<1-\delta$, for some $\delta>0$, where $\mathbf{A}_t(u)$ is given in ((ref)) and the parameters are constant.
rmk(A7) is equivalent to the necessary and sufficient condition for the existence of a fourth moment of the standard GARCH$(p,q)$ process in He and Teräsvirta (1999) and LingMcaleer.
thmUnder (A1)--(A7), \begin{equation} \sqrt{T}\left(\widehat{\boldsymbol{\theta}}_T-\boldsymbol{\theta}_0\right) \overset{D}{\to} N\left(\mathbf{0},\mathbf{B}^{-1}\mathbf{A}\mathbf{B}^{-1}\right), as T\to\infty. \end{equation} The expressions for $\mathbf{A}$ and $\mathbf{B}$ are integrals of the stationary approximation of the expected score and Hessian at $u$, $u \in [0,1]$, respectively: $$ \mathbf{A} = \int_0^1 \mathbb{E}(\mathbf{A}(u)) \ \text{d}u, $$ where $$ \mathbf{A}(u) = \text{Var}(\widetilde{\mathbf{s}}_{0} (u,\boldsymbol{\theta}_0)), $$ and $$ \mathbf{B} = \int_0^1 \mathbb{E}(\widetilde{\mathbf{H}}(u,\boldsymbol{\theta}_0))\ \text{d}u. $$
proofSee Appendix A.5.
rmkTo obtain asymptotic normality, we show that a condition of the type ((ref)) with $n=2$ holds for the score. Further, the proof requires a moment condition and a mixing condition on the stationary approximation $\widetilde{\mathbf{s}}_{t}(u,\boldsymbol{\theta})$ of the score.
rmkTo show that Theorem 2 remains true for the truncated estimator $\bar{\boldsymbol{\theta}}_{T}$, it suffices to show that $|\widehat{\boldsymbol{\theta}}_{T}-\bar{\boldsymbol{\theta}} _{T}|=O(T^{-1})$ almost surely. Asymptotic normality of $\sqrt{T}(\bar{\boldsymbol{\theta}}_{T}-\boldsymbol{\theta}_{0})$ is then an immediate consequence of Theorem 2 (cf. BHK, proof of Theorem 4.4, p. 226).

Simulation study

In order to examine the approximations in Theorems 1 and 2, we carry out a simulation study. We use time series lengths $T = 3000$ and $6000$. The number of Monte Carlo replications is $10000$. The GARCH$(1,1)$ parameters in the data-generating processes (DGPs) are $\alpha _{0}=0.05$, $\alpha _{1}=0.1$ and $\beta _{1}=0.8$. We consider three different DGPs for $g(t/T;\boldsymbol{\theta }_{1})=\alpha_{01} G(t/T; \gamma, c)$, with $\alpha _{01}=0.15$. This implies that the intercept increases by a factor of four. The value of the linear parameter $\alpha_{01}$ is empirically relevant. In the empirical example in the next section, the intercept is reduced by a factor of five. We experimented with other values for $\alpha_{01}$ and found that the approximations work better in finite samples when the transition is a pronounced feature of the data. We choose $\gamma $ such that $G(a;\gamma ,c)=0.01$ and $G(b;\gamma ,c)=0.99$. The values $a=0.1$ and $b=0.9 $ in DGP 1 define a slow transition with $80\%$ of the observations affected by the transition. The values $a=0.25$ and $b=0.75$ in DGP 2 define a moderate transition with $50\%$ of the observations affected by the transition. Finally, the values $a=0.4 $ and $b = 0.6$ in DGP 3 define a rapid transition with $20\%$ of the observations affected by the transition. The value for $c$ is $c=0.5$. The parameter values are contained in Table (ref). Figure 3 plots the transition functions of DGPs 1--3. The errors $\{\varepsilon_t\}$ are NID$(0,1)$. To reduce the impact of the starting values, we use a burn-in period of $500$ observations.

figure[figure omitted — 183 chars of source]

ML estimation is carried out using the solver solnp by Ye1987, implemented in the R package Rsolnp by Rsolnp. The recursion for the conditional variance ((ref)) is truncated at $200$ observations. This has a negligible impact on the conditional variance, but it drastically speeds up estimation. In the estimation we impose $\alpha _{1}+\beta _{1}<1$. For large values of $\gamma $, its impact on the shape of the transition function becomes small and the log-likelihood function is flat in the direction of this parameter. To improve the numerical accuracy of the estimate of $\gamma$, we apply the transformation $\gamma =\eta (1-\eta )^{-1}$ proposed by ekner2013parameter. We set the maximum value for $\eta$ equal to $0.999$, which was reached in a small number of replications. They were discarded and new draws obtained until the number of replications reached $10000$. For each DGP, we report in Table (ref) the frequency with which this happens. The starting values for the parameters $(\alpha _{0},\alpha _{1},\beta _{1},\gamma,c,\alpha _{01} )$ are set to the true values in the DGPs.

table[table omitted — 521 chars of source]

Table 2 reports the means and standard deviations of the parameter estimates. The estimates converge to the true parameter values when $T$ increases. The results for $\widehat{\alpha }_{1}$ show that the ARCH parameter $\alpha _{1}$ is accurately estimated in moderate samples. The estimates $\widehat{\alpha }_{0}$ and $\widehat{\alpha }_{01}$ converge from above, whereas the estimate $\widehat{\beta }_{1}$ converges from below. In small samples, the effect of the overestimated intercept parameters $\alpha _{0}$ and $\alpha _{01}$ on the unconditional variance is offset by an underestimated $\beta_1.$ Consequently, the time-varying intercept is more accurately reproduced than suggested by the individual parameter estimates. The location parameter $c$ and the slope parameter $\gamma $ are well estimated in moderate samples. For large values of $\gamma $, the estimate $\hat{\gamma}$ is less precise. As $\gamma $ increases, the first derivative of the logistic transition function with respect to $\gamma $ increases around $c$, making the Lipschitz constant larger. This means that the number of observations affected by the transition gets smaller.

Figures 4--9 show the distributions of the standardised parameter estimates. It is seen that the empirical distributions of the standardised parameter estimates of the GARCH parameters $(\alpha _{0},\alpha _{1},\beta _{1})$ are close to normal for moderate sample sizes. For the estimates of the nonlinear parameters $\gamma $ and $c$, longer time series are required for the normal approximation to be good. As the increasing Lipschitz constant suggests, the quality of the normal approximation deteriorates with an increasing $\gamma $. In some replications the algorithm gets stuck at the initial value for $\gamma $, which results in a spike in the simulated distribution of the standardised parameter estimates.

The ATV-GARCH model is a model for long financial time series which exhibit gradual change in the unconditional volatility. A consequence of this is that long time series are required to provide enough observations affected by the transition, so that the normal approximation to $\widehat{\gamma }$ is a good approximation. From a practical point of view, the logistic transition function is an approximation of a non-linear structure. In empirical applications of the ATV-GARCH model the shape of the transition function implied by the parameter estimates of $\gamma $ and $c$ is more important than the exact estimates and their standard errors.

table[table omitted — 2,244 chars of source]
figure[figure omitted — 234 chars of source]
figure[figure omitted — 234 chars of source]
figure[figure omitted — 234 chars of source]
figure[figure omitted — 234 chars of source]
figure[figure omitted — 234 chars of source]
figure[figure omitted — 234 chars of source]

Empirical example

As an example of the use of the ATV-GARCH model we consider Oracle Corporation (ticker: ORCL) daily closing prices on the New York Stock Exchange from the beginning of its listing on 12 March 1986 until 13 February 2023, 9306 observations in all. The data are downloaded from Yahoo Finance. The series and the corresponding logarithmic returns are depicted in Figure 10. As can be seen, up until the early 2000s the amplitude of the volatility clusters is large, and the descent that follows corresponds to the time when the dot-com bubble bursts. Summary statistics for the log returns appear in Table 3, Panel A. They show that the robust skewness measure based on quartiles does not indicate any skewness. This suggests that the observed nonrobust skewness is mainly caused by a small number of large negative returns being greater than their positive counterparts. As may be expected, both measures of kurtosis indicate leptokurtosis in the returns.

The results of fitting a standard GARCH$(1,1)$ model to the logarithmic returns can be found in Table 3, Panel B. The estimated persistence is $\widehat{\alpha }_{1}+\widehat{\beta }_{1} > 0.999$, which indicates that the unconditional variance may not be constant over time. This tentative conclusion is supported by the LM-type tests discussed in Section 2.1. The table contains results from both the nonrobust and robust (robustified against departures from distributional assumptions as in wooldridge1990) versions of the tests. The statistics are asymptotically $\chi^{2}$-distributed with three degrees of freedom under the null hypothesis of constant unconditional variance. Both tests strongly reject the null hypothesis.

The parameter estimates and standard errors of the fitted ATV-GARCH(1,1) model can be found in Table (ref), Panel C. The results show that the persistence, compared to the standard GARCH, decreases from above $0.999$ to $0.94$. This lends credence to the hypothesis in Diebold (1986). The fitted variance and transition function for the period $1998-2005$ are shown in Figure (ref). We note that the fitted model agrees with the informal discussion in the beginning of this section. The smoothly time-varying intercept starts out high and exhibits a precipitous but discernibly gradual decline during the years $2001-2004$, which coincides with the bursting of the tech bubble. To put the value $\widehat{\eta} =0.994$ of the slope parameter into perspective, the reverse transformation yields $\widehat{\gamma } =166$. Very small changes in $\eta$, when it is large, translate into substantial changes in $\gamma $. From Table (ref) it is seen that the standard deviation of $\eta $ equals 0.003. The range $\widehat{\eta } \pm 0.003$ corresponds to the interval $(110, 332)$ for $\widehat{\gamma }$.

As can be seen from Table 3, Panel B, testing one transition against two does not lead to a rejection of the null hypothesis. Our conclusion is that the ATV-GARCH$(1,1)$ model with a single transition is an adequate first-order GARCH-type description of the volatility of the Oracle return series.

figure[figure omitted — 642 chars of source]
figure[figure omitted — 596 chars of source]
table[table omitted — 4,328 chars of source]

Conclusion

We have introduced a new GARCH model with a deterministic time-varying intercept, called the ATV-GARCH model. In this model, volatility is mean-reverting towards a time-varying mean. The model captures structural change that slowly affects the amplitude of a time series while keeping the short-run dynamics constant over time. The idea is that short-run fluctuations in volatility are stationary, but long-run change makes the assumption of stationarity inappropriate. The model is particularly well suited for situations in which the volatility of an asset or index is smoothly increasing or decreasing over time.

The ATV-GARCH model is globally nonstationary but can be locally approximated by a stationary GARCH process. The stationary approximation enables the derivation of asymptotic results. We show that the QMLE of the parameters of the ATV-GARCH model is consistent and asymptotically normal. The asymptotic theory of QML estimation of the ATV-GARCH model relies on assumptions that are straightforward and comparable to the ones in strictly stationary GARCH processes.

We apply our model to Oracle Corporation daily stock returns. As LM tests find strong evidence of a smooth transition in the intercept, we fit an ATV-GARCH$(1,1)$ model to the series. The persistence implied by the standard GARCH parameter estimates is substantially reduced. We conjecture that, by using an ATV-GARCH model, the estimated persistence in volatility is reduced in many situations. This reduction would likely impact the results of empirical studies in which GARCH models are fitted to long financial time series with the purpose of forecasting volatility. It is possible that the slowly changing intercept may generate spurious evidence of long memory in volatility, leading to an IGARCH or a FIGARCH model. This is an empirical problem which is left for future research.