Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
71,579 characters · 8 sections · 0 citation commands
Statistical inference for autoregressive models under heteroscedasticity of unknown form
Consider a $p$-th order heteroscedastic auto-regressive (AR) model:
where the error $\varepsilon_{t}$ satisfies
for $t=1,\cdots,n$. Here, $g(\cdot)$ is a positive bounded scalar function with unknown form, $u_{t}$ is a re-scaled error with unknown form and a finite variance for all $t$, and $n$ is the sample size. In model ((ref)), $y_{t}$ is allowed to be non-stationary with a finite variance due to the heteroscedasticity of $\varepsilon_{t}$, but it can not be a unit root process, since the coefficients \{$\phi_{i0}\}_{i=1}^{p}$ are required to satisfy the stationarity condition of AR($p$) model. In model ((ref)), $\{u_{t}\}$ need not be independent and identically distributed (i.i.d.), nor even a martingale difference sequence (m.d.s.), allowing for model mis-specification. This setting is important, in view of that the literature generally needs a non-i.i.d. sequence $\{u_{t}\}$ (see, e.g., Drost and Nijman (1993)). Under certain identification condition on $u_t$, $g_t$ is identifiable. Particularly, when $u_{t}$ is stationary, the (conditionally) heteroscedastic structure of $\varepsilon_{t}$ is allowed to change abruptly, gradually, or periodically according to the un-specified form of $g(\cdot)$. Similar specifications as for $g_{t}$ have been widely adopted in the literature; see, e.g., Robinson (1989, 2012), Fryzlewicz, Sapatinas, and Subba Rao (2006), Zhou and Wu (2009), and Chen and Hong (2016) to name a few.
Model ((ref)) is one of the often used models of empirical macroeconomics and statistics. Conventional statistical inference methods for this model are designed for homogenous error $\{\varepsilon_{t}\}$ (i.e., $E\varepsilon_{t}^{2}=\mbox{a constant}$), which is either a sequence of i.i.d. random variables or an m.d.s.; see, e.g., Gon\c{c}alves and Kilian (2004), Yao and Brockwell (2006), Zhu and Ling (2015) and references therein. However, homogenous error can be a crucial weakness in applications, since heteroscedasticity has been widely demonstrated by Tsay (1988) for social science data, Watson (1999) for interest rates data, Busetti and Taylor (2003) and Sensier and van Dijk (2004) for macroeconomic data, Amado and Ter\"{a}svirta (2013, 2014) for stock index data, and many others. As shown in Diebold (1986), Mikosch and St\v{a}ric\v{a} (2004), and St\v{a}ric\v{a} and Granger (2005), the presence of heteroscedasticity could mislead the conventional time series analysis procedure resulting in erroneous conclusions. Hence, it is necessary to develop a valid statistical inference procedure for model ((ref)) under heteroscedasticity.
So far few works have centered around this topic, and most of them are based on the least squares estimator (LSE); see Nicholls and Pagan (1983) for earlier works and Phillips and Xu (2006) for more recent ones. The seminal work in Carroll (1982) and Robinson (1987) demonstrated that the LSE in regression models is less efficient than the adaptive LSE (ALSE), which takes the unknown heteroscedastic form of the error term into account; see also Cragg (1983) for the study of the instrumental variable LSE. Motivated by this, Xu and Phillips (2008) constructed a feasible ALSE for model ((ref)), and showed that this feasible ALSE is more efficient than the LSE. However, their theory does not allow $u_{t}$ to be conditionally heteroscedastic, and their methodology may be limited in practice since no valid statistical inference tools (e.g., $t$-test, Wald test and model diagnostic checking) is provided. Moreover, a heavy-tailed error is an often observed feature in fitted model ((ref)), and the LSE-based inference methods in this case are undesirable from the viewpoint of robustness.
This paper provides an entire statistical inference procedure for model ((ref)) based on the weighted least absolute deviations estimator (LADE). We first establish the asymptotic normality of the weighted LADE. Second, since the asymptotic covariance matrix of the weighted LADE can not be estimated directly from the sample, we develop the random weighting (RW) method to estimate this covariance matrix, leading to the implementation of the Wald test. The RW method initiated by Jin, Ying, and Wei (2001) has been widely applied to provide statistical inference; see Zhu (2016). Third, we construct a portmanteau test for model diagnostic checking, and use the RW method to obtain its critical values. As a special weighted LADE, the infeasible adaptive LADE (ALADE) with weight equal to $g_{t}$ is desired in terms of efficiency, and its advantage in efficiency is demonstrated by a formal comparison. To circumvent the un-observable $g_{t}$, we further propose a feasible ALADE, prove that it has the same efficiency as its infeasible counterpart, and show that our entire inference procedure is still valid based on this feasible estimator. Under the identification condition $E|u_{t}|=1$ for $g_{t}$, this feasible ALADE is constructed with weight equal to a kernel estimator of $g_{t}$. Unlike Xu and Phillips (2008), our entire inference procedure based on the weighted LADE does not rule out the conditional heteroscedasticity of $u_{t}$. The importance of our methodology is illustrated by simulation results and the real data analysis on three U.S. economic data sets.
We emphasize that the idea of the weighted LADE has been adopted for the linear regression model with heteroscedasticity of either known form in Gutenbrunner and Jure\v{c}kov\'{a} (1992) or unknown form in Zhao (2001). But our technique is different from theirs, since they neither considered a general heteroscedastic time series model as models ((ref))-((ref)), nor provided an entire valid inference procedure. When $\varepsilon_{t}$ is stationary and conditionally heteroscedastic with unknown form, Zhu and Ling (2015) provided the statistical inference method for model ((ref)) based on either the self-weighted LADE or the usual LADE. Our weighted LADE and their self-weighted LADE are not compatible due to different weighting mechanisms. The weight $w_{sw,t}$ for the self-weighted LADE is a random variable in the filtration generated by $\{y_{k}\}_{k\leq t-1}$, and it is useful to deal with the infinite variance AR model. For the finite variance AR model (as in our setting), $w_{sw,t}$ is not needed any more, since the self-weighted LADE is always less efficient than the usual LADE in this case. On the contrary, the weight $w_{t}$ for our weighted LADE is a deterministic function (or its kernel estimate based on the whole sample $\{y_{t}\}_{t=1}^{n}$), and it is used to take the unknown heteroscedastic form of $\varepsilon_{t}$ into account. As a result, the weighted LADE can be more efficient than the usual LADE by choosing an appropriate weight (e.g., the weight for the feasible ALADE) in many circumstances. Compared to the usual LADE in Zhu and Ling (2015), our weighted LADE has two more incremental contributions besides the advantage in efficiency. First, the results based on the weighted LADE allow $\varepsilon_{t}$ to be heteroscedastic with unknown form, and hence the scope of applications is much wider in dealing with the finite variance AR model. Second, although the RW method is motivated by Zhu and Ling (2015), new proof techniques are used to handle the heteroscedasticity of $\varepsilon_{t}$, especially for the feasible ALADE.
The remainder of the paper is organized as follows. Section 2 obtains the asymptotic normality of the weighted LADE. Section 3 presents the RW method to estimate the asymptotic covariance matrix. Section 4 constructs a portmanteau test for model diagnostic checking. Section 5 considers the choice of the weight function, and discusses the asymptotic efficiency. Section 6 proposes the feasible ALADE. Simulation results are reported in Section 7. Concluding remarks are offered in Section 8. Proofs of all theorems are relegated to Appendixes. Additional simulation results, applications, some technical lemmas and the remaining proofs are provided in the supplementary material (Zhu, 2018).
Throughout the paper, $|\cdot|$ and $\|\cdot\|$ denote the absolute value and the vector $l_{2}$-norm, respectively, $A'$ is the transpose of matrix $A$, $\|\xi\|_{\kappa}=(E|\xi|^{\kappa})^{1/\kappa}$ is the $L^{\kappa}$-norm of random variable $\xi$, $\operatorname*{plim}$ denotes the convergence in probability, $\to_{d}$ denotes the convergence in distribution, $I(\cdot)$ is the indicator function, and $\mbox{sgn}(a)=I(a>0)-I(a<0)$ is the sign of any $a\in\mathcal{R}$.
Let $\theta=(\mu,\phi_1,\cdots,\phi_p)'\in\Theta$ be the unknown parameter of model ((ref)), and $\theta_0=(\mu_0,\phi_{10},\cdots,\phi_{p0})'\in\Theta$ be its true value, where $\Theta$ is the parameter space, and it is a compact subset of $\mathcal{R}^{p+1}$ with $\mathcal{R}=(-\infty,\infty)$. Given the observations $\{y_{t}\}_{t=1}^{n}$ and the initial values $\{y_{t}\}_{t=-p}^{0}\equiv 0$, we consider the weighted least absolute deviations estimator (LADE) for $\theta_{0}$:
where $Y_{t-1}=(1,y_{t-1},\cdots,y_{t-p})'$ and $w_{t}:=w(t/n)$ are weights with $w(\cdot)$ being a positive scalar function. Particularly, when $w(\cdot)\equiv 1$, $\widehat{\theta}_{wn}$ becomes the usual LADE, denoted by $\widehat{\theta}_{n}$. To study the asymptotic theory of $\widehat{\theta}_{wn}$, we need the notation of near-epoch dependent (NED) random variables.
{\sc Definition 1.} {\it\,\, $\{Z_{nt}\}$ is near-epoch dependent in $L^{\kappa}$-norm ($L^{\kappa}$-NED) on $\{V_{nt}\}$ if $\|Z_{nt}-E(Z_{nt}|\mathfrak{F}_{t-m}^{t+m})\|_{\kappa}\leq d_{nt}\psi_{m}$, where $\mathfrak{F}_{t-m}^{t+m}:=\sigma(V_{ns}; t-m\leq s\leq t+m)$ is a filtration with $\sigma(\cdot)$ being the $\sigma$-field operator, $\{d_{nt}\}$ are positive constants, and $\psi_{m}\to0$ as $m\to\infty$. }
The definition of near-epoch dependence was first introduced in Billingsley (1968) and has been widely used in the literature. NED processes allow for considerable heterogeneity and also for dependence and include the mixing processes as a special case. Many non-linear models are shown to be NED (see, e.g., Davidson (2002, 2004) for overviews), and hence the NED concept makes possible a convenient theory of inference for these models. Besides near-epoch dependence, the physical dependence in Wu (2005) is a parallel tool to depict the dependence of a dynamic process. It might be interesting to apply physical dependence to our setting, and we leave this for future study.
Let $f_{t}(x)$ be the conditional density of $u_{t}$ given $\mathcal{F}_{t-1}$, where $\mathcal{F}_{t}:=\sigma(u_{s}; s\leq t)$ is a filtration. We make the following five assumptions throughout the paper, where the definition of $f_{t}(x)$ is essential only for the last one.
We offer some remarks on the aforementioned assumptions. Assumption (ref) is the usual stationarity condition for AR($p$) model, and it implies that
where $\rho=\mu_0/(1-\phi_{10}-\cdots-\phi_{p0})$ and $\sum_{i=0}^{\infty}|\alpha_i|<\infty$. Assumption (ref) is sufficient to guarantee that both $g(\cdot)$ and $w(\cdot)$ are integrable on the interval $[0, 1]$ up to any finite order, and similar conditions have been used in Cavaliere (2004), Cavaliere and Taylor (2007), Phillips and Xu (2006), Xu and Phillips (2008), and many others. Assumption (ref) is the identification condition for $\theta_{0}$ based on the weighted LADE. Assumption (ref) is adopted from Kuersteiner (2002) and Gon\c{c}alves and Kilian (2004) in a similar way. Assumption (ref)(i) is stronger than the condition that $Eu_{t}^{2}<\infty$, which is sufficient for the asymptotic normality of the LADE if $\varepsilon_{t}$ is stationary; see, e.g., Zhu and Ling (2015) and references therein. We resort to a stronger moment condition of $u_{t}$ in Assumption (ref)(i) to handle the heteroscedasticity of $\varepsilon_{t}$. Assumption (ref)(ii)-(iii) are less restrictive than the condition that $u_{t}$ is covariance stationary with mean zero. Assumption (ref) is required only for the LADE but not for the LSE, and it is stronger than the one in Bassett and Koenker (1978), Davis and Dunsmuir (1997), Pan, Wang, and Yao (2007), and Zhu and Ling (2012, 2015) due to the heteroscedasticity of $\varepsilon_{t}$. When $f_{t}(0)\equiv f(0)$ is independent of $t$, Assumption (ref)(ii)-(iii) hold automatically, and Assumption (ref)(iv)-(v) become Assumption (ref)(ii)-(iii), respectively. When $u_{t}$ follows a stationary autoregressive conditional heteroskedasticity (ARCH) type model, the NED condition in Assumption (ref)(ii) can be checked or even be un-needed, and the moment conditions in Assumption (ref)(iii)-(v) can be satisfied (see the discussions for Corollaries 2.1-2.3 below).
We shall point out that Assumption (ref)(ii)-(iii) and Assumption (ref)(iii)-(v) need $m_{r}^{(1)}$, $m_{r,s}^{(2)}$, $\tau_{0}$, $\tau_{r}^{(1)}$ and $\tau_{r,s}^{(2)}$ to be independent of $t$. In the general framework, these terms may depend on $t$. A detailed study for this general framework might be done in a separate paper.
Let $d_{1}=\int_{0}^{1}\frac{1}{w^{2}(x)}dx$, $d_{2}=\tau_{0}\int_{0}^{1}\frac{1}{g(x)w(x)}dx$, and
With these notations, define
where $K_{l}$ is the $p\times 1$ column vector with $r$-th element $\zeta_{r}^{(l)}$, and $\Omega_{l}$ is the $p\times p$ matrix with $(r,s)$-th element $\xi_{r,s}^{(l)}$. Our first main result is given as follows:
In the finite variance AR model, the technique in Zhu and Ling (2015) for the usual LADE requires the stationarity of $y_t$, which is against our setting for a either time-varying $g_t$ or non-stationary $u_t$. Our proof techniques for the weighted LADE in Theorem (ref) or its related inferential methods below rely on Theorem 1 of Andrews (1988), and they require a key technical condition that $\{f_{t}(0)\}$ is $L^{2+\delta_1}$-NED on $\{u_t\}$ (see Assumption (ref)(ii)), which seems new to the study of the LADE. Moreover, in the present of $g_t$, we further propose a special weighted LADE with $w_t$ equal to a kernel estimator of $g_t$; see Section 6 below. In many situations, this special weighted LADE is “optimal” in terms of efficiency, but its related proof techniques (particularly for Propositions A.3-A.4 in Appendix A) are new and not available in Zhu and Ling (2015).
In the infinite variance AR model, Zhu and Ling (2015) constructed a self-weighted LADE, which is $\sqrt{n}$-consistent and asymptotically normal when $y_t$ is stationary. Meanwhile, they pointed out that the convergence rate of the usual LADE is slower than $n^{-1/2}$ when $\varepsilon_{t}$ follows the first order (G)ARCH model, while Davis (1996) proved that the convergence rate of the usual LADE is faster than $n^{-1/2}$ when $\varepsilon_{t}$ is i.i.d. This implies that the self-weighted LADE may not always outperform the usual LADE (or the weighted LADE) when $y_t$ is stationary with an infinite variance. In general, when $y_t$ is allowed to be non-stationary (as in our setting) with an infinite variance, the asymptotic properties of the weighted LADE and the self-weighted LADE are not clear at this stage, and we leave this topic for future study.
Next, we relax the technical conditions in Theorem (ref) for some special cases of model ((ref)). We first consider the case that $u_{t}$ has the ARCH-type structure (Engle, 1982; Bollerslev, 1986; Giraitis, Kokoska, and Leipus, 2000; Francq and Zako\"{i}an, 2010):
where $u_{t}$ is stationary, $\{\eta_{t}\}$ is a sequence of i.i.d. innovations with median zero, and $\sigma_{t}\in\mathcal{F}_{t-1}$ satisfies that $\sigma_{t}\geq \underline{c}$ for some $\underline{c}>0$. This condition on $\sigma_{t}$ directly holds for most of ARCH-type models as long as their intercept term has a positive lower bound. Since $u_{t}$ is stationary, Assumption (ref)(ii)-(iii) are not needed any more if Assumption (ref)(i) holds.
Let $f_{\eta}(\cdot)$ be the density of $\eta_{t}$. Under model ((ref)), $f_{t}(0)=f_{\eta}(0)/\sigma_{t}$. Since $\sigma_t\geq\underline{c}$, Assumption (ref)(ii) holds if $\{\sigma_t\}$ is $L^{2+\delta_{1}}$-NED on $\{u_{t}\}$, and Assumption (ref)(iii)-(v) hold by Assumption (ref)(i). Hence, Assumption (ref) can be relaxed to Assumption (ref) in this case.
Since $\sigma_t\geq\underline{c}$, by a minor extension of lemma 2.1 of Davidson (2002), we can show that Assumption (ref)(ii) holds if $\{\sigma_{t}^{\kappa_1}\}$ is $L^{2+\delta_{1}}$-NED on $\{u_{t}\}$ for some $\kappa_1\geq 1$. When $\sigma_{t}$ in ((ref)) satisfies that
with $\psi_0>0$ and $\psi_{i}\geq 0$ for all $i\geq 1$, it is straightforward to see that $\{\sigma_{t}^{\kappa_1}\}$ is $L^{\kappa_2}$-NED (for some $\kappa_2\geq1$) on $\{u_{t}\}$ if
Hence, Assumption (ref)(ii) holds under ((ref))-((ref)) with $\kappa_2=2+\delta_1$, which allow a general class of ARCH-type models, including (G)ARCH, ARCH($\infty$) (Robinson, 1991), power GARCH (Ding, Granger, and Engle, 1993), hyperbolic GARCH (Davidson, 2004), and many others. Note that our NED condition in ((ref)) is different from the one in Davidson (2004), since our filtration is based on $\{u_{t}\}$ instead of $\{\eta_{t}\}$.
Second, we consider the case that $f_{t}(0)$ depends on finite lags of $u_{t}$. In this case, Assumption (ref)(i) and Assumption (ref)(ii) can be further relieved as shown in Assumption (ref), which does not need $\{f_{t}(0)\}$ to be NED.
Clearly, Assumption (ref)(ii) holds under model ((ref)) with $\psi_i\equiv 0$ for all $i\geq i_0$ and some integer $i_0\geq 1$. This implies that when $u_{t}$ follows a stationary finite-lag ARCH model with $E|u_{t}|^{2+\delta_{0}}<\infty$ and $Eu_{t}^{4}=\infty$, $\widehat{\theta}_{wn}$ is asymptotically normal. However, the LSE in this case is asymptotically non-normal with a slower convergence rate than $n^{-1/2}$ as shown in Zhang and Ling (2015). Hence, the weighted LADE seems to be more convenient and efficient than the LSE to tackle the heavy-tailedness of $\varepsilon_{t}$.
Third, we consider the case that $\phi_{i0}\equiv 0$, under which model ((ref)) becomes a location model. For this location model, Assumptions (ref) is redundant, and Assumptions (ref)-(ref) can be replaced by Assumption (ref) below.
As argued before, Assumption (ref)(ii) holds under ((ref))-((ref)) with $\kappa_2=1$. It is worth noting that the weaker technical assumptions in Corollaries (ref)-(ref) not only suffice for Theorem (ref) but also for other asymptotic results subsequently.
Since the asymptotic variance of $\widehat{\theta}_{wn}$ in Theorem (ref) depends on the unknown forms of $g(\cdot)$ and $u_{t}$, it can not be estimated directly by its sample counterpart. To solve this problem, we use the random weighting (RW) method to approximate the limiting distribution of $\widehat{\theta}_{wn}$ in Theorem (ref). Let $\{w^{*}_{t}\}_{t=1}^{n}$ be a sequence of i.i.d. nonnegative random variables, with mean and variance both equal to 1. Define
Based on Assumption (ref), we can show that the distribution of $\sqrt{n}(\widehat{\theta}_{wn}-\theta_{0})$ can be approximated by the resampling distribution of $\sqrt{n}(\widehat{\theta}_{wn}^{*}-\widehat{\theta}_{wn})$.
According to Theorems (ref) and (ref), we can approximate the asymptotic covariance matrix of $\widehat{\theta}_{wn}$ by the following procedure:
1. Generate $J$ replications of the i.i.d. random weights $\{w_{t}^{*}\}_{t=1}^{n}$ from the standard exponential distribution, which has mean and variance both equal to one.
2. Compute $\widehat{\theta}_{wn}^{*}$ for $i$-th replication, and denote it as $b_{i}$.
3. Calculate the sample variance-covariance matrix of $\{b_{i}-\widehat{\theta}_{wn}\}_{i=1}^{J}$, denoted by $\widehat{V}_{wn}$, which provides a good approximation for the asymptotic covariance matrix of $\widehat{\theta}_{wn}$ for large $J$.
Based on $\widehat{V}_{wn}$, we can construct a Wald test statistic
to test the following linear null hypothesis:
where $\Gamma$ is an $s\times (p+1)$ constant matrix with rank $s$ and $r$ is an $s\times1$ constant vector. If $W_{wn}$ is larger than the upper-tailed critical value of $\chi^{2}_{s}$, the null hypothesis $H_{0}$ is rejected. Otherwise, $H_{0}$ is not rejected.
Model diagnostic checking is an important step in applications. Most of the existing methods, including the time domain correlation-based tests and frequency domain periodogram-based tests, require that the observed data are stationary with a homogenous innovation; see, e.g., Hong (1996), Hong and Lee (2005), Escanciano (2006), and Zhu and Li (2015) for an overview. Hence, they are invalid for model ((ref)), which calls for a new method for its diagnostic checking. In this section, we construct a sign-based Ljung-Box portmanteau test as in Zhu and Ling (2015) for this purpose.
Let $\varepsilon_{t}(\theta)=y_{t}-\theta'Y_{t-1}$. The idea of our sign-based portmanteau test is based on the fact that $\{\mbox{sgn}(\varepsilon_{t}(\theta_{0}))\}$ is a sequence of un-correlated random variables under Assumption (ref). Hence, if model ((ref)) is correctly specified, it is expected that the sample autocorrelation function of $\{\mbox{sgn}(\varepsilon_{t}(\widehat{\theta}_{wn}))\}$ at lag $k$, denoted by $\widehat{r}_{wn, k}$, is close to zero, where $$\widehat{r}_{wn,k}=\frac{\sum_{t=k+1}^{n}\left[ \mbox{sgn}(\varepsilon_{t}(\widehat{\theta}_{wn}))-\overline{\varepsilon}(\widehat{\theta}_{wn})\right] \left[ \mbox{sgn}(\varepsilon_{t-k}(\widehat{\theta}_{wn}))-\overline{\varepsilon}(\widehat{\theta}_{wn})\right]} {\sum_{t=1}^{n}\left[ \mbox{sgn}(\varepsilon_{t}(\widehat{\theta}_{wn}))-\overline{\varepsilon}(\widehat{\theta}_{wn})\right]^2}$$ with $\overline{\varepsilon}(\theta)=\frac{1}{n}\sum_{t=1}^{n}\mbox{sgn}(\varepsilon_{t}(\theta))$. Let $\widehat{r}_{wn}=(\widehat{r}_{wn,1},\widehat{r}_{wn,2},\cdots,\widehat{r}_{wn,M})'$ for some integer $M\geq1$. To study the joint distribution of $\widehat{r}_{wn}$, we need one more assumption below.
Let $d_{3}=\int_{0}^{1}\frac{1}{g(x)}dx$ and $$ K_{3k}= \left(
\right), \,\,\,\, \Omega_{3k}= \left(
\right) $$ for $k=1,\cdots,M$. With these notations, define
where $K_{3}=\big(\int_{0}^{1}\frac{g(x)}{w(x)}dx\big)\dot{K}_{3}$ with $\dot{K}_{3}=(K_{31},K_{32},\cdots,K_{3M})\in\mathcal{R}^{(p+1)\times M}$, $\Omega_{3}=(\Omega_{31},\Omega_{32},\cdots,\Omega_{3M})' \in\mathcal{R}^{M\times (p+1)}$, and $\overline{p}=p+M+1$. The following theorem gives the joint distribution of $\widehat{r}_{wn}$.
Since the forms of $g(\cdot)$ and $u_{t}$ are not specified, the asymptotic covariance matrix of $\widehat{r}_{wn}$ can not be directly estimated from its sample counterpart. We use the similar RW method as in Section 3 to avoid this difficulty. Define $$\widehat{r}_{wn,k}^{*}=\frac{\sum_{t=k+1}^{n}w_{t}^{*}\left[ \mbox{sgn}(\varepsilon_{t}(\widehat{\theta}_{wn}^{*}))-\overline{\varepsilon}(\widehat{\theta}_{wn}^{*})\right] \left[ \mbox{sgn}(\varepsilon_{t-k}(\widehat{\theta}_{wn}^{*}))-\overline{\varepsilon}(\widehat{\theta}_{wn}^{*})\right]} {\sum_{t=1}^{n}\left[ \mbox{sgn}(\varepsilon_{t}(\widehat{\theta}_{wn}^{*}))-\overline{\varepsilon}(\widehat{\theta}_{wn}^{*})\right]^2}.$$ Let $\widehat{r}_{wn}^{*}=(\widehat{r}_{wn,1}^{*},\widehat{r}_{wn,2}^{*},\cdots,\widehat{r}_{wn,M}^{*})'$. The following theorem indicates that the distribution of $\sqrt{n}\widehat{r}_{wn}$ can be approximated by the resampling distribution of $\sqrt{n}(\widehat{r}_{wn}^{*}-\widehat{r}_{wn})$.
According to Theorems (ref) and (ref), we can approximate the asymptotic covariance matrix of $\widehat{r}_{wn}$ by the following procedure:
1. Generate $J$ replications of the i.i.d. random weights $\{w_{t}^{*}\}_{t=1}^{n}$ from the standard exponential distribution.
2. Compute $\widehat{r}_{wn}^{*}$ for $i$-th replication, and denote it as $c_{i}$.
3. Calculate the sample variance-covariance matrix of $\{c_{i}-\widehat{r}_{wn}\}_{i=1}^{J}$, denoted by $\widehat{U}_{wn}$, which provides a good approximation for the asymptotic covariance matrix of $\widehat{r}_{wn}$ for large $J$.
Based on $\widehat{U}_{wn}$, we can construct a portmanteau test statistic
to detect the adequacy of model ((ref)). If $S_{wn}(M)$ is larger than the upper-tailed critical value of $\chi^{2}_{M}$, the fitted model ((ref)) is not adequate. Otherwise, it is adequate.
By choosing the weight function $w_{t}\equiv1$, a systematical statistical inference procedure for model ((ref)) based on the usual LADE $\widehat{\theta}_{n}$ is already available in Sections 2-4. However, this choice of the weight function may lead to a deficient weighted LADE in terms of efficiency. In this section, we are interested in finding the “optimal” weight $w_{t}$ to minimize the asymptotic covariance matrix in Theorem 2.1. Given $w_{t}=g_{t}$, we consider a weighted LADE defined by
Clearly, $\widehat{\theta}_{an}$ is infeasible in practice as $g_{t}$ is not observable. However, Corollary (ref) below shows that $\widehat{\theta}_{an}$ is the desired one under three different scenarios.
Under (S1), the asymptotic covariance matrix of $\widehat{\theta}_{an}$ is adaptive to the unknown form of $g(\cdot)$. From this point of view, we follow Xu and Philips (2008) to call $\widehat{\theta}_{an}$ the infeasible adaptive LADE (ALADE). Obviously, Corollary (ref) indicates that $\widehat{\theta}_{an}$ is more efficient than $\widehat{\theta}_{n}$ under (S1), (S2) or (S3). Our finding that $g_{t}$ is the optimal weight is similar to those in Gutenbrunner and Jure\v{c}kov\'{a} (1992) and Zhao (2001) for the linear heteroscedastic regression model.
Next, we make a formal comparison of efficiency among the LADE $\widehat{\theta}_{n}$, the infeasible ALADE $\widehat{\theta}_{an}$, the LSE $\widecheck{\theta}_{n}$ in Phillips and Xu (2006), and the infeasible adaptive LSE (ALSE) $\widecheck{\theta}_{an}$ in Xu and Phillips (2008), where
We assume that $\mu_0\equiv0$ as in Xu and Phillips (2008), and that $\{u_{t}\}$ is an i.i.d. sequence with $\mbox{median}(u_{t})=0$, $Eu_{t}=0$, and $Eu_{t}^{2}=1$ to make the comparison feasible. As $\{u_{t}\}$ is an i.i.d. sequence, the moment condition that $E|u_{t}|^{2+\delta_{0}}<\infty$ for some $\delta_{0}>0$ is sufficient for the asymptotic normality of $\widehat{\theta}_{n}$ and $\widehat{\theta}_{an}$ by Corollary 2.2, and this is also the case for $\widecheck{\theta}_{n}$ and $\widecheck{\theta}_{an}$ by a slight modification of the proof in Xu and Philips (2008).
Define
where $f(\cdot)$ is the density function of $u_{t}$. Under the aforementioned conditions, result ((ref)) implies that as $n\to\infty$,
respectively; and Phillips and Xu (2006) and Xu and Phillips (2008) showed that as $n\to\infty$,
respectively. The asymptotic efficiencies of these four estimators can be formally compared by just looking at the values of $b_{i}$ under different scenarios of $u_{t}$ and $g(\cdot)$. Below, Example 1 considers the case that $g(\cdot)$ has an abrupt change in the variance, when $u_{t}$ is chosen to be an i.i.d. $\mbox{standardized }t_{3}$ ($\mbox{ST}_{3}$) or $\mbox{standardized Laplace}(0,1)$ ($\mbox{SL}(0,1)$) sequence with $E(u_{t}^{2})=1$. For the cases that $g(\cdot)$ has gradual and periodical change in the variance, one can refer to Examples 2 and 3 in Zhu (2018).
{\sc Example 1.} ({\it An abrupt change in the variance}) Let $\tau\in[0,1]$ and $g(\cdot)$ be the step function
where $s\in[0,1]$, $e_{0}>0$, and $e_{1}>0$. Under ((ref)), the variance of $\varepsilon_{t}$ is $e_{0}^{2}$ before the break point $[n\tau]$, and $e_{1}^{2}$ afterwards. Let $\delta=e_{1}/e_{0}$. Then, some algebra shows that
Fig\,(ref)(a)-(f) plot the values of all $b_{i}$ in terms of $\delta$ for $\tau=0.1, 0.5$, and 0.9, respectively. From this figure, our findings are as follows:
(i) $\widehat{\theta}_{an}$ is more efficient than $\widecheck{\theta}_{an}$ under all considered cases. Meanwhile, $\widehat{\theta}_{n}$ can be more efficient than $\widecheck{\theta}_{an}$ when the break happens earlier with a large $\delta$, the break happens later with a small $\delta$, or the break happens in the middle with $\delta$ not largely deviating from 1. All of these findings imply the efficiency advantage of the ALADE (or LADE) over the ALSE when $u_{t}$ is heavy-tailed.
(ii) $\widehat{\theta}_{an}$ (or $\widecheck{\theta}_{an}$) is always more efficient than $\widehat{\theta}_{n}$ (or $\widecheck{\theta}_{n}$) as expected, but the efficiency advantage is less significant when the break happens earlier with a large $\delta$, the break happens later with a small $\delta$, or the break happens in the middle with $\delta\approx 1$ (see, Xu and Philips (2008, p.270) for a similar phenomenon and explanation). This indicates that the ALADE with an appropriate choice of the weight function can have a pronouncing efficiency advantage over the usual LADE.
Overall, Examples 1-3 demonstrate that (i) $\widehat{\theta}_{an}$ and $\widehat{\theta}_{n}$ can be more efficient than $\widecheck{\theta}_{an}$ when $u_{t}$ is heavy-tailed; (ii) $\widehat{\theta}_{an}$ is more efficient than $\widehat{\theta}_{n}$ in all considered cases. However, $\widehat{\theta}_{an}$ is infeasible in practice, and this problem will be solved in the next section.
The objective of this section is to give a feasible ALADE, prove that this feasible ALADE calculated with estimated weights is equivalent to the infeasible ALADE $\widehat{\theta}_{an}$ asymptotically, and show that our entire statistical inference procedure in Sections 2-4 is still valid based on this feasible ALADE.
We use a similar non-parametric method as in Xu and Phillips (2008) to achieve this goal. Let $\widehat{\varepsilon}_{t}=y_{t}-Y_{t-1}'\widehat{\theta}_{n}$ be the residual from the LADE. Denote
where $k_{ti}=(\sum_{i=1}^{n}K_{ti})^{-1}K_{ti}$ with
where $K(\cdot)$ is the kernel function defined on $\mathcal{R}$, and $b>0$ is the bandwidth. Of course, the role of $\widehat{g}_{t}$ is to deputize for $g_{t}$, which is the mean of $|\varepsilon_{t}|$ based on the following assumption:
Assumption (ref) is used for the identification of $g_{t}$, and it does not rule out the conditional heteroscedasticity of $u_{t}$. For the feasible ALSE, Xu and Phillips (2008) resorted to the identification condition $E[u_{t}^{2}|\mathcal{F}_{t-1}]=1$; however, their identification condition is restrictive, since it does not allow $u_{t}$ to be conditionally heteroscedastic. Moreover, we shall mention that besides ((ref)), other methods could also be used to estimate $g_{t}$ (see, e.g., Fan and Yao, 1998; Yu and Jones, 2004), and this is a promising direction for the future study.
By using $\widehat{g}_{t}$ in ((ref)), our feasible ALADE is defined as follows:
To obtain the asymptotic theory of $\widetilde{\theta}_{an}$, we need two more assumptions.
Assumption (ref) from Shao and Yu (1996) is a technical condition for some moment inequalities of the partial sums of strong mixing sequences. An exponentially fast decaying $\alpha_{u}(k)$, which holds for large classes of processes (see, Carrasco and Chen, 2002), is sufficient for the validity of this assumption. Assumption (ref)(i) holds for commonly used kernels such as uniform, Epanechnikov, biweight, and Gaussian. Assumption (ref)(ii) requires that $b$ converges to zero at a slower rate than $n^{-1/(5+\delta_{4})}$, and this rate can be improved under some stronger conditions; see Remark (ref) below.
The following theorem shows that $\widetilde{\theta}_{an}$ has the same asymptotic variance as $\widehat{\theta}_{an}$, and hence it is the desired weighted LADE we are looking for.
Since the implementation of $\widetilde{\theta}_{an}$ depends on the bandwidth $b$, we use a similar cross-validation (CV) method as in Xu and Phillips (2008) to select $b$. That is, $b$ is chosen to be the value of $b^{*}$ which minimizes
In theory, it is unclear whether the bandwidth $b^{*}$ chosen by this CV method satisfies Assumption (ref)(ii). In practice, we can set $b=Cn^{-1/(5+2\delta_{4})}$ for a small $\delta_4$, and find $C^{*}$ by a grid search such that $b^*=C^*n^{-1/(5+2\delta_{4})}$ minimizes $\widehat{\mbox{CV}}(b)$. With $b^{*}$ computed by this data-driven method, simulation studies in Section 7 show that our $\widetilde{\theta}_{an}$ has a good finite sample performance.
Next, we consider the estimation of $g(\cdot)$ as an interest in its own.
Corollary (ref) shows that $\widehat{g}_{[n\tau]}$ converges in probability to $g(\tau)$ for interior points $\tau$ when $g(\cdot)$ is continuous, but in general to a point between $g(\tau-)$ and $g(\tau+)$ depending on the shape of the kernel function $K(\cdot)$. It is worth noting that the inconsistency of $\widehat{g}_{[n\tau]}$ at points of discontinuities has no impact on the asymptotic theory of $\widetilde{\theta}_{an}$ as shown in Theorem (ref).
Third, since the forms of $g(\cdot)$ and $u_{t}$ are not specified, we use the RW method as before to estimate the asymptotic covariance matrix of $\widetilde{\theta}_{an}$. Define
The following theorem indicates that the distribution of $\sqrt{n}(\widetilde{\theta}_{an}-\theta_{0})$ can be approximated by the resampling distribution of $\sqrt{n}(\widetilde{\theta}_{an}^{*}-\widetilde{\theta}_{an})$.
According to Theorems (ref) and (ref), we can approximate the asymptotic covariance matrix of $\widetilde{\theta}_{an}$ by $\widetilde{V}_{an}$, where $\widetilde{V}_{an}$ is calculated in the same way as $\widehat{V}_{wn}$ with $\widehat{\theta}_{wn}$ and $\widehat{\theta}_{wn}^{*}$ replaced by $\widetilde{\theta}_{an}$ and $\widetilde{\theta}_{an}^{*}$, respectively. Then, we can construct another Wald test statistic $W_{an}$ based on $\widetilde{V}_{an}$, to test the linear null hypothesis $H_{0}$ in ((ref)), where
If $W_{an}$ is larger than the upper-tailed critical value of $\chi^{2}_{s}$, the null hypothesis $H_{0}$ is rejected. Otherwise, $H_{0}$ is not rejected.
In the end, we consider the sign-based portmanteau test based on $\widetilde{\theta}_{an}$. Define $\widetilde{r}_{an,k}$ and $\widetilde{r}_{an,k}^{*}$ in the same way as $\widehat{r}_{wn,k}$ and $\widehat{r}_{wn,k}^{*}$, with $\widehat{\theta}_{wn}$ and $\widehat{\theta}_{wn}^{*}$ replaced by $\widetilde{\theta}_{an}$ and $\widetilde{\theta}_{an}^{*}$, respectively. Let $\widetilde{r}_{an}=(\widetilde{r}_{an,1},\widetilde{r}_{an,2},\cdots,\widetilde{r}_{an,M})'$ and $\widetilde{r}_{an}^{*}=(\widetilde{r}_{an,1}^{*},\widetilde{r}_{an,2}^{*},\cdots,\widetilde{r}_{an,M}^{*})'$. The following theorem indicates that the distribution of $\sqrt{n}\widetilde{r}_{an}$ can be approximated by the resampling distribution of $\sqrt{n}(\widetilde{r}_{an}^{*}-\widetilde{r}_{an})$.
According to Theorem (ref), we can approximate the asymptotic covariance matrix of $\widetilde{r}_{an}$ by $\widetilde{U}_{an}$, where $\widetilde{U}_{an}$ is calculated in the same way as $\widehat{U}_{wn}$ with $\widehat{\theta}_{wn}$ and $\widehat{\theta}_{wn}^{*}$ replaced by $\widetilde{\theta}_{an}$ and $\widetilde{\theta}_{an}^{*}$, respectively. Then, we can construct another sign-based portmanteau test statistic $S_{an}(M)$, based on $\widetilde{U}_{an}$, to detect the adequacy of model ((ref)), where
If $S_{an}(M)$ is larger than the upper-tailed critical value of $\chi^{2}_{M}$, the fitted model ((ref)) is not adequate. Otherwise, it is adequate.
In this section, we first assess the performance of the LADE $\widehat{\theta}_{n}$, the feasible ALADE $\widetilde{\theta}_{an}$, and the corresponding RW approach in finite samples. We generate 1000 replications of sample size $n=100$ and $200$ from the following model:
where $\theta_0=0.5$, the error generating process for $\varepsilon_{t}$ is designed as follows:
and $g_{t}=g(t/n)$ with $g(\cdot)$ satisfying one of the following structures:
Here, we use $(\alpha_{\dag},\beta_{\dag})=(0, 0)$ or $(0.1, 0.8)$, choose $\eta_{t}\sim$ i.i.d. $\mbox{SL}(0, 1)$, $\mbox{N}(0, 1)$ or $\mbox{ST}_{3}$ in model ((ref)), and set $\delta\in\{0.2, 5\}$ in models ((ref))-((ref)), $\delta\in\{2\pi, 4\pi\}$ in model ((ref)). Although $u_{t}$ does not satisfy $E|u_{t}|=1$ in the aforementioned set-up, the estimation property of $\theta_0$ is unaffected by noting that $\varepsilon_{t}=g^{\dag}_{t}u_{t}^{\dag}$ with $g^{\dag}_{t}=g_{t}E|u_{t}|$ and $u_{t}^{\dag}=u_{t}/E|u_{t}|$. To save space, we only report the results for model ((ref)) in what follows, and the results for models ((ref))-((ref)) are similar and can be found in Zhu (2018).
Table (ref) reports the sample root mean squared error (SE) and the average bootstrapped sample root mean squared error (AE) of $\widehat{\theta}_{n}$ and $\widetilde{\theta}_{an}$, based on 1000 replications. In all calculations, $\widetilde{\theta}_{an}$ is obtained by using Gaussian kernel and the CV method in ((ref)) to choose $b$; and the bootstrapped sample root mean squared errors for $\widehat{\theta}_{n}$ and $\widetilde{\theta}_{an}$ are computed by using the RW method with the bootstrap sample size $J=500$. From Table (ref), we can find that (i) the disparity between SE and AE in each case is small, indicating that our RW approach is reliable; (ii) the SE of $\widetilde{\theta}_{an}$ is always smaller than the one of $\widehat{\theta}_{n}$ as expected; (iv) the SE of each estimator becomes small as the sample size $n$ increases; (v) the SE of each estimator in the case of heavy-tailed $\eta_{t}$ (i.e., $\eta_{t}\sim \mbox{SL}(0, 1)$ or $\mbox{ST}_{3}$) is smaller than the corresponding one in the case of light-tailed $\eta_{t}$ (i.e., $\eta_{t}\sim \mbox{N}(0, 1)$). Besides SE, we also compute the sample median and sample median of absolute deviation of $\widehat{\theta}_{n}$ and $\widetilde{\theta}_{an}$ to see their variability, and the details are relegated to Zhu (2018).
Furthermore, we examine the performance of our CV method by calculating the value of the following ratio:
where SD stands for the sample standard deviation based on 1000 replications. If our CV method works well, the value of $R_{1}$ should be close to one. Meanwhile, we compare the efficiency of $\widehat{\theta}_{n}$, $\widetilde{\theta}_{an}$ and the LSE $\widecheck{\theta}_{n}$ with the infeasible ALSE $\widecheck{\theta}_{an}$ by looking at the values of the following three ratios:
Clearly, $\widehat{\theta}_{n}$, $\widetilde{\theta}_{an}$ and $\widecheck{\theta}_{n}$ are more efficient than $\widecheck{\theta}_{an}$ when the values of $R_{2}$, $R_{3}$ and $R_{4}$ are smaller than one, respectively. Also, the most efficient estimator among $\widehat{\theta}_{n}$, $\widetilde{\theta}_{an}$ and $\widecheck{\theta}_{n}$ is related to the smallest value of $R_{2}$, $R_{3}$ and $R_{4}$.
Table (ref) reports the values of these four ratios. From this table, we first find that the value of $R_{1}$ is close to one, and hence it implies that our CV method works satisfactorily. Second, we can see that $\widetilde{\theta}_{an}$ is more efficient than $\widecheck{\theta}_{an}$ (with values of $R_{3}$ less than one) in general when $\eta_{t}$ is heavy-tailed, and this efficiency advantage in the case of $(\alpha_{\dag},\beta_{\dag})=(0, 0)$ is less substantial than the one in the case of $(\alpha_{\dag},\beta_{\dag})=(0.1, 0.8)$; on the other hand, as expected, $\widecheck{\theta}_{an}$ is generally more efficient than $\widetilde{\theta}_{an}$ when $\eta_{t}$ is light-tailed. For $\widehat{\theta}_{n}$, it can still be more efficient than $\widecheck{\theta}_{an}$ (with values of $R_{2}$ less than one) in most of examined cases especially when $\eta_{t}\sim\mbox{SL}(0, 1)$, and this is not the case when $\eta_{t}\sim\mbox{N}(0, 1)$. Third, we note that $\widecheck{\theta}_{n}$ is always less efficient than $\widetilde{\theta}_{an}$ and $\widecheck{\theta}_{an}$, and it is more efficient than $\widehat{\theta}_{n}$ only when $\eta_{t}\sim\mbox{N}(0, 1)$.
In summary, our numerical results in Tables (ref)-(ref) show that the RW method can provide reliable estimators of the standard deviations for both $\widehat{\theta}_{n}$ and $\widetilde{\theta}_{an}$, and $\widetilde{\theta}_{an}$ calculated by the CV method shall be recommended for the heavy-tailed $\eta_{t}$.
Next, we examine the performance of the Wald tests $W_{wn}$ and $W_{an}$, the sign-based portmanteau tests $S_{wn}(M)$ and $S_{an}(M)$, and the corresponding RW approach in finite samples. We generate 1000 replications of sample size $n=100$ and $200$ from the following model:
where $(\phi_{10},\phi_{20})=(0.5, \kappa)$ with $\kappa=0, 0.2$ or $0.4$, and the error generating process for $\varepsilon_{t}$ is given by model ((ref)). We fit each replication by an AR($2$) model with the LADE (or the feasible ALADE) of $(\phi_{10},\phi_{20})$, and then apply $W_{wn}$ in ((ref)) (or $W_{an}$ in ((ref))) to test the hypothesis of $\phi_{20}=0$ (i.e., $\Gamma=(0,1)$, $\theta_0=(\phi_{10},\phi_{20})'$ and $r=0$ in ((ref))). Furthermore, we fit each replication by an AR($1$) model with the LADE (or the feasible ALADE) of $\phi_{10}$, and then apply $S_{wn}(M)$ in ((ref)) (or $S_{an}(M)$ in ((ref))) to check whether this fitted AR(1) model is adequate. In all cases, we set the significance level $\underline{\alpha}=0.05$ and $M=6$. Based on 1000 replications, the empirical power of all the tests are reported in Table (ref) for the case that $\eta_{t}\sim\mbox{SL}(0, 1)$, and their sizes correspond to the results for the case that $\kappa=0$. Since the results for other two distributions of $\eta_{t}$ are similar, they are not provided here for saving the space. From Table (ref), it is clear that the sizes of the two Wald tests are close to their nominal ones, and the sizes of the two portmanteau tests are conservative especially when $n$ is small. For the power of all tests, we find that all the power becomes large as the value of $n$ or $\kappa$ increases; the power of $W_{an}$ and $S_{an}$ based on the feasible ALADE is larger than that of $W_{wn}$ and $S_{wn}$ based on the LADE, respectively, and this power advantage is more distinct for the Wald test. Overall, all tests based on the RW approach have a good performance especially when the sample size is large, and we recommend $W_{an}$ and $S_{an}$ based on the feasible ALADE in practice.
This paper provides an entire statistical inference procedure for the AR($p$) model under (conditional) heteroscedasticity of unknown form by establishing the asymptotic normality of the weighted LADE, developing the RW method to estimate its asymptotic covariance matrix, and constructing a portmanteau test for the model diagnostic checking. This entire procedure can either be based on the usual LADE or the feasible ALADE, and as demonstrated by the simulation results, the ALADE is the better choice in terms of the estimation efficiency and testing power. Applications to three U.S. economic data sets are considered in the supplementary material, and the results indicate that after the financial crisis in 2007-2008, the monetary policies may become less prudent, and their control on the PPI and CPI seems to be weaker. The methodology developed in this paper provides a new way to handle the heteroscedastic time series data and shall have a large applicable scope in practice. Extensions to other time series models (e.g., autoregressive moving average models) and estimation methods (e.g., M- and quantile estimations) could be potentially interesting future works.
\setcounter{equation}{0}