Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
70,515 characters · 14 sections · 44 citation commands
Testing for threshold effects in the TARMA framework
\affil[1]{ Department of Statistics and Actuarial Science, University of Iowa, Iowa City, USA} \affil[2]{Department of Statistical Sciences, University of Bologna, Italy} \affil[3]{School of Mathematical Science, University of Electronic Science and Technology, Chengdu China} \affil[4]{Center for Statistical Science, Tsinghua University, China} \affil[5]{London School of Economics and Political Science, U.K.}
Threshold autoregressive models have gained popularity in Economics, Biology, and many other fields, Ton90,Cha17b,Cha09,Ton11,Han11. In particular, TARMA models, introduced in Ton78 and Ton80, are non-linear models with a regime-switching mechanism specifying an ARMA sub-model in each regime. They include two particular models of independent interest: the threshold autoregressive (TAR) model and the threshold moving-average (TMA) model. By combining both the TAR model and the TMA model, TARMA models are parsimonious and yet rich models for non-linear time series analysis (see e.g. Gor20,Gor21). Li11b developed the theory for least squares estimation of parameters of the general TARMA$(p,q)$ model by assuming stationarity and ergodicity. Nevertheless, the conditions for ergodicity derived by Lin99 are quite restrictive. A full characterization of the long-run probabilistic behaviour of TARMA models was not available until Cha19, which derived the necessary and sufficient conditions for the ergodicity of the first-order TARMA model. Moreover, they provided a complete parametric classification of the first-order TARMA into regions where it is (geometrically) ergodic, null recurrent and transient. Considerable efforts have been produced to test whether a threshold model provides a better fit with respect to its linear counterpart. Most contributions focus on AR-type models. For instance, Pet86 developed a portmanteau test based on cumulative sums of standardized residuals from an autoregressive fit. Tsa98 studied a variation of such test. Luu88 proposed a Lagrange Multiplier test for linearity against a large class of non-linear models that includes the TAR specification. A Lagrange Multiplier test was also developed in Won97, Won00 for TAR models with conditional heteroscedasticity. Quasi-likelihood ratio tests were studied in Cha90a,Cha90b,Cha91 up to the recent test for threshold diffusion of Su17. For a review see also Ton11. The framework of threshold models that includes a moving-average component has been under-investigated probably due to the mathematical difficulties that arise when the moving-average component is incorporated in a non-linear setting. However, since data are almost always affected by measurement error, the threshold ARMA framework is more appropriate than a pure autoregressive approach. Indeed, it is known that a AR process plus measurement error becomes a ARMA process. Likewise, it can be proved that a TAR process of order $p$ corrupted with additive measurement noise may be approximated by a TARMA model of order $(p, p)$. The adoption of the TARMA framework is not a minor point since a high autoregressive order may be needed to approximate the moving-average component, at the expense of loss in power of the test. In the framework of MA-type models, Lin05 investigated a quasi-likelihood ratio test for the MA model against its threshold extension. They proved that, under the null hypothesis, the test statistic converges weakly to a functional of the centered Gaussian process. Their results were extended in Li08 to the case with GARCH errors. More recently, Li11 developed a quasi-likelihood ratio statistic to test the presence of thresholds in ARMA processes. They use a stochastic permutation device to build the distribution of the statistic under the null hypothesis. In this paper we extend the work of Cha90a and Lin05 and propose supremum Lagrange multiplier test statistics (supLM) to determine whether a TARMA model fits a stationary time series significantly better than an ARMA model. One of the main advantages of using a Lagrange Multiplier approach over likelihood-ratio tests is that it does not need estimating the model under the alternative hypothesis. We prove that both under the null hypothesis and contiguous local alternatives the asymptotic distribution of the test statistics reduces to the same functional of a Gaussian process that is centered under the null and non centered under the alternative. The results extend the work of Lin05 on the weak convergence of linear marked empirical processes with infinitely many markers to the case where the underlying process is an ARMA$(p,q)$. Moreover, we prove the consistency of our tests and show that they have non-trivial power against local alternatives. In order to test the ARMA$(p,q)$ specification against its TARMA extension we propose two supLM statistics: in the first one, denoted by $\operatorname{sLM}$, only the autoregressive part is tested for threshold non-linearity whereas in the second statistic, denoted by $\operatorname{sLM^{\star}}$, both the autoregressive and the moving-average part are tested. As it will be clear, the two statistics are different; in particular, the $\operatorname{sLM^{\star}}$ statistic does not reduce to the $\operatorname{sLM}$ when the moving-average part is either absent or does not change across regimes. This is reflected on the different finite sample behaviour of the tests. We explore systematically the performance of our supLM tests and compare them with the quasi-likelihood ratio test of Li11 (qLR from now on): the extensive simulation study shows clearly that our tests have better size and power while enjoying a much lower computational burden. Furthermore, the two supLM tests are robust against model mis-specification and the performance of the tests is not adversely affected if the order of the ARMA process is unknown and is selected through the Hannan-Rissanen procedure. Lastly, we apply our test to the time series of standardized tree-ring growth indexes. We show that the TARMA(1,1) specification can provide a better fit with respect to the accepted ARMA model. We believe that the TARMA framework can lead to a better understanding of the tree-ring dynamics and lead to novel directions of research where the econometric approach is properly adopted in climate studies. The rest of the paper is organized as follows: in Section (ref) we present our setting and the tests; in Section (ref) we derive the distributions under the null hypothesis and tabulate the empirical quantiles. In Section (ref) we derive the asymptotic distribution of the statistics under local contiguous alternatives and prove consistency of the tests. Section (ref) contains a Monte Carlo study to assess the finite sample performance of our proposals. We also investigate the behaviour of the tests under model mis-specification and when the order of the tested model is unknown. In Section (ref) we apply our tests to a tree-ring time series whereas some discussion and the conclusions are reported in Section (ref). All the proofs are detailed in the Supplementary Material, that also contains additional results regarding both the simulation study and the tree-ring data analysis.
Let the time series $\{X_t:t=0,\pm 1,\pm2,\dots\}$ follow the threshold autoregressive moving-average model defined by the difference equation:
In the case where the moving-average parameters are fixed across regimes, the sum involving $\Psi_{2s}$, $s=1,\dots,q$ is absent. The innovations $\{\varepsilon_t\}$ are independent and identically distributed (iid) random variables such that, for each $t$, $\varepsilon_t$ has zero mean, finite variance $\sigma^2$ and is independent of $X_{t-1}$, $X_{t-2}$, \dots . {Note that the iid assumption can be relaxed to a stationary ergodic martingale difference sequence with respect to the $\sigma$-algebra generating by $\varepsilon_s$ with $s<t$. In such a case the proofs do not change.} The positive integers $p$ and $q$ are the autoregressive and moving-average orders, respectively; $d$ is the delay parameter that takes positive integer values. We assume $p,q,d$ to be known. Moreover, $I(\cdot)$ is the indicator function and $r\in\mathds{R}$ is the threshold parameter. For notational convenience, we abbreviate $I(X_t\leq r)$ by $I_r(X_t)$. All the results are derived conditionally upon the $p$ initial values of $\{X_t\}$. Let
and $\boldsymbol{\Psi}$ be the vector containing the parameters that are tested. The true parameters of the model are $ \boldsymbol\eta=\left(\boldsymbol{\zeta}^\intercal,\sigma^2,\boldsymbol\Psi^\intercal\right)^\intercal $, where
We test whether a threshold ARMA$(p,q)$ model provides a significantly better fit than the linear ARMA$(p,q)$ model. To this end, we develop two Lagrange multiplier test statistics for the hypothesis $$
$$ where $\boldsymbol 0$ is the vector with all zeroes. The statistic for testing the threshold effect in the AR parameters is denoted as $\operatorname{sLM}$, whereas $\operatorname{sLM^{\star}}$ is the statistic for the general test where both the AR and the MA parameters change across regimes. Under $H_0$ the process follows a linear ARMA$(p,q)$ model:
To derive the asymptotic features of the test, we assume the model to be ergodic and invertible both under the null and the alternative hypothesis. Let the autoregressive and moving-average polynomials be defined as follows:
Assumption A1 below ensures the ergodicity and invertibility of model ((ref)) and also avoids certain degeneracy. For more details see Cha10 and Cha19.
From now on we fully develop the theory for the general statistic $\operatorname{sLM^{\star}}$. Unless otherwise specified, the results hold also for the statistic $\operatorname{sLM}$. Suppose we observe $X_1,\dots,X_n$. We develop the Lagrange multiplier test based on the Gaussian likelihood conditional on the initial values $X_0,X_{-1},\ldots,X_{-p+1}$:
where, by an abuse of notation, we set
and $\varepsilon_0,\varepsilon_{-1},\ldots,\varepsilon_{-q+1}$ are set to be zero. Clearly, $\varepsilon_t$ is a function of $\boldsymbol\eta$ and $r$, but we omit the arguments for simplicity. As in Eq ((ref)) for the $\operatorname{sLM}$ test, the sum involving $\Psi_{2s}$, $s=1,\dots,q$ is absent. {Let $\partial\ell/\partial\boldsymbol\eta$ be the score vector, whose components are
and $\partial\ell/\partial\boldsymbol{\Psi}$ be the derivatives of the log-likelihood with respect to $\boldsymbol{\Psi}$.} Moreover, let
and
Finally
The TARMA$(p,q)$ model under the null hypothesis can be estimated by using the method of the maximum likelihood. Let $\frac{\partial \hat{\ell} }{\partial \boldsymbol\Psi}(r)$ and $\hat{\mathcal{I}}_{n}(r)$ be equal to $\frac{\partial \ell }{\partial \boldsymbol\Psi}(r)$ and $\mathcal{I}_n(r)$, respectively, evaluated at the maximum likelihood estimates for the ARMA part and with $\boldsymbol{\Psi}=\boldsymbol 0$. Under the null hypothesis, the threshold parameter $r$ is absent thereby the standard asymptotic theory is not applicable. To cope with this issue, we firstly develop the Lagrange multiplier test statistic as a function of $r$ ranging in a set $\mathcal{R}$. Then, for all the values $r \in \mathcal{R}$, we compute the test statistic and, finally, we take the overall test statistic as the supremum on $\mathcal{R}$. We set $\mathcal{R}=[r_L,r_U]$, $r_L$ and $r_U$ being some percentiles of the data. This approach has become widely used in the literature of tests involving threshold models, see, for instance, Cha90a, Lin05, Li11 and Cha20. Our test statistic is
In this section we derive the asymptotic distribution of $T_n$ under the null hypothesis that $\{X_t\}$ follows an ARMA$(p,q)$ process. Unless stated otherwise, all the expectations are taken under the true probability distribution for which $H_0$ holds. Also, $o_p(1)$ denotes the convergence in probability to zero as $n$ increases and $\|\cdot\|$ is the $L^2$ matrix norm (the Frobenius' norm, i.e. $\|A\|=\sqrt{\sum_{i=1}^{n}\sum_{j=1}^{m}|a_{ij}|^2}$, where $A$ is a $n\times m$ matrix). Moreover, let $\mathcal{D}_\mathds{R}(a,b)$, $a<b$ be the space of functions from $(a,b)$ to $\mathds{R}$ that are right continuous with left-hand limits. $\mathcal{D}_\mathds{R}(a,b)$ is equipped with the topology of uniform convergence on compact sets, see Bil68 for more details. In the following two lemmas, we rewrite $\partial\varepsilon_t/\partial\boldsymbol\phi$, $\partial\varepsilon_t/\partial\boldsymbol\theta$ and $\partial\varepsilon_t/\partial\boldsymbol\Psi$ as functions of the roots of the characteristic moving-average polynomial $\theta(\cdot)$ and provide a uniform approximation of the matrix $\mathcal{I}_n(r)$.
Now, define $ \nabla_n(r)=\left(\nabla^\intercal_{n,1},\nabla^\intercal_{n,2}(r)\right)^\intercal$, where
Moreover, define
where the $\Lambda_{ij}(r)$ have same dimension as the $\mathcal{I}_{n,ij}(r)$ of Eq. ((ref)) and
with the $\alpha$'s defined as in Lemma (ref) and where, for the $\operatorname{sLM}$ statistic, the components from position $p+2$ to $p+q+2$ and the last $q$ of $D(r)$ are absent. In the following Lemma, we show some properties of the $\nabla$'s and $\Lambda$'s under the null hypothesis and Assumption A1.
Note that in the above proposition, $\left(\nabla_{n,2}(r)-\Lambda_{21}(r)\Lambda_{11}^{-1}\nabla_{n,1}\right)$ is a linear marked empirical process with infinitely many markers. Similarly to Lin05,Li11, we rely on Assumption A2:
In the following theorem we derive the null asymptotic distribution of the Lagrange Multiplier test statistic $T_n$.
In Table (ref) we tabulate the empirical quantiles of the null asymptotic distribution of our supLM statistics at levels 90%, 95%, 99% and 99.9% for autoregressive orders from 1 to 4 and moving-average orders from 1 to 2. The threshold is searched between the 25th and the 75th percentiles of the sample distribution. For each order, the results have been obtained from 10000 simulated series of length 1000 and are presented in Table (ref). The quantiles of the asymptotic distribution of the $\operatorname{sLM}$ statistic do not depend upon the moving-average parameters and are in good agreement with those of Cha91, table 1, that refer to testing the AR against the TAR model (In that case, the length of the series was 200 and the number of replications 1000). The rightmost part of Table (ref) contains the quantiles for the $\operatorname{sLM^{\star}}$ statistic. Indeed, even when the moving-average parameters do not change across regimes, the $\operatorname{sLM^{\star}}$ statistic does not reduce to the $\operatorname{sLM}$ statistic. Furthermore, the asymptotic distribution of the $\operatorname{sLM}$ statistic is equivalent to that of Cha91 since the vector $\nabla_{n}(r)$ of Eq. ((ref)) does not contain the partial derivatives w.r.t. to the moving-average part. Notably, the asymptotic behaviour of the two statistics depends only upon the dimension of the parameter vector $\boldsymbol{\Psi}$, irrespectively of its components being either autoregressive or moving-average, see the Supplement for more details and an assessment of the similarity of the supLM statistics Gor21SM. Finally, note that the tabulated values match those of table 1 of And03 where $\pi_0=0.25$.
In this section, we derive the asymptotic distribution of $T_n$ under a sequence of local alternatives and prove the consistency of the associated tests. For each $n$, the null hypothesis $H_{0,n}$ states that $\{X_t, t=0,\dots,n\}$ follows the model: $$ X_t=\phi_{10}+\sum_{k=1}^{p}\phi_{1k}X_{t-k}-\sum_{s=1}^{q}\theta_{1s}\varepsilon_{t-s}+\varepsilon_t. $$ The alternative hypothesis $H_{1,n}$ states that $\{X_t, t=0,\dots,n\}$ follows the model:
where $\mathbf{h}=\left(h_{10},h_{11},\cdots,h_{1p},h_{21},\cdots,h_{2q}\right)^\intercal\in\mathds{R}^{p+q+1}$ is a fixed vector and $r_0$ is a fixed scalar. In the case of the $\operatorname{sLM}$ statistic, the rightmost summation term within square brackets is absent so that $\mathbf{h}=\left(h_{10},h_{11},\cdots,h_{1p}\right)^\intercal\in\mathds{R}^{p+1}$. Let $P_{0,n}$ and $P_{1,n}$ be the probability measure of $\left(X_0,X_1,\dots, X_n\right)$ under $H_{0,n}$ and $H_{1,n}$, respectively. In the following proposition we prove the asymptotic normality of the loglikelihood ratio and the contiguity of $P_{1,n}$ to $P_{0,n}$. As in C5 of Chan (1990), 3.1 of Ling and Tong (2005), A5 of Li and Li (2011), we assume
Next, we derive the asymptotic distribution of $T_n$ under a sequence of local alternatives $H_{1,n}$:
Finally, we prove the consistency of our tests.
Note that Proposition (ref) and Theorem (ref) can be proved for the $\operatorname{sLM}$ statistic without assumption A3. The two proofs are reported separately in the Supplementary Material.
In this section, we investigate the finite sample performance of our supLM tests ($\operatorname{sLM}$, $\operatorname{sLM^{\star}}$) and compare them with the quasi-likelihood ratio test developed in Li11 (qLR). Hereafter $\varepsilon_t$, $t=1,\dots,n$ is generated from a standard Gaussian white noise, the length of the series is $n=100,200,500$, the nominal size is $\alpha=0.05$ and the number of Monte Carlo replications is 1000. For our tests we have used the tabulated values of Table (ref). For the qLR test we have used $B=1000$ resamples. In Section (ref), we study the size of the tests; Section (ref) shows the power of the tests in scenarios where $i)$ only the autoregressive parameters change across regimes, $ii)$ only the moving-average parameters change across regimes and $iii)$ both the autoregressive and the moving-average parameters change across regimes. Then, we assess the behaviour of the tests in presence of model misspecification (Section (ref)) and when the order of the ARMA process tested is treated as unknown and is selected by means of the Hannan-Rissanen method (Section (ref)).
We have generated time series from 25 different simulation settings of the following ARMA$(1,1)$ model:
where $\phi_{11}=0,\pm 0.3, \pm 0.6$ and $\theta_{11}=0,\pm 0.4, \pm 0.8$. Table (ref) shows the rejection percentages for the three sample sizes in use. Note that the case $\theta_{11} = 0$ corresponds to testing an AR versus a TAR model. For $n=100$ the qLR test is biased in almost all settings and reaches 53% of false rejections for the case $\theta_{11}=\phi_{11}=0$. The size of the $\operatorname{sLM}$ test is always acceptable as it is slightly greater than 10% only in three cases and its maximum value is 15.5%. The size of the $\operatorname{sLM^{\star}}$ is slightly more biased than that of the $\operatorname{sLM}$ test. When $n=200$ the bias of the $\operatorname{sLM}$ test reduces further and its size is not far from the nominal 5% in most situations. This also holds for the $\operatorname{sLM^{\star}}$ test, except for the case $\theta_{11}=0.8$ where the size is still around 10%. This is not the case for the qLR test whose size is close to 40% in three simulation settings. When $n=500$ both our supLM tests achieve a size which is close to the nominal 5% level, whereas the qLR test is still severely biased for some cases when near cancellation occurs, particularly when $\theta_{11}=\phi_{11}=0$ . One may argue that it is not appropriate to apply these tests to a realization of a white noise process and it would be more sensible to apply other kinds of tests in first place. Nevertheless, it may occur that a threshold process is mistaken for a white noise if the piecewise linear structure is such that the parameters of a linear ARMA fit result non-significant. Indeed, some sort of mis-specification is always present and this aspect will be investigated in Section (ref).
In this section we study the power of the supLM tests and highlight the differences between them. Note that the parameter vector $\boldsymbol{\Psi}$ (see Eq.(ref)) represents the departure from the null hypothesis and in all the simulations below we take sequences of increasing distance from $H_0$ in all of its components. We simulate from three different TARMA$(1,1)$ models where $i)$ only the autoregressive parameters change across regimes, $ii)$ only the moving-average parameters change across regimes, and $iii)$ both the autoregressive and the moving-average parameters change across regimes. As for the first case, we simulate from the following model:
where $\Psi_{10}$,$\Psi_{11}$ are as in Table (ref), first two columns. We combine these with $\theta_{11}=0,\pm 0.4, \pm 0.8$ as to obtain 20 different parameter settings. Table (ref) presents the size-corrected power of the tests (in percentage). Clearly, the supLM tests outperform the qLR test uniformly (except for a single case). Note that the power depends upon the true value of $\theta_{11}$ and the case $\theta_{11}=0$ seems to impinge most negatively. In such instance, the qLR test has no power even for $n=200$ whereas the supLM tests show power loss due to the size correction only for $n=100$. Overall, starting from $n=200$ both the supLM tests present a good power in almost every situation. As expected, the $\operatorname{sLM}$ test is slightly superior to the $\operatorname{sLM^{\star}}$ test since the moving-average parameter is fixed across regimes.
The case $ii)$ where only the moving-average parameter changes across regimes is studied by simulating from the following model:
where $\theta_{11}=-0.6$, whereas $\phi_{10}$, $\phi_{11}$ and $\Psi_{21}$ are as in Table (ref), that shows the rejection percentages. In this case the behaviour of the tests depends on different factors. When the departure from the null hypothesis is mild, the qLR test has an advantage over supLM tests. The situation is reversed when $\Psi_{21}$ is large (e.g. $\Psi_{21}=1.2$): in such case both supLM tests are more powerful than the qLR test. When $n = 500$ and $\phi_0=0.6$, $\phi_1=0.7$ the $\operatorname{sLM^{\star}}$ test is always more powerful than the qLR test, whereas in the remaining cases there is not a clear winner and the results are comparable.
As concerns case $iii)$ we simulate from the following model:
where $\phi_{10}=-0.5$, $\phi_{11}=-0.5$, $\theta_{11}=0.5$ and $\Psi_{10}$, $\Psi_{11}$ and $\Psi_{21}$ are as in Table (ref), first three columns. When $n=100$ the $\operatorname{sLM^{\star}}$ and the qLR tests are comparable and more powerful than the $\operatorname{sLM}$ test. However when $n=200$ the $\operatorname{sLM^{\star}}$ test is more powerful than the qLR and $\operatorname{sLM}$ tests, whose power is comparable. When $n=500$ the $\operatorname{sLM^{\star}}$ test is always the most powerful of the three. Also, on average the $\operatorname{sLM}$ test has 6% less power than the qLR test. Note that, in principle, this setting is more favourable to the $\operatorname{sLM^{\star}}$ and qLR tests since, even if the sequence of departures from $H_0$ is monotonically increasing with respect to all the components, the rate is faster along the moving-average component $\Psi_{21}$ and slower on the autoregressive part $\Psi_{10}$ and $\Psi_{11}$ (see the first three columns of Table (ref)). When all the components are distant from the null hypothesis, then, the power of the tests are either comparable or the $\operatorname{sLM}$ test is even more powerful (results not shown here).
The results for higher order TARMA models confirm the above conclusions and some of these are reported in the Supplement Gor21SM.
In this section we assess the impact of model mis-specification upon the performance of the tests. The sources of mis-specification can be diverse: as above, we focus on testing the ARMA$(1,1)$ versus the TARMA$(1,1)$ specification but the data generating process is not encompassed within the two models. Loosely speaking, we are investigating the capability of the test to detect general departures from linearity beyond the direct comparison of two specific models. Ideally, if the data generating process falls within the class of linear processes we would want the test not to reject the null hypothesis. Likewise, if the data generating process is non-linear in some of its components, then we expect the test to reject the null hypothesis. In Table (ref) we show the list of linear and non-linear data generating processes used. The seven linear processes are not ARMA$(1,1)$ since they contain higher-order autoregressive or moving-average terms. In the second part of the table we show non-linear processes that cannot be encompassed within the two-regime TARMA$(1,1)$ specification. In particular, we simulate from TAR models with both higher autoregressive order and more than two regimes. Lastly, we generate from six non-linear models that do not belong to the TARMA class, such as non-linear moving-average (NLMA), bilinear (BIL), exponential autoregressive (EXPAR) and deterministic chaos (NLAR).
The rejection percentages are reported in Table (ref). As discussed above, the first seven rows should reflect the empirical size at nominal level $5\%$ under mis-specification. Consistently with the results of Table (ref), the $\operatorname{sLM}$ test is well behaved in terms of size even for $n=100$ whereas the $\operatorname{sLM^{\star}}$ test has acceptable size starting from $n=200$. The qLR test presents acceptable size just for $n=500$, except for the ARMA21.1 case with a 27.8% of false rejections. The lower panel of Table (ref) shows the rejection percentages for the 9 non-linear processes that do not belong to the TARMA$(1,1)$ class. Here, the two supLM tests show higher power in almost every situation even for $n=100$ and with a consistent increase over the sample size. The qLR test has good power in several instances but for the TAR3, the 3TAR1 and the NLAR processes the power decreases as the sample size increases. The causes of this phenomenon are not clear and deserve further investigation, we offer some discussion in the Conclusions section. The general conclusions that can be drawn are that the supLM tests are robust with respect to modelling mis-specifications, both in terms of size and power. The same cannot be said for the qLR test that, in some cases presents oversize and power loss even for $n=500$.
Testing the ARMA against the TARMA specification requires selecting a specific order beforehand. In this section we show that there is virtually no loss incurred in using our supLM tests when no previous information on the order is available, provided a proper model selection procedure is adopted. We advocate the use of the consistent ARMA order selection proposed in Han82 (see also Cho92). In Table (ref) we present the empirical size of the supLM tests at nominal level 5% for 6 parameterizations of an ARMA$(2,2)$ process (see the first four columns). The upper panel of the table refers to the $\operatorname{sLM}$ test whereas the lower panel refers to the $\operatorname{sLM^{\star}}$ test. In both cases, the subscript HR indicates that the order of the ARMA process has been selected through the Hannan-Rissanen procedure; the true order has been used otherwise. The results indicate that not only the model selection step does not produce a size bias, but it also seems to reduce it in some instances. The impact of model selection on the power of the tests is shown in Table (ref) where we simulate 12 parameter settings of the TARMA$(1,1)$ model of Eq. ((ref)). Clearly, the power loss produced by using model selection is minimal and lies within 2% for the $\operatorname{sLM}$ test and 3% for the $\operatorname{sLM^{\star}}$ test. Finally, also in presence of mis-specification the results of the HR model selection poses no problems. The results are shown in the Supplementary material.
The Monte Carlo study has shown that supLM tests have good finite sample properties. They are also robust against model mis-specification and their performance is not affected if the order of the tested process is unknown, provided a consistent order selection procedure is used. The tests do not suffer from some of the drawbacks that affect the quasi-likelihood ratio test. The reasons can be diverse. First and foremost, supLM statistics only require fitting an ARMA model whereas the qLR test is bound to estimating a full TARMA model. We remind that there are no theoretical results regarding the sampling properties of the maximum likelihood estimators for the parameters of a TARMA model. Moreover, the qLR test by Li11 uses a representation in terms of a quadratic form but it is only valid asymptotically and this can impinge on the rate of convergence of the statistic towards its asymptotic distribution. The statistic $\operatorname{sLM}$, tests the ARMA($p,q$) against the TARMA($p,q$) model when only the $p$ autoregressive parameters change across regimes. However, the results show that such test has power also when only the moving-average parameters change. This could be ascribed to the duality between MA and AR processes and indicates a capability to detect general departures from linearity as also witnessed by the results of Section (ref). The results also show that, as expected, when either only the $q$ MA parameters or all the $p+q$ parameters of the ARMA model change across regimes, then the $\operatorname{sLM^{\star}}$ test is more powerful. The price to be paid for this superior power is the increased size bias in small samples. In general, we expect the two tests to behave similarly but in case of small samples the $\operatorname{sLM}$ statistic is recommended and can be used in conjunction with the test based upon $\operatorname{sLM^{\star}}$ .
In this section we present an application of our test to the time series of the tree-ring standardized growth index. Tree rings provide a measure of the responses of tree growth to past climate variation and this information is very useful in climate studies. Despite the recognition that the climatic factors affecting tree growth form a complex network, according to the literature, the best model adopted is either the AR$(1)$ or the ARMA$(1,1)$, see table 3 in Fox01. Usually, the indexes of many trees from the same site are used to cross date the rings and are finally averaged into a single index as to obtain a chronology that covers a long time span. Here we focus on the tree-ring chronology of a Pinus aristata var. longaeva (California, USA) from year 800 to 1979 ($n=1180$), for more details on the data see ca535. We test the ARMA$(1,1)$ specification against the following TARMA$(1,1)$ model
With the threshold searched between the 10th to 90th percentiles, the $\operatorname{sLM}$ test statistic is 23.45, while the $\operatorname{sLM^{\star}}$ statistic is 25.21 and both of them correspond to a $p$-value smaller than 0.001, suggesting that tree-ring growth is regulated by floor. Table (ref) reports a TARMA$(1,1)$ model parameterized in the form of ((ref)) with common moving-average parameter but with unconstrained $\phi_{i,1}, i=0,1,2$, and an ARMA$(1,1)$ model fitted to the data. The estimated autoregressive parameters point to a threshold effect and the normalized AIC and BIC indicate an improvement with respect to the ARMA$(1,1)$ model. The estimated TARMA$(1,1)$ model is invertible and geometrically ergodic. The estimated threshold is $\hat r = 0.97$ which is close to 1, the mean of the process, and identifies an upper regime where the growth is accelerated with respect to the lower regime. Finally, model diagnostics reported in the {Supplementary Material} indicate that the TARMA$(1,1)$ model provides a good fit to the data whereas an unaccounted dependence structure is present in the residuals of the ARMA$(1,1)$ model.
In this paper we have presented consistent supremum Lagrange Multiplier tests to compare a linear ARMA specification against its TARMA extension. Our proposal extends previous results, such as Cha90a,Lin05 and enjoys very good finite-sample properties in terms of size and power. Moreover, being based upon asymptotic theory, it has a low computational burden. From the tabulated quantiles of the asymptotic distributions it seems that these depend only on the numbers of parameters tested and match those of And03 and this prompts interesting further theoretical investigations. The Monte Carlo study has shown that supLM tests are also robust against model mis-specification and their performance is not affected if the order of the tested process is unknown, provided a consistent order selection procedure is used. Our supLM tests do not suffer from some of the shortcomings that affect the quasi-likelihood ratio test so that they can be used for small samples. In such a case, the $\operatorname{sLM}$ statistic has less power than the $\operatorname{sLM^{\star}}$ statistic but it is better behaved in terms of size so that it is recommended. For sample sizes from 200 onwards, the two tests can be used in conjunction. The theoretical framework of our supLM tests is valid for the innovation process being a martingale difference sequence but does not take into account GARCH-type innovations. A possible solution would be to adopt a wild-bootstrap scheme similar to that used in Cha20. While the implementation is straightforward, to the best of our knowledge, the validity of the bootstrap in a threshold framework has not been proven, even for TAR models, and constitutes and interesting challenge for future investigations. The analysis of the tree-ring time series shows that TARMA models can provide a new insight into all those problems that make use of dendrochronological data. The TARMA(1,1) fit improves considerably over the commonly accepted linear specification that did not account for a short term non-linear effect.
The supplemental document contains additional results from both the simulation study and the tree-ring data analysis.