EconBase
← Back to paper

Testing the order of fractional integration in the presence of smooth trends, with an application to UK Great Ratios

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

64,255 characters · 12 sections · 96 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Testing the order of fractional integration when smooth deterministic trends are possibly present

abstractThis paper introduces a test for fractional integration in a model that possibly contains smooth deterministic trends. We model the trend component using a Chebyshev polynomial and specify the short-run dynamics semi-parametrically, accommodating a broad class of possibly nonlinear processes, including those with conditional heteroskedasticity. We use a local Whittle approach for constructing a Lagrange multiplier test statistic and for constructing a frequency-domain information criterion for the selection of the order of the Chebyshev polynomial. We show that widely used time-domain information criteria are generally inconsistent for the true order, whereas our frequency-domain criterion remains robust under both short- and long-memory behaviour. Monte Carlo simulations and an empirical application to the UK Great Ratios support our theoretical findings. Keywords: Lagrange multiplier tests, spurious long memory, smooth trends, Chebyshev polynomial, local Whittle likelihood, information criterion. JEL Codes: C12, C14, C22.

Introduction

The intricate interplay of structural change and long memory is a well-documented phenomenon. Early research by diebold2001long, granger2004occasional and sibbertsen2004long illustrates how a short memory process with (random) regime changes can be mistaken to be a long memory process. In essence, tests for fractional order of integration based on models that erroneously do not account for structural change are generally uninformative. iacone2022semiparametric illustrate this by explicitly modelling structural change by step-indicators and showing that, for instance, the Lagrange multiplier test discussed in lobato1998nonparametric is uninformative under structural breaks because its test statistic diverges under the null. This limitation motivates iacone2022semiparametric to extend the model of lobato1998nonparametric by using the procedure suggested by lavielle2000least to model and locate the structural breaks. They use standard information criteria for determining the number of breaks and a local Whittle approach for avoiding the need to explicitly model any short-range dependence in the data.

Another possibility for modelling structural change is via smooth trends instead of step-indicators, see e.g.\ harvey2010testing. cuestas2016testing use a Chebyshev polynomial to model smooth trends when assessing the fractional integration order. bierens1997testing and bierens2010time had shown that these trigonometric functions are particularly effective in approximating any square integrable and differentiable function of time arbitrarily closely. Additionally, trigonometric functions are useful for capturing regular fluctuations in time series data, referred to as cyclical trends by anderson1994statistical and discussed further by perron2020trigonometric. They are also employed to represent seasonal, trend, and irregular components of time series, as described by harvey1993time. cuestas2016testing employ a general-to-specific testing procedure for determining the order of the Chebyshev polynomial and model the short-run dynamics fully parametrically. Yet the difficulty to disentangle smooth trends and long memory is also well-documented, cf., for instance, Bhattacharya_Gupta_Waymire_1983, Kunsch_1986 and giraitis2001testing.

A challenge in iacone2022semiparametric and cuestas2016testing lies in identifying the number of structural breaks and the order of the Chebyshev polynomial in the face of fractional integration, respectively. We present a solution to this challenge. In particular, we combine the methods proposed by cuestas2016testing and iacone2022semiparametric by using a Chebyshev polynomial for modelling structural change and a semi-parametric setup for describing the short-run dynamics. We use a local Whittle objective for constructing a Lagrange multiplier (LM) test for fractional integration and for setting up a novel information criterion for selecting the order of the Chebyshev polynomial.

We establish the asymptotic theory for the proposed testing procedure. In particular, we show that choosing too low an order for the Chebyshev polynomial leaves part of the smooth trend in the residuals, contaminates the spectrum at the origin and causes the LM-statistic to diverge under the null, rendering the resulting test uninformative. This low-frequency contamination is closely related to that discussed by iacone2022semiparametric for structural breaks. On the other hand, when the polynomial order is correctly specified or over-specified, the LM statistic based on the residuals have the same limiting behaviour as the infeasible procedures based on the true unobserved errors. More specifically, under alternatives drifting to the null at the standard local-to-null rate, the LM-statistic converges to a non-central chi-squared distribution, just as in the infeasible benchmark, cf.\ also lobato1998nonparametric, shao2007local, and iacone2022semiparametric. This implies that estimating the deterministic component does not entail any asymptotic loss of local power, provided that the selected polynomial order is not smaller than the true one. This asymptotic robustness to over-specification should not, however, be interpreted as saying that over-fitting is innocuous in practice. As noted by bierens1997testing, introducing superfluous Chebyshev terms may adversely affect finite-sample test performance, so that accurate order selection remains essential.

Indeed, our simulation evidence and empirical findings both indicate that the finite-sample consequences of over-specification can be substantial. We therefore theoretically analyse the behaviour of information criteria for selecting the order of the Chebyshev polynomial. We show that such conventional time-domain criteria as the BIC and HQ, which are valid in analogous settings with $I(0)$ errors, see hall2013inference, are generally unsuitable in the present environment. For, in the presence of positive long memory, they tend to interpret persistence as additional deterministic terms and therefore over-select the polynomial order. lavielle2000least consider penalty modifications to information criteria to account for persistence in the stochastic component. These modifications, however, are infeasible because the penalty depends on the unknown memory parameter. In addition, our simulation results indicate that these modified criteria may exhibit unsatisfactory finite-sample performance, underestimating for strong negative persistence and overestimating when persistence is strong. As a solution to these problems, we show that our novel frequency-domain local Whittle information criterion consistently selects the true polynomial order and thus provides a valid foundation for feasible inference on the fractional integration parameter.

We illustrate the methodology using the UK Great Ratios. kaldor1961capital first introduced the concept of Great Ratios, emphasising stable relationships among key macroeconomic variables. These ratios have since attracted substantial empirical interest, including recent analyses by kapetanios2020time and chudik2023revisiting, particularly regarding their integration properties. While kapetanios2020time find that UK Great Ratios display $I(1)$ behavior but can be represented as $I(0)$ processes when incorporating smooth deterministic trends, we extend this analysis to fractional integration. Our empirical findings reveal that the Great Ratios exhibit long memory. Furthermore, we show that inference can be very sensitive to the choice of the polynomial order, underscoring the important role of the proposed local Whittle information criterion.

The structure of the remaining paper is as follows. Section (ref) and (ref) introduce the model and the test of fractional integration we propose. Section (ref) studies the asymptotic properties of time-domain information criteria and introduces the new frequency-domain criterion for selecting the order of the Chebyshev polynomial. In Section (ref) we conduct a simulation study to examine the finite sample properties of our proposed tests and our local Whittle information criterion. Section (ref) presents an illustrative empirical example applying the methods developed in this article to analyse the UK Great Ratios. Section (ref) concludes. Proofs of the theorems are relegated to the appendix while the supplement contains the results of an extensive Monte Carlo simulation study.

Model specification, testing and order selection

The model

For the observable time series $y_t$, $t = 1,\ldots,T$, we consider a stationary fractional model supplemented by a Chebyshev polynomial, given by

align[align omitted — 126 chars of source]

with

align[align omitted — 93 chars of source]

In this notation, each of the $k+1$ components $P_t(n)$ is referred to as Chebyshev function, while the sum $\sum_{n=0}^k \beta_n P_t(n)$ will be called a Chebyshev polynomial of order $k$. Note that $P_t(0)$ is a constant term. The polynomial serves as a flexible smooth deterministic trend component that can approximate any square integrable and differentiable function of time arbitrarily closely, see bierens1997testing and bierens2010time.

The so-called fractional integration operator in (ref) is defined by $\Delta^{-\delta} = (1-L)^{-\delta} = \sum_{i=0}^{\infty} \pi_i(\delta) L^k $ with the binomial coefficients $\pi_0(\delta) = 1$ and $\pi_i(\delta) = \delta (\delta + 1) \ldots (\delta + i -1)/ i !$ for $i \geq 1$, see for instance hassler2019time. The error term $\eta_t$ is assumed to be a mean-zero stationary, causal non-linear process. Following shao2007local,shao2007local1 and iacone2022semiparametric we impose the following assumptions on $\eta_t$, so as the be able to use the lemmas they establish for this class of processes. Let $\delta_0$ denote the true value of $\delta$.

assumptionLet $$\eta_t = F(\ldots, \varepsilon_{t-1}, \varepsilon_t),$$ where $\{\varepsilon_t\}_{t \in \mathbb{Z}}$ are independent and identically distributed (IID) random variables and $F$ is a measurable function such that $\eta_t$ is well-defined as a stationary, causal, ergodic process. For a random variable $\xi$ and $p > 0$, write $\xi \in L^p$ if $\|\xi\|_p = \left(E(\left|\xi\right|^p)\right)^{1/p} < \infty$. Let $\{\varepsilon^*_t\}_{t \in \mathbb{Z}}$ be an IID copy of $\{\varepsilon_t\}_{t \in \mathbb{Z}}$, $\mathscr{F}_t = (\ldots, \varepsilon_{t-1}, \varepsilon_t)$, $\mathscr{F}_0^* = (\mathscr{F}_{-1}, \varepsilon_0^*)$, $\eta_{\kappa}^* = F(\mathscr{F}_0^*, \varepsilon_1, \ldots, \varepsilon_{\kappa})$, and define the dependence measure $\theta_q({\kappa}) = \|\eta_{\kappa} - \eta_{\kappa}^*\|_q$, for some $q>\max\left\{4,\frac{1}{1+2\delta_0}\right\}$. Define the projection operator $\mathcal{P}_{\kappa} \xi = \mathbb{E}(\xi \mid \mathscr{F}_{\kappa}) - \mathbb{E}(\xi \mid \mathscr{F}_{{\kappa}-1})$. Then: \begin{enumerate}[noitemsep] • $\eta_t \in L^q$, $\sum_{\kappa_1,\kappa_2,\kappa_3}\left|\operatorname{cum}\!\left(\eta_0,\eta_{\kappa_1},\eta_{\kappa_2},\eta_{\kappa_3}\right)\right|<\infty$ and $\sum_{\kappa=0}^{\infty}\|\mathcal{P}_0\eta_{\kappa}\|_q<\infty$, where $\operatorname{cum}(\cdot)$ denotes the joint cumulant of its arguments. • $\sum_{{\kappa}=1}^{\infty} {\kappa} \theta_q({\kappa}) < \infty$. • The spectral density of $\eta_t$, $f_{\eta}(\lambda)$, satisfies $f_{\eta}(\lambda) = G(1 + O(\lambda^2))$ as $\lambda \to 0^+$ for some $G \in (0,\infty)$. \end{enumerate}

Part $(i)$ imposes finite moments together with weak higher-order dependence through the cumulant and projection summability conditions. shao2007local note that fourth-order cumulant summability is a standard condition in spectral analysis, while shao2007local1 show that, in the case of $F$ being a linear function, the corresponding projection and dependence conditions reduce to weighted summability restrictions on the linear coefficients. Part $(ii)$ captures how quickly the effect of a single shock dies out over time in terms of the dependence measure. As emphasised by shao2007local1, this measure is especially appealing because it is directly tied to the data-generating mechanism. Part $(iii)$ requires the spectral density of $\eta_t$ to be smooth and strictly positive at the origin, so that $\eta_t$ behaves as a short-memory component in the local Whittle approximation, see shao2007local1. Assumption (ref) covers a broad class of nonlinear models of $\eta_t$, including bilinear, threshold, GARCH, and ARMA-GARCH processes. The requirement on $q$ is needed to apply a functional central limit theorem later, see iacone2022semiparametric.

The parameter space for our model is defined as follows: $k \in \{0, 1, \ldots, K\}$, $\beta = (\beta_0, \beta_1, \ldots, \beta_k)' \in \mathbb{R}^{k+1}$ and $\delta \in (-0.5, 0.5)$. Finally, the true but unknown parameter values are denoted by $k_0$, $\beta_0 = (\beta_{0,0}, \beta_{1,0}, \ldots, \beta_{k_0,0})' $, $\delta_0$, respectively. The value $K$ represents the maximum order of the Chebyshev polynomial and must be larger or equal to the true value $k_0$. It can, in principle, be set arbitrarily large as long as it is less than $T-1$. In line with the literature, we impose the restriction $|\delta_0|<1/2$ , see e.g.\ iacone2022semiparametric. If $\delta_0$ were to exceed 0.5, the process $ u_t $ would become non-stationary and would dominate the deterministic component, entailing the inconsistent estimation of the coefficients $\beta_n$, $n = 0, 1,\ldots,k$, of the Chebyshev functions, see e.g.\ hualde2020truncated. However, non-stationary data could be made stationary by first-differencing, allowing us to still continue working within the framework of the model above.

The next two subsections contain the two main contributions of the paper. In Section (ref), we develop an LM test for $H_0:\delta=\delta_0$ within the local Whittle framework and derive its asymptotic behaviour. This approach does not require preliminary estimation of $\delta$, nor does it require explicit modelling of the short-run dynamics, making it robust to misspecification of the latter, see shao2007local,iacone2022semiparametric,lobato1998nonparametric. In Section (ref), we then address the selection of the Chebyshev order $k$, since misspecification of the deterministic component adversely affects the testing problem. Because the LM test is itself based on the local Whittle objective, it is natural to base order selection on the same frequency-domain approach. This leads us to propose a novel local Whittle information criterion that provides a consistent basis for selecting $k$. We show that existing time-domain procedures, as opposed to that, are unsuitable in our setting.

Testing the order of fractional integration

We are interested in testing the null hypothesis

align[align omitted — 56 chars of source]

for $\delta_0 \in (-0.5, 0.5) $. lobato1998nonparametric study this problem in the context of model (ref) when the Chebyshev polynomial only consists of a constant term, i.e.\ when $k = 0$. They base the test on the Lagrange multiplier (LM) principle and the local Whittle objective function of robinson1995gaussian. The latter is given by

align[align omitted — 184 chars of source]

Applying the LM principle to (ref) yields the test statistic

align*[align* omitted — 240 chars of source]

Equivalently, the LM statistic can be expressed as the square of a $t$-type statistic, that is of interest in itself:

align[align omitted — 279 chars of source]

where $I_u(\lambda) = \left|w_u(\lambda)\right|^2$ is the periodogram of the error term $u_t$, $w_u(\lambda) = \frac{1}{\sqrt{2\pi T}} \sum_{t = 1}^T u_t e^{i \lambda t}$ is the discrete Fourier transform of $u_t$, $\lambda_j = \frac{2\pi j}{T}$ are the Fourier frequencies, and $v_j = \ln(j) - \frac{1}{m} \sum_{j = 1}^m \ln(j)$, where $m = m(T)$ is the bandwidth. While the construction of the test statistic follows from lobato1998nonparametric, the large-sample conditions we impose are tailored to the broader nonlinear framework in Assumption (ref). Accordingly, following shao2007local and iacone2022semiparametric, we impose the following conditions on the relative divergence rates of $m$ and $T$:

assumptionAs $T \rightarrow \infty$, the bandwidth $m = m(T) \rightarrow \infty$ such that $T^{\epsilon}/ m + m/T^{2/3} \rightarrow 0$ for some fixed $\epsilon > 0$.

lobato1998nonparametric show that, when $\delta_0 = 0$ and under conditional homoskedasticity, the test statistics in (ref) and (ref) are asymptotically $\mathcal N(0,1)$ and $\chi^2_1$ under $H_0$, respectively. shao2007local extend these results to conditional heteroskedasticity. iacone2022semiparametric further extend the theory to any $|\delta_0|<1/2$.

As a first contribution of this paper, we consider the model in (ref)-(ref) when the Chebyshev polynomial comprises not only a constant term. We solve the problem of testing $H_0$ in (ref) when $u_t$ is unobserved, as will in practice usually be the case. In essence, we replace the unobservable $u_t$ in (ref) by its sample counterpart, estimating $\beta$ by ordinary least squares (OLS) for a given value of $k$. Neglecting the short-run dynamics in $u_t$ is of course inefficient yet, under the assumptions specified above, the OLS estimator is consistent. The fact that the true order $k_0$ of the Chebyshev polynomial is in fact unknown makes the tests depend on whether $k$ is under- or overspecified in the model. This problem will be remedied in Section (ref) below, where an information criterion is suggested for the consistent selection of $k$.

For a given $k$, let $y = (y_1,\ldots,y_T)'$, $x_t(k)$ $=$ $(P_t(0),\ldots, P_t(k))'$ and $X(k) = (x_1(k),\ldots,x_T(k))'$. Then the OLS estimator of $\beta$ is given by $\hat{\beta}(k)= (X(k)'X(k))^{-1}X(k)'y$. Define the corresponding residuals,

align[align omitted — 73 chars of source]

The statistics in (ref) and (ref) based on $\hat{u}_t(k)$ instead of $u_t$ are denoted by $ t_{\hat{u}(k)}(\delta;m)$ and $\textit{LM}_{\hat{u}(k)}(\delta;m)$, respectively. In the remainder of this section, we will establish the limiting behaviour of $t_{\hat{u}(k)}(\delta;m)$ and $\textit{LM}_{\hat{u}(k)}(\delta;m)$, with Theorem (ref) covering the under-specified case ($k < k_0$) and Theorem (ref) addressing the correctly specified ($k = k_0$) and over-specified ($k > k_0$) cases. These asymptotic results hinge on the fact that the coefficients $\beta_n$, $n = 0, 1,\ldots,k$, of the Chebyshev functions can be consistently estimated at rate $T^{1/2-\delta_0}$, even when $k$ is under- or over-specified, cf.\ Lemma (ref). This is achieved via the orthogonal property of the Chebyshev polynomial and the application of the functional central limit theorem, which, as mentioned earlier, holds in our setting.

The following theorem shows the limiting behaviour of $t_{\hat{u}(k)}(\delta;m)$ and $\textit{LM}_{\hat{u}(k)}(\delta;m)$ under $H_0$ when $k$ is under-specified, under an additional restriction on the bandwidth $m$.

theoremLet $y_t$ be generated according to (ref)-(ref) and let Assumption (ref) hold. The bandwidth $m$ satisfies the conditions in Assumption (ref) and $T^{1-2\delta_0} \ln(m) m^{- 1/2} \rightarrow \infty$. Then for any $k < k_0 $ with $k_0 > 0$ and any $\delta_0 \in (-0.5,0.5)$, \begin{align*} t_{\hat{u}(k)}(\delta;m) &\overset{p}{\rightarrow} \infty, \\ LM_{\hat{u}(k)}(\delta;m) &\overset{p}{\rightarrow} \infty, \end{align*} under $H_0 : \delta = \delta_0$. If, moreover, $T^{1-2\delta_0} m^{-1} \rightarrow \infty$, then \begin{align*} t_{\hat{u}(k)}(\delta;m) &\overset{p}{\sim} C \sqrt{m} \ln(m), \\ LM_{\hat{u}(k)}(\delta;m) &\overset{p}{\sim} C^2 m\ln^2(m), \end{align*} under $H_0 : \delta = \delta_0$, where $C$ is a positive finite constant.

Three observations can be made about Theorem (ref): First, if the DGP includes Chebyshev functions beyond the constant term, i.e.\ if $k_0 > 0$, yet these functions are ignored in the model by setting $k = 0$ then both test statistics will diverge under the null. Therefore, in this case, both tests are not informative. Secondly, even if Chebyshev functions beyond the intercept are included in the model ($k > 0$) but the order of the Chebyshev polynomial in the model is smaller than the true number ($k < k_0$), the test statistics will still diverge. Thirdly, the rate of divergence of both statistics in the second part of Theorem (ref) is the same as in the setting studied by iacone2022semiparametric, who address level breaks instead of smooth trends. Hence, neglecting smooth trends and neglecting level breaks affects these tests through a similar low frequency contamination mechanism.

The assumptions on the bandwidth in Theorem (ref) are best interpreted as relative strength conditions: they guarantee divergence of the test statistics by making the contribution of the omitted Chebyshev component dominate the error term, cf.\ mccloskey2013memory. The weaker restriction in the first part of the theorem is sufficient to make the contamination component dominate in the numerator of the $t$-type statistic, which implies divergence of both test statistics. The stronger restriction in the second part is imposed to ensure that the omitted component also dominates the denominator, which is what delivers the explicit divergence rates in the theorem.

Essentially, therefore, Theorem (ref) says that both the $t$- and the LM-test are uninformative under the null when the Chebyshev polynomial is underspecified. Theorem (ref) rectifies this by establishing that both tests are well-behaved under the null and under local alternatives as along as the Chebyshev polynomial is correctly or overspecified relative to the true value $k_0$. This result highlights the robustness of the proposed testing procedures to potential overfitting in the model and ensures that the limiting distributions remain valid provided that $k \geq k_0$.

theoremLet $y_t$ be generated according to (ref)-(ref) and let Assumption (ref) hold. The bandwidth $m$ satisfies the conditions in Assumption (ref). Then for any $k \geq k_0$ and any $\delta_0 \in (-0.5,0.5)$, \begin{align} t_{\hat{u}(k)}(\delta_0;m) &\overset{d}{\rightarrow} \mathcal N(2c,1), \\ LM_{\hat{u}(k)}(\delta_0;m) &\overset{d}{\rightarrow} \chi_1^2(4c^2), \end{align} under $H_c :\delta = \delta_0 + c m^{-1/2}$.

The theorem establishes that, under local alternatives converging to the null at the $\sqrt{m}$-rate, the $t$-statistic converges in distribution to a normal distribution with mean $2c$ and variance one, while the LM-type statistic converges to a non-central chi-square distribution with one degree of freedom and non-centrality parameter $4c^2$. These limiting distributions are identical to those obtained using the infeasible versions of the statistics based on the true errors $u_t$, as shown in iacone2022semiparametric. This equivalence implies that, asymptotically, the estimation of $u_t$ by $\hat{u}_t(k)$ does not lead to any loss of local power, provided that the Chebyshev order is sufficiently large. Thus, the proposed procedures are not only consistent but also asymptotically efficient in the sense that they achieve the same local asymptotic power as the infeasible tests based on the true errors $u_t$. Note that Assumption (ref) on the bandwidth is sufficient for the asymptotic theory in Theorem (ref).

Selecting the Chebyshev order

A crucial issue in the testing procedures suggested in Section (ref) is of course the fact that the true $k_0$ is unknown. From Theorem (ref), it follows that, if $k < k_0$, both statistics diverge under the null hypothesis. On the other hand, Theorem (ref) shows that if $k \geq k_0$, the tests have a well-defined asymptotic distribution under both the null and local alternatives. Therefore, the choice of $k$ is important. bierens1997testing, too, emphasises the importance of $k$, noting that the power and size of the unit root test he examines is highly sensitive to the choice of $k$. The literature contains a number of procedures for selecting $k$, three of which will be discussed in the next subsection. We will derive their properties and argue that they all have drawbacks. We subsequently suggest a new solution to the selection problem.

Existing selection procedures

First, cuestas2016testing propose a method to address this challenge by selecting $k$ via a sequential testing procedure. Their general-to-specific approach starts with a high-order polynomial and gradually removes statistically insignificant Chebyshev coefficients based on their $t$-statistics, continuing until only significant coefficients remain. However, the method has limitations: First, the distribution of the estimated Chebyshev coefficients depends on the unknown true fractional integration order $\delta_0$ which must be estimated to perform the $t$-tests, cf.\ Lemma (ref). Additionally, the long-term variance estimator required to studentise the coefficient estimates typically also depends on $\delta_0$. A detailed discussion is provided in abadir2009two.

Information criteria offer an alternative approach for conducting inference about the order of the Chebyshev polynomial. These criteria are frequently employed in the context of structural breaks in models with $I(0)$ errors, see bai2003computation and hall2013inference, as well as in models with $I(\delta)$ errors where $|\delta| < 1/2$, see lavielle2000least. The criteria in this literature all take the form

align[align omitted — 71 chars of source]

where $B(k)$ is some estimate of the in-sample fit and $A(T)$ is a positive deterministic penalty term. The estimate of the order of the Chebyshev polynomial, denoted $\hat{k}$, is the value that minimises $\textit{IC}(k)$, that is,

align[align omitted — 100 chars of source]

As the second of the selection procedures we discuss, consider the most natural specifications of $B(k)$ and $A(T)$, viz.\ the Bayesian Information Criterion and the Hannan-Quinn information criterion. We refer to these conventional time-domain criteria as BIC and HQ. These criteria both set $B(k) = \ln \hat{\sigma}^2(k)$ where

align[align omitted — 124 chars of source]

with $\hat{u}_t(k)$ defined in (ref), while their penalty terms are respectively, $A(T) = \log(T)/T$ and $A(T) = 2c\log(\log(T))/T$, where $c > 1$. It can be verified that these penalty terms satisfy the following assumption:

assumptionAs $T \rightarrow \infty$, $A(T) = o(1)$ and $TA(T)\rightarrow \infty$. If $\delta_0 > 0 $ assume that, in addition, $T^{1-2 \delta_0} A(T)= o(1)$.

This assumption is akin to that in hall2013inference and stipulates that the rate at which $A(T)$ converges to zero must be at most $T$ and, if $\delta_0 > 0$, at least $T^{1 - 2\delta_0}$. While, in linear models with structural breaks and $I(0)$ errors, hall2013inference show that a penalty term $A(T)$ satisfying Assumption (ref) yields a consistent estimate of the number of breaks, we show in the following theorem that, in our model with $I(\delta)$ errors, the consistency of $\hat{k}$ depends on $\delta_0$.

theoremLet $y_t$ be generated according to (ref)--(ref) and let Assumptions (ref) and (ref) hold. If $\delta_0 > 0$, then $P(\hat{k} > k_0) \rightarrow 1$. If $\delta_0 \leq 0$, then $\hat{k} \overset{p}{\rightarrow} k_0$.

This theorem shows that $\hat{k}$ leads to overestimation when $\delta_0 > 0$, implying that long memory is misinterpreted as Chebyshev functions. As noted by bierens1997testing, such overestimation is undesirable because it introduces unnecessary Chebyshev functions into the model. This, in turn, distorts significantly the size and power of the tests in Section (ref), as the inclusion of superfluous Chebyshev functions adds noise and does not improve the model's accuracy. This issue is further demonstrated in the simulation study in Section (ref). However, Theorem (ref) also shows that $\hat{k}$ is consistent when $\delta_0 \leq 0$, confirming that the analogous result of hall2013inference for models with structural breaks and $I(0)$ errors also holds for some parameter values in our setting.

The reason for the overestimation of the Chebyshev polynomial order when $\delta_0 > 0$ is that the BIC, the HQ and other criteria satisfying Assumption (ref) do not impose a sufficiently strong penalty when persistence is present in the time series. In the context of structural breaks, this has motivated lavielle2000least to modify the penalty term $A(T)$ such that it converges to zero at a slower rate than in Assumption (ref) when $\delta_0 > 0$. Following lavielle2000least, we thus replace Assumption (ref) by

assumptionAs $T \rightarrow \infty$, $A(T) = o(1)$ and $T^{1-2\delta_0}A(T) \rightarrow \infty$, for any $\delta_0 \in (-0.5, 0.5)$.

Consequently, $A(T)$ must now converge to zero at a slower rate than $T^{1 - 2\delta_0}$. In particular, lavielle2000least modify the BIC penalty term by making the rate in the denominator depend on $\delta_0$, defining $A(T) = \log(T)/T^{1-2\delta_0}$. Similarly, one could modify the HQ criterion by setting $A(T) = 2c\log(\log(T))/T^{1-2\delta_0}$. We refer to these infeasible $\delta_0$-dependent information criteria, the third procedure type in our categorisation, as $\delta$-BIC and $\delta$-HQ. With $B(k)$ still given by $\ln \hat \sigma^2_u (k)$ in (ref), we can then prove the following behaviour of $\hat k$ in (ref):

theoremLet $y_t$ be generated according to (ref)--(ref) and let Assumptions (ref) and (ref) hold. Then, $\hat{k} \overset{p}{\rightarrow} k_0$.

Theorem (ref) shows that the estimated order of the polynomial is consistent when $A(T)$ satisfies Assumption (ref). The modification of the penalty of course poses a challenge: since $\delta_0$ is unknown the $\delta$-BIC and $\delta$-HQ cannot be computed in practice.

All three procedures for determining $k$ described in this section are thus dependent on $\delta_0$, making them ineffective for the use with the tests proposed in Section (ref). In the following, we will suggest a solution out of this impasse.

A new selection procedure

As a second contribution of this paper, we now introduce a new selection procedure for $k$ to avoid the dependence on $\delta_0$ of the information criteria discussed above. Relative to those approaches, it re-defines both the estimator $B(k)$ and the penalty term $A(T)$ in (ref), taking into account the estimation of the memory parameter $\delta$ as well as the non-parametric nature of the short-run dynamics. The main idea is to move order selection to the frequency domain and use the local Whittle (LW) objective function as the measure of fit.

Specifically, an estimator of $B(k)$ for a given candidate model with $k$ Chebyshev functions is obtained by a 2-step procedure. In the first step, the long-memory parameter is estimated semi-parametrically using the local Whittle estimator $\hat \delta (k)$ applied to the residuals $\hat{u}_t(k)$ defined in (ref), i.e.\

align*[align* omitted — 89 chars of source]

where $m = m(T)$ again satisfies the bandwidth rate condition in Assumption (ref). The parameter space is $\mathcal D = [ \Delta_1,\Delta_2 ]$, with $-1/2 < \Delta_1 < \Delta_2 < 1/2$, $\delta_0 \in \mathcal D$. In the second step, define the goodness-of-fit term $B(k)$ in (ref) as

align[align omitted — 77 chars of source]

i.e.\ the local Whittle objective in (ref) evaluated at the first-step estimate.

Turn now to the penalty term $A(T)$ in (ref). In contrast to Assumptions (ref) and Assumption (ref) above, it is defined in terms of the bandwidth $m$ used by the first step local Whittle estimator of $\delta$:

assumptionAs $T \rightarrow \infty$, $A(T) = o(1)$ and $A(T) m (T)\rightarrow \infty$.

According to Assumption (ref), the penalty $A(T)$ should decay to zero sufficiently fast in order to avoid under-selecting the Chebyshev order, while not decaying too quickly to prevent over-selection. Natural choices for $A(T)$ mirror the standard BIC or HQ penalties with the effective sample size $m = m(T)$ in place of $T$, namely

align[align omitted — 161 chars of source]

with $c > 1$. We refer to the resulting criteria as LW-BIC and LW-HQ, respectively.

Given a model with $k$ Chebyshev functions, we hence base the information criterion on the local Whittle fit term $R(\hat{\delta}(k);m)$ in (ref) and on the penalty $A(T)$ in (ref) or (ref), such that the optimal polynomial order is selected as in (ref).

The following theorem shows the consistency of $ \hat{k}$.

theoremLet $y_t$ be generated according to (ref)-(ref) and let Assumptions (ref), (ref) and (ref) hold. Furthermore, assume that $P(\hat{\delta}(k)\leq \delta_0) \rightarrow 0$ holds for all $k < k_0$. Then \begin{align*} \hat{k} \overset{p}{\to} k_0. \end{align*}

The mechanism behind Theorem (ref) is straightforward. When $k<k_0$, the low-frequency contamination left in $\hat u_t(k)$ worsens the local Whittle fit. As a result, $B(k)>B(k_0)$ with probability tending to one. Since $A(T)=o(1)$, the penalty component $(k+1)A(T)$ is asymptotically negligible relative to this deterioration in fit, and the criterion rules out underfitting. When $k> k_0$, the contamination is removed and $\hat\delta(k)$ is consistent with the $\sqrt{m}$-usual rate, cf.\ Lemma (ref). Yet, the resulting difference in fit components is asymptotically negligible, i.e.\ $B(k)-B(k_0)=O_p(m^{-1})$, and no longer separates correctly-specified from overspecified models. Assumption (ref) is then intrumental: $A(T)m\to\infty$ ensures that the penalty term dominates the $O_p(m^{-1})$ fit differences, thereby preventing over-selection.

The additional condition $P(\hat{\delta}(k)\leq \delta_0) \rightarrow 0$ for all $k < k_0$ requires that, should the Chebyshev polynomial be underspecified, the estimator of $\delta$ is sufficiently large and thereby compensates for the low frequency contamination in $\hat{u}(k)$. If this spurious long memory is not present then the local Whittle goodness-of-fit term in (ref) may fail to penalise under-specified models. In that case, the information criterion may favour $k < k_0$, leading to under-selection. Assumptions of this type are standard in related problems, see e.g.\ qu2011test and SibbertsenEtAl18.

While we use the local Whittle estimator in the first step for logical consistency, the consistency result in Theorem (ref) does in fact not hinge on this particular choice. More generally, our consistency proof uses only two properties of the first–step estimator $\hat\delta(k)$. First, under under-fitting ($k<k_0$), it yields spurious long memory in the sense that $P(\hat{\delta}(k)> \delta_0)\rightarrow 1$. Second, under correct specifcation or over-fitting ($k\geq k_0$), it is $\sqrt{m}$-consistent, i.e.\ $m^{1/2} (\hat{\delta}(k)-\delta_0) = O_p(1)$. Consequently, the local Whittle estimator can be replaced by any other semi-parametric estimator that satisfies these two conditions. Examples may include the exact local Whittle estimator of exactShimotsu and the fully extended local Whittle estimator of abadir2007nonstationarity. Similarly, we use the same notation $m$ for the value of the bandwidth in the construction of both the test statistics in (ref)-(ref) and the LW information criteria. These quantities could in principle, however, be based on different bandwidths, provided that they each meet Assumption (ref).

Monte Carlo simulation

In this section, we investigate the finite-sample properties of the tests discussed in Section (ref). Section (ref) considers a setting in which the order of the Chebyshev polynomial is fixed at some value $k$ that is not necessarily the true order. Subsequently, Section (ref) evaluates the tests when the order is selected via the information-criterion approaches described in Section (ref), which is more realistic for empirical work. We conclude Section (ref) by a set of practical recommendations based on the simulation evidence reported in the Supplementary Appendix. All computations are performed in MATLAB 2019a, see MATLAB.

Tests with a fixed Chebyshev order

Figure (ref) compares the asymptotic behaviour of the LM-type statistic as predicted by Theorem (ref) with finite-sample simulations. The DGP given in (ref)--(ref) is assumed to contain a constant term only, so the true order of the Chebyshev polynomial is $k_0 = 0$, and we set $\beta_{0,0} = 0$ without loss of generality. The short-run dynamics are generated by $\eta_t \sim \text{NIID}(0,1)$, while the fractional parameter is $\delta_0 \in [-0.499, 0.499]$. The baseline setup of the simulation is a sample size $T = 512$ and a bandwidth $m = \left \lfloor T^{0.65} \right \rfloor$. Following the literature, we also consider a larger bandwidth of $m = \left \lfloor T^{0.80} \right \rfloor$. So as to match the empirical sample in Section (ref), we examine a sample size of $T = 256$, too. The null hypothesis of interest is $H_0 : \delta = 0$ and the nominal asymptotic significance level is 5%. Rejection frequencies of $H_0$ are reported as a function of $\delta_0$, averaged over 5,000 replications. The LM-type statistic $\textit{LM}_{\hat u \left( k \right)} $ is based on the model whose Chebyshev polynomial is either correctly specified ($k = 0$) and overspecified ($k = 1$). In addition, the asymptotic local power curve from Theorem (ref) is shown as a benchmark.

For all values of $T$ and $m$, the rejection frequencies under $H_0$ approximate the nominal size well. Simulated power is almost indistinguishable between the correctly specified and the overspecified model for negative values of $\delta_0$. For positive values of $\delta_0$, power in the correctly specified model is uniformly higher than in the overspecified model. It is closer to the asymptotic benchmark for positive than for negative values of $\delta_0$. The difference between the correctly specified and overspecified models as well as the difference to the asymptotic benchmark shrink with larger samples and higher bandwidths. Overall, the finite sample patterns match the asymptotic predictions of Theorem (ref), especially for larger bandwidths.

Figure (ref) extends the simulation evidence to include an illustration of Theorem (ref). The DGP is as for Figure (ref), yet the true order of the Chebyshev polynomial is now $k_0 = 1$, so the DGP contains an intercept and one additional Chebyshev function. The true coefficient values are set to $\beta_{0,0} = 0$ and $\beta_{1,0} = 1$, respectively. The LM-type statistic $\textit{LM}_{\hat u \left( k \right)} $ is again based on the model whose Chebyshev polynomial is of order $k = 0$ or $k = 1$, meaning that the model is now underspecified or correctly specified, respectively.

A clear distinction is now apparent between these two model specifications. When $k = 1$ such that the model is correctly specified, the finite-sample size and power of the LM-type test mirror the corresponding conclusions in Figure (ref): The simulated size is close to the nominal size for all values of $T$ and $m$, while the simulated power is substantial and approaches its asymptotic benchmark for increasing sample sizes and bandwidths. However, when $k=0$ such that the model is underspecified the test rejects 100% of the time when $\delta_0 = 0$, confirming the prediction of Theorem (ref). Indeed, the test always rejects when the null is not true.

figure[figure omitted — 1,051 chars of source]
figure[figure omitted — 1,075 chars of source]

Tests with Chebyshev order selected by information criteria

We now examine the behaviour of our LM test when the order of the Chebyshev polynomial is selected by the information criteria discussed in Section (ref). Specifically, we consider {\it (i)} the standard time-domain HQ criterion, {\it (ii)} the infeasible $\delta$-HQ criterion proposed by lavielle2000least and {\it (iii)} our local Whittle HQ information criterion LW-HQ. In all cases, we set the penalty tuning constant in (ref) equal to $c=1.0001$. The experimental setup follows that of the previous section, involving a Chebyshev polynomial of true order $k_0=1$ with coefficient $\beta_{0,0}=0$ and $\beta_{1,0}=1$. We omit the case $k_0=0$ since it leads to qualitatively identical conclusions. The true fractional parameter is $\delta_0 \in \mathcal D$, with $\mathcal D = [-0.499, 0.499]$. The LM-type statistic $\textit{LM}_{\hat u \left( k \right)} $ for a model with correctly specified polynomial order $k = k_0$ is included as a benchmark. The parameter space to estimate $\delta$ over is taken to be $\mathcal D$. The LW-HQ criterion is computed using the bandwidth $m=\lfloor T^{0.65}\rfloor$.

Figure (ref) illustrates simulated finite-sample power functions for the LM-type test. When the time-domain HQ criterion is used for order selection, the test exhibits a substantial loss in power for $\delta_0>0$, as predicted by Theorem (ref), particularly for smaller bandwidths $m$ and smaller sample sizes $T$. This power loss decreases as $m$ increases, but it remains visible even for comparatively large $m$. For the commonly used bandwidth $m=\lfloor T^{0.65}\rfloor$, the loss in power is especially pronounced. By comparison, the LM test using the true order ($k=k_0$) achieves reasonable power across all scenarios. The LM test based on the infeasible $\delta$-HQ also performs well and often tracks the benchmark closely. In some cases it even delivers higher rejection rates than the benchmark. However, as we show forthwith, this reflects systematic mis-selection of the polynomial order rather than a genuine improvement in performance. Importantly, our feasible LW-HQ yields power functions that are very close to those of the true-order benchmark.

Figure (ref) displays histograms of the selected Chebyshev order for the DGP in (ref)--(ref), where $\eta_t \sim \text{NIID}(0,1)$ and $\delta_0\in\{-0.3,0,0.35,0.45\}$. The true order of the Chebyshev polynomial is $k_0 = 1$, and the true coefficient values are set to $\beta_{0,0} = 0$ and $\beta_{1,0} = 1$ without loss of generality. The sample sizes considered are $T\in\{256,512\}$. It turns out that the time-domain HQ criterion is accurate when $\delta_0\le 0$ (see panels (a),(b),(e),(f)), but for $\delta_0>0$ it tends to increasingly overselect large orders, in line with Theorem (ref). This upward bias also becomes more pronounced as $T$ increases (cf.\ panels (c)-(d), vs.\ (g)-(h)). The $\delta$-HQ criterion of lavielle2000least behaves reasonably at $\delta_0=0$ and $\delta_0= 0.3$, but under negative persistence ($\delta_0=-0.3$) it underpenalises and selects overly large $k$, whereas under strong persistence ($\delta_0=0.45$) it overpenalises and collapses to $k=0$ in the vast majority of replications. By contrast, our LW-HQ criterion concentrates its mass at the true order across all values of $\delta_0$ and both sample sizes, indicating stable order selection throughout.

Extensive simulations reported in the Supplementary Appendix (see Tables S.1-S.6 for size, Tables P.1-P.4 for power, and Tables O.1-O.6 for order selection) examine the performance of the LM test, the associated left- and right-tailed $t$-tests, and the order-selection procedures for a variety of other DGPs and model specifications: we allow the short-run component $\eta_t$ to follow IID, AR(1), and ARCH(1)-type processes, so as to better match features of empirical data. Throughout, the deterministic component is modelled by a Chebyshev polynomial with true order $k_0 \in \{0,1\}$ and varying magnitude of the corresponding coefficients. We let the bandwidth $m=\left\lfloor T^{\alpha}\right\rfloor$ vary with $\alpha \in \{0.40,\ldots,0.80\}$ and let the sample size $T \in \{256,512,1024\}$. In the following, we summarise the main findings.

First, the specification of the short-run dependence in $\eta_t$ matters for the finite sample performance of the LM test and for order selection. While the procedures are generally well behaved under IID and ARCH(1)-type specifications, the simulations show that AR(1)-type of dependence can lead to marked size distortions in the test, especially when the bandwidth is chosen too large. For empirically relevant one-sided alternatives, the right-tailed $t$-test behaves well yet can become oversized under $AR$-type short-run dependence. For order selection, the impact of short-run dependence is most visible for the time-domain criteria, which perform markedly worse in the AR(1) design than the LW-based procedures. LW-BIC, in particular, remains highly accurate once a moderate bandwidth is used.

Second, our simulation results demonstrate that the bandwidth $m$ has a significant impact on the finite sample properties of the LM test and of the frequency-based order selection. Regarding the test, we recommend a bandwidth choice that is moderate, i.e.\ between $m=\left\lfloor T^{0.50}\right\rfloor $ and $m=\left\lfloor T^{0.60}\right\rfloor$, which typically delivers the best size and power balance across the different designs. It is important to avoid larger bandwidths (e.g.\ between $m = \left\lfloor T^{0.65}\right\rfloor$ and $\left\lfloor T^{0.80}\right\rfloor$) as they can generate severe size distortions when the DGP contains short-run serial dependence beyond the fractional component. On the other hand, a very small $m$ is often conservative and can significantly reduce power. As for order selection, our additional simulations suggest that the probability of selecting the correct Chebyshev order under the LW-based criteria improves monotonically with $m$, while over-selection declines sharply as the bandwidth increases. In particular, values around $m=\lfloor T^{0.60}\rfloor$ already deliver highly accurate order selection and, at the same time, satisfy the bandwidth condition in Assumption (ref). The LW-BIC performs best, with LW-HQ closely behind. The time-domain BIC or HQ criteria should be avoided because they are not consistent under long memory. Although the $\delta$-BIC and $\delta$-HQ criteria also perform well in some cases, they are infeasible in practice and therefore are best viewed as benchmark procedures.

Overall, we recommend that either the LM-test or the $t$-test be combined with LW-BIC and that a moderate bandwidth be used in practice. In particular, $m=\lfloor T^{0.60}\rfloor$ provides a satisfactory overall compromise across the designs considered in our simulations, performing well in terms of both Chebyshev order selection and finite-sample properties of the tests, while also satisfying the bandwidth condition in Assumption (ref). Accordingly, in the empirical application in Section (ref), we implement LW-BIC with $m=\lfloor T^{0.60}\rfloor$.

figure[figure omitted — 1,295 chars of source]
landscape\begin{figure}[p] \subfloat[$T=256,\ \delta_0=-0.3$] \subfloat[$T=256,\ \delta_0=0$] \subfloat[$T=256,\ \delta_0=0.3$] \subfloat[$T=256,\ \delta_0=0.45$] \subfloat[$T=512,\ \delta_0=-0.3$] \subfloat[$T=512,\ \delta_0=0$] \subfloat[$T=512,\ \delta_0=0.3$] \subfloat[$T=512,\ \delta_0=0.45$] \caption{Histograms of the estimated Chebyshev polynomial order selected by three HQ-type information criteria. The DGP is as in (ref)--(ref) with $u_t = \Delta^{-\delta_0} \eta_t$, where $\eta_t \sim \text{NIID}(0,1)$ and $\delta_0 \in \{-0.3, 0, 0.35, 0.45\}$. The true Chebyshev polynomial is specified by $k_0 = 1$, $\beta_{0,0} = 0$ and $\beta_{1,0} = 1$. The selection frequencies are averaged across 5,000 replications. The results for the time domain HQ criterion are shown in red, those for the infeasible $\delta$-HQ criterion proposed by lavielle2000least in green, and for our LW-HQ criterion in magenta. For LW-HQ we fix the bandwidth $m = \left\lfloor T^{0.65} \right\rfloor$.} \end{figure}

Empirical illustration: UK Great Rations

kaldor1961capital presents a series of “stylized facts” delineating constant relationships among certain macroeconomic variables. klein1961some essentially refer to them as Great Ratios. The concept of Great Ratios has given rise to a considerable empirical research body, which includes recent studies by kapetanios2020time and chudik2023revisiting. Of particular interest has been the question of whether the Great Ratios follow integrated processes. kapetanios2020time show that when analysing several ratios computed with UK data they display $I(1)$ behaviour although they can be described as $I(0)$ when a slowly changing deterministic component is included in the model. We broaden this debate by considering fractional alternatives. kapetanios2020time base their testing approach on the KPSS statistic. A KPSS-type test is, however, inappropriate in our setting since shao2007local establish that it has limited power against local fractional alternatives, namely of order $\log(T)^{-1}$. As opposed to that, our tests possess substantial power against local fractional alternatives, as shown in the Monte Carlos simulations in Section (ref).

We construct the nominal Great Ratios using quarterly UK data from 1955Q1 to 2019Q4, i.e.\ up until the onset of the COVID-19 pandemic, yielding $T=260$ observations. To facilitate direct comparison with kapetanios2020time, we also consider the subsample 1955Q1--2015Q4. Variable definitions follow kapetanios2019time. All ratios are in logs. As recommended by our simulations in Section (ref), we use the LW-BIC criterion for selecting the order of the Chebyshev polynomial. Both our LM-type statistic and our LW-BIC are computed using a bandwidth of $m=\lfloor T^{0.6}\rfloor$. We set the maximal Chebyshev polynomial order equal to $K = 10$ and restrict the fractional parameter to $\delta\in[\Delta_1,\Delta_2]$ with $\Delta_1=-0.499$ and $\Delta_2 = 0.499$. Figure (ref) plots the data series with the estimated deterministic component $\sum_{n=0}^{\hat{k}}\hat\beta_n P_t(n)$ superimposed: time-domain BIC selections are shown in black and frequency-domain LW-BIC in red. The $\delta$-BIC of lavielle2000least is infeasible. Under time-domain BIC, the fitted component follow the data closely because the selected polynomial orders are near the upper bound, producing plots that resemble those in kapetanios2020time, see also the values of $\hat k$ in Panel A of Table (ref). On the other hand, LW-BIC selects parsimonious polynomial orders, yielding fewer Chebyshev functions, cf.\ $\hat k$ in Panel B of Table (ref). The table also reports $t_{\hat{u}(\hat{k})}(\delta_0;m)$ for testing $H_0:\delta=0$ against $H_1:\delta>0$. With $k$ selected by the time-domain BIC, the $t$-statistics are small (i.e.\ less than 1.66) in both the full sample (1955Q1–2019Q4) and the subsample (1955Q1–2015Q4), so the $t$-test cannot reject $H_0$ at conventional levels. This reflects the tendency of time-domain BIC to overestimate $k$ and is consistent with Theorem (ref). Using LW-BIC for the selection of $k$, however, we reject $H_0$ at the 1% level for all seven ratios in both samples, indicating pronounced long memory. Our findings therefore differ from those of kapetanios2020time, who conclude that the nominal Great Ratios are $I(0)$ once slowly varying deterministic components are allowed: Our fractional test continues to reject $H_0 : \delta=0$ at the 1% level for all seven ratios, even when we use the larger sample. This effectively implies that the macroeconomic time series can be seen as fractionally cointegrated.

table[table omitted — 1,732 chars of source]
figure[figure omitted — 964 chars of source]

Conclusion

This paper addresses the challenge of disentangling smooth deterministic trends from the order of fractional integration. Building on the frameworks proposed by cuestas2016testing and iacone2022semiparametric, we develop a semi-parametric testing procedure for the order of fractional integration that models smooth trends using a Chebyshev polynomial. A local Whittle approximation is employed to mitigate issues arising from unmodelled short-range dependence. A key contribution is a novel frequency-domain information criterion that consistently selects the polynomial order, even under long memory. The criterion measures model fit using the local Whittle objective computed from OLS residuals and controls for model complexity through a penalty that depends only on the bandwidth $m$.

We prove the asymptotic validity of the proposed testing and selection procedures and our simulation study confirms their favourable finite-sample performance. Our empirical application to the UK Great Ratios provides new evidence suggesting the presence of fractional cointegration among these important macroeconomic variables.

Beyond our specific setting, the frequency-based information criterion we propose has broader potential. First, it could be employed to address similar model selection challenges in the framework of iacone2022semiparametric, where determining the number of structural breaks is complicated by the presence of long memory and where standard time-domain information criteria select too many breaks when $\delta_0 > 0$. Using our local Whittle information criterion could mitigate this tendency. Secondly, the same approach could be adapted to alternative orthogonal representations of deterministic smooth trends, such as those considered by perron2020trigonometric and abadir2011d.

\printbibliography[title={References}]