EconBase
← Back to paper

Nonfractional Memory: Filtering, Antipersistence, and Forecasting

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

46,414 characters · 12 sections · 39 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Nonfractional Memory: Filtering, Antipersistence, and Forecasting

abstractThe fractional difference operator remains to be the most popular mechanism to generate long memory due to the existence of efficient algorithms for their simulation and forecasting. Nonetheless, there is no theoretical argument linking the fractional difference operator with the presence of long memory in real data. In this regard, one of the most predominant theoretical explanations for the presence of long memory is cross-sectional aggregation of persistent micro units. Yet, the type of processes obtained by cross-sectional aggregation differs from the one due to fractional differencing. Thus, this paper develops fast algorithms to generate and forecast long memory by cross-sectional aggregation. Moreover, it is shown that the antipersistent phenomenon that arises for negative degrees of memory in the fractional difference literature is not present for cross-sectionally aggregated processes. Pointedly, while the autocorrelations for the fractional difference operator are negative for negative degrees of memory by construction, this restriction does not apply to the cross-sectional aggregated scheme. We show that this has implications for long memory tests in the frequency domain, which will be misspecified for cross-sectionally aggregated processes with negative degrees of memory. Finally, we assess the forecast performance of high-order $AR$ and $ARFIMA$ models when the long memory series are generated by cross-sectional aggregation. Our results are of interest to practitioners developing forecasts of long memory variables like inflation, volatility, and climate data, where aggregation may be the source of long memory. {\it Keywords:} Nonfractional memory, long memory, fractional difference, antipersistence, forecasts. {\it JEL classification:} C15, C22, C53.

Introduction

Long memory has been a topic of interest in econometrics since Granger1966's (Granger1966) study on the shape of the spectrum of economic variables. Granger found that {\it long-term fluctuations in economic variables if decomposed into frequency components are such that the amplitudes of the components decrease smoothly with decreasing period}. As shown by Adenstedt1974, this type of behaviour implies long lasting autocorrelations. The presence of long memory in the data can have perverse effects in estimation and forecasting methods if not included into the modelling scheme, see Beran1994, and Beran2013.

The autoregressive fractionally integrated moving average $(ARFIMA)$ class of models has become one of the most popular methods to model long memory in the time series literature. They have the appeal of bridging the gap between the stationary autoregressive moving average $(ARMA)$ models and the nonstationary autoregressive integrated moving average $(ARIMA)$ model. $ARFIMA$ models rely on the fractional difference operator to introduce long memory behaviour. Nonetheless, there are currently no economic or financial reasonings implying the fractional difference operator.

One of the reasons behind the reliance of the time series literature on generating long memory by the fractional difference operator is the existence of efficient algorithms for their simulation and forecasting. In general, these type of algorithms are not available for other long memory generating schemes. Thus, in this paper, we develop algorithms for memory generation and forecasting by cross-sectional aggregation \`a la Granger1980.

We then contrast the properties of cross-sectionally aggregated processes to the ones obtained by the fractional difference operator. It is shown that cross-sectional aggregated processes are more flexible than fractionally differenced ones, in the sense that more short term dynamics may be included.

Moreover, we show that the antipersistent property of negative autocorrelation for negative degrees of memory does not apply to long memory generated by cross-sectional aggregation. We show that this has repercussions for long memory estimators in the frequency domain.

Finally, this paper evaluates the forecasting power of the $ARFIMA$, and high-order $AR$ models when forecasting long memory generated by cross-sectional aggregation. It finds that high-order $AR$ models beat a pure fractional difference process, $I(d)$, in terms of forecasting performance for some cases. Nonetheless, allowing for short term dynamic in the form of an $ARFIMA(1,d,0)$ model produces comparable forecast performance as high-order $AR$ models while relying on fewer parameters.

This paper proceeds as follows. In Section (ref), we present the fractional difference operator commonly used to model long memory in the time series literature. Section (ref) discusses cross-sectional aggregation as the theoretical explanation behind the presence of long memory in the data and develops a fast algorithm for its generation. Moreover, it contrasts properties of the fractional difference operator against the cross-sectional aggregation scheme. Section (ref) discusses the antipersistent property in the context of cross-sectionally aggregated processes. Section (ref) constructs minimum squared error forecasts, and studies the theoretical performance of high-order $AR$, and $ARFIMA$ models when forecasting long memory generated by cross-sectional aggregation. Section (ref) concludes.

The Fractional Difference Operator

The $ARFIMA$ specification due to Granger1980b, and Hosking1981 has become the standard model to study long memory in the time series literature. They extended the $ARMA$ model to include long memory dynamics by introducing the fractional difference operator

equation*[equation* omitted — 45 chars of source]

where $\varepsilon_t$ is a white noise process, and $d\in(-1/2,1/2)$. Following the standard binomial expansion, they decompose the fractional difference operator, $(1-L)^d$, to generate a series given by

equation[equation omitted — 86 chars of source]

with coefficients $\pi_j=\Gamma(j+d)/(\Gamma(d)\Gamma(j+1))$ for $j\in\mathbb{N}$. Using Stirling's approximation, it can be shown that these coefficients decay at a hyperbolic rate, $\pi_j\approx j^{d-1}$, which in turn translates to slowly decaying autocorrelations. We write $x_t\sim I(d)$ to denote a process generated by equation ((ref)), that is, an integrated process with parameter $d$.

The properties of the $ARFIMA$ model have been well documented in, among others, Baillie1996, and Beran2013. Moreover, fast algorithms have been developed to generate series using the fractional difference operator, see Jensen2014. Thus, the $ARFIMA$ model has become the canonical construction for modelling and forecasting long memory in the time series literature.

Even though the $ARFIMA$ provides a good representation of long memory, and bridges the gap between the stationary $ARMA$ models and the non-stationary $ARIMA$ model, to the best of our knowledge, there is no theoretical reasoning linking the $ARFIMA$ model with the memory found in real data. That is, in contrast with the complete market hypothesis implying that stock prices follow a random walk, or capital depreciation suggesting an autoregressive process, there are no economic or financial arguments for the fractional difference operator to occur in real data. In this regard, the next section presents the most common theoretical motivation behind long memory in the time series literature, cross-sectional aggregation.

Cross-Sectional Aggregation

Granger1980, in line with the work of Robinson1978 on autoregressive processes with random coefficients, showed that aggregating $AR(1)$ processes with coefficients sampled from a Beta distribution can produce long memory.

Define a cross-sectional aggregated series as

equation[equation omitted — 79 chars of source]

where the $N$ individual series are generated as $$x_{i,t} = \alpha_i x_{i,t-1}+\epsilon_{i,t}\ \ \ i=1, 2, \cdots, N;$$ where $\varepsilon_{i,t}$ is an $i.i.d.$ process with $E[\epsilon_{i,t}^2] = \sigma_\epsilon^2$. Moreover, $\alpha_i^2$ is sampled from a Beta distribution with parameters $a,b>0$, with density given by $$\mathcal{B}(\alpha; a, b) = \frac{1}{B(a,b)} \alpha^{a-1}(1-\alpha)^{b-1}\ \ \ \ \text{for}\ \ \ \alpha\in(0,1),$$ where $B(\cdot,\cdot)$ is the Beta function.

Granger showed that, as $N\to\infty$, the autocorrelation function of $x_t$ decays at a hyperbolic rate with parameter $d=1-b/2$. Thus, $x_t$ has long memory.

The cross-sectional aggregation result has been extended in several directions, including to allow for general $ARMA$ processes, as well as to other distributions. See for instance, Linden1999, Oppenheim2004, and Zaffaroni2004. As argued by Haldrup2017, maintaining the Beta distribution allows us to have closed form representations. Furthermore, Beran2010 proposed a method to estimate the parameters from the Beta distribution from the individual series, $x_{i,t}$ above.

In applied work, cross-sectional aggregation has been cited as source of long memory for many series. To name but a few, it has been proposed for inflation, output, and volatility; see Balcilar2004, Diebold1989, and Osterrieder2015.

Even though a cross-sectional aggregated process has long memory, Haldrup2017 show that the resulting series does not belong to the $ARFIMA$ class of processes. Pointedly, they show that the fractionally differenced cross-sectional aggregated process does not follow an $ARMA$ specification.

The next subsections expands on the cross-sectional aggregation literature by developing fast algorithms for its generation, and contrasting its theoretical properties to the ones of the fractional difference operator.

Nonfractional Memory Generation

As can be seen by its definition, Equation ((ref)), generating cross-sectionally aggregated processes is computationally demanding. For each cross-sectionally aggregated process, we need to simulate a vast number of $AR(1)$ processes, which are in turn computationally demanding. Haldrup2017 suggest that to get a good approximation, the cross-sectional dimension should increase as the sample size. The computational demands are thus particularly large for Monte Carlo type of analysis on cross-sectionally aggregated processes. In what follows, we use the theoretical autocorrelations of the cross-sectionally aggregated process to present an algorithm that makes generation of long memory by cross-sectional aggregation comparable to fractional differencing in terms of computational requirements.

Denote $x_t\sim CSA(a,b)$ to a series generated by cross-sectional aggregation of autoregressive parameters sampled from the Beta distribution $B(a,b)$, Equation ((ref)). The notation makes explicit the origin of the memory by cross-sectional aggregation and its dependence on the two parameters of the Beta distribution. Theorem (ref) develops a fast algorithm to generate cross-sectionally aggregated processes.

teoLet $x_t\sim CSA(a,b)$ defined as in ((ref)), then $x_t$ can be computed as the first $T$ elements of the $(2T-1)\times 1$ vector $$T^{-1} F(\bar{F}\tilde{\phi}\odot \bar{F}\tilde{\varepsilon}),$$ where $\bar{F}$ is the discrete Fourier transform, $T^{-1}F$ is the inverse transform, and '$\odot$' denotes multiplication element by element. Furthermore, $\tilde{z}$ is a $(2T-1)\times 1$ vector given by $\tilde{z}:=[z_0, z_1, \cdots, z_{T-1},$ $0, \cdots, 0]$ for $z_t=\phi_t,\varepsilon_t$ where $\varepsilon_j\sim i.i.d. N(0,\sigma^2_\epsilon)$, and $\phi_j = \left(B(p+j,q)/B(p,q)\right)^{1/2},\ \forall j\in\mathbb{N}$.

Proof: See appendix (ref).

The theorem is an application of the circular convolution theorem for the coefficients associated with the cross-sectionally aggregated process. In this sense, it is in line with the algorithm of Jensen2014 for the fractional difference operator and thus achieves equivalent computational efficiency. The difference in computation times between the standard simulation of a cross-sectional aggregated times series, Equation ((ref)), and the fast implementation in Theorem (ref) are very large for all sizes. For small series $(T\approx 10^2)$ the gains are in the hundreds of times faster, while in the thousands for medium sized series $(T\approx 10^4)$. This is of course not surprising given that, for a sample of size $T$, the number of computations needed for the standard implementation is of order $NT^{2}$, with $N$ the cross-sectional dimension, which, as argued before, should increase as $T$ does. Meanwhile, the computational requirements for the fast implementation are of order $T \log{(T)}$. Codes implementing the algorithm for memory generation by cross-sectional aggregation are available in Appendix (ref), and on the author's github repository at \href{https://github.com/everval}{github.com/everval/Nonfractional}.

To get a better understanding of the dynamics of the long memory by cross-sectional aggregation, the next sections compare it against the pure fractional noise, that is, a fractionally differenced white noise. However, note that we can allow more short-term dynamic by adding $AR$ and $MA$ filters to both specifications.

Nonfractional Memory and Fractional Difference Operator

One notable difference between a cross-sectionally aggregated process and a fractionally differenced one is the number of parameters needed for its generation. Fractionally differenced processes rely on only one parameter to model the entire series, whilst a cross-sectionally aggregated process uses two. As we will see below, this gives the cross-sectional aggregation procedure more flexibility.

Let $\gamma_{I(d)}(\cdot)$, and $\gamma_{CSA(a,b)}(\cdot)$ be the autocorrelation functions of an $I(d)$, and a $CSA(a,b)$ process, respectively. That is,

equation[equation omitted — 149 chars of source]
equation[equation omitted — 127 chars of source]

which show that both processes have hyperbolic decaying autocorrelations. As argued by Granger1980, the asymptotic behaviour of the autocorrelation function for a cross-sectionally aggregated process, Equation ((ref)), only depends on the second argument of the Beta function. In this context, both processes show the same rate of decay for the autocorrelation function for $b=2(1-d)$.

Moreover, the first argument of the Beta function allows us to introduce short term dynamics. Figure (ref) shows the autocorrelation function of cross-sectional aggregated processes for different values of `$a$', the first parameter of the Beta distribution. The figure shows that as the first parameter gets larger, so does the autocorrelation function for the initial lags. Thus, `$a$' acts as a short memory modulator.

figure[figure omitted — 234 chars of source]

Section (ref) will look at the forecasting performance of $ARFIMA$ models when working with cross-sectionally aggregated processes, but for now suppose we are interested in looking in the other direction. That is, assume we have a process generated by the fractional difference operator and we want to approximate it with a cross-sectionally aggregated process. We show that we can use the first argument to generate such a series.

Consider the loss function

equation[equation omitted — 129 chars of source]

which measures the squared difference between autocorrelations at the first $k$ lags for cross-sectionally aggregated and fractionally differenced processes with the same long memory dynamics, hence we set $b=2(1-d)$ for the cross-sectionally aggregated process.

Minimizing $(\ref{loss:csa2fi})$ with respect to `$a$', allows us to find the cross-sectional aggregated process that best approximates an $I(d)$ one up to lag $k$, while having the same long term dynamics. Given the different forms of the autocorrelation functions, ((ref)), and ((ref)), there is in general not a value of `$a$' that minimizes $(\ref{loss:csa2fi})$ for all values of $k$; for instance, $\min_{a}\mathcal{L}(2,a,0.2)=0.118$, while $\min_{a}\mathcal{L}(30,a,0.2)=0.121$. Moreover, it will depend on $d$.

Nonetheless, selecting a medium sized $k$, say $k\approx 10$, the approximation turns out to be quite good in general. In Figure (ref), we present a filtered $\{\varepsilon_t\}_{t=1}^{10^4}\sim N(0,1)$ vector using the fractional difference operator with long memory parameter $d=0.2$, and using the cross-sectional aggregated algorithm with parameters $a=0.12$, and $b=1.6$, so that they show similar short and long memory behaviours.

figure[figure omitted — 262 chars of source]

The figure shows that the filtered series are quite similar. This behaviour can further be seen on the autocorrelation functions showing similar dynamics. Thus, the figure shows that it is possible to generate cross-sectionally aggregated processes that closely mimic ones due to fractional differencing.

Nonfractional Memory and Antipersistent Processes

It is well known in the long memory literature that the fractional difference operator implies that the autocorrelation function is negative for negative degrees of memory, $d\in (-1/2,0)$. This can be seen in ((ref)) where the sign of $\gamma_{I(d)}(k)$ depends on $\Gamma(d)$ in the denominator, which is negative for $d\in (-1/2,0)$. Furthermore, the behaviour of the spectral density for a fractionally differenced process near the origin is given by

equation[equation omitted — 109 chars of source]

where $c_0$ is a constant. Thus, $f_{I(d)}(\lambda) \to 0$ as $\lambda\to 0$, that is, the fractional difference operator for negative degrees of memory imply a spectral density collapsing to zero at the origin. These properties, among other related ones, has been named {\it antipersistence} in the literature.

We argue that cross-sectionally aggregated processes do not share these features. First, Equation ((ref)) shows that the autocorrelation function for the cross-sectionally aggregated process only depends on the Beta function, which is always positive. Figure (ref) shows the autocorrelation function for both fractionally differenced and cross-sectional aggregated processes for a negative degree of memory, $d=-0.2$. The figure shows that, even though both processes show the same rate of decay in their autocorrelation functions, they have opposite signs.

figure[figure omitted — 306 chars of source]

Then, given the positive sign for all autocorrelations, Theorem (ref) shows that the spectral density for the cross-sectionally aggregated process for negative degrees of memory converges to a positive constant.

teoLet $x_t\sim CSA(a,b)$ defined as in ((ref)) with $b\in(2,3)$ so that the long memory parameter is in the negative range, $d\in(-1/2,0)$, then, the spectral density of $x_t$ at the origin is positive. That is, \begin{equation} f_{CSA(a,b)}(0)=c_{a,b}>0, \end{equation} where $c_{a,b}$ depends on the parameters of the Beta distribution.

Proof: See appendix (ref).

Figure (ref) shows the periodogram, an estimate of the spectral density, for $CSA(a,b)$ and $I(d)$ processes of size $T=10^4$ averaged for $10^4$ replications. The figure shows that the periodogram for both processes show similar patterns for positive degrees of memory, both diverging to infinity at the same rate. Nonetheless, for negative degrees of memory, the periodogram collapses to zero as the frequency goes to zero for the fractionally differenced process, while it converges to a constant for the cross-sectionally aggregated process.

figure[figure omitted — 334 chars of source]

Moreover, the latter property has implications for estimation and inference. In particular, tests for long memory in the frequency domain will be affected. These tests are based on the rate to which the periodogram goes to zero as an estimator for long memory using the log-periodogram regression, see Geweke1983, and Robinson1995a.

The log-periodogram regression is given by $$\log(\hat{f}_X(\lambda_j)) = a-2d \log(\lambda_j)+u_j,$$ where $\hat{f}_X(\cdot)$ is the periodogram, $a$ is a constant, and $u_j$ is the error term. From ((ref)), note that the log-periodogram regression provides an estimate of the long memory parameter, $d$, for fractionally differenced processes. Tests in the frequency domain use this expression to estimate the degree of memory. Nonetheless, as Theorem (ref), and Figure (ref) show, these tests will be misspecified for long memory by cross-sectional aggregation for negative degrees of memory.

To illustrate the misspecification problem, Table (ref) reports the degree of long memory estimated by the method of Geweke1983, $GPH$ hereinafter, for several degrees of memory for both fractionally differenced, and cross-sectionally aggregated processes.

table[table omitted — 842 chars of source]

The table shows that the estimator is relatively close to the true memory for both processes when the memory is positive, if slightly overshooting it for the cross-sectionally aggregated process, as reported by Haldrup2017. This contrasts to the case of negative memory, where the table shows that the estimator remains precise for the fractionally differenced series, while it is incapable of detecting the long memory in the cross-sectionally aggregated processes, with estimators not statistically different from zero. This is of course not surprising in light of Theorem (ref).

In sum, the lack of the antipersistent property in cross-sectionally aggregated processes shows that care must be taken when estimating long memory for negative degrees of memory if the long memory generating mechanism is not the fractional difference operator. Further analysis of negative degrees of memory and antipersistence is a line of inquiry open for further research.

Nonfractional Memory Forecasting

In the time series literature, some effort has been directed to assess the performance of the $ARFIMA$ type of models when forecasting long memory processes. For instance, Ray1993 calculates the percentage increase in mean-squared error ($MSE$) from forecasting $I(d)$ series with $AR$ models. She argues that the $MSE$ may not increase significantly, particularly when we do not know the true long memory parameter. Crato1996 compare the forecasting performance of $ARFIMA$ models against $ARMA$ alternatives and find that $ARFIMA$ models are in general outperformed by $ARMA$ alternatives. Moreover, Man2003 argues that an $ARMA(2,2)$ model compares favourably to an $I(d)$ for short-term forecasts of long memory time series with fractionally differenced structure.

One thing that these forecasting comparison studies have in common is the underlying assumption that long memory is generated by an $ARFIMA$ process. In this context, a priori the forecasting exercises assume that a fractionally differenced process is the correct specification. As previously discussed, even though cross-sectional aggregation does generate long memory, the series does not follow an $ARFIMA$ specification. The question remains whether an $ARFIMA$ model serves as a good approximation for forecasting purposes.

The next subsection computes the minimum square error forecasts for cross-sectional aggregated processes, $CSA(a,b)$. With those as benchmark, the following subsections evaluate the forecasts performance of high order $AR$, and $ARFIMA$ models on $CSA(a,b)$ processes.

Minimum Square Error Forecasts

Theorem (ref) computes the minimum mean square error forecasts for a series generated by cross-sectional aggregation.

teoLet $x_t\sim CSA(a,b)$ defined as in ((ref)), and let $h\in\mathbb{N}$, then the minimum mean square error forecast $h$ periods ahead, $\hat{x}_{t+h}$, can be computed as\footnote{The Theorem assumes a Type II process analogous to samples of fractionally differenced processes, see Davidson2009. In particular, it assumes $\nu_{j}=0,\ \ \forall j <0$.} \begin{equation} \hat{x}_{T+h} = \sum_{j=h}^{T}\phi_{j}\nu_{T-j+h}, \end{equation} where $\nu_{i}$ is computed as \begin{equation*} \nu_{i} = x_i - \sum_{j=1}^{i}{\phi_j \nu_{i-j}} \ \ \ \forall i\in{0,1,\cdots,T}, \end{equation*} with $\phi_j$ as in Theorem (ref).

Proof: See Appendix (ref).

The theorem relies on the $MA(\infty)$ representation to construct the forecasts. Algorithms for computing the minimum mean square forecasts, ((ref)), are presented in Appendix (ref), and are available on the author's github repository at \href{https://github.com/everval}{github.com/everval/Nonfractional}.

Theorem (ref) allows for the construction of forecasts for the correct theoretical specification. Note that in the theorem we have assumed that we know the parameters of the the Beta distribution for the cross-sectionally aggregated process; for empirical applications we can use Beran2010's (Beran2010) method to estimate them. In the following, we will use these computations to assess the forecasting performance of high-order $AR$, and $ARFIMA$ models when working on long memory generated by cross-sectional aggregation.

Forecasts With AR(p) Models

This subsection computes the forecasting efficiency loss of an $AR(p)$ model when working with a $CSA(a,b)$ process. Theorem (ref) obtains the parameter estimates for an $AR(p)$ model fitted to a $CSA(a,b)$ process, and computes the efficiency loss for one-step ahead forecasts.

teoLet $x_t\sim CSA(a,b)$ defined as in ((ref)), and estimate an $AR(p)$ given by $(1-\alpha_1 L - \alpha_2 L^2 -\cdots -\alpha_p L^p )x_t=u_t$, then \begin{equation*} \begin{bmatrix} \alpha_1 \\ \alpha_2 \\ \vdots \\ \alpha_P \end{bmatrix} = \begin{bmatrix} 1 &\gamma_{CSA(a,b)}(1) &\cdots &\gamma_{CSA(a,b)}(p-1) \\ \gamma_{CSA(a,b)}(1) & 1 &\cdots &\gamma_{CSA(a,b)}(p-2)\\ \vdots &\vdots &\ddots &\vdots \\ \gamma_{CSA(a,b)}(p-1)&\gamma_{CSA(a,b)}(p-2) &\cdots &1\end{bmatrix}^{-1} \begin{bmatrix} \gamma_{CSA(a,b)}(1) \\ \gamma_{CSA(a,b)}(2) \\ \vdots \\ \gamma_{CSA(a,b)}(p) \end{bmatrix}, \end{equation*} with $\gamma_{CSA(a,b)}(\cdot)$ defined as in (ref). Furthermore, the one-step ahead forecast error variance relative to the minimum square error forecasts, $\zeta_{AR(p)}$, is given by $$\zeta_{AR(p)} = \left(\frac{B(a,b-1)}{B(a,b)}\right)\left[ \left(1 + \sum_{i=1}^{p}{\alpha_i^2}\right) + 2\sum_{i=1}^{p}{\gamma_{CSA(a,b)}(i)\left(-\alpha_i+\sum_{j=1}^{p-i}{\alpha_j\alpha_{j+i}}\right)}\right].$$

Proof: See Appendix (ref).

As an example, consider estimating an $AR(1)$ model to forecast a cross-sectionally aggregated process. From Theorem \ref*{teo:arp}, the autoregressive parameter is $\alpha_1 = \gamma_{CSA(a,b)}(1) = B(a+1/2,b-1)/B(a,b-1)$, while the one-step ahead forecast error variance is $\zeta_{AR(1)}=(B(a,b-1)/B(a,b))(1-\gamma_{CSA(a,b)}^2(1))$, which shows that the efficiency loss is a nonlinear function of both parameters of the $CSA(a,b)$ process.

To get a better sense of the nonlinearity, Figure (ref) presents both the autoregressive parameter, $\alpha_1$, and the efficiency loss, $\zeta_{AR(p)}$, while varying the first parameter of the cross-sectional aggregated process, `$a$', with fixed $b=1.8$. On the one hand, the figure shows that as `$a$' increases, so does the autoregressive parameter, this is in line with the discussion above regarding the first parameter as a short memory regulator. On the other hand, the one-step ahead forecast error variance shows a maximum around $a=0.5$, with a forecast error variance of almost 15%. That is, more than double the reported by Man2003 for an $I(0.1)$ process. Thus, the figures suggest that we could expect worse performance using an $AR(1)$ process to forecast a $CSA(a,b)$ process than a pure $I(d)$ one.

figure[figure omitted — 309 chars of source]
table[table omitted — 1,165 chars of source]

Table (ref) shows the one-step ahead forecast error variance of fitted $AR(1)$ models on $CSA(a,b)$ processes for different values of $a,b$. The losses are in line with the ones computed by Man2003 for the $I(d)$ case, increasing as the memory of the process increases. Moreover, as noted by Ray1993 and Man2003, we can reduce the one-step ahead forecast error variance by allowing for more lags in the $AR$ specification. To get a sense of the improvements we can achieve, Table (ref) also presents the one-step ahead forecast error variance of $AR(20)$ models fitted to $CSA(a,b)$ processes for different values of $a,b$. As the table shows, increasing the order of the $AR$ model can greatly reduce the one-step ahead forecast error variance, particularly for larger degrees of long memory.

Forecasts With Fractional Models

This section studies the forecast performance of fractional models when working on cross-sectional aggregated processes. In particular, we compute the one-step ahead forecast error variance of the pure $I(d)$ model and of an $ARFIMA(1,d,0)$ allowing for more short term dynamics.

teoLet $x_t\sim CSA(a,b)$ defined as in ((ref)), then the one-step ahead forecast error variance of the $d=1-b/2$ fractional difference of the series, $\zeta_{I(d)}$, is given by $$\zeta_{I(d)} = \gamma_z(0).$$ Furthermore, estimate an $ARFIMA(1,d,0)$ by fitting an $AR(1)$ model to the $d=1-b/2$ fractional difference of the series, then the autoregressive parameter, $\alpha_{I}$, and the one-step ahead forecast error variance, $\zeta_{ARFIMA(1,d,0)}$, are given by $$\alpha_{I} = \frac{\gamma_z(1)}{\gamma_z(0)},\ \ \ \ \zeta_{ARFIMA(1,d,0)} = \frac{\gamma_{z}(0)^2-\gamma_{z}(1)^2}{\gamma_{z}(0)^2}.$$ Where in both expressions $\gamma_z(k)$ is given by $$\gamma_z(k) = \frac{\gamma^*(k)}{B(a,b)}\left[B(a,b-1)\left(F_{1}(k)-1\right)+B(a+\frac{1}{2},q-1)F_{2}(k)\right],$$ with $$\gamma^*(k) = \sigma_{\varepsilon}^2\frac{\Gamma(1+2d)}{\Gamma(-d)\Gamma(1+d)}\frac{\Gamma(-d-k)}{\Gamma(1+d-k)},$$ and \begin{eqnarray*} F_{1}(k) &:= &F\left[\left\{1,a,\frac{1-d+k}{2},\frac{-d+k}{2}\right\},\left\{a+b-1,\frac{2+d+k}{2},\frac{1+d+k}{2}\right\},1\right]+\\ &&F\left[\left\{1,a,\frac{1-d-k}{2},\frac{-d-k}{2}\right\},\left\{a+b-1,\frac{2+d-k}{2},\frac{1+d-k}{2}\right\},1\right],\\ F_{2}(k) &:= &\frac{-d+k}{1+d+k}*\\ &&F\left[\left\{1,p+\frac{1}{2},\frac{1-d+k}{2},\frac{2-d+k}{2}\right\},\left\{a+b-\frac{1}{2},\frac{2+d+k}{2},\frac{3+d+k}{2}\right\},1\right]\\ &&+\frac{-d-k}{1+d-k}*\\ &&F\left[\left\{1,a+\frac{1}{2},\frac{1-d-k}{2},\frac{2-d-k}{2}\right\},\left\{a+b-\frac{1}{2},\frac{2+d-k}{2},\frac{3+d-k}{2}\right\},1\right],\\ \end{eqnarray*} where $d=1-b/2$, and $F[\cdot]$ is the generalized hypergeometric function.

Proof: See Appendix (ref).

Table (ref) presents the one-step ahead forecast error variance of fitted $ARFIMA(1,d,0)$, and $I(d)$ models on $CSA(a,b)$ processes for different values of $a,b$. The table also shows the estimated autoregressive parameter of the fitted $ARFIMA(1,d,0)$ model.

As the table shows, the pure fractional differenced process can be quite bad at forecasting $CSA(a,b)$ processes, specially for higher values of `$a$'. Once again, relating the first parameter functioning as a short memory regulator. Once we allow for more short term dynamics in the form of an $ARFIMA(1,d,0)$ model, the forecasting performance is much in line with the one from high order $AR$ models. Thus, it greatly helps to allow for some short term dynamics in the modelling scheme.

table[table omitted — 1,598 chars of source]

Comparing the results from Tables (ref) and (ref), we see that, as was the case for long memory series generated by the fractional difference operator, a high order $AR(p)$ can be a good model for forecasting long memory at short horizons. Yet, the bias-variance trade-off has to be assessed given the number of estimated parameters. As an alternative, the $ARFIMA(1,d,0)$ produces similar results while relying in only two parameters.

Conclusions

Even though there is no theoretical argument linking the presence of long memory in the data with the fractional difference operator, fractionally differenced processes remain the most popular construction in the long memory time series literature. This may be due to the existence of efficient algorithms for their simulation and forecasting. Thus, this paper presents fast algorithms to generate long memory by a theoretically based mechanism that do not rely on the fractional difference operator, cross-sectional aggregation.

The cross-sectional aggregated process is then contrasted to the fractional difference operator. In particular, the paper analyses the antipersistent phenomenon. It is proven that, for negative degrees of memory, while the autocorrelations are negative by definition for the fractional difference operator, this restriction does not apply to the cross-sectional aggregated scheme. Furthermore, the paper shows that the lack of antipersistence for cross-sectional aggregated processes has implications for long memory estimators in the frequency domain which will be misspecified in general.

Moreover, this work evaluated the efficiency loss of using high-order $AR$ and $ARFIMA$ models to forecast long memory series generated by cross-sectional aggregation. It finds that, at short horizons, high-order $AR$ models beat a pure fractional difference process, $I(d)$, in terms of forecasting performance. Nonetheless, allowing for short term dynamic in the form of an $ARFIMA(1,d,0)$ model produces comparable forecast performance as high-order $AR$ models while relying on less parameters.

The results of this paper can be used in the context of Monte Carlo simulations of long memory estimators and forecasts. Of particular interest is the analysis of long memory in inflation, one of the primer examples of cross-sectional aggregation producing long memory due to the way the data is computed. Moreover, the results allow for financial econometrics and climate econometrics models to incorporate long memory forecasts consistent with the theoretical argument posed by cross-sectional aggregation.

Acknowledgements

I would like to thank Niels Haldrup for the careful reading of this article and all the insightful comments.

thebibliography\bibitem[\astroncite{Adenstedt}{1974}]{Adenstedt1974} Adenstedt, R. K. (1974). \newblock {On Large-Sample Estimation for the Mean of a Stationary Random Sequence}. \newblock {\em The Annals of Statistics}, 2(6):1095--1107. \bibitem[\astroncite{Baillie}{1996}]{Baillie1996} Baillie, R. T. (1996). \newblock {Long memory processes and fractional integration in econometrics}. \newblock {\em Journal of Econometrics}, 73(1):5--59. \bibitem[\astroncite{Balcilar}{2004}]{Balcilar2004} Balcilar, M. (2004). \newblock {Persistence in inflation: does aggregation cause long memory?} \newblock {\em Emerging Markets Finance and Trade}, 40(5):25--56. \bibitem[\astroncite{Beran}{1994}]{Beran1994} Beran, J. (1994). \newblock {\em {Statistics for long-memory processes}}. \newblock Chapman {&} Hall. \bibitem[\astroncite{Beran et al.}{2013}]{Beran2013} Beran, J., Feng, Y., Ghosh, S., and Kulik, R. (2013). \newblock {\em {Long-Memory Processes: probabilistic theories and Statistical Methods}}. \newblock Springer. \bibitem[\astroncite{Beran et al.}{2010}]{Beran2010} Beran, J., Schtzner, M., and Ghosh, S. (2010). \newblock {From short to long memory: Aggregation and estimation}. \newblock {\em Computational Statistics and Data Analysis}, 54(11):2432--2442. \bibitem[\astroncite{Crato and Ray}{1996}]{Crato1996} Crato, N. and Ray, B. K. (1996). \newblock {Model selection and forecasting for long-range dependent processes}. \newblock {\em Journal of Forecasting}, 15(2):107--125. \bibitem[\astroncite{Davidson}{2009}]{Davidson2009} Davidson, J. (2009). \newblock {When is a time series I(0)?} \newblock In Castle, J. and Shepherd, N., editors, {\em The Methodology and Practice of Econometrics}, chapter 13. Oxford University Press. \bibitem[\astroncite{Diebold and Rudebusch}{1989}]{Diebold1989} Diebold, F. X. and Rudebusch, G. D. (1989). \newblock {Long Memory and Persistence in Agregate Output}. \newblock {\em Journal of Monetary Economics}, 24(2):189--209. \bibitem[\astroncite{Geweke and Porter-Hudak}{1983}]{Geweke1983} Geweke, J. and Porter-Hudak, S. (1983). \newblock {The estimation and application of long memory time series models.} \newblock {\em Journal of Time Series Analysis}, 4(4):221--238. \bibitem[\astroncite{Granger}{1966}]{Granger1966} Granger, C. W. (1966). \newblock {The Typical Spectral Shape of an Economic Variable}. \newblock {\em Econometrica}, 34(1):150--161. \bibitem[\astroncite{Granger}{1980}]{Granger1980} Granger, C. W. (1980). \newblock {Long memory relationships and the aggregation of dynamic models}. \newblock {\em Journal of Econometrics}, 14(2):227--238. \bibitem[\astroncite{Granger and Joyeux}{1980}]{Granger1980b} Granger, C. W. and Joyeux, R. (1980). \newblock {An Introduction to Long Memory Time Series Models and Fractional Differencing}. \newblock {\em Journal of Time Series Analysis}, 1(1):15--29. \bibitem[\astroncite{Haldrup and {Vera Vald{\'{e}}s}}{2017}]{Haldrup2017} Haldrup, N. and {Vera Vald{\'{e}}s}, J. E. (2017). \newblock {Long memory, fractional integration, and cross-sectional aggregation}. \newblock {\em Journal of Econometrics}, 199(1):1--11. \bibitem[\astroncite{Hosking}{1981}]{Hosking1981} Hosking, J. R. M. (1981). \newblock {Fractional differencing}. \newblock {\em Biometrika}, 68(1):165--176. \bibitem[\astroncite{Jensen and Nielsen}{2014}]{Jensen2014} Jensen, A. N. and Nielsen, M. {\O}. (2014). \newblock {A Fast Fractional Difference Algorithm}. \newblock {\em Journal of Time Series Analysis}, 35(5):428--436. \bibitem[\astroncite{Linden}{1999}]{Linden1999} Linden, M. (1999). \newblock {Time series properties of aggregated AR(1) processes with uniformly distributed coefficients}. \newblock {\em Economics Letters}, 64(1):31--36. \bibitem[\astroncite{Man}{2003}]{Man2003} Man, K. S. (2003). \newblock {Long memory time series and short term forecasts}. \newblock {\em International Journal of Forecasting}, 19(3):477--491. \bibitem[\astroncite{Oppenheim and Viano}{2004}]{Oppenheim2004} Oppenheim, G. and Viano, M. C. (2004). \newblock {Aggregation of random parameters ornstein-uhlenbeck or ar processes: Some convergence results}. \newblock {\em Journal of Time Series Analysis}, 25(3):335--350. \bibitem[\astroncite{Osterrieder et al.}{2015}]{Osterrieder2015} Osterrieder, D., Ventosa-Santaul{\`{a}}ria, D., and Vera-Vald{\'{e}}s, J. E. (2015). \newblock {Unbalanced Regressions and the Predictive Equation}. \newblock {\em CREATES Research Paper}, 9. \bibitem[\astroncite{Ray}{1993}]{Ray1993} Ray, B. K. (1993). \newblock {Modeling Long Memory Processes for Optimal Long Range Prediction}. \newblock {\em Journal of Time Series Analysis}, 14(5):511--525. \bibitem[\astroncite{Robinson}{1978}]{Robinson1978} Robinson, P. M. (1978). \newblock {Statistical Inference for a Random Coefficient Autoregressive Model}. \newblock {\em Scandinavian Journal of Statistics}, 5(3):163--168. \bibitem[\astroncite{Robinson}{1995}]{Robinson1995a} Robinson, P. M. (1995). \newblock {Log-Periodogram Regression of Time Series with Long Range Dependence}. \newblock {\em The Annals of Statistics}, 23(3):1048--1072. \bibitem[\astroncite{Zaffaroni}{2004}]{Zaffaroni2004} Zaffaroni, P. (2004). \newblock {Contemporaneous aggregation of linear dynamic models in large economies}. \newblock {\em Journal of Econometrics}, 120(1):75--102.