Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
92,264 characters · 0 sections · 51 citation commands
Modeling Long Cycles
\@startsection{section}{1} \z@{1.0\linespacing\@plus\linespacing}{.8\linespacing}{Introduction}
This paper develops an econometric framework for inference on the cyclical properties of time series. We are particularly interested in stochastic cycles arising from persistent low-frequency oscillatory impulse responses. The period of such cycles spans a substantial fraction of a sample, and the econometrician would be able to observe only a handful of peaks and troughs in the data. We refer to such cycles as “long”.
Long cycles are prevalent in macroeconomic and financial data. In a recent paper, beaudry2020aer estimated that many variables have cycles of approximately 32--40 quarters, corresponding to 15% and even 20% of their observed samples. Using data from 1960 to 2011, DrehmannBorioTsatsaronis2012 estimated that the length of the credit cycle is 18 years or approximately 35% of their sample. The first contribution of our paper is to show that statistical inference may be distorted in such cases. We find that substantial distortions occur when the cycle length exceeds 25% of the sample size.\footnote{See Appendix (ref).}
In our second contribution, we propose a new econometric procedure for inference on the periodicity of cycles. The novel aspect of our methodology is that it is specifically designed to take into account the possibility of long persistent stochastic cycles. Our procedure produces confidence intervals for the cycle length that have the following property: their asymptotic coverage probability is correct regardless of the cycle length. Thus, the confidence intervals are asymptotically valid both when the period is small relative to the sample size and when it spans a substantial fraction of observed data. No other procedure in the existing literature has this property. When a data-generating process (DGP) is acyclical, our procedure is expected to produce empty confidence intervals for the cycle length in large samples. Hence, the procedure can be used to rule out cyclical behavior.
Stochastic cycles arise naturally in the AR(2) model $y_t=\phi_1 y_{t-1}+\phi_2 y_{t-2} +u_t$ with complex roots. For example, sargent1987macroeconomic shows that such a process has a peak in its spectrum in the interior of the $[0,\pi]$ range provided that the autoregressive coefficients satisfy $|\phi_1(1-\phi_2)/4\phi_2|<1$.\footnote{The region with an interior spike in the spectrum is a subset of the region with complex roots sargent1987macroeconomic} The period of such cycles is determined by the autoregressive coefficients through ${2\pi}/{\omega}$, where $\omega=\cos^{-1}(-\phi_1(1-\phi_2)/4\phi_2)$ is the spectrum peak frequency. According to this model, the cycle length would amount to only a negligible fraction of the sample size $n$ in large samples: $\frac{2\pi/\omega}{n}\to 0$. Thus, the long-cycle characteristics of the data are not preserved asymptotically: while in a finite sample the cycle length can represent a substantial fraction of the sample size, it would be negligible in the asymptotic approximation. As a result, conventional asymptotic approximations to the finite sample distributions of estimators and statistics would be inaccurate.
To preserve long-cycle characteristics asymptotically, the spectrum peak frequency must be local to zero in the sense that $n\omega_n$ converges to a positive constant as $n\to\infty$, i.e. the spectrum peak frequency $\omega_n$ “drifts” closer to zero with the sample size. This property can be obtained by modeling the two conjugate complex roots of the AR(2) model as local to one, that is, the real and complex parts of the roots “drift” closer to one with the sample size. Following the literature andrews2020generic, we refer to such specifications and coefficients as “drifting”, while specifications with constant parameters independent of $n$ are referred to as “fixed”.\footnote{Drifting specifications can be used to verify the uniform size properties of inferential procedures andrews2020generic.} The corresponding autoregressive coefficients are too drifting.
Our resulting long-cycle model is a restricted version of the nearly-twice integrated model in perron1996useful, and the restriction is imposed to generate persistent oscillatory behavior. Although the resulting processes are near I(2), they are stationary in finite samples.
Figure (ref) illustrates how drifting specifications allow us to preserve the long-cycle feature asymptotically. It shows the difference between the simulated sample paths of cyclical processes generated using the fixed-coefficient and drifting-coefficient AR(2) DGPs for small and large sample sizes. The figure demonstrates that in the model with fixed autoregressive coefficients, by relying on asymptotic approximations, one would distort the cyclical properties of the data. However, the long-cycle properties are preserved in the limit by relying on asymptotic approximations with drifting coefficients.
The problem is closely related to that in the literature concerned with inference on the largest autoregressive root stock1991confidence,andrews1993exactly,hansen1999grid,elliott2001confidence,mikusheva2007uniform,mikusheva2012one,dou2019generalized.\footnote{Equivalently, the sum of the autoregressive coefficients.} It has been shown in this literature that when autoregressive roots are close to unity, the conventional asymptotic theory does not provide an accurate approximation to the finite sample distributions of estimators and statistics. More accurate approximations can be obtained using the so-called local-to-unity asymptotics developed in phillips1987towards,phillips1988regression. However, despite many similarities, the existing results for local-to-unity processes cannot accommodate long cycles because of the presence of two complex conjugate roots. Our third contribution is to develop a novel asymptotic theory for such processes.\footnote{In a recent paper, dou2019generalized propose a generalized local-to-unity ARMA model with multiple local-to-unity autoregressive roots balanced local-to-unity roots in the moving average component. They do not consider cyclical behavior, and their limiting distributions are different from those arising in our case.} Our results also lay out the foundation for a new econometric framework that, besides the inference on cyclicality, can also be used to study cointegrating long cycles and phase shifts in macro-financial aggregates.
Our paper is also related to the literature on complex unit roots bierens2001complex,gregoir2006efficient. Unlike bierens2001complex, our data generating process is stationary in finite samples, and persistent oscillations are achieved through local-to-unity modeling. Local-to-unity modeling with complex roots has been previously considered by gregoir2006efficient. However, gregoir2006efficient only considers oscillations at fixed frequencies, while we focus on oscillations at local-to-zero frequencies. This crucial feature allows us to accommodate arbitrary long cycles with persistent oscillations at very low frequencies, which is an important attribute of many macroeconomic and financial time series, as demonstrated in Section 7.
The fourth contribution of this paper is empirical, where we implement our procedure to study the cyclical properties of key macroeconomic and financial indicators using U.S. data. Recurrent boom-and-bust cycles are a salient feature of economic and financial history. Long-standing interest in understanding these ups and downs in macro-financial aggregates has led to a vast body of literature on business cycles bry1971front,Harvey1985,AHearnWoitek2001JME,harding2002dissecting,comin2006medium,DrehmannBorioTsatsaronis2012,Aikman_etal2015EJ,Strohsal_etal2019JBF,Runstler2018business. Using our methodology, we find that long cycles cannot be ruled out for macroeconomic series, such as the real GDP per capita, the unemployment rates, and the hours per capita. Our results suggest the possibility of cycles that are much longer than those previously reported in the literature. In addition, we find that financial variables such as credit to the nonfinancial sector and home prices exhibit long cycles that are even longer than those for the macro variables. Our results support the position that financial cycles operate at a lower frequency than business cycles. However, our most striking result is that we decisively reject stochastic cycles for asset market variables such as volatility index, credit risk premium, and equity prices. This suggests that the mechanism for asset market fluctuations is different from that of macroeconomic variables and financial variables such as credit and home prices. Importantly, this finding rejects the view suggested in the macro-finance literature that asset prices and economic fluctuations are driven by the same underlying forces: time-varying risk premiums and risk-bearing capacity cochrane2017macro.
The remainder of the paper is organized as follows. In Section (ref), we present our modeling approach for long cycles. Section (ref) presents our core asymptotic results. The results are extended in Section (ref) to allow for linear time trends and deterministic cycles. Section (ref) describes our procedure for constructing confidence intervals for the cycle length. Section (ref) presents our empirical results. In Appendix (ref), we show that the periodogram is an asymptotically biased estimator in the presence of long cycles. In Appendix (ref), we discuss the size distortion from using conventional $\chi^2$ critical values for inference. In Appendix (ref), we show the consistency of the Bayesian Information Criterion (BIC) for the specification of models with long cycles, and then study the finite-sample properties of our proposed inferential procedures using Monte Carlo simulations in (ref).
\@startsection{section}{1} \z@{1.0\linespacing\@plus\linespacing}{.8\linespacing}{A model for long cycles}
In this section, we present a model for processes that exhibit long cycles. Our objective is to develop a parsimonious modeling approach that allows for cycles with periods spanning nonnegligible fractions of observed samples. More formally, the model should allow the period as a fraction of the sample size $n$ to converge to a nonzero constant as $n\to \infty$.
Following sargent1987macroeconomic, we consider the class of autoregressive models. Since in this class cyclical behavior requires complex roots that come as conjugate pairs, the AR(2) model with serially uncorrelated errors\footnote{We extend the approach later in the paper to allow for serially correlated errors.} is a natural starting point. Thus, consider a process $\{y_t\}$ generated according to
where $L$ denotes the lag operator and $\{u_t\}$ is a mean-zero i.i.d. sequence with a finite variance. Let $\lambda_1$ and $\lambda_2$ denote the roots of the characteristic equation $z^2-\phi_{1} z - \phi_{2} =0$ for the lag polynomial in (ref). When $|\lambda_1|<1$ and $|\lambda_2| < 1$, $\{y_t\}$ has the following MA($\infty$) representation:
Suppose the roots $\lambda_1,\lambda_2$ are complex, and consider their polar form representation:
where $r$ denotes the modulus, $\theta$ is the argument of the complex roots, and $i=\sqrt{-1}$ is the imaginary number.
Given the polar coordinate representation for the roots and the MA($\infty$) representation for the process, we can write $\{y_t\}$ as
According to (ref), the realized value of $y_t$ is a weighted infinite sum of past realizations of the innovation sequence $\{u_t\}$. When the characteristic roots are complex, the weights or impulse responses are given by a damped sine wave: the impulse response of $y_t$ to $u_{t-j}$ is
where the modulus $r$ indicates the rate of decay\footnote{In a more common exponential decay representation, $r^j = e^{\ln(r)j}$. Restricting to processes with non-explosive roots, i.e. $r\leq 1$, we have $\ln(r) \leq 0$ and $-\ln(r)$ is known as the decay constant. } or the persistence of the sine wave, and the argument $\theta$ corresponds to the angular frequency and determines the period of the sine wave. The latter has been used in the literature as a measure of the frequency of the cycle Harvey1985.
The stochastic process $\{y_t\}$ inherits its oscillatory behavior precisely from this damped periodic sine weighting function. The closer $r$ is to one, the more persistent $\{y_t\}$, and the closer $\theta$ is to zero, the lower the oscillating frequency and the longer the length of cycles in $\{y_t\}$. Stochastic cycles are therefore conveniently captured in an AR(2) model with a pair of complex conjugate roots.
A cyclical process generated according to (ref) with roots given by (ref) has the expected cycle length of $2\pi/\theta$. With any fixed parameter value $\theta$, the period as a fraction of the sample size is negligible for large $n$. Hence, asymptotic approximations assuming fixed values for the argument $\theta$ can produce distinctly different cyclical behavior from that observed in finite samples.\footnote{This point is illustrated in Figure (ref).} In other words, conventional asymptotics with a fixed complex root argument $\theta$ can distort the cyclical properties of the process. As a result, such asymptotic theory would provide a poor approximation to the actual behavior of the process in finite samples. Since the expected period of the process as a fraction of the sample size is given by $2\pi/(\theta n)$, to preserve the cyclical properties in the limit as $n\to\infty$, one has to consider a drifting sequence of the arguments $\{\theta_n\}$ and allow for $n\theta_n\to d\in[0,\infty]$.\footnote{The approach can still accommodate conventional asymptotics by allowing $n\theta_n \to \infty$.}
We rewrite the AR(2) model in (ref) as follows:\footnote{As in phillips1987towards,phillips1988regression, the solution to the difference equation (ref) is a triangular array of the form $\{y_{n,t}:t=1,\ldots,n;n\geq 1\}$. However, we suppress the subscript $n$ to simplify the notation.}
where $\phi_{1,n}$ and $\phi_{2,n}$ are now drifting coefficients that can change with $n$. We denote the corresponding characteristic roots by $\lambda_{1,n}$ and $\lambda_{2,n}$ and make the following assumption.
Assumption (ref) excludes positive values of $c$ as they correspond to explosive roots. Negative values of $d$ can be excluded as $d$ and $-d$ define the same pair of roots. The autoregressive coefficients are related to the characteristic roots through the following equations:
Hence, the autoregressive coefficient $\phi_{1,n}$ is local to 2 while $\phi_{2,n}$ is local to $-1$. The sum of the autoregressive coefficients is local to one, and the process can be mistaken for those considered in the local-to-unity literature. As we discuss below, processes defined by (ref)--(ref) exhibit persistent stochastic oscillations and, as a result, their asymptotic properties are different from those considered in the local-to-unity literature.
The expressions for the characteristic roots in (ref) and (ref) are equivalent since one can replace $r$ and $\theta$ with $r_n\equiv\exp(c/n)$ and $\theta_n\equiv d/n$ respectively. The modulus $r_n$ in (ref) has the same representation as the autoregressive parameter in the local-to-unity model in phillips1987towards,phillips1988regression. Since there are two roots near one, the process defined by Assumption (ref) can be viewed as near I(2). The model is a restricted version of the nearly-twice integrated model of perron1996useful, where the restriction is imposed to generate complex roots and persistent cycles.\footnote{perron1996useful are primarily concerned with unit root testing and propose modifications designed to improve the size properties of some commonly used tests.}
The localization parameter $d$ controls the length of the cycle, where long cycles correspond to values of $d$ near zero, while $c$ controls its persistence. Under Assumption (ref), the expected period as a fraction of the sample size is given by
With this parameterization, the length of the cycle as a fraction of the sample size is independent of the sample size, and the resulting asymptotic approximations preserve the cyclical properties of the process.
A process with a cyclical factor in its impulse response coefficients may not have visible cyclical oscillations when the damping effect of $r^j$ in (ref) is too strong. An alternative but related measure of the periodicity of a process can be constructed from its spectral properties kaiser2001measuring. The advantage of this measure is that it also takes into account the persistence properties, unlike those based solely on the argument $\theta_n$ of the complex roots. Let $\omega^*_n$ denote the frequency that maximizes the spectral density of the process in (ref). As in sargent1987macroeconomic,
provided that
The condition in (ref) together with (ref)--(ref) imply that for sufficiently large sample sizes $n$, the spectrum has a peak away from the origin if\footnote{The result in (ref) follows from (ref) and (ref)-(ref) using second-order expansions of $\cos(d/n)$ and $\exp(c/n)$.}
The corresponding period of the process as a fraction of the sample size is given by
The following proposition provides an asymptotic approximation of the length of a cycle as a fraction of the sample size when measured using the spectrum-based approach.
The proposition shows that, when using spectrum-based measures of the period, the length of a cycle relatively to the sample size can be approximated by
Unlike the angular frequency-based measure $\tau_\theta$, the spectrum-based measure $\tau_\omega$ takes into account the persistence of the cycle as captured by the localization parameter $c$. Note that larger negative values of $ c < 0$ produce less persistent cycles. In such cases, the spectrum's peak is closer to the origin, and as a result, less persistent processes may not exhibit any visible cyclical behavior.
In this paper, we consider both $\tau_\theta$ and $\tau_\omega$ since both types of cyclicality measures are used in the literature Harvey1985,sargent1987macroeconomic,kaiser2001measuring. If one is only concerned with the presence of a cyclical factor in the impulse response coefficients, $\tau_\theta$ is appropriate. However, if in addition one is interested in cyclicality with persistence strong enough to produce a visible spectrum spike away from zero, then more stringent conditions on the autoregressive coefficients are required, and $\tau_\omega$ is more appropriate. The restriction $d\geq |c|$, which is required for $\tau_\omega$ to be defined, ensures that the cycle is persistent enough to produce a peak in the spectrum.
In Appendix (ref), we show that the periodogram-based estimation approach produces biased estimates of $\tau_\omega$. Therefore, in the following we develop an inference procedure for $\tau_\theta$ and $\tau_\omega$ based on the estimates of the autoregressive coefficients $\phi_{1,n}$ and $\phi_{2,n}$. For this purpose, we proceed in two steps. First, we develop a procedure to construct asymptotically valid confidence sets for the autoregressive parameters $\phi_{1,n}$ and $\phi_{2,n}$. In the second step, we use projection arguments to build confidence intervals for the $\tau_\theta$ and $\tau_\omega$ measures of the length of a cycle. The theory developed below also allows serially correlated $\{u_t\}$.
\@startsection{section}{1} \z@{1.0\linespacing\@plus\linespacing}{.8\linespacing}{Asymptotics for long-cycle processes} We now provide the asymptotic theory for the process defined in equations (ref)--(ref). The theory will be subsequently used to establish the asymptotic distributions of regression-based statistics involving long-cycle time series. It is also necessary to develop robust and asymptotically valid inference about the cyclical properties of a process.
The specification proposed in equations (ref)--(ref) is similar to the first-order autoregressive local-to-unity root model in phillips1987towards. Assuming that a process $\{x_{t}\}$ is generated according to $x_{t} = a_n x_{t-1} + u_t$ with $a_n = \exp(c/n)$, phillips1987towards shows that the distribution of $\{x_{t}\}$ can be approximated by an Ornstein-Uhlenbeck diffusion process:
where \[J_c(r) \equiv \int_0^{r}e^{c(r-s)}\dd W(s), \] $r\in[0,1]$, $\floor{x}$ denotes the largest integer less or equal to $x$, $W(\cdot)$ is a standard Brownian motion, $\sigma^2$ denotes the limit of the long-run variance of $\{u_t\}$, “$\Rightarrow$" denotes the weak convergence of probability measures, and it is assumed that $\{u_t\}$ satisfies a Functional Central Limit Theorem (FCLT). Note that the distribution of the Ornstein-Uhlenbeck process $J_c(r)$ depends on the localization parameter $c$. In what follows, we build on these insights.
We make the following assumption on the innovation sequence $\{u_t\}$.\footnote{Assumption (ref) holds, for example, when $\{u_t\}$ is a mixing process such that $E(u_t) =0$ for all $t$, $\sup_t E|u_t|^{\beta + \epsilon} \leq \infty$ for some $\beta < 2$ and $\epsilon >0$, and $\{u_t\}$ is $\alpha$-mixing of size $-\beta/(\beta-2)$ phillips1987towards. Alternatively, it holds when $\{u_t\}$ is a linear MA($\infty$) process satisfying the conditions in phillips1992asymptotics.}
In the case of long cycles, a different limiting process arises from that of the local-to-unity case, with cyclicality reflected by the sine function. However, similarly to the results in phillips1987towards, the process can be described as an integral with respect to a Brownian motion and depends on the localization parameters $c$ and $d$. We define:
The next proposition shows that in large samples and after appropriate scaling, the distribution of a long-cycle process can be approximated by that of $J_{c,d}(\cdot)$.
The continuous-time Gaussian process $J_{c,d}(\cdot)$ plays the central role in our analysis. It can be viewed as a continuous-time version of the MA($\infty$) representation in (ref): past shocks are weighted by a damped sine wave. Again, the parameters $c$ and $d$ control the persistence and frequency of the cycle, respectively. Note also that long-cycle processes require stronger scaling than local-to-unity: $n^{-3/2}$ instead of $n^{-1/2}$. This is a reflection of the fact that long-cycle processes are near I(2).
We now turn to the properties of the least-squares estimators and the corresponding test statistics for the second-order autoregressive model with long cycles. Let $\widehat{\phi}_{1,n}$ and $\widehat{\phi}_{2,n}$ denote the least-squares estimator of (ref):
As it turns out, the matrix on the right-hand side is asymptotically singular because all three elements $\sum y^2_{t-1}$, $\sum y^2_{t-2}$, and $\sum y_{t-1}y_{t-2}$ converge to the same random limit when properly scaled. This is because $\sum y_{t-1}y_{t-2}=\sum y^2_{t-1}+$ smaller order terms, which follows formally from Lemma (ref)(b) below. The singularity complicates the derivation of the limiting distributions of the estimators and the corresponding test statistics.
To eliminate the singularity arising in the limit, we consider the following transformation of the equation in (ref):
where $\Delta y_{t-1} = y_{t-1} - y_{t-2}$. Since (ref) is obtained from the original equation through a non-singular linear transformation of the regressors and parameters, the OLS estimator of $\phi_{1,n} +\phi_{2,n}$ is given by $\widehat{\phi}_{1,n} + \widehat{\phi}_{2,n}$. Moreover, the usual Wald test statistic for testing joint hypotheses about $\phi_1$ and $\phi_2$ is the same for both regressions. Thus, we have the following:
As we show below, the matrix on the right-hand side of (ref) is no longer singular in the limit.
It follows from the representation in (ref) that the asymptotic theory of the OLS estimator involves the sample moments of $(y_{t-1},\Delta y_{t-1})$. Therefore, in addition to the asymptotic approximation of $y_{t-1}$, we also need the asymptotic approximation for $\Delta y_{t-1}$. The latter involves two additional continuous-time processes. We define:
Note that the diffusion process $K_{c,d}(r)$ is akin to the process $J_{c,d}(r)$ except that it is defined with a cosine function instead of a sine function. The next proposition shows that in large samples and after scaling, the distribution of $\Delta y_{\floor{nr}}$ can be approximated by that of $G_{c,d}(r)$.
Note that in contrast to the scaling $n^{-3/2}$ applied to $y_{t-1}$, its first difference $\Delta y_{t-1}$ requires scaling by $n^{-1/2}$. Therefore, the first differences of long-cycle processes have convergence rates of $O(n^{1/2})$ tantamount to those of local-to-unity processes. However, due to cyclicality, the large-sample distribution of $\Delta y_{t-1}$ differs from that in the local-to-unity model.
Based on the results of Proposition (ref) and (ref), we can now provide the asymptotic theory for the sample moments of long-cycle processes. Parts of the lemma below require the following ergodicity property for $\{u_t\}$.
Note that in part (e) of the lemma, the limiting distribution of the sample covariance between $\Delta y_{t-1}$ and $u_t$ depends on the difference between the long-run and the average over time variances of $\{u_t\}$. This reflects the serial correlation in $\{u_t\}$ and is standard in the unit root literature. However, despite the serial correlation, the difference $\sigma^2 -\sigma^2_u$ does not appear in the limiting expressions in part (d) for the sample covariance between $y_{t-1}$ and $u_t$. This is due to the stronger scaling factor required for the near I(2) long-cycle process $\{y_t\}$.
To simplify the notation, in the rest of the paper we use $\int J_{c,d}^2 $ to denote $\int_0^1 J_{c,d}^2(r) \dd r$ and $\int J_{c,d} \dd W$ to denote $\int_0^1 J_{c,d}(r) \dd W(r)$. We use the same convention for the integral expressions with $G_{c,d}(r)$ with $J_{c,d}$ replaced by $G_{c,d}$. Lastly, we use $\int J_{c,d}G_{c,d}$ to denote $\int_0^1 J_{c,d}(r) G_{c,d}(r) \dd r$.
Equipped with the results of Lemma (ref), we can now describe the asymptotic distribution of the least-squares estimators of $\phi_{1,n}$ and $\phi_{2,n}$.
According to part (b) of the proposition, the joint asymptotic distribution of the least-squares estimators of $\phi_{1,n}$ and $\phi_{2,n}$ is singular and determined by the same random variable. Furthermore, their convergence rate is $O_p(n^{-1})$ despite that the process $\{ y_t \}$ is near I(2). This is a consequence of the asymptotic singularity in (ref) as previously discussed on page (ref). However, in part (a) of the proposition, the least-squares estimator of $\phi_{1,n}+\phi_{2,n}$ has the faster convergence rate $O_p(n^{-2})$ characteristic of I(2) processes. Note that the limiting distributions depend on the localization parameters $c$ and $d$.
In Section (ref) below and following hansen1999grid, we consider a grid-based approach for constructing confidence sets for the autoregressive coefficients. Since hansen1999grid relies on the $t$-statistic for his procedure for the largest autoregressive root, we construct our procedure around the Wald statistic.\footnote{The approach in elliott2001confidence can be used to construct alternative statistics.}
Consider testing a joint hypothesis $H_0:\phi_1=\phi_{1,0}, \phi_2=\phi_{2,0}$ against $H_1:\phi_1\ne\phi_{1,0}\;\text{or}\; \phi_2\ne\phi_{2,0}$. Construction of grid-based confidence sets involves testing a sequence of hypothesis with different null values $\phi_{1,0},\phi_{2,0}$ and then collecting those that are not rejected. The usual Wald statistic is given by
where
and $\widehat{\sigma}^2_n$ is a consistent estimator of the long-run variance $\sigma^2$ constructed using $\hat u_t =y_t-\widehat \phi_{1,n} y_{t-1} - \widehat \phi_{2,n} y_{t-2}$, see NWK/WKD'87 and ADWK'91. The infeasible estimator of $\sigma^2$ that uses $u_t$ is constructed as $\tilde \sigma^2_n=n^{-1}\sum_{t=1}^n u_t^2 + 2\sum_{h=1}^{m_n} w_n(h) n^{-1}\sum_{t=h+1}^n u_t u_{t-h}$, where $m_n=o(n)$ is the lag truncation parameter, and $w_n(\cdot)$ is a bounded weight function such that $\lim_{n\to\infty}w_n(h)=1$ for all $h$. The feasible estimator $\widehat\sigma^2_n$ is constructed similarly using the estimated residuals $\hat u_t$ instead of $u_t$. We make the following assumption.
The conditions for consistency of the infeasible estimator can be found in NWK/WKD'87 and ADWK'91. Our next result describes the asymptotic null distribution of the Wald statistic for long-cycle processes.
The asymptotic null distribution of the Wald statistic is non-standard and non-pivotal: it depends on the ratio of the average over time and long-run variances $\sigma^2_u/\sigma^2$, and on the unknown localization parameters $c$ and $d$. While the ratio $\sigma^2_u/\sigma^2$ does not play a role when $\{u_t\}$ are serially uncorrelated and can be estimated consistently otherwise,\footnote{See the proof of Proposition (ref).} the dependence on $c$ and $d$ remains. Hence, the quantiles of the limiting distribution can only be simulated given the values of $c$ and $d$.
In Appendix (ref), we discuss the differences between the conventional $\chi^2_2$ critical values and the quantiles of the asymptotic distribution in Proposition (ref). Depending on the values of $c$ and $d$, the differences can be substantial, especially when the cycle length exceeds 25% of the sample size and the model includes deterministic components that are introduced in the next section.
\@startsection{section}{1} \z@{1.0\linespacing\@plus\linespacing}{.8\linespacing}{Extensions to models with deterministic components} For practical applications, it is important to allow the DGP to include nonzero means, trends, and deterministic cycles. We discuss such extensions in this section. As the results below show, the limiting distributions of the regression estimators and test statistics take a similar form to those in Section (ref), but with $J_{c,d}$ and $G_{c,d}$ replaced with their residuals from appropriate continuous-time projections. This property is standard in the unit-root literature and continues to hold in our case.
Formally, we assume that the data $\{y_t: t = 1,\ldots,n\}$ are generated according to
where $D_t$ is non-random, can vary with $t$, and depends on unknown parameters. To control for deterministic regressors, estimation of the autoregressive coefficients requires projecting against the components of $D_t$. The asymptotic distributions of the estimators and test statistics change accordingly. We consider the following three formulations of $D_t$:
The specification\footnote{We do not consider specifications in which deterministic components may dominate the stochastic long-cycle component.} in (i) allows $\{y_t\}$ to have a constant over time with a nonzero mean. The DGP in (ii) can be used, for example, to distinguish between very low frequency fluctuations and long cycles, as many time series in economics exhibit such patterns, see beaudry2020aer.\footnote{We thank Paul Beaudry for pointing our attention to this fact.} The $D_t$ component in (ii) generates cosine and sine oscillations at frequencies $2\pi k/n$. The period of such oscillations relative to the sample size is $1/k$, and they can capture very low-frequency cycles in data that are outside the range of interest of the econometrician. For practical purposes, we consider $k=1,2,3$. Inclusion of such components can be viewed as detrending of data by removing fluctuations at the frequencies corresponding to the values of $k$. The asymptotic results developed in this section can be used to account for detrending in inferential procedures.
The DGP in (iii) allows for linear time trends, and such adjustments have a long history in the unit root literature. The division by $n$ is required to derive the asymptotic properties and can be absorbed into the unknown coefficient $\xi$. Hence, observationally, the model in (iii) is identical to the model with no adjustment by $n$.
The empirical application in Section (ref) also considers the case where $D_t$ consists of seasonal dummies and a constant. However, as shown in phillips2002kpss for unit root testing, the arising asymptotic distributions have the same form as those in the constant mean case.
As in the previous section and to avoid singularities in the limit, we use the transformed version of the model with $y_{t-1}$ and $\Delta y_{t-1}$:
\@startsection{subsection}{2} \z@{.8\linespacing\@plus.7\linespacing}{.7\linespacing}{Constant mean} In this section, we consider case (i) of a constant unknown mean. When $D_t=\mu$, equation (ref) becomes
where $\alpha_n \equiv (1 - \phi_{1,n} - \phi_{2,n})\mu=O(n^{-2})$.\footnote{See Lemma (ref) in the Appendix.} Let $\widehat{\phi}_{1,n}$ and $\widehat{\phi}_{2,n}$ be the least-squares estimator of the corresponding coefficients in (ref), and define $\widetilde{y}_{t-1} = y_{t-1} - \bar{y}$ and $\widetilde{\Delta {y}}_{t-1} = \Delta y_{t-1} - \overline{\Delta {y}}$, where $\bar{y}_n$ and $\overline{\Delta {y}}_n$ denote the sample averages of $y_{t-1}$ and $\Delta y_{t-1}$ respectively. Then, \[
=
^{-1}
. \] and we have the following analogue of Lemma (ref).
The results in Lemma (ref) are parallel to those in Lemma (ref). However, instead of $J_{c,d}$ and $G_{c,d}$, the distributions arising in the limit depend on $\widetilde J_{c,d}$ and $\widetilde G_{c,d}$. Note that the latter processes are obtained from $J_{c,d}$ and $G_{c,d}$ by subtracting their respective continuous-time averages, which matches the construction of $\widetilde y_t$ and $\widetilde {\Delta y}_t$ in finite samples.
\@startsection{subsection}{2} \z@{.8\linespacing\@plus.7\linespacing}{.7\linespacing}{Deterministic cycles} In this section, we consider case (ii) of deterministic cycles. When $D_t$ includes deterministic cycles, equation (ref) takes the form
where the intercept $\alpha_n$ is as defined in the case of a constant mean. The lags of the cosine and sine components can be written as linear combinations of $\cos(2\pi k t/n)$ and $\sin(2 \pi k t/n)$ with coefficients depending on $n$ and, therefore, can be omitted. The least-squares estimators $\widehat{\phi}_{1,n}$ and $\widehat{\phi}_{2,n}$ can be obtained by estimating
where $\widetilde{y}_{t-1}$ and $\widetilde{\Delta {y}}_{t-1}$ are the residuals from the regressions of $y_{t-1}$ and $\Delta y_{t-1}$ respectively on $\cos(2\pi k t/n)$, $\sin(2 \pi k t/n)$, and a constant.
The following result describes the asymptotic distributions of the sample moments of $\widetilde y_{t-1}$, $\widetilde{\Delta y}_{t-1}$, and $u_t$.
Lemma (ref) is the analogue of Lemma (ref) for the model with deterministic cycles. The coefficients $\psi_{1,k}$ and $\psi_{2,k}$ can be viewed as the least-squares coefficients in the continuous time regression of $J_{c,d}(s)$ against $\cos(2\pi k s)$, $\sin(2 \pi k s)$, and a constant with $s$ varying over the interval $[0,1]$. The coefficients $\varphi_{1k}$ and $\varphi_{2k}$ have a similar interpretation with $J_{c,d}$ replaced by $G_{c,d}$. Therefore, the processes $\widetilde J_{c,d}$ and $\widetilde G_{c,d}$ are the residuals from the corresponding continuous-time regressions. They are continuous-time versions of $\widetilde y_{t-1}$ and $\widetilde{\Delta y}_{t-1}$ respectively. Therefore, the results of Lemma (ref) continue to hold with processes $J_{c,d}$ and $G_{c,d}$ replaced by their respective residuals from continuous-time regressions.
\@startsection{subsection}{2} \z@{.8\linespacing\@plus.7\linespacing}{.7\linespacing}{Linear time trend} In this section, we consider case (iii) of a linear time trend. The model in equation (ref) now takes the form
where $\delta_n \equiv \alpha_n + (\phi_{1,n} + 2\phi_{2,n})\xi/n=O(n^{-2})$, and $\beta_n\equiv \xi(1-\phi_{1,n}- \phi_{2,n})=O(n^{-2})$.\footnote{See Lemma (ref) in the Appendix.} Similarly to the previous cases, the least-squares estimators $\widehat{\phi}_{1,n}$ and $\widehat{\phi}_{2,n}$ can be obtained by estimating
where $\widetilde{y}_{t-1}$ and $\widetilde{\Delta y}_{t-1}$ are now the residuals from the regressions of $y_{t-1}$ and $\Delta y_{t-1}$ respectively against $t/n$ and a constant.
Lemma (ref) is the analogue of Lemmas (ref) and (ref) for the case of the linear time trend. The processes $\widetilde J_{c,d}(r)$ and $\widetilde G_{c,d}(r)$ can be interpreted similarly as the residuals from the continuous-time regressions of $ J_{c,d}(r)$ and $ G_{c,d}(r)$, respectively, against a constant and $r$ varying over the interval $[0,1]$.
\@startsection{subsection}{2} \z@{.8\linespacing\@plus.7\linespacing}{.7\linespacing}{Asymptotic distributions of the estimators and test statistics} The results of Lemmas (ref)--(ref) can now be used to describe the asymptotic distributions of the least-squares estimators of the autoregressive coefficients and the corresponding Wald statistics for the models with constant mean, deterministic cycles, and a linear time trend, respectively.
Under the same assumptions as those in Proposition (ref), however with the model in equation (ref) replaced by that in either (ref), (ref), or (ref), the asymptotic distribution of the least-squares estimators of the autoregressive coefficients now satisfies
where the convergence holds jointly with the results of either Lemma (ref), (ref), or (ref) respectively with the correspondingly defined residual processes $\widetilde J^2_{c,d}$ and $\widetilde G^2_{c,d}$.
For all three specifications in Sections (ref)--(ref), the Wald statistic for testing $H_0:\phi_1=\phi_{1,0}, \phi_2=\phi_{2,0}$ against $H_1:\phi_1\ne\phi_{1,0}\;\text{or}\; \phi_2\ne\phi_{2,0}$ takes the same form as in equation (ref). However, $\widehat V_n$ is now given by \[ \widehat{V}_n = \widehat{\sigma}_n^2
^{-1}, \] with $\widetilde{y}_{t-1}$ and $\widetilde{\Delta y}_{t-1}$ defined respectively for each specification. Provided that the assumptions of Proposition (ref) hold with the model in (ref) replaced by that in either (ref), (ref), or (ref), the asymptotic null distribution of the Wald statistic is given by
with the correspondingly defined residual processes $\widetilde J^2_{c,d}$ and $\widetilde G^2_{c,d}$.
As in the base case with no deterministic components, the asymptotic null distributions of the Wald statistics are nonstandard and depend on the unknown parameters $c$ and $d$. Differences between the quantiles of these asymptotic distributions and the critical values of the $\chi^2_2$ distribution are discussed in Appendix (ref). Compared to the base case, the inclusion of deterministic components may result in more substantial deviations from the $\chi^2_2$ critical values.
\@startsection{section}{1} \z@{1.0\linespacing\@plus\linespacing}{.8\linespacing}{Inference for cyclicality}
In this section, we propose a procedure for inference on the cycle length in terms of the angular frequency-based measure $\tau_{\theta}$ and the spectrum-based measure $\tau_{\omega}$ that were introduced in Section (ref). Recall that the two measures can be deduced from the autoregressive coefficients $\phi_{1n}$ and $\phi_{2,n}$ through the relationships in (ref)--(ref). Therefore, we first construct confidence sets for the autoregressive parameters by collecting values $(\phi_1,\phi_2)$ consistent with cyclical behavior and not rejected by data. In the second step, we use projection arguments to construct confidence intervals for $\tau_{\theta}$ and $\tau_{\omega}$. By multiplying the values of $\tau_{\theta}$ and $\tau_{\omega}$ in the confidence intervals by $n$, the length of the cycle can also be expressed in time units instead of fractions of sample size.
The proposed confidence sets have the following property: If the true DGP is indeed cyclical, the coverage probability is at least $1-\alpha$ asymptotically whether the roots of the autoregressive equation are close to or far from one. However, if the true DGP is inconsistent with the cyclical behavior, we expect the confidence sets to be empty in large samples. Therefore, the proposed procedure can be used to detect cyclical specifications consistent with the data or to rule out cyclical behavior. However, our procedure is not designed for inference on acyclical specifications.
When the roots of the autoregressive polynomial are local to unity as in Assumption (ref), the least-squares estimators of the autoregressive coefficients are consistent regardless of whether $\{u_t\}$ is serially correlated or not. This is established in Proposition (ref) for the base case and in (ref) for the cases with deterministic components. The serial correlation in $\{u_t\}$ and the resulting correlation between $(y_{t-1}, \Delta y_{t-1})$ and $u_t$ is reflected by the noncentrality term $0.5(1-\sigma^2_u/\sigma^2)$ in the asymptotic distributions. The non-centrality proliferates from the estimators into the asymptotic null distribution of the Wald statistic. This is standard for the unit root literature and continues to hold in our framework.
However, when the roots of the autoregressive polynomial are sufficiently far from unity, that is, under I (0) specifications, the least squares estimators of the autoregressive coefficients are no longer consistent if $\{u_t\}$ is serially correlated. Although the process is I(0), due to the inconsistency of the least-squares estimators, the null asymptotic distribution of the Wald statistic is no longer a central $\chi^2$. Therefore, to design an inferential procedure that remains valid regardless of the magnitude of the roots, we need to be able to accommodate a potential serial correlation in $\{ u_t\}$.
We proceed as follows. First, in Section (ref) we discuss how to construct confidence intervals for $\tau_\theta$ and $\tau_\omega$ when $\{ u_t\}$ are serially uncorrelated. Then, in Section (ref) we extend the procedure to a serially correlated innovation process $\{ u_t\}$ by assuming that it satisfies an AR($p$) formulation with real roots bounded away from one. We use the BIC selection procedure to choose the appropriate number of lags $p$ and the specification for the deterministic part $D_t$.
\@startsection{subsection}{2} \z@{.8\linespacing\@plus.7\linespacing}{.7\linespacing}{Serially uncorrelated $\{u_t\}$} Suppose that $\{u_t\}$ is serially uncorrelated and, therefore, the least-squares estimators of the autoregressive coefficients $\phi_{1,n}$ and $\phi_{2,n}$ are consistent whether the roots are close to unity or far from it. Recall that the expression on the right-hand side of (ref) with $1-\sigma^2_u/\sigma^2=0$ approximates well the asymptotic distribution of the Wald statistic for any configuration of the localization parameters $c$ and $d$. Moreover, recall that given the sample size $n$, there is a one-to-one relationship between $(\phi_{1,n}, \phi_{2,n})$ and the localization parameters $(c,d)$, and let
where the functions $\Phi_{1,n}(c,d)$ and $\Phi_{2,n}(c,d)$ are defined according to (ref) and (ref) respectively. Because the relationship is one-to-one for any given $n$, confidence sets for $(\phi_1,\phi_2)$ can be equivalently represented as confidence sets in terms of $(c,d)$.
After running the regressions of $y_t$ against $y_{t-1}$ and $y_{t-2}$ with different specifications of the deterministic part $D_t$, the BIC can be used to consistently choose the appropriate specification between the constant mean, the deterministic cycle, or the linear time trend. After selecting $D_t$, consider the corresponding Wald statistic $W_n(\Phi_{1,n}(c,d),\Phi_{2,n}(c,d))$. Let $\mathcal{W}_{1-\alpha}(c,d)$ denote the $1-\alpha$ quantile of the asymptotic distribution in (ref) with $1-\sigma^2_u/\sigma^2=0$. I.e., $\mathcal{W}_{1-\alpha}(c,d)$ is the $1-\alpha$ quantile of the distribution of
where the definitions of $\widetilde J_{c,d}$ and $\widetilde G_{c,d}$ correspond to the selected specification for $D_t$. The confidence set for $(c,d)$ can now be constructed by test inversion as
The confidence set $CS_{n,1-\alpha}$ is bounded as $\mathcal W_{1-\alpha}(c,d)\to\chi^2_{2,1-\alpha}$ when $c\to-\infty$ or $d\to\infty$. In practice, the confidence set can be approximated by choosing a dense two-dimensional grid of values $c$ and $d$. We use a grid with $\underline{c}\leq c\leq 0$, and $2\pi< d < n\pi$, where the lower bound $\underline{c}$ is chosen by the econometrician and the lower bound of $2\pi$ is imposed to rule out cycles longer than the sample size when measured by $\tau_\theta$.\footnote{In our empirical application in Section (ref), the largest considered value of $d$ is 678, and $\underline{c}=-678$. Note also that the grid only needs to cover the stationary region corresponding to complex roots sargent1987macroeconomic.}
The construction of $CS_{n,1-\alpha}$ is similar to the grid bootstrap procedure of hansen1999grid, however, we use the asymptotic critical values instead of their bootstrap approximation. Note that the critical values must be adjusted for every point $(c,d)$ considered. The validity of $CS_{n,1-\alpha}$ is due to the following facts. First, $(c,d)$ is included in the confidence set only if the null hypothesis $H_0:\phi_{1,n}=\Phi_{1,n}(c,d), \phi_{2,n}=\Phi_{2,n}(c,d)$ cannot be rejected by the Wald test with the critical value $\mathcal{W}_{1-\alpha}(c,d)$. Second, the critical values are computed using the same values $(c,d)$ as those specified in $H_0$. Third, the distribution in (ref) nests the $\chi^2_2$ distribution, which arises under the fixed $(\phi_1,\phi_2)$ asymptotics, as a limiting case. Note that having the correct size under both drifting and fixed parameter specifications is required for uniform validity andrews2020generic.
We construct confidence intervals for $\tau_\theta$ and $\tau_\omega$ from $CS_{n,1-\alpha}$ by projection:
The confidence interval for $\tau_\theta$ is bounded as long as the grid of $d$ values used to construct $CS_{n,1-\alpha}$ excludes zero. On the other hand, the confidence interval for $\tau_\omega$ can be unbounded if pairs $(c,d)$ with $c=d$ are included in $CS_{n,1-\alpha}$.
\@startsection{subsection}{2} \z@{.8\linespacing\@plus.7\linespacing}{.7\linespacing}{Serially correlated $\{u_t\}$}
In this section we assume that the innovations process $\{ u_t\}$ is generated as AR($p$):
where $\{\varepsilon_t\}$ are iid $(0,\sigma^2_\varepsilon)$, and the roots of the polynomial $1-\rho_1 L-\ldots -\rho_p L^p$ are real and bounded away from unity. In this case, $\{y_t\}$ is AR($p+2$), and by running the regressions of $y_t$ against different specifications of the deterministic part $D_t$ and $y_{t-1},y_{t-1},\ldots,y_{t-(m+2)}$ for some $m>p$, one can again use the BIC to consistently estimate the specification for $D_t$ and the number of lags $p$ in (ref).
Let $\widetilde y_t$ denote the residuals from the projection of $y_t$ against the components of $D_t$. Under $H_0:\phi_{1,n}=\phi_{1,0}, \phi_{2,n}=\phi_{2,0}$, the values $\phi_{1,n}$ and $\phi_{2,n}$ are known and can be calculated from the values of $c$ and $d$. Let \[ \widetilde u_{t,0}\equiv \big(1-\phi_{1,0}L - \phi_{2,0} L^2\big) \widetilde y_t. \] Using the null-restricted residuals $\widetilde u_{t,0}$, one can estimate the autoregressive coefficients $\rho_1,\ldots,\rho_p$. Let $\widehat\rho_{1,0},\ldots,\widehat\rho_{p,0}$ denote their least-squares estimators. Note that under $H_0$, these estimators are consistent. We can now remove the autoregressive part in $u_t$: \[ \widehat x_{t,0} \equiv(1-\widehat\rho_{1,0} L-\ldots -\widehat\rho_{p,0} L^p) \widetilde y_t. \] Thus, to construct the process $\widehat x_{t,0}$, we have filtered the deterministic part $D_t$ and the serial correlation in $\{u_t\}$. Note that under the null, the population counterpart of $\widehat x_{t,0}$ satisfies:
where $\widetilde\varepsilon_t$ is the residual from the projection of $\varepsilon_t$ against the components of $D_t$.
Now we can use $\{\widehat x_{t,0}\}$ for inference on the cyclical properties of $\{y_t\}$, however, additional adjustments are required to account for the estimation of $\rho_1,\ldots,\rho_p$. The main purpose of the adjustments discussed below is to ensure that the modified Wald statistic has the correct asymptotic null distributions both under the long-cycle asymptotics proposed in the paper and under the standard asymptotics with $\phi_1$ and $\phi_2$ fixed in the stationary range.\footnote{Recall that having correct size under both drifting and fixed parameters specifications is required for the uniform validity andrews2020generic. }
Let $\widehat{\phi}_{1,n}$ and $\widehat{\phi}_{1,n}$ now denote the least-squares estimators of $\phi_1$ and $\phi_2$ respectively from the regression of $\widehat x_{t,0}$ against $\widehat x_{t-1,0}$ and $\widehat x_{t-2,0}$: \[ \widehat x_{t,0}=\widehat{\phi}_{1,n}\widehat x_{t-1,0}+\widehat{\phi}_{2,n}\widehat x_{t-2,0}+\widehat\varepsilon_{t,0}, \] where $\widehat\varepsilon_{t,0}$ denotes the least-squares residuals. The modified Wald statistic takes the form \[ W_{n,p}(\phi_{1,0},\phi_{2,0}) \equiv \frac{1}{\widehat \sigma^2_{\varepsilon,n}}
^\top M_n{\Sigma}^{-1}_n M_n
, \] where $\widehat \sigma^2_{\varepsilon,n} \equiv n^{-1}\sum \widehat \varepsilon_{t,0}^2$, and the matrix $M_n$ is given by \[M_n \equiv
. \] To construct $\Sigma_n$, we first define $\dot x_{t,0}$ and $\ddot x_{t,0}$ as the residual from the least-squares regression of $\widehat x_{t,0}$ and $\widehat x_{t-1,0}$ respectively against $\widetilde u_{t,0},\ldots,\widetilde u_{t-p+1,0}$:
where $\dot\zeta_{1,n},\ldots,\dot\zeta_{p,n}$ and $\ddot\zeta_{1,n},\ldots,\ddot\zeta_{p,n}$ are the OLS estimators. The matrix $\Sigma_n$ is given by \[\Sigma_n \equiv
. \] The next proposition shows that under the conventional stationary asymptotics, the asymptotic null distribution of the Wald statistic is the usual $\chi^2_2$ distribution.
In the case of a long-cycle specification, the null asymptotic distribution of the modified Wald statistic is the same as in (ref).
Using the results of Propositions (ref) and (ref), one can now construct confidence sets for $(c,d)$ using the modified Wald statistic as
Similarly to the construction in (ref) and (ref), the confidence set $CS_{n,p,1-\alpha}$ can be projected to construct confidence intervals for $\tau_\theta$ and $\tau_\omega$.
\@startsection{section}{1} \z@{1.0\linespacing\@plus\linespacing}{.8\linespacing}{Cyclical properties of macroeconomic and financial variables}
Recurrent boom-and-bust cycles are a salient feature of economic and financial history. A long-standing interest in understanding these ups and downs in the macro-financial aggregates has led to a vast body of literature on business cycles and a resurgence of research on financial cycles post the financial crisis-induced Great Recession of 2008. Among these strands of work is the empirical characterization of business and financial cycles. The traditional approach to such a characterization is to identify turning points or peaks and troughs in the time series using the dating algorithms of bry1971front and harding2002dissecting. Based on the turning-point analysis, DrehmannBorioTsatsaronis2012 highlight the importance of medium-term cycles that last 18 years for credit, 11 years for GDP and 9 years for equity prices. These findings are in line with studies using frequency-based bandpass filters Aikman_etal2015EJ,comin2006medium.
The cyclical properties of the data have also been formally examined in the literature using a variety of methods, including direct and indirect spectrum estimation AHearnWoitek2001JME,Strohsal_etal2019JBF and structural time-series modeling Harvey1985,Runstler2018business. However, they rely on the conventional asymptotic approximations that may produce misleading results with long-cycle data, as we argue in this paper. For example, we show in Appendix (ref) that the periodogram-based estimator is asymptotically biased in the case of long cycles.
In this section, we apply our inference procedure to the quarterly series of a set of macroeconomic and financial variables for the U.S. All data are publicly available from FRED, Federal Reserve Bank of St. Louis. A detailed description of the data is summarized in Table (ref) in Appendix (ref). All the series are measured in natural logs except for the credit-to-GDP ratio (for the private non-financial sector), which is in percentage points, and the interest rate spread between Moody's seasoned BAA corporate bond yield and the 10-year treasury constant maturity, which is expressed in levels. For each series, we take the longest and most updated sample ending in 2020. Depending on the series, our samples span periods ranging from 34 to 73 years.
We use the empirical models in (ref). Let $y_t$ denote the observed data series such that
where $y^c_t$ is the latent cyclical part, the innovations $\{u_t\}$ are potentially serially correlated according to an AR($p$) specification with unknown $p$, and $D_t$ may contain linear deterministic trends and deterministic cycles as discussed in Section (ref). In all specifications, the intercept (constant) is included by default. The raw data for the credit-to-GDP ratio are not seasonally adjusted and, therefore, the specification for $D_t$ also allows for seasonal dummies.
We use the BIC to select the appropriate specification for $D_t$ (i.e. whether to include a linear time trend, deterministic cycles, or seasonal dummies). We also rely on the BIC to select the lag order $p$ in the AR($p$) specification for $u_t$. To this end, suppose that $p\leq M$ for some known positive integer $M$. As discussed in Section (ref) of the Supplement, one can choose $p\in\{0,1,\ldots,M\}$ that minimizes the BIC for the regression of $y_t$ against $y_{t-1}$, $\Delta y_{t-1}$, the deterministic components, and the second-order differences $\Delta^2 y_{t-1},\ldots,\Delta^2 y_{t-p}$, where $\Delta^2 y_t \equiv \Delta y_t - \Delta y_{t-1}$.
In our empirical application, most of the time series do not exhibit autocorrelation in $\{u_t\}$, that is $p=0$, except for the credit-to-GDP ratio, as indicated by the BIC. Moreover, the credit-to-GDP ratio is the only series in which we have included seasonal dummies.
Table (ref) presents our results. Note that the last three columns of the table describe the specifications selected by the BIC for $D_t$ and the order of autocorrelation for $\{u_t\}$. For example, the hours per capita contains a linear time trend and deterministic cycles of cosine and sine waves with $k=1$, which corresponds to the periodicity of $n/k = n$. According to the BIC, the errors $\{u_t\}$ are serially uncorrelated.
Columns 1 and 2 of the table report, respectively, the angular frequency-based measure $n\tau_{\theta}$ and the spectrum-maximizing frequency-based measure $n\tau_{\omega}$ for the cycle length. The point estimates for $n\tau_{\theta}$ and $n\tau_{\omega}$ are constructed by backing out the corresponding values of $c,d$ from the OLS estimates of $\phi_{1,n}$ and $\phi_{2,n}$, which can then be used to compute $\tau_{\theta}$ and $\tau_{\omega}$. The point estimates are indicated as “---" when the autoregressive coefficient estimates of $\phi_{1,n}$ and $\phi_{2,n}$ correspond to acyclical processes,\footnote{While the OLS estimates of the autoregressive coefficients may correspond to an acyclical process (no complex roots or no spectrum peak) our confidence sets for $\phi_{1,n}$ and $\phi_{2,n}$ may nevertheless include cyclical values.} or when they are not available as in the case of autocorrelation. The minimum and maximum cycle lengths implied by the 95% confidence intervals of $n\tau_{\theta}$ and $n\tau_{\omega}$ are given in parentheses. For completeness, the table also reports the 95% confidence intervals for $c$ and $d$.
The two alternative measures of cycle length generally produce similar lower bound estimates. Based on the 95% confidence intervals, we are unable to reject the null that macroeconomic variables, such as the real GDP per capita, the unemployment rate, and the hours per capita, contain stochastic cycles with periodicity of at least 5-6 years. Partly due to the projection-based construction of the $n\tau_{\theta}$ and $n\tau_{\omega}$ confidence intervals, the implied range of the cycle length is typically wide. The upper bound confidence estimates usually are large and differ considerably between the two measures. Nevertheless, our results point to the presence of cyclicality among macroeconomic variables, conforming to the view of endogenous business cycles beaudry2020aer. On the financial side, we find that credit to the private non-financial sector as a percent of GDP and the home prices exhibit cycles of at least 10 years in duration, twice as long as the minimum detected cycle length in the macroeconomic variables.
The most striking finding of this section is that for asset market variables (the volatility index, credit risk premium, and equity prices), our procedure returns empty confidence sets. This suggests that the underlying mechanism for asset market fluctuations is different from that of macro variables and financial variables such as credit and home prices. Our results are in favour of the dichotomy between the asset market and the real economy. Moreover, the results do not support the view that recessions are driven by risk perception, risk premiums, and risk-bearing capacity suggested in the macro-finance literature cochrane2017macro. Note that the S&P 100 Volatility Index, a measure of market uncertainty, has a deterministic cycle of approximately 46 quarters in length according to the BIC. However, it is different in nature from the stochastic cycles detected in the other variables.
To better visualize the cyclical dynamics consistent with the data, for each series ${y_t}$, we plot in Figure (ref) the impulse responses to a one-standard-deviation shock to the innovation $u_{0}$ of all cyclical specifications in the 95% confidence sets $CS_{n,1-\alpha}$.\footnote{For the credit to the private non-financial sector, the standard deviation of the innovation is computed assuming no serial correlation.} The dynamics shown in the figure resonate with the results in Table (ref). Financial variables such as the credit to the private non-financial sector and the home prices exhibit much longer cycles than the business cycle variables. The duration from peaks to troughs is at least 25-30 quarters in financial cycles and at least 15 quarters in business cycles. Furthermore, financial cycles are also more pronounced: for a one-standard deviation shock, the initial amplitude of the cyclical response is approximately 3 to 7 times the standard deviation for the credit and the home prices, and about 1.5 to 3 times for the unemployment rate and the hours per capita.
For real GDP per capita, the impulse responses are split into two parts. On the left, the axis corresponds to the set of impulse responses similar to those observed in the unemployment rate and the hours per capita. On the right, the axis maps to the set of cyclical impulse responses with large amplitudes and high persistences. Note that the scale of the axis on the right has increased by 10-fold. Although the possibility of having much longer and highly persistent stochastic cycles cannot be rejected, the real GDP per capita shares a similar dynamic to the unemployment rate and the hours per capita.\footnote{Note also that the hours per capita is much more persistent than the unemployment rate and the real GDP per capita. }
In sum, our results suggest that business cycles (as marked by the expansions and contractions of aggregate economic activity) are not just recurrent but periodic, with an average duration of at least 5-6 years. Furthermore, financial cycles as characterized by the booms and busts in credit and home prices are much longer than business cycles: at least 10 years in duration. Additionally, these financial cycles have more prominent oscillations with amplitudes much larger than those of business cycles. Moreover, we find that equity prices, though commonly included in the characterization of financial cycles, do not exhibit stochastic cycles, and therefore merit separate consideration from credit and home prices. Lastly, our results suggest that asset market fluctuations are a different phenomenon from changes in real economic activities.