Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
83,927 characters · 15 sections · 65 citation commands
The boosted HP filter is more general than you might think
\thispagestyle{empty}
{ The global financial crisis and Covid recession have renewed discussion concerning trend-cycle discovery in macroeconomic data, and boosting has recently upgraded the popular HP filter to a modern machine learning device suited to data-rich and rapid computational environments. This paper extends boosting's trend determination capability to higher order integrated processes and time series with roots that are local to unity. The theory is established by understanding the asymptotic effect of boosting on a simple exponential function. Given a universe of time series in FRED databases that exhibit various dynamic patterns, boosting timely captures downturns at crises and recoveries that follow. }
Key words: Boosting, Business cycle, Machine learning, Macroeconomics, Recession
JEL codes: C22 Time Series Models, C55 Large Data Sets, C43 Index Numbers and Aggregation
\onehalfspacing
\doublespacing
Understanding long-term trends and short-term business cycles in economic activity was a primary pursuit of the foundational researchers of the econometrics profession especially during the years of the Great Depression where these matters figured prominently and prompted the development of new econometric approaches ragnar1933propagation,tinbergen1939business. At that time with very few exceptions data were scarce. By contrast vast datasets are now available to researchers covering multiple decades of quarterly and monthly time series observations across a wide range of economic variables. Long trajectories of data provide rich information about many aspects of economic activity and wellbeing, including the impact of technical progress on growth and the course and consequences of intermittent slowdowns and recessions. Analysis of the information carried in such time series provides useful indicators of the passage and present state of economic activity, which in turn helps to shape assessments of policymakers, regulators, corporate executives, and consumers in guiding decision making. The 2008 global financial crisis (GFC) and its aftermath and most recently the global economic impact of the Covid-19 pandemic are timely reminders that the work launched by ragnar1933propagation and tinbergen1939business remains an ongoing mission for the econometrics community.
Modern econometric approaches often focus on decomposing time series observations of a variable $x_{t}$ into additive components that represent long run trending behavior, $f_t$, and cyclical activity, $c_t$, as
The trend embodies the long run general course or tendency in the data and the cycle reflects periodic fluctuations in economic activity in which businesses, labor markets and consumer behavior alternately expand and contract. Trends in many economic aggregates like a nation's real GDP are primarily determined by the impact of production technologies, the size and quality of labor forces, the accumulation of physical and human capital, and entrepreneurship, whereas business cycles are affected by shorter term internal and external influences including contractionary or expansionary fiscal, monetary, and political policies, combined with the dynamic propagation of these forces within an economy. As such, trend and cycle may be considered latent elements in the data which are to be identified and estimated by econometric methods through decomposition or direct modeling.
Since its introduction the Hodrick-Prescott (HP) filter hodrick1997postwar has become a convenient and highly popular off-the-shelf choice for trend-cycle decomposition. Given an observed time series $x=(x_{1},x_{2},\dots,x_{n})^{\prime}$, the HP filter finds $\widehat{f}^{{\rm HP}}=(\widehat{f}_{1}^{{\rm HP}},\widehat{f}_{2}^{{\rm HP}},\ldots,\widehat{f}_{n}^{{\rm HP}})^{\prime}$ via a penalized least squares criterion
where the second difference $\Delta^{2}f_{t}=\Delta f_{t}-\Delta f_{t-1}=f_{t}-2f_{t-1}+f_{t-2}$ of the trend component provides a measure of fluctuations, whose degree is controlled through the tuning parameter $\lambda \in (0,\infty)$ that governs the extent of the penalty in the second component of the extremum criterion (ref). Penalization plays a key role in determining the outcome of the filter, with larger $\lambda$ imposing a greater penalty on roughness thereby favoring smoothness in the outcome $\widehat{f}^{{\rm HP}}$.
Formally developed a century ago by whittaker1922new in pathbreaking work on penalized estimation, ideas for `graduating' data have a long history in actuarial science and subsequent work in statistics and engineering, as reviewed in phillips2021business. When $\lambda \to \infty$ the solution for (ref) satisfies $\Delta^{2}f_{t}=0$, giving a linear trend function $f_t=a+bt$ for some constant coefficients $a$ and $b$. The conventional choice for quarterly data is $\lambda=1600$, suggested by hodrick1997postwar based on empirical experimentation with macroeconomic time series. Corresponding settings of $\lambda$ for empirical work with monthly and annual data were given in ravn2002adjusting.
A poignant newspaper column posted by krugman2012 concerning the importance of long run trend identification during the GFC raised substantial interest on trend determination and the HP filter, spurring an influx of opinions, theory investigations, novel proposals, and empirical evidence. de2016econometrics, cornea2017explicit, and sakarya2017property explored the HP filter's finite-sample algebraic properties. Arguing that the HP filter is a nonparametric procedure in which tuning parameters are typically sample size dependent to achieve consistent estimation, \citetalias{phillips2021business} analyzed its asymptotic features via operator calculus, showing that the common setting $\lambda=1600$ is too large to completely remove stochastic trends in time series of the length usually encountered in empirical work. Instead, the HP filter smooths those trends into paths that are asymptotically differentiable forms of Brownian motion. Several revised or modified versions based on the HP filter have recently emerged. For example, yamada2020smoothing and yamada2021trend generalize the HP filter to overcome data imperfections, and with the advent of Covid-19 lee2021sparse use a sparse filter to identify the turning points in the inflection rates of the virus.
Retaining the squared penalization scheme of the original HP filter, phillips2021boosting proposed a boosted HP filter (bHP hereafter) designed to upgrade the procedure to a machine learning device that uses the data more intensively to improve its properties and performance. The bHP filter is a repeated application of the HP filter to the residual extracted in the last iteration (see Algorithm (ref) below) in which the number of iterations, $m$, controls the intensity of reusage. In practice, \citetalias{phillips2021boosting} suggested monitoring a stopping criteria to terminate the iteration in a data-driven manner (see Algorithm (ref)), which makes bHP automated in application, as envisaged in phillips2005automated. In three empirical examples, \citetalias{phillips2021boosting} fed 127 individual time series of various lengths and trending patterns through the bHP machine. The trend and cycle estimates returned after a few iterations in this algorithm largely confirmed and refined the existing findings in the literature for these series. Further simulation experiments and empirical applications (e.g., hall2021does,hall2024selecting) have been conducted and the robustness continues to hold.
These outcomes are indicative that, as a machine learning device which is agnostic about the data generating mechanism, the bHP filter satisfactorily accommodates trend processes in applications that are more general than unit root processes. But there remain gaps in the present asymptotic theory that involve persistent time series beyond unit root stochastic trends, such as near unit root processes, for which it would be useful in empirical work to have a theory foundation for the use of boosting in consistent trend determination. This paper offers an affirmative answer with analytic, simulation, and empirical support for this extension.
In doing so we provide a preparatory result (Lemma (ref) in Section (ref)) that characterizes the shrinkage effect of the operational form of the HP filter. We keep using the tuning parameter formula $\lambda=\mu n^{4}$ for some constant $\mu>0$, which is an expansion rate extensively studied in \citetalias{phillips2021business}. For sample sizes that are typical in quarterly economic data, this expansion rate approximates well the actual form of the filter in practical work with the common choice $\lambda=1600$ for the HP filter. In view of the two-sided nature of the filter, the HP operator is a function of lead and lag operators. For a class of complex numbers $a\in\mathbb{C}$ such that $a^{4}$ is a non-negative real number, Lemma (ref) shows that the HP residual operator shrinks the complex exponential $\mathrm{e}^{ax}$ towards zero by the factor $\mu a^{4}/ (\mu a^{4}+1 )\in[0,1)$, which is a pseudo-differential operator extension of the elementary property $D_x^{m}\mathrm{e}^{ax}=a^{m}\mathrm{e}^{ax}$ of the usual differential operator $D_x=\mathrm{d}/\mathrm{d}x$. Repeating the operation $m$ times gives the power factor $(\mu a^{4}/ (\mu a^{4}+1 ) )^{m}\mathrm{e}^{ax}$, which tends to zero as $m \to \infty$.
This lemma enables a unified development of asymptotics of the HP and bHP filters for a variety of nonstationary trend processes that include unit root $I(1)$ time series, higher order integrated $I(q)$ processes with integer $q\geq2$, and local-to-unity (LUR) processes, which are among the most widely used models for nonstationary data. In each case, boosting enables consistent estimation of the trend whereas single implementation of the HP filter is inconsistent, producing a smoothed version of the original trend process. Upon standardization these trends all have asymptotic stochastic process representations in terms of convergent series of trigonometric functions with complex exponential forms that are amenable to analysis by operator methods and thereby deliver the respective asymptotic forms of the filter operation. Similar methods apply to general deterministic trend functions with convergent Fourier series representations. Taken together, the asymptotic results span a wide range of trend models commonly used in econometric practice.
The present paper continues the line of research by \citetalias{phillips2021business} and \citetalias{phillips2021boosting} in understanding of the capabilities of the (boosted) HP filter. While these two papers worked with unit root stochastic trends exclusively, this paper considerably expands the types of stochastic and deterministic trend mechanisms. The unified analytical framework sheds new theoretical insight about the shrinkage effect of the HP operator, which exposes the source of limitations of the HP filter and strengths of the repeated fitting. It complements ongoing work that considers the use of boosting methods for time series with long range dependence biswas2022longrange. In addition, the present work provides further numerical evidence of the robustness and versatility of bHP in simulations and various real data applications. Overall, the findings reveal analytic properties and empirical performance that support the boosted filter as a useful machine learning method for extracting low frequency components of macroeconomic time series, thereby contributing to the comprehension of trend phenomena.
In machine learning terminology the HP and bHP filters are unsupervised learning methods which seek to extract key features of the data but do not use regressors to fit the dependent variables or `labels'. In contrast to this methodology, hamilton2018you firmly advocated that the HP filter should be replaced in empirical work by an autoregression $x_{t+h} = \beta_0 + \beta_1 x_{t} + \beta_2 x_{t-1} + \dots + \beta_p x_{t-p+1} + v_{t+h}$ with specific order $(h,p)=(8,4)$ recommended for quarterly data. The trend component at the period $t+h$ is estimated by the fitted dependent variable $\widehat x_{t+h}$. We call it Hamilton's regression filter (HRF). Regressions of this type fall into the category of simple supervised learning methods in which a few lagged observations are trained to predict a future target. Unlike nonparametric approaches such as HP and bHP where tuning parameters are unavoidable and play a central role in consistent estimation, parametric autoregressions typically bypass tuning parameters, although users still need to decide on the number of lags quast2022reliable. hamilton2018you's paper has stirred considerable discussion and debate. The issues raised bear directly on empirical econometric practice and they affect economic policy analysis in fundamental ways concerning the manner in which observed economic indicators can be interpreted as indicative of long term trend behavior as distinct from cyclical fluctuations, the very considerations that motivated krugman2012's public post. schuler2021cyclical pointed out that HRF fails to reproduce the standard chronology of US business cycles and `emphasizes cycles that exceed the duration of regular business cycles (i.e., longer than 8 years), and completely mutes certain short-term fluctuations.'
Among the commentary, knightboosted provided a frequency domain analysis of both the HP and bHP filters, showing that the latter has some additional `free pass' effects at low frequencies over that of the HP filter, giving it improved recovery properties for trends with frequencies in an interval around the origin. Recent empirical works by drehmann2018you, hall2021does and jonsson2020cyclical compared the HP filter and HRF in real data examples and the accumulated empirical evidence from these studies favors the former. In subsequent work hall2024selecting provided transfer function analysis and studied the empirical performance of the HP and bHP filters with various stopping rules, recommending a `twicing' version of bHP with a second iteration (2HP) for trend and growth cycle analysis with New Zealand quarterly macroeconomic data. lu2023boost extended the discussion of the properties of the HP filter and its boosted version in line with their finite sample weights and linked the necessity of boosting to the patterns observed in the macro time series.
\citetalias{phillips2021boosting} provided a detailed response to hamilton2018you's critiques of the HP filter and advocated the use of an automated bHP filter. The present paper adds further simulation and real data testimony concerning these two different approaches. We apply the HP-based methods and HRF to the open-source FRED-QD database mccracken2020fred and FRED-MD database mccracken2016fred. Our empirical findings provide a fairly consistent message about the performance of the HP and bHP filters in relation to the regression filter. Both a small-scale test drive of the procedures on US real GDP and a large-scale deployment to the entire databases are conducted. In brief, the HP filter is able to capture most historical business cycles although the fitted trends tend to be overly flattened or smoothed. The bHP filter is more adaptive to various generating mechanisms and patterns in the trend. On the other hand, HRFs typically seek to reduce observed series to martingale difference residuals and to offer useful mechanisms for prediction and impulse response analysis, but do not provide a method for identifying and estimating general trend and cycle processes, particularly those for which irregularity is a prominent feature, thereby failing to capture key trend and cycle elements of a macroeconomy.
bHP is an $L_{2}$-boosting method applied to time series, drawing on ideas of boosting by repetitive use in the computer science literature freund1995desicion with statistical roots that go back to tukey1977exploratory's introduction of the twicing technique. The key notion is to gradually `boost' an ensemble of many weak learners into a more powerful fitting machine. Boosting has evolved into a very successful machine learning method, with many variants proposed along the way, for instance adaboost friedman2000additive, componentwise boosting buehlmann2006boosting, and $L_{2}$-boosting buhlmann2003boosting.
In addition to the above, some useful theoretical developments and applications of boosting have occurred in the econometric literature bai2009boosting,shi2016econometric,yousuf2021boosting,kueck2023estimation. \citetalias{phillips2021boosting} was the first paper to employ boosting methods in nonstationary time series, much of its asymptotic analysis being built on the foundation of functional limit theory phillips1986understanding,phillips1987time,phillips1987towards. Careful interpretation of regression findings with nonstationary time series and panels has relied on orthonormal series representations of limiting stochastic processes phillips1998new. These methods offer an understanding of nonstationarity in terms of coordinate basis functions and have, in turn, proved useful in analyzing the asymptotic properties of both the HP and bHP filters. Recent years have also witnessed the proliferation in econometric work of machine learning methods, including cross-sectional studies belloni2014inference,caner2018asymptotically,farrell2021deep,athey2021matrix, time series shi2022L,masini2021counterfactual,babii2022machine, and panel modeling su2016identifying,moon2018nuclear, to name a few.
The rest of the paper is organized as follows. Section (ref) introduces the HP and bHP filters and their operational forms, and then presents the basic asymptotic approximation. It is followed in Section (ref) by three applications of the theory to unit root, higher order integrated, and LUR time series on the smoothing properties of the usual HP filter and the consistency of bHP. The theory is supported by simulation exercises in Section (ref). Section (ref) demonstration with HP, bHP and HRF methods applied to US quarterly GDP, and then provides a `big data' empirical implementation of the methods that explores FRED-QD or -MD database to study trends and cycles in the economy. Proofs and additional simulations are given in the Appendix as an online supplementary document.
The squared penalty in the HP filter ((ref)) leads to the explicit solution \[ \widehat{f}^{\text{HP}}=Sx, \] where $S=(I_{n}+\lambda D_{n}D_{n}^{\prime})^{-1}$, $D_{n}^{\prime}$ is the rectangular $(n-2)\times n$ matrix with the second differencing vector $d=(1,-2,1)^{\prime}$ along the leading tri-diagonals and $I_{n}$ is the $n\times n$ identity matrix. The fitted cyclical component (residual) is $\widehat{c}^{\text{HP}}=(\widehat{c}_{1}^{\text{HP}},\widehat{c}_{2}^{\text{HP}},\ldots,\widehat{c}_{n}^{\text{HP}})'=x-\widehat{f}^{\text{HP}}$. These compact expressions are useful in studying the effects of boosting.
The recursive form of the above algorithm is simply $\widehat{c}^{\left(m\right)} =(I_{n}-S)^{m}x$ and $\widehat{f}^{\left(m\right)} =\left[I_{n}-(I_{n}-S)^{m}\right]x$. The bHP filter is easy to implement. For instance, if the original HP filter is called by the hpfilter function in the R package mFilter, then bHP with $m$ iterations can be carried out with a single line of code in the following tidyverse style
where y is the observed time series, m controls the number of iterations, and lambda sets the tuning parameter.
With the tuning parameter $\lambda=\mu n^{4}$ for some constant $\mu>0$ and lag operator $\mathbb{L}$, the asymptotic approximation (as $n\to\infty$) of the HP filtered trend has the operator form \[ G_{\lambda}=\dfrac{1}{\lambda\mathbb{L}^{-2}(1-\mathbb{L})^{4}+1}=\dfrac{1}{\mu\mathbb{L}^{-2}\left[n(1-\mathbb{L})\right]^{4}+1}, \] and the corresponding residual operator is $ 1-G_{\lambda}=\dfrac{\mu\mathbb{L}^{-2}\left[n(1-\mathbb{L})\right]^{4}}{\mu\mathbb{L}^{-2}\left[n(1-\mathbb{L})\right]^{4}+1}. $ The estimated trends of the HP and bHP ($m$ iterations) filters then have the asymptotic forms $\widehat{f}_{t}^{\text{HP}}=G_{\lambda}x_{t}$ and $\widehat{f}_{t}^{(m)}=\left[1-\left(1-G_{\lambda}\right)^{m}\right]x_{t}$, respectively.
The following lemma reveals the effect of the operator when it is applied to a class of complex exponential functions of the form $\exp\left(at/n\right)$.
Lemma (ref) (a) shows that the operator $\left(1-G_{\lambda}\right)$ works as if $\mathrm{e}^{at/n}$ is multiplied by a real factor $\frac{\mu a^{4}}{\mu a^{4}+1}\in[0,1)$ at each operation, because $a^{4} >0$ for any $a\in \mathcal{A}(n)$. Part (b) extends this approximation to include large $m$, where the factor $\left(\frac{\mu a^{4}}{\mu a^{4}+1}\right)^{m}$ vanishes as $m \to \infty$.
Lemma (ref) shows the asymptotic effect of the operator $\left(1-G_{\lambda}\right)^m$ on a simple exponential function. \citetalias{phillips2021business} used the exponential function as an intermediate step in studying the effect of the HP operator involving $\left(1-G_{\lambda}\right)$. In the numerator of the residual operator $(1-G_{\lambda})$, asymptotically when $n\to\infty$ the scaled differencing operator $\left[n\left(1-\mathbb{L}\right)\right]$ acts like a differential operator and the lag operator $\mathbb{L}^{-1}$ acts like the identity, so the effect of these operations applied to $\mathrm{e}^{at/n}$ is straightforward. The HP operator involves these elementary operators in the nonlinear operator $[\mu\mathbb{L}^{-2}\left[n(1-\mathbb{L})\right]^{4}+1]^{-1}$ and \citetalias{phillips2021business} showed how this complex operator may be analyzed as a pseudo-differential operator. Using this machinery \citetalias{phillips2021boosting} studied repeated applications of the operator to the trigonometric functions in the Karhunen-Lo\`eve (KL) representation of Brownian motion --- see ((ref)) below. This approach is used in developing the asymptotic analysis in the next section.
This section illustrates the use of Lemma (ref) for trends that involve unit roots, higher order integrated time series and local unit root time series.
To demonstrate how the bHP filter moderates the residual component in the trend fitting process, we begin with simple unit root time series that was fully analyzed in \citetalias{phillips2021boosting}. The following discussion is heuristic to reveal the manner in which the moderation operates. Consider an observed time series $x_{t}$ generated as an $I(1)$ stochastic trend from a unit root process $f_{t}=f_{t-1}+u_{t}$ with initial value $f_0=o_p(\sqrt{n})$ plus a stationary cyclical component $c_t$. The scale-normalized series satisfies the functional law phillips1992asymptotics
where “$\rightsquigarrow$” denotes weak convergence in the relevant probability space, $B(\cdot)=BM(\omega^{2})$ is a Brownian motion with (long run) variance $\omega^{2}$, and $\lfloor\cdot\rfloor$ is the integer floor function. The KL representation of this Brownian motion over the interval {[}0,1{]} is
where $\xi_{k}\sim i.i.d.\ N(0,\omega^{2})$ are the random coefficients, $\lambda_{k}=1/[(k-\frac{1}{2})\pi]^{2}$ are the eigenvalues, and $\{\varphi_{k}(r)=\sqrt{2}\sin[(k-\frac{1}{2})\pi r]=\sqrt{2}\sin(r/\sqrt{\lambda_{k}})\}_{k=1}^{\infty}$ is an orthonormal system of corresponding eigenfunctions. The series (ref) converges almost surely and uniformly over $r \in [0,1]$.
When the innovation $u_{t}$ follows a linear process as in
with $C(z)=\sum_{j=0}^{\infty}c_{j} z^j$, $\varepsilon_{t}=i.i.d.\left(0,\sigma_{\varepsilon}^{2}\right)$ and $E\left(\left|\varepsilon_{t}\right|^{p}\right)<\infty$ for some $p>4$, we construct an expanded probability space with a Brownian motion $B(\cdot)$ for which uniform convergence holds almost surely phillips2007unit, viz.,
In this space the convergence ((ref)) takes the strong form
In what follows and unless otherwise stated, we assume that we are working in this expanded probability space. In the original space the results translate, as usual, into weak convergence mirroring ((ref)).
Write $\varphi_{k}(t/n)=\sqrt{2}\sin(\frac{t/n}{\sqrt{\lambda_{k}}})=\sqrt{2}\:\textrm{Im} \left[\exp(\frac{\mathbf{i}t/n}{\sqrt{\lambda_{k}}})\right]$, where $\mathbf{i}:=\sqrt{-1}$ is the imaginary unit and $\textrm{Im}[\cdot]$ gives the imaginary part of the argument. When the operator $\left(1-G_{\lambda}\right)^{m}$ is applied to the $k$th term of ((ref)) for any fixed $k$, Lemma (ref) gives
so that when $m$ is large we have
as $m$ and $n$ pass to infinity. With careful handling of a finite-term approximation to the infinite series in ((ref)), \citetalias{phillips2021boosting} showed that when $\lambda=\mu n^{4}$ the residual of the bHP filter becomes \[ n^{-1/2}\widehat{c}_{\left\lfloor nr\right\rfloor }^{\left(m\right)}\approx\left(1-G_{\lambda}\right)^{m}n^{-1/2}x_{t=\lfloor nr\rfloor}\rightsquigarrow0, \] thereby recovering the consistency of the trend component $n^{-1/2}\widehat{f}_{\left\lfloor nr\right\rfloor }^{\left(m\right)}\rightsquigarrow B(r)$, as there is no regular cycle component in the unit root model.\footnote{Unit root and many other nonstationary models generate stochastic trends. Such trends often produce trajectories with some form of highly irregular cyclical behavior because pure unit root processes, just as Brownian motion, return infinitely often to the origin; but such cycles are certainly not regular dynamic cycles of the type generated by stationary dynamic autoregressive processes or which were, at least in the past, associated with deterministic dynamic systems. Thinking about business cycles has inevitably evolved in recent years, taking into account historical experience where the period, intensity and irregularity of business cycles and recessions are frequently acknowledged and characterized in the popular parlance by which they are commonly distinguished, as discussed in \citetalias{phillips2021boosting}. Stochastic trends of very general forms may provide appropriate generative models that encompass a wide range of time series trajectories, inclusive of such irregular cycles, that are suited to economic series like unemployment, interest rates and inflation that are mildly nonstationary, as well as series like GDP that typically have some near random walk component.}
Many other types of nonstationary time series besides $I(1)$ processes occur in macroeconomic data. For instance, aggregate measures of the money supply and nominal price series are often well modeled by higher order integrated time series, particularly by $I(2)$ processes johansen1995stastistical,haldrup1998econometric. The KL series of the limiting Brownian motion process involves orthonormal series of sine functions but it is equally clear from (ref) and (ref) that similar shrinking factors apply to cosine series and more general trigonometric polynomial functions. Such functions figure in series representations of higher order integrated processes.
Suppose the stochastic trend $f_{t}$ follows a higher order integrated process $I(q)$, for some integer $q\in\left\{ 2,3,\ldots\right\}$, of the form
where $u_{t}$ is a linear process satisfying ((ref)). The repeated summation form of $f_t$ from initialization at $t=0$ gives $f_t = \sum_{j_q=1}^t \sum_{j_{q-1}=1}^{j_q} \cdots \sum_{j_1=1}^{j_2}u_{j_1} + p_{q-1}(t)$ where $p_{q-1}(t)$ is a polynomial in $t$ of degree $q-1$ with constant coefficients. Standard weak convergence methods lead to the following limit process of the observed time series after rescaling
as $n \to \infty$. Uniform convergence in ((ref)) ensures that in the expanded probability space we have the corresponding result for $x_{t}/n^{q-0.5}$, viz.,
The orthonormal series representation of $B_{q}(r)$ is obtained by termwise integration in view of the uniform and almost sure convergence of the KL series for Brownian motion in (ref), giving
The braces in (ref) include two terms: $\sum_{\ell=1}^{\lfloor q/2\rfloor}(-1)^{\ell-1}\lambda_{k}^{\ell}r^{q-2\ell}/(q-2\ell)!$ is a $\left(q-2\right)$th order polynomial, and $\lambda_{k}^{q/2}\text{Im}\left[\left(-\mathbf{i}\right)^{q-1}\mathrm{e}^{(\mathbf{i}/\sqrt{\lambda_{k}})r}\right]$ alternates between a sine (for odd $q$) and a cosine (for even $q$).
It was long held as conventional wisdom that the HP filter removed up to four unit roots, thereby detrending integrated processes up to the fourth order king1993low. This claim was disproved by \citetalias{phillips2021business} for $I(1)$ processes under the expansion rate $\lambda=\mu n^{4}$, which was shown to match common quarterly time series applications in macroeconomics. The next result does the same for $I(2)$ processes, showing that the smoothing property of the HP filter continues to apply in this case as $n\to\infty$, giving a limit representation that is a smoothed version of $B_{2}(r)$ rather than $B_{2}\left(r\right)$ itself.
Under $\lambda=\mu n^{4},$ Proposition (ref) shows that the HP filtered $I(2)$ trend approaches a limiting stochastic process $f^{\mathrm{HP}}(r)$ which deviates from the limiting trend process $B_{2}(r)$ and is therefore inconsistent for this expansion rate of $\lambda.$ Correspondingly, the estimated $\widehat{c}_{t}^{\text{HP}}$ of the cycle component $c_{t}$ has the following limiting functional form
upon standardization in the expanded space. This limit function is a stochastic process that inherits some of the stochastic trend properties of the limiting process $B_{2}(r).$ It is therefore to be expected that with a smoothing parameter that approximates $\lambda=\mu n^{4},$ the HP filter will fail to remove all the trend properties of the $I(2)$ process and the imputed business cycle estimate $\widehat{c}_{t}^{\mathrm{HP}}$ will inevitably carry some of these `spurious' characteristics. Note also that at the origin the filtered series limit function is $f^{\mathrm{HP}}(0)=\sum_{k=1}^{\infty}\frac{\sqrt{2} \mu \lambda_{k}}{\mu+\lambda_{k}^{2}} \xi_{k} = -c^{\text{HP}}(0) \not = 0$, a random, mean zero, initialization that is different from $B_2(0)=0$.
The inconsistency of the HP filter estimate of $B_{2}(r)$ in (ref) is anticipated in view of \citetalias{phillips2021business}'s earlier findings for HP filtering of an $I(1)$ stochastic trend. The next result shows that boosting restores consistency to the HP filter for $I(2)$ and higher order integrated time series. For practical applications, the results for $I(2)$ time series are clearly the most relevant.
When $q=1$ this result includes the unit root $I(1)$ case with $B_1=B$. The generalization for $q\geq 2$ follows by use of Lemma (ref) and the asymptotic representation of repeated applications of the HP operator on the exponential functions.
While models with unit roots provide a prototypical framework for capturing persistence in time series data, these models have modifications designed to capture a wider class of time series behavior in which the autoregressive roots are not restricted to unity as they are with integrated processes. An important subclass of more general models with near unit roots phillips2023estimation is the class of LUR models
in which the autoregressive root ${\rm e}^{c/n} \approx 1+c/n$ is local to unity for some constant $c\in\mathbb{R}$ and large $n$. These models were developed in phillips1987towards,chan1987 and have been used for power analyses and in empirical research to provide robustness against pure unit root specifications. Time series generated by (ref) are nonstationary and, after suitable standardization, the observed time series $x_{t}$ converges to a linear diffusion, or Ornstein-Uhlenbeck (OU), process
as $n \to \infty$, where $B(\cdot)$ is Brownian motion with variance $\omega^2$ as in (ref). When $c<0$ the limiting OU process $J_{c}$ is stationary and mean-reverting; when $c>0$ the process is explosive phillips1987towards,phillips2007limit.
Using Lemma 3.1 of phillips2007unit, the uniform convergence law ((ref)) holds, ensuring that
in the expanded probability space. A convenient series representation of $J_{c}(r)$ is
as derived by (ref) in the Appendix. Note that the first component of (ref) has the form
in which $\sum_{k=1}^{\infty}\sqrt{\lambda_{k}}\varphi_{k}(r)\xi_{k}=B(r)$, corresponding to the leading (non-differentiable) Brownian motion component in the decomposition $J_{c}(r) = B(r)+c\int_{0}^{r}\mathrm{e}^{(r-s)c}B(s)ds$. The remaining terms of (ref) provide a series representation of the smooth component $c\int_{0}^{r}\mathrm{e}^{(r-s)c}B(s)ds$ of $J_{c}(r)$, one of which is the exponential term $\sqrt{2}c\nu\mathrm{e}^{cr}$ with a random Gaussian coefficient $\nu:=\sum_{k=1}^{\infty}\frac{\lambda_{k}}{1+c^2\lambda_k}\xi_{k} \sim N\left(0, \sigma^2_{\nu}\right)$, where $\sigma^2_{\nu} = \omega^2\sum_{k=1}^{\infty}\left(\frac{\lambda_{k}}{1+c^2\lambda_k}\right)^2$.
The following proposition shows that the limit representation of the HP filtered LUR time series is inconsistent, yielding a smoothed version of the diffusion $J_{c}(r).$
Whereas the HP filter fails to fully capture an exponential trend function, this objective is fulfilled by the bHP filter, as confirmed in the next result.
For any finite $c\in\mathbb{R},$ Theorem (ref) shows that boosting the HP filter removes stochastic trend components in the residual and provides consistent recovery of a local to unity trend. With respect to the arguments given above in Remark (ref) regarding the effects of filtering on a finite degree time polynomial trend, the limit theory of the bHP filter in \citetalias{phillips2021boosting} (Theorem 2, p. 555) shows that the bHP filter in $m$ iterations removes a polynomial trend of degree $\left(4m-1\right)$ from the residual cycle. Passing $m$ to infinity ensures that boosting captures all the terms in a power series expansion of an exponential trend, corroborating the consistency of the bHP filter given in Theorem (ref) here.
Time series models sometimes explicitly include deterministic trends and structural breaks as constituent trend components. Such components may be analyzed for the wider class of stochastic trend functions considered in this paper. In particular, following \citetalias{phillips2021boosting} we consider a deterministic trend given by the time polynomial
It is a direct corollary that bHP consistently estimates the trend if the underlying stochastic trend is accompanied by an additive polynomial deterministic trend.
A parallel result holds giving consistency of the bHP filter for an LUR time series with deterministic time polynomial drifts and structural breaks. Simply set $q=1$ and replace $B_{q}\left(r\right)$ by $J_{c}\left(r\right)$ in Corollary (ref). For any stochastic trend considered in this paper, if the deterministic trend is consists of finite piece-wise polynomials, then consistent estimation can be obtained following \citetalias{phillips2021boosting}'s Theorem 3. Full statements are omitted.
Given the consistency with the deterministic trends, it is easy to see that finitely additive normalized combinations of these trends are correspondingly included. For example, suppose \[ x_{t}=\frac{x_{t}^{\left(1\right)}}{n}+x_{t}^{\left(2\right)}+\sqrt{n}\alpha+\frac{\beta}{\sqrt{n}}t \] where $x_{t}^{\left(1\right)}$ is an $I(2)$ trend as in ((ref)), $x_{t}^{\left(2\right)}$ is an LUR trend as in ((ref)), and the respective innovations of these time series $(u_{t}^{(1)},u_{t}^{(2)})'$ are potentially correlated random variables satisfying a bivariate version of ((ref)). Then the boosted HP filter reproduces the asymptotic form of this combined trend process so that \[ n^{-1/2}\widehat{f}_{\lfloor nr\rfloor}^{(m)}\rightsquigarrow B_{2}(r)+J_{c}(r)+\left(\alpha+\beta r\right) \] given $\lambda=\mu n^{4}$ as $m,n\to\infty$. Such results demonstrate the versatility of the boosted filter in dealing with complex forms of nonstationary time series.
Practical implementation of boosting requires two tuning parameters $\lambda$ and $m$ to be selected. For low frequency macroeconomic data, \citetalias{phillips2021boosting} suggested that the conventional choice $\lambda=1600$ be used for quarterly data, so that the first iteration gives the HP filter. This choice provides a benchmark that can be adjusted in the case of annual or monthly data ravn2002adjusting. To determine the boosting number $m$ \citetalias{phillips2021boosting} proposed a data-driven stopping rule based on a Bayesian-type information criterion (BIC)
which is motivated by a bias-variance trade-off.\footnote{In addition to the BIC (ref), \citetalias{phillips2021boosting} also suggested a stopping rule based on an augmented Dickey-Fuller unit root test, according to the notion that cyclical behavior should not exhibit unit root behavior. Simulations in \citetalias{phillips2021boosting} showed that (ref) generally provided better finite sample performance and has therefore been used here. } Unless explicitly stated otherwise, the numerical work of bHP in this paper employs the following algorithm.
Code to implement the above algorithm is provided. In R software simply call
where x is the observed time series, and this function will return the estimated trend and cyclical components, along with the sequence of values of $IC\left(m\right)$ and the number of iterations $\widehat{m}$.
In the following numerical exercises, we generate time series in the general form of ((ref)). Along with various stochastic and deterministic trend processes, we specify a cyclical component $c_t$ satisfying a stationary AR(2)\footnote{ We have experimented with many designs of the stationary component and found that the relative performance of the filters is robust. Due to limitations of space we report here the results of a single illustrative design.} $$ (1 - \beta_1 \mathbb{L} - \beta_2 \mathbb{L}^2)\, c_{t}= e_{t}, \ \ \mbox{where } e_t \sim i.i.d.~N(0,\sigma^2_e). $$ To mimic two-year business cycles, we fix $\beta_1 = 1$ and set $\beta_2 = -0.5469$ for quarterly data to produce a spectral density function\footnote{See shumway2000time for the formula of spectral density functions for ARMA processes.} peaking at $\omega = 1/8$ corresponding to a periodicity of $8$ quarters. Similarly, we set $\beta_2 = -0.3492$ for monthly data so that the spectral density peaks at $1/24$, reflecting a periodicity of $24$ months.
We first design three data generating processes (DGPs) based on an $I(2)$ trend. DGP1 is a simple $I(2)$ process, upon which DGP2 adds another 3rd order polynomial trend, and DGP3 further involves a structural break in the middle. All these cases are covered by the asymptotics in Section (ref).
While we set $\sigma_e=5$ for DGPs 1--3 so that the stationary component $c_{t}$ is non-negligible in finite samples in relation to the $I(2)$ trend, we reduce $\sigma_e$ to 1 for the LUR trend to match the settings in \citetalias{phillips2021boosting} where unit root trends were employed, since LUR processes have less prominent trend behavior than $I(2)$ processes. The following designs in DGPs 4--6 are used with a baseline LUR model parallel to DGPs 1--3. LURs from three groups (near explosive, unit root, near stationary) are used with $c \in \{3,0,-3\}$.
We use “quarterly data” and “monthly data” to refer to the respective sample sizes mimicking the FRED databases. Following Algorithm (ref), we use $\lambda=1600$ for quarterly data. Sample sizes are $n=100$ (25 years) as a baseline, while $n=200$ (50 years) and $300$ (75 years) are comparable in size to the empirical application, where most time series in the FRED-QD database have 258 quarters. For monthly data $\lambda=129600$ is used and sample sizes are scaled up by a factor of three, giving $n \in \{300, 600, 900\}$ which is in line with the FREQ-MD database of 775 months.
The original HP filter, `twicing' $m=2$ iterated HP filter (2HP) hall2024selecting, bHP and the autoregressive method are applied to trended time series based on $I(2)$ and LUR. For each replication, the mean squared error (MSE) is calculated as the sample average of $(\widehat{f}_{t}-f_{t})^{2}$, which is equivalent to the average of $(\widehat{c}_{t}-c_{t})^{2} = \left((x_{t}-\widehat{c}_{t})-(x_{t}-c_{t})\right)^{2} = (f_{t} - \widehat{f}_{t})^{2} $. That is, the square of the fitting errors of the trend and the cycle are identical.
We report the empirical MSE averaged over 5000 replications. Table (ref) displays the MSEs for DGPs 1--3 where the stochastic trends are based on $I(2)$ time series with additional trend and break components. It is evident that 2HP improves the original HP filter, and the bHP obtains the lowest MSEs. The MSEs for bHP for each sample size are all close across DGPs 1--2, suggesting that the additional deterministic trend components are estimated accurately and the estimation errors stem primarily from the $I(2)$ stochastic trend. When the tuning parameter $\lambda$ rises from $1600$ for the quarterly data to $129600$ for monthly data, the MSEs for HP and 2HP are inflated in many cases tenfold or more. Evidently, the conventional choice $\lambda=129600$ for monthly data seems much too large for HP trend detection in monthly $I(2)$ time series. By contrast the same large tuning parameter has a considerably muted effect on the trend detection performance of bHP.
For the HRF, $(h,p)=(8,4)$ is set for quarterly data and $(h,p)=(24,12)$ for monthly data following hamilton2018you. The MSEs from HRF are far larger than those from the HP filters appear increasing substantially across sample sizes, revealing the limitations of regression specifications as a flexible trend detection device.
Similar phenomena and rankings of MSEs are observed in Table (ref) for DGPs 4--6 when the underlying stochastic trend is LUR. When $n=100$ for quarterly data and $n=300$ for monthly data, the near-explosive case with $c=3$ yields larger MSEs for HP. Under other sample sizes the results are stable across the values $c \in \{3,0,-3\}$. The MSEs from bHP are only slightly affected when $\lambda$ is raised from 1600 to $129600$. Again, the bHP filter is uniformly superior to the competitors.
Regarding the number of iterations, under the $I(2)$ trends the average $m$ for bHP ranges from 2.5 to 4.2 for quarterly data, and from 15 to 20 for monthly data. This result is consistent with our theoretical result that a larger $\lambda$ in the HP filter for monthly data leads to an overly smooth estimated trend and thus requires more iterations for consistency. Under the LUR trends the two ranges become 2.2--8 and 30--60.
In summary, the simulation results show that, in terms of MSEs the conventional choices of $\lambda$ for quarterly and monthly data appear too large for good trend detection by the HP filter. In all cases, HRF delivers worse MSEs in trend detection than HP. It is clear from these results that the iterated procedure makes the initial choice of the penalty parameter $\lambda$ less critical for performance, compensating for the shortcomings of HP particularly with more complex trend processes, thereby providing a more robust approach to general trend estimation than the other methods. Results in \citetalias{phillips2021business} did show that when the penalty $\lambda=o(n)$, the HP filter is consistent for a stochastic trend when $n \to \infty,$ but no specific rate was suggested or explored. This finding is therefore not particularly helpful for empirical practice. But coupled with the consistency of bHP when $\lambda=\mu n^4$ and the robustness of bHP to choice of the penalty $\lambda$ in initiating iterations, it suggests the existence of a linkage between an `optimal' penalty $\lambda$ and the boosting iteration number $m$ that assures consistent trend estimation. Possible linkages of this type are being considered in ongoing work by yamada2023 and biswas2022longrange.\footnote{biswas2022longrange show that if $\lambda=c\mu n^4$ and $\widehat \lambda = \widehat c(m) \mu n^4$ is chosen to minimize the $L_2$ norm between the HP filter and the bHP filter after $m$ iterations, then the resulting HP with penalty $\widehat \lambda$ is consistent when $ \widehat c(m) \to 0$ as $m,n \to \infty$. yamada2023 derives a relationship between the smoothing matrices involved in the HP and bHP filters whose properties confirm the finding of \citetalias{phillips2021boosting} that `the asymptotic effect of increasing $m$ is similar to reducing the value of $\lambda$ in the simple HP filter'.}
At the time of writing, US real quarterly GDP in the FRED-QD database consists of 258 quarterly observations from 1959:Q1 to 2023:Q2. In Figure (ref) the shaded time periods mark recessions identified by the National Bureau of Economic Research (NBER).
In the left panels,\footnote{Figures (ref) and (ref) are left-rotated $90^{\circ}$ to landscape mode in the paper. So the left (right) panel of the paper appears in the lower (upper) position in the rotated display. Similarly, the top, center and bottom panels refer to the unrotated graphics, whereas they are shown in $90^{\circ}$ rotated view in the paper. The descriptions in the text refer to the original unrotated orientation of the graphics.} the black dots in the graphics display the raw data, and the colored solid lines are the trends as estimated by the HP filter (red line, top), bHP (blue line, middle), and HRF (violet line, lower).\footnote{The 2HP `twicing' filter was also calculated for all empirical examples but the results are not displayed as they are very close to those of HP.} The HP filter produces a smooth, monotonically increasing trend. bHP is more responsive with a slight dip during the long recession in 2008--09, leaving the residual to capture that long cycle; and the trends from HP do not decline in the short Covid-19 recession despite the enormous single period drop in the raw data in 2020:Q2, again leaving this dynamic to the cycle. By contrast, with $h=8$ the HRF largely shifts the observed data to the right by two years.
More revealing are the estimated cycles shown in the right panels, which measure the gap between the raw data and the estimated trend. In the zoomed-in sub-figures, each data point is discernible. The residuals from HP and bHP exhibit clear business cycles that capture the trough in all recessions. The GFC was the most severe prior to 2020, but it was dwarfed in magnitude by the Covid-19 recession. The cyclical component took 1 year to return to positive after Covid-19.
The estimated cycle by HRF in Figure (ref) shows some peculiarity. In all of the 21st-century three recessions, it take 2 years to recover, as marked by the dotted vertical lines. In particular, we observed an unprecedented peak in 2022:Q2, which is the consequence of the colossal slump in 2020:Q2 entering the 8-period ahead predictive model for the first time. These reversions are symptomatic of compensating over-prediction that can occur with autoregressive data fits,\footnote{See Figure (ref) for $h = 1, 4, 8$ and $12$ for this built-in feature of HRF.} which in turn accentuate the swings in actual economic activity.
This section gives a `big data' application, making full use of the time series in these two databases to draw an overall picture of trend and growth cycles in the US economy. There are 246 quarterly series in the FRED-QD database and 127 series in the FREQ-MD database. This broad range of time series includes a mixture of trend behaviors.
The quarterly data findings are presented first. FRED-QD is a comprehensive collection of macroeconomic time series, covering 14 categories of economic activity concerning, for example, the national income and product accounts, the labor market covering unemployment, earnings and productivity, and financial markets covering interest rates, money and credit. Most of the 246 individual time series are complete, with 258 quarters spanning 1959:Q1 to 2023:Q2. A few sequences have missing values and the shortest one has 121 quarters. Many of the time series show strong trend behavior, such as the real GDP, the consumer price index, and stock market indices such as the S&P 500. Other series have random wandering behavior with a more limited range, such as the unemployment rate and interest rates, although the historical patterns vary considerably over the sample period. Our approach to the analysis of the time series in these databases is agnostic, using the HP, bHP and HRF methodologies to filter each individual time series without any pre-processing. Regardless of the sample size, we maintain the use of the standard setting $\lambda=1600$, just as in the simulations.
These time series have their own distinctive scales of measurement. GDP, for instance, is in trillions of US dollars whereas interest rates are reported in small percentage terms. To make the time series comparable for group assessment we take the filtered residuals and standardize each estimated residual sequence to have unit sample variance. These standardized residuals provide rough individual measures of the nature and form of business cycles over the historical period.
Of particular interest are the shapes of the estimated cycles around periods of recession. Some of the time series require reorientation for this purpose as their natural movements reverse normal directions during recessions, e.g., by uptrending in recession periods. These are manually identified and their signs changed. Thus, while most variables decline as economic activity diminishes, unemployment rates rise and so flipping the sign of the unemployment rate to negative brings the direction of movement in line with the majority of the other variables. Accordingly, upon inspection minus signs are assigned to ID 58-72 and 197 (unemployment) and ID 158-162 (money stocks) in the database.
The mass of thin green lines in Figure (ref) show the cyclical components of the time series as estimated by the HP filter (upper panel), the bHP (middle panel), and the HRF (lower panel). These individual residual series are evidently noisy, and the collection of 246 lines in a single graph merge together into a dense green shading where individual movements are hardly visible. We add solid lines (red for HP, blue for bHP, violet for HRF) to aggregate the individual sequences by simple cross section averaging into ensemble indices. These indices provide summary indicator measures of the business cycle status of the economy that broadly reflect the movements of the individual time series estimates.
The average indices generated by the HP and bHP filters move in a fairly synchronized way, particularly around the recession periods identified by the NBER which are marked by the grey shaded areas. For instance, after the dot-com bubble burst in March 2000 these indices each decline to the trough during the 2001 recession of the early 2000s, matching the contraction phase of the business cycle. The bHP index then begins to rise again with the recovery, whereas the HP index is slower and takes much longer to recover. For the great recession of 2008--2009, these two indices again decline to a trough as the financial sector turmoil spreads through the real economy during the economic contraction and then subsequently go through a slow return to normalcy matching the slow recovery with the HP index again taking longer to recover. The recent Covid-19 recession was of much greater scale and the recovery followed swiftly. Outside of the recession periods, business and growth cycles are still visible in the movement of the indices, such as the long expansion in 1990s and the strong economic activity prior to the Covid pandemic. Overall, from the historical data these two indices appear in line with consensus perceptions of US economic activity. Although the differences between these two indices are subtle, the bHP filtered index seems more responsive to historical movements in the data and more in line with the chronology of the recession phases than the HP filter. The results for the HRF again feature the enormous peak in 2022:Q2, and the lagged recovery after the dot-com bubble and GFC.
To check the robustness of these findings from quarterly data, parallel exercises are conducted for the 127 individual time series in the monthly FRED-MD database. The tuning parameter $\lambda=129600$ is used for HP and bHP, and a regression filter with $(h,p)=(24,12)$ is employed. Negative signs are given to series 25--31 (unemployment) and series 70--73 (money stock). The scale-standardized residuals are plotted in Figure (ref) again in thin green lines and the colored solid lines are simple averages of these. Similar patterns are observed from these monthly data and the filtered indices are now evident with fine grain monthly movements. For example, in the recovery episode after Covid-19, the aggregate HP filtered index are relatively smooth whereas the bHP index shows small fluctuations that accord with waves of new Covid variants and the effects of lockdowns on economic activity.
This paper extends the analysis of the boosted HP filter to higher order integrated processes and time series with roots that are local to unity, possibly coupled with polynomial time trends and structural breaks. The primary asymptotic effect of the HP filter is to smooth the underlying stochastic trend, a property that typically leads to inconsistent estimation but still provides a general picture of how the trend has evolved historically. Boosting the HP filter by repeated application ensures consistent estimation of the full limiting trend process and allows for a rich combination of possible trend behavior including breaks. Practical implementation is facilitated by a data-driven procedure that gives researchers an automated machine learning tool of empirical analysis.
The HP and bHP methods of trend estimation are motivated by the penalized likelihood approach initiated by whittaker1922new which seeks in the context of a model such as (ref) to find the `most probable' trend function $f_t$ over a historical period. Trends, as is now well understood, encompass a vast range of possible behavior in which future behavior is not always predictable from the past. It is for this reason that economists are so often wrong in assessing macroeconomic and financial market tendencies, to wit: whether inflation will be transient or sustained, whether there is stock market exuberance and potentially serious consequences of collapse,\footnote{Recall Queen Elizabeth II's timely and famously insightful comment at the London School of Economics (5 November, 2008) on the 2008 GFC collapse that `It's awful --- why did nobody see it coming?' } when a recession will occur, how extensive it will be, how long a recovery will take and whether an earlier trend path will be resumed. In a model of the form (ref) representing behavior over an historical period, such characteristics are implicitly embodied.
The alternative paradigm of pure predictive modeling was strongly advocated by hamilton2018you and motivated by the following characterization:
This conceptualization acknowledges the possibility of model misspecification in that future paths, including cycles but also trends, may differ from those that may be anticipated from past observations. In doing so, the view implicitly accepts that predictive models such as autoregressions may well be incompatible with historical models such as ((ref)) that seek to graduate the data to understand the trends and cycles that have occurred in the past. The choice of $h$, based on the researcher's prior knowledge of length of the business cycles, plays a fundamental role. In Figure (ref), we plot the estimated cyclical component of the US GDP under $h=1, 4, 8,$ and $12$. When $h=8$, the curve is exactly the same as the bottom-right panel of Figure (ref). The dotted vertical lines are $h$-period after the end of the recessions, and the colors are associated with those of the curves. According to this graph, how long the economy recovers from the recession is not determined by the data, but by the researcher's choice of $h$. This feature is particularly pronounced in the enormous peaks after the Covid-19 recession.
Trends and cycles in real world economies are complex and often evolve in unanticipated ways even though the mechanisms and behaviors that drive them may well have some common characteristics reinhart2009time. As \citetalias{phillips2021boosting} (p. 551) remarked:
Such phenomena are difficult to faithfully capture by a supervised learning method like an autoregression that uses only lagged observations to predict future outcomes. By contrast, bHP is a nonparametric unsupervised machine learning approach that is capable of extracting a wide class of underlying trends and cycles from the data. In doing so, the method accommodates historical decomposition of the type (ref) and recognizes the existence of underlying trend and cycle formations that may take irregular and unpredictable forms. The asymptotic theory, simulations and empirical applications reported here all corroborate these advantages.
{\singlespacing } \processdelayedfloats
\setcounter{footnote}{0} \setcounter{table}{0} \setcounter{figure}{0} \setcounter{equation}{0}
\newtheorem{theorem}{Theorem}[section] \newtheorem{lemma}{Lemma}[section] \newtheorem{corollary}{Corollary}[section]
\theoremstyle{remark} \newtheorem{remark}{Remark}[section]
\onehalfspacing