Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
72,302 characters · 15 sections · 68 citation commands
15pt Business Cycle Synchronization in the EU: A Regional-Sectoral Look through Soft-Clustering and Wavelet Decomposition
Since the introduction and later adoption of the euro as a common European currency in 1999, both academia and market participants have been interested in the phenomenon of business cycles' synchronization among European economies. The outbreak of the global financial crisis in 2008 and the consequent European sovereign debt crisis revived debates about the optimality of the European Monetary Union (EMU) due to observed discrepancies between groups of countries within and outside the euro area. In this context, the most lingering question underlying the Optimal Currency Area (OCA) theory is understanding if the euro area members are similar enough to share the same currency in the long run. If they are, does the homogeneity lead to a core-periphery division, and of what complexity (bengoechea2006useful). The latter question can be addressed through a grouping or clustering perspective, considering the cyclical (dis-)similarities between the core and more `vulnerable' regions (see, among others, fidrmuc2006meta, stanivsic2013convergence, di2016business, ahlborn2018core, rathke2022similar, arcabic2022business, de2022coherence).
Mitigating the risk of the euro area fragmentation, in 2012, the European institutions undertook unprecedented actions. The ECB, responsible for the EMU monetary policy, contributed by using the so-called `quantitative easing' strategy, while the public sector of the core economies -- by applying for purchase programs. At the same time, the periphery countries significantly tightened their fiscal policy and banking sector supervision (COUDERT2020428). The importance of business cycle synchronization comes then into play as the necessary OCA condition. Otherwise, for the monetary union with asynchronous regional business cycles, the single monetary policy will not optimally stabilize the common shocks for some of its members. Besides, if the shapes of synchronized cycles are different, individual policy responses to global economic shocks may be too accommodative for some countries, and too tight for the others. This mismatch implies the smooth state changes extending the duration of recoveries from undesirable overheating-recession states. Furthermore, the economic policies may last too long for countries with shorter business cycles and too fast for countries with lengthier ones (bengoechea2006useful).
Examination of business cycles and their components goes back to the OCA theory (see, among others, mckinnon1963optimum, frankel1998endogenity). In these papers, the authors argued that the lack of an independent monetary policy potentially leads to a significant loss of welfare and even a breakdown of a monetary union if the members exhibit asymmetric or asynchronous output fluctuations. On the one hand, krugman1991increasing argues that increasing integration might lead to regional concentration of industrial activities, which would, in turn, result in sectoral-regional specific shocks, thereby increasing the likelihood of asymmetric shock and diverging business cycles. Furthermore, economic integration through more intense capital flows might lead to a more diverse production structure of the union member states and a trade growth, therefore, improving the business cycle synchronization (kalemli2001economic). On the other hand, frankel1998endogenity suggests that the removal of trade barriers will lead to easier transmission of any demand shocks across countries, implying more symmetric fluctuations in the business cycles. inklaar2008trade add that as the economic policies in the euro area become more aligned, the business cycle synchronization should also improve.
In this regard, the purpose of current paper is to revisit the OCA business cycle synchronization controversy by analysing the sectoral-regional view of EU-27 Gross Value Added (GVA) data. In other words, we seek to elaborate on the regional dimension by scrutinising its sectoral split as the clusters of economic activities, which may hint at two things. First, the policymakers might choose the relevant economic policies, mitigating the negative overheating or downturn impacts of the synchronous European regions and sectors, and predict the extent of contagion effects stemming from various economic shocks. Second, the less synchronous and asynchronous EU member states and economic activities hint to investors which areas are potentially interesting for portfolio diversification as their cyclical components move against the joint European economic cycle. Therefore, the methods described in the paper equip researchers and decision-makers with new data that helps understanding the sectoral-regional composition of (a)-synchronous business cycle groups in Europe.
From methodological perspective, we contribute to the business cycle synchronization literature in several ways. First, we introduce the relevant methodological framework that combines wavelet decomposition with soft-clustering techniques for identifying the growth cycles. As the single euro area monetary policy directly mitigates demand-driven shocks, we focus on European business cycles observed within 2-8 years window. Second, to improve the clustering results and demonstrate the evidence of synchronization, we propose a synchronization-based dissimilarity measure for clustering. Third, we use several cluster cleaning approaches based on silhouette scores and the bootstrapped soft-clustering probabilities allowing us to control the concentration and tightness of the estimated groups in synchronization. We find that the optimal number of clusters is small. For the purpose of the paper it is sufficient to study three groups -- two purified clusters and the least synchronous drop-out group, which supports the core-periphery literature in the context of the regional-sectoral breakdown of the European economic activities. All these proposals eventually lead to an efficient approach for obtaining the clean synchronous and asynchronous clusters of EU regions and economic activities that fit the different needs of policymakers and investors.
The paper is structured as follows. In Section 1, we present the methodology introducing the soft-clustering approach. In Section 2, we discuss the features of the data, the trade-offs when choosing the set of hyperparameters, and elaborate on the main results of applying the soft-clustering approach to EU-27 GVA data for 11 segments of the economy. In this section, we primarily discuss synchronous cycle composition and elaborate on the core-periphery hypothesis. Section 3 provides concluding remarks.
In this paper, we scrutinize quarterly, seasonally and calendar-adjusted, real (in chain-linked volumes sense, index 2010=100) GVA data sourced from Eurostat, spanning over $\mathcal{T} = \{2000Q1, \ldots, 2021Q2\}$, $|\mathcal{T}| = 86$ quarters. Seeking to analyse the regional-sectoral view, the dataset contains 27 EU countries (after Brexit, excluding the UK) and 11 sectors that correspond to A*10 economic activities breakdown taken from NACE rev. 2. The latter breakdown broadly defines the industry as B-E economic activities, excluding construction (F). The sectoral breakdown also narrowly defines the industry as manufacturing activities (C). Since the inclusion of C allows choosing more relevant economic policies or building more diversified investment portfolios, we keep both industry definitions. For simplicity, denoting the EU country indices as $S = \{ 1, \ldots, 27\}$ and the different economic activities as $R = \{1, \ldots, 11\}$, we construct the set of indices $\mathcal{I}:= R \times S$, where $\times$ denotes the Cartesian product and $|\mathcal{I}| = 297$. In sum, throughout the paper, we analyse 297 time series, denoting any such series as $X_{i}(t)$, where the indices $ i \in \mathcal{I}$ correspond to a specific regional-sectoral indicator, measured at a time point $t \in \mathcal{T}$. In this context, $\mathcal{T}$ represents the set of 86 features recorded for each time series.
By definition, any trend-cycle decomposition method decomposes macroeconomic data into low-frequency trend $X^*_{i}(t)$ that may have both stochastic and deterministic parts, a cycle $c_{i}(t)$, likely consisting of sub-cycles of different lengths, and high-frequency irregular noise $\nu_{i}(t)$, typically interpreted as shocks (celovcomunale2021):
The data preparation step then has to isolate the growth cycle by removing the trend and noise with the least possible distortions. For this purpose, we employ the wavelet decomposition approach.
The wavelet transformation is a powerful signal decomposition tool that uses both the time and frequency domains. For applications in the business cycle literature, see, for example, crowley2009fused, bruzda2011business, soares2011business, AGUIARCONRARIA2011477, among others.
Following RePEc:eee:finlet:v:19:y:2016:i:c:p:298-304, we employ the maximal overlap discrete wavelet transformation (MODWT) approach with $J=5$. We apply the Least Asymmetric wavelets with 8 parameters (LA(8)) for the smoother results. Using lengthy wavelets may induce smoothing artefacts at the ends-of-sample, hence, require cautious interpretation of the end-points. Although the MODWT transformation is not orthonormal and is highly redundant, it admits any sample sizes $\mathcal{T}$ that need not be the multipliers of a power of two. Besides, the chosen transformation is invariant to circular shifts and can decompose the data by applying multiresolution analysis. For the latter, using the efficient pyramid algorithm under any base functions (e.g., Haar, Least Asymmetric, or Fejér-Korovkin), the data decomposes to:
where $D_{j}^{(i)}(t)$ is the $j$-th level wavelet and $S_{J}^{(i)}(t)$ is the smooth component.
crowley2009fused argue that, for business cycle frequencies, it is sufficient to analyse $D_3$ and $D_4$ wavelets, spanning over 2-8 year cycle lengths (see Table (ref)). The choice may be viewed as a combination of cyclical fluctuations stemming from variation in inventories (Kitchin) and fixed investments (Juglar) correspondingly. However, we prefer to interpret the business cycle as an aggregate outcome of supply-demand interplay from some global dynamic general equilibrium model that reflects sectoral-regional split of the data and demonstrates responses to a single euro area monetary policy designed to mitigate demand-driven shocks. If one would be interested in studying macro-prudential and fiscal policy impacts, we would also recommend including the credit cycle reflecting component $D_5$ into the common cycle. celovcomunale2021 argue that omission of this component impacts mostly the amplitude of the waves as the business cycle wavelets are placed on top of smoother and longer credit cycle wave. Hence, the local turning points of the common cycle reflect more business cycle characteristics than that of the smoother $D_5$ wavelet.
Using LA(8) base functions, we combine the 297 extracted wavelet series $D_{3}^{(i)}(t)$ and $D_{4}^{(i)}(t)$ from the variables $X_{i}(t)$, $i \in \mathcal{I}, t \in \mathcal{T}$, defining the business cycle frequencies related growth cycle $c_{i}(t)$ by:
The form of the cycle $c_{i}(t)$ returned by wavelet decomposition is sensitive to the appropriate trend elimination procedure applied to processing the data for MODWT. First, following celovcomunale2021, we use the drift-adjusted time series data:
Note, that for notational convenience throughout this paper we will assume that $X_{i}(t)$ denotes the drift-adjusted series by (ref). Second, since the MODWT filter is circular, the way the edges indefinitely attach may significantly distort the end-of-sample behaviour of the estimated cycles. We use the “reflective” approach, which extends the observed data by indefinite mirroring of the data. Unlike the “periodic” extension, a reflection of the latest data should, in theory, keep the signals similar at the ends, while the former may experience spurious structural breaks. “Reflective” connection roughly assumes that the static synchronicity signal goes through possibly unobserved data. However, macroeconomic data may experience time-varying synchronicity between different cycles changing after some global or local events, for instance, due to the financial crisis, COVID-19 outbreak, or after adopting the euro by new euro area members. The drift-adjustment (ref) makes the distortions at the ends-of-sample less apparent.
In this paper, the remainder series $D_5^{(i)}(t) + S_5^{(i)}(t)$ determine the stochastic part of the trend of the original drift-adjusted data that also contains credit cycle associated wavelet, while the first two wavelets define the noise. We observe that the applied definition of a trend is sufficient for most of the analysed GVA data. The trends are similar to those obtained with the Hodrick-Prescott (HP) filter with penalty $\lambda$ optimised for the 8-years cycles. The HP filter is widely used in the applications (see, e.g., dickerson1998business, inklaar2008trade, papageorgiou2010business), the robustness of which is confirmed by artis1997international and dickerson1998business, among others. Hence, the HP filter is an appropriate benchmark for a robustness check. We observe the median correlation of $0.996$ between the series (smallest -- 0.98), with scaled and centred RMSE between the two extracted trends falling within $(0.01, 0.32)$, 90% of which is lower than 0.15. The deviation is more visible at the ends-of-sample, increasing only marginally to 0.2, where both approaches experience larger distortions. Following these observations, we deem the trend elimination using wavelets sufficient in our case.
For the analysis of synchronization, we use the synchronicity measure as in mink2012measuring, defined below by (ref)--(ref). For every time period considered, a value of 1 indicates that the two cycles have the same sign, while a value of -1 denotes cycles of different signs. In other words, for any time period $t \in \mathcal{T}$ and $i, j \in \mathcal{I},\ \ i \neq j$, the synchronicity measure is:
where, by (ref), $c_{i}(t)$ denotes the cycle of indicator $i$ at time moment $t$. Averaging over the total feature space $\mathcal{T}$ defines the through-the-cycle synchronicity measure:
where $\varphi^{(\mathcal{T})}(i,j) \in [-1,1]$ for any pair of indicators $i \neq j \in \mathcal{I}$. The closer the $\varphi^{(\mathcal{T})}(i,j)$ estimates are to 1, the stronger are the co-movements between the indicator pair $(i,j)$ over the period $ \mathcal{T} $.
In addition, to validate the quality of business cycles synchronization for the group of indicators, we consider the measure proposed by RePEc:eee:jmacro:v:37:y:2013:i:c:p:265-284 and defined in the indicator space $\mathcal{I}$ as follows:
where $s_{-i}^{\mathcal{I}}(t)$ denotes the standard deviation within the group $\mathcal{I}$, excluding the indicator $i$, at a fixed time point $t\in\mathcal{T}$, and $s^{\mathcal{I}}(t)$ is the standard deviation including the indicator $i$:
for some $i \in \mathcal{I}, |\mathcal{I}|>1$. The main idea is that the exclusion of any indicator $i$ from some group $\mathcal{I}$ of countries (e.g., euro area or the EU) or sectoral-regional indicators, as in our case, can either increase or decrease the total within-group variance. If we assume that the signals are synchronized and concentrated around the mean cycle, we will observe consistently small standard deviations through $t \in \mathcal{T}$. Thus, the removal of any highly-concentrated indicator $i$ from the group $\mathcal{I}$ should only increase the total variance of the group. On the other hand, if certain signals are far away from the mean consistently, removal of those from the group will decrease the variance, improving the synchronization of the remaining members of the club. Following these ideas, we compute the median statistic over $t \in \mathcal{T}$:
We argue that (ref) is the robust proximity of the synchronization measure. A certain degree of robustness is achieved by taking the median over $t \in \mathcal{T}$, countering against periods with high volatility that may introduce large leverage on the whole series. A negative sign of (ref) suggests that the synchronization improves upon excluding indicator $i$ from the group $\mathcal{I}$, while the positive sign shows that the situation is worse without indicator $i$. In addition, this measure captures the proximity of an indicator $i$ to the mean cycle within the group. Note, that if all indicators within any group are in perfect synchronization (e.g., resulting in $\varphi^{(\mathcal{T})}(i,j) =1$, $i\neq j$), the cycles can still vary in their amplitude. Thus, with sufficiently high variation, the (ref) measure will always find a subset of indicators with negative signs of (ref), unless perfect co-variation occurs. Nevertheless, the measure can be useful to gauge how close are the signals to the mean signal within the group/cluster $\mathcal{I}$. Finally, the underlying idea behind (ref) and (ref), that certain elements of the cluster can reduce the overall synchronicity of the cluster, gave rise to the probabilistic soft-clustering approach, introduced in Section (ref).
In the literature, there are many examples of clustering methods applied in the context of macroeconomic time series or related to the business cycle synchronization (see, e.g., artis2002membership, papageorgiou2010business, ahlborn2018core, COUDERT2020428). In this paper, we consider hierarchical clustering (see, COUDERT2020428 and references within) as the base clustering approach with Ward's minimum variance linkage method that aims at finding compact, spherical clusters (ward2). The main novelty of the paper is the application of synchronicity $\varphi^{(\mathcal{T})}(i,j)$ as the dissimilarity measure together with a soft-clustering approach -- the adjusted version of hierarchical clustering algorithm combined with silhouette scores (see Section (ref)) and bootstrapped probabilistic thresholding (see Section (ref)). For hierarchical clustering, by (ref), we construct a dissimilarity measure
where the closer the $\varphi^{(\mathcal{T})}(i,j)$ is to 1, the less dissimilar is the specific pair of indicators $(i,j)$. Throughout the paper, we define the dissimilarity matrix as:
The use of the whole period $\mathcal{T}$ for $\varphi^{(\mathcal{T})}(i,j)$ allows us to gauge the overall synchronization levels between the considered cycles. However, such approach can miss certain vanishing or very recent co-movements due to averaging. Furthermore, the approach is useful to establish strong co-movements between the cycles' pair $(i,j)$ but is less informative for weaker co-movements. Note, that (ref) and (ref) essentially counts the cycle movements of the same sign throughout the set $\mathcal{T}$ but is not affected by the ordering of the series. In other words, the measure does not reward sequential co-movements of the cycles, losing part of the information from the data. To account for these aspects, in the soft-clustering step, we use dissimilarities $D^{(\mathcal{T})}$ along with a established thresholding rule that enables us to re-weight the cluster probabilities based on the observed synchronization strength (see Section (ref)). The thresholding set of hyperparameters allows us to consider only those indicators that show strong synchronization throughout the larger part of the period $\mathcal{T}$, dismissing the weaker co-moving signals. In fact, in Section (ref) we demonstrate that the excluded cycles fall into the group of asynchronous cycles, the dynamics of which can potentially be very useful, e.g., for investment risk diversification.
Furthermore, seeking to reduce the number of dismissed signals, we could split the time horizon $\mathcal{T}$ into a certain number $s \geq 1$ of smaller time windows $\mathcal{T} = \cup_{j=1}^{s} \mathcal{T}_{j}$ and find the corresponding dissimilarity matrices $D^{(\mathcal{T}_{j})}$, $j = 1, \ldots, s$. In this case, the clustering procedure can be applied for every time window $\mathcal{T}_{j}$. Combining all the resulting clusters through a soft-clustering approach would then introduce a certain reward for those cycles, that show high synchronicity during specified time windows $\mathcal{T}_{j}$. In this paper we consider such a procedure with $s = 3$ time windows, namely, $\mathcal{T}_{1} = \{2000Q1, \ldots, 2007Q4\}$, $\mathcal{T}_{2} = \{2008Q1, \ldots, 2014Q4\}$ and $\mathcal{T}_{3} = \{2015Q1, \ldots, 2021Q2\}$. The idea behind the choice of the splitting dates is to separate the co-movements after introduction of the euro, the shock of the 2008 financial crisis and followed sovereign debt crisis (W-recovery), and the period that ends with COVID-19 outbreak. Similar approach was taken, for instance, by COUDERT2020428, bunyan2020fiscal.
Since clustering is an unsupervised learning approach, it is impossible to use cross-validation. Instead, we propose a soft approach in order to improve the homogeneity of the synchronous cycles containing clusters. For this purpose, we employ silhouette scores, which show promising results and consistently good performance in various applications and simulations (see, e.g., arbelaitz2013extensive). High observed silhouette score values would effectively hint on the level of homogeneity in the synchronous cluster of business cycles.
The general algorithm follows ROUSSEEUW198753, artis2002membership. Given a cluster $C_{m}$, some indicator $i \in C_{m}$, and a dissimilarity function $d_{i,j}$, define
Next, for any $ C_{k}, m\neq k$, define
Any cluster $C_{k}$, corresponding to the smallest value $b_i$, determines the neighbouring cluster for the indicator $i$. Thus, for any $i$, the silhouette score $s_i$ is defined as:
and $s_i=0$ if $|C_{m}|=1$. If, for any indicator $i$, the resulting silhouette score $s_i < 0$, the method suggests that there exists a neighbouring cluster, where the indicator $i$ fits better.
Throughout this paper, we employ the silhouette scores in two ways. First, we use the scores for the choice of the number of clusters when performing hierarchical clustering. For any resulting cluster $C_i$, we estimate the average silhouette score within the cluster, i.e., calculate
For a given number of clusters $k > 1$, we estimate $SC_{i}, ~i = 1,\ldots,k$. The aim is to choose such values of $k$, for which both $\min_{1 \leq j \leq k}SC_{j}$ and $k^{-1}\sum_{j=1}^{k}SC_j$ are the largest.
Second, we use the scores $s_i$ to clean the resulting cluster $C_i$ retaining only the indicators with scores above a specified threshold $\kappa \in \mathbb{R}$. The steps for tidying the data are outlined below:
Throughout this paper we assume that $\kappa = 0$. Therefore, the algorithm stops only when all of the corresponding silhouette scores are positive. However, the downside of such approach is that it can discard too many variables of interest from the analysis due to step (4.) of the algorithm. As a general case, one should consider a range of values for $\kappa$ and fine-tune based on the expected percentage of variables discarded. In our case, setting $\kappa = 0$ was sufficient and resulted in relatively small number of discarded variables from the analysis. Note, that excluded variables do not contribute further for the establishment of synchronizing clusters, however are still valid members of the pool consisting of drop-out indicators, likely capturing the asynchronous growth cycles, which may be of further interest for different investment diversification strategies or economic development policies.
As the main result of the paper, we propose the following soft-clustering algorithm, which requires setting hyperparameters $(\omega_1, \omega_2, \ldots, \omega_6)$:
The probability thresholding works as a penalty allowing to optimize a certain goodness-of-fit statistic for the estimated clusters, given the chosen set of hyperparameters $(\omega_1, \ldots, \omega_6)$. When combined, step (ref) controls the cluster strength under a threshold $\omega_4$, while step (ref) acts as a penalty on the within-cluster dissimilarities. Finally, step (ref) imposes a variance-based restriction. The proposed approach moves uncertain boundary cases into the drop-out pool of business cycles. In this regard, our proposed soft-clustering approach is similar to soft $k$-means (fuzzy) clustering but differs on how we define the gray zone of the drop-out pool.
In the context of this paper, identifying and excluding the “outlier” signals from the remaining less synchronous groups can be an interesting aspect for a deeper investigation. In particular cases, the distinction between specific countries can be increasing over time, as suggested by wortmann2016one and ahlborn2018core. The proposed clustering approach is reasonably flexible when compared with typical “hard” clustering methods. We introduce additional control hyperparameters $(\omega_1, \ldots, \omega_6)$ and a specific dropout function. While $\omega_3, \omega_6$ control the shape of the constructed probabilities, dropout allows cleaning the clusters in the similar sense as Sparse PCA, where the dropped-out series can be inspected separately. Besides, as argued previously, the pool of dropped-out series maybe interesting to the investors who seek to diversify the investment portfolios, since these variables uncover the cycles going against the synchronized segments and regions or do not experience visible cyclical behaviour.
This section summarizes the salient features of the European sectoral-regional business cycles extracted applying the wavelet approach (see Section (ref)). We present an analysis in order to provide a deeper understanding of the extraction and contraction phases, average duration and amplitude of any single cycle and give the initial view of the cycles going to a single cluster.
Figure (ref) demonstrates the single pool of business cycles, together with the first principal component extracted from the scaled data. The principal component analysis produces a low-dimensional representation of the data by finding a sequence of linear combinations of variables, maximizing explained variance. In this case the presented first component explains 32.51% variance proportion. This suggests the existence of a strongly synchronous cluster within the pool of cycles, implying that the soft-clustering approach may improve the core signal recovery.
For the triangular dissection of the business cycle phases, we use bbq quarterly (BBQ) approach as applied in celovcomunale2021. The approach locates the local extrema and applies the censoring rules, ensuring that phases alternate between expansion and contraction; phases have a minimum duration of four quarters; and the complete cycle length is at least eight quarters.
The duration of a phase and the amplitude of changes between alternating peaks and troughs are two sides of the right triangle, the hypotenuse of which shows the average constant growth rate of the phase. The average duration and amplitude allow comparing any pairs of cycles or group averages, examining the symmetry of the stages and computing the average cycle length -- the sum of the arithmetic means of contractions and expansions. The analysis, however, omits the incomplete end-of-sample phases.
Table (ref) summarizes the sample means of the BBQ approach and includes the average scaling factors, denoting the higher raw amplitudes having groups of business cycles. The table also has the average synchronicity measure (ref) with the median business cycle taken as the reference point, in line with mink2012measuring, and the robust proximity of the synchronization measure (ref).
Although the amplitudes, on average, are symmetric, meaning that a similar size recovery follows the contraction phase, the amplitudes scaling factors vary significantly across sectors and regions. The public sector (O-Q) reacts the least to economic shocks in line with the average semi-elasticities of the cyclically-adjusted EU general government balances being well below unity, while construction (F) and agriculture, forestry and fishing (A) respond the most. On the regional level, the Baltic states, Greece, Romania, and Ireland form the EU periphery that responds the most to the economic shocks, while most core countries have relatively small scaling factors.
The analysis of durations demonstrates that the data has, on average, from 1 to 3 quarters longer expansion phases than contraction, with more aligned patterns seen on sectoral and diverse observed on regional levels. In some exceptional cases, the regional level shows symmetric phase durations for Slovenia, Romania, Bulgaria and Belgium, implying higher regional diversity and suggesting that further signal refinement may be achieved by soft-clustering on the sectoral-regional level. The average cycle lengths group is around 17 quarters and, on the sectoral level, ranges from 14 quarters for the A sector to 18.2 quarters for trade, hotels, food and transportation (G-I), while, on the regional level, it varies from 14.8 quarters for Malta to 19.7 for the Netherlands.
Finally, we observe high average numbers on sectoral synchronicity levels for industry (B-E), including manufacturing (C), traditional services (G-I), and R&D-related services (M). These sectors score high on both synchronization measures and suggest that these economic activities potentially form the core of the European synchronous cycle. Regional synchronization diversity hints at the weak evidence for the core-periphery split between the European economies, with the highest scores in Austria, Germany, Spain, Italy, France, and the Netherlands. The segment with the least synchronous economies includes Greece, Luxembourg, Cyprus, and Visegrád Four. Concerning the core-periphery story, we expect a more discernible split recovered by applying the soft-clustering approach.
In this section, we present the soft clustering results considering the full time horizon $\mathcal{T}$, with the number of bootstrapped samples $\omega_1 = 1000$ and $\omega_2 = 1$. Supporting the choice of $\omega_1$, EfroTibs94 recommend the number of bootstrap samples to be above 500 or 1000 and demonstrate that larger number of samples $\omega_1$ only marginally improves the quality of obtained distributions. Hence, we take the larger of the two numbers. Besides, our experiments with three proposed periods' split ($\omega_2 = 3$), when judged jointly, align with the parsimonious total period view suggesting stability of the synchronous cluster composition. Importantly, the resulting view is not the same as when judging the composition utilizing only one period at once -- this might be crucial for detecting dynamic changes in the composition of the synchronous core cycle or providing event-dependent view. From the standpoint of long-lasting economic policy, a stable composition identification is preferable.
In order to find the optimal values for the number of clusters $\omega_3$, the threshold probabilities $\omega_4$ and the dropout rate $\omega_5$, we employ a grid-search approach and examine the resulting cases under different combinations $(\omega_3, \omega_4, \omega_5)$. Appendix (ref) presents the grid-search results, where we compare the average and minimum silhouette scores pre- and post-cleaning of the clusters, as described in Section (ref). We seek to select hyperparameters so that the resulting cluster configuration would result in high overall silhouette scores and the weakest resulting cluster wouldn't be “too weak”. The idea is that the less synchronous cluster can collect all of the noise variables. Such results would infer less synchronization between the variables or that any observed co-movements are complex -- the possible occurrence of phase-shifting or change in frequency dynamics, which fall beyond the scope of business cycle synchronization analysis. Therefore, an alternative solution would be to pool the most asynchronous cycles out of the clusters. Note, there is one caveat about maximizing the silhouette scores. Since, by construction, we drop a chosen fraction of variables based on their observed co-movement strength, the average and minimum scores will increase with increased dropout rates. The trade-off is to eliminate as few variables from the data as possible and find a configuration that retains the highest average silhouette scores. This impact follows from analyzing the resulting cluster sizes in Tables (ref) and (ref) and the corresponding cluster composition of selected cases in Figure (ref).
Based on the silhouette scores analysis we note that the most compromising choice for the number of clusters $\omega_3$ is 2. The split into two underlying clusters and a drop-out pool generates the highest silhouette scores continuously over different values of $(\omega_4, \omega_5)$ (see Tables (ref)--(ref)). Assumption of such signal composition is consistent with the business cycle synchronization literature (e.g., the core-periphery hypothesis). Besides, due to the specifics of the synchronization measure (ref) used to form clustering dissimilarities, separating the cycles into two clusters can be seen as a split of the noise from the signal. In other words, we observe that one group consists of highly synchronized business cycles based on the measure (ref), and another cluster contains less synchronous series as shown in Figure (ref). Further analysis of the “noisy” group (for $\omega_3 > 2$) reveals smaller synchronizing clusters with different phases, durations and amplitudes than the largest synchronizing group of business cycles. The more detailed splits result in the deterioration of average silhouette scores and advocate for higher dropout rates that does not fit the purpose of our research. Primarily interested in European (euro area) business cycle synchronization, below we elaborate on the results based on the most synchronous cluster from the two, paying less attention to the information from the “noisy” group or its further splits. Besides, the clean-up procedure generates the pool of the asynchronous business cycles that may be of further interest for those seeking to recognize potential regions and sectors for optimal investment diversification.
The other two parameters, $\omega_4$ and $\omega_5$, jointly define the number of discarded business cycles, which form the most asynchronous group of business cycles from the dropout pool. By construction, $\omega_4$ controls the threshold probabilities used in forming the distance matrix of the soft-clustering approach (ref), while $\omega_5 \in [0,1]$ denotes the percentile drop-out rate (ref). Note, in certain cases we present values for $(1 - \omega_5)$ when highlighting the upper bound percentage of variables remaining in the corresponding analysis.
Applying the additional adjustments, we discard the noisy series that do not tend to group during the soft-clustering iterations. Setting sufficiently high values for $\omega_4$ reduces the amount of noisy data included in both clusters. From grid-search results provided in Appendix (ref), we observe that the compromise average silhouette scores arise when $0.55 \leq \omega_4 \leq 0.95$, and $0.25 \leq \omega_5 \leq 0.65$ with visible nonlinear dependence between the two parameters. Noteworthy, higher number of clusters $\omega_3 > 2$ tends to shift $\omega_5$ closer to $0.8$ drop-out rate, while keeping $\omega_4$ either at values above $0.8$ or below $0.55$.
Tables (ref) and (ref) show that increasing $\omega_5$ first affects the less synchronous cluster, leaving the number of business cycles in the main cluster almost unaffected, when $\omega_4$ is between $0.55$ and $0.8$. Therefore, when applying the soft-clustering approach we should pay attention to the trade-off between the size of clusters and the silhouette scores that, at some point, improve because of the asynchronous cluster improvements.
Following the above arguments, we argue that setting $\omega_4$ to higher levels, e.g., $\omega_4 = 0.8$ allows us to adjust the clustering results aiming to extract more consistently synchronized clusters throughout the whole feature space $t \in \mathcal{T}$. In other words, setting $\omega_4 = 0.8$ assumes that the remaining series are grouped at least 80% of the time during the iterations of the soft-clustering algorithm, indicating the tightness around the mean cycle and suggesting that any dropped variable from the clusters fails to show strong enough bond with the underlying core cycle. Moreover, varying the values of $\omega_4$ allows us to further separate the resulting clusters based on their grouping strength, for instance, establishing weakly co-moving sectors and separate them from strongly co-moving sectors.
The threshold parameter $\omega_5$ controls the number of dropouts from the total data. Note that the dropout is performed based on the $\omega_4$-induced ordering. Thus, the algorithm should begin dropping the series from the “noisy” group first and foremost, as seen in both Table (ref) and Figure (ref), while only marginally affecting the underlying synchronous cluster. Further analysis of Figures (ref) and (ref) reveals that increasing dropout $\omega_5$ from 0.05 to 0.45 improves the synchronization in all three groups. After soft-clustering, the principal components of the synchronous group correspondingly explain the 53.42% and 60.52% of variance, and the cluster's composition is less sensitive to the dropout parameter increase. The reshuffle between the asynchronous group and the dropout pool is more evident. The principal components are almost unchanged for the asynchronous cluster, while the composition of the dropout pool is less aligned. Around COVID-19 events, the dropout pool shows the strong synchronization among its components and synchronous cycles. At the same time, the asynchronous cluster moves against the common European business cycle and demonstrates the delayed response to the global financial crisis. Based on the grid-search results, we conclude that setting the $\omega_5$ parameter between $0.45$ and $0.65$ is a reasonable trade-off when applied to the European business cycle problem and, in the empirical part of the paper, set $\omega_5 = 0.45$
After carefully choosing the soft-clustering parameters in section (ref), we find that parameter setting (1000, 1, 2, 0.8, 0.45, 2) provides a well-balanced view of the synchronous European sectoral-regional set of business cycles. Figure (ref) summarizes the underlying features of the set, while Table (ref) adds information on which sectors comprise the synchronous, asynchronous clusters and the dropout pool.
The upper left graph of Figure (ref) depicts the stacked view of normalized business cycles, where the bold red solid line corresponds to the first principal component computed from the cluster members. The first principal component defines the European business cycle with the highest 60.52% common variability explained. Judging by the allocation of peaks and troughs, the European business cycle is essentially the same as the signal extracted from the single pool, as demonstrated in Figure (ref). Hence, the crucial difference is the cleaner view of the main driving sectoral-regional business cycles that constitute the synchronous cluster.
The allocation of peaks and troughs of the European business cycle adheres to the most prominent global economic shocks observed in the recent two decades. After 2000, the dot-com bubble burst recovers relatively fast with visibly less synchronous levels among the components and other groups (see Figure (ref)). On the back of the non-prudentially booming credit market, the underlying European cycle slowly moves to the overheating expansion, which speeds up after the enlargement of the EU in 2004. Consequently, the overheated EU market gets a heavy punch from the global financial crisis of 2007-2009, with a double deep recovery due to the subsequent European sovereign debt crisis. Finally, the pressures on the EU labor market resulted in a too hot position for the EU economies before the COVID-19 outbreak. However, since the latest wing is close to the end of the sample, we judge its severity with caution.
The upper right graph of Figure (ref) shows the sectoral composition of the synchronous cluster. Supplemented by the numbers in Table (ref), the diagram demonstrates that the core signal encompasses the business cycle originating from traditional services, including trade, hotels, food and transportation (G-I). In addition, the core contains the industry in the narrow (C) and broad (B-E) sense and R&D related services (M-N). Tourism-related arts, recreation, entertainment and other services (R-U) demonstrate mixed results. However, since the statistical institutions often use this segment as a recycle bin in balancing the GDP definition by production, expenditure and income approaches, the R-U sector requires careful treatment. The large proportion (14 out of 27) of R-U going to the dropout pool supports this view.
The lower left part of Figure (ref) shows the regional contributions to the European business cycle, where the color corresponds to the number (count) of different economic activities of the same country. We argue that the higher number of sectors in the synchronous cluster demonstrates a stronger core signal, hence, increases the contribution of such economies in the synchronous group. Supplemented by the numbers in Table (ref), the observed results suggest that the core signal originates from the main contributors to the EU budget: France, Sweden, Netherlands, Germany, Austria, Italy, Belgium and Spain, followed closely by small and open Estonia's economy. In this context, Estonia's business cycle is a barometer of the European business cycle synchronization with the core countries. On the other end is the broad spectrum of periphery countries, with no contribution from Greece and Luxembourg that are the central players in the asynchronous cluster, while Slovakia adding the most to the dropout pool.
On the other hand, the analysis of the asynchronous group and the dropout pool in Table (ref) equips investors with a set of economic activities and regions least synchronized with the underlying European economic cycle. Both groups contain all business cycles for agriculture, forestry and fishing (A), and the significant part of public services (O-Q) that also experience the lowest fluctuations (see, Table (ref)). The groups include financial and insurance (K) and information and telecommunication (J) services with some additional impact stemming from construction (F). On the regional profile, the less synchronous group are driven by the EU outliers: Luxembourg with the core activities and Greece with the asynchronous cluster-specific activities. The dropout pool is more aligned with the periphery composition, including heavily hit by sovereign debt crisis Cyprus, Portugal, Ireland, Finland, Spain, and Malta, and nearly all data for Slovakia with, however, less evident reasons for the inclusion. In sum, the asynchronous pool is a mixture of safe-haven activities and regions with their opposite -- regions and sectors that are not recommended for investments during recessions.
The results in Figure (ref) support our previous statement that France, Italy, Austria, Germany, The Netherlands, Sweden, and Spain form the regional core of the cluster, and Estonia serves as a barometer. The remaining European countries show weaker synchronization both in counts and in proximity measures, with Malta likely being the least synchronous with the rest of the countries and some countries unable to compute due to zero count of business cycles in the synchronous group.
The core-periphery results presented in the previous section closely align with the observations from the related literature. We find that the approach taken in this paper adds a level of granularity, that helps defining the underlying synchronization structures. In this section we elaborate further on the resulting evidence, confirming certain known results. Furthermore, we distinguish and discuss the arising disparities from the perspective of the sectoral composition of the core.
First, we find evidence that France, together with Spain and Germany, forms the core cycle. These results are similar to belo2001some. France is seen as the leader of the common euro area business cycles, e.g., by aguiar2011oil and kapounek2019historical, among others. However, regarding Germany, the results in the literature are more mixed, with ahlborn2018core, kapounek2019historical, bunyan2020fiscal, celovcomunale2021 observing the decoupling of the German business cycles from the EU area. Our results suggest that this is not the case when we consider the main open economic activities, including broad industry (B-E), traditional services (G-I), IT and telecommunication (J), and R&D (M-N). Hence, we conclude that soft-clustering helps strengthen Germany's core signal by removing other “noisy” economic activities.
Second, we find that Greece is not included in the cluster, suggesting evidence of weak synchronization with the growth cycles of the core. Such observation results are consistent with the literature (see, e.g., bunyan2020fiscal, campos2016core). Similarly to campos2016core, we find that Ireland, Portugal, and Greece, which are argued to form a smaller periphery by the authors, are also weakly present in the core cluster. The only inconsistency is that the authors suggest that Spain should also belong to the smaller periphery with the aforementioned regions, which is contradicted by our soft-clustering view, while missing Cyprus and, especially, Malta. At the same time, Luxembourg economy is another, wealthier end outlier with business cycles not in the core and not widely discussed in the literature.
As expected by landesmann2003ceecs, MONFORT2013689, bandres2017regional, a clear differentiation between regions exists, where old EU member states may diverge from the economies joined after 2004 and experiencing fast catching-up/overheating episodes, for instance, for the Baltic States, and very small representation in the core of Greece, Bulgaria, Romania, and Visegrád Four.
To sum up, the sectoral-regional look at the development of European economies in the recent decades opens up prospects for analysing a broad range of problems. From a macroeconomic point of view, it is crucial to understand which economic activities and regions drive the European business cycle and which lag in responding to the shocks or even move in the opposite direction. In this regard, our paper confirms the existing core-periphery findings, keeping the main contributors to the EU budget in the synchronous core, of which all but Sweden are the members of the euro area. The core regions are strongly supported by open-to-trade industrial production, trade, food, accommodation and transportation services, and R&D economic activities -- a hidden aspect when just looking at the regional level.
Having a stable synchronous core is a necessary condition for successful OCA development. However, we find that not all euro area members enter the synchronous cluster. The latest adopters of the euro, small and open economies of the Baltic states, are improving in their catch-up with the synchronous core in recent decades. Furthermore, Central European entrants, Greece, Cyprus, Malta, Ireland, and Luxembourg contribute to the asynchronous cluster and the dropout data pool. These findings support the previously observed core-periphery regional division by adding agriculture, public sector, IT, finance and insurance, real estate, and other services to the asynchronous group of economic activities and the dropout pool.
From an econometric point of view, we argue the importance of tidying the synchronous core signal. Definitely, both the wavelets used for the trend-cycle decomposition and the new soft-clustering approach proposed in the paper have opportunities for further development.
For a more thorough analysis and as a robustness check, the interested reader may consider alternative trend-cycle decomposition methods, for instance, RePEc:fip:fedcwp:9906 for the Christiano-Fitzgerald filter, phillips2019boosting for a boosted HP filter and LEE2020 for sparse HP filter, among many others. celovcomunale2021 provides a comprehensive Monte-Carlo simulation experiment comparing various trend-cycle decomposition methods, including the wavelet approach. In the earlier versions of the paper, alongside the MODWT approach, we also worked with continuous wavelets, Singular Spectrum Analysis (SSA) and Complete Ensemble Empirical Mode Decomposition with Adaptive Noise (CEEMDAN, 5947265). These methods are decent alternatives to the wavelet approach due to their flexibility and a good fit for the frequency domain analysis. However, the focus of the current version of the paper is on improving the soft-clustering approach, leaving the aforementioned alternative approaches for future work. Furthermore, when considering further possible avenues for wavelet application, we find promising results by considering half-peak/trough reflection in order to better gauge the end-of-sample problems.
The soft-clustering approach is already pretty flexible to adjust for tidying the cluster of interest in a data-rich environment. The method employs bootstrapped sampling, probabilistic thresholding and silhouette scores for cleaning the cluster and removing any unstable, in the chosen dissimilarity metrics sense, observations. As the results depend on the dissimilarity metric, we conclude that its choice must stem from the problem analyzed. In this sense, business cycle synchronization perfectly matches the use of synchronicity measures.
In the paper, we attack the stability aspect from several angles. Besides the traditional choice of the number of clusters, bootstrapped sampling facilitates reshuffling the data and highlights the elements in the stable core, together with the boundary cases. These boundary data points tend to migrate from sample to sample, reducing the probability of sitting together with all other data in the same cluster. In this regard, probabilistic thresholding is the crucial clean-up procedure that improves the visibility of the synchronous core and highlights sectors and regions that must come first when focusing on the European economic policy. For future work, it would be interesting to delve deeper into the estimation of optimal probabilistic thresholding parameter values. Although the soft-clustering method belongs to the realm of unsupervised learning approaches, we deem that for a particular problem (convergence, synchronization), it is possible to find the problem-specific objective function. Solving the problem under particular (e.g., sparsity-synchronicity) trade-off restrictions, we could find the optimal combination of the parameters. Finally, we can scrutinize the dynamic composition of the synchronous core by applying, for instance, the rolling-window approach.
This research has received funding from the European Social Fund (project No. 09.3.3-LMT-K-712-01-123) under a grant agreement with the Research Council of Lithuania (LMTLT).
The authors would like to thank Svatopluk Kapounek, Jesus Crespo Cuaresma and all participants of “Euro4Europe” workshops for comments and suggestions.
}