Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
113,976 characters · 18 sections · 95 citation commands
Large datasets for the euro area and its member countries and the dynamic effects of the common monetary policy
\setstretch{1.5}
\footnotetext{Department of Economics, University of Bologna, Italy.}
\footnotetext{Department of Economics, Management and Quantitative Methods, University of Milan, Italy.\\ All authors gratefully acknowledge financial support from the Italian Ministry of Education, University and Research (PRIN 2020, Grant 2020N9YFFE_003).}
The recent and ongoing \enquote{data revolution} has enabled researchers to collect and process large amounts of information, thereby broadening the scope of macroeconomic research. The availability of high-dimensional datasets—where the number of series, $N$, can be comparable to or even exceed the time dimension, $T$—allows researchers to exploit the rich information contained in large sets of macroeconomic and financial indicators. This, in turn, facilitates addressing relevant policy questions using econometric methods specifically designed to handle, and in some cases benefit from, such large datasets.
In this paper, we present and document a new large-scale dataset for macroeconomic analysis, denoted as EA-MD-QD. The dataset comprises time series for both the euro area (EA) as a whole and its ten largest member countries, providing an accessible and comprehensive resource for policy analysis of EA-wide outcomes. It also enables comparisons across data vintages and empirical studies. Overall, EA-MD-QD includes $1136$ time series, at either monthly or quarterly frequency depending on data availability. The first available vintage spans the period January 2000-January 2024. After that date the data has been and is updated every month. The countries covered are Austria, Belgium, France, Germany, Greece, Ireland, Italy, the Netherlands, Portugal, and Spain. All the series are retrieved from institutional sources, namely Eurostat, ECB, OECD and FRED. The EA-MD-QD is made publicly available online and updated every month and come jointly with Matlab and Python codes for data transformation, missing-value imputation, and the treatment of outliers and the COVID period.\footnote{\url{https://zenodo.org/doi/10.5281/zenodo.10514667}}
Although the cost of data collection has substantially decreased in recent years, several sunk costs still hinder data availability and, consequently, the reproducibility of economic research. First, the manual updating of time series becomes increasingly burdensome as the dimensionality of the dataset grows, highlighting the need for systematic procedures to facilitate data acquisition. Second, once collected, data are often not immediately suitable for analysis due to issues such as seasonality, non-stationarity, and missing values, which require additional pre-processing. In this work, we aim to reduce, if not entirely eliminate, these costs by providing a dataset that is periodically updated and accompanied by codes enabling users to generate, within seconds, a research-ready dataset.
Given that data are updated on a monthly basis, then real-time monthly vintages are available. Specifically, updates occur at the end of each calendar month, either on the last working day or on the day when all updates scheduled for that month by the data providers have been released. Making vintages publicly available serves two main purposes. First, it enables users to account for the effects of data revisions across different periods, which have been shown to matter in many macroeconomic applications orphanides2002unreliability. Second, it fosters the reproducibility of research findings, as users can identify and retrieve the exact data vintage employed in a given analysis.
Building on the seminal contribution of sw1996, several large-scale datasets have been developed to foster macroeconomic research. Among the most prominent examples are FRED-MD and FRED-QD by fredMD,fredQD, two publicly available collections of monthly and quarterly macroeconomic and financial time series for the United States. Similarly, stevanovic constructed a large macroeconomic dataset for Canada.
To the best of our knowledge, this is the first attempt to provide a comprehensive dataset for macroeconomic research that encompasses both the EA as a whole and its largest member countries. The closest reference to our work using EA data is the real-time database developed by giannone2012area, which relies on the information contained in the ECB's Monthly Bulletins, i.e., the rawest information set available to policymakers on the data of the ECB governing council. Our dataset differs from theirs along several dimensions. First, the relatively short amount of vintages of our data prevents a fully real-time analysis. Second, both the timing of data collection and the information content of our dataset do not correspond to any specific policy meeting or information set available to policymakers. Instead, we provide a snapshot of macroeconomic conditions in the EA and its largest member countries at a given point in time. Hence, the purpose and scope of the two datasets is different. Finally, unlike giannone2012area, our dataset also covers the largest EA countries, with the aim of offering comparable information at both the national and EA aggregate levels. Moreover, we also provide practitioners with all the codes needed to make the dataset readily usable for empirical macroeconomic research.
As a second contribution, we employ EA-MD-QD to study the Impulse Response Functions (IRF) of a common monetary policy shock in the EA and, particularly, across individual EA member countries. The availability of country-level data enables policymakers to account for cross-country heterogeneity, which may ultimately influence both the effectiveness and the transmission of the common monetary policy. In this context, understanding whether, and to what extent, heterogeneous dynamics arise across countries is essential for designing and implementing more effective EA-level policies.
To this end, we employ the recently proposed Common Component Vector Autoregression (CC-VAR) methodology by forni2020common, which relies on the natural assumption of an underlying factor structure in the data. In a first step, standard Principal Component Analysis (PCA) is employed to extract the common factors and compute the common components of all variables. In a second step, a VAR model is estimated on some selected common components of interest along with other observables. Intuitively, this allows to study IRFs to common shocks, as the EA monetary policy one, without any contamination from idiosyncratic dynamics. More precisely, provided we include in the VAR step at least as many variables as common factors, we can fully identify the space spanned by the common shocks, by focussing on their common components only. This implies that the IRFs to common shocks for the variables considered in the VAR step do not change, even if some variables are substituted by others. Such an invariance property motivates our preference for the CC-VAR over more classical alternatives as, e.g., a standard VAR or a FAVAR bernanke2005measuring.
In our baseline specification we identify an EA-wide common monetary policy shock using monthly data and the Instrumental Variables (IV) approach proposed by stock2012disentangling and mertens2013dynamic, in combination with the High-Frequency Identification (HFI) strategy of gertler2015monetary. We test several instruments from the pool provided by altavilla2019measuring and select the 1-year overnight index swap (OIS), which is associated with the highest first-stage $F$-statistic.
We find four key results. First, looking at the share of variance explained by the common factors, we uncover non-negligible cross-country heterogeneity, particularly in unemployment and interest rate dynamics, whereas industrial production, stock prices, and prices display a higher degree of homogeneity. This findings provide a preliminary assessment of the degree of comovement among EA countries across different economic dimensions.
Second, the IRFs are consistent with standard economic theory. In particular, we observe a decline in the IRFs of prices in all EA economies following a contractionary monetary policy shock--although with different magnitudes, while at the EA level their magnitude aligns with other results in the literature jarocinski2020deconstructing.
Third, we find evidence of a moderate yet meaningful heterogeneity in impulse responses across countries for several key variables. In particular, our estimates point to a core-periphery pattern in price and interest rate dynamics. For prices, core countries’ responses are on average aligned with--or even more pronounced than--the EA aggregate, whereas peripheral countries appear less affected. Symmetrically, interest rates responses in peripheral countries display stronger medium-term effects on average. No clear pattern emerges for real indicators, though some countries--such as Italy and Greece--show notable deviations from the aggregate. By contrast, stock prices exhibit a relatively homogeneous behavior across countries. On average, Germany’s responses closely mirror the EA aggregate, while Greece shows the most pronounced deviations.
Finally, to explore potential drivers of this heterogeneity, we examine the correlations between the peak country-level IRFs and selected structural characteristics related to labor markets, households, firms, and national economic structures corsetti2022one. While no causal inference is implied, the evidence suggests that monetary policy transmission is strongly related to several economic channels--including homeownership rates and savings behavior--that either dampen or amplify the effects of common shocks across countries.
Over the years, a large body of literature has studied the effects of common monetary policy in the EA jarocinski2020deconstructing, andrade2021delphic. In this paper, we also analyse potential asymmetries in the transmission mechanism of monetary policy shocks across EA member countries. In particular, by employing a novel large-dimensional dataset, more recent data vintages, and a novel econometric framework, we extend and update the previous studies by barigozzi2014euro, georgiadis2015examining, burriel2018uncovering, corsetti2022one, mandler2022heterogeneity.
Despite differences in sample coverage, our results broadly align with those of barigozzi2014euro, who document similar cross-country patterns in prices and unemployment rates, partly driven by a core–periphery dynamic within the EA. Differently from our results, corsetti2022one report price increases for several countries in response to monetary tightening. Nevertheless, despite this difference in sign, the relative cross-country patterns we identify are broadly consistent with their findings. Conversely, mandler2022heterogeneity obtain an inverted cross-country ranking in price responses compared with our results.
Section (ref) describes the EA-MD-QD dataset. Section (ref) introduces the baseline specification adopted for the empirical application. Section (ref) discusses the factor analysis, while Section (ref) examines the comovements across variables and countries explained by the factors. Section (ref) describes the estimation and identification of the EA and country-specific IRFs to a common monetary policy shock. Section (ref) reports the estimated IRFs, as well as the analysis of cross-country heterogeneity and its potential drivers. Section (ref) concludes. In a Supplementary Appendix we provide a detailed description of the data, a step-by-step guide to data preparation and estimation of the comovements, and additional empirical results, based on different data and identification approaches, showing the robustness of our results.
The dataset is constructed according to three guiding principles: (i) accessibility: all series are sourced from public, institutional databases; (ii) timeliness: the dataset is fully web-scraped, enabling monthly updates of all series in the panel according to their respective release calendars; and (iii) coverage: it includes variables representing the most relevant sources for modern macro-financial analysis. Importantly, the monthly updates provide practitioners with a new vintage of the dataset each month, incorporating all data releases and revisions to previously published series, if any. All monthly vintages are available from the initial release of EA-MD-QD in January 2024.\\ Euro area. The dataset for the EA comprises $N = 118$ series, of which $N_M = 47$ are monthly and $N_Q = 71$ are quarterly. All series span from 2000:M1 (2000:Q1) to the most recent available observation. The selection of variables follows established datasets for the US fredMD,fredQD and covers: (1) National Accounts, (2) Labor Market Indicators, (3) Credit Aggregates, (4) Labor Costs, (5) Exchange Rates, (6) Financial Markets, (7) Industrial Production and Turnover, (8) Prices, (9) Confidence Indicators, and (10) Monetary Aggregates. Among the 118 series, 104 are sourced from Eurostat, 10 from the ECB Data Warehouse, 3 from the OECD Statistical Database, and 1 from the FRED portal.
Countries. In addition to the EA dataset, we provide datasets for the ten largest EA countries: Austria, Belgium, France, Germany, Greece, Ireland, Italy, the Netherlands, Portugal, and Spain. Each country-specific dataset is constructed according to the same principles used for the EA data, with a few exceptions. For instance, variables related to Monetary Aggregates are not included at the country level, as they cannot be straightforwardly attributed to individual countries. Other groups of variables (e.g., Credit Aggregates or Industrial Production and Turnover) are smaller for some countries, such as Ireland, due to data unavailability. The country specific variables are sourced from the same providers as those of the EA dataset.
Table (ref) provides a summary of the series included in the dataset, along with a brief description, their frequency (column F), and provider (column P). For each series, the last columns of Table (ref) indicate the countries for which the series is available. A checkmark denotes availability, while a dash indicates that the series is not available for that country. A more detailed description of the series, with additional informations, is in Table A2 in Appendix A.
Table (ref) provides a detailed breakdown of the numerosity of series by country, organized according to the ten macro-categories included in the dataset. Broadly, each category is similarly represented across countries, and the frequency of individual series is generally uniform across countries. Specifically, for both the EA and each individual country, approximately 60% of the series are quarterly, while the remaining 40% are monthly. There are two notable exceptions: indicators of Industrial Production and Turnover are unavailable for Ireland, and Producer Prices are missing for both Ireland and Portugal.
Due to their mixed-frequency nature, the dataset is inherently unbalanced, with quarterly series recorded in the first month of each corresponding quarter and missing values in the remaining months. The provided codes allow users to handle this unbalanced structure. Practitioners can choose to: (i) retain the mixed-frequency format; (ii) subset the dataset to include only monthly or only quarterly variables; or (iii) aggregate the monthly data to the quarterly frequency. In the latter case, monthly series are aggregated to the quarterly level by summing the values of monthly flow variables and taking the mean of monthly stock variables.\footnote{Alternatively, one could take the value from the last month of the reference quarter as the quarterly observation. However, this approach would overlook intra-quarter dynamics that may be informative. For instance, if 2020:Q2 were represented solely by June 2020 data, much of the COVID-19 shock observed in April 2020 would be disregarded.}
Besides its mixed frequency nature, the dataset is unbalanced also because of both data availability and ragged edges arising from the asynchronous release of different series. Regarding data availability, while most series begin at or before 2000:Q1 (2000:M1), some start few periods later (see Appendix A for details). As for release timing, the series are updated at different intervals across categories: some variables (e.g., National Accounts) are published roughly one to two months after the end of the reference quarter, while others (e.g., Credit Aggregates) may be released up to four months later. The companion codes allow practitioners to impute these missing values by means of two different procedures described in Section (ref).
Finally, while most series are already seasonally adjusted at the source, some are only available in raw form. For these, we retrieve the unadjusted series and apply seasonal adjustment ex post. In particular, we employ standard seasonal filtering methods findley1998new for monthly variables, while quarterly variables are adjusted using a simple dummy-variable approach.\footnote{It is well known that filtering techniques can suffer from end-of-sample issues, requiring many observations to obtain reliable estimates of the trend component. Given the limited time span at the quarterly frequency, end-of-period estimates for seasonally adjusted quarterly variables obtained with these filters are therefore unreliable.} Only Credit Aggregates and Producer Prices are not seasonally adjusted at the source, accounting on average for about 30% of the total series. Starting from the October 2025 release, we also provide users with raw data, where these series remain unadjusted.
Being composed of variables that capture a wide range of economic and financial aggregates, the EA-MD-QD includes both stationary and non-stationary series. Although recent advances in econometric methods allow for the presence of non-stationarity in large-dimensional settings barigozzi_large-dimensional_2021, many empirical applications still require stationary data. For this reason, together with the raw data, we provide users with transformation codes designed to achieve stationarity.
We consider three possible sets of transformations. First, statistical transformations which remove any $I(1)$ or $I(2)$ according to standard unit root tests dickey1979distribution,phillips1988testing. Overall, both across variable categories and countries, only a small number of series exhibit $I(2)$ dynamics. Most of these belong to the group of Credit Aggregates, distributed relatively evenly across countries, with a few notable exceptions related to Unit Labor Costs and Labor Market Indicators. Conversely, the majority of the series in the panel are $I(1)$. A smaller subset of variables is instead $I(0)$, primarily concentrated among confidence indicators and, for some countries, national accounts variables. The transformations are applied uniformly across countries, with only minor deviations reflecting country-specific idiosyncrasies that generate dynamics differing from those of the corresponding EA series. We refer to Table A2 in Appendix A for details.
Second, we consider the statistical set of transformations described above with the exception of interest rates which are kept in levels. This choice is in line with the literature on EA monetary policy transmission which is also the focus of the second part of the paper corsetti2022one. Third, we consider the statistical set of transformations described above with the exception of interest rates and unemployment rates in order to keep all rates in levels.
The companion codes allow the users either to keep the data in levels or to apply either of the three set of transformations described above.
As discussed above, regardless of the dataset frequency, missing values remain due to ragged edges arising from data availability, the asynchronous timing of data releases, as well as from removals of outliers. Moreover, the transformations applied to the series, as described in Section (ref), introduce additional missing values because of observation losses due to differencing. While the untreated data are always provided, the companion codes to the dataset allow practitioners to choose among two different strategies for imputing missing values.
Specifically, missing values can be imputed using either the Expectation Maximization (EM) algorithm proposed by stock2002macroeconomic and also employed by fredMD, or the one by banbura2014maximum. The former represents the most straightforward approach commonly used in the literature for imputing both outliers and missing values, though not necessarily the most sophisticated. The latter is particularly suitable when the data exhibit a factor structure with autocorrelated factors.
Following fredQD, an observation is considered an outlier if it deviates from the sample median by more than ten interquartile ranges. Once outliers are detected and removed we can impute the corresponding missing values by either of the methods described in Section (ref). This applies to all sample except the Covid period.
Unlike \enquote{standard} outliers, the COVID shock is pervasive, affecting most series in the panel--particularly those representing the real economy--with substantial heterogeneity even within these series. Moreover, it is unclear how much of the COVID shock should be attributed to economic versus exogenous forces ng2021modeling. In light of these considerations, treating the COVID period as composed of sporadic outliers would result in the loss of potentially important economic information relevant for understanding the current economic stance.
In the companion codes, we provide users with two alternatives related to the treatment of the COVID period. In the first option, data for variables representing the real side of the economy, as, e.g., Industrial Production, are treated as missing in 2020-2021 and imputed via the Kalman smoother using information from financial and nominal variables, as, e.g., the Stock Price Index and Prices. Alternatively, users can retain the transformed data without any specific treatment for COVID.
As discussed above the user has many possible choices about the data to be analysed and their treatment treating. Clearly, the specific choices depend on the aim of the empirical analysis. In the rest of the paper, we exploit the EA-MD-QD database to study the transmission of EA monetary policy across countries, with the goal of highlighting the advantages the database offers for structural macroeconomic analysis. To this end, hereafter, we adopt the following baseline specification.
Specifically, we construct a balanced panel $\mathbf{x} = \{x_{it}, i=1,\ldots,N ; t=1,\ldots,T\}$ consisting of 47 EA-wide monthly variables and 200 national series, all transformed as described above, for a total of $N = 247$ time series, and covering the period 2002:M1-2023:M10, corresponding to $T=263$ monthly observations.
These choices are based on a series of practical considerations. First, regarding the chosen sample, we begin the analysis in 2002:M1, following the literature on the identification of EA monetary policy shocks via instrumental variables (IV) altavilla2019measuring, andrade2021delphic, since liquidity in the Overnight Index Swap (OIS) market was limited before then, hindering identification. The sample ends in 2023:M10, which is the latest date for which monetary policy surprises from altavilla2019measuring—used to identify the shocks—are currently available.
Second, when the goal is extracting common factors, as in our empirical analysis, it is well documented that adding more series to the dataset does not necessarily help in recovering the factors boivin2006more. Indeed, one should retain only those series that are most likely to be driven by the same common shocks. Not surprisingly these are the most aggregated series. Hence our choice of appending to the EA monthly dataset the following national variables: Industrial Production Indexes (IPMN, IPCAG, IPCOG, IPDCOG, IPNDCOG, IPING, IPNRG), Harmonized Indexes of Consumer Prices (HICPOV, HICPNEF, HICPG, HICPIN, HICPSV, HICPNG), Producer Price Indexes (PPICAG, PPICOG, PPINDCOG, PPIDCOG, PPIING, PPINRG), 10-years Interest Rates (LTIRT), Stock Price Indexes (SHIX), and Unemployment Rates (UNETOT).\footnote{Indeed, when including all the monthly variables in the dataset, the variance of the idiosyncratic components of the most relevant variables included in our analysis grows by 20% on average.}
Third, although for consistent estimation of the factors we do not require any constraint between the cross-sectional dimension $N$ and the sample size $T$, still it is advisable to have a panel where these quantities have a comparable value. As pointed out by onatski2010determining if $N$ is much larger than $T$ it is harder to recover consistently the number of factors by studying the behavior of sample eigenvalues. Since the considered time span is relatively short, then we prefer working with just a subset of the whole EA-MD-QD data.
A step-by-step description of the procedure adopted in this section and Sections (ref) and (ref) to prepare and analyze the data through a factor model is given in Appendix B.
We assume that the $N$-dimensional vector of observed data at time $t$, $\mathbf x_t$, follows a factor model:
where $\bm\Lambda$ is the $N\times r$ vector of loadings associated to the $r$-dimensional vector of zero-mean common factors $\mathbf{f}_t$, $\bm\xi_t$ is a $N\times 1$ vector of zero-mean idiosyncratic components, and $\bm\mu$ is a $N \times 1$ vector of constants, hence $\mathbb E[\mathbf x_t]=\bm \mu$.
Both the factors and the idiosyncratic components are allowed to be serially correlated. Moreover, when $N$ is large is reasonable to allow the idiosyncratic components to be also (weakly) cross-sectionally correlated, and, in this case, we say that (ref) is an approximate factor model. We refer to bai2003inferential for the formal assumptions.
The loadings and the factors in (ref) are not identified, since the factor model can equivalently be expressed with $\bm\Lambda \mathbf H$ as the loadings matrix and $\mathbf H^{-1}\mathbf f_t$ as the factors vector, for some invertible $r\times r$ matrix $\mathbf H$. To identify the factors additional assumptions would be needed bai2013principal, but since in this paper our interest is only in estimating the common component $\bm\chi_t=\bm\mu+ \bm\Lambda \mathbf f_t$, we do not explore this path further. Indeed, $\bm\chi_t$ is always identified once we determine the number of factors $r$, so that we can disentangle it from the idiosyncratic component.
The factors and loadings are estimated via the classical PCA approach. So we estimate the loadings, denoted as $\widehat{\bm\Lambda}$, as $\sqrt N$ times the $r$ normalized eigenvectors corresponding to the $r$-largest eigenvalues of the sample covariance matrix of the standardized data $\mathbf x_t$, i.e., of $T^{\,-1}\sum_{t=1}^T \widehat{\bm\Omega}^{-1/2}(\mathbf x_t-\widehat{\bm\mu})(\mathbf x_t-\widehat{\bm\mu})'\widehat{\bm\Omega}^{-1/2}$, where $\widehat{\bm\mu}$ and $\widehat{\bm\Omega}$ are, respectively, the $N$ dimensional vector of sample means and a diagonal $N\times N$ matrix with entries the $N$ sample variances of each element of $\mathbf x_t$. Then, the factors, $\widehat{\mathbf f}_t$, are obtained by projecting the estimated loadings onto the data, i.e., $\widehat{\mathbf f}_t=(\widehat{\bm\Lambda}'\widehat{\bm\Lambda})^{-1}\widehat{\bm\Lambda}'\widehat{\bm\Omega}^{-1/2}(\mathbf x_t-\widehat{\bm\mu})= N^{-1}\widehat{\bm\Lambda}'\widehat{\bm\Omega}^{-1/2}(\mathbf x_t-\widehat{\bm\mu})$ (due normalization of the eigenvectors). The estimated common components are then de-standardized and de-centered by multiplying them by the standard deviation of the original series and adding the corresponding mean. The resulting common component $N$-dimensional vector is denoted as $\widehat{\bm\chi}_t=\widehat{\bm\mu}+\widehat{\bm\Omega}^{1/2}\widehat{\bm\Lambda}\widehat{\mathbf f}_t$.
From the results in stock2002forecasting and bai2003inferential, it immediately follows that $\widehat{\bm\chi}_t$ is a consistent estimator of $\bm\chi_t$, as $N,T\to\infty$. This in practice shows the necessity of working with a high-dimensional panel in order to consistently disentangle the common components capturing all main comovements from the idiosyncratic ones.
In Table (ref) we report the number of common factors obtained by employing the following standard methods: (i) the log-information criterion (IC2) of bai2002determining, implemented also when (ii) tuning the penalty as suggested by alessi2010improved, (iii) the eigenvalue-ratio criterion by ahn2013eigenvalue, and (iv) the test by onatski2010determining.\footnote{For those methods requiring it, we set the maximum number of factors to $r_{\max}=15$.} Hereafter, we set $r= 6$.
For a given variable, the share of total variance explained by the common component offers a straightforward measure of cross-country comovement. Specifically, for a given country and variable, a high proportion of variance explained by the common component indicates that the variable is primarily driven by EA-wide common factors rather than by country-specific idiosyncratic dynamics. To provide an intuition of the explanatory power of the common factors, Table (ref) reports the share of variance explained by the common component for selected key variables. This analysis offers a preliminary assessment of the degree of synchronization across EA countries. Indeed, if country dynamics were perfectly aligned, we would expect relatively similar levels of comovement across countries for each variable. Confidence intervals for the explained variance are obtained using 1000 replications of the bootstrap procedure by barigozzi2018simultaneous and described in Appendix B.
Comovement across variables is relatively high at the EA level. At the country level, however, some heterogeneity emerges across indicators. For industrial production, Belgium, Greece, the Netherlands, and, to a lesser extent, Portugal, deviate from the high commonality observed at the EA level. Hence, larger industrial economies are more closely aligned with the EA, whereas smaller economies exhibit a higher degree of idiosyncratic variation in production. In contrast, prices display a relatively more homogeneous pattern across countries, but the average share of variance explained is lower than for industrial production. Indeed, nominal variables have been shown to comove less than real variables in standard large macroeconomic datasets ahn2025common,lissona2025heterogeneous. The degree of commonality for interest rates exhibits a clear core-periphery pattern: Northern European countries display a high level of commonality, whereas interest rate dynamics in Southern countries appear more driven by country-specific factors. However, even within this group, the extent of idiosyncrasy varies considerably: Italy and Spain comove more strongly with the EA, while Greece is almost entirely idiosyncratic. Stock prices show a relatively high and homogeneous degree of commonality across countries, with the exceptions of Greece and Portugal.
The variable exhibiting the highest degree of heterogeneity is the unemployment rate. Recall that this variable is taken in first difference under our baseline specification. As such its degree of comovement is not very large, but still there are considerable differences across countries and these are unaffected by the chosen transformation.\footnote{If the unemployment rate were taken in levels, then its high-persistence would result in an anomalously large explained variance of its common component. The choice of taking first differences is precisely made to avoid such “overfitting” phenomenon.}
Overall, these results provide preliminary insights into the degree of heterogeneity across variables and EA countries. While no structural claims can be made at this stage, they motivate a deeper analysis to determine whether this heterogeneity persists conditional on a monetary policy shock (see Section (ref)).
Finally, we acknowledge that a more rigorous analysis of comovements across EA variables would require explicitly accounting for lag and lead relationships both across variables and countries, as in d2016nowcasting and cascaldi2024back. However, given the large dimensionality of our data--both within and across countries--such an approach would require modifications of the employed methodology which are beyond the scope of this paper.
It is widely acknowledged that by using large datasets we can retrieve structural shocks via factor analysis GiannoneReichlin2006,forni2009opening,Forni2014. In this spirit, here we adopt the CC-VAR by forni2020common in order to estimate and identify the IRFs to the EA monetary policy shock. This approach simply consists in fitting a VAR on a vector $\mathbf{Y}_t$ of $n$ endogenous variables with $n \ll N$, and containing $n^*$ estimated common components of selected variables, $\widehat{\bm\chi}_t$, with $r\le n^*\le n$, along with any possible additional observable of interest.
The CC-VAR has three main advantages. First, since the common components are always identified, identification of the shocks can be achieved by means of any existing method borrowed by the traditional structural VAR literature. Second, if, in presence of $r$ latent factors, we include $n^* = r$ common components in the VAR, then the space spanned by the structural shocks driving all $N$ common components coincides with the space spanned by the reduced form shocks, i.e., the VAR innovations. This is because all common components are generated by the same underlying shocks through the factors, which are common to all $N$ considered variables. It follows that, third, when considering a VAR for the common components and substituting one of these with another, the IRFs of the retained variables do not change, i.e., they are invariant with respect to the choice of the other variables included in the model.
Given the above described properties and our task of identifying the common EA monetary policy shock, the CC-VAR then seems to be a more suitable choice than the FAVAR bernanke2005measuring. Indeed, while by augmenting a classical VAR with latent factors, the FAVAR allows us to recover the space spanned by structural shocks, this space is also contaminated by idiosyncratic dynamics. Hence, we cannot guarantee the invariance of the IRFs which characterizes the CC-VAR. This is because in the FAVAR we use the observed variables and not their common components.\footnote{Note that if we augmented the CC-VAR with factors, as for example in the original application by forni2020common, the invariance of the IRFs would obviously still be preserved, and we could think of it as a FAVAR for common components.} Moreover, since the factors are not identified (see Section (ref)), including them in the VAR makes identification less straightforward. Clearly, a classical VAR is also inferior to the CC-VAR since, not only it lacks the invariance property of IRFs, but also it does not guarantee that we can recover the space of the structural shocks alessi2011non.
Our specification of the vector $\mathbf Y_t$ is given in Table (ref). First, we include the observed EA 2-years Interest Rate $R_t$, which is also our policy rate (see Section (ref) for its motivation). Then, we include the common component for five key EA variables: the Industrial Production growth in manufacturing (IPMN), the Overall Harmonized Consumer Price Index inflation (HICPOV), the 10-years Interest Rate (LTIRT), the Stock Price Index growth (SHIX), and the monthly change in the Unemployment Rate (UNETOT). Finally, we include, one at a time, the common components of various national variables of interest, for which we aim to study the IRF. In particular, we consider a total of 49 different national variables resulting in 49 possible choices for $\mathbf Y_t$.\footnote{The five key variables (IPMN, HICPOV, LTIRT, SHIX, UNETOT) for the ten countries covered by the EA-MD-QD, and recalling that Industrial Production for Ireland is unavailable.} For any of those choices, $\mathbf Y_t$ has always dimension $n=7$ and $n^* =n-1=6$ so that $n^*=r$, and, as a consequence, the IRFs for the first six variables in $\mathbf Y_t$ are unchanged in all 49 VARs (see the results in Section (ref)).
Specifically, for any choice of $\mathbf Y_t$, we estimate the following reduced-form VAR:
where $\mathbf c$ is a $n \times 1$ vector of reduced-form constants, $\mathbf B_i$, $i=1,\ldots p$, are $n \times n$ matrices of reduced-form coefficients and $\mathbf u_t$ is the $n \times 1$ zero-mean vector of reduced-form errors, with zero mean and covariance matrix $\mathbf \Sigma$.
The scaling factor $\sigma_t$ is included to address the significant fluctuations observed during the Covid period by adjusting the model residual volatility, as suggested by lenza2022estimate. In particular, $\sigma_t$, takes a value of 1 for all periods preceding the Covid onset period, while from 2020:M3 until the end of the sample, at $T=$ 2023:M10, $\sigma_t$ is estimated in each period by maximum likelihood.\footnote{Our approach differs slightly from the approach of lenza2022estimate as their approach is based on US data. Nevertheless, the variations introduced do not significantly affect our results. Moreover, adjusting the IRFs for Covid introduces slight deviations from this invariance due to changes in the volatility parameter $\sigma_t$, leading to unwanted dispersion in the EA IRFs. To address this and preserve the invariance property of the CC-VAR, we first estimate $\sigma_t$ for each specification differing only in the last variable. We then take the median of the estimated $\sigma_t$ vectors across specifications to obtain an average volatility, which is subsequently used to compute the national IRFs as described.}
Estimation of the VAR in (ref) gives the estimated coefficients $\widehat{\mathbf B}_i$ and by VAR inversion we obtain the estimated reduced form IRFs. Throughout, we choose $p=8$ lags in the VAR. Consistency, as $N,T\to\infty$, is proved in forni2020common (see also forni2009opening, and bai2006confidence, for similar results).
The structural counterpart of ((ref)) is:
where the reduced-form errors $\mathbf u_t$ are related to the structural errors $\bm \varepsilon_t$ through the following relationship:
Hereafter, let $\mathbf S \equiv \mathbf A_0^{-1}$ for simplicity of notation.
As is well known in the VAR literature, the matrix $\mathbf S$ is unobserved and we need an identification strategy to estimate it and give an economic interpretation to the elements of $\bm{\varepsilon}_t$. For our application, focused on the monetary policy shock only, it is sufficient to identify the associated elements in the column of the matrix $\mathbf S$. We denote such column as $\mathbf s$ and the corresponding monetary policy shock as $\varepsilon_t^p$, while all other structural shocks are collected into the vector $\bm\varepsilon_t^q$. The entry of $\mathbf s$ corresponding to $\varepsilon_t^p$, which is the contemporaneous impact of the monetary policy shock on the policy rate, is denoted as $s^p$, while all other entries are collected into the vector $\mathbf s^q$. Once the EA wide monetary policy shock is identified the corresponding IRFs are then computed from the, truncated, VMA($\infty$) representation of the estimated structural VAR defined in ((ref)).
To identify $\mathbf s$ and $\varepsilon_t^p$, we employ the high-frequency Proxy-SVAR method from gertler2015monetary.\footnote{In Appendix H we consider also sign restrictions as an alternative identification strategy.} This approach requires jointly specifying two types of variables: a policy indicator and a monetary policy instrument. The policy indicator, denoted as $R_t$, is a variable capturing the central bank's monetary policy stance, it is included in $\mathbf Y_t$ and thus enters directly into the CC-VAR model. The monetary policy instrument, denoted as $Z_t$, must satisfying two characteristics. First, it must be correlated with the monetary policy shock:
Second, it must be exogenous to all other shocks, collected in the vector $ \bm \varepsilon_t^q$, i.e.,
For a given choice of $R_t$ and $Z_t$, after estimating the reduced-form VAR in ((ref)), and its vector of estimated reduced-form residuals $\widehat{\mathbf u}_t $, we identify $\mathbf s$ using a two-step procedure. First, we project the residual of the policy indicator $R_t$, which we denote as $\widehat{u}_t^p$, on the instrument $Z_t$. That is, we estimate the linear regression:
Then, letting $\widehat\alpha$ and $\widehat\beta$ be the OLS estimates of $\alpha$ and $\beta$ and letting $\widehat{\widehat{u}}_t^p=\widehat{\alpha}+\widehat{\beta} Z_t$, we estimate the linear regression,
where $\widehat{\mathbf u}^q_t $ contains all other reduced-form residuals. The OLS estimate of ${\bm\gamma}$, denoted as $\widehat{\bm\gamma}$, is then an estimate of ${\mathbf s^q}/{s^{p}} $. This identifies $\mathbf{s}$ up to a scaling factor. To fully identify $\mathbf{s}$, we normalize the impact of the monetary policy shock on the interest rate to one, i.e., we set $s^p=1$, so that $\widehat{\bm\gamma}=\widehat{\mathbf s}^q$. This results into an identified monetary policy shock which increases at impact the policy rate $R_t$ of one percentage point.
As in gertler2015monetary, we test different combinations of instrument-policy indicators. Specifically, our selection of policy indicator $R_t$ candidates includes EA the 1-, 2-, and 3-years rates interest rates downloaded from the Eurostat database, and our choice for the pool of instrument candidates is drawn from the work of altavilla2019measuring.\footnote{To convert high-frequency data into monthly frequency, daily observations are cumulated over the last 30 days and then averaged over the corresponding month, following gertler2015monetary. This procedure substantially mitigates the temporal aggregation bias kilian2024construct.} We then select the pair $(R_t,Z_t)$ that exhibits the highest $F$-statistic related to Equation (ref). The $F$-statistic is used to test the null hypothesis of instrument irrelevance (low $F$-statistic) against the alternative hypothesis of instrument relevance (high $F$-statistic), with a threshold of 10 commonly used to distinguish between strong and weak instruments stock2002survey. Based on this test we choose $R_t$ as the EA 2-years interest rate, while for $Z_t$ we choose the 1-year OIS. This choice gives a standard $F$-statistic equal to 16.6 and a robust $F$-statistic, computed following olea2013robust, equal to 11.3.\footnote{We obtained similar results using other instrument-policy variable combinations. Specifically, we also tested the 1-year OIS together with the 1-year interest rate ($F$-statistic 13.9), the 1-year OIS together with the 3-year interest rate ($F$-statistic 12.0) and the 2-year OIS together with the 2-year interest rate ($F$-statistic 12.8). All these specifications yielded similar results.}
The OIS price is measured immediately before and after both the ECB's press statement and press conference. The policy surprise measure is derived by summing these two price differences. The rationale is that, before the ECB policy announcements, the swap price incorporates market expectations regarding the future path of the OIS. Therefore, any adjustment in the price immediately after the policy announcements is interpreted as a recalibration of market expectations in response to an unforeseen monetary policy surprise. This satisfies the correlation requirement in (ref). Regarding the exogeneity condition in (ref), we assume that no other significant macroeconomic shock occurs in the time span between the press statement and the press conference. Additionally, we distinguish conventional monetary policy shocks from information shocks jarocinski2020deconstructing, andrade2021delphic by retaining only those high-frequency observations for the OIS price that are associated with a simultaneous movement of opposite sign in the swap price on the Euro Stoxx 50 index.
Finally, to quantify uncertainty around the estimated IRFs, we employ a standard Wild boostrap gonccalves2004bootstrapping with $1000$ bootstrap repetitions, also accounting for the well-known bias due to the estimation of the VAR in finite samples kilian1998small.
Figure (ref) presents the estimated IRFs of the EA variables included in the CC-SVAR in (ref) in response to a 100 basis-point (bps) shock to the 2-years interest rate, along with their one-standard-deviation, i.e., 68%, confidence intervals. As already highlighted in Section (ref), we notice that the same EA IRFs are obtained for any of the 49 VARs considered, each with a different national common component included. When comparing these results with classical VAR or FAVAR models (see Figures E1 and E2 in Appendix E), we see that neither of those approaches gives EA IRFs which are invariant with respect to the use of different national variables. Indeed, both for the VAR and the FAVAR we obtain 49 different IRFs for each of the six EA variables. This is what motivates our choice of the CC-VAR.
A contractionary monetary policy shock leads to an increase in both short-term and long-term interest rates, with the latter rising to roughly half the magnitude of the former at impact. This pattern is expected, as monetary policy shocks typically have a stronger effect on shorter maturities than on the upper end of the yield curve altavilla2019measuring. As anticipated, industrial production, prices, and the stock price in the EA decline. In contrast, the unemployment rate rises with a lag of a few months, although the response is not statistically significant.
The magnitudes of the responses of prices and the stock price are comparable to those reported by jarocinski2020deconstructing (after rescaling the interest rate increase to 100 basis points therein), but contrasts with corsetti2022one, who document a muted reaction of prices. This difference is likely to reflect our explicit control for so-called information shocks when constructing the instruments, consistently with the findings of andrade2021delphic and jarocinski2020deconstructing. The pointwise response of the unemployment rate is similar in magnitude to the one reported by corsetti2022one.
These results, and those in the next sections, are coherent with those obtained using alternative sets of transformations (Appendices F and G), monthly or quarterly data and sign restrictions (Appendices H and I), as well as when augmenting the monthly data considered so far with quarterly EA and national GDPs (Appendix J).
Figure (ref) presents the IRFs for the national variables. At each horizon, responses that are significant at the one-standard-deviation, i.e., 68%, confidence level are shown with solid lines, while non-significant responses are depicted with dotted lines. Individual responses with their confidence intervals can be found in Appendix C.
As with the EA IRFs, the direction of the IRFs for each country is consistent with standard economic theory. Moreover, the IRFs exhibit broadly similar dynamics across countries, with the main differences appearing in the magnitude of the response. For industrial production, the peak decline in the IRFs ranges from approximately 6% for Italy to around 2% for Greece, which stands out both in terms of magnitude and the shape of its response relative to other countries. These results highlight how the largest manufacturing economies in the EA are more strongly affected by the adverse effects of a contractionary shock. Price responses are particularly homogeneous across countries, with a slightly more muted reaction in Greece. The maximum difference between the IRFs is approximately 1.5 percentage points, indicating only minor misalignments across countries.
Similar patterns are observed for stock prices, although Greece exhibits a particularly strong response to the contractionary shock, with its stock price declining by roughly 20%.
On the interest rate side, Greece again stands out with an exceptionally strong response, peaking at around 120 bps. More generally, interest rates increase more in other peripheral countries (i.e., Portugal, Ireland, Italy, and Spain) compared to core countries. This core-periphery distinction is less clear for the unemployment rate. While Greece and Spain experience notably higher responses, Italy shows a counterintuitive, though mostly non-significant, effect. Overall, most unemployment responses are borderline significant at best.
These results suggest only moderate heterogeneity across countries, which is mostly concentrated in a few selected cases. However, this analysis alone does not allow us to fully assess this claim, as these observed differences may be entirely obscured by estimation uncertainty.
To quantify heterogeneity across EA countries, for each variable of interest we compute the difference between the national IRFs and the corresponding EA IRF. This procedure allows for a direct comparison across individual countries. Specifically, at each horizon, a positive (negative) value indicates that the national response lies above (below) the corresponding EA response. Importantly, this relative positioning does not necessarily imply a stronger or weaker response in absolute terms, as it depends on the sign and magnitude of both IRFs. For instance, a positive difference may arise either when both IRFs are positive and the national response is larger, or when both are negative and the national response is smaller. By repeating this procedure for each bootstrap repetition, we can also quantify the uncertainty around these point estimates.
Figure (ref) presents these differences along with one-standard-deviation, i.e., 68%, confidence intervals. For industrial production, France and Italy stand out from the other countries, as they are the only ones exhibiting a stronger downward response relative to the EA aggregate. This pattern differs for Germany, which, despite being one of the main manufacturing economies in the EA like Italy and France, shows no substantial deviation from the EA aggregate response. All other countries display weaker responses compared with the EA, particularly core economies such as Austria, Belgium, and the Netherlands.
In contrast, for prices, some differences emerge between core and peripheral countries. Core countries are either closely aligned with the EA aggregate (e.g., Austria) or exhibit a stronger response (e.g., Germany), whereas peripheral countries generally show a more muted reaction, with the exception of Spain, where the difference is not statistically significant. This pattern seems to suggests that the transmission of monetary policy to prices differs between core and peripheral countries barigozzi2014euro.
Stock prices present a relatively homogeneous picture, with only Germany, Italy, and Greece showing statistically significant deviations from the EA aggregate. However, the magnitude of these differences varies markedly: Germany and Italy deviate by at most approximately 2%, whereas Greece exhibits a substantially stronger response, with a difference exceeding 5% compared to the EA counterpart.
An interesting pattern emerges for interest rates. First, for nearly all countries, the peak differences relative to the EA occur between 6 and 8 months, coinciding with the horizon at which most responses are statistically significant. Second, there is a hint of a core-periphery pattern: responses in peripheral countries tend to lie above the EA aggregate at medium horizons, whereas those of core countries peak below the EA at the same horizons. Notably, the magnitude of the response for Greece is particularly strong. On the labor market side, the most pronounced differences are observed for Greece and Spain, whose responses exceed the EA aggregate, and for Italy and France, whose responses are comparatively lower. These pattern is coherent with the one observed by barigozzi2014euro, despite the different samples considered.
Overall, our results indicate the presence of moderate, yet non-negligible, heterogeneity across EA countries. The estimates point to a core-periphery pattern in price and interest rate dynamics. In terms of prices, core economies react broadly in line with--or somewhat more strongly than--the EA aggregate, whereas the responses of peripheral countries are more muted. Conversely, interest rates in peripheral countries tend to display comparatively stronger effects over the medium term. No systematic pattern emerges for real variables, although some countries, such as Italy and Greece, deviate more visibly from the aggregate dynamics. In contrast, stock prices behave in a largely uniform manner. Among all members, Germany’s reactions most closely match the EA aggregate, while Greece exhibits the largest departures.
These results do not signal major concerns for the conduct and transmission of monetary policy in the EA. However, they highlight the importance of closely monitoring within-country dynamics to further smooth policy transmission and mitigate adverse effects on domestic production, labor market outcomes, and--most importantly--financing costs, particularly in light of the recent post-Covid surge in public spending.
To investigate cross-country differences in responses to monetary policy shocks, we follow the methodology outlined in corsetti2022one. For each country, we compile data on selected institutional characteristics, including--but not limited to--labor market features, the financial conditions of firms and households, wage and price dynamics, and other country-specific factors. Whenever possible, these variables are averaged over the sample period; otherwise, we use data up to the most recent available period. We then compute the correlation between each explanatory variable and the peak response of the country-specific IRFs. The results of this analysis are reported in Table (ref).
We begin by examining labor market indicators, including wage flexibility and job security. For wages, we consider three levels of wage adjustment frequency: more than once a year, once a year, and less than once a year. We find no statistically significant relationship between downward wage rigidity and heterogeneous responses across countries. Nevertheless, the signs of the correlations suggest that higher wage rigidity is associated with a stronger response in real activity and a more muted response in prices, consistent with a flatter aggregate demand curve christiano2005nominal. Results for the employment protection index are weaker: as expected, higher employment protection corresponds to a smaller increase in unemployment, but its effect on other variables is relatively muted.
Turning to firm characteristics, lower price stickiness is associated with a stronger transmission of monetary policy to prices, along with a more muted response in real activity alvarez2016real. Regarding the financial structure of firms, the results suggest that a higher leverage ratio of non-financial corporations relative to GDP is linked to a weaker response in real activity and stock prices, but a stronger response in prices. Similar findings are reported by dedola2005monetary, as higher leverage proxies for borrowing capacity and, consequently, firm performance.
For households, we focus on their housing situation and consumption behavior. Our results indicate that homeownership per se provides little information on cross-country heterogeneity. In contrast, the method of financing housing appears more relevant. Countries with a higher share of homeowners with a mortgage experience a stronger impact on prices (and a weaker impact on stock prices), whereas the opposite holds for homeowners without a mortgage. This finding is consistent with recent literature, which, building on rational inattention models, shows that households with mortgages are more attentive to central bank communication and interest rate decisions, thereby enhancing monetary policy transmission ahn2024effects. We do not find any meaningful relationship between the type of mortgages (i.e., fixed versus floating) and country-specific responses. However, in countries where households have higher loan-to-value ratios on their mortgages, the impact on stock prices is smaller, further reinforcing the mechanism described in dedola2005monetary. Unlike corsetti2022one, we do not find any statistically significant results related to households’ consumption habits. However, the direction of our findings is consistent with theirs. Specifically, hand-to-mouth households tend to dampen the transmission of monetary policy to prices, production, and the labor market, while increasing its impact on interest rates and having no effect on stock prices. Similar patterns are observed for wealthy hand-to-mouth consumers.
From a broader perspective, higher saving rates are associated with a stronger impact on prices. In contrast, long-term interest rates exhibit a less pronounced increase, while the negative effects on stock prices are mitigated.
Finally, we examine how the degree of cross-sectional comovement for each variable, as quantified in Section (ref), influences dynamic responses across countries. We find that countries whose industrial production is more closely aligned with the rest of the panel are notably more affected by the contractionary shock. This finding is consistent with Table (ref), which shows the largest comovements in industrial production among the main manufacturing countries in the EA. In contrast, the impact on long-term interest rates is smaller for countries that are more aligned with the panel. This reflects the fact that countries with more idiosyncratic dynamics experience the strongest transmission of monetary policy to sovereign yield movements.
Overall, the results highlight several interesting channels through which a common monetary policy is transmitted across countries. However, these findings should be interpreted with caution, as they are based on variability across a relatively small cross-section of countries and should be regarded primarily as descriptive.
We introduce a new high-dimensional dataset, EA-MD-QD, covering both the euro area (EA) and its ten largest member countries. The dataset, including all its vintages, is publicly available online and offers a freely accessible, continuously updated resource for macroeconomic analysis. We employ this novel dataset to study the effects of a monetary policy shock on EA- and country-level variables using a Common Component SVAR. We then examine the presence and magnitude of cross-country heterogeneity in the dynamic responses to a common monetary policy shock.
Our analysis reveals moderate but meaningful heterogeneity in the transmission of monetary policy across EA countries. Price and interest rate dynamics display a core-periphery pattern. while real variables exhibit less systematic cross-country variation, and stock prices, by contrast, behave relatively uniformly across countries.
Finally, an exploratory analysis relating peak country-level impulse responses to country-specific characteristics suggests that heterogeneity in the transmission of monetary policy may be linked to factors like homeownership and saving behavior. While these correlations should not be interpreted causally, they indicate that domestic economic factors can affect the strength of common monetary shocks across EA countries.