Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
140,562 characters · 12 sections · 85 citation commands
-1.4cm bred Maximally Forward-Looking Core Inflation
\sloppy \\ \color{dpd} \fontfamily{phv}\selectfont \quad \quad \quad Universit\'{e} du Qu\'{e}bec \`{a} Montr\'{e}al \quad \quad \quad \newline \and Karin Klieber \\ \color{dpd} \fontfamily{phv}\selectfont \quad \quad \quad Oesterreichische Nationalbank \quad \quad \quad \newline \and Christophe Barrette \\ \color{dpd} \fontfamily{phv}\selectfont \quad \quad \quad \enskip \phantom{..}Universit\'{e} du Qu\'{e}bec \`{a} Montr\'{e}al \enskip \quad \quad \quad \newline \and Maximilian G\"obel \\ \textbf{\color{dpd} \texttt{\fontfamily{phv}\selectfont \quad \quad \quad \quad \enskip \quad Bocconi University} \quad \quad \quad \quad} }
\setstretch{0.98}
\center \setstretch{1.2}
\thispagestyle{empty}
\setcounter{page}{1}
\newgeometry{left=2 cm, right= 2 cm, top=2.3 cm, bottom=2.3 cm}
Ideally, monetary policy should be forward-looking. However, it can only be as forward-looking as the warning lights on the dashboard. Given the extended delay between policy impulse and the economy's response, the benefits of more timely inflation gauges cannot be understated. Core inflation measures, which distill noise and amplify signal, are crucial inputs guiding monetary policy making, investment strategies, and other consequential economic decisions. But what is noise, what is signal, and what kind of signal are we interested in?
Core inflation measures abound, and all answer the above in their own way. The most well-known are built from heuristics. Some permanently exclude components (e.g., food and energy) that are known to be subject to large (mostly) transitory shocks that can obscure from the deeper underlying trend gordon1975alternative. Others identify the components experiencing the most extreme (positive or negative) growth in each month, and exclude those before aggregating the remainder bryan1991median,BryanCecchetti1993. Then, there are more formal alternatives based on factor models where the key idea is that core inflation is a latent variable related to observable data through some assumed statistical structure, and can be extracted as such StockWatson2016,BanburaBobeica2020. In a subsequent step, core inflation series' empirical merits are assessed according to a variety of criteria. One of the most desirable properties to support timely monetary policy decision-making is that the core series should be as indicative as possible of future inflation conditions clark2001core,Cogley2002. Thus, from this perspective, all the above qualify as unsupervised learning. Model design and evaluating forward-looking qualities are mostly independent affairs. Significant volatility reduction with respect to headline inflation can come at the cost of the resulting product being a lagging indicator. This paper proposes to focus on the forward-looking criterion, and writes a simple supervised learning algorithm that delivers a core series that satisfies it.
\vskip 0.15cm {\sc Assemblage Regression.} We abide by the "train for what you aim" principle. {Our first proposition assembles components so that the resulting aggregate is maximally predictive of the hard indicator that the central bank intends to bring (or keep) at 2%. } This is achieved through a generalized nonnegative ridge regression where the dependent variable is future headline inflation and the regressors inflation subcomponents. At aggregated levels, where components are scarce and (additional) regularization is expendable, this is constrained nonnegative least squares. \textcolor{black}{However, regularization is necessary when considering highly disaggregated levels, an environment where our penalized regression is most free to discover non-trivial predictability patterns.} Our second offering extracts the supervised summary statistics of the realized distribution of price growth rates that best achieve the same objective. This is done by running the assemblage regression and replacing components data by the order statistics time series of the price growth rates distribution. \textcolor{black}{Coefficients then reflect which part of the distribution gets trimmed out, which is retained, and, in the second case, what are the respective weights.} We name the newly constructed indicator Albacore, for adaptive learning-based core inflation, which summarizes both the philosophy (the construction should adapt to the objective) and the methodology. As we will see, both Albacore$_\text{\scriptsize comps}$\xspace and Albacore$_\text{\scriptsize ranks}$\xspace (respectively, weighting in components and rank space) will reaffirm parts of the common wisdom and bring new stylized facts to the table.
\vskip 0.15cm {\sc Related Work.} Econometrically, assemblage regression's closest relatives come from the regression-based forecast combination literature wang2023forecast and autoregression-based filtering Hamilton2018. In the case of Albacore$_\text{\scriptsize ranks}$\xspace, it is coupled with elements of functional regression morris2015functional.
First, the problem yielding Albacore$_\text{\scriptsize comps}$\xspace shares various similarities with regularized forecast combination schemes such as that of diebold2019machine. Namely, we are looking to build an optimal weighted average incorporating regularization inspired from what the default averaging should be in the specific application. A key distinction between assemblage regression and combination schemes following from granger1984improved is that, as the name suggests, we are putting parts of something together. Forecast combination exercises, on the other hand, construct an optimal weighted average using noisy estimates of the same thing.\footnote{An example, where forecasts are not the base material yet, the combination is built upon estimates of the same concept, is aruoba2012improving, who combine GDP by revenues and GDP by expenses (two noisy estimates of GDP) to construct a more reliable GDP indicator.} \textcolor{black}{Following from} this distinction between supervised ensembling and assembling, there are few design choices mirroring the different environments, like the construction of the target and the whole regularization apparatus. Nonetheless, the connection with forecast combination regressions is useful to locate where Albacore$_\text{\scriptsize comps}$\xspace stands in the vast core inflation measures landscape. Among other things, granger1984improved's regression approach to forecast combinations constructs an optimal portfolio of forecasts taking into account their covariance structure, which simpler "plug-in" approaches using bates1969combination's formula do not.
In the unrealistic yet informative case where inflation components would be uncorrelated, reflecting on the least squares formula for weights offers some intuition. Albacore$_\text{\scriptsize comps}$\xspace would downweight components which do not covary significantly with headline at a pre-specified forecasting horizon, have high variance, or both. Clearly, this matches the attributes of a component that should be excluded or downweighted in a core inflation building context: it is not persistent and highly volatile. Moreover, if a component exhibits negative autocorrelation patterns that, on average, have been self-absorbed within a time span corresponding to the forecast horizon, then those transitory shocks will also be left out since they have hardly any predictive power at that horizon. Things inevitably get less trivial when components are heavily cross-correlated, constraints are brought in, and regularization is biting. Still, at a conceptual level, focusing on maximal predictability is not a radical conceptual departure from approaches improving gordon1975alternative's popular suggestion by either excluding or reweighting items according to their overall volatility dow1993measuring,clark2001core,Acosta2018, their cyclical volatility dolmas2009excluding, their persistence BilkeStracca2007, or even their sensitivity to the economic business cycle shapiro2017cyclical,ehrmann2018supercore,stock2020slack. In a way, Albacore$_\text{\scriptsize comps}$\xspace encompasses most of these objectives, but does so in a framework that keeps an eye on the prize (forecasting headline) and accounting for the fact that we are aggregating correlated objects, which comes with interesting hedging possibilities if properly accounted for. There have been rare instances in the literature where predictability for headline played a more direct role in informing the weighting in the components space. However, the proposed aggregation schemes overlooked the covariance structure RavazzoloVahey2009 or lacked, among other things, necessary nonnegativity restrictions gamber2019constructing.
Core inflation estimation can be framed as a signal extraction problem and accordingly, many variants of factor analysis have been suggested cristadoro2005core,morana2007structural,khan2013common,StockWatson2016,BanburaBobeica2020. Compared to those methods, especially when relying on Principal Component Analysis (PCA), assemblage regression retains necessary shrinkage, but makes sure, through supervision, that the extracted signal is one we should care about. Obviously, by discarding some noise, one expects the ensuing factors to incorporate relevant signals to forecast headline inflation. However, it is not optimizing for it. Factors are extracted as the linear combination of components that best explain the variation in all components. The resulting product's correspondence with the "true" concept of core inflation inevitably comes from consistency under unverifiable assumptions embedded in the specification of the model. \textcolor{black}{Therefore, the empirical evaluation is usually conducted ex-post through the backchannels of its implications for a proper core series (e.g., predictive power for headline and association with the strength of real activity). } Albacore$_\text{\scriptsize comps}$\xspace and Albacore$_\text{\scriptsize ranks}$\xspace surely feature their own set of assumptions, but they enjoy the distinct advantage that the optimization objective and the measurable notion of success perfectly coincide.
The proposed method is also not entirely unrelated to approaches attempting to uncover trend inflation, usually through compact state-space models chan2016,eo2023understanding or dynamic factor models StockWatson2016. The link can be understood through the lenses of Hamilton filtering Hamilton2018. Indeed, as an alternative to Hodrick-Prescott filtering (a smoothing splines problem), Hamilton's suggestion is to use the residuals of direct autoregressive forecasting equation (at a user-specified horizon, a tuning parameter) as the detrended series. \textcolor{black}{Assemblage regression in the component space can be seen as a constrained multivariate-to-univariate Hamilton filter, where instead of regressing the target on its own lags, we regress the targets on its lagged subcomponents. Furthermore, the target is the average path between the forecast date and the forecast horizon.} The usage of the average path, typical in forecasting studies and particularly natural for inflation MDTM, leads it to take an implicit mean over many "horizon tuning parameters", a strategy that has been documented to be successful in output gap extraction applications quast2020reliable. Last but not least, one should note that unlike Hamilton2018 and the ensuing literature, the interest here lies in the analysis of the extracted "trend" rather than deviations from it.
Albacore$_\text{\scriptsize ranks}$\xspace is our proposal for a maximally forward-looking core inflation based on temporary exclusion. It comes with its own distinct strand of literature bryan1991median,BryanCecchetti1993,bryan1997efficient,dolmas2005trimmed. Traditionally, this is achieved by systematically ranking individual price changes from their lowest to highest values at each point in time and excluding the upper and lower tail of the distribution (either symmetrically or asymmetrically), with median inflation bryan1991median being the extreme case of keeping only the midpoint rank. Trimming methods have received limited attention in the forecast combination literature wang2023forecast, with most approaches zooming in on permanently eliminating the least accurate contributors diebold2019machine,WangEtAl2022 or trying a few cutting points configurations for a trimmed mean estimator stock2004combination.
Albacore$_\text{\scriptsize ranks}$\xspace features greater flexibility. It assigns optimized weights to each rank rather than solely optimizing trimming points. In terms of mechanics, Albacore$_\text{\scriptsize ranks}$\xspace is a constrained functional regression morris2015functional, where predictors (i.e., order statistics series) are a function of the empirical distribution of price growth rates. Applications of functional regression analysis and related methods in macroeconometrics are still scarce. Recent ones include meeks2023heterogeneous using functional PCA to summarize the distribution of inflation expectations distribution, and chaudhuri2016forecasting who study the autoregressive properties of different parts of the inflation distribution. As it was the case for the above discussion of PCA in components space, the advantage of Albacore$_\text{\scriptsize ranks}$\xspace over functional PCA and related methods is (i) supervision, and (ii) constraints on the kinds of features we want to be extracted.
Finally, there have been numerous papers forecasting inflation using machine learning methods (see medeiros2019, HNN, and references therein). Recently, there has been some attention in coupling those with components-level data barkan2023forecasting,boaretto2023forecasting,joseph2024forecasting. We differ from this strand of literature by our objective, to create a core inflation measure with desirable properties. This leads us to consider constrained linear forecasts based on components, as opposed to any functional form channeling in any dataset.
\vskip 0.15cm {\sc Empirical Results.} As main applications, we consider US and euro area (EA) inflation. We also study the Canadian case more compactly in the appendix. In all instances, we find that Albacore succeeds in being highly predictive of future headline inflation at various horizons. It outperforms benchmark models in terms of predictive accuracy and yields low bias and variance compared to the official headline inflation rate. In terms of forecasting results, Albacore$_\text{\scriptsize ranks}$\xspace is clearly our best performing model. Its gains mount to sizable margins, especially when taking a medium- to long-term perspective. A look at the assembled ranks reveals a highly asymmetric trimming, removing entirely the lower part of the monthly price growth distribution and putting the emphasis on the \textcolor{black}{60$^{\text{th}}$ to 75$^{\text{th}}$ percentile (depending on the application)}. Unlike existing trimmed mean measures, it keeps a non-trivial portion of the upper tail by assigning it a moderate weight. This proactive usage of the skewness in the monthly price growth rates distribution helps in staying alert to accumulating inflationary pressures, like those that were occurring in early 2021. Albacore$_\text{\scriptsize comps}$\xspace's improvements are rather concentrated within the evaluation sample spanning the months following the Covid-19 pandemic. Similarly to its supervised trimming counterpart, outperformance is more substantial for higher-order forecasts (i.e., 6 and 12 months ahead). The optimized weighting scheme results in low weight assigned to highly volatile components, such as energy and food, and high weight on components related to the services sector, health, and housing.
For the run-up of inflation, it is well known that developments in goods prices played an outsized role in the US. Accordingly, we find high contributions of goods inflation to the aggregate of Albacore$_\text{\scriptsize ranks}$\xspace. The officially reported trimmed mean, on the other hand, puts almost no weight on goods components, instead it is driven by shelter, a notoriously lagging component bils2004sticky,cotton2023forecasting. Given that this pattern manifests around inflation turning points, it can prove useful to downweight contributions from shelter using an asymmetric trimming scheme, as is done by Albacore$_\text{\scriptsize ranks}$\xspace. Alternatively, one can assign a downsized weight to problematic subcomponents (such as tenant rent) and compensate by upweighting others, an opportunity which is leveraged by Albacore$_\text{\scriptsize comps}$\xspace, thanks to its high-dimensional regression setup. This is less radical than excluding the shelter component entirely, as in US supercore inflation powell2022,leduc2023will. In the EA, sustained energy and food price shocks played a larger role than in the US, which makes the official core inflation rate a lagging indicator. Albacore$_\text{\scriptsize comps}$\xspace and Albacore$_\text{\scriptsize ranks}$\xspace avoid this predicament via positive contributions from food goods. Regarding the early indication of the disinflationary phase, both Albacore attribute it to easing goods price pressures for the US and the EA.
We explore a suite of extensions. First, we consider building core measures specialized for either upside or downside inflation risks by changing the squared error loss in assemblage regression to a quantile loss. These can be equivalently interpreted as the optimal core inflation measure for an analyst facing asymmetric forecast error costs. Interestingly, we find more harmony between traditional core inflation measures and ours when customizing the latter for disinflation risk. \textcolor{black}{Thereby, gains from our proposed approach are highest in the upper tail, especially so for our quantile extension of Albacore$_\text{\scriptsize ranks}$\xspace post-Covid.} Second, we consider forward-looking yearly time-aggregation of headline inflation for the US. We find that the autoregression in rank space decisively outperforms traditional autoregressions for this task. It does so by eliminating large negative realizations of the month-over-month growth rate within the lookback window, and including positive ones with a moderate weight. This suggests that, in such simple models, it is a more effective strategy to pay attention to the position of an observation in the recent realized distribution rather than the exact location in time, as is typically the case in standard autoregresssions. Third, we consider assembling headline inflation from euro area member states to forecast the aggregate. Albeit these do not necessarily qualify as a traditionally defined core inflation, our results show that a downweighted average of the 4-5 highest state-level inflation rates at a given point in time offers substantial forecasting accuracy gains.
\vskip 0.15cm {\sc Outline.} We proceed as follows. In Section (ref), we introduce the Assemblage Regression. Section (ref) presents our empirical application to core inflation. Section (ref) focuses on applications of Albacore$_\text{\scriptsize ranks}$\xspace to other inflation-related aggregation problems. Section (ref) concludes and proposes various avenues for future research. Appendix (ref) contains additional material, like our more condensed Canadian application.
In this section, we introduce the Assemblage Regression, which, in a nutshell, is a generalized nonnegative ridge regression where the dependent variable is future headline inflation and the regressors are either inflation subcomponents or transformations of them. Nonnegativeness comes from the desire of the output being a weighted average and "generalized" comes from the use of altered regularization schemes. There will be two declinations of our ideas. The first and most conceptually straightforward, Albacore$_\text{\scriptsize comps}$\xspace, aggregates and weights \textcolor{black}{inflation} components directly. In essence, it searches the space of basket weights for which consumers' today inflation is most indicative of the average consumer inflation in the near-future. The second, marginally more sophisticated one performs optimal trimming of inflation components (i.e., temporary exclusion/weighting). Albacore$_\text{\scriptsize ranks}$\xspace is the resulting measure -- it runs the assemblage regression in the rank space of the components data.
\vskip 0.2em
{\sc Ridge Regression Primer.} A ridge regression solves what is more generally known as a penalized linear regression problem ESL. It is a traditional least squares problem plus a penalty on coefficients so to regularize the outcome of minimization. Ridge regression coefficients are obtained via $$\hat{\boldsymbol{\beta}}_\text{Ridge} = \operatorname*{arg\,min}_{\boldsymbol{\beta}} \sum_{t=1}^{T-1} ( y_{t+1} - \boldsymbol{\beta}'\boldsymbol{X}_{t})^2 + \lambda ||\boldsymbol{\beta}||_2 $$ where $ || . ||_2$ is the $l_2$ norm. The latter is equivalent to $\sum_{k=1}^{K} \beta_k^2$ in summation notation, where $K$ \textcolor{black}{denotes the number of components}. The penalty term provides regularization by bringing in the mix the a priori that each coefficient should contribute to the fit, but modestly. In other words, it is shrinking coefficients towards 0. With a suitable $\lambda >0$, a ridge regression curbs overfitting that plagues OLS (i.e., $\lambda=0$) when $K$ is large relative to $T$ and the signal-to-noise ratio is low. \textcolor{black}{Everything good comes at a cost, and here, the cost is bias. $\hat{\boldsymbol{\beta}}_\text{Ridge}$ blends information from the data and our prior knowledge on $\boldsymbol{\beta}$. Thus, the specification of the penalty must be chosen wisely, as must its strength.} $\lambda$ is typically tuned via cross-validation, which is, in essence, a pseudo-out-of-sample evaluation metric. The amount of restrictiveness, which will impact both forecasts and the coefficients (and thus the interpretation of the model), is chosen so to maximize the model's predictive accuracy on unseen data. Changing the $l_2$ norm for $l_1$ makes it the equally well-known LASSO problem. In fact, there is a constellation of possible penalty terms one can choose from depending on the application, and we shall leverage that on the way to our main model hastie2015statistical. Lastly, penalized regression problems are not scale-invariant, so either the predictors have to be scaled to exhibit the same variance, or the penalty should be adjusted accordingly.
\vskip 0.2em
{\sc Supervised Weighting.} We now provide the details of our first assemblage regression. The permanent exclusion version, or supervised weighting of basket components, is obtained via
where $h$ is the forecasting horizon, $T$ is the last training observation, $\pi_{t+1:t+h}$ is average headline inflation between $t+1$ and $t+h$, and $ \boldsymbol{\Pi}_{t}$ is a matrix of component time series at a user-specified level of aggregation. $\lambda ||\boldsymbol{w}-\boldsymbol{w}_{\text{\tiny headline}}||_2$ is the penalty term, and it shrinks the solution towards $\boldsymbol{w}_{\text{\tiny headline}}$, which are headline inflation weights for that level. \textcolor{black}{We constrain weights to be nonnegative ($\boldsymbol{w} \geq 0$) and sum to 1 ($\boldsymbol{w}'\iota=1$).} \textcolor{black}{$ \boldsymbol{\Pi}_{t}$ are expressed as price growth rates and in our applications, it is the 3-months-over-3-months growth rate. This allows for minimal time-smoothing while retaining timeliness BanburaBobeica2020, and does not prevent for the measure to rightfully qualify as core inflation now. }
Albacore$_\text{\scriptsize comps}$\xspace is defined as $ \pi^*_{\text{comps}, t} = \hat{\boldsymbol{w}}_c '\boldsymbol{\Pi}_{t}$ and is, as such, a constrained forecast -- both in terms of the information set (only inflation rates) and how it is synthesized. It is the linear aggregation, or weighted average, of components that is most closely related to future headline inflation as captured by $\pi_{t+1:t+h}$. The adaptiveness is tied to the objective, and it is natural to expect core measures to vary in composition along $h$ and depend on the loss function itself (see Sections (ref) and (ref), respectively). Core measures that investment bankers, monitoring market reactions to inflation number releases, should care about can be different from those of interest to central bankers dealing with the long lags of monetary policy. Depending on the application, the user may choose the forecasting horizons $h$ for the target and the aggregation level of components. While we consider various numbers of $h$ in our empirical applications to showcase the versatility of the approach, the main analysis will focus on $h=12$ months which is widely regarded as the most relevant time frame for the use of core inflation measures in monetary policy decision-making. Further details are given in Section (ref).
Shrinking to $\boldsymbol{w}_{\text{\tiny headline}}$ is intuitive as it represents the official statistical aggregation of the components and implies under $\lambda \rightarrow \infty$ a random walk forecast (if the moving average length of the target matches that of regressors).\footnote{Note that in this context, opting for the traditional ridge penalty $\lambda ||\boldsymbol{w}||_2$ would imply shrinking to equal weights. {The other two constraints (i.e., nonnegative weights which sum to 1) prevent the fit from collapsing to $\boldsymbol{0}$ and thus, shrink what is left to an identical value.} Nonetheless, this {regularization scheme} may suffer from the fact that certain subcomponents of an upper level are simply more granular and numerous, and would push the model to favor items that are highly numerous rather than those with higher consumption weights. } We resist the temptation of shrinking to core inflation (ex. food and energy) weights, which surely implies a less volatile "null" model than what we consider, and perhaps better empirical results. This is motivated from the desire to see, as a basic check, whether Albacore$_\text{\scriptsize comps}$\xspace can discover this everlasting empirical wisdom without assuming half of the discovery first. It is also plausible that some forward-looking core measures at shorter horizons could significantly differ from the typical core measures, which have historically been focused on medium- to long-term horizons. \textcolor{black}{Taking on a different perspective, the assemblage regression can be viewed as an algorithm searching the space of basket weights for which consumers' inflation today is most indicative of the average consumer inflation tomorrow. Although, the resulting assemblage does not explicitly fulfill any requirement to match a representative agent per se, it would be surprising if those two were vastly unrelated. Yet, if that were to prove insufficient according to pseudo-out-of-sample performance, the penalty $\lambda ||\boldsymbol{w}-\boldsymbol{w}_{\text{\tiny headline}}||^2$ shrinks coefficients to the official weights (instead the usual 0).}
Lastly, one needs to think about the strength of that regularization, as captured by $\lambda$. This can be set automatically, but it needs to be done the right way. The target time series is persistent, particularly starting from $h>6$, and so are some components. Thereby, choosing $\lambda$ with standard cross-validation designed for independent data will deliver an overconfident assessment of the generalization error and a downward biased $\lambda$. As recommended in GCFK and others, we opt for a non-overlapping blocks approach. We divide the training sample in 10 contiguous segments and use those in 10 fold cross-validation. The length of blocks thus depends on that of the training sample (e.g., an estimation sample of 20 years implies blocks of two years). While longer blocks could be desirable for $h=12$ and above, one must remember that too {few} blocks and lacking randomization likely would outweigh the benefits.
\vskip 0.2em
{\sc Supervised Trimming.} The second, supervised trimmed inflation, needs further thought in order to be cast within the assemblage regression apparatus described above. After all, trimming implies components jump in and out of the index every month, implying a kind of time variation in their weights based on how each component growth rate ranks compared to others in a given month. Surely, going straight at it in such a fashion would push us away from the "keep it sophisticatedly simple" principle diebold1998elements. The opposite direction, i.e., that of simplicity at the expense of generality and flexibility, would bring things closer to bryan1997efficient's trimmed mean PCE inflation. There, components are either out (with a weight of 0), or in (with a weight proportional to their basket weight versus that of other non-trimmed components at time $t$). The size of trimmed bands on both tails can then be optimized over two parameters using some criterion. The resulting series is a robust weighted average. However, the mean (or the median for that matter), may or may not \textcolor{black}{align with} where our eyes should be to extract leading signals for future headline inflation.
Instead, one can consider, more generally, retrieving the summary statistics of the current realized price growth distribution that best fulfill a specific statistical objective. Or equivalently, components get weights as a function of their location in the empirical distribution at time $t$. A closer examination of what such time variation implies reveals that, in fact, we can run a very similar assemblage regression as described above, but using the empirical order statistics of $\boldsymbol{\Pi}_{t}$ as regressors. To see this, we can simply walk from the components space to the "rank space", by writing the formula for fitted values in summation notation
where $O_{r,t}$ is the $r^{\text{th}}$ order statistic (i.e., the value attached to rank $r$) at time $t$. If $K$ were to be equal to 100, $O_{1:100,t}$ would be a set of poor man's percentiles. The above derivation informs us that this kind of time variation in the component space implies time-invariant coefficients in the rank space (and vice versa). We will later use this duality to map back the implication of $w_r$ in the component space in Section (ref). Thus, ease of implementation can be restored from running the model in rank space, i.e., using $\boldsymbol{O}_{t}$, which is effectively just sorting the components at each $t$ and stacking them in a matrix. Precisely, $ O_{1:K,t}=\texttt{sort}(\Pi_{1:K,t}) \enskip \forall t$. The switch to rank space-based weighting can also be formalized in our preferred matrix notation, with $\boldsymbol{O}_{t}= \boldsymbol{A}_{t}\boldsymbol{\Pi}_{t}$ where $\boldsymbol{A}_{t}$ is a $T \times T$ allocation matrix where entries are 0s and 1s following $ I \left(\text{rank}(\Pi_{k,t})=r \right) \enskip \forall (r,k) $. Thus, we run
where $D$ is the difference operator and Albacore$_\text{\scriptsize ranks}$\xspace is defined as $ \pi^*_{\text{ranks}, t} = \hat{{\boldsymbol{w}}}_r '\boldsymbol{O}_{t}$. Rather than learning which subcomponents to include, the problem will now be learning which ranks (and with which weight) to include or exclude. $||D\boldsymbol{w}||_2 $ (or equivalently $\sum_{r=1}^K \left(w_r - w_{r-1}\right)^2 $ in summation notation) is a fused ridge penalty hastie2015statistical. It favors a smooth weighting scheme, and in the $\lambda \rightarrow \infty$ limit, pushes the model towards the sample mean solution where each rank gets a weight of $\sfrac{1}{K}$.\footnote{ Note that, in the $\lambda \rightarrow \infty$ limit, solutions are equivalent in components and rank space ( $ \iota ' \boldsymbol{O}_t= \iota ' \boldsymbol{A}_{t}\boldsymbol{\Pi}_{t} = \iota ' \boldsymbol{\Pi}_t$). } In the other limit ($\lambda=0$), the solution will be sparse for reasons we will come back to below.
The equality constraint has been changed from $\boldsymbol{w}'\iota=1 $ to $\bar{\pi}_{t+1:t+h} = \bar{ \pi}^*_{ranks, t}$. In the rank space, $\boldsymbol{w}'\iota=1 $ would imply that the average of all weights is $\sfrac{1}{K}$, which does not enforce symmetric trimming, but rather would impose equivalent masses above and below $\sfrac{1}{K}$, and thus blocking highly asymmetric outcomes which we will find to be successful in nearly all our experiments. The $\bar{\pi}_{t+1:t+h} = \bar{ \pi}^*_{\text{ranks}, t} $ constraint brings back discipline by forcing residuals to have mean 0. They could deviate marginally because the model has no intercept, and to minimize the sum of squared residuals, some bias could be traded for variance reduction. We block this possibility by imposing the constraint that Albacore$_\text{\scriptsize ranks}$\xspace has the same long-run mean as headline inflation (over the training sample).
{This stands in contrast to most trimming approaches, which opt for a preselected band within which ranks get assigned their official (components space) weights, and those outside of it have a weight of 0.} For instance, the FRB Cleveland trimmed mean CPI sets to 0 the first and last 16% of ranks, and then a weighted average of the center band is reported. In a similar vein, the FRB Dallas' Trimmed Mean PCE trims out 24% from the lower tail and 31% from the upper tail dolmas2005trimmed. \textcolor{black}{Then, there is the} Cleveland Fed Median CPI, which keeps only one rank, the middle one. Those trimming approaches can be seen as knife-edge cases of the above, where $\hat{\boldsymbol{w}}$ is chosen based on heuristics. It is noteworthy that the fused penalty could be changed to $\lambda ||D\boldsymbol{w}||_1 $ (fused lasso) if one wanted to favor a sharp (boxy shaped) trim. We do not opt for such a possibility given that our framework needs not be limited to such shapes. Having this in mind, it is important to note how our proposition differs from what is commonly done with trimmed measures. It does not reweight included components according to some rule based on their headline inflation weights, nor does it enforce that weights sum to 1 at every point in time.
\vskip 0.2em
{\sc Additional Methodological Remarks.} Note that there is no intercept included in our regressions, which has a few implications. First, the fitted values are mechanically pushed (but not forced unless otherwise specified) towards having the same unconditional mean as $\pi_{t+1:t+h}$ through $w_{0} = \bar{\pi}_{t+1:t+h} - \bar{ \pi}^*_{t} = 0$, therefore making the fitted values already in headline inflation units by construction. Second, it implies a latent random walk hypothesis, which is in synchronicity with the idea of a core inflation measure being the summary of inflation conditions right now that is most indicative of future developments in headline inflation.
Additionally, excluding the intercept $w_{0}$ is necessary to properly identify the elements of $\boldsymbol{w}$ in Equations (ref) and (ref) at longer horizons, like $h=24$. As $h$ grows, the unconstrained regression solution puts a very small weight on current inflation conditions -- and a much larger one on the intercept. This implies a very limited passthrough of information from $ \pi_{t+1:t+h}$ to $\boldsymbol{w}$, which job is to effectively bundle the ensemble of components/ranks. \textcolor{black}{ Thus, with little variance from the supervisor $ \pi_{t+1:t+h}$ being canalized into the learning of $\boldsymbol{w}$, its elements can only be estimated imprecisely, and, going towards the limit of an "intercept only" model, are identified by the prior (i.e., the penalties).} Evidently, this problem does not occur when $\boldsymbol{w}$ is predetermined using heuristic rules, like for most traditional core measures. Interestingly, we will find that excluding the intercept also improves the benchmark regressions in most cases and time periods, but not nearly as much as for Albacores.
The combination of various constraints has implications of its own. In both regressions, the exclusion of the intercept -- implying $w_{0} = \bar{\pi}_{t+1:t+h} - \bar{ \pi}^*_{t} = \bar{\pi}_{t} - \bar{ \pi}^*_{t} \approx 0$ -- combined with the nonnegativity constraint implies adaptive $l_1$ regularization. \textcolor{black}{To see this, note that in an attempt to mitigate variance, $\bar{ \pi}^*_{t}$ may fall marginally below $\bar{\pi}_{t}$, unless equality is forced as in Albacore$_\text{\scriptsize ranks}$\xspace, and ergo classifies as an implicit soft constraint. In other words,} we have $\bar{ \pi}^*_{t} \leq \bar{\pi}_{t}$, which can, by definition, be rewritten as \textcolor{black}{$\sum_{k=1}^K w_k \bar{ \pi}_{k,t} \leq \bar{\pi}_{t}$}. Bringing in the nonnegativity constraint and denoting $\gamma_{k} \equiv \frac{\bar{ \pi}_{k,t}}{\bar{\pi}_{t}}$, we get
which is the constrained optimization form of an adaptive LASSO constraint with a "budget" of 1 -- technically, a tuning parameter. In the case of Albacore$_\text{\scriptsize ranks}$\xspace, Equation (ref) holds with equality. \textcolor{black}{In that of Albacore$_\text{\scriptsize comps}$\xspace sum-to-one constraint, the implications are Equation (ref) holding with equality with $ \gamma_{k} = 1 \phantom{.} \forall k$ , a plain LASSO constraint in addition to Equation (ref) itself.} Thereby, unless other constraints are brought in, the resulting model will likely be sparse. Is this desirable? It depends. In the high-dimensional Albacore$_\text{\scriptsize comps}$\xspace application, this is undesirable given that we expect the forward-looking core measure to be dense, i.e., based on more than a handful of highly disaggregated components. At higher levels of aggregation, sparsity becomes reasonable again, but may lead to a measure that is far from mapping into any realistic consumer, which may impede the successful communication of the index. In rank space, sparsity can make sense, because order statistics themselves are already dense combinations of underlying components through the rule underlying the allocation matrix. A well-known sparse core inflation measure in rank space is \textcolor{black}{the weighted median inflation}. If $\lambda=0$ is the outcome of cross-validation, both Albacore$_\text{\scriptsize comps}$\xspace and Albacore$_\text{\scriptsize ranks}$\xspace can embrace sparsity by letting the cocktail of constraints act as the single source of ($l_1$) regularization. We will see that most often, dense solutions are favored.
\vskip 0.2em
{\sc Implementation Details.} We use the CVXR package in R, which provides fast solutions for linear convex programming problems. Simplified versions of the above could be implemented via the well-known glmnet package. However, necessary equality constraints (such as $ \boldsymbol{w}'\iota=1$ in Albacore$_\text{\scriptsize comps}$\xspace) are beyond their functionalities, and one cannot simply rearrange regressors as in OLS. We provide the R package assemblage, which automatically implements the above given user-provided target and components. Running estimations implying fewer than 20 components takes about 0.14 seconds on a Macbook Pro (with M2 chip), and a full-patch cross-validation of $\lambda$ takes 5.53 seconds (with a $\lambda$ grid of length 20 and 10 folds). Larger models with above 200 components take 0.26 seconds for a single run and 7.91 seconds for cross-validation. These numbers assume running cross-validation in parallel (coded as an option in the assemblage package) on 11 cores. Using a single core implies the tuning of $\lambda$ now takes 23.21 seconds for the small model ($K=20$) and 39.17 seconds for the bigger one ($K=200$).
We construct Albacore for the US, the EA\textcolor{black}{, and Canada} with monthly price indices at different levels of disaggregation. \textcolor{black}{Results and implementation details for the latter are relegated to Appendix (ref).} For the US we focus on the price index for Personal Consumption Expenditure (PCE), which is taken from the Bureau of Economic Analysis (BEA). We base our core inflation measure on level 2, 3, and 6 including 15, 50, and 215 subcomponents, respectively. For the euro area, we use the Harmonized Index of Consumer Prices (HICP) from Eurostat amd choose the two-, three-, and four-digit COICOP level, which comprises a number of 12, 39, and 92 subindices.\footnote{We exclude item CP0735 (Combined passenger transport) from level 4 in the EA due to the item's heavy irregularities during the Covid-19 pandemic.} All disaggregated series are seasonally adjusted\footnote{Note that for the EA and Canada (unlike the US) there is no seasonally adjusted data published by statistical institutions. We provide details on the seasonal adjustment for both applications in Appendix (ref).} and expressed as 3-months-over-3-months changes, which balances timeliness (as opposed to 12 months trailing averages) and noise reduction (versus the month-over-month growth rate). Rank positions are calculated using month-over-month growth rates (as done for trimmed mean inflation), and then the order statistics time series are smoothed using the 3 months moving average.
Constructing Albacore for each level involves predicting the average path of headline inflation for 1, 3, 6, 12, and 24 months ahead ($h \in \{1,3,6,12,24\}$). Albacore based upon $h \in \{6,12,24\}$, indicating medium-term developments, is of interest for policymakers. An index targeting shorter horizons ($h \in \{1,3\}$) can be useful to hedge funds performing reallocations based on anticipation of central banks rate decisions, especially in times where there is sizable monetary policy uncertainty.
We choose two out-of-sample test sets spanning periods before the Covid-19 crisis (i.e., 2010m1 to 2019m12) and the post-Covid period (i.e., 2020m1 to 2023m12) and evaluate the point forecasting performance with root mean squared errors (RMSEs). These are two vastly different regimes, with the first evaluation sample being characterized by stable low inflation and the second by unstable high inflation. We base the estimation on a rolling window of 20 years, allowing for mild structural changes in the composition of the maximally forward-looking inflation series. Note that for the EA this choice inevitably results in an expanding window since our disaggregated data set starts in 2003.
We compare the forecasting accuracy to a set of benchmark models. For each country, this set includes the officially reported core inflation series (i.e., PCE/HICP excluding energy and food, henceforth, core PCE/HICPX), trimmed mean inflation (i.e., FRB Dallas' Trimmed Mean PCE for the US, 30% trimmed mean inflation for the EA) as well as established model-based concepts of underlying inflation measures. Given that the series address the problem from various angles, the information they convey may differ. Hence, instead of focusing on a single measure of core inflation we enhance our benchmarks by combining them and, in this way, perform an ex-ante weighting of the best currently available measures Cogley2002.
Similar to our main setup, we include the different series in a nonnegative regression. Our first benchmark (and numéraire for all RMSEs in our tables) is comprised of headline, core as well as trimmed mean inflation including an intercept ($\boldsymbol{X}_t^{\text{bm}}$), which allows the model to include the long-run average. The same set of variables without the intercept ($\boldsymbol{X}_t^{\text{bm}}$, $(w_0=0)$) is used as another competitor. The former is more akin to a standard forecasting regression, while the latter is more in tune with traditional usage of core inflation measures and the inherent random walk hypothesis. Accordingly, we expect $\boldsymbol{X}_t^{\text{bm}}$ to be a tough benchmark during the stable inflation out-of-sample period, and $\boldsymbol{X}_t^{\text{bm}}$, $(w_0=0)$ to be more competitive when facing the rapidly evolving inflation conditions of our second test sample. We also consider a more comprehensive set of benchmarks by combining 7 publicly available inflation measures for the US, and 8 for the EA ($\boldsymbol{X}_t^{\text{bm+}}$).\footnote{For the US we use PCE services other than housing, FRB Dallas' Trimmed Mean PCE, FRB Cleveland Median CPI, FRB Cleveland 16% Trimmed Mean CPI, the Atlanta Fed Sticky CPI, core PCE (ex. food and energy) and PCE headline. The set for the EA includes HICP, HICP excluding food and energy, HICP excluding energy, HICP excluding energy and unprocessed food, Supercore, PCCI, PCCI excluding energy and the 30% trimmed mean inflation.} Again, we estimate it with and without an intercept ($\boldsymbol{X}_t^{\text{bm+}}, (w_0=0)$).
First, Table (ref) contains the out-of-sample forecasting exercise results. Next, Figure (ref) shows the resulting time series as well as the weights assigned to ranks and components. Finally, in Figure (ref), we zoom onto the post-2020 period and decompose Albacore and two classical core measures into the main aggregates. This provides an understanding of their different assessments for both the surge and the slowdown.
\vskip 0.2em {\sc Forecasting Performance.} For the US data, we find that Albacore$_\text{\scriptsize ranks}$\xspace is the best performing model regardless of the forecasting horizon, level of disaggregation, and evaluation sample. In most cases, it outperforms benchmarks by an appreciable margin, with the highest gains achieved for higher-order forecasts. For the sample ending before the pandemic, Albacore$_\text{\scriptsize ranks}$\xspace yields RMSEs well below all benchmarks with highest performance gains for medium-term predictions ($h \in \{12,24\}$) and higher numbers of assembled components ($K \in \{50,215\}$). Lower-dimensional versions fare better post-2020, and the middle-ground "level 3" option comes out as the most polyvalent. Albacore$_\text{\scriptsize comps}$\xspace's outperformance is more local than that of its supervised trimming counterpart. For the low inflation era, it does not eclipse convex combinations of the usual benchmarks. However, its outperformance for the second test set is notable, particularly for key horizons such as 6 and 12 months ahead.
We also note from Table (ref) that dispensing with the intercept is, unsurprisingly so, playing a non-trivial role in creasing performance for the second out-of-sample. Indeed, both the sparse and more comprehensive benchmarks deliver better results when excluding it. Albacore$_\text{\scriptsize ranks}$\xspace and Albacore$_\text{\scriptsize comps}$\xspace surely benefit from being part of the no-intercept family for the post-2020 evaluation period, but they comfortably distance the benchmarks. Of the $w_0=0$ cluster of models, only Albacore$_\text{\scriptsize ranks}$\xspace surpasses the performance of models featuring an intercept during the low/stable inflation era.
Interestingly, Albacore$_\text{\scriptsize ranks}$\xspace also provides appreciable improvements for short-run forecasts ($h \in \{1,3\}$), and does so for both test periods at nearly all levels. This is notable given that short-run predictive accuracy is hard-earned -- beating headline itself for these targets is no small feat. Most core inflation measures are designed with "monetary policy horizons" in mind, and as such, are suboptimal for such needs. Thus, Albacore$_\text{\scriptsize ranks}$\xspace can also be of use to institutions interested in monitoring a single inflation series that has an edge in foreseeing headline's next release ($h=1$) or its developments over the next 3 months.
\vskip 0.2em {\sc A Look at Time Series.} We focus on $h=12$ and the most granular level (i.e., level 6) for Albacore$_\text{\scriptsize comps}$\xspace and level 3 for Albacore$_\text{\scriptsize ranks}$\xspace. This choice is based on \textcolor{black}{the importance of medium-term horizons for monetary policy decisions and forecasting results discussed above. Acknowledging the remarkable performance of lower levels of disaggregation, we present deeper insights on Albacore for level 2 in Appendix (ref).} Figure (ref) presents Albacore estimated on data through 2019 by plotting the time series of the aggregate \textcolor{black}{and benchmarks in yearly percentage changes} (upper panel) and the weights of subcomponents/ranks in the lower panels. Albacore$_\text{\scriptsize ranks}$\xspace indicates inflation trends close to the 2% target for the periods up to the Great Financial Crisis (GFC). After showing slight upward pressures in 2009, it falls below target in 2010 and remains low and stable up to 2020. Albacore$_\text{\scriptsize comps}$\xspace deviates from both Albacore$_\text{\scriptsize ranks}$\xspace and the target already in 2004. It results in elevated inflation until 2008 before signaling downward pressures during the GFC. From 2010 to 2019, all core inflation measures move in accordance before diverging ahead of the Covid-19 pandemic. Indeed, permanent exclusion metrics (i.e., core PCE and Albacore$_\text{\scriptsize comps}$\xspace) follow headline downward in 2019 whereas the trimming-based approaches remain at the target.
The subsequent periods, covering the pandemic-era inflation surge, demonstrate Albacore's remarkable forecasting performance out-of-sample. We find that Albacore$_\text{\scriptsize ranks}$\xspace shows upward pressures on inflation as early as mid-2020. Albacore$_\text{\scriptsize comps}$\xspace, on the other hand, is not as timely for the initial surge but captures the turning point earlier (already after its peak in 2022m3). Albacore$_\text{\scriptsize ranks}$\xspace remains more persistent during 2022 but indicates a faster disinflationary process thereafter. It lands at the lowest level (at 2.6%) compared to the other measure at the end of our sample.
\vskip 0.2em {\sc Comparing Weights.} The lower panels of Figure (ref) reflect the importance of the different subcomponents (left panel) and ranks (right panel) for the aggregate. For the left panel, level 6 estimates have been re-aggregated back to level 2 for ease of communication. In line with the official core inflation rate, Albacore$_\text{\scriptsize comps}$\xspace excludes energy. It assigns lower weight to food goods as well as clothing, housing, and recreational goods compared to the official headline rate. Unlike core PCE, which completely turns off food goods, Albacore$_\text{\scriptsize comps}$\xspace shrinks it by half. As we will see in Section (ref), only if we wish to specialize Albacore$_\text{\scriptsize comps}$\xspace for predicting low inflation risk will it be desirable to completely exclude food goods.
Services, on the other hand, get at least the weight they would get in the official series, with health and other services being even more important. Given that prices of services tend to be rather persistent and less prone to transitory shocks, they prove valuable in indicating the medium-term developments in inflation bils2004sticky. Moreover, it is plausible that upweighting these low-volatility components compensates for the inclusion of more volatile ones bearing leading signals, like food goods.
The high-dimensional regression setup allows to investigate the importance of components at the disaggregated level, which brings to light some eye-catching elements. For better or worse, beer is the top-weighted component in food goods. In the group of vehicles, we find a significant role of the secondary market with increased weights for used cars and trucks prices. Personal computers, tablets, and television gain in importance when it comes to recreational goods. So do medical insurances and prescription drugs for goods and services regarding health. Downsizing happens for shoes and footwear, garments, gasoline, sporting equipment and air transportation. At first sight these components do not have much in common. However, analyzing the persistence of each component's inflation rate reveals that upweighted items tend to yield higher persistence coefficients than downweighted ones.\footnote{We follow the literature on persistence weighted core inflation rates and use an AR(1) model to classify each component cutler2001core,BilkeStracca2007.}
Looking at the weights of Albacore$_\text{\scriptsize ranks}$\xspace reveals a highly asymmetric trim (see lower left panel in Figure (ref)). \textcolor{black}{Compared to the asymmetric trimming solution of the FRB Dallas' Trimmed Mean PCE (see green shaded area in Figure (ref)), Albacore$_\text{\scriptsize ranks}$\xspace is much more aggressive in removing the lower part of the distribution.\footnote{Note that in commonly used trimming approaches subcomponents are reweighted with their official PCE weights, while Albacore$_\text{\scriptsize ranks}$\xspace is not.} } In Table (ref) (in the appendix), we see this helps in reducing both volatility and bias. The absence of the sum-to-one constraint in Equation (ref) is instrumental to this result. By allowing rank weights to sum to any positive number, Albacore$_\text{\scriptsize ranks}$\xspace can leverage leading signals from a segment of the price growth distribution (here, centered around the 75$^{\text{th}}$ percentile) which by construction, grows in every period at a faster pace than 2%, and multiply them by a constant smaller than 1 to bring back the index to having the same long-run mean as headline inflation. Note that this focus on the upper tail does not prevent Albacore$_\text{\scriptsize ranks}$\xspace from flipping sign and entering deflation territory if need be. This, however, would require for approximately 75% of the level 6 components' prices to contract.
The benefits of highly asymmetric trimming can only be reaped if the shape of the price growth distribution is itself asymmetric and time-varying. Otherwise, there cannot be any comparative advantage versus simply monitoring (robust) location and scale. The pandemic-era inflation episode affected the skewness of the short-term distribution of price changes, which shifted from negative to positive skewness. As a results, existing trimming approaches understated (early) trends rich2022corebias, while Albacore$_\text{\scriptsize ranks}$\xspace is particularly well equipped to catch those. Moreover, positive price changes are typically more persistent than negative ones. As shown by the extensive literature on price setting behaviors of firms, prices tend to be adjusted faster when costs increase than when they decrease ball1994asymmetric,nakamura2008five,gautier2022new. Thus, Albacore$_\text{\scriptsize ranks}$\xspace significantly upweights relative price changes that are more likely here to stay.
\vskip 0.2em {\sc Analysis of the Post-Covid Inflation Surge.} We provide an in-depth analysis based on a decomposition of Albacore, the official core inflation rate and the trimmed mean (Figure (ref)). We focus on the contributions of energy, food, goods, services, shelter, and others, where the latter is comprised of transportation and health. A detailed description of the included categories in each group is given in Appendix (ref).
In the US, a large bulk of the early acceleration in inflation can be attributed to goods prices. The pandemic-induced distortions in several key sectors combined with strong aggregate demand stemming from shifts in consumer spending, accommodating monetary policy, and fiscal stimulus led to high inflation and contributed to its exceptional nature ball2022understanding,guerrieri2022macroeconomic,blanchard2023caused,digiovanni2023quantifying. Clearly, Albacore$_\text{\scriptsize ranks}$\xspace benefits from its focus on the upper half of the distribution. Contrary to Albacore$_\text{\scriptsize comps}$\xspace and core PCE, there is no downward pressure from goods inflation in 2020. Instead, Albacore$_\text{\scriptsize ranks}$\xspace already spots signs of upward tendencies in mid-2020 that build up as more and more sectors face negative impacts from the mix of supply chain disruptions and sustained consumer spending. This stands in contrast to the trimmed mean, which is lagging. Contributions from goods prices are muted until mid-2022 and remain comparably small throughout the sample. Instead, it attaches high weight to shelter; a notoriously lagging indicator bryan2010some,cotton2023forecasting. \textcolor{black}{The stock of rental leases included in the index typically captures new and existing rentals, with the latter being sluggish in adjusting to market rents. Based on these grounds but without a priori excluding shelter components, as in US supercore inflation powell2022,leduc2023will, Albacore$_\text{\scriptsize ranks}$\xspace benefits from the minimal contribution of shelter to the aggregate, on account of its highly asymmetric trimming scheme. }
Although, a substantial part of the pandemic-era inflation surge also reflects shocks in food and energy prices blanchard2023caused,gagliardone2023oil, their contributions in all core inflation measures are small. Contrary to core PCE, neither Albacore$_\text{\scriptsize comps}$\xspace nor Albacore$_\text{\scriptsize ranks}$\xspace exclude food or energy prices a priori, and yet, they both assign low weights to corresponding components, resulting in negligible contributions from energy and small, positive ones from food (given their exceptionally strong dynamics).
The deceleration of core inflation metrics in 2022 reflects fading supply chain disruption, accompanied by a moderation of goods inflation alcedo2022commerce,gagliardone2023oil. All measures, except for the FRB Dallas' Trimmed Mean PCE, peak in 2022m3. Yet, Albacore$_\text{\scriptsize comps}$\xspace is the only measure indicating a clear turning point with goods inflation significantly decreasing thereafter. Core PCE as well as Albacore$_\text{\scriptsize ranks}$\xspace are persistent throughout 2022 before trending downward towards the end of the year. While goods inflation is decreasing in importance, pressures from the labor market become more dominant amiti2022pass,benigno2023pc,blanchard2023caused. This is indicated by the rise of services inflation and its high persistence (since service sectors are typically labor-intensive), which is mainly observed for Albacore$_\text{\scriptsize comps}$\xspace and core PCE. Given that components related to services are getting higher weights in Albacore$_\text{\scriptsize comps}$\xspace, it proves to be a useful indicator for the second phase of the post-pandemic inflation era leduc2023will.
\vskip 0.2em {\sc Robustness to Moving Average Choice.} The choice of 3 months averages is motivated as being the maximal amount of time series averaging one can indulge in when building a measure of current conditions in a forward-looking manner. Nonetheless, as a robustness check, we present results for different moving average transformations of the regressors (from month-over-month to year-over-year) in Tables (ref) and (ref) in the appendix. As one would expect, \textcolor{black}{overall, we observe} that longer trailing averages perform better during calmer periods and worse during the post-2020 era. However, improvements of year-over-year averages, when applicable, are rather quaint vis-à-vis the mechanically more timely 3 months average. \textcolor{black}{Interestingly, Albacore$_\text{\scriptsize ranks}$\xspace yields remarkably stable performance across transformations (from month-over-month to year-over-year) for both test sets, as long as the level of disaggregation is deep enough ($K \geq 50$). So does Albacore$_\text{\scriptsize comps}$\xspace for all cases but the highly volatile month-over-month transformation for the pre-Covid sample.} This suggests that when there is enough leeway for variance reduction in the "cross-section" \textcolor{black}{of components}, additional time-smoothing has limited effects. We see this as desirable, as averaging over long windows by necessity, from a forward-looking perspective, errs on the increased bias side of the time aggregation bias-variance trade-off. Variation in benchmarks performance is limited before 2020 and larger thereafter, with a general preference for less timely averages. All in all, we find Albacore$_\text{\scriptsize comps}$\xspace and Albacore$_\text{\scriptsize ranks}$\xspace in their original 3 months average setup to fare well overall, even in the presence of a deluge of benchmarks and alternative specifications.
In this section, we construct Albacore for the euro area. Again, we compare the model's forecasting performance with that of the most prominent core inflation series (Table (ref) in the appendix), provide shape and features of the proposed measure (Figure (ref)) and delve into the recent surge (Figure (ref)).
\vskip 0.2em {\sc Forecasting Performance.} Similar to our findings for the US, benchmarks are more difficult to beat in the short run. Table (ref) reveals that the convex combination of existing core inflation measures ($\boldsymbol{X}_t^{\text{bm+}}$) features competitive predictive power that either Albacores can match, but struggle to surpass. For medium-term forecasts, we find that Albacore improves upon its competitors and yields the highest predictive accuracy for both evaluation samples. However, it does so by smaller margins than it did for the US, which is not surprising given that a more significant part of EA headline inflation was attributable to, from the perspective of the model, fundamentally unpredictable events. For short horizons, Albacore$_\text{\scriptsize comps}$\xspace tops the list with the lowest level of disaggregation (level 2). When it comes to predictive accuracy for medium-term horizons, we find larger gains for Albacore$_\text{\scriptsize ranks}$\xspace assembling a higher number of components (level 4).
\vskip 0.2em {\sc Comparing Weights.} In light of the results discussed so far, our in-depth analysis will be based on $h=12$, the most granular level (i.e., level 4), and the training sample ending in 2019m12. We find that both Albacore$_\text{\scriptsize comps}$\xspace and Albacore$_\text{\scriptsize ranks}$\xspace show a rather smooth path, moving close to the 2% target. In response to the two major crises in our sample, the GFC and the sovereign debt crisis, both series follow a downward trend with Albacore$_\text{\scriptsize comps}$\xspace indicating this drift earlier and moving below target for several years. Albacore$_\text{\scriptsize ranks}$\xspace remains slightly higher (close to target) for the low inflation era preceding the Covid-19 crisis. From 2020 onward, Albacore$_\text{\scriptsize ranks}$\xspace again showcase its quality as a leading indicator. It does not give in to the deflationary tendencies at the onset of the pandemic and heeds high inflation warnings at an early stage. Even though, Albacore$_\text{\scriptsize comps}$\xspace follows the common downward trend in 2020, it catches up quickly and both Albacore series plateau end-2022 (from 2022m12 to 2023m2), providing timely indication of the slowdown.
Similar to the US, Albacore$_\text{\scriptsize comps}$\xspace assigns low weight to energy and only half the weight to food and alcoholic beverages compared to the official HICP aggregate (see lower left panel in Figure (ref)). Components getting higher weights are found in housing (in particular, rents, services and goods related to the dwellings), other services, and communication, which all feature persistent behavior. Albacore$_\text{\scriptsize ranks}$\xspace shows an asymmetric trim tilted to the upside of the distribution (see lower right panel in Figure (ref)). This prevents Albacore$_\text{\scriptsize ranks}$\xspace from decreasing in the aftermath of the initial Covid-19 shock and signals upward pressures before all other core inflation measures do.
\vskip 0.2em {\sc Analysis of the Post-Covid Inflation Surge.} We focus on Albacore versus HICPX and the 30% trimmed mean and provide the contributions of energy, food, non-energy industrial goods (NEIG), and services to the aggregate in the decomposition exercise.\footnote{These groups are based on the “Special Aggregates” as suggested by Eurostat in the COICOP classification. Details are given in Appendix (ref).} Commodity prices, most importantly food and energy, accounted for the majority of the post-pandemic inflation acceleration in the EA. The standard argument that large swings in food and energy prices are mainly due to transitory shocks and tend to be quickly reversed, did not apply for the exceptional periods in 2021/2022. Since many existing core inflation measures ignore commodity price shocks by construction, they captured very little of the sizable and persistent increase in prices that ensued. For both Albacore$_\text{\scriptsize comps}$\xspace and Albacore$_\text{\scriptsize ranks}$\xspace we find positive contributions of food components throughout the sample (see Figure (ref)), which aligns with, e.g., peersman2022foodinflation, highlighting the importance of global food commodity price shocks for euro area inflation dynamics in the medium term.
However, focusing solely on energy and food would be too narrow nickel2022inflation,ascari2023euro,banbura2023drives. Soaring goods inflation as a response to constrained supply and excess demand, put sizable upward pressures on inflation. Albacore$_\text{\scriptsize comps}$\xspace captures the increase via high weights on housing (including heavily affected durable goods) and transport goods, while Albacore$_\text{\scriptsize ranks}$\xspace accounts for the upward pressure due to its focus on the middle-upper part of the price distribution. The downward trend indicated by Albacore from 2023m1 onwards forestalls that of HICPX and the trimmed mean. It is based on the easing of energy price pressures as well as goods inflation. Services inflation, on the other hand, remains strong for all measures under study, a reflection of mounting domestic price pressures banbura2023drives. For Albacore$_\text{\scriptsize comps}$\xspace, these persistent pressures mainly stem from rents, restaurants, and hotel services, as well as other services (e.g., insurance and financial services).
A natural worry with trimming-based measures is that an item that we would deem important in component space, shows up rarely in the trimmed indicator. This can happen if an important component separates from the pack (e.g., rent in Canada for 2022-2023), potentially opening a rift between trimming-based (or median) inflation and headline. We have seen, from earlier derivations, that $$ \pi_{\text{ranks},t}^* = \hat{\boldsymbol{w}}_r'\boldsymbol{O}_t = \underbrace{{\hat{\boldsymbol{w}}_r' \boldsymbol{A}_t}}_{\boldsymbol{w}_{c,t}} \boldsymbol{\Pi}_{t} \quad \quad \text{and conversely} \quad \quad \pi_{\text{comps},t}^* = \hat{\boldsymbol{w}}_c'\boldsymbol{\Pi}_t = \underbrace{{\hat{\boldsymbol{w}}_c' \boldsymbol{A}_t^{-1}}}_{\boldsymbol{w}_{r,t}} \boldsymbol{O}_{t}$$ and thereby, we can study Albacore$_\text{\scriptsize ranks}$\xspace in components space and Albacore$_\text{\scriptsize comps}$\xspace in rank space -- as time-varying weights models.
\textcolor{black}{We report results in Figure (ref). In the left panel, we conduct the comparison in components space. This includes Albacore$_\text{\scriptsize comps}$\xspace, the official PCE rate, and Albacore$_\text{\scriptsize ranks}$\xspace. For the latter, we present $\boldsymbol{w}_{c,t}$ through its median (bullet) as well as the 16$^{\text{th}}$ and 84$^{\text{th}}$ quantiles (lines). Overall, we find that in the EA both Albacore versions are broadly consistent regarding weights assigned to subcomponents, while in the US the two models show a higher degree of disagreement. Most notably, Albacore$_\text{\scriptsize ranks}$\xspace frequently excludes elements of health and shelter, likely due to their stickiness and their concentration in the middle of the distribution bils2004sticky. As pointed out in Section (ref), this supports forward-looking qualities of Albacore$_\text{\scriptsize ranks}$\xspace since it devotes less attention to indicators with a rearview perspective.}
\textcolor{black}{Higher weights are assigned to recreation services and other durable goods. Both components exhibit high cyclical variation (where for the latter this is mainly attributable to jewelry and watches) and thereby provide early indication of domestic inflationary pressures dolmas2009excluding,stock2020slack. In the EA, we get similar results with recreational services and goods proving to frequently show up in the included part of the distribution. In line with findings in the US, their cyclical sensitivity is found to be strong and indicate price changes based on domestic business cycle movements frohling2011supercore. In both regions energy gets trimmed out most of the time given its tendency to be located in the tails.}
The right panels of Figure (ref) present Albacore$_\text{\scriptsize ranks}$\xspace, Albacore$_\text{\scriptsize comps}$\xspace, and the core PCE/HICPX rate in rank space. This, interestingly, provides a view of the implied trimming of components space-based estimators. As expected, Albacore$_\text{\scriptsize comps}$\xspace and the official core inflation rate cover the whole distribution of price changes over time, as they are not explicitly designed to remove a particular range of ranks. Notably, both solutions are left-skewed, with the extent of it being greater for the US. Thus, permanent exclusion choices from Albacore$_\text{\scriptsize comps}$\xspace and core PCE (like the absence of energy) implies a more aggressive downweighing of the lower part of the price growth distribution -- and a subtle focus on the 60-70 percentiles region. For the EA, Albacore$_\text{\scriptsize ranks}$\xspace accentuates the peak of the distribution at its original location, whereas, in the case of US, Albacore$_\text{\scriptsize ranks}$\xspace provides a concentrated mass located slightly to the left of the core PCE's highest weighted rank. Therefore, Albacore$_\text{\scriptsize ranks}$\xspace seizes the opportunity to significantly amplify what components-space measures are doing only in a modest fashion.
The composition of the maximally forward-looking core inflation measure will differ depending on the chosen forecasting horizon. In this paper, we have focused on horizons typically of interest for the conduct of monetary policy. But shorter (or longer) horizons are of great interest to many institutions. In this subsection, we investigate how Albacore$_\text{\scriptsize comps}$\xspace and Albacore$_\text{\scriptsize ranks}$\xspace definitions adjust as we vary $h$. As we can see in Tables (ref) and (ref), assemblage regression performs well at various forecasting horizons.
Figure (ref) reports weights of Albacore$_\text{\scriptsize comps}$\xspace and Albacore$_\text{\scriptsize ranks}$\xspace as $h$ increases from 1 to 24. In the left panel, we report Albacore$_\text{\scriptsize comps}$\xspace's weights (as a function of $h$) relative to those of headline PCE. First, let us focus on energy and food goods, which are excluded from prototypical core measures. We see a rapid decline in the importance of energy. Energy components get close to 0 weight across all horizons, being entirely excluded for higher-order forecasts. Food goods get initially upweighted at short horizons, before decreasing steadily and converging towards 0 for $h>18$ months. This suggests that food goods, despite their volatile behavior, are at least indicative of short-run price trends.
The services share of Albacore$_\text{\scriptsize comps}$\xspace is ever increasing in $h$. This is visible for other services, and also financial services which takes off after $h=12$. Some components experience the reverse pattern, being more critical in the short term and becoming less significant over longer periods. This symmetric lineup change is obviously inevitable by the sum-to-one constraint: if some components come in, others must go. Among the latter, we have other nondurable goods and transportation, which experience the largest relative increase for the first few forecasting horizons and then eventually decrease to the same level as the official rate (other nondurable goods) or below (transportation). \textcolor{black}{The rationale for the weight curves of transportation and other nondurables can be drawn from their connection to commodity prices. Among PCE categories, energy prices mainly influence items related to fuel and transportation, while higher food commodity prices (e.g., crop) feed through to nondurables based on their high input shares hobijn2008commodity,gao2014oil.\footnote{Note that crops are not only an important input for food goods but also for other nondurables since the latter include tobacco products, flowers, pots, and plants hobijn2008commodity.} This rather fast passthrough of commodity prices makes transportation and nondurables better predictors for near-term developments rather than longer-term trends. Overall, we trace the slope of the weight curves to the components' characteristics with more persistent components being relevant for higher-order forecasts and faster-growing and/or more volatile ones contributing to the predictability in the short run.}
The right panel of Figure (ref) presents the evolution of rank weights across horizons for Albacore$_\text{\scriptsize ranks}$\xspace by grouping them into 6 "buckets" for visibility. Results reveal that the lower percentiles are trimmed out early on, echoing the exclusion of energy in the left panel. The bucket covering the 34$^{th}$ to 50$^{th}$ percentile gets small positive weight for short horizons before approaching 0 from $h \geq 5$ onwards. Highest weights throughout are assigned to the 68$^{th}$ to 84$^{th}$ percentiles followed by the 51$^{th}$ to 67$^{th}$ percentiles. While the latter decreases in importance with increasing $h$, the former shows an upward sloping weight curve. Similarly, the right tail of the distribution gets positive weight, which grows in $h$ and eventually converges to nearly twice its share from $h=1$. This corroborates the view that its a-priori exclusion leads to a sizable loss of information about longer-run inflation trends.
Overestimating and underestimating core inflation have different costs. Therefore, it can be of interest to build core measures specialized at foreshadowing either of the two key risks: deflation and high inflation. We do so by running quantile assemblage regression, which, as the name suggests, now assembles components (or ranks) to minimize a quantile loss.\footnote{Following from a long tradition in financial risk management, it is now widespread to forecast tail risks to macroeconomic variables. adrian2019vulnerable initiated a literature using quantile regression methods for forecasting tail risks to the real activity outlook giglio2016systemic,ferrara2022high,delle2023modeling,clark2024gar. Of course, some have also build analogous models for inflation, which own palette of risks is of interest to central bankers, investors, governments, and others korobilis2017quantile,ghysels2018quantile,pfarrhofer2022modeling,lenza2023density.} Focusing on rare events inevitably leads to a reduced effective number of training observations. So that this number is not further reduced by significant overlapping of long averages for the target (e.g., $h \in \{12, 24\}$), we now turn our attention to 3 and 6 steps ahead forecasts instead. Moreover, we drop the 20 years rolling window in favor of an expanding window from 1990. We predict the 15$^{th}$ quantile for the lower tail risks and the 85$^{th}$ quantile for the upper tail risks (i.e., $\tau \in \{0.15, 0.85\}$) for level 6 in case of Qualbacore$_\text{\scriptsize comps}$\xspace and level 3 in case of Qualbacore$_\text{\scriptsize ranks}$\xspace. We modify Equations ((ref)) and ((ref)), respectively, to
where the quantile loss function is given by $\rho_{\tau}(u) = (\tau - \mathbbm{1}\{u \leq 0 \}) u$ with $\mathbbm{1}\{u \leq 0 \}$ defining an indicator function, which returns the value 1 if $u \leq 0$ and 0 otherwise and $u$ denoting the error term gneiting2011making. {$Q_{\tau}\left(\cdot\right)$ calculates the empirical value of the unconditional quantile $\tau$.} Some modifications to our original equality constraints are necessary, since they are not suitable for their quantile extensions. Summing weights to 1 in case of Qualbacore$_\text{\scriptsize comps}$\xspace and setting the long-run average of Qualbacore$_\text{\scriptsize ranks}$\xspace to that of headline inflation would prevent the algorithm from upweighting tail observations. Rather, we now want those constraints to hold up to a prefixed multiplicative constant $\left( \frac{Q_{\tau}({\pi}_{t+1:t+h})}{\bar{\pi}_{t+1:t+h}}\right)$, reflecting the difference in the long-run level of quantile $\tau$ versus the mean. Thus, we modify the equality constraint for Qualbacore$_\text{\scriptsize comps}$\xspace such that weights sum to our target's respective unconditional empirical quantile ($Q_{\tau}({\pi}_{t+1:t+h})$) relative to its sample mean. For Qualbacore$_\text{\scriptsize ranks}$\xspace, this boils down to its training sample average being equal to that of the respective unconditional quantile of headline inflation $Q_{\tau}({\pi}_{t+1:t+h})$ instead of $\bar{\pi}_{t+1:t+h}$ as used in Albacore$_\text{\scriptsize ranks}$\xspace.
We start our evaluation of Qualbacore by contrasting its forecasting performance to the main set of benchmarks in Table (ref). We present results for each quantile (i.e., $\tau_{15}$ and $\tau_{85}$) and include, for reference, the original loss function (i.e., MSE). Aligned with our main findings in Section (ref), Qualbacore$_\text{\scriptsize ranks}$\xspace yields a remarkable performance regardless of the forecasting horizon, quantile, and evaluation sample. Pre-Covid results reveal that gains achieved by Qualbacore$_\text{\scriptsize ranks}$\xspace are rather evenly distributed across the risk spectrum for $h=3$ and tilted to the upper tail of $h=6$. The post-2020 sample confirms that large improvements are to be found in the upper tail, while those for deflation or low inflation risk are negligible. We observe improvements by sizable margins for $\tau_{85}$, mounting to over 60% reduction in quantile loss for both forecasting horizons.
The predictive accuracy of Qualbacore$_\text{\scriptsize comps}$\xspace is lower, mostly in line with the benchmarks for the first evaluation sample. In contrast to Qualbacore$_\text{\scriptsize ranks}$\xspace, forecasting results of Qualbacore$_\text{\scriptsize comps}$\xspace in this period (as well as that of the remaining benchmarks) tend to be inferior in the upper tail. This finding reverses when analyzing our second evaluation sample, which encompasses the pandemic and post-pandemic periods. Qualbacore$_\text{\scriptsize comps}$\xspace and benchmarks perform well for the 85$^{th}$ quantile but fall short in the lower quantile. \textcolor{black}{This underperformance can be localized to the disinflationary shock in mid-2020 driven by the Covid-19 induced energy price shock.} In general, the tendency for substantial outperformance to be localized in the upper tail is not entirely surprising: key disinflationary events of the last 30 years are almost always due to unpredictable (and often momentaneous) oil price shocks.
To probe deeper, Figure (ref) presents the resulting time series of Qualbacore (right panels) as well as the corresponding weights of assembled components and ranks (left panels) for the 6 steps ahead horizon. Details for the shorter horizon ($h=3$) are relegated to the appendix (Figure (ref)). Studying the results from Qualbacore$_\text{\scriptsize comps}$\xspace's upper tail shows that food goods and food services get high weights as well as components related to transportation and other services. Items assigned lower weights are mostly related to services and durable goods (including clothing, house goods, and other durable goods). For the lower tail, the opposite holds true. Shelter, health, and financial services, as well as durable goods enter the spotlight while food goods and food services get almost no weight. Energy, comprising the most volatile items, is excluded from both tails. \textcolor{black}{Given the usual wisdom, one might have expected a similar outcome for food components. Yet, their assigned weights for the upper tail is nearly that of headline. As shown in the literature, underestimation of future inflation is often linked to positive food price shocks bowles2007ecb,cecchetti2008commodity. They tend to exhibit stronger second-round effects than energy price shocks and are more likely to propagate to headline eventually de2012commodity,de2016macroeconomic. Thus, it is less surprising that upweighting (with respect to core PCE) developments of food prices in the upper tail contributes to the predictive power of Qualbacore$_\text{\scriptsize comps}$\xspace.}
For Qualbacore$_\text{\scriptsize ranks}$\xspace, the resulting rank weights differ significantly when focusing on either the upper or lower quantile. $\tau_{85}$ yields a rather sparse solution centered at the upper part of the distribution of monthly price changes while the weights for $\tau_{15}$ are distributed densely covering most of the distribution except for the lowest part. Hence, when specializing our supervised trimming measure to disinflation/deflation risk, it is much more aligned with the specification of traditional trimmed mean inflation, as reflected by the shaded regions in Figure (ref). An even clearer instance of this phenomenon can be visualized in Figure (ref), where Qualbacore$_{\text{ranks}}(\tau_{15})$ is almost literally median inflation. This leads us to conclude that traditional trimming-based core inflation measures, with a sharp focus on the center of the price growth distribution, are better equipped to foresee low inflation, which was, from 1990 to 2020, seen as the main source of risk. Our methods can capture this equally well, but highlight that when balancing both types of risks within a single loss function (i.e., the MSE), deviating from putting the accent on midpoint ranks is preferable.
In this brief section, we explore additional inflation-related applications of the assemblage regression, with a particular focus on the rank space variant. We deviate from aggregating inflation subcomponents. First, given the widespread use of year-over-year measures, we consider aggregating lags into a supervised (trimmed) moving average of 12 months and compare it with \textcolor{black}{the usual headline rate in yearly growth rates} (Year-over-year, YoY). Second, we devise a forward-looking combination of country-level headline inflation rates in the euro area.
In both applications, we find that assembling in rank space is more successful than the more common components space version. This suggests that, when deciding on heterogeneous aggregation weights, it can be preferable to draw more attention to the location of a realization in the overall distribution rather than its exact position in time or space. This finding is obviously, for now, bounded to an environment where we are devising easily-interpretable forward-looking linear aggregations. More work is needed to see if this holds as strongly elsewhere, e.g., in more sophisticated data-rich nonparametric models or for other targets than inflation.
Year-over-year measures, or equivalently, 12 months averages of headline inflation are a common way to visually inspect price pressure metrics. In this section, we answer the following question: if one wishes to construct some average of the last 12 months of headline inflation that is most predictive of the next 6, 12, and 24 months, what would such an average look like? Given the focus on time-smoothing (with the prefixed 12-month window) and the usage of only autoregressive information, the outcome is conceptually closer to a trend inflation extraction, again, by acknowledging the connection to Hamilton2018's filtering.
We deploy assemblage regression in components and rank space. We focus on the US, and the “components” are now 12 lags of month-over-month PCE growth rates. \textcolor{black}{Of course, in components space, this is simply a constrained autoregression (AR).} The interest rather lies in the rank version, which, instead of permanently upweighting or downweighting lags, performs temporary exclusion. Its usage of past realizations of the time series differs from that of a standard autoregression. \textcolor{black}{The latter orders these lagged observations based on the distance in time between the observation and the target, and allows for corresponding heterogeneous weights.} Instead, the rank version fixes a lookback window (here, 12 months) and weights past realizations based on their location in the window's realized distribution. \textcolor{black}{We thereby shift} the focus away from timing with potentially sophisticated dynamics with lag-specific coefficients to consider each lag as a draw and ask ourselves which average of such draws should we be paying attention to. The elements discussed in Section (ref) apply, {implying} in this context that the order statistics-based autoregression is a time-varying autoregression in lags space.
The usage of the term "realized" to describe month-to-month inflation rates here is deliberate, as what we are building is a summary statistic from high(er)-frequency data designed to be used as a predictor for a lower frequency target. Thus, there is a natural connection to the construction of realized volatility (RV) from intra-day data in finance andersen2003modeling,mcaleer2008realized. A key methodological distinction is that here, the nature of the statistic itself is estimated via rank coefficients rather than being predefined (and proven) to be a consistent estimator for a variance, a quantile dimitriadis2022realized, a mean, or else. The theoretical environment backing RV is that true volatility is fundamentally latent and can be estimated consistently from high-frequency data under various assumptions. Inevitably, we are within the realm of unsupervised learning because the desired supervisor's "true volatility" is unobserved.\footnote{Nonetheless, in the spirit of this paper, it is thinkable that such measures could be defined in a "supervised" fashion by maximizing usefulness in predicting another observed variable (or maximizing general economic value) with rank weights being shrunk to what would be implied by traditional RV.} In our rank AR inflation context, there is no quest for a particular summary statistic\footnote{For instance, if inflation were hypothetically defined as a yearly process, the intra-year squared observations could not be averaged directly to build a variance estimator because month-to-month observations are not uncorrelated mcaleer2008realized.}, but rather, a search of any that will fulfill the empirical task of being closely associated with future headline inflation, as measured by statistical agencies. Validity comes from predictive power holding both in- and out-of-sample.
Table (ref) presents the performance of the ranked (AR$_{\text{ranks}}$) and constrained (AR$^{\text{+}}_{\text{lags}}$) autoregression compared to a simple AR$_{\text{lags}}$(12) with monthly growth rates, an AR(1), and a random walk, both using year-over-year data (the latter as the numéraire). During periods of low and stable inflation, putting the emphasis on persistence is an established strategy clark2014evaluating,StockWatson2016. Accordingly, AR$_{\text{ranks}}$ procures good performance during the pre-2020 sample by leveraging little information. What is more striking is that we find decisive improvements from AR$_{\text{ranks}}$ over AR alternatives for medium- and long-run horizons -- and for both out-of-sample periods. At horizon 3, AR$_{\text{ranks}}$ and AR$^{\text{+}}_{\text{lags}}$ report similar performance, with $\approx 5 \%$ improvements over the numéraire. Only at $h=1$ does AR$^{\text{+}}_{\text{lags}}$ prevail marginally, suggesting that the gains of paying attention to timing with lags (as apposed to distributional aspects with AR$_{\text{ranks}}$) are circumscribed to very short horizons.
The successful AR$_{\text{ranks}}$ uses a downweighted average of the 6 highest-valued lags, \textcolor{black}{i.e. those 6 out of the 12 lags that show the highest month-over-month increase} (see lower right panel in Table (ref)). The inclusion of several lags equips the measure with adequate smoothness, \textcolor{black}{while} the exclusion of lower realizations allegedly mitigates positive "base effects" following negative price shocks. The resulting average puts a high bar for the index to show signs of lasting deflation, as it would require that order statistics corresponding to ranks $R_5$ and above showcase negative growth, implying that a large number of (nearly) consecutive months witnesses deflation. This is consistent with recent history: we find episodes of deflation (2009) or near deflation (2020) in YoY due to a handful of months showing abnormal negative shocks, but those quickly reversed. Depending on the precise event itself, one may attribute this reversal to the very nature of these (oil) shocks or the expansionary monetary policy often triggered in the aftermath. Whatever the cause might be, acknowledging that such shocks have little predictive power for headline inflation conditions in 6, 12, or 24 months, AR$_{\text{ranks}}$ trims them out. Disinflation is not as exceptional of an outcome for AR$_{\text{ranks}}$ as deflation is. If the upper order statistics showcase lower growth than they usually do (by construction they are trending above the target), then AR$_{\text{ranks}}$ will turn in a disinflation forecast, as is the case for 2014 to 2021. This asymmetric outcome is not a major surprise given our main results in Section (ref). The prevalence of temporary negative shocks with no predictive power for future headline inflation can be identified among components at a fixed point in time, or \textcolor{black}{alternatively}, by observing repeatedly an aggregate with a time-varying level of exposition to those \textcolor{black}{negative shocks}.
While AR$_{\text{ranks}}$ does well in our post-2020 sample and is clearly the best {among the models using this limited information set}, it is no match for those using components-level data (Table (ref)). 12 months proves too long of a lookback window and solely focusing on autoregressive information is insufficient, motivating the usage of the more involved Albacore$_\text{\scriptsize ranks}$\xspace of Section (ref), which fares well in all regimes. Thus, we find AR$_{\text{ranks}}$ and Albacore$_\text{\scriptsize ranks}$\xspace to share similarities before 2020, but less so thereafter. Nonetheless, given its decisive merits within the class of data-poor models, the autoregressive inflation model in rank space is a handy tool to keep in one's arsenal. It may be of even greater interest for countries where component level data is less reliable or not available for an extended period.
Although sharing the same currency, and thus a single monetary policy, inflation rates of EA member states are far from being homogenous. As a result, the predefined weighted average of HICP rates forming the EA aggregate may or may not be the most timely indicator for the whole area $h$ months from now. A natural curiosity that arises thereof is whether an alternative aggregation of member states inflation rates forms a better leading indicator for HICP than HICP itself. To answer this question, we introduce Albacore$^\text{G}$\xspace, a geographic assemblage of Albacore, with the two representatives Albacore$^\text{G}_\text{\scriptsize ranks}$\xspace and Albacore$^\text{G}_\text{\scriptsize comps}$\xspace. Different from Albacore$_\text{\scriptsize comps}$\xspace, which assembled the subcomponents of HICP in the predictor matrix, $\boldsymbol{\Pi}_{t}$ is now composed of the headline inflation numbers of each of the $M$ EA member states: $\boldsymbol{\Pi}^G_{t} = \left[HICP^1, ..., HICP^M\right]$. Similar to Albacore$_\text{\scriptsize ranks}$\xspace, the order statistics matrix $\boldsymbol{O}^G_{t}$ for Albacore$^\text{G}_\text{\scriptsize ranks}$\xspace, is then just the ranked version of $\boldsymbol{\Pi}^G_{t}$.
The geographic versions of Albacore$_\text{\scriptsize comps}$\xspace and Albacore$_\text{\scriptsize ranks}$\xspace are designed as follows: the weights of Albacore$^\text{G}_\text{\tiny comps} $ are shrunk towards the annual expenditure shares for each country.\footnote{Expenditure shares are the percentage of total household final monetary consumption expenditure based on national accounts data EURegulation20201148. Data is taken from Eurostat (prc_hicp_cow).} Weights are not constrained to sum to unity, dropping the random walk hypothesis is hardly appropriate when only combining headlines. Sum-to-one can be restored by interpreting the resulting (lower) sum of weights as an AR(1) coefficient, which, being smaller than 1, reduces variance. The loosening of this constraint, which was introduced to discipline the model when assembling a large number of components, prevents the solution from converging to HICP itself. Furthermore, we demean each country rate and add back to it the EA average (i.e., a country fixed effect). Variance reduction in Albacore$^\text{G}_\text{\scriptsize ranks}$\xspace is obtained through different channels, and thus, neither of these modifications to our baseline assemblage regression proposition is needed. It keeps the fused ridge penalty, while constraining the resulting conditional mean to match the unconditional mean of our target.
From Table (ref), it is apparent that Albacore$^\text{G}_\text{\scriptsize ranks}$\xspace trumps all competitors in both subsamples and across all forecasting horizons, whereas the component version is mostly aligned with the performance of the best core benchmarks. We show corresponding time series and the weights for Albacore$^\text{G}_\text{\scriptsize comps}$\xspace and Albacore$^\text{G}_\text{\scriptsize ranks}$\xspace in the panels above. Similar to Figure (ref), Albacore$^\text{G}_\text{\scriptsize ranks}$\xspace anticipates the broiling inflationary pressures at a much earlier stage than HICPX or the trimmed mean. Furthermore, it calls the peak and the following downward trend precisely. Albacore$^\text{G}_\text{\scriptsize ranks}$\xspace showcases strong asymmetric trimming: the three lower quartiles are entirely discarded, and weight is given only to the very right tail of the distribution. While it might appear like a radical outcome, we find reasonable arguments from both a statistical as well as an economic point of view. From a time series perspective, one should keep in mind that we are trimming highly aggregated indices, which feature much less dramatic movements in the tails than subcomponents-level series. Empirically, results suggest that higher inflation rates across countries are more likely to be contagious and generalize to the whole area. This also means that low inflation will be expected in the 6, 12, or 24 months only if the highest inflation rates over the last 3 months are lower than they typically are. Alternatively, this may be linked to tails reacting more quickly to shocks and an intensified spillover of inflation during turbulent times, as is found in the literature aharon2022infection,bouri2023global. Thus, regardless of the size of EA member states, their expenditure intensity, or other country-specific characteristics, a forward-looking approach may want to pay attention to those with higher inflation rates at a given point in time. In particular, as a thumb rule, a downweighted average of the 5 highest headline inflation rates among member states is found to be a good tracker of forthcoming EA-wide price pressures.
In terms of RMSE, Albacore$^\text{G}_\text{\scriptsize comps}$\xspace outperforms the official HICP for longer horizons, but not the rank version discussed above. It moves rather in lockstep with the other core measures (HICPX and trimmed mean) at the outset of the pandemic-induced recession, and is late to calling the mounting inflationary pressures. For the turning point, it is close to Albacore$^\text{G}_\text{\scriptsize ranks}$\xspace and resembles that of a compressed headline. This behavior can be traced back to the allocation of country weights, which deviates substantially from the official ones (see Table (ref)). Albacore$^\text{G}_\text{\scriptsize comps}$\xspace essentially overweights Italy, Finland, Portugal, and Greece, and keeps France and the Netherlands in the set with close to official weights. While Greece and the Netherlands experienced the post-pandemic inflationary pressures late, they were among the first member states to start the disinflationary process. France and Finland, on the other hand, contribute to the smoothness of the series given their comparably moderate increases. At first sight, it might seem surprising to see Albacore$^\text{G}_\text{\scriptsize comps}$\xspace excluding Germany, which constitutes the biggest EA economy with most influential trade relations auer2019international,tiwari2019analysing. However, when taking on a forward-looking perspective, inflation rates of smaller (or peripheral) countries may provide relevant signals since they are more susceptible to shocks. In our sample, Greece stands out as it was heavily affected by the GFC, triggering the subsequent sovereign debt crisis. It experienced low (and negative) inflation before the whole area itself, which likely contributes to its high weights. This is, obviously, quite contextual to the training sample including 2010s, and highlights why Albacore$^\text{G}_\text{\scriptsize ranks}$\xspace, with its temporary inclusions and exclusions scheme, fares better overall than Albacore$^\text{G}_\text{\scriptsize comps}$\xspace, which focuses on a fixed portfolio of countries.
We introduced the assemblage regression to build a maximally forward-looking measure of core inflation. While predictive power for future headline inflation is generally an afterthought for the construction of core inflation measures, our Albacore$_\text{\scriptsize comps}$\xspace and Albacore$_\text{\scriptsize ranks}$\xspace are explicitly optimized to perform well with respect to this criterion. {Both yield favorable forecasting results coupled with economic insights of high relevance.} The most striking {finding} is the asymmetric trimming obtained from Albacore$_\text{\scriptsize ranks}$\xspace, which is particularly pronounced in the US. This allows Albacore$_\text{\scriptsize ranks}$\xspace to perform well in quieter times and capture early signs of the post-pandemic surge months in advance of trimmed mean and other alternatives. {In terms of supervised weighting of components, we find a special role assigned to food goods. It is not totally excluded from Albacore$_\text{\scriptsize comps}$\xspace (as typically done in core PCE), and when it comes to upside inflation risks its overall weight even matches that of headline. These results, and additional insights from our quantile regression extension, suggest that the core inflation measures we should monitor to flag disinflation/deflation risks may differ from what is needed to detect the risks that came to materialize from 2021 onward. }
There are numerous avenues for future research. First, and closest to the analysis conducted in this paper, there are a few natural extensions to Albacore that one may wish to investigate. While we focused on aggregating components in a single space at a time, one may think of combining components and rank selection using more flexible nonlinear optimization and a layered approach (e.g., a structured neural network). Alternatively, one could consider a "reservoir" approach tanaka2019recent, which would boil down to an extremely high-dimensional ridge regression.
Second, assemblage schemes considered in \textcolor{black}{our work} are built to fulfill a set of simple statistical requirements. One can depart from "forward-looking" \textcolor{black}{objective}, or, at least, include a broader set of \textcolor{black}{objectives}. One possibility is using MACE's algorithm to construct a core inflation indicator that is maximally reactive to monetary policy and real activity. Combining these supervised aggregation ideas with the forward-looking criterion could lead to {maximizing} the overall coherency in a vector autoregressive setting (i.e., its likelihood). In the latter case, this would imply that the assembled series would be both highly reactive to key aggregates and predictive of them. This could involve assembling a wider basket of indicators -- like GDP, which is seldomly exposed to sectoral shocks that are a distraction from the general macroeconomic trend.
Third, the econometric idea of running assemblage regression in rank space to build a maximally predictive summary statistic can be exported beyond inflation applications. \textcolor{black}{For example, supervised trimmed forecast combinations might extend trimming methods used in the literature, which only involve choosing the endpoints in a similar vein to trimmed mean inflation (see wang2023forecast for a survey).} While traditional regressions in forecasters space upweight and downweight individuals based on their past skills, running it in rank space allows to detect which section of the crowd's distribution has more wisdom (whatever its time-varying composition in terms of individuals is). Regressions using order statistics may also be better equipped to deal with the inevitable entry and exit of forecasters from the panel. Another, more operationally ambitious avenue, is supervised temporal aggregation of intraday financial returns data to maximize the economic value derived from the resulting series. Possible outcomes, depending on the observable target, include some realized volatility series, one its many refinements and extensions mcaleer2008realized, a quantile, or something else entirely.
All in all, the supervised aggregation ideas developed in this paper are exportable to a variety of settings where the official aggregate series (or manual alterations of it) offers room for improvement. This is, essentially, the foundational idea of deep learning. Rather than manually building predictors from disaggregates, one can endogenously construct the features from the input through multiple layers which are jointly optimized along with the predictive model using them. In an econometric modeling environment, there are obvious constraints on such constructions, so that the extracted features make economic sense. Nonetheless, this paper has shown -- within the context of an elementary forecasting model -- that such cohabitation is possible.
\setlength\bibsep{5pt}
\setstretch{0.75}