Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
526,726 characters · 67 sections · 347 citation commands
\pagenumbering{roman}
\chapter*{Abstract}
Effective economic policymaking requires timely assessments of current economic conditions rather than relying solely on past-quarter data, especially in a dynamic and open economic system like Singapore. Promptly identifying turning points becomes critical, considering how international fluctuations can swiftly influence domestic conditions. This real-time assessment, often called "nowcasting," is particularly challenging in rapidly changing economic conditions.
The econometric approach for GDP growth nowcasting has dominated the forecasting landscape for years. Dynamic Factor Models (DFMs) applications in macroeconomic studies have guided substantial research on reducing high-dimensional datasets into a few latent factors that evolve over time. Nonetheless, these factors may introduce noise when distilling key signals from a vast pool of predictors, and they may not consistently guarantee the most substantial forecasting capabilities—ultimately affecting model interpretability.
Traditional forecasting frameworks often encounter notable difficulties in handling structural breaks and crisis periods, emphasizing the need for flexible methods suited to volatile economic settings. Machine Learning (ML) models offer a highly effective path toward more accurate GDP growth nowcasting. Because these methods can approximate intricate nonlinear relationships, they frequently surpass older econometric techniques. In addition, recent advancements in explainable artificial intelligence (xAI) grant policymakers and researchers a clearer perspective on the internal mechanics of these models, enabling more targeted policy decisions. By integrating sub-period analyses, researchers can capture the heterogeneity of economic shocks, enhance real-time surveillance, and refine policy responses for smaller, highly exposed markets.
This research investigates how ML algorithms can improve the accuracy of nowcasting Singapore's GDP growth while clarifying the reasoning behind forecasts. Specifically, we introduce a structured pipeline designed to mitigate look-ahead bias, estimate model reliability via block bootstrap intervals, and highlight key features driving predictions. We also incorporate model combination techniques—such as simple averaging and weighted approaches—to boost forecast stability and monitor the dynamic weights assigned to individual models, thereby offering an interpretive lens for understanding which learners dominate at different points in time. These strategies are complemented by sub-period breakdowns that evaluate performance under varying market phases, revealing a broader perspective on robustness.
We analyze a set of approximately seventy variables, encompassing economic fundamentals (such as prices, production, trade, and investment), demographic factors (like population growth and age distribution), and societal indicators (e.g. employment rates). Data from the Singapore Department of Statistics and additional sources, including The Bank of Italy for external reference indices, span from 1990 Q1 to 2023 Q2. We implement penalized linear models (Lasso, Ridge, Elastic Net), dimensionality reduction approaches (Principal Component Regression, Partial Least Squares), ensemble learning methods (Random Forest, XGBoost), neural architectures (Multilayer Perceptron, Gated Recurrent Unit). Benchmarks such as a naïve Random Walk, an AR(3), and a Dynamic Factor Model are also included for comparison. A Bayesian optimization procedure is employed to select optimal hyperparameters within an expanding-window walk-forward validation framework, accommodating the temporal dependencies of macroeconomic time series.
Our findings provide reassurance about the stability and reliability of ML-based models in economic forecasting. These models consistently outperform benchmark forecasts, including official projections by major financial institutions, and demonstrate resilience across crisis episodes. On average, penalized linear models, dimensionality reduction approaches, and neural networks consistently achieve substantial reductions in nowcast errors—typically ranging from 40% to 60% compared to established benchmarks such as Random Walk, AR(3), and Dynamic Factor Models. When these strong learners are combined, forecasts achieve further improvements in nowcast errors, as shown by performance metrics.
In addition to GDP growth nowcasting, we employ explainability approaches appropriate to each family model—such as coefficient evolution, Variable Importance in the Projection, Gini-based and permutation-based measures, and Integrated Gradients—to enhance interpretability. By combining these techniques with block bootstrap procedures, we pinpoint which features exert the most significant impact on GDP variations and gauge the inherent uncertainty surrounding them—thereby illustrating, for instance, how changes in foreign trade components or prices can trigger significant shifts in growth projections. Although xAI techniques are gaining traction in fields like image recognition or consumer analytics, they remain relatively unexplored in macroeconomic contexts. This approach enables policymakers to have a deeper understanding of how shocks traverse the economic landscape.
Our study contributes to the broader discourse on ML-based macroeconomic forecasting by showing how sub-period analysis, uncertainty quantification, combined models, and xAI can operate in tandem. Through this integration, we achieve more substantial predictive accuracy and foster greater confidence and clarity in the outputs—factors of particular value for small-scale, internationally interconnected economies. Future work could further refine our methodology by encompassing high-frequency data or trialing innovative ML models that reinforce our framework's versatility and reliability.
\thispagestyle{empty} \cleardoublepage
\chapter*{Acknowledgments}
This thesis was made possible thanks to the institutional and personal support I received over the past years.
I am deeply grateful to the Università degli Studi di Macerata for funding and believing in this research project, and to ESSEC Business School, Asia-Pacific, for hosting me as a Visiting PhD Fellow and providing an intellectually stimulating environment in Singapore.
I wish to thank Dominique Lepore (Università Mercatorum) for immediately believing in my initial idea and for her invaluable advice in transforming it into a coherent research project. I am also grateful to the professors at Università Politecnica delle Marche---Maria Cristina Recchioni, Marco Gallegati (my master’s thesis supervisor), Emanuele Frontoni, Donato Iacobucci, and Mauro Gallegati---for their support and for the references that made my doctoral application possible.
A special word of thanks goes to my year in Hong Kong in 2019: Andrea Croci and the late Eric Szeto were true mentors and role models, and they profoundly shaped my relationship with the Asia-Pacific region.
My heartfelt thanks go to my doctoral supervisors, Rosaria Romano (Università degli Studi di Napoli Federico II) and Jamus Jerome Lim (ESSEC Business School Asia-Pacific), for their guidance, patience, and constant support throughout this work. I am equally indebted to Pierre Alquier (ESSEC Business School Asia-Pacific) and Andrea Bucci (Università degli Studi di Macerata) for their mentorship and expert guidance on econometrics and machine learning.
This thesis was conceived and developed as an autonomous research endeavor, with a strong orientation toward real-world applicability: a robust approach to nowcasting, a focus on uncertainty quantification, and interpretable methods that can be transferred beyond academia. I am grateful to all those, in Italy and in Singapore, who have supported me over the years, even simply through moral encouragement; your help has been invaluable and will not be forgotten.
\listoffigures \listoftables
\cleardoublepage \pagenumbering{arabic} \chapter{Introduction and Literature Review}
Timely and accurate macroeconomic indicators are fundamental for effective policymaking and private-sector investment decisions. However, official economic statistics---including key measures such as Gross Domestic Product (GDP), inflation rates, and employment levels---typically exhibit considerable publication delays. These delays introduce an "information gap", imposing significant challenges to policymakers, central banks, and financial institutions that must operate under heightened uncertainty. Crucial economic information frequently becomes publicly accessible only weeks or even months after the respective reference period, thereby limiting the effectiveness, precision, and responsiveness of economic interventions GiannoneReichlinSmall2008,BokEtAl2018.
Economic researchers have progressively embraced the methodological paradigm known as "nowcasting" to address this gap. Originating from meteorological nowcasting and subsequently adapted to macroeconomic applications, nowcasting aims explicitly to generate contemporaneous assessments of economic conditions. Unlike traditional nowcasting---which predominantly targets future periods---nowcasting leverages timely and readily available high-frequency data to estimate current economic activity, effectively bridging the information lag before official releases BaumeisterEtAl2021,BanburaRunstler2011. Consequently, nowcasting substantially reduces uncertainty, enhancing the accuracy and timeliness of economic decisions in both public and private sectors.
The global COVID-19 pandemic underscored the practical necessity and utility of robust nowcasting frameworks. During this period, traditional macroeconomic reporting frequencies proved insufficient, unable to capture rapid shifts in economic trajectories. As a result, prominent institutions such as the U.S. Federal Reserve heavily utilized advanced nowcasting methodologies, notably the GDPNow model from the Atlanta Fed Higgins2014 and the New York Fed Staff Nowcast BokEtAl2018. These innovative frameworks successfully integrated unconventional real-time indicators—such as weekly unemployment insurance claims, mobility data, and financial market indices—to facilitate prompt, informed policy responses and clear communication with market stakeholders CajnerEtAl2020,ForoniMarcellinoStevanovic2020.
The benefits of nowcasting extend beyond public policy domains, providing substantial advantages for private investors, financial analysts, and corporate decision-makers. The ability to promptly assess economic conditions enables market participants to proactively manage risks, adjust investment portfolios, and refine strategic operations, thereby reducing informational asymmetries and improving overall market efficiency.
Nevertheless, despite their broad adoption and demonstrable benefits, nowcasting approaches face critical methodological limitations. Notably, these frameworks often lack sufficient interpretability and robust measures of predictive uncertainty. Policymakers and analysts increasingly emphasize the necessity for transparent and interpretable nowcasts, demanding clear insights into the underlying economic drivers of predictive outcomes. Therefore, further methodological advancements that explicitly incorporate model interpretability and rigorous uncertainty quantification are urgently required to enhance nowcasting's practical utility PetropoulosMakridakis2020.
In response to these methodological challenges, our research investigates advanced machine learning methodologies tailored to Singapore. Given Singapore's economic openness, the exceptional quality of its macroeconomic data, and heightened sensitivity to global economic fluctuations, it represents a uniquely suitable empirical setting. This research aims to improve nowcasting accuracy, interpretability, and uncertainty management, providing robust tools for policymakers and financial analysts in contexts characterized by structural uncertainties and rapid economic developments. The rationale and potential of these machine learning-based techniques are detailed in the following sections.
Machine learning (ML) techniques have rapidly gained prominence in macroeconomic nowcasting, driven by their distinctive ability to handle complex, high-dimensional data and effectively adapt to structural changes in economic dynamics. Classical econometric frameworks such as autoregressive (AR), vector autoregressive (VAR), and dynamic factor models (DFM) commonly encounter challenges related to multicollinearity, dimensionality, and non-linear dynamics, particularly under structural breaks and volatile economic conditions MedeirosEtAl2021,CoulombeEtAl2022. In contrast, ML approaches effectively address these limitations by leveraging advanced computational frameworks and data-driven flexibility, significantly enhancing nowcasting accuracy and robustness in high-dimensional environments KimSwanson2018.
A key methodological advantage of ML lies in its ability to efficiently manage high-dimensional, heterogeneous datasets through built-in regularization, adaptive model complexity, and non-linear approximations. Specifically, penalized regressions (LASSO, Ridge, Elastic Net), dimensionality reduction methods (PCR, PLSR), ensemble algorithms (Random Forest, eXtreme Gradient Boosting), and neural networks (MLPs, GRUs) inherently adapt to complex patterns, accommodating data sparsity and non-linear interactions typically overlooked by linear models ZouHastie2005,Breiman2001,ChenGuestrin2016,HewamalageBergmeirBandara2021.
Recent empirical evidence underscores ML's superior performance relative to traditional econometric benchmarks, particularly during periods of economic instability. For example, MedeirosEtAl2021 demonstrated enhanced nowcasting accuracy of ML techniques for inflation in volatile environments, while CoulombeEtAl2022 confirmed ML's advantages in adapting dynamically to structural economic shifts. The COVID-19 pandemic notably exemplified ML's responsiveness to rapid regime changes, effectively integrating unconventional, high-frequency data sources such as mobility metrics, financial indicators, and digital transactions to achieve timely and accurate nowcasts ForoniMarcellinoStevanovic2020,BarbagliaEtAl2023.
The methodological robustness of ML approaches principally originates from three core attributes:
Nevertheless, despite their empirical successes, ML models frequently suffer from limited interpretability and challenges related to uncertainty quantification. Policymakers and financial analysts increasingly prioritize transparency and robust uncertainty estimates, where traditional econometric methods still hold clear advantages. Consequently, integrating explainable artificial intelligence (XAI) techniques, such as Integrated Gradients or permutation importance measures, becomes critical for enhancing the transparency and credibility of ML-based nowcasts LundbergEtAl2020,SundararajanTalyYan2017. Indeed, recent scholarship emphasizes interpretability as a fundamental prerequisite for deploying ML models effectively in high-stakes economic and policy environments Rudin2019. Furthermore, the adoption of non-parametric uncertainty quantification methods, notably bootstrap procedures (e.g., stationary and block bootstrap), is essential to provide rigorous prediction intervals and quantify prediction uncertainty reliably PolitisRomano1994,Li2021.
Addressing these interpretability and uncertainty challenges represents a key advancement in ML-based nowcasting research. The methodological framework proposed in this study explicitly tackles these critical dimensions, aiming to strengthen the robustness, transparency, and practical relevance of nowcasts in highly dynamic macroeconomic environments.
Given its economic openness, sensitivity to global shocks, and exceptional data quality, Singapore provides an ideal empirical setting for methodological research on machine learning-based macroeconomic nowcasting.
Given its pronounced openness to trade and financial flows, Singapore rapidly transmits global shocks into domestic indicators like GDP and financial markets, highlighting the urgency for timely nowcasting methodologies.
Singapore's rigorous statistical framework ensures highly accurate and timely publication of key macroeconomic indicators SingStatPortal, significantly facilitating validation of advanced econometric and ML methodologies.
Despite these notable advantages, Singapore remains underrepresented in the nowcasting literature, especially regarding sophisticated ML frameworks. While advanced economies such as the United States or Eurozone have been extensively analyzed within empirical ML-based nowcasting studies, Singapore's case has scarcely been investigated, a gap implicitly confirmed by recent comprehensive reviews of nowcasting methodologies ForoniMarcellinoStevanovic2020. By explicitly addressing this gap, our research provides original insights into ML performance within an economic context uniquely sensitive to global disruptions.
Existing nowcasting practices within Singapore, i.e., the GDP Nowcasting model regularly published by DBS Bank DBSNowcast2025, predominantly rely on simpler year-over-year (YoY) methods and econometric frameworks lacking in sophistication. These models largely neglect advanced, interpretable ML approaches and do not adequately exploit the granular, quarterly-frequency data available. Introducing comprehensive ML methodologies---including penalized regressions, ensemble algorithms, and neural networks---to nowcast Singapore's quarterly GDP growth directly addresses this methodological shortcoming, promising substantial improvements in nowcasting accuracy, responsiveness, and transparency. Our research uniquely combines Singapore's economic sensitivity and exceptional data quality with advanced, interpretable ML methodologies---explicitly addressing critical gaps unaddressed by current practices and studies.
From a methodological perspective, Singapore's economic structure, combined with its high-quality data, offers an advantageous setting for rigorously evaluating ML models' predictive robustness and adaptability under varied economic conditions. The occurrence of recent economic shocks---such as the COVID-19 pandemic and geopolitical tensions affecting global supply chains---provides an ideal context to empirically assess how ML methodologies manage abrupt regime shifts and structural breaks. Such empirical evaluations yield vital methodological insights, extending applicability beyond the Singaporean context.
Its global financial prominence further enhances our research's practical and international relevance, benefiting policymakers, investors, and multinational firms through improved decision-making and risk management strategies. Indeed, addressing interpretability and uncertainty quantification is increasingly recognized as essential for the practical deployment of nowcasting models in policy-making contexts PetropoulosMakridakis2020.
By explicitly focusing on Singapore---an economically strategic yet understudied setting---this research bridges a crucial gap in existing nowcasting literature. It advances robust methodological frameworks with significant implications for global economic nowcasting practices.
The literature on macroeconomic nowcasting features a diverse range of methodologies, broadly categorized into traditional econometric models and emerging machine learning (ML) techniques. We critically synthesize these approaches, highlighting their strengths and limitations, and position our research clearly within this established literature.
Dynamic Factor Models (DFM), extensively developed by StockWatson2002,StockWatson2011,BaiNg2008, represent one of the primary econometric tools for nowcasting. DFMs reduce data dimensionality by extracting common latent factors from large macroeconomic datasets. This approach enhances interpretability and nowcasting efficiency by summarizing numerous correlated variables into a few representative factors. Despite their widespread adoption, DFMs exhibit substantial methodological limitations. In particular, the selection of latent factors is inherently arbitrary and relies heavily on subjective econometric criteria, often leading to model misspecification or omitted variable bias. BanburaEtAl2010 further underlines the complexities of factor selection, especially when dealing with large-scale datasets. Additionally, DFMs predominantly assume linear relationships among variables, failing to capture non-linear dynamics particularly pronounced during episodes of structural instability, such as the global financial crisis or the COVID-19 pandemic.
Other econometric approaches---autoregressive (AR) and vector autoregressive (VAR) models---are similarly prevalent in the nowcasting literature, particularly due to their straightforward implementation and clear interpretability. However, these methodologies suffer considerably from the "curse of dimensionality\footnote{The "curse of dimensionality" refers to the statistical and computational challenges arising from high-dimensional datasets, where the inclusion of numerous variables can deteriorate model performance and nowcasting accuracy due to sparsity, multicollinearity, and parameter instability Bellman1961.}", as the incorporation of numerous variables significantly diminishes model performance, resulting in parameter instability and nowcasting inefficiencies. Consequently, traditional econometric models often lack the robustness and adaptability to reflect rapid economic shifts adequately.
In response to these methodological limitations, recent literature increasingly explores ML-based nowcasting techniques. ML methodologies possess inherent advantages in managing complex, high-dimensional datasets characterized by intricate non-linear interactions and structural instabilities Varian2014,MedeirosEtAl2021,CoulombeEtAl2022. This methodological shift has been systematically documented by KimSwanson2018, who underscore ML's comparative advantages in predictive robustness. The principal ML methods in current nowcasting applications can be classified as follows:
Recent methodological discourse increasingly highlights the importance of interpretability and uncertainty quantification in nowcasting applications. Traditional econometric models provide straightforward inferential frameworks (e.g., hypothesis tests, confidence intervals), enabling clear interpretability and robust uncertainty assessment. In contrast, ML methodologies often lack explicit statistical inference, necessitating alternative procedures---primarily non-parametric resampling methods such as stationary or block bootstrap---to deliver rigorous predictive uncertainty measures PolitisRomano1994,Li2021.
Moreover, recent literature emphasizes integrating explainable artificial intelligence (XAI) techniques into ML-based nowcasting frameworks. Popular methodologies such as Integrated Gradients, SHAP values, or permutation feature importance have been applied extensively to elucidate ML predictions, thereby increasing transparency and credibility for policymakers and financial analysts LundbergEtAl2020,SundararajanTalyYan2017,Rudin2019.
Our research explicitly acknowledges and addresses these methodological gaps, proposing a robust ML-based nowcasting framework tailored to the Singaporean economic context. The study systematically implements a comprehensive range of ML methodologies---penalized regressions, dimensionality reduction techniques, ensemble learning algorithms, and neural network models---rigorously evaluated through transparent interpretability methods and robust bootstrap-based uncertainty quantification. This approach substantially improves existing Singaporean practices, currently dominated by traditional econometric and simplified nowcasting techniques, thereby contributing original empirical insights and methodological advancements to the broader nowcasting literature.
Despite significant methodological advancements, traditional econometric and emerging machine learning (ML) approaches present noteworthy limitations in macroeconomic nowcasting. Critically examining these constraints provides valuable insights into their practical applicability and robustness, particularly under turbulent economic conditions.
Traditional econometric methods—dynamic factor models, autoregressive (AR), and vector autoregressive (VAR)—typically assume linear relationships and stationarity within data-generating processes. Such assumptions frequently become restrictive during pronounced non-linearities and abrupt structural shifts, limiting predictive accuracy and flexibility StockWatson2011,BanburaEtAl2010,NgWright2013. The global financial crisis, the COVID-19 pandemic, and geopolitical disruptions exemplify periods when traditional models often experienced pronounced nowcast deterioration due to their limited ability to dynamically adapt to extreme circumstances StockWatson2011, BanburaEtAl2010. Additionally, the subjectivity in selecting latent factors and variables, particularly within dynamic factor models, introduces risks of omitted-variable bias and model misspecification, thus adversely impacting nowcast stability BaiNg2008.
While conceptually superior in managing high-dimensional data and non-linear dynamics, ML methodologies exhibit their distinct limitations. Although theoretically adept at capturing complex economic interactions, empirical evidence on ML models' robustness remains mixed, especially during rapid regime shifts. Techniques such as Random Forest, eXtreme Gradient Boosting, and neural networks have demonstrated substantial predictive instability under abrupt economic changes MedeirosEtAl2021, CoulombeEtAl2022, raising concerns about their consistency across diverse economic scenarios MedeirosEtAl2021,CoulombeEtAl2022,GiannoneLenzaPrimiceri2021.
Interpretability also remains a prominent challenge for ML approaches. Methods such as ensemble learners and deep neural networks function predominantly as "black boxes," limiting transparency regarding internal decision processes. This opacity significantly diminishes their direct applicability in policy contexts, where the clarity of nowcasts and accountability are paramount. Recent advancements in explainable artificial intelligence (XAI), including Integrated Gradients, permutation importance, and SHAP values, partially mitigate these interpretability issues; however, achieving comprehensive transparency in ML-driven nowcasting remains an ongoing methodological challenge Rudin2019, LundbergEtAl2020.
Robust quantification of nowcast uncertainty presents another substantial obstacle within ML frameworks. Unlike traditional econometric methods, which offer well-established inferential techniques (confidence intervals, hypothesis testing), ML models primarily rely on algorithmic optimization, lacking direct statistical inference procedures. Thus, rigorous uncertainty estimation in ML necessitates alternative methods, such as non-parametric stationary or block bootstrap procedures. These bootstrap techniques, essential for valid uncertainty measures, must explicitly accommodate temporal dependencies and structural instabilities, complicating their methodological application and computational feasibility PolitisRomano1994,Li2021.
Moreover, predictive reliability in ML methods significantly depends on meticulous hyperparameter tuning, extensive computational resources, and rigorous cross-validation. ML models are prone to overfitting, highlighting the necessity of systematic validation procedures and careful preprocessing of heterogeneous high-frequency data to prevent misleading inferences and spurious relationships HastieTibshiraniWainwright2015.
While existing literature underscores ML's potential benefits, empirical evidence consistently indicates substantial heterogeneity in predictive robustness across economic contexts and methodological implementations. Systematic empirical analyses under diverse macroeconomic conditions, especially during structural breaks, remain indispensable for thoroughly understanding the comparative advantages and limitations of ML and traditional econometric methodologies.
Recognizing these gaps, our research explicitly incorporates advanced bootstrap methods for robust uncertainty quantification and sophisticated XAI techniques for enhanced interpretability. This methodological integration directly addresses the limitations identified in traditional econometric and emerging ML approaches, thereby significantly advancing the robustness, transparency, and practical utility of macroeconomic nowcasts under complex economic dynamics.
This study advances the methodological frontier in macroeconomic nowcasting by addressing the following targeted research question: "How can machine learning (ML) and advanced econometric methods significantly enhance the quality, interpretability, and robustness of quarterly GDP growth nowcasts for Singapore, particularly during periods of structural instability?" Motivated by clearly identified gaps within both traditional econometric and emerging ML frameworks, we propose an original, rigorous multi-model nowcasting framework tailored explicitly to Singapore's highly sensitive and globally interconnected economic context.
Specifically, our contributions extend the existing nowcasting literature along five principal dimensions:
These original methodological and empirical advancements significantly extend existing knowledge, providing robust, transparent, and operationally relevant nowcasting solutions to policymakers, central banks, and financial stakeholders in conditions of pronounced economic volatility and structural uncertainty.
\cleardoublepage
\chapter{Methodology and Pipeline}
We relied on a dual‐language and multi‐script workflow to execute our research, combining Python (version 3.x) and R (version 4.x) to exploit each environment's distinctive strengths. All scripts are managed within Spyder and RStudio integrated development environments. This section outlines the libraries employed, the families of nowcasting models under scrutiny, and the pipeline structure devised to accommodate predictive performance, prediction interval estimation, feature‐importance analysis, and full‐sample and sub‐period evaluations.
\phantomsection
We analyzed four major families of nowcasting algorithms:
\phantomsection
Common and General Libraries for Python-based Tools and Scripts
Specific Python Libraries for Scripted Model Families and Diagnostic Tests
Common and General Libraries for R-based Tools and Scripts
Specialized R Routines Implemented in the Project
All scripts implement partial logging, memory cleanup, and a seed fix (e.g., np.random.seed(42), os.environ["PYTHONHASHSEED"]="42", or set.seed(42))\footnote{The value "42" is often chosen as a "random seed" for reproducibility reasons as an Easter egg: in the humorous science fiction novel "The Hitchhiker's Guide to the Galaxy," British writer Douglas Adams identifies "42" as "the answer to life, the universe and everything" Adams1979. Since then, the number has become a recurring cultural reference in programming languages and technical documentation.} to guarantee consistency throughout the entire pipeline.
Our entire research pipeline comprises the following steps:
The overall workflow, summarized in Figure (ref).
This sequential approach ensures that each subtask---from forecast generation and intervals, through the MCS selection and final aggregation to the formal out‐of‐sample testing---remains consistent and reproducible across Python and R environments. The subsequent sections detail the core methodological elements, including data transformations, hyper‐parameter optimization, sub‐period evaluations of predictions and their reliability, feature‐importance analyses, and interpretability of the aggregated results.
In this section, we detail the phases related to the construction of the dataset and the preprocessing techniques adopted to ensure consistency, uniformity, and reliability of the variables in the subsequent one-step-ahead nowcasting process. The central objective is twofold: on the one hand, to have a database adequately refined and free from distortions deriving from non-stationarity phenomena or non-uniform frequencies; on the other, to ensure that the transformation and standardization procedures strictly consider the temporal sequentiality, to avoid look-ahead bias (ex-post information contamination). At the end of the initial acquisition operations, the original dataset had 95 candidate variables (features), with coverage from the first quarter of 1990 (1990 Q1) to the second quarter of 2023 (2023 Q2), for a total of 134 quarterly observations. Following the procedures of filling, cleaning, feature engineering, removal of features excessively incomplete or not plausibly linked to economic activity, and exclusion of series that proved to be non-stationary according to the augmented Dickey–Fuller (ADF) test, a core of 58 features has been reached. Five dummy variables were then added to these ones---three for seasonality ("Seasonality Quarter 1", "Seasonality Quarter 2", "Seasonality Quarter 3") and two to identify positive or negative economic shocks ("Past Negative Shocks of Singapore's GDP Quarter over Quarter Growth", "Past Positive Shocks of Singapore's GDP Quarter over Quarter Growth")---bringing the total number of explanatory variables to 63. Together with the target variable (i.e., the deflated GDP growth rate), the final dataset has 64 columns, continuously covering over 134 quarters. This configuration, the result of the balance between maximizing information and minimizing methodological uncertainty, constitutes the solid basis for the nowcasting analyses in the following chapters.
\phantomsection
To build the dataset, we collected data mainly from statistics produced by the Singapore Department of Statistics SingStatPortal. In particular, we used quarterly and monthly historical series relating to the main indicators of the Singapore economy: industrial production, foreign trade, price dynamics, real estate sector, household consumption, transport and airport movements, labor cost indices, balance of payments, exchange rates, business confidence indicators, as well as demographic variables. Each series required a preliminary acquisition, cleaning, and integration process into a single dataset, temporally indexed from the first quarter of 1990 to the most recent period covered by the official archives (second quarter of 2023, in our case). In addition to the SingStat sources, we had to integrate two specific exchange rate series (Euro/SGD and Renminbi/SGD) from the Bank of Italy BancaItaliaExchange since the corresponding observations in SingStat were incomplete for the period of interest. In particular, for the Euro/SGD, we used a series that also covered the years before the formal introduction of the euro, using data referring to the ECU (European Currency Unit) and the one-to-one correspondence between ECU and euro starting from 1 January 1999. Similarly, in the case of the Renminbi, the SingStat source had gaps up to June 1993, so the series provided by the Bank of Italy offered a more extensive and homogeneous coverage, allowing us to preserve the historical coherence necessary for our analysis horizon. Before proceeding with the transformations, we executed a preliminary check regarding the denomination, the frequency of release, and the correct association with the reference quarter of all the variables to avoid duplication or time misalignment errors, especially in the case of monthly series subsequently converted to quarterly frequency.
\phantomsection
Once the variables were consolidated, we performed transformation and processing operations to make the series comparable and manage the different frequencies.
\phantomsection
Singapore's GDP growth at constant prices is our forecasting target in the nowcasting analysis. We obtained the variable from the quarterly GDP data at current prices and the corresponding GDP deflator (base 2015 = 100).
First, we deflated the nominal values of GDP to obtain a series expressed in real terms and adjusted for variations in price levels. Subsequently, the deflated series was transformed into quarterly growth rates (quarter-over-quarter), calculating the percentage difference between the current and previous quarters.
\phantomsection
The collected series had heterogeneous frequencies: some were available with monthly periodicity (e.g., price indicators, exchange rates, meteorological variables), while others were already processed in quarterly format (e.g., composite confidence indices or GDP data). Since the target variable is on a quarterly basis, we decided to standardize all monthly series to the same frequency to ensure consistency in the subsequent modeling process. The wide availability of data and the relative ease of converting them made it unnecessary to use specific methodologies for managing mixed frequencies (such as MIDAS models\footnote{MIDAS (Mixed Data Sampling) models allow the combination of data observed at different sampling frequencies (e.g., monthly and quarterly) in one regression framework, without the need to convert all variables to the same frequency. For a detailed discussion, see GhyselsEtAl2006}) without compromising the robustness of the analysis. In detail, the choice of the aggregation method (sum, average, or end-of-period value) depends not only on the nature of the variable (e.g., stock vs. flow) but also on the original format of the SingStat publications. In the case of series already expressed as monthly cumulative (e.g., "Contracts Awarded" or "Duty-Paid Releases Of Tobacco"), the sum of the three corresponding months provides the quarterly value. The arithmetic mean over the quarter was calculated for variables released as a monthly average (such as exchange rates or specific price indices). For data cataloged at the end of each period (end-of-period), including "Currency In Circulation," the value recorded in the last month of each quarter was assumed. Although simplified, this approach adequately reflects the sources' structure and ensures a fully quarterly dataset, free from misalignments and gaps.
\phantomsection
Most of the initial variables reflect stocks or indicators on a nominal basis. We decided to convert these stocks into quarterly percentage changes to analyze the rates of change and the characteristics of macroeconomic nowcasting approaches. This practice not only grants comparability between variables with different reference scales but also allows for mitigating possible non-stationarity phenomena. The only exception concerns the availability of some series that already expressed quarterly differences in absolute terms (such as "Changes in Employment"). Since the original levels to calculate a percentage change were unavailable, we constructed an "Acceleration of Net Employment Flow" measure to track the speed with which net job creation varies. Although it is a second difference, we may interpret this indicator as the intensity of employment expansion or contraction.
\phantomsection
Some variables began on dates after 1990 Q1. To limit the introduction of noise and preserve the length of the time series, we eliminated those features with ten or more missing observations in the initial part of the sample (e.g., "Landed Properties Supply Change" or "Foreign Trade Index Change"). We opted for the forward-filling technique for variables with a few missing initial values (less than ten), i.e., translating the last known value to the subsequent quarters in which it was missing. We may justify this imputation by the hypothesis that, in the initial phase of the analysis period, there were no international or internal shocks to the Singapore economy that could cause structural changes that would make this technique inadequate. Finally, we removed some variables that were poorly correlated with the performance of Singapore's real economy, such as elementary demographic data (change in live births, change in deaths) and meteorological variables (changes in air temperature and sunshine, relative humidity and rainfall) whose influence was neither plausible nor theoretically justifiable in a GDP growth forecasting context. This filtering allowed us to focus only on indicators intrinsically linked to macroeconomic developments.
\phantomsection
Although the conversion to growth rates helps reduce the presence of trends or seasonality, we further verified the stationarity of each variable using the Augmented Dickey-Fuller (ADF) test, adopting a conventional significance level (5%). We eliminated those series that remained non-stationary even after the transformation (including changes in personal saving rate and electricity generation). This approach reflects the need to build reliable predictive models, avoiding the persistence of trends or unit-roots generating misleading results in the estimation phase.
\phantomsection
To present the non-artificial nature of the features and the target variable and allow the model to recognize seasonal patterns through the associated coefficients, we chose to handle seasonal effects by creating dummy variables (i.e., "Seasonality Quarter 1", "Seasonality Quarter 2", "Seasonality Quarter 3") rather than proceeding with an ex-ante seasonal adjustment.
Regarding macroeconomic shocks, we defined two additional dummies (i.e., "Past Negative Shocks of Singapore's GDP Quarter over Quarter Growth", "Past Positive Shocks of Singapore's GDP Quarter over Quarter Growth") to mark quarters characterized by strongly negative GDP changes or particularly pronounced rebounds. To this end, we adopted a threshold-based criterion (–2.5% to identify significant declines and +5% to recognize robust recoveries), in line with the literature that uses percentage thresholds to distinguish episodes of contraction or exceptional growth ClaessensKoseTerrones2011,HausmannPritchettRodrik2005,KannanScottTerrones2009.
Integrating these values with the history of significant economic and financial events for Singapore (e.g., the Asian financial crisis of the late 1990s, the Great Recession of 2007–2009, the COVID-19 pandemic) helps avoid ambiguous labeling between normal oscillations and shocks of significant magnitude, thereby more clearly distinguishing growth peaks from a context of statistical normality.\footnote{For an in-depth analysis of threshold-based approaches to identifying recessions and expansions, see also BarroUrsua2008 and BryBoschan1971.}
\phantomsection
The last phase of preprocessing is represented by the iterative standardization of continuous variables so that each forecast window reflects the actual information availability of the moment and avoid look-ahead bias (i.e., it does not benefit from future temporal knowledge). Specifically, at each nowcasting step, we separated the portion of data considered "training" up to that point and calculated the mean and standard deviation for each variable. These parameters were then applied to the validation and test subset to normalize the data consistently with the conditions present during training. The dummy variables were not involved in the standardization process, as they did not need to be scaled. We repeated this mechanism iteratively as the training time horizon expanded, ensuring a correct simulation of the actual practice of nowcasting (in which models are regularly updated as new observations become available) and preventing look-ahead bias as a source of overestimation of the predictive performance.
The phases described in this section ensure that our input dataset's economic variables are consistent, stationary, and comparable while preserving historical sequentiality. The choices made regarding exclusion, imputation, and transformation of the variables constitute a balanced compromise between maximizing the amount of information available and safeguarding the statistical robustness of the nowcasting experiments illustrated in the following chapters.
\phantomsection
The integrated implementation of the Expanding Window and Rolling Window methodologies forms the basis of our Singapore GDP quarterly nowcasting framework Tashman2000,InoueRossi2012,HyndmanAthanasopoulos2021. This solution allows us to draw on the entire history of the data, maintaining a dynamic validation criterion that, at each iteration, provides updated feedback on the model's performance. At the same time, it is consistent with the need for a limited forecast horizon and accurate calibration of the hyperparameters, enhancing the robustness of the entire econometric analysis framework.
We adopt a walk-forward evaluation scheme with an expanding training window and a 12-quarter rolling validation window, followed by a one-step-ahead test quarter. This design preserves temporal ordering and avoids look-ahead bias while continuously assessing performance on the most recent data. The basic idea is to build a data window at each step that includes the training period and an updated validation set and then perform the estimate for a future quarter. This scheme preserves the flexibility needed to apply to different families of algorithms Tashman2000. A schematic of the Expanding+Rolling Window procedure is shown in (ref).
The choice of a progressively expanding training window (Expanding Window) is based on the principle of valorizing the entire information heritage from the first available quarter to the last, without discarding remote data that might contain relevant structural signals. In a macroeconomic context, this approach is sometimes preferable to a pure Rolling Window since it facilitates the identification of long-term patterns in the presence of external shocks or cyclical changes that characterize open economies StockWatson1996,ClementsHendry1998,PesaranTimmermann2007. In our implementation, we initially allocate approximately 80% of the available quarters to the training stage (ending at Q4 2016). \footnote{An 80--20 division of the dataset is a common heuristic in time-series forecasting HyndmanAthanasopoulos2021, balancing the need for a sufficiently large training set with the requirement to retain enough observations for robust validation and testing.} Subsequent iterations expand the training boundary by one quarter each time, preserving the entire economic history of Singapore GiannoneReichlinSmall2008.
In parallel, we applied the Rolling Window concept to define a validation subset that "slides" forward by one quarter at each iteration. In this way, the model is fine-tuned on an expanded training set but always validated on the last 12 quarters (3 years) immediately preceding the forecast PesaranTimmermann2007,BanburaRunstler2011. Choosing a 12-quarter window reflects a balance between reactivity to recent changes and stability of the validation error, a common practice in short-horizon macroeconomic forecasting InoueRossi2012. \footnote{Shorter validation windows (e.g., 4--8 quarters) often yield higher variance in the validation metric, while significantly longer windows may dilute the impact of recent structural shifts. Three years is frequently adopted by institutions and empirical studies, as it provides enough recent data without overwhelming the validation with older observations ClarkMcCracken2009.}
The integration between the Expanding Window and the Moving Window thus represents a trade-off between maximizing historical information and evaluating the model on data consistent with the recent economic situation. Extending the training window allows us to capture long-term patterns, while the moving validation portion highlights any instability or regime changes ClarkMcCracken2009.
Once hyperparameters are selected based on the last 12 quarters of validation, the model is then re-fitted on the entire training-plus-validation window to incorporate all available data. This final fit forms the basis for generating the one-step-ahead forecast on the immediately following quarter, thus ensuring that the model parameters reflect the most up-to-date information collected prior to the test horizon.
At the end of each iteration, the forecast is made for the immediately following quarter (the Test Set Window). The result is a one-step-ahead forecasting mechanism that maximizes the use of historical information and isolates a subset of data to fine-tune the hyperparameters. Furthermore, this nowcasting procedure aligns with the operational practice of economic research centers and the mainstream literature on quarterly forecasts since it systematically updates the estimates for each newly available quarter GiannoneReichlinSmall2008,BanburaRunstler2011.
\phantomsection
In the first iteration, we set 2016 Q4 as the final date of the initial raining Set Window (green+yellow area), covering about 80% of the initial observations. Each subsequent iteration extends the training subset boundary (green area) by an additional quarter, including progressively newer observations. This incremental approach leverages the entire economic history available for Singapore while ensuring that the model receives fresh data at each step GiannoneReichlinSmall2008.
\phantomsection
We used the last 12 quarters (three years) of the Training Set Window to fine-tune the hyperparameters during each iteration. This Validation Subset Window (yellow area) shifts one quarter forward after each step, ensuring continuous performance examination of recent data PesaranTimmermann2007,InoueRossi2012. Twelve quarters are frequently cited in the macro-forecasting literature as a window length that balances the need to detect short-run regime shifts without discarding essential cyclical information ClarkMcCracken2009,BanburaRunstler2011.
\phantomsection
Once the training-validation block is built, we use the immediately following quarter as the Test Set Window (pink area) for the out-of-sample evaluation. Adopting a one-step-ahead horizon is standard in nowcasting exercises: it reduces error propagation across multiple steps and focuses on the next immediate quarter, which is typically of greatest policy and practical interest Tashman2000,MarcellinoStockWatson2006.
\paragraph The combination of Expanding Window in training and Rolling Window in validation meets the requirements of a quarterly nowcasting system, where the progressive availability of new information consolidates the model's robustness GiannoneReichlinSmall2008. The 12-quarter Validation Window offers a balance between continuity of estimates and attention to the most recent signals, while the single-step forecast horizon reflects the immediate macroeconomic outlook BanburaRunstler2011. It is worth recognizing that exceptional shocks or sudden regime shifts could attenuate the framework's ability to capture new conditions promptly PesaranTimmermann2007. In such circumstances, the progressive accumulation of data may be insufficient to compensate for structural changes ClarkMcCracken2009. Nonetheless, the iterative mechanism illustrated here retains a high level of adaptability, as it incorporates the most recent observations at each round and, if necessary, allows for the reconsideration of variables or model configurations GiraitisKapetaniosPrice2013.
\phantomsection
Validation and selection of hyperparameters are crucial steps in our nowcasting framework, as the optimal configuration of the prediction model's hyperparameters significantly affects forecast accuracy.
In our setup, we integrate hyperparameter optimization into the iterative Expanding+Rolling data scheme described in the previous subsection. This procedure unfolds in several steps, ranging from data loading and validation-split creation to the final choice of the hyperparameter set that minimizes the mean squared error (MSE) on the validation set.
\phantomsection
Hyperparameter fine-tuning in a predictive model (e.g., the hidden dimension of a neural network, the regularization strength, or the learning rate) includes various strategies, for example:
While each method has its merits, they can become computationally intensive in higher-dimensional scenarios. Grid Search suffers from combinatorial explosion, Random Search fails to incorporate past trial outcomes systematically, and Genetic Algorithms demand multiple generations and specialized parameter tuning. By contrast, Bayesian Optimization updates its search distribution after every trial, directing the search toward the most promising hyperparameter regions and reducing the computational overhead in our nowcasting setting SnoekEtAl2012,GustafssonEtAl2023.
To thoroughly explore potentially optimal model configurations, we adopt broad hyperparameter ranges for all methods tested\footnote{Recall that these methods are: LASSO, Ridge, Elastic Net, Principal Component Regression, Partial Least Square Regression, Random Forest, eXtreme Gradient Boosting, Multilayer Perceptron, Gated Recurrent Unit.}. Theoretically, wide intervals help prevent overlooking unusual but beneficial hyperparameter values in macroeconomic contexts BergstraEtAl2011,ShahriariEtAl2016. Narrow ranges can bias or restrict the model to local optima, especially when relationships are nonlinear or the predictor set is large SnoekEtAl2012. By contrast, allowing hyperparameters to span orders of magnitude (e.g., for regularization or hidden dimensions) gives the algorithm the flexibility to discover diverse solutions. Practically, macroeconomic nowcasting often involves heterogeneous data, from short, low-frequency series to many high-dimensional indicators BanburaRunstler2011,Varian2014. Broader ranges increase robustness, letting TPE and pruning mechanisms find "extreme" values that might be useful under regime shifts or strong nonlinearities CerqueiraEtAl2020. Bayesian optimization also prunes less promising configurations early SnoekEtAl2012,AkibaEtAl2019, focusing the search on promising regions without high computational costs. That is especially relevant for rolling or expanding-window procedures, where limiting hyperparameters could exacerbate misspecification risks Tashman2000,ProbstWrightBoulesteix2019.
In the specific case of quarterly nowcasting for Singapore's deflated GDP, with a dataset of 63 explanatory variables and 134 total observations, we need a method that efficiently uses each evaluation and quickly prioritizes promising hyperparameter configurations, thus preventing excessive computational burden. Recent evidence from large Bayesian VARs highlights how careful hyperparameter selection (e.g., shrinkage priors) can markedly improve macroeconomic forecast accuracy\footnote{See also the foundational work on BVARs in Litterman1986.} BanburaEtAl2010,CarrieroClarkMarcellino2015a,GiannoneEtAl2015,CimadomoEtAl2022.
Bayesian Optimization, specifically the Tree-structured Parzen Estimator (TPE) method from the Python library Optuna, offers an attractive compromise since it updates a surrogate model at each new trial (where a trial, within each iteration, corresponds to a training and validation attempt using a given set of hyperparameters, here represented by the vector \(\boldsymbol{\theta}\)). This surrogate model steadily focuses on the most promising regions of the hyperparameter space.
\phantomsection
Bayesian optimization aims to progressively locate those regions in the hyperparameter space that yield superior results without resorting to an exhaustive or purely random search. In general terms, consider defining an objective function:
where \(\boldsymbol{\theta} \in \mathcal{H}\) is a hyperparameter configuration (e.g., the hidden dimension of a neural network, a regularization coefficient, or the learning rate), and \(\mathcal{H}\) is the space of all admissible hyperparameter combinations. In \(\hat{y}_{i}^{(\mathrm{val})}(\boldsymbol{\theta})\), the model configured by \(\boldsymbol{\theta}\) yields the prediction for the \(i\)-th observation in the validation set. In contrast, \(n_{\mathrm{val}}\) is the number of observations in that validation subset. Bayesian optimization iteratively constructs a surrogate model of \(\mathcal{L}(\boldsymbol{\theta})\), updating it whenever we obtain a new MSE value for a particular set of hyperparameters.
We employ the Tree-structured Parzen Estimator algorithm in Optuna\footnote{The TPE algorithm was originated in the Hyperopt library and subsequently integrated into other frameworks, including Optuna.}. TPE splits trials into two classes: those with a "good" error (below a threshold \(\gamma\)) and those with a "less good" error (above or equal to \(\gamma\)). The threshold \(\gamma\) is a dynamic threshold used by TPE to classify trials (i.e., the different hyperparameter sets tested), distinguishing those with a validation error under \(\gamma\) from those with worse performance. Unlike a static approach, \(\gamma\) is often determined by selecting some fraction (e.g., 10%) of the lowest errors so far, and it updates as new results are generated. TPE then constructs two conditional probability distributions:
and selects the next hyperparameter configuration \(\boldsymbol{\theta}\) by maximizing the ratio \(\ell(\boldsymbol{\theta}) / g(\boldsymbol{\theta})\). Consequently, the algorithm focuses on upcoming trials in regions of the hyperparameter space that, based on prior observations, appear most promising. This approach is especially effective with a limited dataset (134 quarters, in our case) since each evaluation provides valuable information immediately incorporated into the surrogate model. In practice, the algorithm works in this way:
The algorithm applies a pruning mechanism, i.e., early termination of trials whose partial performance is significantly worse than the median of completed trials at the same training epoch. Pruning avoids expending excessive computational resources on clearly suboptimal configurations. Specifically, the "Median Pruner" in Optuna compares the cumulative validation MSE in the early epochs of each trial to the median MSE recorded among completed trials. If the trial in progress lags well behind that median, Median Pruner truncates the trial prematurely. This strategy proves highly beneficial in our context, where the limited sample (134 quarters) and the large number of features (63) make it necessary to conserve computational resources. Additionally, a predefined number of "warm-up" trials run to completion, gathering baseline statistics for TPE and balancing exploration against potential exploitation of already-discovered good solutions.
Hence, there can be two levels of early termination in each trial:
As explained in Section 2.3.1, at each time-window, iteration the data are partitioned as follows:
In each iteration, TPE explores up to \(n_{\mathrm{trials}}\) hyperparameter configurations internal to that same time window. For each trial:
Afterward, the model is retrained on \(\bigl(X_{\mathrm{train}} \cup X_{\mathrm{val}}, y_{\mathrm{train}} \cup y_{\mathrm{val}}\bigr)\) using \(\boldsymbol{\theta}^{*}\), generally with the same training parameters (up to 200 epochs, early stopping at 30). This final model then produces an out-of-sample forecast on \((X_{\mathrm{test}}, y_{\mathrm{test}})\). Shifting the time window by one quarter completes one iteration; repeating this procedure yields a sequence of out-of-sample forecasts that aligns with real-world nowcasting practices.
Adopting TPE alongside a pruning mechanism (Median Pruner) effectively handles the search for hyperparameter configurations over a broad space, even with relatively few quarterly observations GustafssonEtAl2023. TPE rapidly directs trials toward promising regions, reducing the full training required compared to a fully random or exhaustive search; pruning halts configurations with high error, saving computational resources. Overall, this synergy iteratively improves the out-of-sample forecasts while maintaining a pipeline sufficiently agile to be updated every quarter.
\phantomsection
Modeling macroeconomic shocks in a nowcasting context plays a crucial role in correctly interpreting economic evolution. In such a context, using dummy variables that signal disruptive events of a positive or negative nature---such as sudden financial crises, pandemics, or unexpected production recoveries---allows us to emphasize exceptional circumstances that usual cyclical trends cannot recognize Salkever1976,ClementsHendry1996,CarrieroEtAl2022. However, such variables must be rigorously managed to ensure that the validation and forecasting process does not incorporate future information that is not yet available during the analysis (look-ahead bias) GiannoneLenzaPrimiceri2021,LenzaPrimiceri2022.
During the Expanding+Rolling iterative procedure and in order to prevent any temporal bias, we adopted the following strategy:
By neutralizing shock dummies only in the last validation quarter and in the test quarter, the procedure reflects a genuine forecasting scenario in which new shocks are unknown until they manifest in real data. At the same time, the model can still learn from historically observed shocks, thereby retaining any insights gained from past anomalous events. This combination of selective neutralization and rolling estimation promotes robust and bias-free short-term projections, thus enhancing the credibility of the overall nowcasting framework.
\phantomsection
The absence of an official nowcasting model for Singapore's deflated quarterly GDP growth motivates the adoption of three essential benchmarks, alongside which we apply several families of statistical-mathematical models. On the one hand, we establish a comparison with elementary or standard references (Random Walk, AR(3) selected from various orders, and a specially configured Dynamic Factor Model); on the other hand, we evaluate several types of models that, due to their structure and theoretical assumptions, present characteristics potentially suitable for a small, dynamic and highly integrated economy.
In this section, we describe the fundamental mathematical-statistical principles of benchmarks and nowcasting models, grouped according to their distinct estimation approaches and theoretical premises to provide a solid methodological basis essential to evaluate the effectiveness of each approach:
\phantomsection
{Random Walk} \\ The Random Walk (RW) or "na\"ive forecast" offers a basic yet essential strategy for predicting GDP growth without imposing structural constraints or including exogenous regressors Muth1960, MakridakisEtAl1998, Tashman2000. Its core assumption is that the optimal predictor for period \(t+1\) is simply the observed value at time \(t\), thereby forgoing any additional hypothesis about underlying dynamics.
Formally, the model assumes a unit-root process:
\[ y_{t+1} = y_{t} + \varepsilon_{t+1}, \]
where:
Consequently, the one-step-ahead forecast is given by:
\[ \hat{y}_{t+1 \mid t} = y_{t}. \]
This specification implies that any permanent shock is not reabsorbed and, in the absence of further information, the method cannot distinguish a long-term economic trend from a temporary fluctuation NelsonPlosser1982, StockWatson1996.
The forecast procedure can be summarized in three steps:
The effectiveness of the Random Walk heavily depends on the data-generating process. If the time series predominantly follows a stochastic pattern, the na\"ive forecast can be highly competitive vis-\`a-vis more complex techniques MakridakisEtAl1998, Tashman2000. However, in the presence of pronounced deterministic trends or clearly defined cyclical patterns, this approach may underestimate or overestimate structural changes, as it merely projects the latest observation forward without modeling any trend component. Because it requires no parameter or hyperparameter estimation, the Random Walk serves as a fundamental benchmark in forecasting exercises MakridakisSpiliotisAssimakopoulos2022, MakridakisSpiliotisAssimakopoulos2022. Its minimalism highlights the gains in predictive accuracy offered by more sophisticated methods, making improvements more credible once compared with a technique that, by definition, integrates no additional information beyond the last observed data point.
{Autoregressive Models} \\ Autoregressive (AR) models rank among the most widespread approaches for analyzing economic time series, as they assume that the current value of a variable (e.g., quarterly deflated GDP growth) depends linearly on its past observations. The general form of an AR($p$) model is:
\[ y_{t} = \alpha + \phi_{1} \, y_{t-1} + \phi_{2} \, y_{t-2} + \cdots + \phi_{p} \, y_{t-p} + \varepsilon_{t}, \]
where:
Such a structure enables the model to capture persistence over time in an economic process, especially when prior shocks continue to affect subsequent observations.
An AR(3) model expresses the current value of a time series (here, Singapore's deflated quarterly GDP growth) as a linear combination of its three most recent past values. Formally, this can be written as:
\[ y_{t} = \alpha + \phi_{1}\,y_{t-1} + \phi_{2}\,y_{t-2} + \phi_{3}\,y_{t-3} + \varepsilon_{t}, \]
where:
This specification goes beyond the simplicity of, for instance, an AR(1) or a Random Walk NelsonPlosser1982, as it can capture short-run dynamics over multiple quarters yet remains more parsimonious than higher-order autoregressive models.
The estimation and forecast procedure for an AR(3) typically follows these steps:
In our context, AR(3) offers a middle-ground strategy that often outperforms simpler models, such as AR(1) or the Random Walk NelsonPlosser1982, while remaining more tractable than AR processes of higher order. Its practical effectiveness, however, ultimately depends on the stability of the underlying economic environment and on how frequently large, unanticipated shocks occur.
{Dynamic Factor Models} \\ Dynamic Factor Models (DFM) are a widely used tool in macroeconomic analysis. They allow the synthesizing of a large set of variables that describe the dynamic structure of an economic system in a small number of latent factors.
Starting from the studies of BaiNg2002, which introduce optimal selection criteria for the number of factors, and of StockWatson2002, which show how a limited set of principal components can explain most of the macroeconomic variability, DFMs have evolved by integrating different solutions for the management of incomplete data and heterogeneous frequencies. In particular, DozGiannoneReichlin2011,DozGiannoneReichlin2012 employ a state-space approach \footnote{In the state-space approach, we represent a system by a transition equation for the dynamics of latent variables (states) and an observation equation that links the states themselves to the empirical data. A Kalman filter continuously updates the states, even with missing data or non-homogeneous release timing.} and an Expectation--Maximization (EM) algorithm \footnote{EM procedures alternate an Expectation phase, in which the distribution of unobserved variables is estimated based on the current parameters, with a Maximization phase, in which the parameters are updated to maximize the conditional likelihood. This iteration continues until convergence.}, combined with the Kalman filter, to address problems of ragged edge\footnote{With ragged edge, we indicate the situation in which some variables end before others or have different publication delays, generating an irregular structure in the last points of the sample.} BanburaRunstler2011 or time series with mixed frequency.
This type of approach provides maximum flexibility. However, in contexts where the data are aligned on a single frequency and are complete, these tools may exceed the actual operational needs and be computationally expensive without frequency discrepancies or significant missing data.
From a theoretical point of view, a DFM assumes that the vector
\( x_{t} \in \mathbb{R}^{N} \),
the vector of the variables observed at time \(t\), is composed of common factors and idiosyncratic components according to the equation:
\[ x_{t} = \Lambda f_{t} + \varepsilon_{t}, \]
where:
In this work, we adopt a "light" DFM model with a more essential structure than the state-space one. In particular, we extracted the factors \(f_{t}\) through Principal Component Analysis. This choice reflects the perspective of StockWatson2002, according to which a few principal factors can explain most of the macroeconomic behavior. To allow for dynamic modulation of the unrelated errors, we consider a structure
\[ \varepsilon_{t} = \Phi(L)\,\varepsilon_{t-1} + e_{t}, \]
where:
Therefore, we preferred an AR for the idiosyncratic term\footnote{In this autoregressive model, the \(\varepsilon_{t}\) variables can retain persistences not captured by the factors. Since we have completed and without frequency mismatches in the quarterly series, we considered the adoption of iterative estimates of the latent states unnecessary.}, avoiding using a completely state-space model since the data were already aligned and did not present missing values that needed to be estimated iteratively. To automatically search for the optimal number of factors \(r\) and the order of \(\Phi(L)\) for each quarter of nowcasting, we used the Bayesian-like iterative optimization process on a sliding validation window and pruning described in section (ref). We omitted using Kalman filters and EM procedures since our series were already quarterly and free of missing data. This condition made unnecessary the data reconciliation or imputation mechanisms proposed in DozGiannoneReichlin2011,DozGiannoneReichlin2012. During the analysis, we then directly integrated the dummy variables of seasonality and economic shocks into the regression component that links the factors to the target variable without adopting more sophisticated state-space smoothing techniques.
This approach is consistent with the reference literature. However, it simplifies the management of any temporal misalignments and missing data, exploiting a dataset in which the quarterly frequency is already uniform.
\phantomsection
Penalized linear models offer a structured approach to controlling overfitting in contexts where numerous regressors may exhibit strong collinearity or marginal relevance. In the one-step-ahead quarterly nowcasting setting for a small open economy, introducing specific shrinkage terms in the objective function helps reconcile flexibility and parsimony. This mechanism prevents extraneous variables from overshadowing essential economic signals, enhancing predictive accuracy Tibshirani1996,HoerlKennard1970,ZouHastie2005.
At the core of this family are techniques that impose different regularization strategies. Lasso, for instance, incorporates an L1 penalty, which drives some coefficients to zero and facilitates automatic variable selection. Conversely, Ridge employs an L2 penalty, shrinking all coefficients continuously while retaining each predictor. Elastic Net combines L1 and L2 components, balancing feature selection with robust shrinkage. These frameworks adapt to evolving economic relationships and integrate new indicators without excessively complicating the model's structure. Their adaptability is particularly relevant when policy shifts or global disturbances arise, highlighting their utility in dynamic environments.
Several design principles distinguish penalized linear models from unpenalized regressions:
Penalized linear models thus represent robust tools for macroeconomic forecasting, especially when data availability expands incrementally. They can handle variations in sample size and shifting data patterns without resorting to intricate architectures. While these methods typically employ linear specifications and do not inherently capture highly nonlinear dynamics, they often balance predictive performance and simplicity. With careful and incremental hyperparameter tuning, they yield short-horizon forecasts that absorb shocks yet remain attuned to underlying trends. This robustness underscores their efficacy in capturing the dynamic nature of economic systems, frequently outperforming unregularized alternatives in volatile and high-dimensional contexts SmeekesWijler2018,UematsuTanaka2019.
{Least Absolute Shrinkage and Selection Operator} \\ Least Absolute Shrinkage and Selection Operator (LASSO) is a regression technique for forecasting models when dealing with several explanatory variables Tibshirani1996, ZouHastie2005. It effectively manages many explanatory variables, including some that may not substantially influence the prediction. In this scenario, LASSO is theoretically helpful for selecting the most significant variables and avoiding overfitting\footnote{The term overfitting refers to excessive adaptation of the model to the training data, which compromises its generalization ability SmeekesWijler2018.} the model to the data, thus optimizing real-time forecasts.
The use of L1 regularization distinguishes LASSO HastieTibshiraniWainwright2015. In this context, the parameter $\lambda$ imposes a penalty equal to the sum of the absolute values of the coefficients $\bigl(\sum_{j=1}^{p}\lvert \beta_{j}\rvert\bigr)$. This mechanism, often referred to as "shrinkage penalty," progressively reduces the coefficients of some variables until they are zero, thereby excluding the less relevant ones. By increasing $\lambda$, the extent of this reduction increases, producing a more accentuated "sparseness\footnote{"Sparsity" is the property of L1 penalty to drive some coefficients exactly to zero, thus eliminating unimportant features. This principle was originally introduced by Tibshirani1996, and later extended by ZouHastie2005, demonstrating its practical advantages in high-dimensional regression scenarios.}" effect.
In the case of models with multiple regressors, where there is a high risk of overfitting, the LASSO allows maintaining a good balance between the estimates' precision and the specification's simplicity Zou2006, UematsuTanaka2019. LASSO differs from traditional linear regressions by automatically selecting variables based on coefficient values, thereby highlighting the most predictive variables MeiShi2024.
From a theoretical point of view, LASSO\footnote{In this discussion, we follow the convention used by the scikit-learn library scikit-learn, sklearnLasso, which inserts the term $\tfrac{1}{2n}$ to normalize the sum of squares and estimates an intercept parameter $\beta_{0}$ separately.} is based on minimizing a cost function that combines the estimation error with an L1 penalty, proportional to the sum of the absolute values of the coefficients. In summary form, the function to minimize is:
\[ \hat{\boldsymbol{\beta}} = \operatorname*{arg\,min}_{\boldsymbol{\beta}} \Bigl\{ \,\underbrace{\frac{1}{2n} \sum_{i=1}^{n} \bigl(y_{i} - \beta_{0} - \sum_{j=1}^{p} \beta_{j}\,x_{ji}\bigr)^{2}}_{\text{error term (reduced MSE)}} +\underbrace{\lambda \sum_{j=1}^{p} \lvert \beta_{j} \rvert}_{\text{L1 regularization}} \Bigr\}, \]
where:
More specifically,
Our approach to modeling with LASSO follows these fundamental steps, emphasizing that the penalty parameter $\lambda$ is determined through an iterative process:
Adopting LASSO in this context allows us to choose $\lambda$ in a robust manner, minimizing the validation error and improving the model's predictive ability in future data HastieTibshiraniWainwright2015, UematsuTanaka2019, MeiShi2024.
{Ridge} \\ Ridge Regression is a forecasting method used in regression analysis that involves several explanatory variables HoerlKennard1970,ExterkateEtAl2016,KimSwanson2014,SmeekesWijler2018. It efficiently handles cases where a large number of variables affect the outcome, even when none of them have a particularly strong effect on their own. In this context, using the L2 penalty allows for keeping all the regressors, even the less relevant ones, while ensuring rigorous control over overfitting by reducing the amplitude of the coefficients HastieTibshiraniFriedman2009.
The distinctive element of Ridge lies in the penalty that equals the square of the coefficients \(\bigl(\sum_{j=1}^{p} \beta_{j}^{2}\bigr)\). This approach, known as the "L2 type shrinkage penalty," incrementally lowers the coefficients' values without eliminating them completely. By raising \(\alpha\), the degree of contraction experienced by the coefficients increases, thus helping to reduce the model's variance in situations that may be prone to multicollinearity.
In the presence of multiple regressors, the Ridge maintains a balance between the estimates' precision and the model's structural simplicity since each variable remains included but suffers a penalty proportional to its size. As a result, the influence of less significant predictors is reduced without implying their complete removal.
From a theoretical point of view, the Ridge\footnote{As with other penalties, in scikit-learn we follow the convention of introducing the factor \(\tfrac{1}{2n}\) in front of the sum of the squared residuals and of estimating the intercept \(\beta_{0}\) separately scikit-learn,sklearnRidge.} minimizes a cost function that combines the estimation error with an L2 penalty proportional to the sum of the squared coefficients. In summary form, the function to minimize is:
\[ \hat{\beta} = \arg\min_\beta \Bigl\{ \,\underbrace{\tfrac{1}{2n}\sum_{i=1}^{n} \bigl(y_{i} - \beta_{0} - \sum_{j=1}^{p} \beta_{j}\,x_{ji}\bigr)^{2}}_{\text{estimation error (reduced form of MSE)}} \;+\; \underbrace{\alpha \sum_{j=1}^{p}\beta_{j}^{2}}_{\text{L2 regularization}} \Bigr\}, \]
where:
More specifically,
Our Ridge implementation follows these key steps, emphasizing that the parameter \(\alpha\) is determined through an iterative process:
The calibration and validation mechanism in our model, which uses Ridge regression as a "prediction engine," allows the configuration of the hyperparameter \(\alpha\) while preserving the model's robustness for future data, thanks to effective shrinkage in reducing multicollinearity among the regressors ExterkateEtAl2016,KimSwanson2014.
{Elastic Net} \\ Elastic Net (EN) is a regression methodology that integrates the L1 type penalty and the L2 type penalty in a single framework ZouHastie2005,HastieTibshiraniWainwright2015. This provides a flexible structure that combines the advantages of LASSO regression (partial selection of variables by zeroing some coefficients) with those of Ridge regression (variance containment thanks to the contraction of parameters). This double penalty is particularly effective in models with numerous regressors since it mitigates multicollinearity problems SmeekesWijler2018,UematsuTanaka2019 and gradually eliminates the less significant predictors.
From a theoretical point of view, Elastic Net\footnote{The scikit-learn library scikit-learn,sklearnElasticNet applies by default a \(\tfrac{1}{2n}\) factor to the squared residuals and estimates the \(\beta_{0}\) intercept separately.} minimizes the following cost function:
\[ \hat{\boldsymbol{\beta}} \,=\, \operatorname*{arg\,min}_{\boldsymbol{\beta}} \Bigl\{ \underbrace{\frac{1}{2n}\sum_{i=1}^{n} \Bigl(y_{i} - \beta_{0} - \sum_{j=1}^{p}\beta_{j}\,x_{ji}\Bigr)^{2} }_{\text{error term (reduced MSE)}} \;+\; \underbrace{\alpha\Bigl[ (1-\gamma)\sum_{j=1}^{p}\beta_{j}^{2} \;+\; \gamma \sum_{j=1}^{p}\lvert\beta_{j}\rvert \Bigr]}_{\text{mixed penalty L2 and L1}} \Bigr\}, \]
where:
More specifically:
Our Elastic Net implementation follows these key steps, emphasizing that the parameters \(\alpha\) and \(\gamma\) are determined through an iterative process:
The sequential update strategy allows the Elastic Net to adjust its penalty parameters dynamically, progressively combining the L1 and L2 components. This approach improves the forecasting capacity of future data since it mitigates the risks of overfitting and maintains high performance in contexts characterized by continuous variations in the regressors.
\phantomsection
Dimensionality reduction techniques prioritize parsimony and robustness by mapping a potentially large set of interconnected variables onto fewer latent factors JolliffeCadima2016,StockWatson2002,BoivinNg2006. In quarterly nowcasting for a small open economy in Southeast Asia, such as Singapore, these methods reduce excessive noise and mitigate multicollinearity, thereby improving forecast stability. By transforming high-dimensional datasets into fewer orthogonal or covariance-driven components, dimensionality reduction models limit the risk of overfitting and accommodate unforeseen policy shifts or external shocks.
Principal Component Regression (PCR) and Partial Least Squares Regression (PLSR) exemplify this approach. PCR first identifies a subset of principal components that capture the bulk of variance in the predictor space JolliffeCadima2016, then fits a regression on these uncorrelated axes. This procedure attenuates issues arising from large predictor pools, helping the model focus on essential signals. Meanwhile, PLSR targets components that maximize the covariance between predictors and the target variable Wold1985,MehmoodEtAl2012, making it particularly appealing when the data contain many near-redundant series BoivinNg2006. By directly emphasizing predictive relevance, PLSR can outperform variance-only methods in scenarios where traditional factor-based models might include uninformative directions.
Several design elements characterize these methods:
Dimensionality reduction provides a more parsimonious representation, as superfluous features are filtered out. This streamlined perspective proves advantageous for nowcasting tasks in dynamic environments, where robust forecasts are often required. Moreover, these models do not impose the strict assumptions of purely linear frameworks, enabling flexible adaptations to evolving economic conditions.
{Principal Component Regression} \\ The Principal Component Regression (PCR) constitutes a hybrid methodological framework that combines the dimensionality reduction of a set of predictors with the estimation of a linear regression model. In particular, the PCR extracts a subset of principal components starting from a potentially large and high-dimensional, multicollinearity-prone (i.e., the presence of strong linear correlations among two or more predictors) set of regressors. These components are constructed as orthogonal linear combinations of the original predictors---meaning they are uncorrelated weighted sums of the input variables, each capturing a specific direction of maximal variance in the data. By concentrating most of the original information into a smaller set of uncorrelated components, PCA effectively reduces the problem's dimensionality and mitigates issues such as multicollinearity JolliffeCadima2016,sklearnPCR.
Once the original regressors are projected onto the new orthogonal axes defined by these components, linear regression is performed using the resulting transformed variables, known as principal components (or scores). The classical formulation of PCR uses Ordinary Least Squares (OLS)\footnote{Ordinary regression (OLS) does not include any penalty term; it minimizes only the sum of the squared residuals.}---a method first proposed by Massy1965---which, however, may suffer from overfitting or high variance in the presence of noisy components. In traditional PCR, the components \(\mathbf{Z}\) are used within a simple OLS regression on the dependent variable \(\mathbf{y}\); such approaches have been successfully applied in macroeconomic forecasting StockWatson2002. In our implementation, we replace the OLS estimator with a Ridge regression HoerlKennard1970,sklearnRidge, thereby introducing an L2 penalty term that regularizes the magnitude of the coefficients and extends the standard PCR framework to enhance robustness and generalization. Recent studies have shown that Ridge regression can be a viable alternative to traditional PCR in high-dimensional settings DeMolGiannoneReichlin2008, GiannoneLenzaPrimiceri2021.
The Principal Component Analysis (PCA) itself relies on the spectral decomposition of the regressor matrix \(\mathbf{X} \in \mathbb{R}^{n \times p}\) after appropriate centering or standardization. This process yields eigenvectors and eigenvalues\footnote{Eigenvectors are directions that remain unchanged when a linear transformation is applied. Eigenvalues are scalars that quantify the amount of stretching, shrinking, or mirroring along those directions.} that characterize the data structure. Formally, one begins by computing the sample covariance (or correlation) matrix:
\[ \mathbf{S} = \frac{1}{n-1} \mathbf{X}^\top \mathbf{X}, \]
and solving the eigenvalue problem:
\[ \mathbf{S}\,\mathbf{v}_r = \lambda_r\,\mathbf{v}_r, \]
which yields a set of orthogonal eigenvectors \(\mathbf{v}_r \in \mathbb{R}^{p}\), each associated with an eigenvalue \(\lambda_r\). The eigenvectors define the directions of maximal variance in the data, while the eigenvalues indicate how much variance is explained in each direction.
We retain the first \(k\) components that capture most of the total variance by ordering the eigenvectors according to their associated eigenvalues in decreasing order. Letting \(\mathbf{V}_k \in \mathbb{R}^{p \times k}\) denote the matrix whose columns are the top \(k\) eigenvectors, the matrix of scores---i.e., the transformed data in the reduced component space---is obtained as:
\[ \mathbf{Z} = \mathbf{X}\,\mathbf{V}_k. \]
Each row of \(\mathbf{Z}\) corresponds to an observation expressed in terms of the new orthogonal axes. The columns of \(\mathbf{Z}\) are uncorrelated by construction, facilitating subsequent modeling steps and alleviating problems arising from multicollinearity in the original data.
In "traditional" PCR, the principal components \(\mathbf{Z}\) are used within a simple OLS regression on the dependent variable \(\mathbf{y}\). However, OLS may yield estimates with high variance when the data are noisy. Moreover, even if all variables carry some marginal signal, purely OLS-based estimates can become unstable if each predictor contributes only slightly but is still relevant; in such cases, leaving minor coefficients unshrunk can inflate variance.
Instead, our second stage adopts a Ridge regression in the principal component domain. Specifically, let \(\boldsymbol{\gamma} = (\gamma_1, \dots, \gamma_k)\) be the vector of Ridge coefficients on the principal components \(\mathbf{Z}\). Formally:
\[ \hat{\boldsymbol{\gamma}}_{\mathrm{Ridge}} \;=\; \operatorname*{arg\,min}_{\boldsymbol{\gamma}} \Bigl\{ \underbrace{\tfrac{1}{2n}\,\|\mathbf{y} - \mathbf{Z}\,\boldsymbol{\gamma}\|^2}_{\text{estimation error (MSE form)}} \;+\; \underbrace{\lambda\sum_{j=1}^{k}\gamma_{j}^{2}}_{\text{L2 penalty}} \Bigr\}, \]
where \(\lambda\) controls the intensity of the regularization and \(\mathbf{Z}\) is the matrix of the first \(k\) principal components of \(\mathbf{X}\). Notably, the penalty \(\lambda \sum_{j=1}^{k} \gamma_j^2\) acts directly on these coefficients in the (orthogonal) component space, rather than on the original predictors. The intercept \(\gamma_0\) is estimated separately (and typically excluded from the penalty) if mean-centering is performed. Once \(\hat{\boldsymbol{\gamma}}_{\mathrm{Ridge}}\) is obtained, the corresponding coefficients in the original feature space arise via the loadings \(\mathbf{V}_k\), namely: \[ \hat{\boldsymbol{\beta}}_{\mathrm{PCR}} \;=\; \mathbf{V}_k\,\hat{\boldsymbol{\gamma}}_{\mathrm{Ridge}}, \]
which effectively re-expresses the penalized coefficients back to the domain of the initial predictors.
This variant of PCR (with Ridge regression in the second stage) can limit the estimates' variance while maintaining the benefits of dimensionality reduction. By imposing \(\lambda \sum_{j=1}^{k} \gamma_{j}^{2}\) on the coefficients \(\boldsymbol{\gamma}\), the Ridge regression shrinks large parameter estimates toward zero, thereby mitigating the risk of overfitting that could arise even after dimensionality reduction. Indeed, while the projection onto a reduced space alleviates multicollinearity by extracting orthogonal components, some of those components may still capture variance that is not strictly relevant for prediction. The L2 penalty thus serves as an additional safeguard by discouraging overly large weights and containing the impact of noise. This mechanism is especially beneficial in high-dimensional contexts, where the number of potential regressors---and hence principal components---may be considerable and the sample size is limited. Consequently, combining PCA with a Ridge penalty reduces the dimensional complexity and provides a second layer of shrinkage, promoting more stable and robust forecasts.
In our nowcasting context, PCR is adopted with an iterative approach on time blocks, where at each sample expansion, both the number \(k\) of principal components and the regularization parameter \(\lambda\) of the L2 regression are recalibrated. In each window (train+validation), the procedure follows two fundamental phases:
At the end of each prediction, the training window is expanded to include the last tested observation, and the entire hyperparameter calibration and training cycle is repeated. This "expanding window" strategy allows constant updating of the principal component decomposition and the degree of L2 penalization to reflect any structural changes in the underlying economic relationships BanburaRunstler2011.
PCR emerges as a flexible forecasting engine that leverages dimensionality reduction and linear regression. The adoption of Ridge regression in our context (instead of the traditional OLS estimation) enhances the model's robustness, mitigates the risk of overfitting while maintaining a high predictive capacity, and is in line with recent findings in high-dimensional macroeconomic forecasting GiannoneLenzaPrimiceri2021. Furthermore, the iterative hyperparameter calibration process, which employs Bayesian techniques, enables the model to adapt to changes in data ChinnMeunierStumpner2023.
{Partial Least Squares Regression} \\ Partial Least Squares Regression (PLSR) originates from the work of Wold1985 and serves as a regression strategy that helps reduce the dimensionality of the covariate system while identifying the latent directions that maximize the covariance between the regressors and the dependent variable MehmoodEtAl2012. This procedure is advantageous in the presence of datasets with potential multicollinearity or a high number of variables, since it concentrates the relevant information in a small number of uncorrelated factors while preserving the relationship with the response.
From an applicative point of view, PLSR offers a compromise between the reduction of dimensionality (characteristic of component methods) and the focus on the relationships with the variable of interest. Compared to other linear techniques, PLSR does not only perform a decomposition of the regressor matrix \(\mathbf{X}\)\footnote{The notation \(\mathbf{X}\in\mathbb{R}^{n\times p}\) indicates \(n\) observations and \(p\) regressors; in PLSR, the dependent variable \(\mathbf{y}\) is treated in parallel, identifying the components with maximum covariance between \(\mathbf{X}\) and \(\mathbf{y}\).}, but also builds latent components that aim to explain both the variance of \(\mathbf{X}\) and, simultaneously, the variability of \(\mathbf{y}\).
This feature differentiates PLSR, for example, from Principal Component Regression. In PCR, the extraction of the components occurs by considering only the maximization of the variance of \(\mathbf{X}\), without explicitly taking into account \(\mathbf{y}\). If some aspects of \(\mathbf{X}\) explain a large portion of variance but are poorly correlated with the target variable, they can still be included among the principal components, slowing or hindering the predictive capacity MehmoodEtAl2012. On the contrary, PLSR integrates the variable of interest in the same decomposition procedure, selecting informative directions that correlate with \(\mathbf{y}\). This approach tends to generate latent factors more relevant from the viewpoint of prediction, making PLSR an effective option in situations where the mere explanation of the variance of the regressors does not suffice to guarantee adequate model performance.
Furthermore, PLSR proves particularly robust when regressors exhibit strong collinearity or when the number of variables is high compared to the available observations. In such cases, focusing on components that maximize \(\mathrm{cov}(\mathbf{X}, \mathbf{y})\) helps exclude directions with high variance but low predictive relevance, which could otherwise distort the estimate KraemerSugiyama2011. Consequently, the set of latent factors tends to be better suited for capturing meaningful relationships between the regressors and the output variable, mitigating the risk of overfitting.
Thanks to this mechanism, PLSR can provide a competitive advantage in regression tasks by extracting dimensions that both explain data variance and contribute significantly to estimating the dependent variable. In nowcasting or macroeconomic forecasting contexts, where numerous potentially informative series are often available, PLSR supports a more direct focus on the components that truly matter for the phenomenon of interest KellyPruitt2015, GroenKapetanios2016, Stamer2024, thus allowing the model to concentrate on latent structures that meaningfully enhance predictive ability.
PLSR\footnote{In the implementation based on scikit-learn scikit-learn-pls, the PLSRegression class jointly handles the regressors and the dependent variable, typically excluding the intercept from the penalization.} searches for a sequence of latent factors \(\mathbf{t}_{1}, \mathbf{t}_{2}, \dots, \mathbf{t}_{a}\) such that the projection of \(\mathbf{X}\) and \(\mathbf{y}\) onto these factors is maximally correlated Wold1985. Formally, let \(\mathbf{w}_{j}\) be the weight vectors that define each factor \(\mathbf{t}_{j} = \mathbf{X}\,\mathbf{w}_{j}\). Once \(a\) such factors are determined, the model can be written as:
\[ y \,=\, \beta_{0} \;+\; \sum_{j=1}^{a} t_{j}\,c_{j} \;+\; \varepsilon, \quad \mathbb{E}\bigl[\varepsilon \mid \mathbf{X}\bigr] \,=\,0 \]
where \(\mathbf{c}_{j}\) denotes the coefficient associated with the \(j\)-th factor. The number of components \(a\) is a critical hyperparameter, as it dictates the level of dimensional reduction and the balance between explanatory capability and model complexity. The estimation error is typically evaluated via the Mean Squared Error (MSE):
\[ \mathrm{MSE} \;=\; \frac{1}{n} \sum_{i=1}^{n} \bigl(y_{i} - \hat{y}_{i}\bigr)^{2}. \]
The search for latent factors proceeds iteratively, maximizing the covariance \(\mathrm{cov}(\mathbf{t}_{j}, \mathbf{u}_{j})\) between the projections \(\mathbf{t}_{j}\) (from \(\mathbf{X}\)) and \(\mathbf{u}_{j}\) (from \(\mathbf{y}\)) at each step MehmoodEtAl2012.
Consistent with our other approaches, we employed PLSR as a "prediction engine" to generate one-step-ahead estimates through an expanding window technique. In each time block, the dataset was split into a training+validation subset and a test observation. The selection of the number of components \(a\) was carried out via Bayesian optimization based on TPE BergstraEtAl2011, AkibaEtAl2019, SnoekEtAl2012, which:
After this procedure was completed, we recalibrated the model using the entire span (training+validation) with the optimal number of components, then generated the forecast for the test portion. Upon each sample expansion, the window was extended to include the newly available observation, and the calibration--validation cycle was repeated in full. These progressive hyperparameter updates yielded a flexible framework that adapted to the evolving nature of the data.
Adopting PLSR in this framework allows combining the selection of latent components and the optimization of their number in an iterative way. This methodology provides a model able to capture significant linear relationships even in the presence of potentially collinear regressors and updates the structure of the forecasting system with the arrival of new information.
\phantomsection
Ensemble learning methods combine multiple predictors into a single forecasting framework, leveraging the diversity among base learners to enhance predictive accuracy and stability Breiman2001,ChenGuestrin2016. These approaches offer notable benefits in the context of quarterly one-step-ahead nowcasting for a small open economy such as Singapore. By aggregating distinct trees or boosted sequences of estimators, ensemble models can more effectively accommodate sudden policy shifts, external shocks, and other structural changes than purely linear specifications Yoon2021,ChuQureshi2023. Their adaptability is further heightened by their capacity to handle many potentially correlated variables without imposing rigid functional forms.
A key advantage of ensemble techniques lies in their ability to reduce the variance of individual learners through either bagging or boosting. Random Forest, for instance, grows multiple trees independently on bootstrap samples and averages their outputs, producing stable point estimates with limited sensitivity to outliers Breiman2001. Conversely, XGBoost trains trees sequentially, with each new iteration refining previous residuals through gradient-based error correction ChenGuestrin2016. Although bagging and boosting differ in managing errors and variance, both methods can address intricate, nonlinear patterns frequently encountered in macroeconomic series.
The methodological structure of these models is characterized by:
Ensemble learning strategies adapt systematically to new observations when implemented in a rolling or expanding window setup. The model selection procedure involves identifying optimal hyperparameters by maximizing performance on a validation segment and then retraining on updated datasets. Such continuous recalibration proves particularly relevant for small and open economies, where shifts in global demand, exchange rates, or trade policies can alter traditional forecasting relationships. The combination of variance reduction, nonparametric flexibility, and straightforward interpretability of split structures makes ensemble learning a compelling paradigm for nowcasting tasks, complementing traditional factor models and more recent neural network approaches.
{Random Forest} \\ The Random Forest (RF) is an ensemble learning technique based on bagging Breiman2001,BiauScornet2016, where multiple decision trees are trained independently on bootstrap subsamples of the original dataset. At each tree node, a subset of regressors is randomly selected ProbstWrightBoulesteix2019, thereby ensuring internal diversification among the various trees, reducing the variance of the single model, and mitigating the risks of overfitting. Furthermore, unlike some penalized regressions that require an explicit functional hypothesis, the Random Forest can capture potential nonlinear relationships between macroeconomic variables while maintaining robustness to correlations among predictors. Such methodological flexibility proves advantageous in nowcasting analyses, where multiple potentially correlated series and heterogeneous indicators often render the specification of a strictly linear model problematic.
From a theoretical standpoint, the Random Forest is composed of \(B\) decision trees. Each tree \(T_{b}\) \(\bigl(b=1,\dots,B\bigr)\) is built starting from a bootstrap sample of the training data Breiman2001, and a random subset of available features is considered at each node to find the optimal split. In terms of prediction, each of these \(B\) trees provides an estimate \(\hat{y}_{b}(\mathbf{x}_{i})\) for the observation \(\mathbf{x}_{i}\). Aggregation then occurs by averaging the predictions of the individual trees, i.e.:
\[ \hat{y}_{i} \,=\, \frac{1}{B} \sum_{b=1}^{B} T_{b}(\mathbf{x}_{i}). \]
The training objective can be formalized as the maximization (or minimization with opposite sign) of a function related to the root mean square error (\(\mathrm{RMSE}\)). More precisely, we define:
\[ \mathrm{Obj} = \underbrace{-\,\sqrt{\frac{1}{n}\sum_{i=1}^{n} \Bigl( \underbrace{y_{i}}_{\text{actual value}} \;-\; \underbrace{\frac{1}{B}\sum_{b=1}^{B}T_{b}\bigl(\mathbf{x}_{i}\bigr)}_{\text{ensemble-based forecast}} \Bigr)^{2} }}_{\text{negative RMSE (score function)}}. \]
where:
By combining the outputs of all \(B\) trees, the Random Forest reduces variance compared with a single tree and, through randomization both in the observations (bootstrap) and in the features (max_features), it preserves model stability even in scenarios characterized by high multicollinearity.\footnote{In high-dimensional macroeconomic contexts, this bagging framework and the random subspace selection ($m_{\mathrm{try}}$) can effectively mitigate collinearity and overfitting BiauScornet2016,ProbstWrightBoulesteix2019. Such diversity-promoting mechanisms bear a conceptual resemblance to the feature selection or shrinkage strategies adopted by penalized linear models.}
The main hyperparameters of the Random Forest regulate each tree's complexity and define the sampling strategies ProbstWrightBoulesteix2019. In particular, within our modeling framework\footnote{In brackets we have indicated the names of the hyperparameters used in our model in Python sklearnRandomForestRegressor.}:
Calibrating these parameters enables practitioners to balance bias and variance, flexibly adapt to the dataset's features, and ensure satisfactory predictive performance Breiman2001,ProbstWrightBoulesteix2019.
In our procedure, we implemented the Random Forest as a "prediction engine" within a sequential protocol. In each time block, we split the data into a train+validation segment (to calibrate the aforementioned hyperparameters) and a single test period corresponding to the subsequent observation. We relied on a Bayesian search strategy for hyperparameter optimization SnoekEtAl2012,AkibaEtAl2019,ShahriariEtAl2016, enhanced by early "pruning" of less promising configurations. Once the optimal values were identified, we retrained the model on the same train+validation subset and generated the forecast for the test. The dataset was then expanded to incorporate the newly predicted observation, and the entire process was repeated at every fresh period. This "expandable window" approach enabled the Random Forest to adapt progressively to shifts in macroeconomic conditions, reducing the likelihood of prolonged misspecification and supporting more reliable estimates over time CerqueiraEtAl2020.
Integrating the Random Forest with a Bayesian hyperparameter search permits exploring a vast configuration space, leveraging the trees' inherent capacity to detect nonlinear patterns. By iteratively updating these parameters at each dataset expansion, the model achieves greater stability in its estimates without relinquishing algorithmic flexibility, making the Random Forest a competitive option in nowcasting scenarios that demand dynamic and consistent forecasts.
Despite its advantages, our application uses an i.i.d. bootstrap for sampling, rather than a time-series block bootstrap Kunsch1989\footnote{In mathematical terms, if we indicate with $\{(x_t,y_t)\}_{t=1}^T$ our series, a block bootstrap could consist in building a succession of blocks $\{B_j\}$ of length $L$, where $B_j = \{(x_{(j-1)L+1}, y_{(j-1)L+1}), \dots, (x_{jL}, y_{jL})\}$. By sampling such blocks with replacement (possibly managing the remainder when $T$ is not a multiple of $L$), a new artificial series $\tilde{\mathcal{D}}$ of length $T$ is constructed, preserving part of the internal temporal correlation. Each Random Forest tree would then be trained on a subsample $\tilde{\mathcal{D}}_{b}$ generated via this block logic.}. The latter could better account for autocorrelation in quarterly macroeconomic data but would considerably increase computational times, already high due to ensemble methods and hyperparameter optimization, especially when timely nowcasts are needed on a quarter-by-quarter basis BiauScornet2016,ProbstWrightBoulesteix2019. Implementing a fully time-series-friendly Random Forest, incorporating explicit lag features and a block bootstrap for each tree, further inflates data dimensionality and necessitates more elaborate data manipulations. Given the one-step-ahead nature of our nowcasting exercise, we opt for the simpler i.i.d.\ bootstrap, balancing real-time feasibility with robust predictive performance Varian2014,MedeirosEtAl2021.
We acknowledge that this strategy does not fully preserve temporal dependence in the training process and may underreact to abrupt regime shifts CarrieroClarkMarcellino2015b. Nonetheless, adopting an expanding-window validation protocol and sequential hyperparameter tuning alleviates much of the risk of overfitting past regimes, as newly available data gradually refine the model HyndmanAthanasopoulos2021. Moreover, this streamlined approach has proven effective in multiple forecasting competitions, such as the M5 Competition, where computational overhead is a constraint and many participants successfully employed simpler i.i.d.\ or rolling holdout strategies MakridakisSpiliotisAssimakopoulos2022. Future research could explore whether the gains from a block bootstrap or additional time-series-oriented structures outweigh the added complexity in contexts characterized by significant autocorrelation or frequent structural breaks.
{Gradient Boosting with XGBoost (eXtreme Gradient Boosting)} \\ The Gradient Boosting methodology is an ensemble learning technique based on a sequential boosting approach. This approach combines multiple "weak trees" to reduce the forecast error progressively. The key idea is to correct, at each iteration, the residual errors committed by the previous model, thus reducing the bias and strengthening the predictive capabilities.
In this context, eXtreme Gradient Boosting (XGBoost or XGB) ChenGuestrin2016,XGBoostDocs2023 integrates powerful regularization mechanisms, controlled rate iterative learning, and parallel computing facilities while preserving the ability to capture non-linear relationships between regressors and macroeconomic output. This flexibility presents advantages in datasets with multiple potentially correlated variables, where a purely linear model could prove limiting.
Formally, XGB builds an ensemble of \(M\) regression trees, each denoted as \(f_{m}\). The set of such trees produces an aggregate prediction:
\[ \hat{y}_{i} \,=\, \sum_{m=1}^{M} f_{m}\bigl(\mathbf{x}_{i}\bigr), \]
where \(f_{m}\bigl(\mathbf{x}_{i}\bigr)\) represents the single prediction provided by the \(m\)-th tree for the observation \(\mathbf{x}_{i}\). The optimization focuses on minimizing an error metric, such as RMSE (Root Mean Squared Error), which can be transformed into an objective function to be maximized by placing a negative sign in front of the square root. Concretely:
\[ \mathrm{Obj} \,=\, \underbrace{-\,\sqrt{\frac{1}{n}\sum_{i=1}^{n} \Bigl( \underbrace{y_{i}}_{\text{actual value}} \;-\; \underbrace{\sum_{m=1}^{M}f_{m}\bigl(\mathbf{x}_{i}\bigr)}_{\text{sum of the predictions of the M trees}} \Bigr)^{2} }}_{\text{score function (RMSE with negative sign)}}, \]
where:
In addition to the minimization of the loss function, XGB also incorporates a penalty term on the complexity of each tree, aiming to reduce overfitting. Specifically, the regularization includes an L2 penalty governed by \(\lambda\) (reg_lambda), which penalizes large magnitudes of the leaf weights:
\[ \Omega(f_{m}) \,=\, \gamma\,T_{m} \;+\; \frac{\lambda}{2}\sum_{j=1}^{T_{m}}\bigl(w_{j}\bigr)^{2}, \]
where:
By incorporating \(\Omega(f_{m})\) into the global objective, XGB discourages both an excessive number of leaves and excessively large leaf weights, thereby mitigating variance and helping the model generalize better.
Each tree \(f_{m}\) operates as a step function that partitions the space of variables into distinct regions, assigning a constant estimate \(w_{j}\) to each region (leaf). The structure of these trees—splits, depth, and leaf values—is learned iteratively, correcting, at each iteration, the residual errors of its predecessors. However, the L2 penalty shrinks large leaf coefficients, thus reducing the risk of overfitting to high-frequency noise. As previously discussed, this shrinkage also helps preserve smaller yet potentially informative effects by proportionally reducing leaf weights rather than nullifying them altogether.
In our study, we implement XGB as a "prediction engine" within a sequential protocol analogous to that adopted for the Random Forest ChuQureshi2023,MasiniMedeirosMendes2023. In each time block, the dataset is divided into a train+validation subset (used to calibrate the hyperparameters) and a single test observation corresponding to the subsequent period.
For hyperparameter tuning, we rely on a Bayesian search strategy based on TPEBergstraEtAl2011,AkibaEtAl2019,OptunaDocs2023, augmented by a pruning mechanism that halts evaluations showing limited promise. Specifically, we explore the following main hyperparameters (with the names used in the script in brackets):
Once the best configuration of these hyperparameters is identified, we retrain the model on the combined train+validation subset and generate the forecast for the test observation. The dataset is then expanded to include the newly predicted quarter (or period), and the entire process is repeated at each fresh time step. This expanding window approach enables XGB to adapt gradually to structural shifts in the macroeconomic environment and to reduce the likelihood of long-lasting misspecification, thereby enhancing its predictive reliability.
This XGB approach, therefore, merges the predictive capacity of sequential trees with an accurate iterative error correction mechanism while maintaining control over the complexity thanks to targeted regularization parameters. Furthermore, the continuous updating of hyperparameters, made possible by Bayesian optimization, ensures a constant adaptation of the model to macroeconomic series changes without compromising the overall estimates' robustness.
\phantomsection
Neural networks Hornik1989,GoodfellowBengioCourville2016 offer a flexible framework for capturing nonlinear dynamics, complex interactions among variables, and evolving economic conditions. In the specific context of one-step-ahead nowcasting for a small and open economy such as Singapore, their capacity to adapt to rapidly changing environments proves advantageous. This section focuses on two architectures: the Multilayer Perceptron (MLP) and the Gated Recurrent Unit (GRU) ChoEtAl2014, each equipped with regularization mechanisms to ensure robustness against overfitting.
Neural models differ from linear or factor-based approaches by incorporating activation functions, memory cells, or gating structures that introduce nonlinearity. Such features are crucial when macroeconomic indicators shift abruptly due to exogenous shocks, policy interventions, or other regime changes. Indeed, recent studies highlight the competitiveness of neural architectures in macroeconomic contexts characterized by volatility AlmosovaAndresen2023,HewamalageBergmeirBandara2021. Neural networks respond to these shifts by updating their internal parameters in a rolling or expanding manner, thus preserving the ability to learn new data patterns as they arise.
In designing neural networks for quarterly nowcasting, there are several guiding considerations:
Although neural approaches can be more opaque than penalized regressions, the iterative recalibration of weights and the inclusion of interpretability strategies (as already discussed in this study) facilitate a clearer understanding of which predictors drive the nowcast. Theoretically, the synergy between nonlinearity, sequence modeling, and automatic hyperparameter tuning renders neural models a competitive alternative in scenarios marked by high volatility or nonstationary data-generating processes.
Multilayer Perceptron \\ The Multilayer Perceptron (MLP) is a feedforward neural network architecture\footnote{A feedforward neural network is a class of networks in which signals flow solely from inputs to outputs, with no feedback or internal loops.} Hornik1989 designed to address regression problems that may exhibit nonlinear relationships\footnote{Nonlinear relationships are interactions between variables in which a slight change in a regressor can produce non-constant (or non-proportional) changes in the dependent variable.} between the regressors and the variable of interest. Unlike pure linear models, the MLP introduces nonlinear activation functions\footnote{Nonlinear activation functions (e.g., Rectified Linear Unit (ReLU) or sigmoid) transform the output of a neuron after the linear combination of the inputs, allowing to capture more complex relationships.} GlorotBordesBengio2011 and multiple hidden layers\footnote{Hidden layers are the intermediate levels between the input and the output of the network, where internal representations of the data are elaborated.} (hidden layers), allowing the representation of even complex links between economic inputs.
Adopting an MLP for nowcasting is particularly appropriate when the number of variables is high, and it is suspected that non-trivial interactions may influence the target variable (e.g., GDP growth). Furthermore, including regularization mechanisms (such as dropout\footnote{This technique randomly "turns off" some units during the training process, helping to prevent the model from passively storing the noise in the historical data.} SrivastavaEtAl2014) helps mitigate overfitting, favoring a robust adaptation in the face of unstable macroeconomic dynamics.
In such settings, it is common that some predictors become irrelevant or partially redundant, thus introducing noise and complicating the learning process. Regularization mechanisms such as dropout are advantageous in addressing extraneous or spurious regressors that overshadow the genuinely significant ones. These mechanisms foster a more parsimonious representation of the data.
Formally, the MLP expresses the relationship
\[ y_{i} \,=\, f\!\bigl(\mathbf{x}_{i};\,\boldsymbol{\theta}\bigr) \;+\; \varepsilon_{i}, \quad \mathbb{E}\bigl[\varepsilon_{i} \mid \mathbf{x}_{i}\bigr] \,=\,0, \]
where:
Each neuron\footnote{A neuron is the fundamental processing unit of a neural network, in which an activation function follows a linear combination of inputs.} linearly combines the inputs and applies an activation function (e.g., ReLU) to introduce nonlinearity. If a hidden layer has \(h\) neurons, each neuron \(j\) processes
\[ z_{j} \,=\, \sigma\!\Bigl(\sum_{k=1}^{d} w_{jk}\,x_{k} \;+\; b_{j}\Bigr), \]
where:
Connecting multiple layers in sequence creates a hierarchy of representations useful for capturing complex patterns.
In our empirical setup, we configure the MLP with a single hidden layer (or a limited number of layers). Although deeper architectures can approximate highly intricate relationships, in macroeconomic nowcasting with relatively short samples, additional hidden layers do not necessarily improve generalization and can exacerbate overfitting Hornik1989,GlorotBordesBengio2011. A shallower network is often sufficient to capture the bulk of nonlinearities while retaining computational efficiency.
The estimate of \(\boldsymbol{\theta}\) occurs by minimizing a cost function consisting of the mean squared error (MSE):
\[ \min_{\boldsymbol{\theta}} \Bigl\{ \frac{1}{n}\sum_{i=1}^{n} \bigl(y_{i} - f(\mathbf{x}_{i};\,\boldsymbol{\theta})\bigr)^{2} \Bigr\} \]
The backpropagation process\footnote{Backpropagation is the procedure by which the prediction error is propagated backward, allowing the partial derivatives of the error to be computed with respect to each weight and bias.} RumelhartHintonWilliams1986 calculates the gradients\footnote{The gradient is the vector of the partial derivatives of a function (in this case, the error) with respect to the variables to be optimized (weights and biases). It indicates the direction of maximum growth of the function itself and, by reversing its direction, the model iteratively reduces the error.} of this MSE with respect to each weight and bias, allowing them to be updated iteratively through an optimizer (e.g., Adam) KingmaBa2015.
In deep networks or datasets with a large number of variables, overfitting may emerge. Regularization strategies (i.e., dropout) are used to counteract overfitting, restricting the influence of unhelpful or purely noisy features. Hence, although the MLP is in principle flexible enough to incorporate multiple hidden layers, we typically keep the network shallow in nowcasting tasks to avoid excessive complexity and make training more stable in limited-sample scenarios.
Our empirical study implements the MLP as a "forecasting engine" according to a rolling protocol with an expandable window Tashman2000. During each iteration, we explore the following main hyperparameters (with the names used in the script in brackets):
Also, in this case, the search for optimal parameters occurs during each iteration through the Bayesian optimization approach based on the TPE BergstraEtAl2011,SnoekEtAl2012,AkibaEtAl2019. The algorithm explores the space of configurations and, for each one, estimates the MSE on a validation subset. The search process is accelerated by "pruning" mechanisms that stop the least promising hyperparameter combinations early. At the end of the exploration, the combination of hyperparameters that minimizes the validation error is selected. To prevent excessive overfitting, we halt training when there is no improvement in validation error for a specified number of epochs. Once the optimal parameters are determined, the MLP is recalibrated on the entire training+validation block, then the out-of-sample forecast for the test quarter is generated. The last observation (just forecasted) is incorporated into the dataset, and the next block is moved on, repeating the entire process at the new time horizon. In this way, the model adjusts its weights and hyperparameters at each sample expansion, remaining sensitive to possible shocks or regime changes and reducing the probability of misspecification over time.
MLP, integrated with Bayesian optimization and regularization strategies, combines the ability to capture nonlinear relationships and adapt quickly to macroeconomic framework changes. Despite the lower transparency of network parameters compared to penalized linear models (such as LASSO or Ridge), the iterative selection of hyperparameters and the adoption of mechanisms such as dropout ensure a balance between flexibility and robustness, mitigating the impact of potentially irrelevant regressors. Frequent updating allows the system to respond promptly to structural changes, preserving good generalization to future data or unexpected economic scenarios.
Gated Recurrent Unit \\ Gated Recurrent Units (GRUs) are an architecture belonging to the class of recurrent networks\footnote{Recurrent networks (RNNs) are distinguished from feedforward models by their feedback loops that allow the store of information about the previous state, providing a dynamic structure to the forecast.} ChoEtAl2014,ChungGulcehreChoBengio2014, designed to capture dynamic relationships between economic variables along the time axis. The introduction of the gating mechanism\footnote{In the GRU, the update gates (\(z_{t}\)) and the reset gates (\(r_{t}\)) regulate how the new information is combined with the previous state, mitigating phenomena such as the vanishing gradient typical of traditional RNNs.} allows to efficiently synthesize and transmit the memory between successive states, maintaining relatively low computational costs compared to other more complex recurrent networks JozefowiczZarembaSutskever2015.
A key motivation in using GRUs for regression is the need to contain the influence of any irrelevant variables, which could introduce noise in the estimation phases. In this context, it is appropriate to include forms of regularization, such as the L2 penalty (or weight decay) KroghHertz1992, to penalize the excessive growth of coefficients associated with non-informative predictors. In parallel, adopting dropout guarantees further control over overfitting, randomly switching off some neurons during the training phase SrivastavaEtAl2014,GalGhahramani2016,ZarembaSutskeverVinyals2014. Thanks to these measures, GRUs offer flexibility and simultaneously limit the risk of excessively including superfluous components, favoring a reduction in the complexity of the model MerityKeskarSocher2018.
Formally, the GRU model represents the relation
\[ y_{i} \,=\, f\!\bigl(\mathbf{x}_{i};\,\boldsymbol{\theta}\bigr) \;+\; \varepsilon_{i}, \quad \mathbb{E}\bigl[\varepsilon_{i} \mid \mathbf{x}_{i}\bigr] \,=\,0, \]
where:
In GRUs, at each instant \(t\), the latent output \(\mathbf{h}_{t}\) selectively updates the information coming from the previous state \(\mathbf{h}_{t-1}\) and from the current input \(\mathbf{x}_{t}\).
In particular, the update gate \(\mathbf{z}_{t}\) balances the stored and new components:
\[ \mathbf{h}_{t} \,=\, (1 - \mathbf{z}_{t}) \,\circ\, \mathbf{h}_{t-1} \;+\; \mathbf{z}_{t} \,\circ\, \tilde{\mathbf{h}}_{t}, \]
with \(\tilde{\mathbf{h}}_{t}\) calculated by a linear combination of \(\mathbf{x}_{t}\) and the "reset" part of the previous state. The operator \(\circ\) indicates the element-by-element product. The number of hidden units and the number of layers filled determine the representativeness of the model.
In this study, we are employing a single hidden layer for the GRU.\footnote{A similar design choice applies to the MLP in the previous section, wherein a single hidden layer suffices to capture nonlinear relationships while controlling overfitting in limited-sample environments Hornik1989,GlorotBordesBengio2011}. Although deeper recurrent architectures can be beneficial under extensive data availability, several empirical findings suggest that, for quarterly macroeconomic nowcasting, additional layers do not necessarily improve accuracy and may increase overfitting risk HewamalageBergmeirBandara2021,AlmosovaAndresen2023. A single-layer GRU often suffices to capture short-run dynamics in limited-sample environments, preserving computational efficiency and converging more reliably ZarembaSutskeverVinyals2014.
The set of parameters \(\boldsymbol{\theta}\) is estimated by minimizing a cost function consisting of the mean squared error (MSE) plus an L2 penalty
\[ \min_{\boldsymbol{\theta}} \Bigl\{ \frac{1}{n}\sum_{i=1}^{n} \bigl(y_{i} - f(\mathbf{x}_{i};\,\boldsymbol{\theta})\bigr)^{2} \;+\; \lambda_{\mathrm{L2}}\;\|\boldsymbol{\theta}\|^{2} \Bigr\} \footnote{Similarly to what we did in MLP's section, here we refer to the regularization parameter as $\lambda_{\mathrm{L2}}$ to highlight the neural-network weight decay perspective, but its function is analogous to $\lambda$ in a linear Ridge model.} \]
using backpropagation through time Werbos1990. Backpropagation is the algorithm that computes the gradient of the error with respect to all trainable parameters, systematically propagating the errors backward from the output to earlier layers (and across time steps in a recurrent network).
Backpropagation through time gradually propagates the error through the past recursive iterations of the network and allows weights to be updated using optimizers (e.g., Adam) with an L2 penalty. In addition to the dropout, these regularization elements help to contain the possible proliferation of coefficients relevant only to some data segments, preventing an uncontrolled growth of the effective dimensionality.
In the empirical study adopted here, the estimation of the GRU model followed a rolling iterative forecasting protocol with an expandable window: at each sample expansion, the best hyperparameters were automatically identified thanks to TPE BergstraEtAl2011,AkibaEtAl2019 that evaluated candidate configurations based on the in-sample validation results. The search process was further accelerated by "pruning" mechanisms, which interrupted less promising trials early and thus reduced the computational overhead.
The search variables included are:
The search for hyperparameter configurations occurred at each estimation block, stopping training early (early stopping) if no tangible improvements in validation emerged. Once the optimal values were determined, the model was recalibrated on the entire training+validation set and produced the forecast for the out-of-sample horizon. Subsequently, it was advanced by one quarter, including the new historical information and the entire procedure was repeated. This iterative scheme guaranteed a constant updating of the network, allowing it to adapt to any structural changes and containing the risks of misspecification.
The GRU configured as a "predictive engine" is capable of exploiting temporal persistence, with a lighter internal structure than LSTMs, but equally suitable for modeling non-linear connections between macroeconomic inputs and the target variable. The dynamic optimization approach---through progressive sample expansion and Bayesian selection of hyperparameters---ensures its operational flexibility, maintaining a balance between learning ability and generalization robustness.
Constructing "prediction intervals" in macroeconomic nowcasting allows to quantify the uncertainty associated with each point estimate Chatfield1993, Tashman2000. In contexts where the residuals can be considered approximately linear and stationary\footnote{Residuals are defined as "linear" when they can be described by linear relationships with their lags or with the past values of explanatory variables (the so-called "state variables"). "Stationary" means those processes whose statistical properties (mean, variance, autocorrelation) remain constant over time.}, more traditional methods can be adopted, such as the residual-based Bootstrap (or, if appropriate, the sieve bootstrap in AR structures\footnote{Bootstrap is a method for estimating the distribution of a parameter (or estimate) by resampling the original data with replacement.
"Time series bootstrap" methods consider the autocorrelation present in time series (i.e., subsequent values may depend on previous ones). Therefore, resampling procedures (e.g., "block bootstrap") preserve the order of the data or group contiguous intervals to reproduce the temporal structure more realistically.
In the "residual-based bootstrap," the estimated residuals from a model are sampled and fed back into the equations to generate new simulated paths.
The "sieve bootstrap" is a variant typically used in AR contexts, where the residuals are obtained from an autoregressive process approximating the dynamics of the series.}). However, when the analysis is extended to multiple nonlinear models or structural breaks are suspected, these techniques risk not correctly capturing the data's complexity Kunsch1989. Thus, in the presence of potential nonlinearities and structural changes, the pair block bootstrap Masarotto1990,PanPolitis2016 offers the possibility of generating numerous alternative training trajectories while maintaining the temporal dependence in the data and being agnostic to the shape of the residuals, thus more suitable for a multi-model scenario and subject to possible regime shifts. Furthermore, it allows monitoring the evolution of the estimated uncertainty at each forecasting iteration, thus providing a measure of the predictive robustness even in the presence of endogenous variations and potential structural shocks Grigoletto1998.
Consider any predictive model that, starting from regressors \(\mathbf{x}_t \in \mathbb{R}^p\), produces an estimate \(\widehat{y}_t\) of the macroeconomic variable \(y_t\). Formally, we assume that
\[ y_{t} \,=\, f\!\bigl(\mathbf{x}_{t};\,\boldsymbol{\theta}\bigr) \;+\; \varepsilon_{t}, \quad \mathbb{E}\bigl[\varepsilon_{t}\mid \mathbf{x}_{t}\bigr] \,=\, 0, \]
where:
The overall uncertainty around the forecast \(\widehat{y}_t\) depends not only on the variability of the errors \(\varepsilon_t\) but also on possible limitations of the model \((f)\), on the small size of the available sample and any breakpoints (e.g., crises, shocks, legislative changes) ClarkRavazzolo2015, PesaranPettenuzzoTimmermann2006.
In a preliminary analysis, we applied a Bai-Perron test\footnote{The Bai-Perron test is a procedure to identify multiple breakpoints in a regression model. Using information criteria (AIC, BIC) evaluates whether the regression function should be split into segments with different parameters and determines the optimal number of breaks.} using Akaike Information Criterion (AIC) and Bayesian Information Criterion (BIC) on the entire historical series of the quarterly GDP growth rate (from 1990 Q1 to 2023 Q2). According to these criteria, no significant structural breaks emerged since the model with zero breaks was preferred\footnote{Specifically: for \(m=0\) breaks we found RSS = 998.853 (AIC = 273.175, BIC = 278.971); for \(m=1\) break yielded RSS = 945.221 (AIC = 275.668, BIC = 281.412); increasing \(m\) to 2--5 further raised both AIC and BIC. Consequently, the no-break (\(m=0\)) model remained preferred under both criteria.}.
However, this outcome may reflect specific features of Singapore's economy---such as its relatively small size, high degree of openness, and history of rapid cyclical adjustments---which can limit the power of break tests in detecting short-lived or swiftly reversed shocks GiraitisKapetaniosPrice2013. Despite the Bai-Perron test preferring a model with zero breaks, many studies adopt ex-ante break dates ClementsHendry1996 whenever major global crises are known to have potentially altered economic regimes. Such an approach is economically motivated rather than purely statistical, justified in contexts where some shocks, though real, prove too transient or compensated to manifest as strong, persistent breaks in the data.
Due to the well-known global economic shocks (Asian financial crisis of the late 1990s, the September 11 attacks in 2001, SARS, the 2008 crisis, the advent of COVID-19, and the Russian-Ukrainian conflict), we decided to impose certain structural breaks ex-ante corresponding to the quarters in which such events have likely affected global and Singaporean macroeconomic dynamics, regardless of the statistical test ClementsHendry1996. This choice is justified by the economic interest in assessing the possible impact of such shocks. It also reflects a model design approach intended to preserve relevant historical information, even if purely statistical criteria do not emphasize these breaks.
In our analysis and for each iteration, we divided the procedure into the following steps:
These intervals are estimated within an Expanding+Rolling protocol with hyperparameter selection through TPE and pruning (and early stopping, in the case of neural networks). In particular:
The process continues iteratively, moving forward one quarter until covering the entire range of out-of-sample forecasts, ensuring consistency with the Expanding+Rolling strategy BanburaEtAl2010.
Additionally, to evaluate how predictive uncertainty evolves under different macroeconomic regimes, we compute the average width of the prediction intervals across several sub-periods (Pre-COVID, COVID, Post-COVID, Excluding-COVID) and the entire forecast horizon (Overall). This breakdown highlights how global disruptions (e.g., COVID-19) alter the scale of forecast uncertainty compared to more stable phases, offering a more granular assessment of each model's sensitivity to structural shocks.
The "prediction intervals around nowcasts" thus calculated preserve the series' internal correlation, integrate the effect of historical-economic breaks (even if not confirmed by the Bai-Perron statistical test), and provide empirical uncertainty measures valid on multiple models, reflecting the degree of risk associated with each nowcast. At the same time, these intervals serve as a crucial verification criterion for measuring forecast stability, contributing to a more robust and informed comparative analysis between different methodologies MakridakisSpiliotisAssimakopoulos2022.
\phantomsection
Constructing a coherent interpretability framework for macroeconomic nowcasting requires reconciling each model's internal logic with the temporal dependencies inherent in time-series data. Our primary objective is twofold. First, we aim to align the feature importance measure with the specific structure of every model, whether linear or nonlinear, so that we retain the model's characteristic effects (e.g., shrinkage, gradient-based splits, or latent-factor extractions). Second, we seek to determine whether distinct estimation procedures converge on similar rankings of predictors, thus providing a robust cross-check of our nowcasting results.
We choose not to adopt computationally intensive methods, such as time-series variants of SHapley Additive exPlanations (SHAP)\footnote{SHAP (Shapley Additive Explanations) is a "game-theoretic" approach to model interpretation in which the contribution of each explanatory variable is computed based on "Shapley values," i.e., the average marginal effect of including that variable across all possible coalitions, as derived from cooperative game theory. The so-called "time-series SHAP" methods extend this algorithm by introducing dynamic baselines---which replace a single static reference with time-varying or regime-specific points of comparison, thereby adapting the measure of each variable's impact to evolving macroeconomic conditions---and similar adaptations for capturing temporal dependencies and regime changes typical of macroeconomic forecasting. See Lundberg and Lee LundbergLee2017 for the theoretical foundations of SHAP, and Janizek et al. JanizekTsunodaLundberg2021 for applications to time-series predictions.}, for two main reasons:
Instead, we relied on measures compatible with rolling and expanding time windows, preserving seasonality and potential structural shifts while minimizing computational overhead. That maintained the temporal structure of our data, which is necessary for capturing real-world macroeconomic regimes.
We divided our methodology into a triple analysis to gain an overall view of the variables' importance:
In each of these perspectives (global, sub-periods, and quarterly), we sorted the features by the absolute value of their importance scores. Finally, we identified the top 10 for immediate ranking and interpretability. This final step simplifies the comparison of relevant predictors, facilitating an at-a-glance overview of which variables dominate under various economic conditions or time frames.
This triple perspective extends well beyond a single global viewpoint, enabling in-depth exploration of how variables may fluctuate in importance when subjected to structural breaks or nonlinearities. By retaining the internal rationale of each model's fitting procedure, we capture a finer-grained picture of the drivers behind our forecasting. Consequently, our interpretability framework offers a scalable and consistent method across various model classes and provides valuable insights into the temporal progression of each feature's influence. This approach is especially pertinent in macroeconomic contexts, where economic indicators can shift rapidly due to policy changes, crises, or other exogenous shocks, calling for nuanced and computationally feasible interpretability solutions.
\phantomsection
In penalized linear models (e.g., LASSO Tibshirani1996,sklearnLasso, Ridge HoerlKennard1970,sklearnRidge, and Elastic Net ZouHastie2005,HastieTibshiraniWainwright2015,sklearnElasticNet), the estimated coefficients are direct indicators of the influence exerted by each explanatory variable on the forecasts. From this perspective, we can consider the Coefficient-Based Feature Importance (CBFI) as an interpretability measure for penalized linear models. It consists in interpreting the value (positive or negative) of each coefficient as a signal of the direction of the effect and its magnitude as a measure of the intensity of the impact on the target.
This approach hinges on the idea that, in a linear model
the coefficient $\beta_{j}$ associated with each predictor $\mathbf{X}_{j}$ reflects the sensitivity of the dependent variable to that predictor Hamilton1994,HastieTibshiraniFriedman2009. In the case of penalized models, the estimation process introduces a constraint (of type L1, L2, or hybrid) that tends to reduce---or, in some cases, cancel---the coefficients associated with the less relevant features. Formally, for instance, a LASSO model solves
where $\lambda$ controls the degree of shrinkage, pushing less important $\beta_{j}$ toward zero Tibshirani1996,Zou2006.
Consequently, those regressors characterized by coefficient values $\hat{\beta}_j$ significantly different from zero in terms of amplitude assume a central role in explaining the variability of the target quantity. In contrast, coefficients close to zero suggest a lower relevance. Furthermore, this approach maintains a model-specific connotation by relying entirely on the linear regression structure and the shrinkage effects of regularization SmeekesWijler2018,UematsuTanaka2019.
In our linear models, the procedure for calculating the importance of the variables via CBFI follows these steps:
The choice of using a Coefficient-Based Feature Importance in linear contexts with the regularization logic adopted by penalized models proves advantageous and consistent because:
\phantomsection
Also in PCR Massy1965,Jolliffe1982,JolliffeCadima2016, the estimated coefficients are direct indicators of each explanatory variable's influence on the forecasts, albeit through an intermediate transformation. From this perspective, we can also view the Coefficient-Based Feature Importance as an interpretability tool for PCR models. Specifically, each coefficient's sign signals the direction of the effect, and its magnitude measures the intensity of the impact on the target once those coefficients are re-expressed in the space of the original predictors.
This approach revolves around the idea that, in a standard linear relationship
the coefficient $\beta_{j}$ reflects how sensitive the dependent variable is to changes in predictor $\mathbf{X}_{j}$ Hamilton1994,HastieTibshiraniFriedman2009. However, under PCR, we first reduce $\mathbf{X}$ to $\mathbf{Z}$---a set of orthogonal principal components JolliffeCadima2016---and then fit a regression (a penalized one using Ridge, in our case) of the form:
where:
Only after projecting these $\hat{\gamma}_{k}$ back to the original domain (via the loadings from the PCA) Jolliffe1982, we obtain a set of $\hat{\beta}_{j}$ corresponding to each original predictor. As in penalized models SmeekesWijler2018, large absolute values of $\hat{\beta}_{j}$ reveal the most relevant features, whereas small (or near-zero) coefficients indicate minor relevance. In this manner, PCR retains a model-specific interpretation anchored in a linear structure, though partially mediated by the dimension-reduction step Massy1965.
In our PCR framework, the procedure for calculating the variables' importance via CBFI proceeds through the following steps:
Implementing a Coefficient-Based Feature Importance in the context of Principal Component Regression offers a natural extension of the logic employed in penalized linear models with an added dimension-reduction step. In particular:
This additional projection step does not alter the fundamental meaning of $\hat{\beta}_{j}$ regarding its direction and magnitude; instead, it embeds the penalty-induced shrinkage within the dimensionality-reduced space Jolliffe1982. Therefore, PCR's Coefficient-Based Feature Importance approach remains consistent with its penalized linear counterparts while accommodating the PCA transformation.
\phantomsection
In Partial Least Squares Regression (PLSR) Wold1985,MehmoodEtAl2012,Abdi2010, the extracted latent components are direct indicators of the influence exerted by each explanatory variable on the forecasts, albeit through an intermediate step tied to dimensionality reduction. From this perspective, the Variable Importance in Projection (VIP) AkarachantachoteChadchamSaithanu2014,ChongJun2005 is an interpretability measure for PLSR. It consists of interpreting the VIP value of each feature as a signal of the intensity (magnitude) of its impact on the target. By definition, VIP values are always positive. Hence, they quantify only the magnitude of a predictor's contribution to the latent factors and do not convey any information about the sign of the relationship with the target.
This approach hinges on the idea that, in a PLS framework
the latent factors $\mathbf{T}$ reflect the directions in $\mathbf{X}$ that maximally covary with $\mathbf{y}$ KraemerSugiyama2011. Consequently, those regressors characterized by large VIP scores assume a central role in explaining the variability of the target quantity. In contrast, VIPs close to or below unity suggest a lower relevance AkarachantachoteChadchamSaithanu2014,ChongJun2005. Furthermore, this approach maintains a model-specific connotation by relying entirely on the PLS decomposition, which aligns $\mathbf{X}$ and $\mathbf{y}$ in a shared latent space.
As described in our PLSR forecasting section, this estimation routinely adopts an expanding or rolling window. Thus, we recalibrate the model with newly available observations in each iteration before measuring VIP FuentesPoncelaRodriguez2015. That consistency ensures that the interpretative measure aligns with our one-step-ahead forecast procedure.
In our PLSR setting, the procedure for calculating the importance of the variables via VIP follows these steps:
The choice of using VIP-based feature importance in a PLSR context proves advantageous and consistent because:
The emphasis on maximizing covariance rather than solely capturing variance does not alter the fundamental role of each predictor with respect to the target; instead, it leverages the latent-factor structure of PLSR to highlight the most relevant explanatory directions. Therefore, the VIP-based assessment of feature importance remains fully coherent with the logic adopted in penalized linear models and PCR while incorporating an explicit alignment of \(\mathbf{X}\) and \(\mathbf{y}\) that enhances interpretability and predictive accuracy.
\phantomsection
In Random Forest models Breiman2001,BreimanFriedmanOlshenStone1984, the Impurity-Based Feature Importance (also known as Mean Decrease in Impurity, MDI) directly indicates the influence exerted by each explanatory variable on the forecasts. From this perspective, the Impurity-Based Feature Importance is a model-specific interpretability measure for Random Forests. It consists of interpreting the average reduction in node-level impurity\footnote{In a regression tree, "impurity" measures how heterogeneous the target values are within a node. Typically, for node $\ell$ containing $t$ observations $\{y_1,\dots,y_t\}$, impurity can be defined via the variance of the $y_i$ (equivalent to the MSE). An alternative measure is the mean absolute deviation from the median (MAE), which may be more robust to outliers. In both cases, reducing impurity leads to more homogeneous nodes and enhanced predictive performance. For classification trees, criteria such as Gini or entropy are instead employed BreimanFriedmanOlshenStone1984.} (for instance, mean squared error for regression trees) associated with each predictor, thus capturing the magnitude of its impact on the target MullainathanSpiess2017.
This approach hinges on the idea that, in a Random Forest:
each \(\mathrm{Tree}_b\) is independently built on a bootstrap subsample, and the splitting at each node is chosen to maximize the decrease in impurity. In a regression setting, we can define a succinct formalism for the MDI of the variable \(j\) as:
where:
In regression contexts, Random Forests may adopt either "squared error" or "absolute error" as the node-level impurity measure. In our study, the choice between these two criteria is determined dynamically via the TPE-based hyperparameter optimization ShahriariEtAl2016,BorupChristensenMuhlbachNielsen2023. During each iteration, our model tries both options and selects the one minimizing the validation mean squared error:
Thus, Mean Decrease in Impurity (MDI) is conceptually similar in both cases: each split contributes its local decrease in impurity \(\Delta(\mathrm{Impurity})\) to the variable causing that partition, and we average these contributions across all trees to obtain
\[ \mathrm{MDI}_{j} \;=\; \frac{1}{B} \sum_{b=1}^{B} \sum_{\ell \,\in\,\mathcal{L}_b(j)} \Delta(\mathrm{Impurity})_{\ell}. \]
Here, \(B\) is the total number of trees, and \(\mathcal{L}_b(j)\) denotes the set of splits using \(X_j\) in the \(b\)-th tree. The only distinction is whether impurity is measured as variance-like (squared error) or deviation-like (absolute error), and the hyperparameter-tuning routine selects whichever yields a superior validation performance.
Regressors with high MDI values are central determinants of the overall predictive performance. On the other hand, regressors with near-zero MDI scores indicate low relevance. This method is closely tied to the model's specific features because it depends solely on how decisions are made through the tree-splitting process in the Random Forest Breiman2001.
In line with the principles illustrated in the previous sections (Coefficient-Based Feature Importance in penalized linear models, PCR, and PLSR), the procedure adopted to estimate the importance of the variables via Impurity-Based Feature Importance consists of four phases:
The choice of using Impurity-Based Feature Importance (Mean Decrease in Impurity) in a Random Forest context proves advantageous and consistent because:
\phantomsection
In XGBoost models, the Gain-Based Feature Importance (sometimes referred to simply as "Gain Importance") ChenGuestrin2016, XGBoostDocs2023 directly indicates the influence exerted by each explanatory variable on the forecasts. From this perspective, the Gain-Based Feature Importance is a model-specific interpretability measure for boosting methods LundbergEtAl2020,ZhouHooker2021. It consists of interpreting the average improvement in the objective function (such as a regularized mean squared error) associated with each predictor at split nodes, thus capturing the magnitude of its impact on the target. In line with the interpretative rationale observed in linear and factor-based approaches HastieTibshiraniFriedman2009,Jolliffe1982,Wold1985, this measure offers a more adaptable perspective on variable significance, mainly when the data-generating process includes nonlinearities or intricate interactions. However, like other model-specific importance, it only pinpoints which features most effectively reduce the loss function without disclosing whether their relationship with the outcome is positive or negative StroblEtAl2007,NembriniKonigWright2018,ZhouHooker2021. Consequently, it complements methods such as Coefficient-Based Feature Importance, Principal Component Regression, and Partial Least Squares Regression, especially under conditions of strong nonlinearity.
This approach centers on the concept that, in an XGBoost model:
each \(\mathrm{Tree}_b\) is trained sequentially to reduce a differentiable loss function \(\mathcal{L}\) (often with \(L_2\) regularization), and the splitting at each node is chosen to maximize the gain, which measures the drop in the overall loss once a node is partitioned into two child nodes. Formally, for a node \(\ell\) split on feature \(j\), the gain can be expressed as
where:
Thus, the Gain-Based Feature Importance (GBI) of the variable \(j\) in regression settings with XGBoost can be summarized as:
where:
In regression contexts, XGBoost commonly employs a squared-error objective with \(L_2\) regularization ChenGuestrin2016,XGBoostDocs2023, but it may also adopt alternative loss formulations. For example:
Hence, different objective functions alter the exact \(\mathrm{Score}(\ell)\) calculation but preserve the overarching concept that each split's \(\mathrm{Gain}\) indicates how much that split (and hence the feature used) advances the model's predictive accuracy. The hyperparameter-tuning routine (e.g., TPE-based) can systematically compare these objectives BergstraEtAl2011,AkibaEtAl2019,ShahriariEtAl2016 and retain the one yielding the best validation performance.
In our study, we typically adopt a squared-error loss with \(L_2\) regularization, but alternative objectives (including absolute-error-based functions) may also be employed if indicated by the validation results. The choice of hyperparameters (e.g., number of boosted trees, maximum depth of each tree, learning rate, regularization terms) is similarly handled via TPE-based optimization BergstraEtAl2011,AkibaEtAl2019, mirroring the procedure described for Random Forest. The model determines the best combination of these hyperparameters during each iteration and re-estimates the XGBoost ensemble, yielding updated gain-based importance scores.
High GBI values indicate central determinants for the predictive performance. Conversely, regressors with near-zero GBI scores have a marginal influence. Like other model-specific importances, the Gain-Based measure depends solely on the gradient-based splitting decisions within the ensemble and does not reveal the sign of the underlying relationship.
Following the same principles illustrated in previous sections (Coefficient-Based Feature Importance in penalized linear models, PCR, and PLSR), the procedure adopted to estimate the importance of the variables via Gain-Based Feature Importance consists of four phases:
Adopting a Gain-Based Feature Importance approach in XGBoost offers several advantages:
Impurity-Based Feature Importance (MDI) in Random Forest and Gain-Based Feature Importance (GBI) in XGBoost represent model-specific approaches to interpretability. They quantify how much each variable contributes to improving the model's performance by reducing errors (measured by chosen loss measure---an impurity measure or a gradient-based objective) at different decision points (split nodes). However, these methods do not indicate whether a variable has a positive or negative effect on the target outcome. While they do not show direct cause-and-effect relationships, they can effectively capture complex and nonlinear connections, especially in macroeconomic situations where simpler linear models might be inadequate.
Therefore, MDI and GBI offer an alternative and complementary insight to CBFI in penalized and Principal Component regressions, and VIP in Partial Least Squares Regression into which predictors matter the most under conditions of high complexity or time-varying structures while at the same time requiring careful handling of potential correlation biases StroblEtAl2007,NembriniKonigWright2018.
\phantomsection
In Multilayer Perceptron (MLP) and Gated Recurrent Unit (GRU) models, the Integrated-Gradients-Based Feature Importance (IG) quantifies the contribution of each explanatory variable to the neural network's forecast. From this perspective, the Integrated Gradients (IG) technique is a model-specific interpretability measure for differentiable models. It relies on computing the integral of the gradients of the network output concerning the input features, thus capturing how each predictor influences the final prediction SundararajanTalyYan2017.
This approach hinges on the concept that, in a neural network \( F \):
\[ F(\mathbf{x}) \;=\; \widehat{y}, \]
the partial derivatives \(\tfrac{\partial F}{\partial x_j}\) indicate how small perturbations in the input feature \(x_j\) affect the output \(\widehat{y}\). The Integrated Gradients formula extends this idea by evaluating and averaging those partial derivatives along a straight-line path from a reference baseline \(\mathbf{b}\) to the actual input \(\mathbf{x}\).
In practice, we define a linear trajectory
\[ \mathbf{x}(\alpha) \;=\; \mathbf{b} \;+\; \alpha \,\bigl(\mathbf{x} - \mathbf{b}\bigr), \quad \alpha \,\in\, [0,1], \]
where \(\mathbf{b}\) is often a zero vector or another neutral reference. The Integrated Gradient \(\mathrm{IG}_j(\mathbf{x})\) of feature \(j\) is then given by
\[ \mathrm{IG}_j(\mathbf{x}) \;=\; \bigl(x_j - b_j\bigr) \;\times\; \int_{0}^{1} \frac{\partial F\bigl(\mathbf{x}(\alpha)\bigr)}{\partial x_j} \,d\alpha. \]
Hence, the total importance of predictor \(j\) is derived by multiplying the difference \(\bigl(x_j - b_j\bigr)\) by the average gradient along that path. A numerical approximation typically divides \(\alpha\) into discrete steps and sums the corresponding gradients.
In both MLP and GRU models, this technique quantifies the magnitude of each input's impact on the final output FreeboroughVanZyl2022,PaudelEtAl2023, without indicating its sign (positive or negative effect). Consequently, it complements methods such as Coefficient-Based Feature Importance, Principal Component Regression, or Partial Least Squares by addressing highly nonlinear or interaction-heavy data-generating processes.
A key practical decision involves choosing the baseline \(\mathbf{b}\). In many applications, \(\mathbf{b}\) is set to the zero vector, but alternative references include the average of the training set or any meaningful neutral point in the feature space. Similarly, the discretization of \(\alpha\) (i.e., the step size or number of steps) influences both the computational burden and the numerical precision of the integrated gradients.
In analogy to the Random Forest (MDI) and XGBoost (GBI) frameworks, the neural network model (MLP or GRU) is fitted at each iteration of a rolling or expanding window. Formally, we define:
where each \(\mathrm{NN}^{(t)}\) is either an MLP or a GRU trained on a specific time block. We assume the hyperparameters (e.g., number of layers, hidden dimension, learning rate, regularization) are selected via TPE-based optimization or a similar approach. Once the network parameters are fixed for iteration \(t\), the calculation of IG-based importances proceeds through four main steps:
At the conclusion of iteration \(t\), the vector \(\widehat{I}_{j}^{(t)}\) is available for all non-dummy features. These values measure how much each variable \(\mathbf{X}_j\) contributes, on average, to the neural network's output relative to \(\mathbf{b}\). We preserve these IG values to:
In neural networks, correlated inputs may lead to partial redundancies in IG scores, analogous to how collinearity affects MDI and GBI, albeit via a different allocation mechanism governed by backpropagation.
Consistent with the principles adopted for Random Forest and XGBoost, we aggregate the IG values across all iterations according to three perspectives:
Following the same logic adopted for MDI or GBI, we establish an importance ranking by ordering the features according to their average IG score in the perspectives described above (global, sub-periods, quarterly). Predictors with high \(\overline{I}_{j}\) values consistently appear as principal contributors, whereas variables whose IG scores remain close to zero signify a limited effect on the neural network's output. Because IG measures how the output changes relative to the baseline, these values are strictly model-specific and do not provide information on whether increments in \(\mathbf{X}_j\) associate with a positive or negative effect on the target.
Adopting an integrated-gradients-based feature importance in neural networks offers several benefits:
Similarly to other model-specific metrics (MDI, GBI), IG-based importances do not inherently convey the sign of the underlying relationship, nor are they entirely immune to correlation-related distortions. Nevertheless, this methodology remains a powerful tool for illuminating predictors' intricate and evolving significance in deep or recurrent neural architectures Rudin2019.
\phantomsection
In this study, we calculated confidence intervals for feature importance across all models, including linear ones, instead of relying on p-values. This approach provides a consistent measure of uncertainty for feature importance, regardless of the type of model used. Specifically, p-values are not readily computable for models such as XGBoost or GRU, making it challenging to assess the stability of feature importance in those cases WassersteinLazar2016. By employing confidence intervals for all models, we ensure a common foundation for comparison, allowing us to evaluate the stability of feature importance across a diverse range of model types. This methodological choice strengthens the comparability of our results, offering a unified framework for interpreting feature importance across different predictive models Shmueli2010,Molnar2022.
The construction of "confidence intervals for feature importance" represents an essential step in quantifying the uncertainty surrounding the contribution of each predictor in macroeconomic models. This process provides a more robust interpretation of the role of variables in the model's predictive capacity BreimanTwoCultures2001.
In contexts where residuals exhibit temporal dependencies or structural breaks, advanced methods such as the block bootstrap Kunsch1989,PolitisRomano1994,Lahiri2003,PanPolitis2016,Masarotto1990 estimate these confidence intervals, offering a more reliable measure of feature importance. Specifically, we adopt a pair block bootstrap, meaning that each resampled block contains both the explanatory variables $\mathbf{x}_t$ and the target variable $y_t$, thus preserving the simultaneous temporal structure of input-output data. As previously described in Section (ref), where a pair block bootstrap was used for prediction intervals, this technique ensures that the temporal structure of the series is respected, thereby providing a resampling procedure agnostic to the distribution of the residuals.
The block bootstrap procedure can be broken down into the following steps:
This procedure is also applied iteratively within an Expanding+Rolling framework. The best hyperparameter configuration is determined in each iteration, and the model is retrained on the updated historical data. After the hyperparameters are fixed, feature importance for the next test set is computed, and the block bootstrap is again employed to evaluate the sensitivity of the feature importance estimates to possible breaks or variations in the training data. This iterative process ensures that feature importance measures are continuously refined and accurately reflect changes in the data structure over time.
Furthermore, the average width of the confidence intervals for feature importance is calculated across multiple sub-periods (Pre-COVID, COVID, Post-COVID, Excluding-COVID). That enables an assessment of how structural changes, such as economic crises or pandemics, influence the uncertainty associated with the feature importance measures. The analysis allows us to evaluate the robustness of feature importance estimates under varying macroeconomic conditions and to understand how global disruptions impact the model's sensitivity to predictors. By comparing these intervals across different phases, we gain valuable insights into the dynamic nature of feature importance over time.
The "confidence intervals for feature importance" calculated in this manner provide reliable, valid uncertainty measures across various models. These intervals offer a deeper understanding of the variability in feature importance, serving as an essential tool for assessing the stability of feature importance across different model configurations and historical periods. This process contributes to a more informed and robust analysis of the role of predictors in macroeconomic forecasting.
\phantomsection
The selection of a coherent subset of forecasting models constitutes a critical prerequisite for effective ensemble methods. In particular, macroeconomic forecasting often encompasses multiple candidate specifications (linear and nonlinear, parametric and nonparametric), each potentially suited to different phases of the economic cycle. Consequently, identifying a "Superior Set of Models" (SSM)\footnote{The term "Superior Set of Models" (SSM) refers to the subset of candidate models that, according to the "Model Confidence Set" (MCS), cannot be statistically rejected as inferior at a chosen significance level. In our application, we set this level to $\alpha$ = 0.10. White2000,Hansen2005,HansenLundeNason2011} in a statistically rigorous manner is essential for mitigating forecast uncertainty and avoiding the inclusion of systematically underperforming models. In this section, we adopt the "Model Confidence Set" (MCS) framework HansenLundeNason2011 to fulfill this task. We first outline the theoretical underpinnings of the MCS approach and then detail its step-by-step implementation. Our exposition aims to clarify how the MCS can be applied effectively in a time-series context, ensuring that temporal dependence is duly respected.
A Model Confidence Set is a statistical procedure identifying the best subset of models based on predictive accuracy. It tests, in an iterative and statistically controlled manner, whether any model in a pool is inferior\footnote{"Inferior" here indicates that a model's predictive loss is significantly higher than that of at least one competing model. Specifically, we say that model $M_j$ is "inferior" if the p-value computed from the MCS bootstrap procedure falls below our significance threshold $\alpha$ = 0.10. In other words, "significantly higher" means that the estimated difference in losses (relative to at least one other model) is unlikely to have arisen by chance under the joint null hypothesis of equal predictive ability. HansenLundeNason2011} to the others by a statistically significant margin.\footnote{A "statistically significant margin" is any observed difference in mean losses between two models (or between one model and the group) that, after block bootstrap resampling, yields a test statistic with an associated p-value < $\alpha$. In our empirical implementation, $\alpha$ = 0.10, so significance implies that we reject the null hypothesis of equal performance at the 10% level.} In the presence of $K$ candidate models $\{M_1,\dots,M_K\}$, the procedure proceeds as follows:
In adopting this iterative screening, the MCS obviates the need for ex-ante assumptions about model validity; instead, it relies on evidence from the block bootstrap to identify those models that do not show persistent inferiority relative to others AiolfiTimmermann2006.
Below, we summarize the main conceptual steps followed in our methodological application:
By structuring the procedure in these steps, we ensure that each model is subjected to a consistent, statistically rigorous evaluation. The block bootstrap is especially vital in macroeconomic contexts, where serial correlation in forecast errors is the rule rather than the exception.
The MCS focuses on comparative performance across multiple forecasts: it does not proclaim a single, absolute "best model," but rather identifies those models for which there is insufficient evidence of inferiority\footnote{"Insufficient evidence of inferiority" means that, within the bootstrap-based hypothesis testing framework, the loss differentials do not exhibit a consistent or statistically robust indication that a given model is outperformed by any other model in the set at $\alpha=0.10$. HansenLundeNason2011} relative to the rest of the group. This balanced perspective is aligned with common practices in forecast evaluation, especially in real-time macroeconomic settings characterized by model uncertainty.
Including only the surviving models from the final SSM in subsequent ensemble methods has several advantages:
While the MCS does not claim that any included model is definitively superior in all circumstances, it ensures that none of the retained models shows statistically validated inferiority at the chosen level (i.e., $\alpha=0.10$). In other words, the methodology offers a robust filter, grounded in time-series bootstrap inference, before forming aggregated nowcasts. This systematic screening aligns with our broader principle of reducing model risk by leveraging iterative hypothesis testing and bootstrap-based diagnostics, and it is computationally more demanding than standard pairwise tests but furnishes a more reliable subset of forecast candidates.
In the subsequent sections, we illustrate how these retained models (i.e., the final Superior Set of Models) are combined through various aggregation techniques, including unweighted and weighted averaging. Such a two-stage strategy---MCS selection followed by forecast aggregation---helps balance the competing goals of robust performance (through screening) and informational breadth (by blending multiple specifications).
\phantomsection
The search for robust predictive accuracy often transcends the reliance on a single specification, even when the model under scrutiny emerges from rigorous selection procedures such as the Model Confidence Set (MCS). Indeed, the practice of combining forecasts traces back to seminal studies arguing that a suitable aggregation of predictions from multiple parsimonious models can outperform (in terms of bias or variance reduction) any single best contender\footnote{In the literature, this approach is frequently referred to as "forecast combination methods," "aggregation models," or "ensemble models." In the present work, we adopt the term "combination models" to avoid confusion with standard machine-learning "ensemble methods," such as random forests or gradient boosting, which we already employ as base learners in certain specifications.}BatesGranger1969. Recent refinements reinforce this intuition, suggesting that pooled forecasts mitigate specification risk in dynamic and potentially unstable contexts Timmermann2006.
In macroeconomic scenarios, structural shifts, regime changes, and incomplete information can undermine the stability of individual estimators. Aggregating the predictions of multiple "superior" models---each capturing distinct data features or leveraging different identification strategies---can help offset model-specific fragilities. Furthermore, exogenous shocks or evolutions in data collection processes may cause an abrupt deterioration in the performance of even well-tuned estimators. A pooled strategy, by design, spreads the risk of relying excessively on a single specification. Consequently, these combination models provide a pragmatic safeguard, especially where the sample size is modest, or the pattern of economic indicators is subject to frequent fluctuations AiolfiTimmermann2006.
In the present work, we focus on three principal techniques for combining the forecasts of the MCS-retained models:
While differing in the explicit formulation of the weighting scheme, these three methods share a common architecture: they revisit and aggregate the predictions at each time step, recording residuals and partial metrics (e.g., MSFE, RMSFE, MAFE) in a rolling or iterative fashion. Such a parallel structure enables a coherent comparison of strengths, limitations, and overall forecast accuracy across the aggregation approaches.
The next subsections elaborate on each of the aforementioned combination strategies. For each, we proceed as follows:
These approaches offer a nuanced perspective on balancing multiple models' relative merits and minimizing misspecification's impact on final nowcasts. By systematically updating (or recalibrating) the weighting structure, we aim to accommodate potential breaks in economic relationships while preserving the benefits of diversification. In the subsequent subsections, we detail each technique's design and present the respective empirical findings in a standardized format, thus ensuring a direct comparability of results. We also compare WA forecasts with both simpler and more advanced approaches in order to determine whether partial RMSE-based weighting yields measurable improvements in predictive stability under diverse economic conditions.
\phantomsection
A Simple Average (SA) is the most straightforward strategy for aggregating multiple model forecasts. It assigns each base learner (model) the same weight, thereby avoiding discrimination based on historical performance or model complexity. The underlying rationale is to minimize the risk of over-reliance on a single specification by evenly distributing the contribution of each model in the final forecast BatesGranger1969, Clemen1989, SmithWallis2009. Consequently, the SA approach offers a transparent and stable benchmark, particularly valuable when the pool of candidate models exhibits comparable performance or when no clear ex-ante indication suggests favoring a specific learner. Additionally, in contexts characterized by abrupt regime shifts or data outliers, uniform weighting can provide a robust baseline that reduces the influence of sudden performance degradations from any one component model MakridakisEtAl2020_M4.
From a statistical standpoint, the Simple Average can reduce variance by blending forecasts from distinct estimation methods, data transformations, or structural assumptions Armstrong2001, StockWatson2004, Timmermann2006. This property is especially desirable in situations marked by parameter instabilities. Moreover, the mechanism does not require frequent re-estimation of weights. Once the set of base models has been determined, each forecast is given an identical portion of the overall "voting power." Although the method disregards potential differences in predictive accuracy, it often yields robust outcomes in empirical settings where economic or financial indicators are subject to substantial noise. In principle, it would be possible to introduce a dynamic aspect to the weights, but this would deviate from the uniform philosophy at the core of SA. In our script implementation, partial error metrics (such as MSFE or RMSFE) are tracked at each time step purely for diagnostic or comparative purposes, without feeding back into the weights. As such, any deterioration in the performance of a specific model does not prompt a reallocation of weights HyndmanAthanasopoulos2021, ChenHood2020.
Let the MCS procedure select a subset of $M$ forecasting models, indexed by $m \in \{1,2,\dots,M\}$. At each forecasting period $t$, each model $m$ produces a predicted value $f_{m,t}$. Under the Simple Average aggregator, the final combination $\widehat{y}_{t}$ is defined as a uniform mean of all base forecasts: \[ \widehat{y}_{t} \;=\; \frac{1}{M}\sum_{m=1}^{M} f_{m,t}. \] Hence, all models are assigned the weight $1/M$, regardless of their historical performance or estimation strategy. When the size of the model set $M$ remains fixed over the evaluation window, the aggregation proceeds identically for all time steps GenreEtAl2010, Zhemkov2021.
Partial metrics, such as MSFE or RMSFE, are often computed at each time step to assess accuracy patterns or to compare with more complex weighting schemes. However, these statistics do not feed back into the weights for SA, since the method preserves the same assignment ($1/M$) throughout the sample. Thus, no dynamic mechanism adjusts any coefficient in response to recent forecast errors. This design choice allows a direct comparison with methods such as Weighted Average or Exponentially Weighted Average, which do update their weights as new data arrive MakridakisHibon2000, Hansen2008, White2000.
This weighting scheme offers two principal advantages:
In this study, the SA method is applied offline (batch mode) to the models retained by the preceding MCS-based selection. Specifically, once the MCS step determines which learners are statistically superior under the chosen loss function, each of these $M$ models contributes equally to generate the final aggregated nowcast. The procedure operates as follows:
The Simple Average provides a robust and parsimonious benchmark by obviating the need for a dynamic adjustment mechanism. Its forecasts may, in certain instances, rival or surpass those derived from more elaborate procedures, especially in moderate samples or under conditions where no single approach maintains a lasting predictive edge. Later sections compare this uniform aggregator with adaptive schemes, investigating whether more sophisticated weighting improves accuracy and stability across diverse economic environments.
\phantomsection
The Weighted Average (WA) combination strategy adjusts each model's influence according to its observed performance over a historical window Clemen1989,Timmermann2006. In contrast to a simple average, WA grants proportionally significant weight to models displaying lower cumulative forecasting errors BatesGranger1969,StockWatson2004. It relies on the idea that models with consistently minor deviations from actual outcomes are more likely to capture the prevailing economic dynamics.
The core principle of WA is that the accuracy of each model, when measured and tracked over multiple periods, furnishes a credible basis for guiding future weights Armstrong2001,Atiya2020. In practice, one first defines a performance metric, such as the partial root mean square error (RMSE), which reflects how effectively a model has predicted the target variable from the initial period up to the current time. Models with lower partial RMSE are considered more reliable and receive higher weights.
Let there be \(M\) distinct models and \(d_{m,t}\) the in-sample forecast error (in our case, the squared difference between observed and predicted) for model \(m\) at time \(t\). A partial RMSE, spanning the interval \([1, \dots, t]\), can be inverted to yield a factor \(\alpha_{m,t}\) that increases with model quality\footnote{In this context, "increases with model quality" means that lower partial RMSE values correspond to higher \(\alpha_{m,t}\). Hence, a model exhibiting a smaller RMSE up to time \(t\) is deemed to have performed better historically, thus warranting a proportionally greater weight in the combination.}:
\[ \alpha_{m,t} = \frac{1}{\mathrm{RMSE}_{m,t}} \quad \text{where} \quad \mathrm{RMSE}_{m,t} = \sqrt{\frac{1}{t} \sum_{\tau=1}^{t} d_{m,\tau}}. \]
If \(\mathrm{RMSE}_{m,t}\) is especially small, \(\alpha_{m,t}\) becomes correspondingly large, signifying robust past performance. After computing \(\alpha_{m,t}\) for each model, one applies a normalization step to ensure the sum of the weights to unity:
\[ w_{m,t} = \frac{\alpha_{m,t}}{\sum_{j=1}^{M}\alpha_{j,t}}. \]
The coefficient \(w_{m,t}\) represents the proportion of the total ensemble weight allocated to model \(m\) at time \(t\).
Once these weights are established, the combined WA forecast for time \(t\) is formed as a convex linear blend of the base models:
\[ \widehat{y}_{t} = \sum_{m=1}^{M} \, w_{m,t} \, f_{m,t}, \]
where \(f_{m,t}\) is the prediction generated by model \(m\). Whenever new data becomes available, each model's partial RMSE is updated to include the latest forecast error, and a fresh set of weights is derived. This routine encourages models with deteriorating accuracy to lose weight, while those maintaining a strong alignment with the observed values see their weight increase.
The main advantage of WA lies in its direct, performance-driven rationale: the better a model has performed recently, the more heavily it is relied upon for the ensemble forecast. Should a model's forecasts begin to diverge from actual realizations, its partial RMSE grows, automatically reducing its relative weight. Conversely, a model whose predictions remain close to observed values achieves a smaller partial RMSE, thereby gaining a larger share of the overall combination GenreEtAl2013.
In this study, we apply WA to the base models retained by the Model Confidence Set (MCS) HansenLundeNason2011. The key steps are:
We perform these steps in an offline (batch) fashion Tashman2000, retrospectively simulating how the weights evolve each quarter. Each time new macroeconomic data become available, the entire sample is reprocessed, allowing the method to reflect changes in model performance over differing economic regimes.
The principal strength of WA lies in its simplicity and interpretability: the weights have a tangible meaning, namely that models performing better in the recent past receive proportionally higher emphasis. Moreover, the weights adjust automatically when a model's accuracy degrades, discouraging continued reliance on outdated performance.
Nonetheless, WA might fail to capture abrupt regime shifts if partial RMSE takes too long to incorporate recent, more severe errors. In highly volatile environments, a swifter reweighting mechanism may be necessary: supposing that multiple models exhibit nearly indistinguishable performance, their weights can converge to a near-uniform distribution, effectively approximating a simple average and thereby limiting the added value that a more selective weighting scheme would otherwise provide.
The Weighted Average (WA) combination adopted in this study adheres to these guidelines:
This procedure allows a quarter-by-quarter reconstruction of how partial RMSE and weights would have evolved in real-time. Whenever new macroeconomic data become available, the entire procedure is repeated so that the WA weighting mechanism remains responsive to recent patterns of model accuracy.
\phantomsection
The Exponentially Weighted Average (EWA) constitutes a flexible mechanism for combining multiple forecasting models, each distinguished by its own pattern of predictive errors LittlestoneWarmuth1994,DevaineEtAl2013. At a conceptual level, EWA assigns greater weight to those models that have demonstrated superior accuracy in the most recent past. Formally, at every time step, the method incorporates an exponential decay factor into each model's cumulative loss, progressively attenuating the impact of older forecast errors and allowing the ensemble to adjust promptly when new information arrives.
Macroeconomic series are prone to regime changes and structural breaks that can weaken the predictive capabilities of a single forecast combination with static weights ClarkMcCracken2010,TianAnderson2014. The EWA algorithm addresses this limitation by weighting models according to their latest performance, penalizing recent inaccuracies more heavily than older ones Timmermann2006. It is grounded in the notion that, under evolving economic dynamics, an estimator's most reliable gauge of current accuracy is its error profile in the latest observations Rossi2021.
From a formal perspective, let \(m\in\{1,\dots,M\}\) index the base models selected via the Model Confidence Set (MCS). At time \(t\), each model \(m\) accumulates a loss \(\mathcal{L}_{m,t}\), which could be measured as a sum of squared forecast errors up to period \(t\). An exponential weighting function then transforms these cumulative losses into a set of normalized coefficients:
\[ w_{m,t} \,=\, \frac{\exp\!\bigl[-\,\eta\,\mathcal{L}_{m,t}\bigr]}{\sum_{j=1}^{M} \exp\!\bigl[-\,\eta\,\mathcal{L}_{j,t}\bigr]}, \]
where \(\eta>0\) denotes a learning rate (or decay parameter). It determines the speed at which the EWA scheme adapts to new forecast errors: larger \(\eta\) values amplify the penalty for recent mistakes, enabling quick revisions of weights; smaller values yield a more gradual adjustment. The coefficient \(w_{m,t}\) represents the fraction of the total weight allocated to model \(m\) at time \(t\). In all cases, these weights remain nonnegative (\(0 \,\le w_{m,t}\le 1\)) and sum to unity (\(\sum_{m=1}^M w_{m,t} = 1\)), thus yielding a valid convex combination. Because \(w_{m,t}\) depends inversely on \(\mathcal{L}_{m,t}\) through the exponential term, models that exhibit smaller cumulative losses (i.e., better historical performance) receive higher weights and those with larger losses see their weights diminished.
To generate a forecast at time \(t\), EWA calculates:
\[ \widehat{y}_{t} \,=\, \sum_{m=1}^{M} \,w_{m,t}\,f_{m,t}, \]
where \(f_{m,t}\) is the one-step-ahead forecast from model \(m\). Once the actual \(y_t\) is observed, each model's loss (e.g., squared error, \((y_t - f_{m,t})^2\)) is updated accordingly, and the new cumulative losses feedback into the weighting scheme. EWA proves particularly valuable in contexts where distinct models exhibit comparative advantages in different macroeconomic regimes. An approach that excels during expansions may temporarily lose favor if it underperforms during downturns, allowing more robust models in recessive phases to gain increased weight in the ensemble RafteryEtAl2010,deRooijEtAl2014. Hence, EWA systematically adjusts the relative influence of each model as the economic environment evolves without relying on prespecified thresholds or manual interventions.
In our analysis, we employ EWA to merge the forecasts of those models (also referred to as "experts" CesaBianchiLugosi2006 retained by the MCS procedure. At each quarterly step, we:
While EWA can naturally operate in an online context---updating weights at each incoming observation---we implement it offline (batch mode) for consistency with our quarterly data releases Tashman2000. Specifically, in each experiment, we retroactively "simulate" the time evolution of weights by sequentially introducing historical observations and recomputing the ensemble's forecast at each step.
We further enhance EWA by including a "meta-layer\footnote{We refer to it as "Meta-EWA" because, after computing separate EWA forecasts for multiple values of \(\eta\), we treat these parallel EWA aggregators themselves as "meta-experts" in a second-tier exponential scheme. This meta-layer automatically learns which decay parameter performs better at each stage, attenuating the risk of using a single \(\eta\) poorly suited to the entire sample.}" that accounts for multiple values of \(\eta\). Each distinct learning rate \(\eta_j\) is treated as defining a separate EWA aggregator, wherein the base experts' weights are updated according to \(\eta_j\). Concretely, for each \(\eta_j\in\{\eta_1,\eta_2,\dots,\eta_K\}\):
Subsequently, we interpret these \(K\) aggregated forecasts as "meta-experts," each with its own incurred error (i.e., the difference between its aggregated forecast and the realized outcome). A second-tier exponential weighting framework then assigns mass across these \(\eta\)-based aggregators, using an analogous rule:
\[ \omega_{j,t} \,=\, \frac{\exp\!\bigl[-\,\lambda\,\widetilde{\mathcal{L}}_{j,t}\bigr]}{\sum_{h=1}^{K}\exp\!\bigl[-\,\lambda\,\widetilde{\mathcal{L}}_{h,t}\bigr]}, \]
where \(\widetilde{\mathcal{L}}_{j,t}\) denotes the cumulative forecast loss for aggregator \(j\) up to time \(t\), and \(\lambda\) is an additional learning rate controlling how aggressively we switch among different \(\eta_j\)-forecasters. By combining these \(\eta\)-aggregators via a second-level EWA, the "Meta-EWA" mechanism mitigates the risk of misspecification that arises from fixing a single decay parameter \(\eta\). This layered scheme automatically identifies the \(\eta_j\) that best adapts to the prevailing data pattern, assigning a proportionally higher weight to its associated aggregator Wolpert1992,Breiman1996,MonteroMansoEtAl2020.
Regarding uniform weighting, EWA provides a more dynamic mechanism that promptly favors models with lower recent errors. Consequently, EWA's adaptability may result in superior out-of-sample performance in volatile or rapidly shifting environments BallingsEtAl2019. Because of these qualities, EWA remains widely adopted whenever the real-time responsiveness of weights is deemed essential. Its recursive structure ensures the ensemble accommodates new information without discarding the historical record entirely. As a result, it aligns well with macroeconomic nowcasting contexts, where the timely integration of evolving economic indicators is of primary concern.
The EWA combination in this study adheres to the following protocol:
These steps are repeated over the entire sample in an offline simulation, enabling a quarter-by-quarter reconstruction of how weights and forecasts would have evolved in real-time. Whenever new macroeconomic data become available, we rerun this offline procedure so that the EWA weights remain anchored to recent error patterns. In subsequent sections, we compare EWA outcomes with those from simpler weighting schemes, thereby assessing whether exponential adaptation enhances predictive accuracy under various macroeconomic conditions.
\phantomsection
Weighted Average (WA) and Exponential Weighted Average (EWA) models offer robust frameworks for combining forecasts, capturing the dynamic contributions of individual models over time. These aggregation techniques are especially beneficial in time-series forecasting, where understanding each model's evolving relative importance enhances the accuracy and interpretability of predictions. By focusing on their shared architecture, we can highlight the conceptual transparency of these methods, which allows for a clear explanation of how each nowcasting model's relative importance changes as economic conditions evolve.
Both WA and EWA follow a similar process for weight assignment, which can be broadly outlined in the following steps AiolfiTimmermann2006, GenreEtAl2013, Hansen2008:
In this shared framework, Bates and Granger BatesGranger1969 originally proposed calculating each model's relative importance by inverting its partial RMSE (Weighted Average, WA). A later work introduced exponential decay factors to reduce the influence of older forecast errors (Exponentially Weighted Average, EWA) DevaineEtAl2013, LittlestoneWarmuth1994. Despite these mathematical differences, the interpretability of both methods is rooted in the same principle: if a model performs poorly over a specific period, its weight diminishes, and vice versa.
To explicitly explain how the weights are calculated and evolve at each time step \( t \), we break down the process as follows:
1. Forecast Errors Calculation (Update): the forecast errors for each model \( m \) are computed based on the difference between the model's prediction \( f_{m,t} \) and the observed target \( y_t \). The forecast error is given by:
\[ d_{m,t} = |f_{m,t} - y_t| \]
For WA, the error is simply the root mean square error (RMSE) calculated over a rolling window of observations. For EWA, the errors are accumulated into a cumulative loss function, which is then exponentially decayed as follows:
\[ \mathcal{L}_{m,t} = \sum_{\tau=1}^{t} \exp\left(-\eta(t-\tau)\right) \cdot d_{m,\tau} \]
2. Normalization of Performance Indicators (Normalize): the forecast errors are then normalized to ensure the final weights are on a comparable scale. This step is critical for WA and EWA, ensuring that all weights are positive and sum to 1. For WA, this is done by computing the inverse of the RMSE for each model:
\[ w_{m,t}^{WA} = \frac{1 / \text{RMSE}_{m,t}}{\sum_{j=1}^{M} 1 / \text{RMSE}_{j,t}} \]
For EWA, the cumulative losses are normalized using the exponential loss formula:
\[ w_{m,t}^{EWA} = \frac{\exp\left(-\eta \mathcal{L}_{m,t}\right)}{\sum_{j=1}^{M} \exp\left(-\eta \mathcal{L}_{j,t}\right)} \]
3. Reweighting Forecasts (Reweight): once the normalized weights are calculated, they are applied to the individual forecasts of each model. This reweighting reflects the relative performance of each model, with better-performing models being assigned a higher weight. The combined forecast for the time step \( t \) is then given by:
\[ \widehat{y}_{t} = \sum_{m=1}^{M} w_{m,t} f_{m,t} \]
This methodology ensures that the weights are updated in real-time (or on a rolling basis in a batch setting), reflecting the dynamic contribution of each model's forecast based on its recent accuracy.
Our analysis used "stacked area plots" to visually represent how model weights evolve and WA and EWA's interpretability. These plots allow a clear visual interpretation of the evolution of the weights assigned to each model at every time step as new data becomes available:
Using this visualization technique, we can track how each model adapts to changes in the economic environment. For example, a sudden increase in the weight of a specific model may suggest that it is better suited to the current economic regime. In contrast, a decrease in its weight may indicate underperformance. These visual tools enhance the interpretability of both WA and EWA models, allowing users to gain deeper insights into the underlying decision-making process.
Although both WA and EWA share the core logic of Update $\rightarrow$ Normalize $\rightarrow$ Reweight, they differ in how they penalize recent errors. WA relies on an inversion of errors, where the model's weight is inversely proportional to its RMSE GenreEtAl2013. In contrast, EWA applies an exponential decay factor to cumulative forecast errors, gradually reducing the influence of older data DevaineEtAl2013. This difference makes EWA more sensitive to abrupt changes in model accuracy, allowing it to respond more quickly to deteriorations in predictive performance. On the other hand, WA can appear more stable, requiring a more extended period of consistent underperformance before a significant weight adjustment occurs Atiya2020.
In contrast to WA and EWA, the Simple Average (SA) aggregator assigns an equal weight to each model and does not modify these weights over time. As a result, SA lacks the dynamic dimension that enables the interpretability of WA and EWA. While SA is transparent in its equal allocation of weights across models, it does not provide insight into which models perform better at any given stage since the weights remain fixed Clemen1989, StockWatson2004. Therefore, while SA promotes diversity by giving equal importance to all models, it does not allow for the interpretation of how individual model performance affects the aggregated forecast.
In our framework, WA and EWA models provide a comprehensible method for aggregation model interpretability by revealing a time-dependent set of weights that indicate which models are deemed most reliable at each forecasting period HansenLundeNason2011, SamuelsSekkel2017. These methods enable an understanding of how the combined forecast is formed and offer a quantitative measure of the evolving trust placed in each model. These features foster transparency that is particularly valuable in economic and financial contexts, where understanding the sources of predictive power is as important as achieving high levels of forecast accuracy MakridakisHibon2000, StockWatson2011, WangEtAl2023.
\phantomsection
The robustness of macroeconomic forecasting models, particularly those based on combined approaches such as Simple Average (SA), Weighted Average (WA), and Exponentially Weighted Average (EWA)Clemen1989,GenreEtAl2013,WangEtAl2023, necessitates thorough diagnostic checks. These tests ensure that the models' residuals do not exhibit systematic patterns that would undermine the validity of the forecasts. In this section, we focus on two widely employed diagnostic tests: the Shapiro-Wilk test for normalityShapiroWilk1965 and the Ljung-Box test for autocorrelation in the residualsLjungBox1978. Both tests were implemented in the context of the model combination strategies discussed earlier, offering a systematic evaluation of their predictive performance.
\phantomsection
The Shapiro-Wilk testShapiroWilk1965,BarrowKourentzes2016 is a statistical method used to assess whether a sample of data comes from a normally distributed population. This test is particularly relevant in time-series forecasting, where the assumption of normally distributed residuals often underpins various diagnostic checks and the construction of confidence intervals. In the context of combined models, the Shapiro-Wilk test examines the normality of residuals derived from the aggregate forecasts of the SA, WA, and EWA methods.
The Shapiro-Wilk test evaluates the null hypothesis that the model's residuals are normally distributed. It is based on the calculation of a test statistic \( W \), which compares the observed distribution of the residuals with a theoretical normal distribution. A value of \( W \) close to 1 indicates that the residuals do not significantly deviate from normality, while a value considerably less than 1 suggests non-normality.
In our implementation, the residuals from each combination model—SA, WA, and EWA—are extracted after the predictions are generated. The Shapiro-Wilk test is then applied to these residuals, assessing whether they follow a normal distribution. This check is crucial because many forecasting models, particularly those relying on statistical techniques such as ordinary least squares or maximum likelihood estimation, assume normality in the residuals for valid inference.
\[ W = \frac{\left( \sum_{i=1}^n a_i x_i \right)^2}{\sum_{i=1}^n (x_i - \bar{x})^2} \]
where:
In time-series forecasting, especially in our context of nowcasting (from 2017 Q1 to 2023 Q2), the normality of residuals is often a desirable but not always attainable property. The presence of economic shocks (e.g., the COVID-19 pandemic)CarrieroEtAl2022 can lead to distributions with heavy tails or skewness, which deviates from normality. Despite these challenges, the Shapiro-Wilk test remains useful for preliminary diagnostics. Although it assumes independence of residuals (an assumption that is often violated in time-series data, where residuals can exhibit autocorrelation), its application can still highlight significant deviations from normality.
In this study, we justify using the Shapiro-Wilk test as a first diagnostic tool for evaluating residuals, even though it does not account for serial correlation. However, the test's application is more reliable, considering that the data used in this study were iteratively standardized to avoid look-ahead bias and that stationarity issues were fully addressed. Despite the economic shocks during the period, this preparation ensures the test is a valid preliminary check for normality in the residuals. It provides a simple and effective way to detect non-normality, which could suggest underlying issues such as model misspecification or inadequate handling of non-linearity in the data.
The Ljung-Box testLjungBox1978 is a diagnostic tool used to detect autocorrelation in the residuals of a time series model. Autocorrelation in residuals suggests that the model has failed to capture some of the time series structure, which could lead to biased or inconsistent forecasts. This test is crucial when dealing with time series data, where residual autocorrelation may indicate model misspecification or an inadequate model structure.
The Ljung-Box test evaluates the null hypothesis that no autocorrelation exists at any lagged values up to a specified maximum lag \( q \). The test statistic \( Q \) is computed as:
\[ Q = n(n + 2) \sum_{k=1}^q \frac{\hat{\rho}_k^2}{n-k} \]
where:
Under the null hypothesis of no autocorrelation, \( Q \) follows a chi-squared distribution with \( q \) degrees of freedom. A significant value of \( Q \) (i.e., larger than the critical value from the chi-squared distribution) indicates the presence of autocorrelation in the residuals, suggesting that the model may not have fully accounted for the time-series structure.
We applied the Ljung-Box test to the residuals of the combined forecasts produced by the SA, WA, and EWA models. The presence of autocorrelation in these residuals would indicate that the aggregation process might not be optimal and that further refinements in model selection or combination may be necessary.
For each combination model—SA, WA, and EWA—the residuals are first calculated as the difference between the actual observations and the model forecasts. These residuals are then subjected to the Shapiro-Wilk test for normality and the Ljung-Box test for autocorrelation. The implementation steps are as follows:
The results of these diagnostic tests provide valuable insights into the reliability of the combined models' forecasts. In theory, a failure to meet the assumptions of normality or autocorrelation in the residuals would suggest potential issues in the model structure, necessitating further refinement or reconsidering of the combination approachTashman2000.
Applying the Shapiro-Wilk and Ljung-Box tests to the residuals of the combined models (SA, WA, EWA) ensures that the residuals are approximately normal and uncorrelated—conditions essential for the validity of many econometric techniques used in forecasting. The outcomes of these tests allow for the identification of potential issues within the aggregated models, guiding subsequent refinements, and enhancing the reliability and robustness of the forecasts. This process is crucial for providing accurate and dependable insights and informed macroeconomic decision-making.
\phantomsection
The evaluation of predictive performance is essential in macroeconomic forecasting, where consistently assessing models' relative accuracy and ability to outperform benchmark models is critical. This section discusses the application of the "Giacomini-White test"GiacominiWhite2006 to compare the predictive ability of the best individual models selected through the Model Confidence Set, and the combined models in comparison with our three reference models (Random Walk, AR(3), and Dynamic Factor Model). We aim to evaluate whether our individual and combined models (SA, WA, EWA) significantly improve predictive accuracy over these selected benchmarks.
The Giacomini-White test is a statistical method used to compare the predictive performance of a candidate model against a benchmark, testing whether there are significant differences in forecasting accuracy. Building on earlier work on forecast accuracy tests such as the Diebold-Mariano testDieboldMariano1995 and the Reality CheckWhite2000, it focuses on conditional predictive ability rather than unconditional comparisonsInoueRossi2012.
The null hypothesis of the Giacomini-White test posits that there is no systematic difference in predictive ability between the candidate and reference models. The alternative hypothesis suggests that at least one model outperforms the other statistically. To implement the test, we compute the difference in forecast errors between the candidate and benchmark models and then regress these differences using ordinary least squares (OLS)Hansen2005.
A key component of this test is using a robust covariance estimator to account for heteroskedasticity and autocorrelation in the residuals. Specifically, we employ the "Lumley-Heagerty covariance matrix" (also called WEAVE, Weighted Empirical Adaptive Variance Estimation)\footnote{The Lumley-Heagerty covariance matrix is a robust "sandwich estimator" designed to adjust the standard errors from an OLS regression for both heteroskedasticity and autocorrelation. It achieves that by applying adaptive weights---often derived through isotonic regression to enforce a monotonic decay---to the autocovariance of the residuals. These weights are then used to construct a more reliable covariance matrix estimator, thereby improving inference accuracy when model assumptions regarding constant variance and uncorrelated errors are violated. For the sake of clarity, a "sandwich estimator" is a succinct term for a robust variance estimator that typically appears in the product form bread $\times$ meat $\times$ bread, where "bread" represents the inverse Hessian of the regression model, and "meat" captures the empirical covariance of the residuals, thus accounting for possible misspecifications and error correlations.}LumleyHeagerty1999, which adjusts for autocorrelation and heteroskedasticity, ensuring reliable statistical inference.
The Giacomini-White test performed in this study follows these steps:
The results of the Giacomini-White test for each best individual and aggregated model are summarized by the "Wald statistic" and the associated p-value\footnote{The R-squared values are not reported, as the primary focus of the Giacomini-White test is on the Wald test and the significance of the intercept term. Since only the intercept is used in the regression model to capture the mean difference in forecast errors, the R-squared is expected to be low and is not critical for interpreting the test results.}. These results provide valuable insights into the relative performance of the models:
The outcomes of these tests are used to determine whether the best individual models and the combination models significantly outperform the reference models in terms of average predictive accuracy. A failure to reject the null might also indicate forecast breakdowns or model instability if combined with further analysisGiacominiRossi2009. However, negative, significant intercepts in the Wald test results here indicate that the models generally outperform the benchmarks in forecasting accuracy, offering robust evidence of their predictive superiority.
In summary, the Giacomini-White test for model comparison, applied in conjunction with the Lumley-Heagerty covariance estimator, provides a robust framework for evaluating the predictive ability of different forecasting models. By comparing the selected best individual models and the combined models against benchmark models, we can confidently assess whether these models offer statistically significant improvements in forecasting accuracy. Using the Wald test and the Lumley-Heagerty covariance ensures that our results are robust to heteroskedasticity and autocorrelation, which are common in time-series data.
\cleardoublepage
\chapter{Empirical Results and Discussion} [opening remarks]
In this subsection, we present the results related to the nowcast accuracy of individual models, evaluating both the Mean Square Forecast Error (MSFE), Root Mean Square Forecast Error (RMSFE), and Mean Absolute Forecast Error (MAFE) metrics in absolute value (Table (ref), first part) and the values of RMSFE normalized to the three benchmarks considered (Random Walk, AR(3), Dynamic Factor Model), as illustrated in the second table (Table (ref)) in correspondence with different macroeconomic regimes.
The main objective is to highlight how each approach manages to contain the nowcast error (MSFE, RMSFE, MAFE) in different macroeconomic contexts and quantify the relative benefits (relative gain) compared to the benchmarks. Below, we provide a summary, isolating the most relevant trends in the overall nowcast period (Overall) and the sub-periods (Pre-COVID, COVID, Post-COVID, Excluding COVID).
The mechanisms underlying these differences and the methodological and operational implications will be explored in section (ref).
\phantomsection
The first table (Forecast accuracy metrics across models and sub-periods) reports the main error measures (MSFE, RMSFE, and MAFE) for each family of models. Although multiple metrics are provided, we focus on the RMSFE as it heavily penalizes large deviations and proves particularly relevant in highly volatile phases, typically observed during macroeconomic crises or turbulent periods. Below, we summarize each model family's behavior across the entire prediction horizon (Overall) and the four sub-periods (Pre-COVID, COVID, Post-COVID, and Excluding COVID), maintaining a qualitative perspective consistent with Table (ref).
{ \fontsize{9pt}{10pt}\selectfont
}
The absolute results, therefore, highlight that Ridge and Elastic Net consistently obtain the lowest and most stable error levels in most circumstances, particularly in phases characterized by volatility and shocks (COVID). PCR exhibits remarkable robustness, specifically during the COVID crisis, but deteriorates notably in the subsequent phase, while PLSR behaves oppositely, suffering significantly during COVID but recovering thereafter. Conversely, LASSO, ensemble methods (RF and XGB), and neural networks show significant nowcasting deterioration during turbulent periods. Among neural networks, GRU demonstrates relatively greater resilience compared to MLP.
\phantomsection
The second table (Prediction accuracy relative to benchmarks across sub-periods (RMSFE ratios)), Table (ref), provides a comparison of the ratios between the RMSFE of each model and that of the three reference benchmarks (Random Walk, AR(3), Dynamic Factor Model). A ratio < 1 indicates an improvement to the benchmark, while a value > 1 corresponds to a relative deterioration.
{ \fontsize{9pt}{10pt}\selectfont
}
Penalized linear models (especially Ridge and Elastic Net) and Dimensionality reduction-based models generally achieve RMSFE ratios consistently below benchmarks across most periods. However, PCR shows significant variability: it displays exceptional robustness during COVID but deteriorates substantially in the subsequent Post-COVID phase. In contrast, ensemble learning and neural network models, though frequently outperforming benchmarks, demonstrate higher overall vulnerability to shocks, particularly pronounced during the COVID period. Among neural approaches, GRU systematically outperforms MLP, highlighting better stability and resilience to macroeconomic disturbances.
The reliability of nowcasting methods is inherently linked to their capacity to accurately measure and represent prediction uncertainty, especially in contexts subject to structural changes or unexpected economic shocks. Prediction intervals offer valuable insights, quantifying how model predictions vary under different macroeconomic conditions. Table (ref) presents the average widths of these intervals, computed across the whole prediction period as well as distinct sub-periods (Pre-COVID, COVID, Post-COVID, Excluding COVID), enabling a nuanced assessment of each model family's consistency and resilience.
Below, we outline the key patterns characterizing prediction uncertainty, highlighting differences and commonalities within and across model families over the analyzed sub-periods:
A detailed examination of the actual and predicted GDP trajectories across model families, visualized through prediction interval plots (see Figures (ref) to (ref)), provides additional insights regarding prediction interval coverage and reliability:
The analysis of prediction intervals reveals distinct uncertainty profiles across model families. Penalized linear models---particularly Elastic Net--- consistently achieve narrower and more stable prediction intervals across most periods, indicating greater reliability in prediction uncertainty quantification, especially during phases of heightened volatility (e.g., COVID). Among Dimensionality reduction-based models, PCR exhibits considerable stability during the COVID crisis but performs less consistently in the subsequent Post-COVID phase. In contrast, PLSR intervals widen notably during the crisis but return to lower levels thereafter. Conversely, ensemble learning methods (particularly Random Forest) and neural network approaches (especially MLP) tend to produce substantially wider prediction intervals, signaling pronounced uncertainty and heightened sensitivity to macroeconomic shocks. Within the neural network family, GRU generally outperforms MLP, maintaining comparatively narrower intervals and exhibiting greater robustness across the evaluated sub-periods.
Examining feature importance offers a complementary perspective to predictive accuracy and uncertainty, enhancing interpretability and transparency and providing deeper insights into the economic mechanisms underlying model predictions. To this end, we investigate how explanatory variables contribute to nowcasting outcomes across different model families, highlighting general trends and specific variations observed across macroeconomic regimes (Overall, Pre-COVID, COVID, Post-COVID, and Excluding COVID). These results, complementing previously presented analyses on predictive accuracy and uncertainty, offer additional insights on the relative relevance assigned to economic indicators by each modeling framework within the macroeconomic nowcasting context.
Figure (ref) to Figure (ref) illustrate the top 10 explanatory variables identified by GRU for each analyzed period, together with their 95% confidence intervals.
The graphical representation clearly shows stability and variability in selecting key predictors across macroeconomic regimes, consistently reflecting the descriptive results summarized in Table Y.
The reported patterns highlight differences in feature importance attribution among Penalized linear, Dimensionality reduction-based, Ensemble learning, and Neural network models across varying macroeconomic regimes.
Penalized linear and Dimensionality reduction-based approaches consistently identify industrial production, trade indicators, and labor market variables among their most influential predictors, maintaining stable relevance across various sub-periods.
Ensemble learning and neural network models display variability in feature importance, reflecting differences in how these methodologies select and emphasize predictors across economic regimes. Such variability is particularly evident during significant macroeconomic fluctuations (e.g., COVID).
Figure (ref) provides an additional perspective to the static analyses presented. It demonstrates, for illustrative purposes, the quarterly evolution of the feature importance calculated via Integrated Gradients for the variable industrial production index change in the GRU model. This representation highlights the dynamic reactivity of the model in correspondence with relevant economic events, such as the onset and development of the COVID-19 pandemic.
To further contextualize the relevance of the explanatory variables that emerged from the previous analyses, Figure (ref) shows the Spearman correlation between each predictor variable and the target variable (quarterly Singapore's GDP growth). This representation highlights that some of the variables consistently selected by the models, such as industrial production, air cargo loaded and pawnshop pledges received, show statistically significant correlations with GDP, empirically confirming the robustness and coherence of the selections made by the analyzed models.
Nowcast combination methods enhance prediction accuracy and robustness, mitigating model-specific errors and uncertainties BatesGranger1969, Timmermann2006, MakridakisEtAl2020_M4. In our approach, we combined predictions from models selected through the Model Confidence Set (MCS) procedure HansenLundeNason2011 into three aggregation strategies: Simple Average (SA), Weighted Average (WA), and Exponentially Weighted Average (EWA). WA and EWA explicitly track dynamic model weights, substantially enhancing interpretability.
The SA approa