Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
91,421 characters · 20 sections · 78 citation commands
Non-linear Phillips Curve for India: Evidence from Explainable Machine Learning
\doublespacing
Accurate inflation forecasts are of prime importance to policymakers, especially central banks that are entrusted with the responsibility of managing monetary policy in an economy (bernanke_2003). A foundational framework within the literature on inflation dynamics is the Phillips Curve (PC) model. The Phillips Curve posits a short-term trade-off between inflation and a measure of economic slack, typically proxied by unemployment rate, such that higher inflation is associated with lower slack in the economy and vice-versa. The earliest empirical validation of this relationship, based on wage inflation and unemployment rate was provided by phillips1958relation for the United Kingdom. Since then, the Phillips Curve framework has undergone significant theoretical advancements, culminating in the development of the micro-founded New Keynesian Phillips Curve (NKPC) taylor1980aggregate, CALVO1983383,gali1999inflation as the workhorse model for inflation analysis. Despite its theoretical appeal, the practical application of the NKPC for inflation modelling and forecasting—particularly within central banks—has been fraught with challenges.
Such difficulties stem from structural breaks, state dependencies, and intrinsic nonlinearities in the relationship between inflation and its fundamental determinants, complicating its empirical validity and predictive performance (see cristini2021nonlinear). These issues are particularly pronounced in case of emerging market economies (EMEs), like India, characterized by rapid structural transformations, evolving labor market dynamics, and inflationary pressures influenced by both domestic and global shocks. Unlike advanced economies with relatively stable inflation dynamics, inflation trajectory in EMEs tend to be shaped by factors such as supply-side constraints, weather shocks leading to food price fluctuations, global oil prices and evolving monetary and fiscal policy frameworks. The recent Covid-19 pandemic further added large-scale volatility in the data adding to the existing challenges for inflation forecasting lenza2020estimate, bobeica2023covid, carriero2024addressing.
Traditionally, such issues have been resolved through non-linear econometric models such as threshold and regime-switching models. The recent integration of machine learning (ML) and econometrics, however, has revolutionized the manner in which these challenges can be addressed (see varian2014big, mullainathan2017machine). Machine learning, a subset of artificial intelligence (AI), leverages advanced statistical techniques to discern patterns and relationships in large and complex datasets. Within economics, ML techniques are now increasingly being employed to forecast macroeconomic and financial indicators. These models have delivered newer insights to stakeholders due to their greater predictive power and the ability to capture nonlinear relationships in the data (athey2019machine, coulombe2022machine). Yet ML methods, despite their proven track record of providing superior forecasts, have most often been critiqued on grounds of being a black-box approach. This does not bode well from a public policy or business perspective wherein identification and quantification of a ‘cause-and-effect’ attains utmost importance (see doshi2017interpretable, miller2017explainable).
To address this critique, emerging research has opened the doors for explainable ML techniques -- novel methods designed to open the black box -- to act as a pivotal bridge between the forecasting power of ML models and the interpretability demanded by practitioners.\footnote{While there is no strict definition of interpretation, it can be defined as the ability to present or explain a given model in terms understandable by humans.} To implement such an explanation for an ML model, one must rely on an algorithm that relates the feature (input) values of the given model with its prediction to generate local interpretation i.e., explain an instance-level prediction. Similarly, one can also generate global interpretation of model output at the level of the sample dataset. By allowing a transparent framework to analyze the impact of each variable on model predictions and understand nonlinear relationship between various variables, such techniques equip both model developers and end users to derive better insights from the ML model.\footnote{In simple terms, interpretability is about developing an understanding of the cause-and-effect within an artificial intelligence (AI) system. In the literature, it refers to the degree in which one can measure or estimate the outcome of a model given an input, understand how the predictions may change with changes in input data or algorithmic parameters and finally understand when the model can go wrong. Explainability, on the other hand, refers primarily to the process that helps us understand in a human-readable form how and why the model came up with the predictive outcome(s).} This paper, therefore, seeks to bring together macroeconomics and machine learning by applying a ML-based forecasting and explanation framework to estimate a nonlinear Phillips curve for India.
Like other EMEs, inflation dynamics in India have been influenced by a host of domestic and global factors. Given its high dependence on oil imports and a large share of agriculture sector, crude oil prices and climate-related factors, such as rainfall, affected inflation in India. Thus, for much of its history, inflation in India has remained high, often times in double digits, lowering the macroeconomic stability of the country. Thereafter, in October 2016, India formally adopted a flexible inflation targeting (FIT) framework to guide its monetary policy in targeting a publicly announced inflation target of 4 per cent (+/- 2 per cent). This marked a significant policy shift in one of the most populous and fastest growing economies in the world.\footnote{The RBI Act was amended in March 2016 to institutionalize the Flexible Inflation Targeting (FIT) regime and the first meeting of the six-member Monetary Policy Committee (MPC) was held in October 2016. See chakravarty_2020, dua_2023 and ghate_ahmed_2023 for more information on the monetary policy landscape in India and the recent transition to FIT regime.} However, after the adoption of the FIT framework, inflation in India returned to lower levels and has remained low since then eichengreen2024inflation. Emerging evidence has highlighted the role of credible monetary policy and effective supply management in contributing to keeping inflation low and stable kishorandpratap2023.
Notwithstanding the success of monetary policy in taming inflation in India, forecasting inflation in India still remains a challenge for policymakers owing to the issues highlighted above. For instance, chart (ref) depicts the dynamic nature of relationship between output gap and inflation in India over the last two decades. While economic theory of the Phillips Curve suggests a positive relationship, the actual relationship in this case has oscillated over time, with the mean estimate being 0.11 across the whole 2000Q1 - 2023Q1 sample. Furthermore, Chart (ref) highlights that the relationship between inflation and output gap in India may have undergone a structural break amidst the policy transition. The above evidence suggests that the PC relationship in the Indian context is complex which may not be adequately captured by simplistic, linear regression models that are conventionally used for modelling and forecasting purposes. Given these complexities, a more nuanced approach to inflation forecasting becomes crucial for policymakers to formulate effective monetary policy responses and ensure macroeconomic stability.
Therefore, in this paper, we use quarterly data ranging from Q1:2000 to Q1:2023, to train a host of linear, Tree-based and Neural Network-based models across various specifications of the Phillips curve. We then undertake a pseudo out-of-sample forecasting exercise to demonstrate that Tree-based models achieve significantly lower forecast errors compared to traditional linear regression models, including regularized models such as the LASSO model. This is reported across one- and multi-step ahead forecast horizons and different Phillips curve specifications. Tree-based models deliver superior predictive performance, even surpassing state-of-the-art Neural Network models. While neural networks offer advanced capabilities, their effectiveness is constrained by the requirement for large datasets, making them less suitable for contexts where data is limited. Our findings remain robust to several robustness checks.
Thereafter, we apply several explainable ML techniques to a hybrid Phillips curve model (which includes both backward- and forward-looking components) to explore the nonlinear relationship between inflation and its structural determinants in the Indian case. To address the "black box" criticism of ML models, we use methods such as feature importance (FI), partial dependence plots (PDP), Shapley values and Shapley regressions shapley1953value, breiman2001random, friedman2001greedy, lundberg2017unified, strumbelj2010efficient, joseph2019parametric, buckmann2023interpretable.
In particular, we compare the Random Forest (RF) model with a Linear Regression (LR) benchmark, focusing on key aspects of the NKPC model, namely inflation expectations, lagged inflation, and the output gap. Our analysis suggests that inflation and its determinants, seen within the Phillips curve framework, share a highly nonlinear relationship. In terms of variable (feature) importance, both models highlight inflation expectations and lagged inflation as key drivers, but the RF model emphasizes expected inflation, while the LR model favours lagged inflation. The output gap, though absent from the LR model's top features, is captured by the RF model, suggesting it better identifies the nonlinear relationship between inflation and economic slack that is missed by the LR model.
PDPs support these findings where inflation expectations display a non-linear, step-like pattern indicating thresholds; lagged inflation appears linear; and the output gap shows a state dependence based on the whether the output gap is positive or negative. We find that there exists more than one threshold -- in the range of 4.75 - 7.25 percent -- around which the level of inflation expectations has an asymmetric impact on realized inflation. Similarly, the influence of backward-looking component in the Phillips curve – lagged inflation – also increases sharply when the level of inflation crosses the range of 5.0 - 6.0 per cent. Two-way PDPs reveal that the RF model is able to capture interaction effect among these variables which contributes to its higher accuracy over the LR model. Furthermore, using a Shapley values-based approach, we provide a global interpretation of our model, finding an asymmetric response of headline inflation to high vs. low values of input features or variables. Shapley dependence plots highlight the non-linear relationships captured by the RF model in contrast to the linear LR model. Shapley regressions for the RF model confirm that expected inflation, lagged inflation, rainfall deviation, and the output gap are the primary drivers of inflation in India.
In what follows, a brief survey of related literature on the Phillips curve as well as ML interpretation and inference is presented in Section (ref). Our approach to model training, forecasting along with a brief discussion of explainable ML techniques is presented in Section (ref). The key findings of our paper including the forecasting analysis and model explanation using ML methods are discussed in Section (ref). Robustness analysis is presented in Section (ref). The paper concludes in Section (ref) while drawing contours for further research.
This study builds upon two primary strands of literature: one on the Phillips Curve, encompassing both theoretical and empirical research on inflation dynamics, and second on the application of machine learning (ML) and explainable AI (xAI) techniques for macroeconomic forecasting. By synthesizing these perspectives, we contribute to a more nuanced understanding of inflation dynamics particularly in the Indian context.
The Phillips Curve, which characterizes the trade-off between inflation and unemployment, was first empirically documented for the United Kingdom by phillips1958relation and later extended to the US by samuelson1960analytical. The introduction of inflation expectations into Phillips Curve model by phelps1967phillips and friedman1968role laid the foundation for the expectations-augmented Phillips Curve. This in turn led to the development of the micro-founded New Keynesian Phillips Curve (NKPC) with seminal contributions by taylor1980aggregate, calvo1983staggered and gali1999inflation. Subsequent empirical studies, such as those by roberts1995new, fuhrer1995phillips and sbordone2002prices tested and refined this framework. Empirical validity of the Phillips Curve, however, has been debated. Influential studies such as ball2011inflation and stock2020slack have questioned its robustness, arguing that structural changes in the economy—such as globalization, labor market rigidities, and higher degree of inflation anchoring—have weakened the relationship between inflation and economic slack. More recent studies, such as hazell2022slope and ball2019phillips, have sought to revive the Phillips Curve by incorporating nonlinearities and state-dependent effects in its modelling and estimation. More recently, doser2023inflation, examine nonlinearities in the Phillips Curve using a threshold regression model and find that once inflation expectations—especially those of consumers—are properly accounted for, a linear specification cannot be rejected. The study shows that perceived nonlinearity often arises from mismeasured expectations rather than inherent structural breaks. While some episodes, like the “missing disinflation,” exhibit mild nonlinear effects, the model’s predictive power remains largely unchanged for key historical inflationary periods. Their findings underscore the central role of inflation expectations in shaping the Phillips Curve, challenging the necessity of nonlinear specifications in most economic conditions.
In the Indian context, early studies on the Phillips curve relationship encompass kapur2000price, srinivasan2006modelling, paul2009phillips, mazumder2011stability, patra2012monetary, and kapur2013phillips. Using various measures of prices and economic slack, most of these studies showed the presence of a distinct Phillips Curve in Indian data. Similarly, recent studies such as patra2014post, ball2016understanding, chinoy2016responsible, behera2018phillips, pattanaik2020inflation and jose2021alternative provide revised estimates for different Phillips curve specifications in the Indian context. These studies, in general, have also highlighted the role of supply-side shocks, especially those affecting food and fuel prices, in impacting inflation in India. mohanty2015determinants and dua2021determinants are related studies that offer a detailed analysis of inflation determinants for India. More recent studies, such as patra2021phillips, explored whether the Phillips Curve remains relevant in India, particularly in the post-pandemic period. Their findings indicate that while the Phillips Curve relationship persists, its slope varies with the level of output gap, flattening when the gap is negative or low and steepening when it is high. A significant limitation in the Indian Phillips Curve literature is its heavy reliance on linear econometric models, which may fail to capture complex, nonlinear relationships between inflation and its determinants. Our study addresses this gap by leveraging machine learning (ML) methods, which allow for greater flexibility in modeling nonlinearities and structural breaks in the Phillips Curve relationship.
Recent advances in machine learning have led to the development of model-agnostic techniques for interpreting complex models and forecasts, enhancing their applicability in macroeconomic research guidotti2018survey, vilone2020explainable, molnar2020interpretable). Methods such as Shapley values for local and global interpretation (shapley1953value, strumbelj2010efficient, lundberg2017unified; lundberg2019consistent), Partial Dependence Plots (friedman2001greedy), and Individual Conditional Expectations (goldstein2015peeking) are increasingly used to improve the explainability of machine learning models. In macroeconomics, ML-based forecasting has gained serious traction, especially in case of inflation, GDP and unemployment forecasting nakamura2005inflation, chakraborty2017machine, hall2018machine, almosova2023nonlinear, malladi2023benchmark. Recent advancements in macroeconomic forecasting have emphasized the need for models capable of handling the complexities inherent in macroeconomic data, including non-linearities, long-range dependencies, and structural breaks.
Several recent studies have also applied these methods to forecast inflation in India pratap2019macroeconomic, singh2022inflation, sengupta2024forecasting. In particular, sengupta2024forecasting introduce a novel algorithm, namely the Filtered Ensemble Wavelet Neural Network (FEWNet), specifically designed to address such challenges. Their model decomposes inflation data into high- and low-frequency components using wavelet transforms while integrating additional exogenous factors, such as economic policy uncertainty and geopolitical risk to improve forecast accuracy. By leveraging wavelet-transformed inputs and filtered variables within an ensemble of autoregressive neural networks, FEWNet effectively captures intricate data patterns and enhances its adaptability to macroeconomic forecasting tasks. The study, applied to the BRIC nations, demonstrated that FEWNet consistently outperforms both traditional linear statistical models and state-of-the-art non-linear algorithms in producing robust long-term (24-month-ahead) inflation forecasts.
Beyond macroeconomic forecasting, ML and xAI techniques are also being used broadly in economic and financial domains to derive fresh policy insights sun2024deep. Using artificial neural networks and SHapley Additive exPlanations, they uncover a nonlinear U-shaped relationship between the digital economy and energy productivity, highlighting regional heterogeneity, lagging effects, and spatial dependencies. Their findings demonstrate how ML-driven insights can enhance policy decisions by capturing dynamic interactions often overlooked in traditional econometric models. Additionally, research on green housing policies offers insights into how economic incentives and policy interventions shape behavioral responses, which can be analogously applied to monetary policy and inflation dynamics. Similarly, li2023social develop a four-party evolutionary game model to examine interactions between government, realtors, and residents under different policy scenarios. Their findings demonstrate that government incentives and regulatory measures influence private sector decisions, leading to market-driven sustainability outcomes. This aligns with our study’s emphasis on state-market interactions in inflation dynamics, where monetary policy interventions affect inflation expectations and labor market responses, reinforcing the importance of nonlinear modeling approaches.
This study advances the Phillips Curve literature in India by addressing key limitations of traditional linear models, which often fail to capture structural changes, supply shocks, and thresold effects in the data. While some studies incorporate state-dependent effects, they lack methodological rigor in modelling nonlinearity and time variation. By integrating ML techniques within the New Keynesian Phillips Curve (NKPC) framework, we provide a more flexible and theory-consistent approach to forecasting inflation. Our study further enhances interpretability by employing explainable AI methods such as Shapley values, Partial Dependence Plots, and Individual Conditional Expectations, offering policymakers with deeper insights into the key drivers of inflation and their nonlinear interactions. This synthesis of Phillips Curve theory, ML-based forecasting, and explainable AI strengthens our understanding of inflation dynamics in emerging economies, ensuring both improved predictive accuracy and greater policy relevance.
The empirical analysis is divided into three parts. In the first step, we train (estimate) a mix of linear regression and non-linear ML models on various specifications of the Phillips curve, namely a backward-looking PC, a purely forward-looking PC and a hybrid PC which incorporates both backward- and forward-looking components. In the second step, we generate multi-step ahead forecasts of inflation using each of the model-specification combination on a test (out-of-sample) dataset. We compare the forecast accuracy of these models against a simple Random Walk (RW) benchmark. The last part focuses on applying various explainable ML techniques to explain the predictions generated by our best performing model. More importantly, this part focuses on analyzing the non-linear relationship between inflation and its determinants in the Indian context.
Following the literature on Phillips curve, three specifications are trained and tested – an adaptive PC, a purely forward-looking PC and a hybrid PC as shown below in equations (1), (2) and (3), respectively:
where $\pi_t$ is inflation, $E_t \pi_{t+1}$ is the inflation expected in $t+1$ period at time $t$, $y$ is the output gap and $X$ is a vector of controls variables, which mainly consists of supply-related indicators. Following the literature on India, X consists of changes in exchange rate, crude oil and rainfall deviation in the present case. The adaptive PC assumes the data generating process of inflation to be backward-looking i.e., determined by its own lagged values. Similarly, the forward-looking PC models inflation to be primarily driven by expected level of future inflation. Finally, the hybrid specification of the Phillips curve considers the inflationary process to contain both backward- and forward-looking components.
The Phillips curve, shown in its linear form in equations (1) - (3), can be described in a nonlinear functional form as shown below. In this case, the nonlinear function \(F(.)\) is approximated by a machine learning model.
We employ a diverse set of models for inflation forecasting, categorised into linear models, non-linear models, and neural network-based models. Linear models include Linear Regression, Ridge Regression, and LASSO, each capturing relationships between predictors and inflation. Ridge Regression adds an L2 penalty to shrink the coefficients, thereby reducing variance, while LASSO uses an L1 penalty to perform both regularisation and feature selection by setting some coefficients to zero. Non-linear models, such as Random Forest and XGBoost, improve upon decision trees by either averaging predictions from bootstrapped trees (Random Forest) or building trees sequentially with gradient boosting (XGBoost). These models can capture more complex, non-linear patterns in the data. Neural network-based models include N-BEATS, which uses a fully connected architecture to decompose forecasts into trend and seasonality components, and N-HITS, which extends N-BEATS by incorporating hierarchical interpolation filters. We also use BlockRNN with LSTM cells, which are effective at capturing long-term dependencies in sequential data. These models can also incorporate external covariates, enhancing their ability to account for exogenous factors influencing the forecast. Additional details regarding these models are presented in the Appendix.
In practical macroeconomic forecasting, we are usually concerned with next period forecasts while using information available up to the current period. To recreate this scenario, we conduct an iterative forecasting horse-race between competing models. This is achieved by fitting a model on data up to quarter $q$ and then predicting the target variable for $q+1^{th}$ quarter. In the next step, the model is fitted on data up to quarter $q+1$ and prediction for $q+2^{nd}$ quarter is made and so on until we extinguish the test set. This process is repeated for each model. To make this exercise as close to real time forecasting as possible, output gap and expected inflation are also calculated iteratively by employing information only up to last available quarter in each iteration. The pseudo-algorithm is presented as below. In the present case, our sample consists of 93 quarters out of which the test set comprise of the last 24 quarters (6 years).
Finally, we compute several error metrics namely root mean squared error (RMSE), median absolute relative error (MdRAE), Symmetric Mean Absolute Percentage error (SMAPE) and Theil's U2 statistic for each of the competing models for comparison. Computational details for each of these metrics is provided in Appendix A.
Interpretability is often a big hurdle for using a machine learning (ML) model in economics. Classical econometric models usually assume a data-generating process beforehand, which a priori imposes constraints on the model but allow for easy interpretation of the model, including its coefficients and forecasts. On the other hand, ML algorithms usually employ nonlinear and non-parametric approaches which results in poor interpretability of the model. Understandably, by choosing nonlinearity over simpler linear methods, ML algorithms tend to be complex and not-so straightforward to understand. This section briefly describes the model-agnostic, interpretability techniques used in the paper.
Permutation importance is a model-agnostic method used to estimate the importance of individual features in predictive models by measuring the decrease in performance when a feature's values are randomly shuffled. Given a trained model \( f \) and a performance metric \( M \), the baseline performance \( M(f, X, y) \) is first computed on the original dataset \( X \). For each feature \( X_j \), its values are permuted(i.e. randomly shuffled), generating a new dataset \( X^{\text{perm}}_j \), while all other features remain unchanged. The model performance \( M(f, X^{\text{perm}}_j, y) \) is then evaluated on this permuted dataset. The importance of feature \( X_j \), denoted as \( I(X_j) \), is defined as the difference in model performance before and after permutation:
This method quantifies the contribution of each feature to the model's predictions. Features with larger importance values have a more significant impact on the model, as their randomization leads to a greater reduction in performance. Permutation importance is widely applicable and does not require retraining the model, making it computationally efficient and suitable for complex models, such as random forests or neural networks. However, it can be sensitive to feature interactions and correlations, which may result in an underestimation of a feature's true importance.
Partial Dependence Plots (PDP) and Individual Conditional Expectations (ICE) are model-agnostic techniques used to visualize the effect of one or more features on the predicted outcome of a machine learning model. PDPs help estimate how a feature influences the model's predictions by averaging the effect of that feature over all possible values of other features. This process marginalizes the effects of other features, providing an expected outcome for any given value of the feature of interest.
Mathematically, given a trained model \( f \), feature set \( X \), and target variable \( y \), the partial dependence of a feature \( X_j \) on the model prediction is computed as the expected prediction of the model, averaging over all other features \( X_{-j} \) (i.e., all features except \( X_j \)). This expectation is taken with respect to the marginal distribution of \( X_{-j} \), denoted as \( P(X_{-j}) \). The partial dependence function is defined as:
where \( \mathbb{E}_{X_{-j} \sim P(X_{-j})} \) denotes the expectation over the distribution of the remaining features. In essence, the PDP reflects the average predicted value for different values of \( X_j \), accounting for the variability of all other features.
In contrast, ICE plots provide a more granular view by showing the individual predictions for each data point in the dataset. Instead of averaging over the distribution of other features, ICE plots keep all other features fixed and vary only \( X_j \) for each observation, offering insights into how different instances react to changes in \( X_j \).
The calculation of Partial Dependence Plots (PDPs) assumes that the features are uncorrelated, an assumption that may not always hold in practice. Consequently, when a feature is highly correlated with other features, the computation of its PDP often involves averaging predictions over synthetic data points that may be implausible or rare in real-world scenarios. This can introduce significant bias in the estimation of the feature’s true effect, potentially leading to misleading interpretations.
When considering interactions between two features \( X_j \) and \( X_k \), a 2-way PDP can be employed. This plot visualizes the joint effect of two features on the model prediction by calculating the expected predictions over a grid of values for \( X_j \) and \( X_k \), while marginalizing over all remaining features \( X_{-j,k} \):
This method helps reveal whether and how the interaction between the two features influences the model's predictions, which cannot be captured by single-feature PDPs.
In summary, PDPs provide a global view of feature effects averaged over the distribution of other features, while ICE plots reveal individual variations. The combination of both techniques helps in understanding both global trends and instance-specific behaviour in complex machine learning models.
Originally introduced in cooperative game theory, Shapley values provide a method to allocate a joint payoff obtained in a cooperative game to the individual players of a coalition based on their contributions shapley1953value. The concept of Shapley values was introduced to machine learning by strumbelj2010efficient as a way of attributing predictions in a supervised model to its constituent features. In the context of ML models, shapley values $\phi_k(x_i;f)$ for a given feature $k$ and prediction $x_i$ measure the marginal contribution of feature $k$ to the prediction, averaged over all possible subsets of the features. Formally, for a model $f$, the Shapley value for feature $k$ in prediction $x_i$ is given by:
where $N \setminus \{k\}$ represents all subsets of features excluding $k$, $S$ denotes the number of variables included in that subset and $n$ is the total number of features.
Intuitive Description: The formula for Shapley values can be understood by imagining each feature as a player in a cooperative game. The goal is to fairly distribute the payout (the model's prediction) among all the features based on their contributions. For any feature $k$, we calculate how much adding that feature changes the model's prediction across all possible subsets $S$ of the other features. This change is weighted by the number of subsets in which feature $k$ could potentially appear, ensuring a fair contribution is assigned to each feature, regardless of the order in which they are considered. The Shapley value $\phi_k(x)$ is essentially the weighted average of these marginal contributions across all subsets.
Shapley regression, introduced by joseph2019parametric, is a statistical inference framework that leverages Shapley values to evaluate the importance of features in machine learning models. The approach allows for conventional parametric inference in nonlinear models by performing a regression of the outcome variable on the Shapley values of the features. This makes it possible to test the statitical significance of individual features in machine learning predictions often treated as "black boxes". The regression model is formulated as:
where \( y_i \) represents the observed outcome for the \( i \)-th observation, while \( \beta_0 \) denotes the intercept of the model. The term \( \phi_k(x_i) \) refers to the Shapley value for the \( k \)-th feature in the \( i \)-th observation, capturing the contribution of feature \( k \) to the model's prediction. The corresponding regression coefficient, \( \beta_k \), quantifies the strength of the relationship between the Shapley value and the outcome.
This regression allows us to estimate the significance of each feature by testing the hypothesis:
where $\Omega \in \mathbb{R}^n$ is the model input space, where \( n \) represents the number of input features.
The key idea of Shapley regression is to capture the importance of each feature in a nonlinear ML model in a way that is analogous to the coefficients in a linear regression model. The Shapley value \( \phi_k(x_i) \) quantifies the contribution of feature \( k \) to the prediction for the \( i \)-th observation. By regressing the observed outcome on these contributions, we are able to make inferences about the importance of features as if we were working with a linear model. The advantage of this framework, as shown by joseph2019parametric, is that it allows for parametric statistical inference on machine learning models, which is typically difficult due to their nonlinearity and complexity.
Additionally, Shapley regression provides a mechanism for calculating the Shapley share coefficients $\Gamma_k$, which measure the relative contribution of each feature across the entire dataset. The Shapley share for feature \( k \) is defined as:
where $sign$ is the sign of coefficients when regressing $y$ on $x$ and (*) indicate the confidence level with which we can reject $H_0$ in equation ((ref)). The interpretation of $\Gamma_k$ is also similar to that of regression coefficient, as it measures the strength, direction and confidence in alignment with the target variable. However, it should be noted that $\Gamma_k$ cannot be interpreted as the marginal effect, unless the model is linear.
Shapley values possess desirable properties such as consistency and efficiency, making them a powerful tool for feature attribution in machine learning models (lundberg2017unified). Furthermore, Shapley regressions bridge the gap between predictive modeling and statistical inference, providing insights into feature importance and significance in a rigorous manner.
The data employed in the study begins in 2000Q1 and ends in 2023Q1. Table (ref) presents a brief description of the data.
We compute the output gap using an Unobserved Component Model (UCM) on seasonally adjusted quarterly GDP series\footnote{Seasonal adjustment was made using the X-13 ARIMA model.}. Similarly, trend inflation is calculated by fitting a UCM model to the quarterly inflation series.\footnote{See watson1986univariate and harvey1993detrending for more details on the UCM model.} The trend component of inflation – shifted forward one-time step – is used as a measure of inflation expectations.
Given that inflation expectations are unobserved, there is no consensus in literature about the best measure of inflation expectations. Market based methods, derive inflation expectations using inflation-indexed bonds, inflation swaps and inflation options. However, such measures may be noisy and reflect risk premia. Further, there doesn’t exist a liquid market for such instruments in the context of India. Secondly, survey measures may also be employed to derive inflation expectations. cecchetti2007understanding provided evidence that survey inflation forecasts are correlated with future trend inflation, measured using a stock2007why UCSV model. However, in the context of India, a long, continuous time-series of survey-based data on inflation expectations is not available or is available at an uneven frequency. Further, several studies have highlighted the need for debiasing inflation expectations derived from responses in the Survey of Households data (muduli2022assessing). Amid these data issues, researchers are forced to look at “second best” options. For example, patra2010inflation fit an ARMA model to observed inflation and use the one-period ahead forecasts as a measure of inflation expectations. Some others like bicchal2019rationality have used Google Trends data to derive inflation expectations for India. We select our measure of inflation expectations keeping in mind the issues presented above. Chart (ref) plots the data series employed in the study.
The analysis of the statistical properties of key macroeconomic variables, summarized in Table 2, reveals several critical characteristics that provide insights into the underlying dynamics of the data and inform the selection of appropriate forecasting methodologies. The data exhibits significant variability, particularly in supply-side factors such as crude oil prices, rainfall deviations, and the output gap, as indicated by high coefficients of variation and entropy values. Most variables, including quarterly inflation, crude oil prices, USD/INR exchange rate, and rainfall deviations, demonstrate the presence of long-range dependence, as identified by the Hurst exponent. In particular, the output gap is the only variable that does not exhibit long-range dependence, reflecting its distinct structural characteristics within the data set. Tests for non-linearity, including Tsay’s test (tsay1986nonlinearity) and Keenan’s one-degree test (keenan1985tukey), indicate that the majority of the variables, except for crude oil prices and trend inflation, exhibit non-linear dynamics. This non-linearity emphasizes the importance of utilizing flexible forecasting models capable of capturing these complex relationships. Additionally, all variables are confirmed to be trend-stationary based on the Kwiatkowski–Phillips–Schmidt–Shin (KPSS) test (kwiatkowski1992testing), ensuring their suitability for econometric modeling. The Ollech and Webel test (ollech2023random) identifies no seasonality in any of the variables, further simplifying the structural assumptions required for modeling. Skewness and kurtosis values reveal that variables such as inflation and output gap are positively skewed, reflecting asymmetric distributions, while crude oil prices exhibit negative skewness. Outlier detection using the Bonferroni outlier test (weisberg1982residuals) based on Studentized residuals reveals the presence of outliers in the supply-side factors (crude oil and exchange rates) and the output gap, indicating the occurrence of sudden and significant shocks in these variables. These findings collectively highlight the complexity and heterogeneity of the dataset, necessitating advanced forecasting approaches, particularly those that can accommodate non-linearities, long-range dependence, and the presence of outliers. This comprehensive analysis serves as a foundation for developing robust econometric and machine learning-based forecasting models tailored to the intricate dynamics of these macroeconomic variables.
As described earlier, we undertake a pseudo real-time out-of-sample forecasting exercise and compute the forecast errors across different models and specifications. The root mean square error (RMSE) for the five models and three specifications for 1- and 4-quarter ahead forecast horizons\footnote{Details forecast error metrics are presented for 1 up to 4-quarter ahead forecasts in Appendix A: Table A1 through A4.} presented below in Chart (ref). Further, RMSE of a benchmark Random walk model is also presented for comparison.
The results reveal that machine learning models, particularly Random Forest and XGBoost, consistently outperform both linear models (Linear Regression, Ridge Regression, and LASSO) and complex neural network models (NBEATS, NHits, and BlockRNN) in forecasting CPI inflation across short-term (1-quarter ahead) and longer-term (4-quarter ahead) horizons. Chart 4 provides a detailed comparison of root mean squared error (RMSE) values across different forecasting approaches for the three Phillips Curve specifications. For 1-quarter ahead forecasts, Random Forest achieves the lowest RMSE under the backward-looking Phillips Curve (1.05) and hybrid Phillips Curve (1.06), while XGBoost delivers competitive performance under the forward-looking Phillips Curve (1.16). The 4-quarter ahead results exhibit a similar pattern, with Random Forest achieving the lowest RMSE under the hybrid Phillips Curve (0.89) and backward-looking Phillips Curve (0.92), and XGBoost leading under the forward-looking Phillips Curve (0.96). These findings underscore the exceptional performance of Random Forest and XGBoost in capturing the complex, non-linear dynamics of inflation, surpassing both simpler linear models and more complex neural network-based approaches. Additionally, the Random Walk model, serving as a baseline, exhibits strong short-term accuracy but is less robust for multi-step forecasts, further highlighting the predictive power of the machine learning approaches.
The superiority of ensemble ML models suggests that their ability to capture non-linearities and complex interactions between variables provides a significant advantage for macroeconomic forecasting. This result highlights the potential for machine learning techniques to enhance the accuracy of macroeconomic models, particularly in contexts where relationships between variables may be more complex than what linear methods can adequately represent.
The Random Walk model, however, outperforms all other models for one-quarter ahead predictions with the lowest RMSE (0.94). Despite this, its performance deteriorates substantially for the 4-quarter ahead forecasts, where it has the highest RMSE (1.74). While the Random Walk model can provide accurate short-term predictions, in the current case, it neither offers any insights into the structural determinants of inflation nor help explain the factors driving inflation forecasts generated at any given time. This limitation makes it less useful for policy analysis where understanding the underlying drivers of inflation or unemployment is critical. This is true even for other univariate time-series models, such as ARIMA models. In contrast, models based on the Phillips curve framework not only provide forecasts but also offer interpretability, helping policymakers understand the economic forces at play, which is essential for informed decision-making. Overall, Random Forest model outperforms in the one-quarter ahead forecasts while XGBoost stands as the best model for 4-quarter ahead forecasts. Specifically, Random Forest model achieves 20-40 percent improvement over the linear and regularised linear regression model. Overall, the hybrid NKPC estimated using the Random Forest algorithm provides the best forecasts for one-quarter ahead forecasts.
Based on the forecasting results, we select the hybrid Phillips curve specification with the Random Forest (RF) model for further analysis and comparison with Linear Regression (LR).
Feature Importance - Feature importance results using Shapley values and Permutation Importance are presented in Chart (ref). Both models highlight inflation expectations and lagged inflation as significant drivers of inflation dynamics. However, the RF model places greater emphasis on expected inflation, as shown by its dominance in both Shapley values and permutation importance analyses. This suggests that the RF model is better equipped to capture forward-looking inflation expectations, which are critical in a policy-driven macroeconomic context. It also highlights the rising role of inflation expectations in determining the inflation process in India.
In contrast, the LR model assigns more importance to lagged inflation, indicating that it relies more heavily on historical inflation data to make predictions. The output gap, a key element in understanding the inflation-economic slack relationship, is notably absent from the LR model's top 10 features but is captured by the RF model. This implies that the RF model is better suited to detect the potential nonlinear interactions between inflation and economic slack, a relationship that the LR model, due to its linear nature, fails to capture. Overall, the RF model demonstrates superior feature identification, particularly in its ability to weigh inflation expectations and the output gap more effectively than the LR model, aligning with its improved forecasting performance under the hybrid Phillips curve specification.
Partial Dependence Plot - The Partial Dependence Plots (PDPs) estimate the marginal effect of a given feature on the predicted outcome of a model, providing insight into the nature of the relationship between the predictor and the target variable. PDPs can reveal whether these relationships are linear, monotonic, or exhibit more complex patterns, allowing for a clearer understanding of modelled relationship. Chart (ref) presents the PDPs for expected inflation, lagged inflation and output gap for both Random Forest and Linear Regression models. While the dotted blue line depicts the mean effect of the variable i.e., the partial dependence plot, the grey lines show the ICE plots for the given variable (feature).
The RF model captures non-linear relationships in the data, as evidenced by the distinct, step-like patterns seen in its PDPs for expected inflation and lagged inflation. These results suggest that the RF model identifies thresholds at which the impact of these variables on realized inflation shifts sharply, a characteristic that cannot be captured by the LR model. For instance, as inflation expectations increase, the prediction for inflation increases but this increase is subject to two thresholds – around 4.75 percent and 7.25 percent – suggesting that expected inflation has a nonlinear impact on actual inflation. Likewise, the ICE plots also presents a visual cue on the level of variation in model predictions when we fix the value of a feature. For example, lagged inflation of around 7.0 percent leads to a prediction that is 0.5 to 2.0 percent higher than the average prediction but such an effect is absent at lower values of lagged inflation. Similarly, the PDP for output gap, generated with the RF model, reflects a more complex interaction indicating potential non-linear relationship between output gap and inflation. Notably, the RF model correctly captures the positive relationship between the output gap and inflation, such that positive (negative) output gap leads to higher (lower) inflation, aligning with the Phillips curve theory and evidence. On the other hand, LR estimates a negative relationship in the sample.
The PDPs for the LR model, unsurprisingly, demonstrate linear relationship between inflation and all other variables in the model. The LR model assumes a constant, additive effect for each predictor, leading to straight-line PDPs that cannot account for any non-linear interactions in the data. This limitation underscores the fundamental difference between the two models: while LR provides a simple, interpretable structure, it fails to capture the more intricate relationships that are often present in macroeconomic data, such as those detected by the RF model. As such, the Random Forest's ability to identify non-linearities provides a more nuanced understanding of the inflation process.
The 2-dimensional Partial Dependence Plots (PDPs) provide insights into the interaction effects between key variables of interest, in this case, expected inflation, lagged inflation, and output gap on inflation forecasts. The Random Forest (top row) in Chart (ref) captures complex, non-linear interactions between these variables. For instance, in the first plot (expected inflation vs. lagged inflation), inflation forecasts increase sharply when both expected and lagged inflation cross certain thresholds, indicating a compounding effect. This is possible when inflation expectations can get de-anchored in the face of already high inflation leading to an even higher surge in prices. Similarly, the second and third plots reveal non-linear interactions between output gap and both expected and lagged inflation, demonstrating that inflation is likely to be higher when the output gap is positive, particularly when combined with elevated inflation expectations.
In contrast, the Linear Regression (bottom row) shows strictly linear, additive interactions, as expected. The PDPs exhibit uniform, parallel contours, indicating that the LR model does not capture any interaction effects between the variables. Inflation predictions in the LR model are driven by independent, additive contributions of each variable, missing the non-linear dynamics that the Random Forest captures. This inability to model interactions may lead to less accurate predictions in complex macroeconomic environments, where the relationship between variables like inflation expectations and the output gap is inherently non-linear.
Shapley Values and Shapley Regression - Chart (ref) shows the Shapley Summary Plot highlighting the contribution of the top-10 model features to the model's predictions. Each point on the plot represents a single observation's SHAP value for a given feature. The x-axis indicates the SHAP value, reflecting the feature's impact on the model's output, while the colour gradient represents feature values (blue for low, red for high).
Expected inflation and lagged inflation are the most influential features. Higher values of these features (shown in red) strongly increase inflation forecasts (positive SHAP values), while lower values (shown in blue) reduce inflation. The plot also highlights an asymmetric effect for both variables: high values of expected and lagged inflation tend to have a much larger positive impact on predictions compared to the negative impact from lower values. Other features, such as rainfall deviations and exchange rate changes, exhibit smaller and more symmetric influences on predictions, contributing to moderate positive or negative adjustments depending on their values.
This plot underscores the strong influence of inflation-related variables and the asymmetric nature of their impact on the model’s predictions, particularly at higher values of expected and lagged inflation.
Chart (ref) illustrates how the Random Forest (RF) and Linear Regression (LR) model capture the relationships between inflation and its key determinants.
Notably, being an ML model, the RF model is able to discern a non-linear pattern in the data. For expected inflation, both models exhibit a positive relationship, but the RF model is able to learn a non-linear relationship, wherein inflation forecasts become more sensitive to expected inflation with higher levels of future inflation. In other words, the higher values of inflation expectations have an asymmetrically higher impact on actual inflation vis-a-vis lower levels of expected inflation. This contrasts with the constant, linear relationship estimated by the LR model. Similarly, in case of past inflation, the RF model again captures inherent non-linearities in the data, such that forecasts tend to plateau as past inflation stabilizes at lower values of inflation, says around the target rate of inflation.
The most distinctive difference is seen in the relationship between inflation and output gap. The RF model identifies a complex, asymmetric relationship where positive and negative values of output gap influences inflation in different manner. This underscores that the level of inflation tends to be sensitive with respect to the level of slack in the economy. The LR model oversimplifies this, assuming a constant, linear relationship over time.
Finally, Table (ref) presents the Shapley Regression coefficients ($\beta_s^k$) and Shapley Share coefficient ($\Gamma$) for our ML-based Phillips curve model. The results provide insights into the alignment of key variables with the inflation as the target variable. The coefficients \( \beta_S \) measure the alignment of a variable with the target, where values close to 1 indicate perfect alignment. Values greater than 1 show underestimation of a variable’s effect, while values less than 1 show overestimation. The model’s hyperplane tilts towards a variable’s Shapley component when \(\beta_s^k\) is greater than 1 and away when it’s less than 1. The significance of a variable decreases as \( \beta_S \) approaches zero.
The shapely share coefficient has three parts: the sign shows the direction of alignment, the magnitude measures how much of the model predictions is explained by the variable, and the significance level indicates confidence in rejecting the null hypothesis of zero/negative coefficients. This is similar to the interpretation of coefficients in a linear regression analysis. Expected inflation shows near-perfect alignment ($\beta_s^k$ = 1.033) and contributes significantly ($\Gamma$ = 0.159 or 15.9%), meaning the model accurately captures its effect on inflation. Lagged inflation is also well-aligned ($\beta_s^k$ = 0.951 or 9.51%), though slightly overestimated, with a meaningful contribution ($\Gamma$ = 0.075). In contrast, the model underestimates the impact of rainfall deviation ($\beta_s^k$ = 1.621) and the output gap ($\beta_s^k$ = 1.958), despite their statistical significance ($p < 0.01$). The positive contributions of the output gap ($\Gamma$ = 0.004) and rainfall ($\Gamma$ = 0.005) suggest their relevance for modelling inflation dynamics in India. However, the underestimation implies that the model may not fully capture their true influence in the given sample.
To ensure the robustness of our findings, we expand our analysis in several ways. First, to evaluate the predictive accuracy of ML models, we test them against a broader set of baseline models. These baselines include univariate autoregressive (AR) models and multivariate vector autoregression (VAR) models. The results, presented in Table (ref), indicate that RF consistently outperforms other univariate models for two and three quarter-ahead forecasts, achieving the lowest RMSE values of 1.286 and 1.544, respectively. In case of four quarter-ahead forecasts, RF remains highly competitive, with an RMSE of 1.813, closely aligning with the performance of AR(4) (1.803). VAR-based forecasts perform poorly across horizons. It should be noted that AR and VAR forecasts were generated using a recursive forecasting approach, whereas the RF model employed a direct forecasting methodology. This methodological distinction further underscores the superior performance of RF, demonstrating its effectiveness in capturing complex inflation dynamics across varying forecast horizons.
Chart (ref) compares the forecasting performance of Random Forest, XGBoost, and Lasso Regression for predicting one-quarter ahead inflation.
Second, to further evaluate the relative effectiveness of ML models, we employ a model-agnostic Multiple Comparisons with the Best (MCB) procedure (koning2005m3). The MCB test ranks models based on their ex-ante accuracy across datasets and determines statistical significance using critical distances (CD). This non-parametric approach identifies the model with the lowest mean rank as the "best" performer and calculates CDs as $\Theta_{\alpha} \sqrt{\frac{\mathcal{M}\left(\mathcal{M} + 1 \right)}{6 \mathcal{D}}}$, where $\Theta_{\alpha}$ is the critical value from the Tukey distribution at level $\alpha$. Chart (ref) summarizes the results of the MCB test for the RMSE metric. The results demonstrate that the Random Forest (RF) model is able to achieve the lowest mean rank of 1.33, establishing itself as the top-performing forecasting method. This is followed by XGBoost with a mean rank of 2.25. Other models, including Lasso Regression (3.42), Ridge Regression (4.58), and Random Walk (4.83), trail further behind. Interestingly, deep learning models such as BlockRNN, NHits, and NBEATS rank even lower with mean ranks of 6.08, 7.67, and 8.25, respectively. The shaded region in the figure represents the critical distance for the RF model serving as the benchmark for statistical significance. Models with intervals that overlap this critical distance, such as the XGBoost model, exhibit performance close to RF. On the other hand, other models including neural networks fall significantly short. This further underlines the superiority of RF in forecasting CPI inflation over both short-term and long-term horizons.
Third, in general, machine learning models demonstrate superior adaptability and performance in rapidly evolving environments. To examine this, we train and measure the forecasting performance of our ML models across different sample periods. This is pertinent in the current context given the macroeconomic volatility induced by the Covid-19 pandemic. When we examine forecasts in the pre- and post-Covid period, a clear pattern emerges (see Table (ref)). During the relatively stable pre-Covid period, characterized by moderate inflation and steady economic growth, all models perform well. In this case, linear approaches like the LASSO model prove adequate. However, the sample period after the outbreak of Covid-19 pandemic was marked by large-scale disruptions, supply chain bottlenecks, evolving consumer behaviour, and geopolitical shocks such as the Russia-Ukraine conflict. In this environment, ML models like the Random Forest and XGBoost outperform LASSO Regression by effectively capturing nonlinear patterns and rapid shifts in inflation dynamics. These findings underscore that flexible machine learning models can provide better forecasts to navigate periods of high volatility and complex structural changes.
Fourth, in order to analyze the forecasting performance of our best performing model across time in a more formal manner, we employ the Giacomini and Rossi (GR) test to assess the accuracy of the RF model relative to the RW and XGBoost model. Proposed by giacomini2010forecast, the GR test provides a formal statistical approach to assess forecast accuracy over time by employing a rolling evaluation windows. Thus, the test allows a researcher to identify if one model consistently outperforms other models across different time periods or if performance fluctuates between models. For brevity, the analysis is restricted to a four quarter-ahead forecast horizon. The null hypothesis of the test assumes equal performance between models across all time points against the alternative which identifies specific instances of divergence. Unlike traditional forecast evaluation tests, such as those proposed by diebold2002comparing and rossi2016forecast, the GR test is uniquely suited for analyzing time-varying model performance under unstable conditions.
The results of the GR test are shown in Chart (ref). The chart shows the critical values (CVs) represented by thick black lines to delineate thresholds for statistical significance. Parameter $\mu$ defines the rolling window size relative to the evaluation sample. As the results showcase, the RF model exhibits consistent improvements over baseline models, with notable divergence observed in the post-Covid sample marked by heightened economic uncertainty and structural disruptions. For the backward-looking PC specification, RF model significantly surpasses both XGBoost and RW in 2021. A similar pattern emerges in case of both the forward-looking and hybrid PC models. The increased variability in the relative performance of the RF model during the post-pandemic period highlights the impact of structural changes on the underlying data relationships. These findings further confirms the robustness and superior predictive accuracy of ML models, particularly in the volatile post-pandemic environment.
Finally, we explore three methodological variations within the Phillips curve framework: (a) employing the HP filter to estimate trend inflation and the output gap, (b) constructing a composite measure of supply shocks using the principal components approach (PCA) to account for aggregate supply-side dynamics, and (c) integrating alternative measures of inflation expectations derived from internet search volume data, as proposed by bicchal2019rationality. These variations allow us to rigorously assess how different proxies and assumptions related to inflation drivers influence forecasting outcomes. Across all specifications, ML models consistently demonstrate strong predictive performance, with Random Forest and XGBoost maintaining competitiveness across scenarios. LASSO Regression also performs well in specific cases, showcasing its potential in certain contexts. Detailed results of these robustness checks are provided in Appendix A (Tables A5 through A7), further substantiating the efficacy of ML models under varying assumptions.
A critical aspect of forecasting, particularly in complex and non-linear settings, is quantifying the uncertainty associated with predictions. In this section, we leverage the conformal prediction interval (CPI) framework to assess the uncertainty of forecasts generated across the three Phillips Curve specifications—backward, forward, and hybrid. The conformal prediction approach, first introduced by vovk2005algorithmic, provides a non-parametric methodology for constructing robust prediction intervals. Unlike traditional confidence intervals, which primarily quantify the uncertainty of model parameters in a parametric setup, conformal prediction intervals are designed to capture the uncertainty of the forecast or predictive outcome, making them highly relevant for forecasting applications. For our study, we construct CPIs using the sequential nature of time series data that leverages its temporal ordering to build prediction intervals (sengupta2024forecasting).
Consider a training data set $\left\{Y_t, \bar{X}_t \right\}_{t=1}^N$ where $Y_t$ represents the target variable to be forecasted, and $\bar{X}_t$ denotes the set of predictor variables or features. To quantify the uncertainty associated with predictions, we fit the Random Forest (RF) model alongside an uncertainty model $\widehat{\Xi}$ on $\bar{X}_t$ which maps $\bar{X}_t$ to a scalar uncertainty measure. The conformal score $\mathcal{S}_{t}$ can be calculated as:
Given the sequential structure inherent in time-series data, conformal scores are adjusted using weights to reflect temporal dependencies. Specifically, a fixed $\kappa$-sized sliding window is applied, where the weights $\omega_{t^{'}}$ are defined as $\omega_{t^{'}} = \mathbb{1}\{ t^{'} \geq t - \ \kappa), \forall\ t^{'} < t$. The time-adjusted quantile values $\widehat{\mathbb{Q}}_t$ are then computed as:
These quantiles are critical in constructing prediction intervals that adjust dynamically for data heterogeneity and reflect the uncertainty around predictions. Using the weight-adjusted quantiles $\widehat{\mathbb{Q}}_t$, the conformalised prediction interval for each time step $t$ is given by:
We analyze CPIs derived from the Random Forest (RF) model for 4-quarter-ahead forecasts of headline inflation over 21 quarters from 2017:Q2 to 2022:Q2. The RF model is selected as the best-performing approach based on the MCB test discussed earlier. It demonstrated superior predictive accuracy compared to both linear models (linear regression, lasso regression, and ridge regression) and other non-linear models (XGBoost, NBeats, NHits, and BlockRNN). Chart (ref) illustrates the CPIs produced by the RF model along with point forecasts generated using the RF model, the XGBoost model (second-best) , and the Random Walk (RW) model. Including XGBoost forecasts provides a benchmark for evaluating the RF model against a competitive alternative, while the RW model serves as a baseline to assess performance against a simplistic forecasting strategy.
Though it often serves as a simplistic yet effective baseline, the RW model fails to capture nuanced variations in headline inflation, particularly during periods of high volatility. This limitation is evident in its inability to align closely with the ground truth data across all three model specifications. In contrast, the RF model not only demonstrates competitive point-forecasting capabilities but also offers a rigorous mechanism for quantifying predictive uncertainty through CPIs. A notable feature of CPIs is the variability in width across the three specifications reflecting differences in forecast uncertainty. For instance, the average interval width is approximately 3.48 for the backward-looking PC, 3.08 for the forward-looking PC (FPC) regime, and 2.84 for the hybrid PC. This suggests that forecasts from the RF model exhibit lower uncertainty in a hybrid specification. This is likely because the hybrid model combines both backward- and forward-looking elements that may be better at capturing the structural dynamics of inflation.
Therefore, conformal prediction approach allows us to quantify forecast uncertainty in our models and assess their robustness over a multi-step ahead forecast horizon. Being a non-parametric, model-agnostic method for quantifying forecast uncertainty, CPIs can be a valuable addition to the forecasting toolkit, especially in the context of non-linear and complex economic relationships such as that encapsulated in the Phillips Curve framework for India.
Machine learning (ML) and artificial intelligence (AI) have been around for decades, but they have only recently achieved ubiquity due to rapid advances in computing power and data availability. ML/AI algorithms excel at prediction tasks, but they have traditionally struggled to explain their predictions. This has been a major reason for their slow adoption in economic policy.
Therefore, in this paper, we propose machine learning (ML) models as a credible alternative to classical econometric approaches commonly used in empirical macroeconomics, specifically within the context of inflation modelling and forecasting. We employ a New Keynesian Phillips Curve (NKPC) framework to examine inflation dynamics and its determinants in the context of India. Our findings suggest that ML techniques not only outperform conventional linear models but also offer a nuanced approach for analysing complex, nonlinear relationships in the data through explainable machine learning methods.
While ML has traditionally been criticised for its black-box nature, the application of techniques such as feature importance, partial dependence plots, and Shapley values allow us to peep into this black box, providing interpretable insights into the structural determinants of inflation in India. These tools enable a deeper understanding of the inflationary process, revealing the importance of forward-looking elements i.e., inflation expectations and the presence of threshold effects in variables such as the output gap. This allows us to model the non-linearities and structural breaks that conventional models often fail to capture.
However, certain limitations in our analysis should be acknowledged. First, ML models are inherently data-intensive, and our study could benefit from a larger sample size to enhance robustness and potentially uncover further non-linearities in the data. Second, while our results show that expected inflation is a crucial determinant of realised inflation, the choice of the inflation expectations measure remains open to debate. Using alternative measures, such as household or professional forecasts, may yield different outcomes. Similarly, the output gap, which proxies marginal cost in the Phillips curve, is a latent and often imprecisely estimated variable. Variations in its measurement could influence the model's conclusions. Lastly, for large datasets both training ML and Deep learning models as well as calculating shapley values discussed in the paper may be prohibitively compute intensive which the researcher/practitioner must keep in mind while trying to implement this framework in other prediction tasks.
Notwithstanding its limitations, this study demonstrates the potential of machine learning techniques for modeling complex economic relationships, particularly in settings characterized by structural breaks and volatility, such as inflation forecasting in developing economies. The application of Shapley values facilitates the disaggregation of next-period inflation into supply-side (e.g., weather or oil shocks) and demand-side (e.g., output gap) components which can have important implications for monetary policy. This research establishes a foundation for subsequent investigations into macroeconomic relationships through machine learning, addressing the constraints of traditional models in economic forecasting and policy formulation. Future studies may assess the generalizability of these findings across emerging market economies and extend similar methodologies to other macroeconomic variables.
\onehalfspacing \setlength\bibsep{3pt}
The views expressed in the paper are those of authors and do not necessarily reflect the views of the institutions to which they belong.
The authors declare that no funds, grants, or other support were received during the preparation of this manuscript.
The authors have no relevant financial or non-financial interests to disclose.
The authors have no conflict of interest to declare.