Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
98,959 characters · 15 sections · 50 citation commands
Panel Data Nowcasting: The Case of Price-Earnings Ratios
{\it Keywords:} Corporate earnings, nowcasting, data-rich environment, high-dimensional panels, mixed frequency data, textual news data, sparse-group LASSO. \\ \thispagestyle{empty}
\setcounter{page}{0}
Nowcasting is intrinsically a mixed frequency data problem as the object of interest is a low-frequency data series --- observed say quarterly --- whereas real-time information --- daily, weekly or monthly --- during the quarter can be used to assess and potentially continuously update the state of the low-frequency series, or put differently, {\it nowcast} the series of interest. Traditional methods being used for nowcasting rely on dynamic factor models which treat the underlying low-frequency series of interest as a latent process with high-frequency data noisy observations. These models are naturally cast in a state-space form, and inference can be performed using standard techniques (in particular the Kalman filter, see banbura2013now for a recent survey).
Things get more complicated when we are operating in a data-rich environment {\it and} we have many target variables. Put differently, we are no longer interested in nowcasting a single key series such as the GDP growth where we could devote a lot of resources to that particular series. A good example is corporate earnings nowcasting for a large cross-section of corporate firms. The fundamental value of equity shares is determined by the discounted value of future payoffs. Every quarter investors get a glimpse of firms' potential payoffs with the release of corporate earnings reports. In a data-rich environment, stock analysts have many indicators regarding future earnings that are available much more frequently. ball2018automated took a first stab at automating the process using MIDAS regressions. Since their original work, much progress has been made on machine learning (ML) regularized mixed frequency regression models.
In the context of earnings, we are potentially dealing with a large set of individual firms for which there are many predictors. From a practical point of view, this is clearly beyond the realm of nowcasting using state space models. In the current paper, we significantly expand the tools of nowcasting in a data-rich environment by exploiting panel data structures. Panel data regression models are well suited for the firm-level data analysis as both the time series and cross-sectional dimensions can be exploited. In such models, time-invariant firm-specific effects are typically used to capture cross-sectional heterogeneity in the data. This is combined with regularized regression machine learning methods which are becoming increasingly popular in economics and finance as a flexible way to model predictive relationships via variable selection. We focus on the panel data regressions in a high-dimensional data setting where the number of covariates could be large and potentially exceed the available sample size. This may happen when the number of firm-specific characteristics, such as textual analysis news data or firm-level stock returns, is large, and/or the number of aggregates, such as market returns, macro data, etc., is large.
Our paper relates to several existing papers in the literature. khalaf2020dynamic consider low-dimensional dynamic mixed frequency panel data models but do not deal with high-dimensional data situations in the context of nowcasting or forecasting. Similarly, fosten2019panel consider nowcasting with a mixed-frequency VAR panel data model, but not in the context of a high-dimensional data-rich environment that we are interested in here. babii2022machine introduce the sparse-group LASSO (sg-LASSO) regularization machine learning methods for heavy-tailed dependent panel data regressions potentially sampled at different time series frequencies. They derive oracle inequalities for the pooled and fixed effects models, the debiased inference for pooled regression, and consider an application to the Granger causality testing. In this paper, we explore how to use their framework for nowcasting large panels of low-frequency time series.
We focus on nowcasting current quarter firm-specific price-earnings ratios (henceforth P/E ratios). This means we focus on evaluating model-based within-quarter predictions for very short horizons. It is widely acknowledged that P/E ratios are a good indicator of the future performance of a company and, therefore, are used by analysts and investment professionals to base their decisions on which stocks to pick for their investment portfolios. Typically investors rely on consensus forecasts of earnings made by a pool of analysts. We, therefore, choose such consensus forecasts as the benchmark for our proposed machine learning methods. ball2018automated and carabias2018real documented that analysts tend to focus on their firm/industry when making earnings predictions while not fully taking into account the impact of macroeconomic events. babii2022machine tested formally in a high-dimensional data setting the hypothesis that systematic and predictable errors occur in analyst forecasts and confirmed empirically that they {\it leave money on the table}. The analysis in the current paper is therefore an logical extension of this prior work. In addition, we also compare our proposed new methods with the MIDAS regression forecast combination approach used by ball2018automated as well as a simple random walk model.
Our high-frequency regressors include traditional macro and financial series as well as non-standard series generated by textual analysis of financial news. We consider structured pooled and fixed effects sg-LASSO panel data regressions with mixed frequency data (sg-LASSO MIDAS). By “structured” we mean that the ML procedure is set up such that it recognizes the time series and panel structure of the data. This is a departure from standard ML which is rooted in a tradition of i.i.d.\ covariates and therefore time series and panel data structures are not recognized. For the purpose of comparison, we include elastic net estimators in our analysis, as a representative example of standard ML.
In our empirical analysis we study nowcasting the firm-level P/E ratio for a large set of firms. Moreover, we decompose the (log of) the P/E ratio into the return for firm $i$ and analyst prediction errors. Therefore, nowcasting the log P/E ratio could also be achieved via nowcasting its two components. The decomposition corresponds to the distinction between analyst assessments of firm $i$'s earnings and market/investor assessments of the firm.
Our empirical results can be summarized as follows. Predictions based on analyst consensus exhibit significantly higher mean squared forecast errors (MSEs) compared to model-based predictions. These model-based predictions involve either direct log P/E ratio nowcasts or their individual components. The MSE for the random walk model and analysts' concensus are quite similar, and therefore random walk predictions are outperformed by the model-based ones as well. A substantial proportion of firms (approximately 60%) exhibit low MSE values, indicating a high level of prediction accuracy. However, there are a few firms for which the MSEs are relatively larger, suggesting lower prediction performance for these specific cases. Comparing direct log P/E ratio nowcasts versus those based on its components, we observe a substantial improvement in prediction accuracy when using the individual components. This improvement is consistently evident across individual, pooled, and fixed effects regression models. Moreover, the sparsity patterns differ significantly across the direct versus component prediction models.
Our framework allows us to go beyond providing quarterly nowcasts and generate daily updates of earnings series. Leveraging the daily influx of information throughout the quarter, we continuously re-estimate our models and produce nowcast updates as soon as new data becomes available. We report the distribution of Mean Squared Errors (MSEs) across firms for five distinct nowcast horizons: 20-day, 15-day, 10-day, and 5-day ahead, as well as the end of the quarter and show that as the horizons become shorter, both the median and upper quartile of MSEs decrease. The sg-LASSO estimator we employ in our study is well-suited for incorporating grouped fixed effects. This approach involves grouping firm-specific intercepts based on either statistical procedures or economic reasoning, as outlined in bonhomme2015grouped. In our analysis, we utilize the Fama French industry classification to form 10 distinct groups for grouping fixed effects. Our findings suggest that grouped fixed effects strike a better balance between capturing heterogeneity and pooled parameters, resulting in more accurate nowcast predictions. These results support the notion that incorporating group fixed effects enhances the overall performance of our forecasting model.
Next we address the challenge of missing earnings data, which can complicate the analysis. We examine the performance of parameter imputation methods in computing nowcasts, see, e.g, brown2023nowcasting, even when earnings and/or earnings forecasts are missing for certain observations in the sample. The results obtained through parameter imputation outperform the analyst consensus nowcasts in terms of prediction accuracy.
The paper is organized as follows. Section (ref) introduces the models and estimators. A simulation study reporting the finite sample nowcasting performance of our proposed methods appears in Section (ref). The results of our empirical application analyzing price-earnings ratios for a panel of individual firms are reported in Section (ref). Section (ref) concludes. All technical details and detailed data descriptions appear in the Appendix and the Online Appendix.
In this section, we describe the methodological approach of the paper. Motivated by our application, we will refer to the cross-sectional observations as firms, the low-frequency observations as quarterly while the high-frequency observations are daily or monthly. However, the notation presented in this section is generic and can correspond to other entities and frequencies. The objective is to nowcast $\{y_{i,t}:i\in[N],t\in[T]\}$ (where for a positive integer $p$, we put $[p]=\{1,2,\dots,p\}$), in our case a panel of P/E ratios (or its decomposition into returns and analyst forecast errors) for $N$ firms observed at $T$ time periods. The covariates consist of $K$ time-varying predictors measured potentially at higher frequencies
where $n_k^H$ is the number of high-frequency observations for the $k^{\rm th}$ covariate in a low-frequency time period $t$, and $n^L_k$ is the number of low-frequency time periods used as lags. For instance, $n^L_k=1$ corresponds in our application to a quarter of high-frequency lags used as covariates and $n_k^H=3$ corresponds to monthly data with 3 month of data available per quarter. Note that we can think of mixtures of say annual, quarterly, monthly and weekly data, and therefore $n_k^H$ represents different high frequency sampling frequencies and associated lags $n_k^L n_k^H.$
In our empirical analysis we examine three types of regression model specifications: (a) regularized single equation regressions for each individual firm, (b) regularized panel regressions with pooling, and (c) regularized panel regressions with fixed effects. Hence, in (a) we do not explore the panel structure of the data, whereas in (b) and (c) we do. To discuss the model specifications, we focus here on (b) and (c), keeping in mind that the single regression case is a straightforward simplification of the panel regression models.
Consider the mixed frequency panel data regression for $y_{i,t|\tau},$ that is observation $i$ for low-frequency nowcasting $y$ at time $t$ using information up to $\tau:$
where $\alpha_i$ is the entity-specific intercept (depending on $\tau$ but we suppress this detail to simplify notation), and
where $k_{max}$ is the maximum lag length which may depend on the covariate $k,$ and for each high frequency covariate $x_{i,\tau,k}$ we have the most up to date information available at time $\tau.$ This may imply that for some high frequency regressors this is stale information as they have not been updated yet, but presumably at least some of the high frequency data are fresh real-time information at the time $\tau$ the nowcast is being made. For instance, in our quarterly/monthly application we can have $\tau$ = $(t - 1) + 1/3$ in which case we nowcast quarter $t$ with information available at the end of the first month of that quarter. In this example, some high frequency series for the first month may be available while some may not due to say publication lags. Likewise, with $\tau$ = $(t - 1) + 2/3$ we can revise the previous nowcast with one extra month of information, which taking into account publication lags may include observations from the first month as the most recent releases. It should parenthetically be noted that for $\tau$ $\leq$ $t - 1,$ we are dealing with a forecasting situation and therefore our analysis applies to both nowcasting and - ceteris paribus - forecasting.
To reduce the dimensionality of the high-frequency lag polynomial, we follow the MIDAS ML literature, see babii2020inference,babii2020machine, and estimate a weight function $\omega$ parameterized by a relatively small number of coefficients $L$
where the MIDAS weight function is $\omega(s;\beta_k)$ = $\sum_{l=0}^{L-1}\beta_{l,k}w_l(s),$ $(w_l)_{l\geq 0}$ is a collection of $L$ approximating functions, called the dictionary, and $\beta_k\in\ensuremath{\mathbf{R}}^L$ is the unknown parameter. An example of a dictionary used in the MIDAS ML literature is the set of orthogonal Legendre polynomials. To streamline notation it will be convenient to assume, without loss of generality, a common lag length, i.e.\ $\bar{k}_{max}$ = $k_{max}$ $\forall$ $k$ $\in$ $[K].$ The linear in parameters dictionaries map the MIDAS regression to a standard linear regression framework. In particular, define $\mathbf{x}_i = (X_{i,1}W,\dots,X_{i,K}W)$, where for each $k\in[K]$, $X_{i,k} = (x_{i,\tau-j/n_k^H,k},j = 0, \ldots, \bar{k}_{max} - 1)_{\tau\in[T]}$ is a $T\times \bar{k}_{max}$ matrix of covariates and $\bar{k}_{max} W$ = $(w_l(j/n_k^H; \beta_k)_{0\leq l\leq L-1, 0 \leq j\leq \bar{k}_{max}}$ is a $\bar{k}_{max} \times L$ matrix corresponding to the dictionary. In addition, let $\mathbf{y}_i$ = $(y_{i,t|\tau}, t, \tau \in [T])^\top$ and $\mathbf{u}_i$ = $(u_{i,t|\tau},t, \tau \in [T])^\top.$ The regression equation after stacking time series observations for each firm $i\in[N]$ is as follows
where $\iota\in\ensuremath{\mathbf{R}}^T$ is the all-ones vector and $\beta\in\ensuremath{\mathbf{R}}^{LK}$ is a vector of slope coefficients. Lastly, put $\mathbf{y} = (\mathbf{y}_1^\top,\dots, \mathbf{y}_N^\top)^\top$, $\mathbf{X}=(\mathbf{x}_1^\top, \dots, \mathbf{x}_N^\top)^\top$, and $\mathbf{u} = (\mathbf{u}_1^\top,\dots,\mathbf{u}_N^\top)^\top$. Then the regression equation after stacking all cross-sectional observations is
where $B=I_N\otimes\iota$, $I_N$ is $N\times N$ identity matrix, and $\otimes$ is the Kronecker product. Given that the number of potential predictors $K$ can be large, additional regularization can improve the predictive performance in small samples. To that end, we take advantage of the sg-LASSO regularization, suggested by babii2020machine.
The fixed effects sg-LASSO estimator $\hat\rho=(\hat\alpha^\top,\hat\beta^\top)^\top$ solves
where $\Omega$ is the sg-LASSO regularizing functional. It is worth stressing that the design matrix $\mathbf{X}$ does not include the intercept and that we do not penalize the fixed effects which are typically not sparse. In addition, $\|.\|_{NT}^2 = |.|^2/(NT)$ is the empirical norm and
is a regularizing functional. It is a linear combination of the $\ell_1$ LASSO and $\ell_{2,1}$ group LASSO norms. Note that for a group structure $\mathcal{G}$ described as a partition of $[p]=\{1,2,\dots,p\}$, the group LASSO norm is computed as $\|b\|_{2,1}=\sum_{G\in\mathcal{G}}|b_G|_2$, while $|.|_q$ denotes the usual $\ell_q$ norm. The group LASSO penalty encourages sparsity between groups whereas the $\ell_1$ LASSO norm promotes sparsity within groups and allows us to learn the shape of the MIDAS weights from the data. The parameter $\gamma\in[0,1]$ determines the relative weights of the $\ell_1$ (sparsity) and the $\ell_{2,1}$ (group sparsity) norms, while the amount of regularization is controlled by the regularization parameter $\lambda\geq 0$.
In Section (ref), we called our approach structured ML because the group structure allows us to embed the time series structure of the data. More specifically, these structures are represented by groups covering lagged dependent variables and groups of lags for a single (high-frequency) covariate. Throughout the paper, we assume that groups have fixed size, and the group structure is known by the econometrician. Both are reasonable assumptions to make in the context of our empirical application.
For pooled regressions, we assume that all entities share the same intercept parameter $\alpha_1=\dots=\alpha_N=\alpha$. The pooled sg-LASSO estimator $\hat\rho=(\hat\alpha,\hat\beta^\top)^\top$ solves
Pooled regressions are attractive since the effective sample size $NT$ can be huge, yet the heterogeneity of individual time series may be lost. If the underlying series have a substantial heterogeneity over $i\in[N]$, then taking this into account might reduce the projection error and improve the predictive accuracy.
babii2022machine provide the theoretical analysis of predictive performance of regularized panel data regressions with the sg-LASSO regularization, including as special cases (a) standard LASSO, (b) group LASSO regularizations as well as (c) generic high-dimensional panels not involving mixed frequency data. Finally, babii2022machine also develop the debiased inferential methods and Granger causality tests for pooled panel data regressions.
It is not clear that the aforementioned theory is of practical use in the context of nowcasting using modestly sized samples of data. For this reason, we investigate in this section the finite sample nowcasting performance of the machine learning methods covered so far. We consider the standard (unstructured) elastic net with UMIDAS (called Elnet-U), where UMIDAS refers to unconstrained MIDAS proposed by FMS15 in a classic non-ML context, and sg-LASSO with MIDAS. Both methods require selecting two tuning parameters $\lambda$ and $\gamma$. In the case of sg-LASSO, $\gamma$ is the relative weight of LASSO and group LASSO penalties while in the case of the elastic net $\gamma$ interpolates between LASSO and ridge. In both cases we report results on a grid $\gamma \in \{0, 0.2, \dots, 1\}$.
In addition to evaluating the performance over the grid of $\gamma$ tuning parameter values, we need to select the $\lambda$ tuning parameter. To do so, we consider several approaches. First, we adapt the $K$-fold cross-validation to the panel data setting. To that end, we resample the data by blocks respecting the time-series dimension and creating folds based on cross-sectional units instead of the pooled sample. We use the 5-fold cross-validation both in the simulation experiments and the empirical application. We also consider the following three information criteria: BIC, AIC, and corrected AIC (AICc) of hurvich1989regression. Assuming that $y_{i,t}|x_{i,t}$ are i.i.d.\ draws from $N(\alpha_i + {x}_{i,t}^\top\beta, \sigma^2)$, the log-likelihood of the sample is
Then, the BIC criterion is
where $df$ denotes the degrees of freedom, $\hat\sigma^2$ is a consistent estimator of $\sigma^2$, $\hat\mu=\hat\alpha\iota$ for the pooled regression, and $\hat\mu=B\hat\alpha$ for fixed effects regression. The degrees of freedom are estimated as $\widehat{df} = |\hat\beta|_0+1$ for the pooled regression and $\widehat{df} = |\hat\beta|_0+N$ for the fixed effects regression, where $|.|_0$ is the $\ell_0$-norm defined as a number of non-zero coefficients; see zou2007degrees for more details. The AIC is computed as
and the corrected Akaike information criteria is
The AICc is typically a better choice when $p$ is large relative to the sample size. We report the results for each of the tuning parameter selection criteria for $\lambda,$ along the grid choice for $\gamma.$
To assess the predictive performance of pooled panel data models, we simulate the data from the following DGP with a quarterly/monthly frequency mix in mind and $\bar{k}_{max}$ = $k_{max}$ with $n_k^H$ = $n^H$ $\forall$ $k:$
where $i\in[N]$, $t\in[T]$, $\alpha$ is the common intercept, $\bar{k}_{max}^{-1}\sum_{j=0}^{\bar{k}_{max}-1} \omega(j/n_k;\beta_k)$ the weight function for $k$-th high-frequency covariate and the error term is either $u_{i,t|\tau} \sim_{i.i.d.}N(0,1)$ or $u_{i,t|\tau} \sim_{i.i.d.}\text{student-}t(5)$.
We are interested in a quarterly/monthly data mix, and use four quarters of data for the high-frequency regressors which covers 12 high-frequency lags for each regressor. In terms of information sets we start with $\tau$ = $t - 1,$ which corresponds to a prediction setting and then have $\tau$ = $t - 1 + 1/3,$ i.e.\ nowcasting with one month's worth of information. We set the number of relevant high-frequency regressors $K$ = 6. The high-frequency regressors are generated as $K$ i.i.d.\ realizations of the univariate autoregressive (AR) process $x_h = \rho x_{h-1}+ \varepsilon_h,$ where $\rho=0.6$ and either $\varepsilon_h\sim_{i.i.d.}N(0,1)$ or $\varepsilon_h\sim_{i.i.d.}\text{student-}t(5)$, where $h$ denotes the high-frequency sampling. We rely on a commonly used weighting scheme in the MIDAS literature, namely $\omega(s;\beta_k)$ for $k=1,2,\dots,6$ are determined by beta densities respectively equal to $\mathrm{Beta}(1,3)$ for $k=1,4$, $\mathrm{Beta}(2,3)$ for $k=2,5$, and $\mathrm{Beta}(2,2)$ for $k=3,6$; see ghysels2007midas or ghysels2019estimating, for further details. The MIDAS regressions are estimated using Legendre polynomials of degree $L=3$.
We consider DGPs featuring pooled panels and fixed effects. For the pooled panel regression DGPs we simulate the intercepts as $\alpha \sim \text{Uniform}(-4,4).$ For the fixed effects models the individual fixed effects are simulated as $\alpha_i \sim_{\text{i.i.d}} \text{Uniform}(-4,4)$ and are kept fixed throughout the experiment.
For $\tau$ = $t - 1,$ the {\it Baseline scenario}, in the estimation procedure we add 24 noisy covariates which are generated in the same way as the relevant covariates, use 4 low-frequency lags and the error terms $u_{i,t|\tau}$ and $\varepsilon_h$ are Gaussian. In the student-$t(5)$ scenario we replace the Gaussian error terms with a student-$t(5)$ distribution while in the {\it large dimensional} scenario we add 94 noisy covariates. For each scenario, we simulate $N=25$ i.i.d.\ time series of length $T=50$; next we increase the cross-sectional dimension to $N=75$ and time series to $T=100$.
Finally, for $\tau$ = $t - 1 + 1/3$ the thought experiment in the simulation design is one where the first high-frequency observations during low frequency $t$ are available. The nowcaster of course does not know which of the covariates are relevant nor does she know the parameters of the prediction rule. We will call this scheme “one-step ahead” nowcasts.
Tables (ref) and (ref) cover the average mean squared forecast errors (MSFE) for one-step ahead nowcasts for the three simulation scenarios. We report results for sg-LASSO with MIDAS weights (left block) and elastic net with UMIDAS (right block) using both pooled panel models (Table (ref)) and fixed effects ones (Table (ref)). We report results for the best choice of the $\gamma$ tuning parameter.\footnote{Results for the grid of $\gamma\in\{0.0, 0.2, \dots, 1.0\}$ are reported in the Online Appendix Tables (ref)-(ref).}
Firstly, structured sg-LASSO-MIDAS consistently outperforms unstructured Elnet-U for all DGPs and in both pooled and fixed effects cases. The most significant discrepancy between the two methods is observed in situations with small N and small T, specifically when N = 25 and T = 50. As either N or T increases, this gap gradually diminishes. When comparing the results of pooled and fixed effects, it becomes evident that the difference between the two approaches — structured sg-LASSO-MIDAS versus Elnet UMIDAS — widens further in the case of fixed effects with student-t(5) data. This indicates that our structured approach yields higher quality estimates for the fixed effects and thus more accurate nowcasts.
In the case of sg-LASSO-MIDAS, the best performance is achieved for $\gamma\notin\{0,1\}$ for both pooled panel data and fixed effects cases, while $\gamma=0$, i.e.\ ridge regression, seems to be dominated by estimators that $\gamma\notin\{0,1\}$ in both pooled and fixed effects cases. For the student-$t(5)$ and large dimensional DGP, we observe a decrease in the performance for all methods. However, the decrease in the performance is larger for the student-$t(5)$ DGP, revealing that heavy-tailed data have — as expected — a stronger impact on the performance of the estimators.
For the pooled panel data case, increasing $N$ from $25$ to $75$ seems to have a larger positive impact on the performance than an increase in the time-series dimension from $T=50$ to $T=100$. The difference appears to be larger for student-$t(5)$ and large dimensional DGPs and/or for the elastic net case. Turning to the fixed effects results, the differences seem to be even sharper, in particular for student-$t(5)$ and large dimensional DGPs.
When comparing the results across the different model selection methods, i.e., cross-validation and the three information criteria, we find that almost always cross-validation leads to smaller prediction errors in both pooled and fixed effects panel data cases. Notably, the gains appear to be larger for the large $N$ and $T$ values. Comparing BIC, AIC, and AICc information criteria, the results appear to be similar for AIC and AICc across DGPs and different sample sizes, while the BIC performance is slightly worse than AIC and AICc.
ball2018automated, carabias2018real and babii2022machine documented that analysts make systematic and predictable errors in their P/E forecasts. We therefore consider nowcasting the P/E ratios using a set of predictors that are sampled at mixed frequencies for a large cross-section of firms.
A natural question one may ask: should we nowcast P/E ratio directly or it's components. We, therefore, decompose the (log of) the P/E ratio for firm $i$ as follows:
where $r_{i,t+1}$ is the log return from $t + 1$ to $t$ for firm $i,$ $E^a_{i,t+1|t}$ the analyst's prediction at time $t$ pertaining to $t + 1$ earnings, and $e^a_{i,t+1|t}$ $\equiv$ $\log (E_{i,t+1}) - \log (E^a_{i,t+1|t})$ is the log earnings forecast error of analysts pertaining to their end of period $t$ prediction for $t + 1.$ Finally, $\log (P_{i,t}/E^a_{i,t+1|t})$ is perfectly known at time $t.$ The above defines an additive decomposition of the log P/E ratio into the return for firm $i$ and the analyst prediction error. Therefore, nowcasting the log P/E ratio could also be achieved via nowcasting its two components. The decomposition corresponds to the distinction between analyst assessments of firm $i$'s earnings and market/investor assessments of the firm.
There is a considerable literature on using machine learning to predict returns, see e.g.\ rapach2010out, kim2014forecasting, gu2020empirical, d2020artificial, among others. Here we are dealing with a slightly modified setting where we are nowcasting quarterly returns with information during quarter $t + 1.$ Nevertheless, prediction and nowcasting are closely related. The second component, $e^a_{i,t+1|t}$ has been explored by babii2022machine, who revisit a topic raised by ball2018automated and carabias2018real, and confirmed in a rich data setting that analysts tend to focus on their firm/industry when making earnings predictions while not fully taking into account the impact of macroeconomic events. Put differently, one can forecast and nowcast analyst prediction errors.
It should also parenthetically be noted that equation ((ref)) can be rewritten as a decomposition of returns, namely:
which can be viewed as an alternative decomposition of returns compared to ferreira2011forecasting. They propose forecasting separately the three components of stock market returns: (a) the dividend price ratio, (b) earnings growth, and (c) price-to-earnings ratio growth. ferreira2011forecasting argue that predicting the separate components yields better return predictions compared to the usual models producing direct forecasts of the latter. They estimate the expected earnings growth using a 20-year moving average of the growth in earnings per share. The expected dividend price ratio is estimated by the current dividend price ratio. This implicitly assumes that the dividend price ratio follows a random walk. While our application is different in many regards, the arguments being considered are similar. It is worth reminding ourselves that if the nowcast $\widehat{pe}_{i,t+1}$ is constructed from individual component nowcasts, then
Hence, depending on the co-movements between returns for firm $i,$ $r_{i,t+1}$ and analyst earning prediction errors $e^a_{i,t+1|t},$ we are better off to directly predict $pe_{i,t+1}$ or its components. If the latter are positively correlated, then we are better off direct forecasting is preferred.
Given the aforementioned decomposition, we are interested in the following LHS variables: $pe_{i,t+1},$ $r_{i,t+1}$ and $e^a_{i,t+1|t}.$ First, we estimate the individual sg-LASSO MIDAS regressions for each firm $i=1,\dots,N$, namely:
where the firm-specific predictions are computed as $\hat y_{i,t+1} = \hat\alpha_i + x_{i,t+1}^\top\hat\beta_i$. As noted in Section (ref), $ \mathbf{x}_i$ contains lags of the low-frequency target variable and high-frequency covariates to which we apply Legendre polynomials of degree $L=3$.
Next, we estimate the following pooled and fixed effects sg-LASSO MIDAS panel data models
and compute predictions as
Once we compute the forecast for the log of P/E ratio ($pe_{i,t+1}$), log returns ($r_{i,t+1}$) and log earnings forecast error ($e^a_{i,t+1|t}$), we compute the final prediction accuracy metrics by either taking directly log P/E nowcast or the sum of its components, i.e., $\hat{S} = \hat r_{i,t+1} - \hat e^a_{i,t+1|t} + \log (P_{i,t}/E^a_{i,t+1|t})$.
We benchmark firm-specific and panel data regression-based nowcasts against two simple alternatives. First, we compute forecasts for the RW model as
Second, we consider predictions of P/E implied by analysts' earnings nowcasts using the information up to time $t+1$, i.e.\
where the predicted/nowcasted log of P/E ratio is based on consensus earnings forecasts pertaining to the end of the $t+1$ quarter using the stock price at the end of quarter $t.$ To measure the forecasting performance, we compute the mean squared forecast errors (MSE) for each method. Let $\mathbf{\bar y}_i = (y_{i,T_{is}+1},\dots, y_{i,T_{os}})^\top$ represent the out-of-sample realized P/E ratio values, where $T_{is}$ and $T_{os}$ denote the last in-sample observation for the first prediction and the last out-of-sample observation respectively, and let $\mathbf{\hat y}_i = (\hat y_{i,t_{is}+1},\dots, \hat y_{i,t_{os}})$ collect the out-of-sample forecasts. Then, the mean squared forecast errors are computed as
We look at 210 US firms and use 24 predictors, including traditional macro and financial series as well as non-traditional series from textual analysis of financial news. We apply (a) single regression individual firm high-dimensional regressions, (b) pooled and (c) individual fixed effects sg-LASSO MIDAS panel data models and report results for several choices of the tuning parameters. We compare these three type of models with several benchmarks, which include a random walk (RW) model and analysts' consensus forecasts. The remainder of the section is structured as follows. We start with a short review of the data followed by a summary of the empirical results.
The full sample consists of observations between the \(1^{st}\) of January, 2000 and the \(30^{th}\) of June, 2017. Due to the lagged dependent variables in the models, our effective sample starts in the third fiscal quarter of 2000. We use the first 25 observations for the initial sample, and use the remaining 42 observations for evaluating the out-of-sample forecasts, which we obtain by using an expanding window forecasting scheme. We collect data from CRSP and I/B/E/S to compute the quarterly P/E ratios and firm-specific financial covariates; RavenPack is used to compute daily firm-level textual-analysis-based data; real-time monthly macroeconomic series are from the FRED-MD dataset, see mccracken2016fred for more details; FRED is used to compute daily financial markets data and, lastly, monthly news attention series extracted from the {\it Wall Street Journal} articles are retrieved from bybee2019structure.\footnote{The dataset is publicly available at \href{http://www.structureofnews.com/}{http://www.structureofnews.com/}.} Online Appendix Section (ref) provides a detailed description of the data sources.\footnote{In particular, firm-level variables, including P/E ratios, are described in Online Appendix Table (ref), and the other predictor variables in Online Appendix Table (ref). The list of all firms we consider in our analysis appears in Online Appendix Table (ref).}
Our target variable is the P/E ratio for each firm. To compute it, we use CRSP stock price data and I/B/E/S earnings data. Earnings data are subject to release delays of 1 to 2 months depending on the firm and quarter. Therefore, to reflect the real-time information flow, we compute the target variable using stock prices that are available in real-time. We also take into account that different firms have different fiscal quarters, which also affects the real-time information flow.
For example, suppose for a particular firm the fiscal quarters are at the end of the third month in a quarter, i.e.\ end of March, June, September, and December. The consensus forecast of the P/E ratio is computed using the same end-of-quarter price data which is divided by the earnings consensus forecast value. The consensus is computed by taking all individual prediction values up to the end of the quarter and aggregating those values by taking either the mean or the median. To compute the target variable, we adjust for publication lags and use prices of the publication date instead of the end of fiscal quarter prices. More precisely, suppose we predict the P/E ratio for the first quarter. As noted earlier, earnings are typically published with 1 to 2 months delay; say for a particular firm the data is published on the 25$th$ of April. In this case, we record the stock price for the firm on 25$th$ of April, and divide it by the earnings announced on that date.
To simplify the exposition, we denote $y$ as one of the three target variables we consider. The main findings from our analysis are presented in Table (ref). Column $\hat pe_{i,t+1}$ reports results for directly nowcasting the log P/E ratio, column $\hat S$ reports the results of nowcasting and summing up the components, column $r_{i,t+1}$ reports results for the log return component and column $\hat e^a_{i,t+1|t}$ reports results for the log earnings forecast error of analysts component. Row {\it RW} reports results for the random walk, while row {\it Consensus} for the median consensus nowcast. Panels {\it Individual}, {\it Pooled} and {\it Fixed effects} report results for different panel data models relative to the consensus MSE (columns $\hat pe_{i,t+1}$ and $\hat S$) and for the components (columns $r_{i,t+1}$ and $\hat e^a_{i,t+1|t}$) we report ratios relative to the RW MSE since there are obviously no concensus series notably for the analyst forecast errors.
In light of the simulation evidence, we report the empirical results using cross-validation in Table (ref) and provide the full set of results in Online Appendix Table (ref). The entries in the top panel of Table (ref) reveal that predictions based on analyst consensus exhibit significantly higher mean squared forecast errors (MSEs) compared to model-based predictions since all the ratios with respect to the concensus are less than one (see first two columns). These model-based predictions involve either direct log P/E ratio nowcasts (first column) or their individual components (second column). Since the MSE for RW and concensus are quite similar, this also implies that RW predictions are outperformed by the model-based ones. The substantial improvement in the accuracy of model-based predictions compared to analyst-based predictions underscores the value of employing machine learning techniques for nowcasting log P/E ratios. Across various machine learning methods, including single-firm and panel data regressions, we consistently observe enhanced performance.
When comparing the first and second columns, which correspond to direct log P/E ratio nowcasts versus those based on its components, we observe a substantial enhancement in prediction accuracy when using the individual components. This improvement is consistently evident across individual, pooled, and fixed effects regression models. To shed light on these findings, we computed the pooled correlation between returns and earnings for the entire sample, i.e.\ Corr($r_{i,t+1}, e^a_{i,t+1|t}$) = -0.206. The correlation indicates a (weak) negative relationship between returns and earnings. Consequently, the prediction errors of each component tend to offset each other, resulting in more accurate aggregated nowcasts (recall equation ((ref))). The last two columns of Table (ref) present the prediction results for these components. We observe that analyst earnings prediction errors appear to be more predictable than those of log returns. We also report diebold1995comparing test statistic p-values comparing each model against the RW and consensus benchmarks, pooling all the nowcasting errors across firms. Using one-sided test critical values we observe that our models outperform both the RW and consensus benchmarks, particularly when we use the component approach. While we cannot compare the $\hat pe_{i,t+1}$ component with the consensus, judging by the RW benchmark it is clear that the second component is the most important in terms of nowcasting gains. When we use individual MIDAS regressions the evidence is less compelling, underscoring the importance of using panel data models.\footnote{We also experimented with the forecast combination of MIDAS regressions used by ball2018automated and found them to be inferior to the individual MIDAS ML regressions as well as the panel data models. We therefore refrain from reporting the details here. }
Figure (ref) illustrates the sparsity patterns of selected covariates for the most effective methods in predicting either log P/E ratios (Panel a) or their components (Panels b and c). It is worth noting that the sparsity patterns differ significantly across the three panels. For instance, firm volatility is often chosen as a relevant covariate across all targets, albeit not consistently throughout the entire out-of-sample period. In the case of log P/E ratios, news series related to earnings are frequently selected, along with firm and market volatility series. Conversely, for log returns, a denser pattern of covariate selection is observed, distinct from the other two cases. Interestingly, none of the news-based firm series are chosen for this target. Regarding log analyst earnings forecast errors, macroeconomic series such as the unemployment rate, short-term rates, and TED rate are frequently selected. Moreover, unlike log P/E ratios and returns, news-based firm series occasionally appear in the selected covariates for this target. The fact that macroeconomic series are drivers for nowcasting the $e^a_{i,t+1|t}$ component is a confirmation of the findings reported in ball2018automated, carabias2018real and babii2022machine.
Figure (ref) depicts the histogram of mean squared errors (MSEs) across firms. Notably, a substantial proportion of firms (approximately 60%) exhibit low MSE values, indicating a high level of prediction accuracy. However, there are a few firms for which the MSEs are relatively larger, suggesting lower prediction performance for these specific cases. The largest MSE is for Crown castle international corporation (CCI) which appears as a strong outlier.
Removing the single outlier firm has a dramatic impact on the nowcasting performance evaluation as shown in the lower panel of Table (ref). We now have very strong evidence that the panel regression models dominate analyst predictions. Again the component nowcasts are the best, but even the individual regression models do significantly better when the component specification is used.
Our framework allows us to go beyond providing quarterly nowcasts and generate daily updates of earnings series. Leveraging the daily influx of information throughout the quarter, we continuously re-estimate our models and produce nowcast updates as soon as new data becomes available. In Figure (ref), we present the distribution of Mean Squared Errors (MSEs) across firms for five distinct nowcast horizons: 20-day, 15-day, 10-day, and 5-day ahead, as well as the end of the quarter. We report the best model based on Table (ref). Notably, as the horizons become shorter, both the median and upper quartile of MSEs decrease. Therefore, updating nowcasts with daily information appears to significantly enhance the prediction performance of log earnings ratios. The largest errors persist for the same firm, CCI.
The sg-LASSO estimator we employ in our study is well-suited for incorporating grouped fixed effects. This approach involves grouping firm-specific intercepts based on either statistical procedures or economic reasoning, as outlined in bonhomme2015grouped. In our analysis, we utilize the Fama French industry classification to form 10 distinct groups for grouping fixed effects. Rather than assuming a common fixed effect for all firms within a group, we apply a group penalty to the fixed effects of firms belonging to the same industry. This allows us to capture industry-specific heterogeneity while avoiding overfitting.
We present the findings in Table (ref), which highlight several key observations. Similar to previous analyses, our results suggest that predicting individual components of the log price-earnings ratio leads to more accurate aggregate nowcasts compared to a direct nowcast approach. Furthermore, we observe that the use of group fixed effects improves the accuracy of our nowcasts when forecasting individual components. This can be seen in column 2 of both Tables (ref) and (ref). Comparatively, when considering the best tuning parameter choice, grouped fixed effects outperform other panel models, including the pooled panel model. Therefore, our findings suggest that grouped fixed effects strike a better balance between capturing heterogeneity and pooled parameters, resulting in more accurate nowcast predictions. These results support the notion that incorporating group fixed effects enhances the overall performance of our forecasting model.
In Figure (ref), we present the distribution of (MSEs) across firms for five industries, based on the best model specification from Table (ref). The industries we focus on are the ones with the highest number of firms in our sample. The results reveal variations in performance among different industries. Specifically, the firms categorized as {\it Consumer Durables} exhibit the lowest accuracy in terms of the median MSE, although the quartiles are comparatively lower compared to the other industries. On the other hand, the nowcasts for firms in the {\it Consumer Nondurables} and {\it Others} categories demonstrate the highest accuracy at the median. However, it is important to note that the largest errors occur within the firms classified as {\it Others}.
Next we address the challenge of missing earnings data, which can complicate the analysis. We examine the performance of parameter imputation methods in computing nowcasts, see, e.g, brown2023nowcasting, even when earnings and/or earnings forecasts are missing for certain observations in the sample. We identify a subset of 117 firms for which at least one earnings observation is available in our out-of-sample period, and for which we have matched daily news data. To handle missing data, we match these firms with missing observations to firms in our main sample using the Fama French industry classification. We then utilize the parameter estimates obtained from the best group fixed effects model, as shown in Table (ref), to compute the nowcasts of log earnings ratios, either directly or based on its components. The results of this analysis appear in Table (ref).
Firstly, the results obtained through parameter imputation support the conclusion that nowcasting the components of the log earnings ratio yields higher quality predictions. This indicates that incorporating the individual components of the ratio improves the accuracy of the nowcasts. Secondly, the panel models with the parameter imputation method outperform the analyst consensus nowcasts in terms of prediction accuracy. This suggests that employing machine learning panel data models along with parameter imputation could be a straightforward yet effective approach in situations where earnings data is not available. Overall, these findings highlight the potential benefits of leveraging machine learning techniques and imputation methods for improving nowcasting accuracy, particularly in cases where earnings data may be missing.
This paper uses a new class of high-dimensional panel data nowcasting models with dictionaries and sg-LASSO regularization which is an attractive choice for the predictive panel data regressions, where the low- and/or the high-frequency lags define a clear group structure. Our empirical results showcase the advantages of using regularized panel data regressions for nowcasting corporate earnings either directly or using a decomposition which separates stock market return predictions and analyst assessments of a firm's performance. While nowcasting earnings is a leading example of applying panel data MIDAS machine learning regressions, one can think of many other applications of interest in finance. Beyond earnings, analysts are also interested in sales, dividends, etc. Our analysis can also be useful for other areas of interest, such as regional and international panel data settings.
\setcounter{page}{1} \setcounter{section}{0} \setcounter{equation}{0} \setcounter{table}{0} \setcounter{figure}{0}
The full list of firm-level data is provided in Table (ref). We also add two daily firm-specific stock market predictor variables: stock returns and a realized variance measure, which is defined as the rolling sample variance over the previous 60 days (i.e.\ 60-day historical volatility).
We select a sample of firms based on data availability. First, we remove all firms from I/B/E/S which have missing values in earnings time series. Next, we retain firms that we are able to match with CRSP dataset. Finally, we keep firms that we can match with the RavenPack dataset.
We create a link table of RavenPack ID and PERMNO identifiers which enables us to merge I/B/E/S and CRSP data with firm-specific textual analysis generated data from RavenPack. The latter is a rich dataset that contains intra-daily news information about firms. There are several editions of the dataset; in our analysis, we use the Dow Jones (DJ) and Press Release (PR) editions. The former contains relevant information from Dow Jones Newswires, regional editions of the Wall Street Journal, Barron's and MarketWatch. The PR edition contains news data, obtained from various press releases and regulatory disclosures, on a daily basis from a variety of newswires and press release distribution networks, including exclusive content from PRNewswire, Canadian News Wire, Regulatory News Service, and others. The DJ edition sample starts at $1^{st}$ of January, 2000, and PR edition data starts at $17^{th}$ of January, 2004.
We construct our news-based firm-level covariates by filtering only highly relevant news stories. More precisely, for each firm and each day, we filter out news that has the Relevance Score (REL) larger or equal to 75, as is suggested by the RavenPack News Analytics guide and used by practitioners, see for example kolanovic2017big. REL is a score between 0 and 100 which indicates how strongly a news story is linked with a particular firm. A score of zero means that the entity is vaguely mentioned in the news story, while 100 means the opposite. A score of 75 is regarded as a significantly relevant news story. After applying the REL filter, we apply a novelty of the news filter by using the Event Novelty Score (ENS); we keep data entries that have a score of 100. Like REL, ENS is a score between 0 and 100. It indicates the novelty of a news story within a 24-hour time window. A score of 100 means that a news story was not already covered by earlier announced news, while subsequently published news story score on a related event is discounted, and therefore its scores are less than 100. Therefore, with this filter, we consider only novel news stories. We focus on {\it five sentiment indices} that are available in both DJ and PR editions. They are:
\paragraph{\bf Event Sentiment Score} (ESS), for a given firm, represents the strength of the news measured using surveys of financial expert ratings for firm-specific events. The score value ranges between 0 and 100 - values above (below) 50 classify the news as being positive (negative), 50 being neutral.
\paragraph{\bf Aggregate Event Sentiment} (AES) represents the ratio of positive events reported on a firm compared to the total count of events measured over a rolling 91-day window in a particular news edition (DJ or PR). An event with ESS $>$ 50 is counted as a positive entry while ESS $<$ 50 as negative. Neutral news (ESS = 50) and news that does not receive an ESS score does not enter into the AES computation. As ESS, the score values are between 0 and 100.
\paragraph{\bf Aggregate Event Volume} (AEV) represents the count of events for a firm over the last 91 days within a certain edition. As in AES case, news that receives a non-neutral ESS score is counted and therefore accumulates positive and negative news.
\paragraph{\bf Composite Sentiment Score} (CSS) represents the news sentiment of a given news story by combining various sentiment analysis techniques. The direction of the score is determined by looking at emotionally charged words and phrases and by matching stories typically rated by experts as having short-term positive or negative share price impact. The strength of the scores is determined by intra-day price reactions modeled empirically using tick data from approximately 100 large-cap stocks. As for ESS and AES, the score takes values between 0 and 100, 50 being the neutral.
\paragraph{\bf News Impact Projections} (NIP) represents the degree of impact a news flash has on the market over the following two-hour period. The algorithm produces scores to accurately predict a relative volatility - defined as scaled volatility by the average of volatilities of large-cap firms used in the test set - of each stock price measured within two hours following the news. Tick data is used to train the algorithm and produce scores, which take values between 0 and 100, 50 representing zero impact news.
For each firm and each day with firm-specific news, we compute the average value of the specific sentiment score. In this way, we aggregate across editions and groups, where the later is defined as a collection of related news. We then map the indices that take values between 0 and 100 onto $[-1,1]$. Specifically, let \(x_i \in \{\text{ESS}, \text{AES}, \text{CSS}, \text{NIP}\}\) be the average score value for a particular day and firm. We map $x_i \mapsto \bar x_i\in[-1,1]$ by computing $\bar x_i$ = $(x_i - 50)/50.$
{\scriptsize
}