EconBase
← Back to paper

Machine Learning for Zombie Hunting: Predicting Distress from Firms' Accounts and Missing Values

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

87,546 characters · 10 sections · 89 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Machine Learning for Zombie Hunting: Predicting Distress from Firms' Accounts and Missing Values We are grateful to participants to the Bank of Italy/CEPR/EIEF conference on `Firm Dynamics and Economic Growth', to the Bank of England/King's College conference on `Modelling with Big Data and Machine Learning', to the Annual Conference of the JRC Community of Practice in Financial Research organized by the European Commission, and to the workshops on `Data Science for Impact Evaluation' jointly organized by KU Leuven and IMT School for Advanced Studies. We want to thank Tommaso Aquilante, Nicola Benatti, Kristina Bluwstein, Elena Cefis, Giulio Bottazzi, Dimitrios Exadaktylos, Nicol\'o Fraccaroli, Mahdi Ghodsi, Andreas Joseph, Francesca Lotti, Francesco Manaresi, Juri Marcucci, Andrea Mina, Chiara Osbat, Gianmarco Ottaviano, Giacomo Rodano, Andrea Roventini, Gabriele Rovigatti, Abhishek Samantray, Federico Tamagni, Francisco Queiro, Gias Uddin, Nicolas Woloszko and Nicol\'o Vallarano for their valuable comments.

titlepage\and Fabio Incerti. Laboratory for the Analysis of Complex Economic Systems, IMT School for Advanced Studies, piazza San Francesco 19 - 55100 Lucca, Italy.} \and Massimo Riccaboni. Laboratory for the Analysis of Complex Economic Systems, IMT School for Advanced Studies, piazza San Francesco 19 - 55100 Lucca, Italy.} \and Armando Rungi. Laboratory for the Analysis of Complex Economic Systems, IMT School for Advanced Studies, piazza San Francesco 19 - 55100 Lucca, Italy.}} \begin{abstract} In this contribution, we propose machine learning techniques to predict zombie firms. First, we derive the risk of failure by training and testing our algorithms on disclosed financial information and non-random missing values of 304,906 firms active in Italy from 2008 to 2017. Then, we spot the highest financial distress conditional on predictions that lies {\color{black} above a threshold for which a combination of false positive rate (false prediction of firm failure) and false negative rate (false prediction of active firms) is minimized}. Therefore, we identify zombies as firms that persist in a state of {\color{black} financial distress, i.e., their forecasts fall into the risk category above the threshold for at least three consecutive years}. For our purpose, we implement {\color{black} a gradient boosting algorithm (XGBoost) that exploits information about missing values. The inclusion of missing values in our predictive model} is crucial because patterns of undisclosed accounts are correlated with firm failure. Finally, we show that {\color{black} our preferred machine learning algorithm} outperforms (i) proxy models such as Z-scores and the Distance-to-Default, (ii) traditional econometric methods, and (iii) other widely used machine learning techniques. We provide evidence that zombies are on average less productive and smaller, and that they tend to increase in times of crisis. Finally, we argue that our application can help financial institutions and public authorities design evidence-based policies--- e.g., optimal bankruptcy laws and information disclosure policies. \\ \noindentKeywords: zombie firms; machine learning; financial constraints; bankruptcy; missing data\\ \noindentJEL Codes: C53; C55; G32; G33; L21; L25\\ \end{abstract} \setcounter{page}{0} \thispagestyle{empty}

\onehalfspacing

Introduction

In this paper, we propose machine learning techniques as suitable tools to make predictions about business failure when information about firm viability is partially undisclosed. Therefore, we define zombies as firms that persist in a high-risk status because their predicted probability of failure {\color{black} is above a threshold for which a combination of the false positive rate (i.e., predicting a firm that did not fail as failed) and the false negative rate (i.e., predicting a firm that failed as not failed) is minimized}. We prove that the latter is also the distributional segment in which the probability of transitioning to a lower risk of failure is minimal. According to our machine learning approach, the Italian proportion of zombie firms in the analysis period is between {\color{black} 1.5% and 2.5%}. Zombie firms are on average {\color{black} 19%} less productive and {\color{black} 17%} smaller than viable firms. Interestingly, we find that zombies are countercyclical, as their share increases in times of crisis and decreases in times of economic recovery.

The problem of identifying nonviable firms is important to scholars and practitioners, whether to assess the credit risk of an individual firm or to identify the part of an entire economy that is in trouble. {\color{black} Originally, the notion of zombie firms was associated with the phenomenon of “zombie lending", in which banks extend credit to otherwise insolvent borrowers. In some cases, zombie lending is a deliberate strategy to avoid a bank's budget restructuring while only apparently complying with capital standards set by financial regulators bonfim2020site, e.g., the case of Japanese banks in the 1990s peek2005unnatural,caballero2008zombie. More recently, schivardi2021credit studied Italian companies during the 2008 financial crisis and found that the misallocation of credit under zombie lending increased the default rate of otherwise healthy firms while decreasing the default rate of nonviable firms. This was because undercapitalized banks could restrict lending to more viable projects to avoid disclosing nonperforming loans in their portfolios. Paradoxically, severely distressed companies can appear resilient in times of financial crisis thanks to continued access to financial resources.

From a more general perspective, recent studies introduce zombies as firms that have persistent problems meeting their interest payments andrews2017confronting, mcgowan2018walking, banerjee2018rise, andrewspetroulakis, thus expanding the original category to include firms that are in some form of financial distress. Based on proxy indicators from financial accounting, they show that zombies account for a non-negligible share of modern economies--- up to 10% of incumbent firms--- while absorbing up to 15%, 19%, and 28% of the capital stock in countries such as Spain, Italy, and Greece, respectively. When market exit or restructuring delays occur, they drag down aggregate productivity by hindering the reallocation of resources in favor of healthier firms and preventing the entry of potentially more innovative and younger firms. It is therefore argued that identifying non-viable, financially distressed firms can be particularly useful in avoiding the misallocation of productive and financial resources.

Against this background, we argue that the empirical problem of assessing whether a firm is a zombie is closely related to the more general problem of determining its credit risk. Ultimately, a zombie is a non-viable firm that may escape bankruptcy despite its extreme financial distress--- i.e., despite scoring the highest credit risk. From another perspective, healthier firms are the furthest from bankruptcy and zombie status. Traditionally, credit risk has been studied from the perspective of a financial firm, which must assess the health of a company using information, albeit limited, from financial books and public records.\footnote{The seminal reference is to a departure from the modiglianimiller theorem, according to which capital structure should not be relevant to a company's value if there are no market frictions, including bankruptcy costs. Thus, a firm's ability to raise external financing should depend solely on the profitability of its investment projects. However, since financial market frictions cannot be eliminated, a firm's capital structure actually provides information about the profitability of the firm and its assets. See also the discussion of rajanzingales for an international perspective} Thus, for decades, academics and practitioners have attempted to determine a firm's profitability after benchmarking exercises on firm-level indicators of financial constraints--- e.g., in estimating Z-scores altman1968financial, altman2000predicting, Distance-to-Default merton1974pricing, or investment- to- cash- flow sensitivity fazzarihubbardpetersen. However, information on corporate viability may be incomplete due to strategic disclosure of relevant information and simplified financial records for certain categories of unlisted companies.

Thus, we propose a machine learning approach to predict credit risk and zombie status from incomplete financial accounts. To show the potential of our approach, we work with a sample of 304,906 Italian firms over the period 2008-2017. Italy is a compelling example of a country that hosts a relevant share of inefficient firms that hinder the growth potential of the economy calligarisetal.

The underlying intuition is simple: based on the experience of firms that failed in previous periods, we derive predictions about the risk of failure of active firms. Each time we compare the observed outcomes with the predicted outcomes, the algorithm updates and reduces the prediction errors in the next periods after processing new “in-sample" information about the financial accounts. In the end, we obtain a probabilistic measure of the likelihood that a business will fail. Our machine learning framework improves upon existing benchmark models by leveraging a rich set of firm-level economic and financial indicators that potentially contain diverse information about both the firm's core economic activity and its ability to meet financial obligations.

Importantly, we find that emerging patterns of missing financial accounts are correlated with firm failure, possibly due to the fact that managers are more likely to conceal accounts when they are in financial distress. We provide evidence that most missing variables are often those that have been used as proxies for zombies or financial constraints in the previous literature. Therefore, we implement {\color{black} our missing-aware methodologies} incorporating patterns of undisclosed accounts {\color{black} and test whether a substantial improvement in prediction occurs when missing data are properly handled}. Using the missing values information as another predictor of outcome, we show that the best predictive algorithm in our setting, eXtreme Gradient Boosting (XGBoost) chen2016xgboost, explains up to $0.97$ of the Area-Under-the-Curve (AUC), and its Precision-Recall (PR) performance reaches $0.76$. XGBoost outperforms credit scoring models (i.e., Z-score and Distance-to-Default models), standard econometric methods (i.e., logistic regression), and also other machine learning techniques (i.e., Classification and Regression Tree and Random Forests).

We argue that asymmetric and undisclosed (missing) information is ubiquitous in corporate financial accounts. Therefore, simply omitting records with missing values would severely impair the search for predictors of zombies and lead to a loss of precision and bias little2019statistical. If the missing information is correlated with firm failures, excluding these firms would significantly reduce the number of failures in the sample and thus hinder the algorithm's learning for these instances. This problem could potentially be avoided by a missing data imputation approach that would allow the inclusion of these firms in the training sample. On the other hand, the missing information per se constitutes relevant information to learn from failures. This may be the case if nonviable firms have the option of not disclosing information in their financial accounts. As a result, the mechanism in the data is Missing-Not-At-Random (MNAR), and the complete records are not a random sample of the population of interest. In this second case, if some companies consistently avoid disclosing part of their information, imputation of missing data would not be sufficient to solve the problem. Indeed, we know from previous literature that techniques for imputing data, such as mean or median imputation, can fail in the presence of MNAR because they distort the empirical distributions.\footnote{Further limitations of traditional approaches to imputing missing data are discussed in he2010multiple, white2018imputation, little2019statistical.}}tatistical}.}}

{\color{black} Based on the previous considerations, we choose an approach that can simultaneously incorporate missing data into our analysis while being robust to MNAR. We implement two methods specifically designed to deal with non-random patterns of missing data: XGBoost and Bayesian Additive Regression Trees with Missing Incorporated in Attributes (BART-MIA) kapelner2015prediction. Although these methods differ in implementation, they have very similar routines for dealing with MNAR patterns. XGBoost uses default directions, also known as block propagation josse2019consistency, to group all incomplete observations and send them to one side of the tree. The MIA method twala2008good, extended by kapelner2015prediction, JSSv070i04, strengthens block propagation so that missingness can be used as an explicit feature to compute the best splits. According to josse2019consistency, MIA can handle both informative and non-informative missing values. Our results suggest that both XGBoost and BART-MIA effectively capture the direct influence of missing values as either implicit (block propagation) or explicit (MIA) predictors of the response variable. Finally, in the following analyzes, we show that XGboost has a significantly lower computational cost and relatively higher predictive power than BART-MIA. }

To shed more light on the information that contributes most to prediction, we use an interpretable framework that has its roots in game-theoretic Shapley values, introduced by strumbelj2010efficient. Shapley values have recently been proposed to identify the economically meaningful nonlinearities learned by machine learning models buckmann2021interpretable. We find that no single financial indicator predicts failure better than the ensemble of predictors used in our machine learning approach.

Economically informative groups of variables have heterogeneous predictive power. In particular, indicators of firms' financial constraint or previously used indicators of zombies are important. Nevertheless, they are of secondary importance when compared to information on firms' {\color{black} financial accounts, indicators of corporate governance}, and the presence of non-random missing values. Our results confirm that machine learning techniques perform better than single indicators when incorporating as much valuable in-sample information as possible while updating each time there is new out-of-sample information because prediction errors dynamically decrease after independent tests that minimize the discrepancies between realized and predicted outcomes athey2018impact.

Our framework is of particular interest to policymakers in designing optimal bankruptcy laws.\footnote{See also the suggestions by the European Directive 2012/30/EU, and the recent Italian Law on business failures on October 19, 2017, n. 155, which provides the legal basis for early notification of corporate crises to improve targeted interventions.} Tracking a company's bankruptcy risk allows all stakeholders, not just creditors, to understand whether there is an opportunity for restructuring and, if not, to prevent incumbent, albeit non-viable, firms from wasting additional economic resources. Evidence-based, interpretable methods are even more important after the recent pandemic crisis, as we believe that financial support must be targeted to companies that have a real chance of recovering and staying on their feet in normal times to avoid misallocation of resources.\footnote{For a first analysis of the impact of the pandemic crisis on Italian companies, see SchivardiRomano}

Data and preliminary evidence

We obtain the financial accounts from the ORBIS database\footnote{ORBIS firm-level data orbis have become a common source of global financial accounts. For previous use of this database, see gopinath2017capital and cravinolevchenko, among others. Coverage of smaller firms and some financial accounts may change across countries as national business registries impose different filing requirements, as observed in the validation exercises of kalemli2015construct and gal2013measuring. In the case of Italy, the original information provider for Italian financial accounts is CERVED, a credit rating agency. Bureau Van Dijk standardizes and translates the original financial accounts to make them comparable between countries. Note that, unlike other platforms of the same Bureau Van Dijk (e.g., AIDA or AMADEUS), ORBIS does not drop exiting firms in our analysis period. It supplements the financial accounts with other information from various sources on ownership, management and intellectual property rights, which we also use for predictions.}, compiled by the Bureau Van Dijk, for manufacturing firms active in Italy for at least one year from 2008-2017. Italy is a compelling case to study business failure and zombie firms: it is a country where relatively inefficient firms hinder the economy's growth potential calligaris2016italy, bugamelli, perpetuating geographic divergence rungi2019heterogeneous, and are studied extensively by international organizations mcgowan2018walking, andrewspetroulakis.

For our purpose, we use two main variables that help us identify business failure: the status of a firm and the date on which it becomes inactive. Table (ref) shows our sample coverage by firm status over in the period of analysis. We assume that a firm has failed in the first year if it is reported as “bankrupt”, “dissolved”, or “in liquidation”, as in the original data. Overall, the share of exiting firms accounts for about 5.7% of the total sample, which is close to the average official 6.3% obtained by ISTAT, the national statistics office, for the same period.

{\color{black} Figure (ref) shows the proportions of firm failures by NUTS 2-digit region. As expected, we find a higher concentration of failures in the north and center of the country, where economic density is higher. It is noteworthy that we fully represent the whole Italian territory since we detect business failures in every region during our period of analysis.}

table[table omitted — 412 chars of source]
figure[figure omitted — 410 chars of source]

We use a set of economic and financial indicators to train our predictive models. The battery of predictors includes (i) original financial accounts at the firm level; (ii) widely used indicators to proxy firm-level financial constraints; (iii) indicators previously used to detect zombie firms; (iv) indicators included as warnings of corporate crises in the recent Italian bankruptcy law.\footnote{A recent reform of the bankruptcy law (L. 155/2017 and DL. 14/2019) proposes an early warning system based on indicators identified by practitioners, the purpose of which is to identify companies in distress in time to intervene to preserve entrepreneurial capabilities and find a way out of the crisis. It delegated practitioners (in particular the National Association of Chartered Accountants: Consiglio Nazionale dei Dottori Commercialisti e degli Esperti Contabili ) to draw up a list of indicators that could help assess the state of crisis of a company} Each predictor we consider is described in detail in Table (ref) in the Appendix (ref).

Note that many indicators we select as predictors have been used in various frameworks to assess the extent to which a firm is in trouble. From our machine learning perspective, they cannot be interpreted as drivers of failure. It is sufficient that they contribute, albeit in a small way, to the assessment of a company's health. In our predictive framework, they might even border on multicollinearity, e.g., in the case of different measures of efficiency, liquidity, and solvency ratios. Since we are not interested in identifying a causal contribution to firm failure, high collinearity does not pose a problem for our predictions makridakis2008forecasting, shmueli2010explain. On the contrary, we will discuss in Section (ref) how, by construction, one cannot separate the empirical contribution of an indicator from the entire set in the context of a pure prediction problem. For this reason, the predictors we use should be considered as an inseparable ensemble, and a discussion of the statistical significance of the individual predictors is not relevant in our framework.

In addition to mandatory and basic information (volume of activity, profits, location, industry affiliation, ownership, and intellectual property rights), many other financial accounts have different patterns of missing values over time. {\color{black} In the Appendix (ref), Figure (ref), we visualize a map of missing values in our sample. The frequency of missing values affects all firms in the data, both active and failed, and it is more concentrated in the financial accounts of recent years. In addition, liquidations exhibit less pronounced patterns over time than the other failure categories}. After running a series of chi-squared tests (see Table (ref)), we find a positive statistical relationship between the patterns we observe in the sample and the event of a firm's failure. In short, a firm is more likely to fail if a pattern of missing financial accounts is observed. The exercise performed in Table (ref) clearly shows such correlations. We run a simple logistic regression using the observed failure of a firm as the dependent variable. We then add a binary regressor equal to one if the predictor is missing at least once in the three years prior to failure. Fixed effects by NUTS 3-digit region and NACE 2-digit industry are included. We report results for each predictor per row in Table (ref).

table[table omitted — 1,090 chars of source]

The above correlations are particularly relevant to the scope of our analyses. The main problem is sample selection when observations are selectively missing for some categories of firms. In this case, there are two potential sources of sample selection bias: (i) distressed firms vis \`{a} vis firms that are not in distress, as the former may have the incentive to disclose less information than the latter; (ii) smaller companies firms vis \`{a} vis bigger firms because the first are often exempted from a complete financial report, in accordance with Italian Regulation.\footnote{Under Italian civil law, companies that do not list financial activities on the stock exchange have the option to provide more aggregate financial reports if their size does not simultaneously exceed two of the following thresholds in one or two consecutive periods: i) 4,400,000 euros in total assets; ii) 8,800,000 euros in operating revenues; iii) 50 employees. A simplified financial statement always includes the most significant items in the first or second position digit of the aggregation.} Of course, the two sets of firms may overlap, as smaller companies may also be the ones that are proportionally more affected by financial difficulties. Using a missing-aware procedure, as described in Section (ref), allows us to consider both sources of sample selection when patterns of missing financial accounts emerge since the algorithm considers such patterns as another predictor of firm failure.

Interestingly, we note that most of the missing variables are also those used in previous work as proxies for zombies or financial constraints. Take, for example, the case of the Interest Coverage Ratio (ICR), which is derived as the ratio between a firm's earnings before interest and taxes (EBIT) and its interest expense. If the ICR is less than one, bankofengland2013 assumes that a company is a zombie because it is having trouble meeting its financial obligations. In our sample, we find that about 19% of firms have an ICR smaller than one, but at the same time, there are 62.50% firms whose ICR information is not available at all. Moreover, according to bankofkorea2013, negative value added is the most appropriate indicator for evaluating a zombie status, as it indicates that intermediate inputs have a higher market value than the firm's output. In the case of Italy, about 64.27% of enterprises do not report their value added in at least one period, while about 3% of them report a negative value. A negative value added is, of course, a more severe condition than a negative profit, since a company can make no profits without destroying economic value. Indeed, firms' profitability is at the heart of two similar proxies for zombies used by schivardi2021credit when comparing firm-level profits to an external benchmark. Repeating the same exercises, we find that about 3% of Italian firms are distressed, while a large part of the sample (62.50%) does not report any information on the predictor. Finally, both caballero2008zombie and mcgowan2018walking perform another benchmarking exercise, comparing the interest a firm pays to raise external finance with the cost opportunity to invest in alternative, safer investments. In our case, when we try to reproduce the same exercise with the yields of Italian government bonds with a ten-year maturity, we find that there is a high proportion (60.29%) of companies for which we have no information in at least one period in which they were active.

We also include an estimate of Total Factor Productivity (TFP) as a predictor, following the methodology proposed by ackerberg2015identification, to account for the simultaneity bias arising from ex-post adjustments in combining factors of production. In this respect, firm-level TFP allows us to make predictions based on the ability to transform inputs and sell output in the market. Indeed, the relationship between financial constraints and productivity is one of the most debated issues aghionetal, ferrandoetal. The simple assumption for our forecasting models is that less productive firms are the ones that have more difficulty surviving in the market. At the same time, zombie firms have also often been defined in terms of (lack of) productivity mcgowan2018walking, andrews2017confronting, andrewspetroulakis, schivardi2020identifying.\footnote{Please note how white2018imputation recently highlighted the limitations of using missing imputation techniques for the variables used to calculate TFP in the US.}

Empirical strategy

It is difficult to identify non-viable companies for obvious reasons. If financial accounts are in bad shape, one could argue that it is only a matter of time before they become more competitive if conditions are right. If the balance sheets are good, one could argue that the worst is yet to come because bad management decisions will show up later. Trivially, only the already bankrupt companies were certainly not viable at some point. Still, an outside observer will never know when that happened because the manager of a company in trouble has an incentive not to disclose private information.

In principle, an analyst would like to observe the entire event horizon to discount all possible scenarios and understand the value of a company and its investment projects. However, this is not possible because it is the typical double problem of a financial institution facing uncertainty in the presence of information asymmetries. On the one hand, the company has a clear information advantage in its investment plans. On the other hand, both the financial institution and the company have a limited ability to predict future economic shocks, which can have either positive or negative effects.

In their seminal works altman1968financial and ohlson1980financial apply standard econometric techniques---i.e., multiple discriminant analysis (MDA) and logistic regression---to assess the probability of firm bankruptcy. Following these contributions and the Basel Accord II in 2004, default forecasts are based on standard reduced-form linear regression approaches. However, these approaches may fail because their limited complexity precludes nonlinear interactions among predictors, while their ability to handle large sets of predictors is limited due to potential multicollinearity problems.

Machine learning algorithms compensate for these shortcomings by providing flexible models that allow nonlinear interactions in the space of predictors and the inclusion of a large number of predictors without the need to invert the covariance matrix of the predictors, thereby circumventing multicollinearity linn2019estimating. In addition, machine learning models are directly optimized to perform the prediction task, resulting in better prediction performance in many complex situations.

The data science literature has already developed exercises to predict corporate failures using financial accounts, but without a clear economic and financial framework. In this context, bargagli2021supervised provides a comprehensive review of the recent literature on the use of machine learning for the analysis firm dynamics. The authors highlight that the majority of work dealing with the prediction of bankruptcy or financial distress uses either decision tree-based techniques behr2017default, linn2019estimating, moscatelli2019corporate, deyou2020, davies2023predicting, incerti2022two or neural network-based methods alaka2018systematic, bredart2014bankruptcy, hosaka2019bankruptcy, sun2011dynamic, tsai2008using, tsai2014comparative, wang2014improved, lee1996hybrid, udo1993neural.

With this in mind, we propose a machine learning procedure that uses past information about previously failed firms to estimate the probability that another firm in similar condition will go bankrupt. The broader the variety of past experiences on which we can draw, the more accurate the prediction about the health - or lack thereof - of a company kleinberg2015prediction. Ultimately, our perspective is on a firm's (lack of) resilience, using potentially any observable data that might hold information about the firm's viability. In the end, we obtain a probabilistic measure at the firm level, ranging from 0 to 1, which tells us the probability that a firm will exit the market in the next period, given that other firms in a similar situation have done so. As shown in Figure (ref), we can assess the distance of each company from the highest financial distress.

Let us consider a generic predictive model in the form:

equation[equation omitted — 102 chars of source]

where $Y_{i,t}$ is the binary realization of the outcome at time $t$ that takes the value 1 if the $i$th firm exits the market and takes the value zero otherwise, while $\mathbf{X}_{i,t-1}$ is the $P$-dimensional vector of firm-level predictors in the previous time period, where $P$ is the number of predictors included in the model. The functional form linking the predictors to the outcomes is determined by the generic supervised machine learning procedure used to predict out-of-sample information. In short, the generic algorithm chooses the best in-sample loss-minimizing function in the form:

equation[equation omitted — 213 chars of source]

where $F$ is a function class from where to pick $f(\cdot)$, and $R\big(f(\cdot)\big)$ is the generic regularizer that summarizes the complexity of $f(\cdot)$ mullainathan2017machine. In our case, the function $f(\cdot)$ is an element from the family of classification trees or a combination of them. The set of regularizers, $R$'s, will change according to the standards adopted by each method. Ultimately, each algorithm takes a loss function $L({f}(x_i), y_i)$ as input and searches for the function that minimizes the prediction losses.

{\color{black} To determine the region of highest financial distress, we show in the following analysis that a cutoff of 0.9 minimizes the combination of false-positive (false prediction of firm failures) and false-negative (false prediction of active firms) rates. This empirical result supports the choice of the highest decile of the predicted risk distribution as the optimal threshold for identifying cases of critical financial distress. We will also prove that using the highest decile to identify zombie firms is successful, as this is the segment where prediction accuracy is highest, and after which the probability of transitioning back to lower levels of financial distress is minimal.}

figure[figure omitted — 196 chars of source]

We would like to emphasize that we are not interested in identifying the causes of firm failure, since we are dealing with a pure prediction problem,--- i.e., the probability of the failure of firm $i$ at time $t$. Nevertheless, we perform variable selection to evaluate the contribution of the predictors to the estimated risk. We show how the predictors of failure can change over time, under the circumstances in which we make the predictions, each time there is an update with new out-of-sample information.

Decision tree learning for failure prediction

Consistent with the literature on bankruptcy and firm exit predictions presented above, we derive---given new out-of-sample information--- a prediction for the failure of each firm based on its current financial accounts, both for established firms that have operated in previous periods and for firms entering the market for the first time. From another perspective, we interpret this probability range as the degree of risk of an investor who has no information other than that contained in the current financial accounts.

{\color{black} Our study introduces an innovative approach to failure prediction by demonstrating the remarkable effectiveness of tree-based machine learning models, especially in the presence of missing data. While neural networks are widely considered the best option for prediction tasks with homogeneous data such as images and text, they often fail for “problems with heterogeneous features, noisy data, and complex dependencies" prokhorenkova2018catboost. Decision tree-based algorithms are considered suitable tools in these cases due to their flexibility and high performance. The Classification and Regression Tree (CART) algorithm, first introduced by friedman1984classification, is a widely used decision tree algorithm that constructs binary trees where each node is divided into only two branches. Figure (ref) shows how binary partitioning works in practice, using a simple example with only two predictors. Ensemble methods have been developed to improve the stability of decision tree estimates, which combine multiple weak learners into one strong learner. Bagging and boosting are two popular methods for this purpose breiman1996bagging, freund1996experiments. In line with recent developments in the literature, we perform a comparative evaluation of firm failure prediction using two state-of-the-art approaches: XGBoost chen2016xgboost and BART-MIA chipman2010bart. Both statistical models are specifically designed to deal with missing values, which makes them ideal for dealing with severe cases of missing financial data in corporate accounts. As illustrated in Section (ref), the absence of financial data does not follow random patterns, or at least does not follow the failure of a company completely randomly. When a company is financially distressed and/or smaller, it is more likely that data is missing from some accounts. Simply discarding the missing observations would introduce selection bias and exclude certain categories of firms with a higher probability of failure. Crucially, both XGBoost and BART-MIA include patterns of undisclosed accounts as a feature of the model.}

figure[figure omitted — 1,120 chars of source]

{\color{black} XGBoost is an ensemble method based on decision trees with a gradient boosting framework. It has proven to be highly competitive and reliable on various prediction tasks, mainly due to its efficient use of computation time and memory resources gumus2017crude, abbasi2019short, li2020diabetes. It is equipped with various optimization techniques and tools, such as a percentile-based split-finding algorithm, parallelized tree building, a depth-first approach to tree pruning, and efficient handling of missing values through default directions. In addition, it uses regularization to prevent overfitting bentejac2021comparative.

BART-MIA is a robust Bayesian ensemble of trees methodology that combines the traditional BART chipman2010bart with a variant called MIA twala2008good. This variant was specifically designed to deal with MNAR (Missing Not At Random) patterns in the data. BART is a well-established method that has shown excellent performance on various prediction tasks murray2017log, linero2018bayesian1, linero2018bayesian2, hernandez2018bayesian. This can be attributed to two key factors: first, BART contains a noise component that mitigates the overfitting problem common to random forest methods, and second, it has shown consistently strong performance on standard model specifications, avoiding time-consuming hyperparameter tuning procedures. This property is particularly important because it reduces the researcher's dependence on parameter choice and minimizes the computational time and cost associated with cross-validation. See the appendix (ref) for the technical details of the two models.}

Validation against other machine learning techniques

{\color{black} We compare the predictive performance of our methods that account for missing values, XGBoost and BART-MIA}, with the Conditional Inference Tree hothorn2006unbiased, the Random Forest breiman2001random, and the Super Learner van2007super. The Conditional Inference Tree we use is a simple variant of the Classification and Regression Tree (CART) algorithm friedman1984classification, based on a significance testing procedure that avoids bias towards variables with many possible splits oden1975arguments, loh2002regression, hothorn2006unbiased. The Random Forest is an ensemble method that combines different trees to obtain stronger predictive power. Each tree is created by randomly selecting different variables from all possible predictors and randomly selecting a subset of the total number of observations breiman2001random. The Super Learner van2007super is based on a weighted combination of other algorithms. {\color{black} We build it as a convex combination of the following models: Logistic regression, CART, Random Forest, BART, and XGBoost.

To evaluate the performance of the models, we use a standard five-fold cross-validation, where the dataset is divided into five groups and the model is trained on four of them, while the fifth group is used as a validation set, iteratively repeating this process. In this way, we can comprehensively evaluate the predictive performance of the model over the entire time series.

To ensure comparability with the other models that do not consider missing values, we perform a Complete-Case analysis. The latter is a widely used approach for dealing with missing values in models that are not designed to do so. In this approach, all observations with at least one missing value in the predictors are discarded. Although XGBoost and the BART-MIA models could use the entire dataset without excluding missing observations, we perform a Missing-Aware analysis to ensure fair performance evaluation of all models: we train and test XGBoost and BART-MIA using five-fold cross-validation with identical dimensions as the original folds but allow missing observations to be included in the sample. For clarity, we use the terms Missing-Aware XGBoost (MA-XGBoost) and BART-MIA for the models trained on observations with missing values. In this case, MA-XGBoost uses default directions to deal with missing values, while BART-MIA uses the MIA procedure. On the other hand, we will simply refer to the models trained on Complete-Case data as XGBoost and BART.}

table[table omitted — 1,884 chars of source]
figure[figure omitted — 738 chars of source]

Table (ref) shows the results of the horse race of the models. It is clear that the missing-aware models consistently outperform the state-of-the-art methods involved in the Complete-Case analysis. To compare the predictive power of the different methods, we show five different performance measures commonly used for classification problems: the Area Under the receiver operating characteristic Curve (AUC); the area under the Precision-Recall (PR) curve; F1-Score; Balanced ACCuracy (BACC); adjusted $R^2$. Both AUC and PR vary between 0 and 1, with 0 indicating complete misclassification and 1 indicating perfect prediction. The AUC hanley1982meaning is a general measure of predictive power that tells us the extent to which we are able to classify failures vis \`{a} vis non-failures, hence with an accent on the false discovery rate (FDR). PR is particularly useful for our scope because it takes into account both the total proportion of true failures that we are able to predict in the data (i.e., the sensitivity/recall of the predictions) and for the proportion of predicted failures that turn out to be true failures (i.e., the precision of the predictions). Indeed, assessing sensitivity/recall alone could be misleading in a zero-inflation environment such as ours, where the number of non-failures systematically exceeds the number of failures.\footnote{For more details, see saito2015precision, fawcett2006introduction.} Figure (ref) reports the average ROC and PR curves over the five-folds for the Complete-Case Logit and {\color{black} MA-XGBoost}. The F1 score van1979information and the BACC brodersen2010balanced are used for cases of unbalanced data. The former is based on a harmonic mean of precision and \textit{recall}, while the latter is a simple average between the rate of true positives and the rate of true negatives from our predictions.

{\color{black} In the Complete-Case analysis, the best performing model is the Super Learner, which, however, stops at $0.9231$ and $0.5463$ in terms of AUC and PR, respectively.\footnote{It is noteworthy that the Super Learner performs better than the Random Forest in terms of accuracy, but not in terms of F1 score. The reason is that the Super Learner algorithm van2007super is optimized to find a convex combination of algorithms that minimizes the accuracy of the ensemble method. The latter strategy is not optimal in our case because the dataset is unbalanced.} The results of the Missing-Aware analysis differ significantly from those of the Complete-Case ones. MA-XGBoost shows a PR of $0.7591$ and an AUC of $0.9685$, followed by BART-MIA, whose performances are almost identical. Interestingly, the most notable improvement comes from higher precision. This is reflected in the larger departure of the Missing-Aware models with respect to the performance measures that use precision (namely, PR and F1 score). This is critical to our objective as it means that Missing-Aware models have better predictive power in identifying companies that will close within a year based on the predictions. However, training the BART-MIA model is computationally intensive and takes 45 times longer than MA-XGBoost. Due to its significantly lower computational cost and higher predictive power, we choose MA-XGboost as the optimal prediction algorithm for the following analysis.

To ensure the robustness of our results, we perform a number of robustness checks in Appendix F. We investigate whether (i) imputing missing observations could improve the performance of predictive models that do not include information on missing attributes; (ii) whether excluding liquidations from the set of firm failures would affect the performance of the model. For (i), we use out-of-range and median imputation methods josse2019consistency. We find that the performance of traditional models increases significantly when missing values are considered, especially for tree-based methods. This result is consistent with our observation that missing entries signaled by a constant value either over the entire dataset (out-of-range imputation) or over a specific feature (median imputation) are included in tree splitting. As a result, the model can now account for missingness in a similar way to the standard directions and MIA techniques, leading to comparable results. With respect to (ii), we test whether predictive performance depends on the type of failure (e.g., bankruptcy, dissolution, M&A, and liquidation). We find that excluding the largest group of business failures -- i.e., liquidations -- has little effect on the performance of the models, suggesting that the dependence on the type of failure is negligible.}

Validation against proxy models of credit scoring

So far, we have compared the predictions of failures from different econometric and machine learning techniques and concluded that {\color{black} MA-XGBoost is the best choice in terms of predictive power and computational burden.} Here, we compare our baseline predictions with widely known proxies for firm-level credit scores: the Z-scores altman1968financial and the Distance-to-Default merton1974pricing.

Z-scores involve putting a selection of financial ratios (profitability, leverage, liquidity, solvency, and volumes of activity) into an equation with some weights to proxy their relative importance. The weights are taken from the literature and from previous scholarly estimates of the relative importance of these indicators in assessing a firm's distress. In this way, a threshold is obtained, the crossing of which indicates a high probability of future bankruptcy. Unlike Z-scores, the Distance-to-Default (DtD) by merton1974pricing focuses specifically on a firm's ability to meet its financial obligations. The original intuition is that a firm's equity can be modeled as a call option on its assets. Thus, to build such a model, one must combine firm-level accounts (firm assets, debt, market value) and information from financial markets (risk-free interest rate, standard deviation of stock returns). Finally, one puts the variables into an equation that gives the value of a theoretically fair call option.\footnote{Following the insights of the distance-to-default model, black1973pricing developed their widely known model based on the observation that one can eliminate a systemic risk component by hedging an option.}

{\color{black} We evaluate the predictive power of both Z-scores and Distance-to-Default (DtD). The Precision and False Discovery Rate (FDR) are given in Table (ref). Similar to our previous analysis, we performed five-fold cross-validation. We start with the non-missing observations in Z-scores and DtD and randomly split them into five folds. Based on the Z-scores and DtD measures, distressed firms have lower scores. Therefore, we used the first ten percentiles of the in-sample scores as cutoffs for classifying the out-of-sample observations. We found that the DtD predictions have higher precision (0.3314 vs. 0.2239) and lower False Discovery Rate (0.6686 vs. 0.7761) than the Z-scores. We then train a MA-XGBoost model with the same cross-validation routine and use the in-sample percentiles as cutoffs to classify the out-of-sample observations. The MA-XGBoost model considers the highest risk firms to be in the right tail, unlike Z-scores and DtD. Therefore, the first column of Table (ref) represents the bottom ten percentiles (1-10) for DtD and Z-scores, while for MA-XGBoost the top ten percentiles (99-90) symmetrically. It is clear that MA-XGBoost outperforms both DtD and Z-scores in all percentiles.}

table[table omitted — 1,595 chars of source]

Unboxing the black box: Shapley values for predictors

In the previous sections, we validated our proposed {\color{black} MA-XGBoost} model against state-of-the-art financial indicators, traditional econometric models, and widely used machine learning techniques. Despite its excellent predictive performance, the nonlinear relationships learned by {\color{black} MA-XGBoost} are not directly observable, leading to the recurring criticism that machine learning models are black boxes despite their improved accuracy. However, there are a number of tools to shed light on complex nonlinear relationships in machine learning predictions to make models interpretable lundberg2020local.\footnote{Interpretability is a non-mathematical concept, but is often defined as the degree to which a human can understand the cause of a decision or consistently predict the outcomes of the model kim2016examples, miller2018explanation, stoffi2021assessing, lee2020causal, bargagli2022heterogeneous, bargagli2020essays, bargagli2020causal.} Popular methods for assessing predictor importance include permutation importance or Gini importance for tree-based models breiman2001random, LIME ribeiro2016model, DeepLIFT shrikumar2017learning, and Shapley values strumbelj2010efficient.

We chose to use Shapley values because they are not only able to reveal the complex patterns linking the predictors and the outcome, but also guarantee a number of desirable properties: efficiency, null payoffs, symmetry, monotonicity and linearity rozemberczki2022shapley. In the case of a general ML model, these properties state that: the importance of individual variables must add up to the goodness-of-fit of the model trained with the entire set of variables (efficiency)\footnote{This property is crucial because it allows quantifying the contribution of each variable to the overall performance of the model.}; a variable that does not improve the goodness-of-fit of the model is assigned a contribution of zero (null payoff); two variables that make the same marginal contribution to the global goodness-of-fit have the same importance (symmetry); if variable A consistently contributes more to the global goodness-of-fit of the model than variable B, the Shapley value of A must be higher than that of B (monotonicity); for two subsets of the same data set, the global Shapley value of a variable is equal to the sum of the Shapley values calculated separately for the same variable in the two subsets (\textit{linearity}).

An intuitive way to understand how Shapley values for variable importance work is to include each variable in the predictive model in random order. All variables in the model contribute to the final predictions (and thus to the goodness-of-fit of the model). The Shapley value of a variable is the average change in prediction (and goodness-of-fit) experienced by the set of variables already in the model when that variable is included molnar2020interpretable. From a mathematical perspective, assume that $S$ is a $q$-dimensional subset of variables, $m$ is a generic variable in $S$ ($m \subset S$), and $v(T)$ is a generic value function that takes in the subset $S$ and returns real-valued payoff of the model (e.g., the goodness-of-fit) created using $S$ or subsets thereof. Then the Shapley value $\phi_m(v)$ for a generic variable $m$ is:

equation[equation omitted — 164 chars of source]

Using (ref), we see how the Shapley value is computed by calculating a weighted average gain in payoff (read: gain in goodness-of-fit) that the variable $m$ yields when included in all subsets of variables that exclude $m$.

figure[figure omitted — 192 chars of source]
figure[figure omitted — 202 chars of source]

The results of our Shapley value analyses are shown in Figures (ref) and (ref). Figure (ref) reports the results of the average Shapley Values over the time frame of our analysis for each variable used as input. Figure (ref) shows the same value, but this time all variables are aggregated into some economically relevant groups.\footnote{A full description of how variables are aggregated into groups is described in Appendix Table A1. Governance variables are indicated by (G), financial constraints by (FC), financial accounts by (FA), zombie indicators by (ZI), area by (A), innovation by (I), productivity by (PDC), sector by (SE), profitability by (PFT), size by (SI).}

{\color{black} First, we note that no single financial indicator has better predictive power than the aggregate of indicators. On their own, most indicators have relatively low predictive power. The financial indicators with higher Shapley Values -- i.e., that contribute most to the model's predictions -- are the corporate control indicator (Corporate control), the missingness in the financial accounts (Missingness), some profit and loss characteristics (Net income, Taxes, Revenues), the size and age indicators (Employees, Size-age), the resources invested in the company (Shareholders funds) and the long-term physical assets (Fixed assets).\footnote{See Table (ref) for a full description of the predictors and their construction.} When analyzing the groups of indicators, it appears that original financial accounts with no further elaboration carry the greatest explanatory power. This result is reasonable considering that most predictors fall into this category. In addition, the governance, size, and financial constraint indicators also have relatively high predictive power. Although they compete with groups consisting of multiple predictors, missingness remains an important factor. On the other hand, firms' productivity, sector, and geographic location of firms also have high predictive power, but are of secondary importance compared to the aforementioned indicators.

Finally, to check for robustness, we conduct a logit-LASSO analysis to evaluate, using a shrinkage method, the most frequently selected variables that predict a firm's distress. We find substantial overlap, with corporate control, profit and loss characteristics, and firm size-age indicators among the most important predictors. The full analysis can be found in the Appendix (ref).\\

We conclude this Section with a discussion of the possibility that firms manipulate financial accounts and the predictive power of missing values in this case. We argue that there are two reasons for the occurrence of missing values in firms' financial accounts:

enumerate• Smaller firms are often exempt from reporting complete information to national registries if they fall below certain size thresholds because it may be too costly for them to implement complex accounting procedures.\footnote{In footnote 7 of the manuscript, we indicate the combination of size thresholds that the Italian law provides.} • Some financial accounts are optional each year or may be incomplete (e.g., number of employees, management of inventories, cash prospects, etc.).

In any case, there is room for manipulation by firms that know that banks and policymakers do not use random missing values to predict their financial distress. They have two choices: they can provide truthful information and thereby disclose their true financial distress, or they can provide false information and thereby potentially commit accounting fraud. In the first case, the missing values are no longer meaningful, but our algorithm can still rely on the disclosed financial accounts for prediction. In the second case, we would have false negatives, i.e., companies that are predicted to be financially viable but are not. More generally, we conclude that a limitation of our approach is that we assume that companies do not report false information; otherwise, we would have statistical noise that reduces prediction accuracy.}

A case for zombie firms

{\color{black} In our analysis, we showed how we can use machine learning to predict firms' failures.} But how do we spot a zombie firm? There is no consensus on the exact meaning of the zombie status, other than its suggestive power. To date, scholars have merely adopted various thresholds based on a proxy assessment of one or more available financial indicators in the absence of more precise theoretical guidance. Ideally, a company's competitiveness and financial constraints should be viewed from a dynamic perspective that considers the entire horizon of future events. Considering all future threats and opportunities and their impact on profit and net cash flow at the firm level could be the theoretical equivalent of “zombie firms". Without the latter, empirical identification is left to the creativity of academics and practitioners. caballero2008zombie define zombies as firms that receive subsidies in the form of bank loans after observing how interest payments compare to an estimated benchmark of debt structure and market interest rates. mcgowan2018walking assume that zombies are old firms that have persistent problems meeting their interest payments, although they focus a policy discussion on the macroeconomic impact of low-productivity firms. bankofkorea2013 explicitly examines when the interest coverage ratio (ICR) is below one over three years. bankofengland2013 disregards financial management and considers firms that have both negative profits and negative value added, thus focusing on a firm's core activity.

In our view, the direction of the work so far is clear: scholars and practitioners want to infer companies' future viability from their current financial accounts. If they do not appear to be in good shape, it is likely that the company will be in trouble in the near future. From this perspective, the empirical classification of zombie firms is a perfect case study for applying machine learning techniques to firm-level data. It is a call to use in-sample information to predict an out-of-sample event.

Strengthened by the previous intuition, we propose an identification of zombies, starting from the predictions of failure made in Section (ref). We propose to classify as zombies those firms that are at the right end of the risk distribution and for which the chances of recovering from financial distress are minimal for at least three years.

In notation, we first consider the deciles along the predictions ${f}(\mathbf{X}_{i,t-1})$, where each $q_{j,t}$ is the threshold for the $j$th decile of the probability of default at time $t$:

equation[equation omitted — 174 chars of source]

where $Y_{i,t}\neq 1$ indicates that the $i$th firm did not fail yet at time $t$.

table[table omitted — 638 chars of source]

{\color{black} In Table (ref), we report a transition matrix for the firms that did not fail, based on elaborations over the entire period of analyses in 2008-2017. We observe that a significant share of firms that our predictions locate beyond the 9th decile in a representative year $t$ do not have a high chance to improve in $t+1$. In fact, most of them (46%) remain stuck in the same highest-risk category, and only 12% are able to recover and reach a more reasonable level of financial distress, i.e. below the 6th decile. Interestingly, the 9th decile is quite difficult to reach from the bottom of the distribution, as only 11%, 7%, and 3% of firms from lower deciles, respectively, are observed to transit to a situation of highest distress. Moreover, about 78% of companies stay below the 6th decile in the entire analysis period.

In general, we can say that, according to the information we have observed, a viable firm does not easily shift into financial distress, but if it does, it is difficult to recover from it. With this in mind, it makes sense to set an appropriate threshold that realistically reflects the most difficult situations in business life. We obtain this threshold by determining the cutoff that minimizes the combination of false positive and negative rates. We accomplish this task by maximizing the BACC, since it corresponds to a convex combination of the true positive and negative rates, which in turn are complementary to the false positive and negative rates. Moreover, the concept of zombie firms is closely related to the concept of false positives, since these firms are expected to fail but remain active. Therefore, focusing on the false positive and negative rates within this framework is natural. According to our results in Figure (ref), the BACC peaks at about 0.9.

The latter result provides a first justification for using the 9th decile as a cutoff. Nevertheless, we want to give firms a trial period to discount the break-even strategies of some firms in the short run, e.g., in the case of newly formed firms and start-ups. For all the above reasons, we define zombie firms those firms that persist at least three years beyond the 9th decile of the risk distribution:

equation[equation omitted — 70 chars of source]

where the risk distributions are estimated with the MA-XGBoost, as proposed in Section (ref).

figure[figure omitted — 240 chars of source]

Finally, Figure (ref) shows a transition analysis of zombie firms based on our definition from Equation (ref). We calculate the fraction of zombies that either fail, remain in zombie status, transition to a lower (but still significant) risk category, or transition to a non-distressed state in the following year. These results suggest that the most common outcome is remaining in zombie status, which applies to nearly 50% of them over multiple years. In addition, a substantial proportion of zombies transition to a lower-risk category. However, only a small fraction of them completely deviate from their current state, resulting in either failure or absence of distress.}tress.}

figure[figure omitted — 753 chars of source]

Zombies in Italy

{\color{black} In this section we give some coordinates for the phenomenon of zombies in Italy. First, we provide an overview of their evolution along the macroeconomic cycle. Then, we show how they differ from the rest of the viable firms that survive the market. Finally, we provide a comparison across geography and industry taking into account their financial performance and their ability to generate added value.}

In Figure (ref), we plot the share of zombies in our analysis period against the GDP growth rates observed over the same period. Interestingly, we find that a range between {\color{black} $1.48\%$ and $2.55\%$} of manufacturing firms are on the verge of bankruptcy. The share of zombies is higher immediately after the financial crisis in 2011 and then decreased from 2013. In fact, the presence of zombies seems to be related to the business cycle. The latter is an interesting finding for further analysis but beyond the scope of this paper. We presume it makes sense for zombie firms to be countercyclical: many firms can be pushed to the brink of bankruptcy in times of crisis, while a few financially distressed firms can find a way to recover when the recovery begins.

{\color{black} To fully capture their economic importance, we briefly show in Table (ref) what distinguishes zombies from the other firms. We find that they have lower productivity on average, i.e., 19.7% and 45% when measured as TFP and labor productivity, respectively. They are also consistently smaller on average, selling about 200% less and employing 17% less than the otherwise representative healthy firm. {\color{black} The latter results are consistent with the idea that these firms drag down aggregate productivity while hindering the redistribution of resources andrews2017confronting, mcgowan2018walking.}}ing}.}}

figure[figure omitted — 613 chars of source]
table[table omitted — 1,372 chars of source]

{\color{black} Eventually, we lay out how the segment of zombies that we find according to our working definition compares to the segments of firms that:

enumerate• have problems in making their interest payments because their Interest Coverage Ratio (ICR); • have problems in their core economic activity because they are destroying value, i.e., their value added is negative.

Interestingly, both interest coverage ratio and negative value added have occasionally been proposed as indicators of zombieness in previous literature (see Appendix (ref) for more details on these indicators and their previous use).

In Figure (ref), the rays of the radar show for NUTS 2-digit regions the proportions of zombie firms that we detect with MA-XGBoost (circles) compared to firms that have an ICR of less than one (triangles in panel a) and firms that experience negative value added (triangles in panel b). We find that there is indeed common support when firms are financially distressed because their ICR is too low, and at the same time they are also classified as zombies according to our algorithm. This is evident from a common core region in the radar graph bounded by square nodes. However, the two segments do not coincide. If we consider only the ICR, we may have zombies that are obviously not in distress but are classified as such by our algorithm. It is noticeable that this occurs mostly in southern and central Italy. Abruzzo (ITF1), Basilicata (ITF5), Calabria (ITF6), Campania (ITF3), Lazio (ITI4), Molise (ITF2), and Puglia (ITF4) host a relatively large proportion of zombies according to MA-XGBoost.\footnote{Not surprisingly, patterns of missing values are also more pronounced in southern and central Italy. Between 2008 and 2016, ICR values are missing for 81.79% of firms in central Italy and for 76.24% of firms in southern Italy on average.} In these cases, the cause of distress may lie in other aspects of business activity, as our methodology allows us to use information from a wide range of financial predictors covering different aspects of firms' economic life.

At the same time, panel (a) of Figure (ref) also identifies firms that have liquidity problems because their ICR is less than one, but our algorithm does not classify them as zombies, possibly because they are still solvent thanks to profitable economic activity that allows them to eventually meet their financial obligations and thus overcome temporary liquidity constraints shortages. This is particularly evident in North and Insular Italy, including Emilia-Romagna (ITH5), Friuli-Venezia Giulia (ITH4), Lombardy (ITC4), Piemonte (ITC1), Sardinia (ITG2), Tuscany (ITI1), Valle d'Aosta (ITC2) and Veneto (ITH3).

figure[figure omitted — 1,263 chars of source]

On the other hand, in panel (b) of Figure (ref), we show that a relatively small segment of firms destroy economic value because what they sell has less value than what they buy as inputs. They are problematic; in fact, most of them are classified as zombies. However, there is a significant portion of zombies that show positive value added, but are still struggling financially and are therefore detected as zombies by our algorithm. This scenario extends over the entire country, but peaks in some central and southern regions: Abruzzo (ITF1), Basilicata (ITF5), Calabria (ITF6), Campania (ITF3), Lazio (ITI4), Molise (ITF2) and Puglia (ITF4). The descriptive statistics in Figure (ref) confirm our intuition that only a comprehensive battery of predictors, including missing values, can provide a clear picture of the relevance of the phenomenon of zombie firms. Finally, in the Appendix (ref), we report in Figure(ref) and Table(ref) a snapshot of the proportion of zombies in different industries (2-digit digits NACE Rev. 2). As largely expected, firms that have problems with interest payments ($ICR 1$) are always a larger segment than firms that we identify as zombies. On the other hand, the share of firms with negative value added is generally smaller than the share of zombie firms.}ombie} firms.}

Conclusions

In this contribution we show how statistical learning can infer non-trivial information on a set of financial indicators, and we successfully classify firms into risk categories after training on past failures. Our preferred algorithm is {\color{black} XGBoost}, which outperforms other well-known econometric and machine learning methods, especially in the presence of missing patterns of financial accounts.

Thanks to our machine learning approach, we can also reduce prediction errors compared to traditional credit scoring tools such as the Z-scores and the Distance-to-Default. Therefore, we propose to classify as zombie firms those firms that remain in high-risk status because they are at the right end of the distribution of our predictions for at least three years, beyond the 9th decile of risk, where we find that the chances of recovery to a smaller risk of distress are minimal.

The preceding evidence suggests that identifying zombies may be crucial for financial institutions to avoid wasting credit resources on insolvent firms when the latter survive the market only thanks to some adverse selection mechanisms that arise from imperfect financial markets. However, from a more general perspective, we believe that our exercise can be useful to identify a part of an economy that is in trouble. The issue is all the more critical because recent studies discuss how zombies can hinder the growth potential of many countries by impeding the redistribution of productive resources in modern economies andrews2017confronting, banerjee2018rise, andrewspetroulakis.

Indeed, we find that Italian zombies have lower levels of total factor productivity and are mainly found in smaller firms. Interestingly, we note a possible relationship with the business cycle that would be worth investigating further in the future: the share of zombies was higher after the two financial crises in 2008 and 2011, but then declined in recent years as the recovery took hold.

It is beyond the scope of this article to make predictions for the post-pandemic scenario. However, we anticipate that the problem of separating companies that can stand on their own two feet from those that hide their inability to pay will become all the more urgent now that policymakers have finally withdrawn financial support programs. The challenge is to avoid further misallocation of resources, which would slow down the much-needed economic recovery.

\singlespacing \setlength\bibsep{3pt}

\onehalfspacing

center[center omitted — 35 chars of source]

\setcounter{table}{0} \onehalfspacing