EconBase
← Back to paper

Sustainable Investing and the Cross-Section of Returns and Maximum Drawdown

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

108,213 characters · 40 sections · 61 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Sustainable Investing and the Cross-Section of Returns and Maximum Drawdown

abstractWe use supervised learning to identify factors that predict the cross-section of returns and maximum drawdown for stocks in the US equity market. Our data run from January 1970 to December 2019 and our analysis includes ordinary least squares, penalized linear regressions, tree-based models, and neural networks. We find that the most important predictors tended to be consistent across models, and that non-linear models had better predictive power than linear models. Predictive power was higher in calm periods than in stressed periods. Environmental, social, and governance indicators marginally impacted the predictive power of non-linear models in our data, despite their negative correlation with maximum drawdown and positive correlation with returns. Upon exploring whether ESG variables are captured by some models, we find that ESG data contribute to the prediction nonetheless.

Introduction

We apply a variety of supervised learning models to forecast returns and maximum drawdown---the largest decline in cumulative return over a fixed period. As predictors, we use standard accounting ratios and factors, sector information, and environmental, social, and governance (ESG) indicators. We look specifically at whether or not ESG indicators augment the predictive power of our methods.

Motivation

The search for factors that drive the return and risk of individual stocks has been the focus of the asset pricing theory for decades. While the capital asset pricing model (CAPM), see Treynor1961, Sharpe1964, Lintner1965 and Mossin1966, is probably the first model to address this question and states that stocks' returns are explained by their market risk through the betas, empirical findings suggested that this model is incomplete. Despite decades of research and hundreds of papers, a thorough understanding of the drivers of cross-sectional variation of expected returns eludes us. The Fama and French three-factor model, see FamaF1992, and the Carhart four-factor model, see Carhart1997, are well-accepted academic models used to capture this variation. They demonstrate that size, book-to-market ratio, and momentum are significant drivers of asset returns, and they complement the CAPM. Nevertheless, new empirical studies continue to emerge, and document a growing number of factors, see FamaF2008 and HarveyLZ2015 for an overview, or GreenHZ2013 and FengGX2017 for extensive factors mining. In recent papers, GuKX2018, chen2020deep investigate which firm characteristics predict the cross-section of expected return using machine learning and deep learning techniques, and find an advantage in using non-linear models.\\ \\ Firm characteristics are also used to forecast risk. For example, Barra models, which are standard in industry, use firm characteristics to estimate stock covariance matrices, see barra2014. They decompose the covariance matrix of security returns using the estimation of sensitivities of the securities to the common factors, the covariance matrix of factors, and the variances of security-specific returns. Vozlyublennaia2013 investigate the effect of firm characteristics in the dynamics of idiosyncratic risk. They find that firm characteristics can be used in the analysis of the differences in risk across securities. HerkosvicKL2013 show that a firm's idiosyncratic volatility obeys a strong factor structure. In addition to returns, our paper is interested in an under-looked but yet very important metric to assess risk by portfolio managers; maximum drawdown.\\ \\ Maximum drawdown has also received extensive attention in the literature on financial markets. However, most of the papers are theoretical. Taylor75 elucidates properties of maximum drawdown under the assumption of a Brownian diffusion for the underlying process, MagdonIsmailAAP2004 on the other hand give a series representation of the maximum drawdown distribution and estimate its expected value. A second body of literature addresses portfolio construction and formalizes some risk measures based on maximum drawdown, see ChekhlovUZ2004, HeidornKR2009, or GoldbergM2016. \\ \\ In many cases, investors are forced to liquidate positions after a large loss. A body of literature uses ranking rules to momentum-style portfolio construction. In this concept, tail risk is considered in risk-adjusted performance measure like the Calmar ratio. rachev2007 report that tail behavior can predict future price directions of equities. They report that by using risk-adjusted performance measure, namely the Stable-Tail Adjusted Return ratio. jaehyung2021 use maximum drawdown and its consecutive recovery as a stocks selection rule and find that it outperforms other alternative momentum portfolios. Such strategies resulted in a better risk-adjusted returns. Close to this line of inquiry is DanielM2016, who look at how momentum strategies perform following drawdown, and Plastira14, who studies performance rankings of popular size, value, reversal and momentum portfolios using drawdown-based portfolio performance measures.\\ \\ This project began with the question of whether sustainability metrics predict stock performance. Sustainability metrics include a set of scores that specialized agencies assign to companies based on their environmental, social, and governance (ESG) strengths and weaknesses. Although sustainable investing was a niche topic a few years ago, it has become mainstream for many asset managers and mutual funds. According to the US SIF Foundation’s 2020 Report on US Sustainable and Impact Investing Trends, as of year-end 2019, one out of every three dollars under professional management in the United States—\$17.1 trillion—was managed according to sustainable investing strategies. It is no surprise that it has also become a key topic in financial research. SRI can be roughly defined as “an investment process that integrates social, environmental, and ethical considerations into investment decision making”, see Renneboog2008. This involves screening companies for corporate social responsibility on the basis of environmental (E), social (S), and governance (G) criteria. Whether or not ESG and other non-financial criteria have a material economic impact is the subject of an ongoing debate. \\ \\ The main question that animates the dialog around ethical investing is whether companies that “do good" also “do well." This statement suggests that companies with superior ESG performance generate higher financial performance and/or have a lower risk. In that sense, a wide range of literature explores the link between sustainability metrics and several dimensions of performance and risk. PrestonO1997 explore the possible empirical association between social and financial performance (through return on assets and other measures) in longitudinal data. Their sustainability data are based on a Fortune magazine reputation ratings for individual companies. They find that positive synergies explain social-financial performance correlations. KhanSY2016 explore the question of materiality, i.e., a classification that maps different sustainable indicators as material for different industries. Their analysis relies on KLD\footnote{Founded in 1988, KLD Research & Analytics, Inc. provides performance benchmarks, corporate accountability research, and consulting services. The company offers environmental, social, and governance research for institutional investors.} data and the “materiality” mapping provided by the Sustainability Accounting Standards Board (SASB)\footnote{The Sustainability Accounting Standards Board (SASB) was founded in 2011 to develop and disseminate sustainability accounting standards.}. Using both portfolio and firm-level regressions, they find that firms with good ratings significantly outperform firms with poor ratings. Immaterial sustainability issues, on the other hand, do not improve financial performance. Using MSCI ESG KLD STATS data between 2000 and 2016 on the US market, brogi2019environmental also find evidence of the positive impact of ESG on the return of assets (ROA), especially in the banking sector. giese2019foundations find a positive impact on companies' valuation and performance through reduced capital costs, higher valuations, higher profitability, and lower exposure to tail risk. The analysis of over 2200 individual stocks by friede2015esg highlights that about 90% of them show a non-negative relationship between ESG and corporate financial performance. halbritter2015wages show that the direction of the overperformance of ESG portfolios is strongly dependent on the rating providers.\\ \\ An opposing point of view is “doing good but not well". This view is linked to the “managerial opportunism hypothesis," which suggests that managers tend to maximize their gains and that socially responsible activity may cost resources to the firm, putting it at a disadvantage relative to firms that invest less in sustainability issues, see AupperleCH1985. In that sense, BrammerBP2006 and HongK2009 show a higher performance for portfolios with low sustainability performance compared to their peers with high sustainability performance. In a recent paper, bruno2022honey finds that most of the outperformance of ESG-based strategies can be explained by their exposure to equity-style factors that are mechanically constructed from balance sheet information. madhavan2021 separate the impact of ESG variables into a factor-related effect and an idiosyncratic effect. They found that the factor-related effect drives the alpha and active returns, but failed to reject the lack of relationship between idiosyncratic ESG components and performance. madhavan2020factor find similar results for fixed-income mutual funds. They conclude that ESG funds derive a significant portion of their performance from traditional factors, such as quality and low volatility. Other studies on the European market find similar results, see su12166387. billio2021inside analyze the ESG rating criteria used by known agencies. They find that the disagreement in the scores among agencies disperses the effect of the preferences of ESG investors and dissipates their impact on financial performances. As a result, a new body of literature studies the rating disagreement and incorporates it in the performance analysis, see gibson2021esg, avramov2022sustainable.\\ \\ Our paper, on the other hand, explores the question of how sustainability issues impact financial risk through maximum drawdown. We ask whether ESG data enhance our ability to predict future performance and risk. Previous literature explores similar questions by testing for a relationship between sustainability measures and firm systematic risk. In this line of inquiry, JoN2012 and BenlemlihSQG2018 use the CAPM beta as the measure of a firm's systematic risk. While the first of the two papers focuses on `controversial industries', both studies find a strong negative link between corporate social performance and systematic risk. Chollet2018 consider, in addition to systematic risk, specific and total risk, translated respectively by the standard deviation of residuals from the CAPM model and the volatility of stocks. Using Thomson Reuters ASSET4 to measure ESG scores, they find that a firm's good social and governance performance reduces its financial risk.

Contributions

While this research began with the assessment of ESG's impact on stock performance, the contributions of this paper go beyond that. We identify determinants of the cross-section of middle-term returns and drawdown using firms' quantitative and qualitative characteristics. We perform a pooled regression analysis using firm-level factors and nine regression algorithms in supervised statistical learning and compare the performance of these methods based on their out-of-sample performance.\\ \\ Stock returns are difficult to predict in general and particularly in the short term and incorporate a low signal-to-noise ratio but long-term investors rely on fundamentals for their stock selection. Maximum drawdown, which measures a firm's risk of successive negative returns over a period of time, has higher predictability across firms. The first analysis seeks to find if stock level characteristics explain long-term returns and maximum drawdown over a large data set that includes 2008 financials without ESG variables. We then split our data to account for ESG variables and repeat the analysis. We find that maximum drawdown gives high predictability results in both cases compared to returns. In both cases, we can obtain both high predictability and economically meaningful results.\\ \\ Our empirical analysis clarifies what features are most valuable to forecasting a company's future return and drawdown. By applying nine regression methods, we can identify a consistent set of variables that effectively forecast maximum drawdown and, to a certain extent, returns.

Empirical findings

{We rely on monthly stock data from January 1970 until December 2019 with a total of 600 months. After some data processing and focusing on companies listed on Russel 3,000, we obtain an average of 2,717 individual stocks per month. Our predictions are based on 119 lagged firm characteristics and an additional 7 ESG-related variables. Our first analysis uses only non-ESG variables for a long test period. Our second analysis is constrained by the ESG data history, so we adjust the training and test sets accordingly. Due to the co-linearity between OWL Analytics ESG scores, we group them into three sets; one where we include granular indicators, a second where we include E, S, and G, and a third with the single, fully aggregated ESG score but provide results for the E, S, G, and ESG scores only. TruValue Labs's ESG indicators are also added, both separately and in combination with OWL data. Finally, we compare the forecasts of a set of linear and non-linear supervised learning algorithms. The empirical findings of our analysis are summarized in the following.} Machine learning offers new tools to address regression problems and account for non-linear relationships in asset pricing theory. Supervised learning approaches usually aim to predict an output from several input variables. Linear regressions, either ordinary or penalized, and dimension reduction methods like principal components and partial least square regressions are limited to linear relationships. The advantage of more complex models used in machine learning like tree-based models and neural networks is that they can overcome this limitation and account for non-linear relationships without increasing the dimension of the problem (by introducing new variables). Our results show that non-linear models improve performance. However, linear models are also a match when data are primarily processed, generally offer comparable performance, and are robust when little data are available. \\ \\ The ranking of one-year returns and maximum drawdown across stocks was predictable in our out-of-sample test period. Maximum drawdown has a better prediction performance than the returns. We report as a performance metric the mean squared error (MSE) and Kendall rank correlation, and we analyze decile portfolios based on predicted returns and maximum drawdown. While MSE assesses how small the error of the prediction to the true value is, the Kendall correlation, which measures the concordance between the prediction and the realization, gives a better results interpretation. In particular, a large positive value would suggest that your strategy is not ranking stocks randomly. Kendall correlation on the test period is close to 50% for maximum drawdown and over 10% for log excess returns (logER). The cross-sectional correlation reached its lowest value during the 2008 crisis, $\sim 20\%$ for logER and, $\sim 25\%$ for the maximum drawdown. During mild market conditions, the value exceeded over 20% for logER and $55\%$ for MDD. The positive sign of the Kendall correlation suggests that even during low-performance periods, there was a positive relationship between the realized and predicted values of $MDD$ while return prediction can lead to bad investment choices.\\ \\ The set of dominant features was largely consistent across models. Based on a sensitivity analysis that ranks variables by their predictive power, when excluding ESG, the top five variables in most models were volatility-based measures (one-year, one-month, and idiosyncratic), bid-ask spread, and size. Among important variables, we also find momentum and earnings-to-price.\\ \\ Some combinations of environmental, social, and governance scores improved the predictability of returns and maximum drawdown, although marginally, when added to the top-performing non-ESG variables over the test period, January 2015 to December 2019. We find that penalized linear regression models discarded these variables but that non-linear models captured them. In fact, through exploring variable importance using the Shapley value, we find that the G score and S are among the top features in the prediction and have a favorable relationship to risk (negative correlation with MDD) and performance (positive correlation with returns). The results we find suggest that the correlation fades away when controlling for firm characteristics. \color{black}

Methodology

One of the goals of this paper is to compare the ability of different supervised learning methods to predict next-period return and maximum drawdown, or more precisely, their ranking across stocks, from an observed set of quantitative and qualitative firm attributes. We summarize the (well-known) methods used in the analysis here for completeness.\\ \\ All the methods we consider follow the general equations:

align[align omitted — 96 chars of source]

and

align[align omitted — 76 chars of source]

where $f^*$ defines a general mapping function, $y_{i t+1}$ denotes the dependent variable, i.e., log excess return or maximum drawdown, of firm $i$ over the period from time $t$ to time $t+1$, $x_{i, t} = [x_{i, t}^1, ..., x_{i, t}^J]$ is a vector of realizations of firm characteristics for a stock $i$ at time $t$, and $\varepsilon_{i, t+1}$ an error term.\\ \\ The relationship expressed in formulas ((ref)) and ((ref)) mirrors the asset pricing relationships specified in countless papers in the finance literature. In those papers, the dependent variable is usually excess return over the risk-free rate or another benchmark.

The Models

Linear models

The first set of methods we use are linear regression models, in which the mapping function $f$ is a linear combination of entries $\boldsymbol{x_{i, t}} = [x_{i, t}^1, x_{i, t}^2, ..., x_{i, t}^J]$:

align*[align* omitted — 129 chars of source]

where the $\beta$ is the parameter vector and $ \boldsymbol{x_{i, t}'}$ is the transpose of $ \boldsymbol{x_{i, t}}$ the vector of the firm characteristics for stock $i$ at time $t$.

Ordinary Least Square Regression

To find the regression parameters $\beta_0, \beta_1, ..., \beta_J$, linear regression minimizes the sum of squares errors between the observation $y_{i, t+1}$, and the prediction $f(x_{i, t})$:

align*[align* omitted — 135 chars of source]

When the number of input features is large or when the regressors are correlated, linear regression can lead to overfitting and leads to spurious coefficients.

Penalized regression techniques: Lasso, Ridge, and Elastic Net

Penalized regression models aim to create a more robust output model in the presence of a large number of potentially correlated variables. They are a simple modification of the ordinary linear regression that introduces a regularization term in the optimization problem - a technique in Machine Learning that aims to reduce the out-of-sample forecasting error by penalizing the coefficients. In this paper, we explore the three main ones: Lasso, Ridge, and Elastic Net.

itemize• The first penalized regression technique, called Lasso, for “least absolute shrinkage and selection operator”, see Tibshirani1996, is based on a penalty term equal to the absolute value of the beta coefficients: \begin{align*} \beta = \operatorname*{arg\,min}_{\beta}\frac{1}{NT}\sum_{t = 1}^{T}\sum_{i = 1}^{N_t}\big(y_{i, t+1} - f(\boldsymbol{x_{i, t}}; \beta)\big)^2 + \lambda \sum_{j = 1}^{J}\mid\beta_j\mid \end{align*} where $\lambda$ is a non-negative hyperparameter. • Ridge regression, see Hoerl1962, adds a penalty related to the square of the magnitude of the coefficients called $\ell_2$ regularization and solves the following objective function: \begin{align*} \beta = \operatorname*{arg\,min}_{\beta}\frac{1}{NT}\sum_{t = 1}^{T}\sum_{i = 1}^{N_t}\big(y_{i, t+1} - f(\boldsymbol{x_{i, t}}; \beta)\big)^2 + \lambda \sum_{j = 1}^{J}\beta^2_j \end{align*} where $\lambda$ is a non-negative hyperparameter. • Elastic Net, see ZouH2005, uses an intermediate objective function between the Lasso and Ridge: \begin{align*} \beta = \operatorname*{arg\,min}_{\beta}\frac{1}{NT}\sum_{t = 1}^{T}\sum_{i = 1}^{N_t}\big(y_{i, t+1} - f(\boldsymbol{x_{i, t}}; \beta)\big)^2 + \lambda_1 \sum_{j = 1}^{J}\mid\beta_j\mid + \lambda_1\sum_{j = 1}^{J}\beta^2_j \end{align*} where $\lambda_1$, $\lambda_2$ are two non-negative hyperparameters.

The ordinary least squares objective is obtained by setting the parameter $\lambda=0$ (resp. $\lambda_1 = \lambda_2=0$). Moreover, as $\lambda$ (resp. $\lambda_1$ and $\lambda_2$) increases, we choose a smaller set of predictors by decreasing the value of the coefficients and shrinking the least relevant ones.

Dimension Reduction

Dimension reduction techniques aim to decrease the number of features in a data set without discarding salient information. In contrast to penalized regression methods, which discard weak regressors by setting their loadings to zero, dimension reduction techniques form an uncorrelated linear combination of the predictors for the purpose of reducing noise and concentrating signal. In our analysis, we rely on two widely used methods, principal component regression, and partial least squares.\\ \\ Principal Component Regression (PCR) is a two-step procedure. The first step is a principal component analysis (PCA) that combines the independent variables into a set of leading components ranked by their explained variance. The predicted variable plays no role in this step. The second step is a simple linear regression on the leading components. PCA is one of the most widely used dimensionality reduction techniques, and it dates back to Pearson1901. The key idea is to find a new coordinate system in which the input data can be expressed with fewer variables without a significant error.\\ \\ Unlike PCR, in which the two steps are performed separately, a partial least squares (PLS) regression combines dimension reduction and regression by directly taking into account the covariance of the predictors with the target prediction. This is carried out by estimating, for each predictor $p$, a univariate return prediction coefficient via OLS. This coefficient $\varphi_p$ represents the partial sensitivity of returns to a given predictor $p$. Predictors are then averaged into a single aggregate component with weights proportional to $\varphi_p$, which would give the strongest univariate predictors the highest weights. Then the target and all predictors are orthogonalized with respect to previously constructed components, and the procedure is repeated on the orthogonalized set. The procedure stops when the desired number of components is obtained.\\ \\ More formally, we write the linear model in its vectorized version,

align*[align* omitted — 30 chars of source]

where $R$ is the $NT\times 1$ vector of $r_{i, t+1}$, $Z$ is the $NT\times P$ matrix of stacked predictors $z_{i, t}$, and $E$ is a $NT\times 1$ vector of residuals $\varepsilon_{i, t+1}$.\\ \\ The linear model given above is re-written for the set of reduced predictors:

align*[align* omitted — 42 chars of source]

where $K$ is the number of reduced predictors corresponding to a linear combination of the initial ones, $\Omega_K$ is the $P\times K$ matrix with columns $w_1, w_2, ..., w_K$, where $w_j$ for $j \in 1,..., K$ is the set of linear combination weights used to create the $j^{th}$ predictive components, and $Z\Omega_K$ is the reduced version of the original predictor set.

Regression Trees and Random Forests

Decision trees are among the simplest non-linear models that rely on a tree structure to approach the outcome. We can express the prediction of a tree with $M$ terminal nodes and depth $L$ as:

align*[align* omitted — 125 chars of source]

where each $C_m(L)$ is one of the $M$ partitions of the data.\\ \\ The algorithm that fits the decision sequentially splits subsets of the data in two on the basis of one of the predictors. The split at each step is chosen to optimize a loss function defined in terms of impurities in the child nodes. Impurity is typically measured in terms of the Gini index or entropy. To prevent overfitting and to ensure the tree is interpretable, different criteria can be used like the maximum depth of the tree or node size.\\ \\ A random forest averages the output of many decision trees. Each decision tree is fit on a small subset of training examples or constrained to use a subset of input features. Doing so increases the bias relative to a simple decision tree, but decreases the variance; see Breiman2001 for more details.\\ \\ Formally, if a regression tree has $M$ leaves and accepts a vector of size $n$ as input, then we can define a function $q: \mathbb{R}^n\rightarrow\{1, ..., T\}$ that maps an input to a leaf index. If we denote the score at a leaf by the function $w$, then we can define the $k$-th tree (within the ensemble of trees considered) as a function $f_k(x) = w_{q(x)}$, where $w \in \mathbb{R}^T$. For a training set of size $n$ with samples given by $(x_i, y_i)$, $x_i \in \mathbb{R}^m$, $y_i \in \mathbb{R}$, a tree ensemble model uses $K$ additive functions to predict the output as follows:

align*[align* omitted — 103 chars of source]

Extreme Gradient Boosting

The term `boosting' refers to the technique of iteratively combining weak learners (i.e., algorithms with weak predictive power) to form an algorithm with strong predictive power. Boosting starts with a weak learner like a regression tree algorithm, and records the error between the learner's predictions and the actual output. At each stage of the iteration, it uses the error to improve the weak learner from the previous iteration step. If the error term is calculated as a negative gradient of a loss function, the method is called `gradient boosting.' Extreme gradient boosting (or XGBoost) refers to the optimized implementation in ChenG16. \\ \\ Formally, the model uses $K$ additive functions to predict the output as follows:

align*[align* omitted — 103 chars of source]

where we take $f_k(x) = w_{q(x)}$ ($q : \mathbb{R}^m \rightarrow T, w \in \mathbb{R}$) from the space of regression tree. The function $q$ represents the structure of each tree that maps an example of the data set to the corresponding leaf index, $T$ is the number of leaves in the tree, and each $f_k$ corresponds to an independent tree structure $q$ and leaf weight $w$.\\ \\ To learn the set of functions in the model, the regularized objective is defined as:

align*[align* omitted — 102 chars of source]

where $\Omega(f) = \gamma T + \frac{1}{2}\lambda\left\lVertw\right\rVert^2$.\\ \\ The model is then optimized in an additive manner. If $\hat{y}_i^{(t)}$ is the prediction for the $i$-th training example at the $t$-th stage of boosting iteration, then we seek to augment our ensemble collection of trees by a function $f_t$ that minimizes the following objective:

align*[align* omitted — 107 chars of source]

The objective function is approximated by a second-order Taylor expansion and then optimized (see ChenG16 for details and calculation steps). To prevent overfitting, XGBoost uses shrinkage and feature sub-sampling.

Artificial Neural Networks: The Multi-Layer Perceptron

Artificial neural networks (ANN) are a set of algorithms inspired by the human brain and designed to recognize patterns. The idea behind such a framework is to represent complex non-linear functions by connecting simple processing units into a neural network, each of which computes a linear function, possibly followed by a non-linearity.\\ \\ A neuron-like processing unit is given by:

align*[align* omitted — 52 chars of source]

where the $x_j$'s are the inputs to the unit, the $w_j$'s are the weights, $b$ is the bias, $\phi$ is a non-linear activation function, and $a$ is the unit's activation. An activation function is a function used to transform a trigger level of a unit (neuron) into an output signal. Examples of such activation functions are:

itemize• The identity activation function: $\phi(x) = x$. • The logistic activation function: $\phi(x) = \frac{1}{(1 + e^{-x})}$. • The hyperbolic tan function `Tanh': $f(x) = \tanh(x)$. • The rectifier linear unit function `ReLu': $f(x) = \max(0, x)$.

The combination of a set of these units is what forms the neural network. Each unit performs a simple function, but in aggregate, the units can do more complex computations. In our analysis we apply a variant of the simple feed-forward neural network, that is, the multi-layer perceptron. In such a model, the units are arranged in an acyclic graph, and calculations are done sequentially (in contrast to a recurrent neural network, where the graph can have cycles).\\ \\ The multi-layer perceptron (MLP), as shown in Figure (ref), is composed of a set of layers, each of which contains identical units. In an MLP, the network is fully connected, i.e., every unit in one layer is connected to every unit in the next layer. The {\it input} layer takes the values of the predictors. The last {\it output} layer has a single unit in the case of regression. The intervening {\it hidden} layers are mysterious because we do not know ahead of time what their units should compute. The number of layers is known as the depth, and the number of units in a layer is known as the width. “Deep learning" refers to training neural networks with many layers.\\ \\

figure[figure omitted — 1,218 chars of source]

We denote the input units by $x_j$, the output unit by $y$, and the $\ell$-th hidden layer by $h_i^{(\ell)}$. Since MLP is fully connected, each unit receives input from all the units in the previous layer. This implies that each unit has its own bias and a weight is associated with every pair of units in consecutive layers:

align*[align* omitted — 233 chars of source]

where $\phi^{(1)}$ and $\phi^{(2)}$ are the activation functions (which can be different for different layers).

The Data

The independent variables

The independent variables in the analysis include firm characteristics taken from the CRSP/Compustat database. The construction of firm characteristics and the notation we use to refer to them are taken from Wharton Research Data Service (WRDS) with some further cleaning. Descriptions of the firm characteristics are provided in the appendix. We work with CRSP stocks, identified by their PERMNO code. We restrict the analysis to shares within the Russel 3000 to limit the impact of micro and small caps. CRSP firm characteristics are then merged with ESG data from two different data providers, OWL Analytics and TruValue Labs.\\ \\ OWL Analytics aggregate data from different providers to generate their ESG scores, which are updated on a monthly basis and cover 12 primary categories:

itemize• 3 environment scores: pollution prevention (E1), environmental transparency (E2), resource efficiency (E3). • 6 social scores: compensation & satisfaction (EMP1), diversity & rights (EMP2), education & work condition (EMP3), community & charity (CIT1), human rights(CIT2) and sustainability integration (CIT3). • 3 governance scores: board effectiveness (GOV1), management ethics (GOV2), and disclosure & accountability (GOV3).

The scores of each category are aggregated into the main three scores E (for environment), S (for social), and G (for governance), before being averaged into a single ESG metric (ESG score). Stocks in the OWL Analytics database are identified by their ISIN, and their coverage starts in 2009-07. We limit the current analysis to the pillars E, S, and G and the aggregated ESG score from Owl Analytics. \\ \\ TruValue Labs (TVL), on the other hand, relies on public data, mainly news, and filters the data through the SASB “materiality map” to assess their ESG impact. Their data coverage starts in 2007. In our analysis, we focus on four of their scores:

itemize• Trailing 12-month volume: Number of articles tagged to SASB categories during the past 12 months. • Insight: Exponentially-weighted moving average (EWMA) of pulse score as a long-term sentiment indicator. • Momentum: Slope of Insight score over past twelve months intended to identify companies with improving or deteriorating ESG.

ESG data and firm characteristics are merged to form our data set. WRDS data are usually available by the end of the month and made public with a lag. OWL Analytics ESG scores require a two months lag, while TruValue Labs are lagged by one day since they are updated daily. Those sets of predictors are then aligned with the one-year forward maximum drawdown. \\ \\ We exclude stocks and dates with a lot of missing data and winsorize variables beyond the median by 3 interquartiles to the 1% and 99% percentile value cross-sectionally. Finally, we keep only stocks for which returns are available for the subsequent twelve months (necessary to calculate one-year logER and MDD without any missing dates). The resulting data set covers 564 months (from 1970-01 until 2019-12) with a total of $14,488$ stocks and an average of $2722$ stocks per date. Figure (ref) shows the number of stocks available in the analysis for each time period, along with the number of stocks with OWL Analytics and TruValue Labs data.

figure[figure omitted — 172 chars of source]

\color{black}

The dependent variable

As reported to Benjamin Graham\footnote{Warren Buffet claims his mentor Benjamin Graham wrote it, however, the quote is not found in either of Graham's books, Security Analysis or The Intelligent Investor.}; “in the short run, the market is a voting machine, but in the long run, it is a weighing machine”. Motivated by the importance of long-horizon predictions of stocks risk and return, and probably a tendency to be predictable by stocks' fundamental data, we orient the exercise of factor models using statistical regression methods to a one-year horizon.\\ \\ Because of the desirable properties of log returns, we use log excess returns to the risk-free rate. Also, rather than using the volatility or variance as the risk measure, we focus on an under-looked, but nonetheless important risk measure, which is maximum drawdown.\\ \\ The maximum drawdown of a stock is defined as its largest cumulative loss from peak to trough over a fixed period $\tau$. Letting $P$ denote stock price, maximum drawdown can be calculated through the following formula\footnote{While maximum drawdown is a negative return, we take the form that gives a positive value/loss.}:

align*[align* omitted — 97 chars of source]

By construction, the maximum drawdown is unique for each period (and each stock), however, it depends on the observation frequency. For example, using daily observations, the calculated maximum drawdown would miss the flash crash that happened in 2010. Intraday observations on the minutes or seconds level in this case would detect a larger maximum drawdown than what daily observations do. However, for several reasons such as the availability of intraday data, and the unlikelihood that such uncommon market activity is driven by fundamentals, we stick to daily observations for maximum drawdown calculations in this paper. We rely on a one-year forward maximum drawdown using the same moving window we use to calculate log excess returns. We display in Figure (ref) two representations of MDD: its evolution in time for two stocks and its cross-sectional distribution for two periods.

figure[figure omitted — 467 chars of source]

Training, validation, and test data

In financial economics, it is standard practice to split a data set into in-sample and out-of-sample subsets and to evaluate performance based on the latter. For machine learning models, it is more common to split the data into three subsets: training, validation, and test. The training set is used to fit a number of models with the same general architecture but different hyperparameters.\footnote{In machine learning, a hyperparameter is a parameter whose value is set before the learning process begins, and is specific to the model's architecture. Examples of hyperparameters are the penalization term for a penalized regression and the number of hidden layers and number of units in a neural network. Hyperparameters are specified through a validation process and are chosen to minimize model error over the validation set.} The validation set is used to tune the hyperparameters. The test set is used to evaluate model performance. This more elaborate procedure addresses the complexity of machine learning models, and their tendency to overfit or underfit when hyperparameters are poorly chosen\footnote{$K$-fold cross-validation is a standard method for tuning hyperparameters. The method splits training data into $k$ equally sized subsets. A model with a fixed set of hyperparameters is fit on the union of $k-1$ subsets and evaluated on the $k$th subset. This process is repeated $k-1$ times and the evaluation scores are averaged, yielding an overall score. This practice is, however, very costly and does not respect the time ordering of events, which is essential to our analysis.}.\\ \\ The avoidance of look-ahead bias, the goal of subjecting our models to the most challenging tests possible, and the availability of data guide the specification of our training, validation, and test sets. To select the parameters of the models, we omit ESG variables and split the data into a training set that runs from 1970-01 until 1990-12 and a validation set that runs from 1991-12 until 2001-01. Once the best model parameters selected, we then retrain the model and test it. First we do so without incorporating ESG variables due to their short history. In this specific case, our training data runs between 1970-01 and 2000-01-01.\\ \\ The short history of ESG data (which begin on 2009-07 for OWL and 2007-01 for TVL) forces us to make two compromises). First, to limit the effect of missing data we dismiss data points without any ESG information. We also opt for a test set of one year at a time and increase the training set (fixed starting data and moving ending date with retraining at each beginning of the year). Therefore, we train the models 5 times with a starting date on 2009-01 and an end date from 2014-01 to 2018-01 with one year increment. Each trained model will be used to predict one year length test set between 2015 and 2019\footnote{Since we evaluate maximum drawdown over one-year periods and stagger start dates by a month, a period of $N$ years contains $12\times (N-1)$ observations.}.

Preprocessing

Data preprocessing is the technique of transforming raw data into a machine-understandable format. Some machine learning algorithms are particularly sensitive to ranges of input values and stationarity of the variables, and rescaling is a common way to overcome this issue. However, there is no consensus on a rescaling method to apply to these problems although it is very common to see z-scoring and uniform quintile scaling which mimics common portfolio sorting techniques (see GuKX2018). After comparing different methods on the validation set, we settled for z-scoring as it maintains the distribution shape of the variables in contrast to uniform quantile scaling. It is important however to stress that different scaling methods may lead to different results and on our observations we found that some models might perform better using other scaling methods. The mathematical formula corresponding to such transformation is as follows:

align*[align* omitted — 84 chars of source]

where $\bar{x}_{., t}$ and $\sigma(x_{., t}$ are the cross-sectional mean and standard deviation of the observations $(x_{1, t}, x_{N_t, t})$. Note that we proceed with the rescaling after winsorizing extreme values (beyond 3 inter-quartiles of the median since variables are not normally distributed). When omitting missing observations, the transformation standardizes all variables to have a mean 0 and unit variance. For the predicted variables, log excess returns and maximum drawdown, we would leave the variables unchanged.\\ \\ Before running the models, we need to deal with missing data. This issue is pervasive and important in machine learning applications. Removing stocks with missing information is not a solution as it will result in very few observations. Therefore, we rely on imputation methods. The most common imputation methods are mean, median, mode, and K-nearest neighbors. In our case, replacing missing values for categorical variables by their median and those in numerical variables by their mean gave better results on the validation set. Once again, the median and mean are taken over stocks on a given date.

Performance and feature importance

Performance measures aim to evaluate a statistical model and are usually reported on out-of-sample (OOS) data, i.e., on data that are not used to calibrate the model. In the best-case scenario, a model's prediction should be as close as possible to the real outcome. A measure of the error between the forecast and the actual value can be used to interpret the results. Since regression models typically minimize the mean squared error, we display it as one of the performance measures. However, as an investor, the value is not necessarily informative and a mean squared error could be small but the trading strategy resulting from the prediction could still be bad. Therefore, we display the Kendall rank correlation, which measures the ordinal association between two quantities, as a second performance measure as it has better information on the performance of decile ranked portfolios\footnote{The coefficient of determination, or R-squared, which measures the proportion of the variance in the dependent variable that is predictable from the independent variable, is commonly used for regression analyses. Our reason for choosing an alternative measure is our focus on the ranking of stocks by their level of risk.}. \\ \\ We recall the formulas for the mean squared error (MSE) for our pooled data:

align*[align* omitted — 108 chars of source]

where $NT=\sum_{t\in \mathcal{T}}N_t$ denotes the number of observations in the test set. For the MSE, the model is better when the MSE value is smaller.\\ \\ The Kendall rank correlation ($\tau$) for a sample of size $N$ is given by the following formula:

align*[align* omitted — 109 chars of source]

We apply the formula for each prediction date $t$:

align*[align* omitted — 146 chars of source]

where $N_t$ is the number of stocks available at $t$.\\ \\ To calculate the performance measure over the entire test period, we average the time series of Kendall $\tau$ over the test set $\mathcal{T}$:

align*[align* omitted — 63 chars of source]

In theory, the overall Kendall correlation ranges from -1 to 1. A value of 1 means that the predicted MDD ranks the stocks exactly as realized MDD does. A correlation close to 0 can be interpreted as no relationship between the prediction and realization, while a negative correlation means a discordance between the prediction and realized MDD rankings.\\ \\ Note that any other correlation measure (e.g., Spearman or Pearson) works as well. The choice of the Kendall rank correlation is mainly due to its simplicity to interpret. \\ \color{black} The idea of feature importance is to measure how much a performance metric decreases when a feature is not available. A potential way to measure feature importance would be to remove the feature from the data set, re-train the model with the other estimators, and measure the change in performance. On the one hand, this solution requires re-training and can be computationally intensive. On the other hand, it shows what feature is important within the data set rather than within a given trained model\footnote{Unlike simple models like linear regressions, where the “beta” coefficients and t-statistic give information on the importance of a regressor, more complex models, like neural networks or random forest, don't have such parameters and thus require to resort to more advanced feature importance analyses.}.\\ \\ One goal of the paper is to find characteristics that enable one to predict future stock risk and performance to a certain extent. While asset pricing theory relies on linear regression models and their coefficients. Our multiple models, some of which are nonparametric, require a unified framework. Luckily, a large range of literature on machine learning tackles this question. RibeiroTG2016 propose the “local interpretable model-agnostic explanations". Their method relies on local surrogate models that are trained to approximate the predictions of the underlying black box model. A surrogate model can be any interpretable model such as linear regression or a decision tree. WeiLS2015 review and compare a large set of methods for variance importance models (VIM) including difference-based VIMs, hypothesis test techniques, and variance-based VIMs. The intuition behind these methods is that the more a model's decision criteria depend on a feature, the more the predictions change as a function of perturbing this feature. Breiman2001 and AltmannTSL2010 suggest a permutation importance method, which measures how a score decreases when the empirical values of a feature are replaced by random noise drawn from the distribution of feature values. In their analysis of the predictability of monthly returns, GuKX2018 measures the importance of a variable by the reduction in $R^2$ obtained by setting all the values of the selected variable to zero and keeping the values of the other variables fixed. \\ \\ In choosing the best method for the application, one needs to consider a trade-off between efficacy and application. Permutation methods can be time-consuming for large data sets and like other methods, ignore the correlation between the variables for certain models. Since we are not interested in inference in this paper, we opt for the same approach as GuKX2018 by setting one column to zero and measuring the change in the Kendall correlation when doing so.\\ \\ For the second part of the analysis where ESG data are incorporated, we use a more novel approach in addition to the feature importance from the first part. The new approach is called Shapley additive explanations (SHAP)LundbergL17. SHAP stems from game theory and assigns importance to variables based on their contribution to improving the model from different collaborative combinations. The complexity of Shapley value is NP-hard which makes applying it to all the models very time-consuming. There are approximations that are polynomial in time for tree-based models. Therefore, we focus on XGBoost to conduct the analysis using Shapley values.

Empirical Study

We collect monthly data from WRDS, OWL Analytics, and TruValue Labs for stocks listed in the Russel 3000 index. Our data set starts in January 1970 and ends in December 2019. We account for 119 non-ESG characteristics in total and 20 ESG-related variables, 16 from OWL Analytics and 4 from TruValue Labs.\\ \\ The regression is performed by considering the log excess returns and maximum drawdown (without any transformation as opposed to the explanatory variables) as the dependent variable and the matrix of rescaled firm characteristics as the predictors. The models trained along with their parameters are given in Table (ref) and (ref). We compare the results based on out-of-sample performance

A first analysis without ESG data

Because ESG data begin in 2007, we omit ESG scores for a first analysis in which we test the models using the other 111 variables. This enables us to carry out the analysis on a larger test set that includes the 2008 crisis. The training set starts in January 1970 and ends in December 2000, while the test set starts in January 2001 and ends in December 2019.

The cross-section of returns

The models were trained after tuning the hyperparameters given in Table (ref).

table[table omitted — 1,202 chars of source]

Performance

In Figure (ref), we report the MSE and average Kendall correlation for all the stocks, and the largest and smallest 1000 stocks by market value. XGBoost achieved the best correlation performance ($9.56\%$), while MLP achieved the smallest MSE (0.244). RF and PCR are slightly below other models. Overall, the values of the performance measures are very close from one model to another.

figure[figure omitted — 1,959 chars of source]

The previous results are aggregated over the entire test period, and they do not inform us about when the models performed poorly and when they gave a sound estimation. Also, because the MSE is less informative regarding the impact on investment strategies, we focus on the cross-sectional correlation for each date of the test period, leading to a time series of out-of-sample Kendall correlation. The results are shown in Figure (ref). All models performed poorly during 2003 and the 2008 financial crisis without exception. The cross-sectional Kendall correlation reached its minimum (around $-20\%$) in both periods, which suggests that the models would have failed to provide a useful signal for investors. To inspect this claim further, we divide our stocks for each date into five quintiles based on their predicted maximum drawdown. We then plot the empirical distributions of the stocks' realized drawdown.

figure[figure omitted — 205 chars of source]
figure[figure omitted — 518 chars of source]

Figure (ref) displays the relationship between predicted and realized log excess returns over two year-long periods. For each period, stocks are sorted into five quintiles on the basis of predicted log excess returns, and the distributions of realized log excess returns are plotted. The first period (between 2008-01-01 and 2008-12-31) includes the financial crisis, while calmer market conditions prevailed during the second period (between 2016-01-01 and 2015-12-31). During the calmer period, the densities of different quintiles are more distinct. During the crisis, the bellies of the distributions tended to overlap with an average that decreases as the quintiles increase suggesting.

\color{black}

Feature importance

To explore the predictive power of different firm characteristics, we calculate a feature importance score for all the variables over the test set. We display variables that don't decrease the performance which accounts for 72. Reported in Figure (ref), these results show that, volatility-based and liquidity variables, i.e., idiosyncratic volatility (ivol), short-interest (sio), operating profitability (oprof), book-to-market (btm) and beta. Momentum (momentum36) and size (mve) are also among the important features. Overall, the models agree on several variables and disagree on others. We find without surprise that volatility and idiosyncratic volatility outclass the other variables.

figure[figure omitted — 452 chars of source]

Performance of decile portfolios

We investigate the performance of quintile portfolios based on predicted log excess returns with annual rebalancing. We form portfolios based on deciles of the predicted quantity each first of June of each year (06-01)\footnote{We adopt this convention to be consistent with common practice relying on new companies information being published around that period, and also to mitigate market conditions around the new year.}. The portfolio is held for one year and rebalanced on the next first of June based on new stock deciles. We display in Figure (ref) the cumulative returns of the equally-weighted portfolios and report in Table (ref) some statistics for these portfolios. We report the results for XGBoost but other models exhibit similar results\footnote{We also performed the analysis on market value-weighted portfolios, and the results were consistent.}. Sharpe and Calmar ratios are mostly increasing as a function of the quintiles (as the predicted logER increases) which is desirable for a portfolio manager.

figure[figure omitted — 511 chars of source]
table[table omitted — 1,189 chars of source]

Figure (ref) shows how the high logER quintile (decile 10) outperformed other deciles. We also note that the low logER quintile (decile 1) is the one with the lowest Sharpe and Calmar ratios.

The cross-section of maximum drawdown

In the following section, we repeat the same set of analyses to maximum drawdown using the same periods, models, and performance measures. The models were trained after tuning the hyperparameters given in Table (ref).

figure[figure omitted — 1,964 chars of source]
table[table omitted — 1,211 chars of source]

Performance

Figure (ref) reports the MSE and average Kendall correlation. The best performance was achieved by XGBoost ($47.71\%$ correlation and 0.028 MSE), followed by linear regression (OLS) ($47.14\%$ and 0.028 MSE). Random forest and principal component regression are again slightly underperforming the other models ($44.41\%$ and $45.26\%$ respectively). Overall, the values of the performance measures are very close from one model to another. We observe again the time series of Kendall correlation for the period 2001-2019. Those results are shown in Figure (ref). Without exception, all models performed poorly during the 2008 financial crisis between 2008 and mid-2009. The cross-sectional Kendall correlation reached its minimum but remained positive (slightly above $25\%$). This value indicates that, despite the poor performance, the concordance between the prediction and realized maximum drawdown is persistent throughout the period.

figure[figure omitted — 201 chars of source]
figure[figure omitted — 513 chars of source]

Figure (ref) displays the relationship between predicted and realized maximum drawdown over two year-long periods. For each period, stocks are sorted into five quintiles on the basis of predicted maximum drawdown, and the distributions of realized maximum drawdown are plotted. The first period (between 2008-01-01 and 2008-12-31) includes the financial crisis, while calmer market conditions prevailed during the second period (between 2016-01-01 and 2016-12-31). During the calmer period, the densities of different quintiles are more distinct. During the crisis, the middle quantiles tended to overlap while extreme quintiles remained well distinguished.

Feature importance

We explore again the power of different firm characteristics and their contribution to the performance of the models out-of-sample. We show the ranking of the predictors that contributed to the performance of the models on average. Reported in Figure (ref), these results show that, volatility-based and liquidity variables, i.e., idiosyncratic volatility (ivol), realized volatility (vol_y), and beta are the top predictors. We also find among those variables short interest and option-based variables (sio, smirk). Overall, the models exhibit the same predominant features but also have differences among other weaker predictors. We find without surprise that volatility and idiosyncratic volatility outclass the other variables.

figure[figure omitted — 460 chars of source]

Performance of decile portfolios

We investigate the performance of decile portfolios based on predicted MDD with annual rebalancing. We display in Figure (ref) the cumulative returns of the equally-weighted portfolios and report in Table (ref) some statistics of these portfolios with XGBoost as a prediction model. The results are consistent with other models\footnote{Market-value-weighted portfolios exhibit the same pattern as equally weighted portfolios}. Sharpe and Calmar ratios are mostly decreasing as a function of the quintiles which is desirable for a portfolio manager.

figure[figure omitted — 523 chars of source]
table[table omitted — 1,164 chars of source]

Figure (ref) shows how the low MDD decile (decile 1) outperformed other deciles and that it has a lower risk. This result is in line with feature importance, which suggests that volatility is the main factor and that a low MDD portfolio is more likely to be a low volatility portfolio, and vice versa. When we look at Table (ref), we can see that the returns of equally-weighted portfolios were higher for lower MDD deciles and that the volatility, maximum drawdown, and Sharpe and Calmar ratios were better for low MDD deciles\footnote{Qualitatively similar results are observed for market-cap-weighted portfolios.}.

The impact of ESG scores

Having investigated the ability of non-ESG firm characteristics to predict the ranking of returns and maximum drawdown on a large test set, we repeat the analysis, adding ESG to the mix. Our goal is to determine whether indicators of a company's sustainability improve our ability to predict returns or maximum drawdown. Since the ESG data set begins only in 2009 for OWL Analytics and 2007 for TVL, we are forced to start our training set later and extend it beyond the financial crisis, providing the models with enough observations with ESG scores. As such, we drop observations without any ESG score and update the models each year to account for more recent data. Using the most recent trained model, we predict one year of out-of-sample data. We also keep variables that improved the predictive models from the previous section which sums up to 72 covariates for log excess returns, and 67 variables for the maximum drawdown. We made sure this validation process uses data before our out-of-sample to avoid look-ahead bias.\\ \\ We compare the results based on the out-of-sample MSE and Kendall correlation as before. For this case, we compare the base case, which uses only non-ESG characteristics, to cases where we add ESG variables. We separate the addition of 4 ESG variables from OWL into two subcases\footnote{This is because the ESG score is constructed as a weighted average of the three pillars E, S, and G scores, it is .} and consider 3 TVL scores. Notations and definitions of these cases are as follows:

itemize• Base: Top firm characteristics except for ESG scores. • E/S/G: E, S, and G scores from OWL Analytics. • ESG: Single OWL ESG score. • TVL: volume, insight, and momentum from TruValue Labs data.

Using the firm characteristics of the base case and adding different combinations of OWL ESG variables and TVL scores results in six cases of interest. We train these six cases with all the models again but retrain after each year of prediction to increase the number of observations and use more recent market conditions.

The case of returns

Performance

Table (ref) reports the MSE and Kendall correlation for the six different combinations. As indicated above, these results are based on an out-of-sample period between 2015-01 and 2019-12. Without detailing the results, which are reported in the table, we point out the three following insights:

itemize• Due to the short training period (at least at the beginning of the period), machine learning methods such as random forest, XGBoost, or MLP underperform linear models at times. • Linear models don't seem to account for ESG variables as seen by the unchanged performance measures. • XGBoost performed slightly better than linear models and MLP, but this outperformance is very small ($\sim$1%). • ESG variables have a mixed effect on non-linear models (XGBoost and MLP), sometimes improving the prediction and sometimes causing it to deteriorate.
table[table omitted — 2,167 chars of source]

To explore any potential improvement over time of ESG data, we report the time series of Kendall correlation and compare the base model with the enhanced ESG models between 2015-01 and 2019-12. The time series of the performance metric for XGBoost, which seemed to be the model that capture ESG data the most, are shown in Figure (ref).

figure[figure omitted — 285 chars of source]

The change in the performance when including ESG variables varies around zero and it is therefore difficult to associate ESG variables with an improvement in the prediction. \color{black}

Feature importance

We repeat the feature importance analysis for “base + E/S/G + TVL” for all nine models. Figure (ref) reports the ranked variables by the average importance score across the nine models. We find again the same top 3 variables, namely, idiosyncratic volatility, realized volatility, and short interest. Among ESG scores, we find TVL's insight score in the top 20, mainly captured by XGBoost. The G score is ranked 23rd among the 72 variables, and the S score is 25th.

figure[figure omitted — 498 chars of source]

Portfolio performance

We investigate the performance of decile portfolios based on predicted logER, when ESG variables are included in the set of predictors. We display in Figure (ref) the cumulative return of equally-weighted decile portfolios, and we report in Table (ref) some statistics for those portfolios. \\

figure[figure omitted — 590 chars of source]

The inclusion of ESG variables made very little difference in the performance of decile portfolios (see Figure (ref)). This is consistent with the results shown above that ESG variables did not significantly improve the prediction of the logER ranking. Table (ref), which reports the average returns, volatility, maximum drawdown, and Sharpe and Calmar ratios, gives an extensive comparison of different cases for one illustrative model. We see a slight difference in decile portfolios. Moreover, including the ESG score seems to increase the gap between the performance of decile 1 and decile 10 which supports the finding that “base + ESG” performs better than “base” for the XGBoost model.

table[table omitted — 5,420 chars of source]

The case of maximum drawdown

Performance

Using the same out-of-sample period between 2015-01 and 2019-12, we report in Table (ref) the MSE and Kendall correlation for the six different combinations and nine different models. We point out the following insights:

itemize• Results for the new test sample period (2015-2019) were better than for the previous test period (2001 to 2019). This is not surprising because of the absence of the 9 crisis and the shorter time period. • XGBoost performed slightly better than other models when looking at the correlation but this outperformance is very small ($\sim$2-4% relative) but underperformed in terms of MSE. Again, this result is not surprising given that the training set is relatively small. • ESG variables marginally improved the results for non-linear models (<2% relative). One explanation for that could be that the information for ranking the stocks by their MDD is already explained by the other variables, especially volatility-based characteristics, or that they simply have a very low predictive power of MDD.

Although these results show an absence of improvement of maximum drawdown prediction using ESG variables, we emphasize that this result is bound to the set of characteristics used along with ESG data as well as the universe of stocks used. It is important to also stress that these variables are quite recent and missing for multiple data points.

table[table omitted — 2,360 chars of source]

We explore how the models performed on different dates through the time series of cross-sectional Kendall correlation between 2015-01 and 2019-12 (see Figure (ref)). We find that including ESG variables led to a significant improvement in some dates (almost 10% relative). The difference in performance between the base case and “base + E/S/G + TVL” is mostly positive throughout the period.

figure[figure omitted — 276 chars of source]

\color{black}

Feature importance

We compute feature importance for the case “base + E/S/G + TVL” in order to assess the predictive power of ESG variables for all nine models. Figure (ref) reports the variables ranked by the average importance scores. The leading dominant variables across the models were still the same, i.e., one-year volatility (vol_y), idiosyncratic volatility (ivol) and beta. The G score comes 30th out of a total of 74, 47 of which don't decrease the performance.

figure[figure omitted — 506 chars of source]

Portfolio performance

We finally investigate the performance of decile portfolios based on predicted MDD, where ESG variables are included in the prediction. We display in Figure (ref) the cumulative return to equally-weighted quintile portfolio, and we report in Table (ref), some statistics for both equally-weighted and market-cap-weighted quintile portfolios. \\

figure[figure omitted — 585 chars of source]

\newline The inclusion of ESG variables made visually no difference in the performance of decile portfolios (see Figure (ref)). This is consistent with the results shown above that ESG variables did not significantly improve the prediction of the MDD ranking. This is also consistent with the results in Table (ref) which report the average returns, volatility, maximum drawdown, and Sharpe and Calmar ratios. There is very little change in the decile portfolio performance when including ESG variables. \\

table[table omitted — 5,422 chars of source]

ESG scores contribution to the prediction

We conclude our analysis by exploring the contribution of ESG variables to the prediction. For this purpose, we use a novel approach to variables' importance, namely the Shapley Additive Explanation (SHAP) (see LundbergL17). SHAP relies on Shapley value which is a solution concept in cooperative game theory introduced by shapley1951notes. In the context of the model's prediction, explanatory variables are the players in this cooperative game, and the model $f$ plays the role of the coalition whose payoff is the model's prediction.\\ \\ Let us consider a permutation $P$ of the set of indices $\{1,2,\ldots, p\}$ corresponding to an ordering of $p$ explanatory variables included in the model $f$. Denote by $\pi(P, j)$ the set of the indices of the variables that are positioned in $P$ before the $j$-th variable. Note that, if the $j$-th variable is placed as the first, then $\pi(J,j)=\emptyset$. Consider the model's prediction $f(\underbar{x}_{*})$ for a particular instance of interest $\underbar{x}_{*}$. The Shapley value is defined as follows:

align*[align* omitted — 104 chars of source]

where the sum is taken over all possible permutations $p!$ (ordering of explanatory variables). The variable-importance measure $\Delta^{j|J}$ for the j-th is the change between the expected prediction, when setting the values of the explanatory variables with indices from the set $J\cup \{j\}$ equal to their values in $\underbar{x}_{*}$, and the expected prediction conditional on setting the values of the explanatory variables with indices from the set $J$ equal to their values in $\underbar{x}_{*}$, i.e. $\Delta^{j|J}=\mathbb{E}[f(\underline{X})|X^{j_1}=\underline{x}_*^{j_1},\ldots, X^{j_K}=\underline{x}_*^{j_K},\ldots, X^{j}=\underbar{x}^{l}_{*}] - \mathbb{E}[f(\underline{X})|X^{j_1}=\underline{x}_*^{j_1},\ldots, X^{j_K}=\underline{x}_*^{j_K}]$.\\ \\ Calculating Shapley value is NP-hard and the time increases exponentially with the number of explanatory variables, however, LundbergL17 provides an efficient implementation of computations of Shapley values for tree-based models which we will rely on for our application. In our analysis we apply SHAP to the XGBoost model for the reason above, but also because XGBoost is the nonlinear model that seemed to capture the effect of ESG variables through our analysis.

figure[figure omitted — 466 chars of source]

We can see from Figure (ref) that TVL's insight score is ranked among the top 20 variables with a positive impact on return, the S score is ranked 40th and the G score is 51st. All these scores are colored in red which suggests a positive impact on returns. For maximum drawdown, the S score is ranked 12th, the G score 15th, and the E score 28th. S and G score have a negative impact on maximum drawdown (negatively correlated with the prediction), however, the E score has a positive correlation with maximum drawdown prediction which suggest that companies with a high E score tend to have a high maximum drawdown. Upon verifying the surprising result, the E score exhibits a small negative correlation with maximum drawdown which explains the feature importance result. As reported by literature, ESG data are divergent among data providers (see Dimson75). This complicates such analysis seeking to find the relationship between ESG data and stock risk and performance.

Conclusion

This paper explores the cross-section of returns and maximum drawdown in the US equity market. We apply supervised learning methods to forecast the one-year forward returns and maximum drawdown using current information on firm-level characteristics. We find a cross-sectional correlation between the predicted and realized target ranges between 25% and 60% with an average of $\sim 45\%$ for the maximum drawdown. For excess returns, the correlation goes between -25% and 40% for an average of $\sim 9\%$. This resulted in a good performance of the best decile portfolios and a weak performance of the worse ones in most of the periods of the backtests.\\ \\ In terms of model comparison, the results indicate that non-linear methods slightly outperform linear methods. Some firm characteristics, like idiosyncratic volatility, beta, option volatility, or short interest, can explain the cross-section of both returns and maximum drawdowns. Some factors are more important to returns, like operational profitability or book-to-market. In contrast, others are more relevant to maximum drawdowns, like price-to-sales or profit-margin-to-price. A number of these characteristics are consistently dominant across models, namely one-year volatility and idiosyncratic volatility, beta, and short interest. Finally, the agreement between predicted and realized maximum drawdown persists even in periods of turmoil.\\ \\ The addition of environmental, social, and governance scores to the set of predictors failed to improve our models' performance significantly. We identified two variables, the Governance score (g) and the Article Volume (articlevolumettm), captured mainly by the MLP model. These two scores were notably higher in low maximum drawdown/high logER quintile portfolios, consistent with the correlation between the target and those ESG variables. In terms of prediction, ESG variables failed to improve the models in a significant way in our data. This is consistent with other empirical analyses. We might not expect to find significant contributions by ESG variables due to the disparate methodologies by different ESG data providers, and the high correlation between ESG variables and some firm characteristics like size or profitability. To explore the question further, we conducted a Shapley value analysis to measure ESG scores' contribution in cooperative game theory. We found that the contribution of ESG scores, particularly the Social score, was not insignificant on average. We also found that higher scores are associated with higher returns and lower maximum drawdown.\\ \\ Our empirical conclusions about the association between ESG indicators and returns or maximum drawdown must be framed in terms of the data set. Currently, ESG indicators have relatively short data histories, and the growth and evolution of markets and ESG may provide a new perspective in the near term.