EconBase
← Back to paper

Robust Knockoffs for Controlling False Discoveries With an Application to Bond Recovery Rates

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

72,591 characters · 11 sections · 57 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Robust Knockoffs for Controlling False Discoveries With an Application to Bond Recovery Rates

\thispagestyle{empty}

abstractWe address challenges in variable selection with highly correlated data that are frequently present in finance, economics, but also in complex natural systems as e.g. weather. We develop a robustified version of the knockoff framework, which addresses challenges with high dependence among possibly many influencing factors and strong time correlation. In particular, the repeated subsampling strategy tackles the variability of the knockoffs and the dependency of factors. Simultaneously, we also control the proportion of false discoveries over a grid of all possible values, which mitigates variability of selected factors from ad-hoc choices of a specific false discovery level. In the application for corporate bond recovery rates, we identify new important groups of relevant factors on top of the known standard drivers. But we also show that out-of-sample, the resulting sparse model has similar predictive power to state-of-the-art machine learning models that use the entire set of predictors. Keywords: Knockoffs, Weighted FDR Control, Recovery Rates, Group-Variable Selection. JEL Classification: G33, G17, C52, C58

\footnotetext[3]{Corresponding author: Konstantin G\"{o}rgen; email: [email removed]; phone: +49721-608-43793}

Introduction

In large-scale financial and economic but also complex natural systems as e.g. weather, with many potentially influencing variables for a target quantity, there is a key interest in detecting the relevant driving factors in a data-driven way. Such a fully data-adaptive choice of factors yields transparency in the identification of important channels avoiding biases from insufficient pre-specification but also inspiring and complementing future model building.

We contribute to this literature by proposing a new knockoff-type methodology building on e.g. Candes2018 that offers control over the rate of falsely selected variables. With the false discovery rate (FDR) as the key hyperparameter, the selection and prediction performance based on the proposed technology is dominated by the FDR. It is thus transparent and interpretable while being data-driven since the FDR can be directly estimated as the empirical proportion of false discoveries (FDP). Note that this is in contrast to e.g. LASSO-type approaches Tibshirani1996, where penalty parameters can be chosen adaptively but have no stand-alone interpretation and meaning, which often creates a black-box connotation. Moreover, our technique gains robustness from simultaneously taking several nominal FDR-levels into account. In this way, we mitigate hyperparameter pre-selection effects and obtain robustness of results in the presence of time-dependent data. Both points are key for valid selection results in practice. We show that the proposed methodology provides interesting insights in detecting novel relevant factors for corporate bond recovery rates which might be important from a business but also regulatory perspective. In particular, we study the recovery rates of 2,079 U.S. corporate bonds that defaulted between 2001 and 2016 depending on industry and stock specific information from Bloomberg Financial Markets and 144 macroeconomic market variables from the Federal Reserve Economic Data (FRED). For this, we also document superior out-of sample performance of the resulting sparse model using only relevant factors comparing them to state-of-the-art machine learning models on the entire and the selected set of predictors and to LASSO-type specifications. We confirm our point-wise ranking with results from model confidence sets Hansen2011.

In particular, the proposed robustification technique works for the entire set of different knockoff baseline procedures from model-$X$ Candes2018 over deep knockoffs Romano2019 to group versions Dai2016 and mitigates the influence of hyperparameter input levels and data dependence challenges. We address the hyperparameter influence problem by proposing several weighted aggregation schemes for variable selection rates of different FDR-levels. By considering different weighting schemes, we account for vanishing scrutiny of the procedures in size of FDR-levels but extract the information from each level for the overall selection result. Secondly, we use a repeated subsampling scheme to control for the variability of the knockoff procedures, which themselves are random. While this shares similarities with Ren2021, we employ subsampling Meinshausen2010, which provides robustness in the presence of correlated observations and high outliers. This is of key importance for the determination of relevant factors of the considered corporate recovery rates. In this empirical study, we additionally employ principal components analysis (PCA) on groups of macroeconomic variables of similar type to reduce cross-sectional correlation of the knockoff input factors while retaining interpretability on the group level. We investigate the performance of the proposed methodology for different baseline procedures and show that an ensemble yields superior out-of-sample prediction results.

Our paper belongs to a recent quickly emerging literature that provides extensions and robustifications of the original model-$X$ knockoff framework Candes2018, for example ultra-high-dimensional approaches employing pre-screening steps liu2022model,PAN2022107504 or resampling approaches focusing on reducing variability of the knockoffs Ren2021. Other contributions diverge from the model-$X$ framework to provide better knockoffs for heavy-tailed distributions bates2021metropolized, or by changing the process of generating the knockoffs, either via neural networks Romano2019 or by using different minimization functions spector2022powerful. Recently, a rapidly growing literature has also applied ML-based approaches in financial applications in particular for empirical asset pricing Chen_2019, Freyberger2020, Kelly2019, where variable selection methods serve as a dimension reduction device and are mostly based on the lasso or other penalization frameworks with respective sparsity assumptions kozak2018,CHINCO_2019,FENG_2020,Freyberger2020. There has also been considerable research on recovery rates and loss given default (LGD), see e.g. Qi2011, Leow2012, Kaposty2020, Kellner2022 for LGD and Nazemi2018,Nazemi2021, and Bellotti2021 for recovery rates.

The rest of the paper is structured as follows. Section (ref) introduces our proposed methodology in a general setting, while Section (ref) focuses on the application. Therein, Section (ref) introduces the data set of corporate bond recovery rates and Section (ref) presents the main results of our analysis. Finally, we conclude in Section (ref).

Model Selection with Knockoffs

Setup

We propose a novel robustified knockoff-type procedure Candes2018 that offers direct control of the false discovery rate (FDR) while being robust to correlations in the data. The key hyperparameter FDR corresponds to the number of coefficients determined as non-zero while being truly zero relative to all obtained non-zero factors. FDR is well-known from the multiple testing literature as the type II error but has also been shown to directly link to the size of error rates in model estimation and prediction of after pre-selection with knockoffs Barber2019. Thus setting the acceptable (nominal) FDR-level for knockoffs directly controls estimation and prediction performance of the resulting model, while the performance of other model selection techniques depends on parameters that lack a direct interpretation. Interpretable hyperparameters, however, are key for adequate tuning and eventual transparency of the results. Moreover, the suggested knockoffs work irrespective of the type of underlying sparsity and in particular for high-dimensional cases, results are not dependent on a specific form of sparsity. We contribute a robust knockoff version which works in particular in the presence of strong cross-sectional and time dependence. This is key in many economic and financial applications and of peculiar importance for our application of the determination of relevant factors of corporate recovery rates.

We work in the following setting, where $y_i \in \mathbb{R}$ and $X_i\in \mathbb{R}^{p}$ are observed for $i=1, \dots, n$ but only some unknown subset of the $p$ components in $X$ is relevant for $y$ and forms the so-called active set $\mathcal{S}$ of $X$, i.e. for $j\notin \mathcal{S}$, $y$ is independent of component $X_{(j)}$ conditional on $(X_{(k)})_{k\in \mathcal{S}}$. These components $\mathcal{S}$ should be selected in

align[align omitted — 46 chars of source]

with an error term $\epsilon_i \in \mathbb{R}$ and a function $f(\cdot)$ that describes the impact of $X_i=(X_{i1},\dots,X_{ip})=(X_{ij},X_{i-j})$ for any $j=1, \dots p$ on $y_i$. In general, we assume that $f(X_i)=X_i \beta$, with $\beta\in \mathbb{R}^p$ for an easily interpretable structure, but we also include unknown non-parametric versions of $f$ in our application. We set as usual $Y=(y_1,\dots,y_n)'$ and $X=(X_1',\dots,X_n')'=(X_{(1)}, \dots,X_{(p)})= (X_{(j)},X_{(-j)})$ with $X_{(j)}\in\mathbb{R}^n$ for all $j= 1, \dots, p$.

In the literature, there exist different procedures for the construction of knockoffs such as model-$X$ knockoffs Candes2018, deep knockoffs Romano2019, and group-knockoffs Dai2016. They all build on the same main idea to compare the regressors of interest $X$ with randomly generated knockoffs $\Tilde{X}=(\Tilde{X}_{(1)},\dots,\Tilde{X}_{(p)})$ that fulfill two properties:

enumerate• pairwise exchangeability, i.e. the distribution of\\ $(X_{(j)}, X_{(-j)}, \Tilde{X}_{(j)}, \Tilde{X}_{(-j)})$ and $(\Tilde{X}_ {(j)}, X_{(-j)}, {X}_{(j)}, \Tilde{X}_{(-j)})$ is identical • $Y$ is independent of $\tilde{X}$ conditional on $X$.

When regressing $Y$ on $(X,\Tilde{X})$ jointly, only those regressor components in $X$ which fundamentally differ from their corresponding ones in the random $\Tilde{X}$ according to a variable importance measure are judged as relevant and are part of the active set $\mathcal{S}$. The variable importance depends on model and estimation techniques, but for the linear case, e.g. the difference of absolute lasso coefficients of $X_{(j)}$ and $\Tilde{X}_{(j)}$ must be large enough.

For model-$X$ knockoffs Candes2018, the two knockoff conditions (i) and (ii) are addressed by matching first and second moments of of $X$ and $\Tilde{X}$ in the construction of $\Tilde{X}$ subject to independence of $Y$ and $\Tilde{X}$ conditional on $X$. Matching expectations is straightforward and the second order construction leads to a convex optimization problem minimizing pairwise correlations of $X$ and $\Tilde{X}$ under the constraint that $Cov(X,\Tilde{X})$ be positive semi-definite. This effectively targets the off-diagonal elements of the covariance of $X$ and $\tilde{X}$ and leads to the approximate semi-definite program algorithm (ASDP) of Candes2018. The obtained model-$X$ knockoffs $\Tilde{X}$ approximately fulfill conditions (i) and (ii), and the construction is exact if $(X,\Tilde{X})$ are normal. For a more detailed description of the construction of the more general deep knockoffs Romano2019 which fully operationalize the distributional form of (i) and (ii) and of group-knockoffs that use a pre-specified group-structure Dai2016, see Appendix (ref). In general, the construction principle of all knockoff techniques is based on the standard Gaussian results in Barber2015 and the non-Gaussian, high-dimensional extension in Candes2018 to the model-$X$ knockoff filters.

Once the knockoffs have been constructed, they can be used as a filtering device to select the active set. For this, each knockoff feature $\Tilde{X}_{(j)}$ is compared to its true counterpart ${X}_{(j)}$ via a feature-statistic $W_j$ for all $j=1,\dots, p$. In the linear case, for a lasso regression of $y$ on the joint $(X,\Tilde{X})$ over a grid of penalty parameters $\lambda$ with corresponding coefficients $\hat{\beta}_j(\lambda)$ we work with $\lambda_j=\sup{\{\lambda| \hat{\beta}_j(\lambda) \neq 0\}}$ as the largest $\lambda$ for which variable $j$ is in the active set and define

align[align omitted — 240 chars of source]

where $\lambda_0$ is chosen according to some global criterion like cross-validation and the $\operatorname*{sgn}(\cdot)$ function returns the sign of the input. Note that the $W_j$ from equations (ref) and (ref) correspond to the lasso coefficient difference (LCD) of the model-$X$ knockoffs and the lasso signed max (LSM) as described in Barber2015, respectively. In practice, we mostly rely on the LCD measure which was shown to be preferable and robust to highly correlated features as in our application Candes2018. Only for group knockoffs, we use the proposed adapted version of the LSM as suggested by Dai2016. We only select variable components $j$ as part of the active set $\mathcal{S}$ if $W_j$ is greater or equal to some threshold $T$ with

align[align omitted — 111 chars of source]

where $\alpha \in [0,1]$ is the pre-specified level of acceptable (nominal) false discovery rate $FDR=\mathbb{E}[FDP]$, where $FDP=\dfrac{\vert \hat{S} \setminus {S} \vert }{\vert \hat{S}\vert}$ is the false discovery proportion, with $\hat{{S}}$ as the set of selected variables\footnote{In case $\hat{{S}}$ is empty, we set $FDR=0$ as in Candes2018.} and $S$ as the set of truly relevant variables. In practice, we calculate this proportion as $\widehat{FDP}=\dfrac{\#\{j: W_j\leq -t\}}{\#\{j: W_j\geq t\}}$. Note that for the definition of $T$ we rely on Candes2018 of the original suggestion of Candes2018.

Robustified Knockoffs

First, we propose an adapted version of the baseline knockoff procedures that can deal with time-dependence and an unknown, possibly non-standard covariate distribution due to high correlations among $X$. Secondly, we examine the full grid of possible nominal FDR-levels for the knockoff procedure to uncover dependence of selections on certain specific FDR-levels. We do this by repeating each robustified baseline procedure $K$ times over a grid of $K$ nominal FDR-values and combine the results by weighting the selection probabilities depending on the FDR-level. We call this procedure weighted FDR selection (wFDR).\\ With covariate distributions different from normality and possible time-dependence, the standard assumptions from the Knockoff framework are violated, which could lead to strong variability of knockoff selections that are per definition random. We therefore suggest repeated subsampling to stabilize the selection procedure, motivated by the stability selection of Meinshausen2010 and Ren2021, who suggest a similar procedure, where the knockoff procedure is repeated without subsampling.. We repeat the full knockoff procedure $B=100$ times only using a subsample of the full $n$ observations, with subsampling rate $\theta$. The subsampling ensures that large outliers and data artifacts do not majorly affect the selection, while the repetition of the knockoff procedure controls the randomness of the knockoffs. This randomness would also allow no subsampling at all as in Ren2021. For our application with a substantial amount of outliers in finite samples, however, we choose to use $\theta=0.9$. For a fixed FDR-level $\alpha$, the procedure ranks variables in decreasing order according to their selection frequency, i.e. empirical selection probability. The variable with the highest selection frequency receives rank $p$, the second most selected variable gets rank $p-1$, up to the least selected variable receiving rank 1. Alternatively, we also directly work with the selection probability instead of the ranks, which puts a larger emphasis on the variability of selection probabilities\footnote{Technically, it would also be possible to run a standard knockoff machine without subsampling and report either zero (no selection) or one (selection) for each variable. We refrain from such an approach due to the data challenges stated above.}. See also Method (ref) for details.

algorithm[algorithm omitted — 1,063 chars of source]

Note that by construction the proposed subsampling adapted knock-off procedure keeps the fixed FDR-level $\alpha$ but robustifies the selection result.

Moreover, we propose to conduct each knockoff-baseline selection (Method (ref)) over a grid of $K$ different FDR values $\alpha_k$ jointly. Thus, for each fixed-level $\alpha_k$, we detect whether a variable is relevant or not and determine the corresponding selection probability via subsampling using our methodology. Over the grid of different $\alpha_k$-values, the corresponding selection probabilities are then weighted depending on the level $\alpha_k$, where higher values of $\alpha_k$ receive lower weights corresponding to the definition of the FDR. The final selection probability for each variable is then obtained as the weighted sum of all selection probabilities for this component over all $\alpha_k$. Since the number of selected variables varies depending on the respective $\alpha$-level, our procedure prevents situations where results crucially depend on the pre-setting of one specific $\alpha$-level. We show explicitly that such situations happen in our empirical example where high correlations between variables exist and solve this issue by combining the results from distinct weighting schemes that control the influence of selections over the grid of possible FDRs. With that, we transparently control the FDR influence on selections while maintaining flexibility by not restricting the baseline procedure for selections too strongly. An overview of our method is given in Method (ref).

algorithm[algorithm omitted — 1,134 chars of source]

We suggest two different weighting schemes that all depend on the following observation. By definition, a low nominal FDR implies that the number of falsely selected variables is small compared to the number of selected variables. This suggests that weighting should be conducted in a way that low FDRs, for which selection probabilities are thus more informative, should receive higher weight. Imagine we have two selection probabilities $pr_j^{low}$ and $pr_k^{high}$ for variable $j$ and $k$ at nominal levels $\alpha_{low}=0.1$ and $\alpha_{high}=0.95$, respectively. For $\alpha_{low}$, less than $10\%$ selected variables should be false selections, while for $\alpha_{high}$ less than $95\%$ of selections should be false. Obviously, when $pr_j^{low}$ and $pr_k^{high}$ are very similar, one would give variable $j$ a higher weight in being a true influencing variable compared to variable $k$. To formalize this intuition, we propose two weighted averages and compare them with an unweighted baseline. One average is just based on weights that decay linearly, while the other one uses weights that decay exponentially and are equidistant on the log-scale. More specifically, for the value (probability or rank) at $FDR_k=\alpha_k$, where $k=1,\dots, K$, and $FDR_k \in (0,1)$ on an equidistant grid, the linear weight for position $k$ on the FDR-grid is given by

align[align omitted — 64 chars of source]

and the exponentially decaying weight is given by

align[align omitted — 167 chars of source]

We compare these weights with an unweighted average, where we expect the weighted averages to be more informative and thus give better indications of true influencing variables than their unweighted counterparts.

Empirical Study: Corporate Recovery Rates

Data

Our empirical study uses a data set consisting of 2,079 U.S. corporate bonds that defaulted between 2001 and 2016 obtained from S&P Capital IQ-similar. We retrieved industry and stock variables from Bloomberg Financial Markets. Moreover, we collected 144 macroeconomic variables that were used in previous credit risk studies from the Federal Reserve Bank of St. Louis (FRED, Federal Reserve Economic Data). We classified these macroeconomic variables into 20 groups as detailed in Appendix C. We structured the groups according to financial conditions (Loans, Bank Credit and Debt), monetary measures (Savings, CPIs, Money Supply), corporate measures (Cash Flow and Profit), business cycle (Unemployment, Industrial Production, Private Employment, Housing, Income, Real GDP, Inventories), stock market (Index Returns and Volatilities), international competitiveness (Exchange Rates, Trade), and micro-level factors (Producer Price Index). These groups are tailored to yield interpretable factors and are more granular than in Nazemi2021, who consider a prediction-focused analysis. As can be seen from Table (ref), variables within groups are often highly correlated, which makes it hard to directly analyze them without transforming the data. Although there are still a few highly dependent groups, the correlation between groups is much smaller in general, which can be seen in Figure (ref). There, the median correlation across groups is shown, which the diagonal indicating the median correlation in-group, corresponding to column $50\%$ in the Table on the left of Figure (ref). \\

table[table omitted — 2,577 chars of source]

The recovery rate is defined as the mean of trading price between the default day and data 30 days after default, which we retrieve from Capital IQ. The data is originally from the Trade Reporting and Compliance Engine (TRACE). In our analysis, all corporate bonds have debt values, at the time of default, of greater than \$50 million. The mean value of the recovery rate for the 2,079 U.S. corporate bonds in our sample is 45.57 percent, and the sample standard deviation is 35.04 percent. The empirical distribution of the recovery rates of defaulted US corporate bonds naturally peaked in the financial crisis from 2008-2010\footnote{A more detailed figure on the distribution of defaults over time can be found in Figure (ref) in Appendix (ref). A large share of defaults was caused by both the Lehman Brother bankruptcy in September 2008, and the CIT Group Inc. bankruptcy in November 2009.}. Around 30 percent of defaulted bonds have recovery rate less than 10 percent. There is another distribution peak in the range of values between 60 percent and 70 percent, which is visualized in Figure (ref).

In addition, our bonds consist of four seniority levels for the bonds: (i) senior secured, (ii) senior unsecured, (iii) senior subordinated, (iv) subordinated, and junior subordinated. An overview over the distributions over the different bond types can be found in Appendix (ref) in Figure (ref). Since most defaults (82.5%) occurred in the class of senior unsecured bonds, which is driving the distribution of recovery rates, we decided to not distinguish between groups of seniority levels in our analysis. Additionally, the sample size would be too small for such a sub-analysis.

figure[figure omitted — 271 chars of source]

Empirical Results

In this subsection, we use our methodology to identify and quantify the recovery rates of corporate bonds in a data-driven way. This is key in practice for investments, hedging, and supervision but also for model building and interpretation. As an extension in a comprehensive out-of-sample forecasting study, we also demonstrate that simple linear predictions based on variables selected by the knockoff procedures can compete with and often improve upon non-parametric methods that employ the full set of variables. Moreover, we show that the obtained knock-off selection of variables is robust to discarding certain time subperiods, e.g. after the financial crisis.

Identification of Important Groups and Effects on Recovery Rates

To determine the driving factors of corporate bond recovery rates, we use our proposed methodology on the full sample from 2001 to 2016 for the data-driven selection of relevant components. We show results of our suggested combined subsampling and weighted FDR selection technique (see Method (ref) and (ref)) across all types of different baseline-knockoff methods, i.e. in particular, we study model-$X$ knockoffs, deep knockoffs with two different neural network architectures, and group knockoffs.

Since our data is highly correlated within groups (see Table (ref)) and we are mainly interested in group effects and selections, we transform our data using principal component analysis (PCA)\footnote{See e.g. Hastie2009.}. To retain interpretability on a group level, we conduct one PCA per group and only use the most important principal components (PC) to describe that specific group, i.e. a maximum of four PCs that explain at least 90% of the variability in the group. This helps to reduce high correlations among variables and break down large groups of variables to one or two components to see their main effects, avoiding multicollinearity issues in post-selection linear models. Additionally as a robustness check for the model-$X$ knockoffs, we employ a smaller version using a maximum of two PCs ("2comp").This group-PCA step greatly reduces the dimensionality of the data and serves as a viable alternative to other pre-screening procedures such as omitting variables with high pairwise correlations. With that, the variables are scaled by their standard deviation and centered around zero. Similar approaches in reducing dimensions have also been taken by Kelly2019 in an asset pricing application.

Figure (ref) graphically shows the most important PCA-features for model-$X$ knockoffs over all possible nominal FDRs, while similar figures for the other procedures can be found in Appendix (ref) (Figure (ref)). Most prominently, the selection probabilities of important features are rather high (with selection frequency of 0.6) already at a nominal $FDR=0.2$ and rise to 0.8 at nominal $FDR=0.4$ for the model-$X$ procedures, where other, less important factors only attain similar levels from a nominal $FDR=0.8$ onwards. Such levels of FDR are clearly undesirable, but investigating the entire grid of FDR-levels jointly and with appropriate weighting is beneficial and yields robustness due to the rather high variability of selection probabilities for minimal changes at a considered specific nominal FDR-level. For the deep knockoff procedures, however, we see high selection probabilities for relevant features throughout all FDR levels (see Figure (ref) (bottom)). Comparing different structures for the neural networks in the deep knockoffs, this effect is more pronounced for narrower networks with only 5 neurons per layer. Since a wider network can learn more complex structures, it can build more accurate knockoffs and with that, identify variables that are less likely to be true influencing variables, especially for cases where nominal FDR-levels are low. This highlights the importance of considering multiple methods for generating knockoffs and combining their insights to identify the most important groups.

figure[figure omitted — 602 chars of source]

To finally select appropriate variables over all five methods and FDRs, we compare two different ranking procedures for variables at each FDR-level. To combine the different rankings from each FDR-level to a final selection (probability), we employ three distinct weighting approaches. First, we distinguish between using ranking of variable selections (from 20 to one) or selection probabilities from our procedures. The subsequent weighting of ranks/probabilities for each variable over all nominal FDR-values is conducted as in Section (ref) using either equal weighs, linear decaying weights ($\omega_k^{lin}$, Lin-decreasing), or exponentially decaying weights ($\omega_k^{exp}$, Log-decreasing). The weights assign the highest weights to low FDR-values and decrease with higher FDR-levels (see Section (ref) for details). Figure (ref) shows boxplots over selections from different methods by both ranking procedures and the three weighting schemes. In general, we can see that group 14, 11, and 12 are always among the top four in each procedure, while group 5 only appears as important when looking at probabilities, and group 20 vice versa only when looking at ranks. Table (ref) highlights the influence selection probabilities and ranks, where both group 5 and group 20 are given higher relative importance in the weighted schemes that focus on low nominal FDRs. Otherwise, results for the most selected groups are mostly stable over both schemes and all weights. This is largely in line with the literature that also determines the factors in group 14, 11, 20, and 5 as relevant with a more simplistic and less robust data-driven selection technology Jankowitsch2014,Nazemi2018\footnote{Other, less important groups that have been selected often by our approach include high yield default rates, defaults in the respective industry, and GDP measurements. These findings are also in line with Nazemi2018 and Jankowitsch2014, who report similar factors to be important.}. Though different from the existing studies, however, our methods additionally also detect group 12 as important, that consists of exchange rates.

figure[figure omitted — 583 chars of source]
table[table omitted — 1,438 chars of source]

Note that group 14 describes bond yields of major bonds and rates of different general indicators such as mortgage, treasury, and loans, which might be considered naturally predictive for the state of the economy and thus of bonds. The data-driven selection therefore confirms the intuition that these indicators have an influence on recovery rates. Similarly, the selection of the other groups can be explained. These describe international exchange rates against the USD (group 12) and stock market indicators such as returns and volatilities of the most important indices (group 11). Furthermore, capacity utilization of industries and change in inventories of private and businesses play an important role (group 20), as well as corporate measures such as profit of the firm and cash flow (group 5). More specifically, in the light of the financial crisis in 2008 and the following euro crisis, exchange rates were strongly affected McCauley2009,Kohler2010, which might explain their connection to recovery rates in regard to globally active firms. The capacity utilization group, on the other hand, measures how much of total potential output is actually utilized by industry and additionally contains information about inventories (and their change over time). This is highly relevant for recovery rates when thinking of a firm's business model in general and inventories of firms that could indicate how much can be recovered given default. Naturally, returns and volatilities of the most important indices such as the S&P 500 or the NASDAQ 100 describe the general situation of the economy and the value of companies, which again is an indicator of recovery rates of these firms. Finally, profit and cash flow are probably the most direct factors for short-term companies finances, and are thus a good predictor for the default of a firm.

As a robustness check, we also computed variable importance measures from a random forest model using mean variance reduction as a measure for the importance of a group\footnote{To give the nonparametric random forest maximum flexibility, we used the the raw data as input instead of using the aggregated PCA-groups.}. Computational details of this model can be found in Appendix (ref). Table (ref) in Appendix (ref) shows the mean values aggregated on a group level of variance reduction and corresponding p-values from the PIMP procedure of Altmann2010. There, we can confirm that especially group 12, 14, and also 11 have high importance, while group 20 and group 5 are not deemed important. This can be explained by the predictive nature of the measure and method that favors groups such as group 15, which measures the bond defaults within the industries. Interestingly, group 15 is also in the top 6 groups of many of the other procedures, although mostly ranked below all the other selected groups.

Post-Selection Performance

To obtain an unbiased quantification of the effect of each factor, we re-estimate a linear model using only those groups that were selected in Section (ref). We show the effect of the most important principal components (PCs) for each group in Table (ref). While the effects of selected variables appear mostly significant, it has to be noted that assuming a linear model might be too optimistic and effects between groups might be affected by some remaining multicollinearity between similar groups.

Using the PCs allows using and working with the strong correlation within groups, and facilitates the interpretation and comparison of effects between different groups since variables in each group are centered and scaled. The results of Table (ref) show that most of the selected coefficients seem highly significant, where focus should lie on the first and second PC which capture the largest share of the variance in each group. Group 14 has a largely negative impact on recovery rates, although especially the last PC is affected by inclusion of more variables switching signs. Group 11 has a primarily positive impact (in the first two PCs), while adding large explanatory power (Adjusted $R^2$), It is, however, also correlated with group 12, which adds little explanatory power, but also has a significant positive first PC. Group 20 and 5 have a negative impact but appear to affect the coefficients of other components, which is why we consider their inclusion rather as a robustness check, since they also do not add as much to an increase in $R^2$ as for example group 14 and 11. Group 12 is a special case, having both positive (PC1) and negative (PC2) significant coefficients. Taking a closer look at the weights of the PCs\footnote{A complete list of the PCA-weights can be found in Table (ref) in the Appendix.}, the first PC assigns a large negative weight to the exchange rates of Canadian dollar, Swiss franc against one USD, and the real broad effective exchange rate for the US, while the second PC gives large negative weight to the rate of USD against one British pound.

table[table omitted — 2,855 chars of source]

It is in line with intuition that higher bond yields and rates such as mortgages have a negative impact on recovery rates (group 14), since they might indicate a riskier environment. The positive impact of group 11 can be attributed to the fact that higher returns and smaller volatilities in the stock indices result in larger recovery rates. At the same time, the positive impact of exchange rates (group 12) in the first PC indicates that when the USD is weak against other major currencies, recovery rates are higher, while the opposite effect is observed in PC2. This effect could be explained by defaulted companies holding assets in foreign currencies that are more valuable when the USD is weak. On the other hand, we cannot fully rule out that this effect is caused by the USD exchange rate dropping against other major currencies Kohler2010 because of large crash events that are connected to bond defaulting.

Extension: Out-of-Sample Prediction Performance

In addition to the identification and interpretation of important factors explaining recovery rates, we also assess the out-of-sample forecasting performance of the reduced models in various scenarios. Here, we distinguish between two main cases: firstly, we check the infeasible forecasting scenario as reference point, where we use the determined models from Section (ref) employing information from the full data comprising 2001-2016 in the model selection step when forecasting for the year 2012-2016 (Table (ref)). Additionally, we also provide results for the “completely” out-of-sample forecasting case , where we re-determine all models on a limited time period from 2001-2011 and predict 2012-2016 (see Table (ref)). For the post-selection estimation step, we use a wide variety of models ranging from simple linear methods, standard and penalized, up to flexible fully non-parametric methods such as random forests (see Appendix (ref) for implementation details). Moreover, we consider both cases with the full raw data and group-PCA-transformed data in different settings. In all settings, we clearly confirm that knockoff pre-selection improves prediction performance. The results also highlight that this cannot be achieved with lasso pre-selection, thus confirming the importance of our robust approach. After knockoff pre-selection, simple (penalized) linear forecasting models often achieve quite competitive forecasting performance with only slight improvements by a non-linear fit. This highlights that forecasting with a data-driven selection of important predictors pays-off, while maintaining easy interpretation in comparison to their fully non-parametric counterparts.

We assess the forecasting performance by calculating the root-mean-squared forecasting error $RMSE=\sqrt{\dfrac{1}{K}\sum_{\tau=k}^T (\hat{y}_\tau-y_\tau)^2}$ and the mean-absolute error $MAE=\dfrac{1}{K}\sum_{\tau=k}^T \vert\hat{y}_t-y_\tau \vert$ for a prediction $\hat{y}_\tau$ of $y_\tau$ at forecasting time $\tau=k,\dots,T$, and forecast length $K=T-k$. We use different forecast constructions with fixed, expanding and rolling windows on annual and daily horizons. For the fixed window type, we set the training data to 2001-2011 and provide daily predictions for 2012-2016. In the expanding window case, we use data from 2001 up to a certain year $\tau$ in the set $\{2011, 2012, 2013, 2014\}$ and predict daily values in $\tau+l$ where $l\in \{1,2\}$\footnote{We use an expanding window here to account for the difficulty of predicting two full years at once.}. For daily rolling windows, we set the training length to $10$ years as in the initial expanding case and the fixed window setting and predict one corporate default observation ahead (Daily)\footnote{This does not necessarily mean that this is one day ahead ahead, as some defaults occurred on the same day. We chose to always jump to the next day containing a default to maintain a realistic time structure in that scenario (see also Nazemi2021).} We estimate either cross-validated elastic nets (mixing parameter $\alpha=0.5$)\footnote{More specifically, the penalty in the objective function is specified as $\sum_{j=1}^p (\alpha \vert b_j \vert + (1-\alpha) b_j^2)$}, cross-validated lasso regressions, or simple linear models, and use random forests as non-parametric benchmarks. For each window construction, we employ the post-selection methods either with the full raw data or the group-PCA transformed data (as described in Section (ref), we use as many PCs to explain 90% of the variance in each group, see also Table (ref)). We either use the above data without pre-selection or employ the pre-selected set of variables according to the different setups in Section (ref). This comprises using our proposed weighted FDR selection (wFDR)\footnote{For the wFDR, we use all PCA-components from the three most-selected groups, i.e. 14,12,11. See also Table (ref) for comparison.} combining all baseline knockoff procedures or using only our repeated-subsampling procedure (see Procedure (ref) in Section (ref)) in combination with the baseline-methods. These are either model-$X$ knockoffs (MX, or MX 2 Comp. using a maximum of 2 PCs) or deep knockoffs with 5 (Narrow) and 25 (Wide) neurons per layer. For the infeasible reference scenario in Table (ref) and the group-PCA-transformed data, we use PCs that are estimated over the full data set. In the “completely out-of-sample” forecasting case in Table (ref), the out-of-sample PCs for time points after 2011 are created using the weights from the PCs with only data up to the end of 2011\footnote{For comparability, scaling/centering uses information of the entire sample. But scaling with weights from data up to the end of 2011 does not substantially change the prediction performance. Results are available from authors upon request.}.

table[table omitted — 2,716 chars of source]
table[table omitted — 2,268 chars of source]

Table (ref) shows that generally simple linear models with limited pre-selected variables from our proposed weighted FDR procedure that combine different knockoff selections works best for forecasting both longer and shorter time horizons. Moreover, as a single selection techniques, also, the model-$X$ procedure within our robustified framework yields excellent results with a simple linear post-selection fit. Determining variables with the Deep Knockoff robustified framework generally performs slightly worse also for nonlinear post-selection models, with the relative best performance for shorter forecasting horizons. This is not unexpected since for the deep knockoffs, selection probabilities were generally much higher, meaning they could contain more noise variables that would bias predictions for longer time horizons (i.e. they do not generalize as well as the weighted FDR counterparts). The baseline linear models using the full raw data perform poorly for large time horizons, especially for fixed windows, which can be explained by potential overfitting on noisy data. Interestingly, this cannot be fully countered by regularization using elastic nets for forecasting tasks. Only for very short time horizons as the daily rolling window, the baseline procedures can compete. These findings are highly in favor of using our proposed statistical model selection techniques also for forecasting tasks. The machine learning (ML) benchmarks with selection on the full raw data perform similarly, but always slightly worse than the knockoff counterparts with the generally downside of lacking transparency and interpretability of the influence of certain groups and factors. Using the PCs instead of raw data only significantly improves the random forest model for the fixed window. Note that for this study, we used the standard recommended data-driven choice of tuning parameters for the machine learning benchmarks but did not additionally fine-tune from there in order to maintain comparability between the simple baselines, the knockoff procedures, and the benchmark models.\footnote{While tuning all hyperparameters cautiously could improve ML-forecasts to some extent, previous studies for recovery rates show that expected changes are minor (see e.g. Nazemi2021 with additional news-based variables).}.

In the “completely” out-of-sample scenario, we repeat the model selection step from Section (ref), but only use data up to the end of 2011 to determine the relevant variables with our proposed methodology for all different knockoff-baseline procedures. This represents the most realistic but also most challenging scenario for the knockoff procedure, where post-crisis recovery rates are not contained in the training set but must be predicted. The resulting stability in selections and in forecasting performance therefore indicates that our choices are important also for non-crisis periods. Table (ref) summarizes the new results\footnote{We did not include the Deep Knockoff procedures here since the sample size is significantly reduced for selections.}. As expected, our proposed methodology (“wFDR") is on par with the top-performing machine learning methods in this case as well, although the random forest with the full raw data is slightly better for the daily window. In comparison to the infeasible full-sample selection results in Table (ref), the single model-$X$ knockoffs perform slightly worse, with less variability between the different knockoffs employed within the subsampling knockoff framework. This can be explained by the smaller data set, where fewer variables are selected in general and difference between the different model-X knockoffs is smaller. Selections in general are the same compared to the full data case for our proposed weighted FDR procedure, and very similar for the single “baseline" methods, with only a few minor changes (see Table (ref) in Appendix (ref) for details). In terms of forecasting performance, the ranks for the different procedures are stable, meaning that our proposed weighted FDR methodology together with the random forest are still performing best, while using no or no robust selection still performs worst overall. The margin between the latter and our procedure is smaller only for the daily rolling window and the MAE, where very bad predictions (e.g. in cases of large default events) are not punished as heavily as with the MSE. Using the usual MSE-measure, the raw data with an elastic net (or lasso) still perform considerably worse.

table[table omitted — 3,335 chars of source]
table[table omitted — 2,503 chars of source]

As an additional robustness check for the predictive performance of our methods, we compute model confidence sets Hansen2011 for the same forecasting combinations as before and for both the infeasible scenario and the “pure forecasting case”. As suggested in Hansen2011, we use $B=5000$ bootstrap replications, the TMax test-statistic, and test-level $\alpha=0.15$. Please see Appendix (ref) for details. The displayed results are robust across different $\alpha$ test levels in the standard range $[0.1; 0.2]$. \footnote{Results are omitted here but are available upon request from the authors.} Although the model confidence sets differ over the various forecasting schemes, we can identify models that consistently fall into the model confidence set. Please see Table (ref) and Table (ref) for full-sample selection and completely out-of-sample forecast results, respectively. We want to highlight that our proposed wFDR procedure always belongs to the model confidence set, together with the random forest procedure in the full-sample selection scenario. Over the full sample selection, the other model-X procedures can still compete, while in the “completely” out-of-sample case, they are outperformed by our proposed wFDR. This is also confirmed by the corresponding test statistics, where a lower value indicates better performance. More specifically, a negative value of the test statistic indicates that the average loss is smaller compared to all other methods in the confidence set, where our suggested method clearly outperforms the other procedures in general for the full-sample selection, while for the “completely” out-of-sample scenario, the raw random forest is slightly better when comparing the test-statistic. Using no or non-robust selection methods (i.e. lasso, elastic net) is always worse, and also the plain group knockoff does not perform well on our data set, which might be caused by our specific data structure, where we still have some strong correlations between groups.

Conclusion

In this paper, we demonstrate the benefit of connecting the flexibility of the knockoff framework with repeated subsampling and techniques controlling the proportion of false discoveries over the full spectrum of possible values. We employ a comprehensive set of distinct knockoff machines and illustrate that a transparent combination of their results yields optimal ensemble results.

With the proposed methodology, we are able to uncover important macroeconomic factors of corporate bond recovery rates while maintaining excellent forecasting performance. In particular, predictive power in various settings using linear models with just the identified groups is significantly higher than using the full set of variables in similar models. Furthermore, our procedure outperforms other model selection procedures and performs similar to flexible machine learning methods. The latter are developed for prediction tasks but lack easy interpretation and identification of important factors in contrast to the proposed methodology.

For future research, the proposed technique shows high-potential in other data-rich environments, such as asset pricing or climate modelling. In a separate paper, it would be of interest to derive conditions for FDR-level components and optimized forms of parsimonious weighting schemes to theoretically achieve and derive optimal in-sample or out-of-sample fits and respective statistical rates.