EconBase
← Back to paper

A Pipeline for Variable Selection and False Discovery Rate Control With an Application in Labor Economics

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

62,869 characters · 13 sections · 35 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

A Pipeline for Variable Selection and False Discovery Rate Control With an Application in Labor Economics

\thispagestyle{empty}

abstractWe introduce tools for controlled variable selection to economists. In particular, we apply a recently introduced aggregation scheme for false discovery rate (FDR) control to German administrative data to determine the parts of the individual employment histories that are relevant for the career outcomes of women. Our results suggest that career outcomes can be predicted based on a small set of variables, such as daily earnings, wage increases in combination with a high level of education, employment status, and working experience.

Keywords: Variable selection, aggregated FDR control, knockoff-filter, inference, penalized logistic regression, female employment careers

\doublespacing

Introduction

Titles like “Economics in the age of big data” Einav2014, “Big data: new tricks for econometrics” Varian2014 and “Beyond prediction: using big data for policy problems” Athey2017 highlight a shift in the scale of economic data towards Big Data. The term Big Data refers to settings, where we observe information on a large number of units, or many pieces of information on each unit, or both, and often in a complex setting with several crosssections per unit. Nowadays, empirical research in economics increasingly relies on newly available large-scale administrative data sets, which offer new possibilities for fitting models.

In particular, Big Data applications bring the opportunity to fit high-resolution models. Complex models arise naturally in Big Data, where the total number of measurements is large. In the extreme case, fitting sophisticated models can mean to deal with parameter spaces that are comparable ($p \approx n$) to the number of samples or even larger ($p \gg n$). Parameter spaces that are too high-dimensional ($p \approx n$ or $p \gg n$) can not be examined with standard estimation methods and therefore rely on tools from high-dimensional statistics. High-dimensional statistics comes into play whenever one fits complex models ($p$ is large). In high-dimensional settings the number of variables is substantial, but there is often a sense that many of the variables are of minor importance or completely irrelevant. The basic concept is to focus on the most relevant parts of the parameter space, which are identified by leveraging prior information. High-dimensional statistics complement classical estimators with so-called penalty or prior functions that formulate mathematically how “likely” or “promising” certain models are. The data-driven calibrated tuning parameters weight the prior function and thus determine the degree of regularization, i.e. for example, the degree of feeding the bare measurements with additional information.

The notion most commonly imposed on the prior functions in high-dimensional statistics is sparsity. Sparsity means that the data generating process can be modelled accurately by using only a small number of variables even though the actual number of variables at hand is large. The probably most discussed and prominent sparsity-inducing penalty function, such as in the Lasso literature Tibshirani1996, is the $\ell_1$-prior, which we also focus on. If the data generating process can be approximated by a sparse model, it makes sense to speak of variable selection and false discovery rate (FDR) control Benjamini1995. Variable selection means that we are interested in teasing apart the “relevant variables” (variables with non-zero-valued entries in the parameter vector) from the “irrelevant variables” (variables with zero-valued entries in the parameter vector). FDR control means that we want to ensure that the expected fraction of false discoveries among all discoveries is not to high in order to guarantee that the selected variables are indeed the true ones.

Roughly speaking, sparsity is the main ingredient of variable selection, since we are often interested in finding a few of the important covariates, that is to extract the relevant variables from those which are scientifically not extremely useful for understanding the dependence between the response variable and the few truely relevant covariates.

The final goal is to draw inferences about the estimates $\hat{\boldsymbol{\beta}}$ of the selected variables. In most cases, it is impossible to recover the true subset $\mathcal{S}$ of variables with no error. Hence, we are naturally interested in procedures that keep the resulting error small. This is typically measured in terms of false positives and false negatives. It is also known as the type I and type II errors, respectively. False positives are falsely selected variables, and false negatives are falsely omitted variables. A suitable measure to evaluate the estimator's variable selection accuracy is, for example, the hamming distance, the sum of false positives and false negatives.\footnote{The hamming distance is defined as $\text{hd}:= \text{fp}+\text{fn}$, where $\text{fp}:=|\{j \in \{1,\dots,p\}:\hat{\beta}_j \neq 0 \ \text{and} \ \beta^*_j = 0 \}|$ and $\text{fn}:=|\{j \in \{1,\dots,p\}:\hat{\beta}_j=0 \ \text{and} \ \beta^*_j \neq 0 \}|$. } The smaller this number, the better the estimator's performance.

How does variable selection and FDR control work? In classic hypothesis testing we are commonly interested in controlling the Type I error at a certain significance level $\alpha$ by individually testing $p$ null-hypotheses of the form $\mathfrak{H}_{0,j}: \beta_j^* =0$, for $j \in \{1, \dots, p\}$. Extending the Type I error control to multiple testing, we are able to control the FDR. The FDR is the expected proportion of false discoveries and total discoveries. However, accurate hypothesis testing also requires that the (statistical) power, the proportion of correctly selected hypotheses and total number of true hypotheses, is large. The power is typically measured in terms of true positives. True positives are truely selected variables. Accordingly, the number of true positives is tp$:=|\{j \in \{1,\dots,p\}:\hat{\beta}_j \neq 0 \ \text{and} \ \beta^*_j \neq 0 \}|$.

Therefore, in addition to the FDR, we also control the aggregated false discovery rate (AFDR) recently proposed by Xie2019. It maximizes the power while simultaneously guaranteeing FDR control. This aggregation scheme is an improvement on the original FDR and has the advantage of having the same theoretical guarantees as the FDR. Consequently, we can use the same methods as for the FDR to control the AFDR.

Indeed, we want to find as many variables as possible while at the same time not having too many false positives. We focus here on controlling the (A)FDR with the knockoff filter Barber2015, a data-driven procedure. There is a variety of other procedures to control the (A)FDR, but the generality and flexibility of the knockoff-filter makes it so worthwhile to introduce and prepare this procedure to economists. The knockoffs are designed to mimic the correlation structure of the design matrix $X$, in a way that allows for accurate (A)FDR control. The knockoff filter can be adapted for a broad class of models and for a variety of test statistics Candes2016,Romano2019. Moreover, the knockoff procedure works under any fixed design matrix $X$ and does not require the strong assumptions for the covariates, usually required for most variable selection techniques in high-dimensional statistics Zhao2006. In fact, high-dimensional tools often have to limit the correlations between the covariates. Typically the irrepresentability condition is also required which limits the correlations between the “relevant” and “irrelevant” variables.

Why is variable selection and (A)FDR control for economists of interest? First of all, variable selection and (A)FDR control answer a simple, but very fruitful, question: which variables does an outcome of interest depend upon? Economists often want to know which demographic or socioeconomic variables affect future economic outcomes such as income. Answering this question is not easy, as economic data sources become more and more detailed, providing a flood of potential explanatory variables, often knowing full well that the outcome of interest only depends on a small fraction of it. Especially in administrative data, the main workhorse in labor economics, there is due to the longitudinal nature a deluge of explanatory variables possibly interacting in many different ways. So far, many economic papers have shied away from including many explanatory variables. They subconsciously assumed that the data generating process is characterized by only a few variables that matter for the outcome of interest. Important covariates are included based on economic reasoning. As consequence, relevant explanatory variables which enter the model with complex and very flexible functional form are usually not accounted for. Even though economists may believe that an economic outcome depends on a small set of variables, they have a priori little or no clue about which ones are relevant. Therefore, modern high-dimensional tools aimed to controlled variable selection can help economists to tease apart the relevant from irrelevant variables and to achieve valid inference.

So far, in the empirical research, economists limit the number of explanatory variables by hand, rather then choosing them in a data-driven manner. Only for prediction tasks we already have many insights from the economic literature on data-driven techniques. In recent years, scholars adopted machine learning (ML) tools in economics. Indeed, most recent influential reviews of ML methods aimed to economists (for example Mullainathan2017,Athey2019) introduce ML as a powerful tool to solve problems around prediction. There have been successful applications of the prediction methodology to policy problems, where ML tools have been embeeded in the context of economic decision-making Kleinberg2015,Kleinberg2018. However, in the economic literature, there is less knowledge about variable selection and certainly not about (A)FDR control. Variable selection with (A)FDR control solves problems around parameter estimation, the main goal in applied economics. The (A)FDR control is a new framework of controlled variable selection. It achieves correct inference in such a broad setting by constructing so-called knockoff variables which serve as a kind of control group for the covariates. Therefore, we prepare this framework for economists and show that we can gain new insights for the empirical work.

In this paper, we demonstrate the potentials of model-X (MX) knockoffs Candes2016 in the field of economics based on an empirical application towards the labor market. MX knockoffs is a new data-driven tool that allows to link a large number of potential covariates to an outcome of interest in a nonlinear fashion. It identifies a subset of important variables from a large set while controlling the (A)FDR.

In particular, we are interested in which variables from individual employment and wage histories affect future professional careers of women. Due to the longitudinal nature of the administrative employment records there is complex information on individual employment histories (for example types of employment, wages, skill level, occupation, age, professional experience) for the entire elapsed working live possibly interacting in many different ways. Despite of having a large set of potential covariates, we assume that the binary career outcome only depends on a small fraction of it. The main challenge is to search for the few truely important variables which are linked to the response in a nonlinear fashion. We use a binary choice model within the class of generalized linear models and apply $\ell_1$ penalized logistic regression. To achieve valid inference in our setting, we apply the MX knockoffs to a data set from the Sample of Integrated Employment Biographies (SIAB). Correct inference in our controlled variable selection problem means that we effectively control the (A)FDR even in logistic regression with a large set of covariates. Thus, we can be pretty confident that the MX knockoff procedure correctly teases apart important from irrelevant variables while maximizing power and guaranteeing FDR control.

It is well known from the economic literature that a higher educational level is associated with better career outcomes in terms of higher earnings. There is a huge economic literature showing the positive labor market returns of a high educational attainment (for example Boeckerman2019,Card1999,Heckman2018). There is also some indication in the literature that the occupational choice is related to the career outcomes Almquist1970, but there seems to be no extensive analysis of which of the several hundred explanatory variables from individual employment and wage histories are truely associated with future professional careers of women.

Our results suggest, first, that individual employment histories have predictive power with regard to professional careers and second, that only a relatively small subset of information gained from employment records is crucial.

The remainder of this paper is organized as follows. Section (ref) introduces the model set-up and the methodology. Section (ref) describes the data, after which Section (ref) presents the main empirical results. The final section concludes the paper.

Methodology

In our empirical application, we use a large-scale model that links a large set of potential explanatory variables to the response variable in a nonlinear fashion. We want to discover which variables from individual employment histories are truely associated with future professional careers of women. The data consist of vector-valued samples, each of them describing the employment and wage situation of women over the past five years. The design matrix $X \in \mathbb{R}^{79\,782 \times 327}$ is prepocessed and scaled such that it follows a multivariate normal distribution. The outcome variable $\boldsymbol{y} \in \mathbb{R}^{79\,782}$ is an indicator that takes the value $1$ if the woman makes a professional career in 2010, otherwise 0.

Model and assumptions

We consider data in form of a real-valued deterministic design matrix $X \in \mathbb{R}^{n \times p}$ and a binary response $\boldsymbol{y} \in \mathbb{R}^n$. We denote the rows of $X$ by $\boldsymbol{x}_1, \dots,\boldsymbol{x}_n \in \mathbb{R}^p$ and the columns of $X$ by $\boldsymbol{x}^1, \dots,\boldsymbol{x}^p \in \mathbb{R}^n$ . The design matrix $X$ is assumed to be scaled such that $(X^{\top}X)_{jj}=n$ for $j \in \{1, \dots, p\}$.

We define the vector of residuals as $\boldsymbol{u} =(u_1, \dots, u_n)^{\top}$ with entries $u_i= y_i - \mathbb{P}(y_i=1 | \boldsymbol{x}_i)$ for $i \in [n]$. The vector $\boldsymbol{u}$ is random noise with mean zero, i.e. $\mathbb{E}(y_i - \mathbb{P}(y_i=1|\boldsymbol{x}_i)) = 0$ . Further, we assume $\mathbb{E}(u_i|\boldsymbol{x}_i) =0$ and $\mathbb{E}(u_i u_j) =0 \quad \forall i \neq j$, i.e. our models contain only exogenous variables and the error terms are assumed to be independent and identically distributed (i.i.d.).

The design matrix $\ensuremath{X}$ and the response vector $\ensuremath{\boldsymbol{y}}$ are linked to standard logistic regression model

align[align omitted — 200 chars of source]

where $\boldsymbol{\beta}^* \in \mathbb{R}^p$ is the unknown regression vector.

We assume known that our data generating process is sparse, that is, only a small and a priori unknown number of predictors is relevant for predicting the careers of women ($|\{j:\beta_j\neq 0 \}|\ll \min\{n,p\}$). Sparsity can be motivated on economic grounds in situations where a researcher believes that the economic outcome can be modeled accurately by using only a small number of variables (relative to the sample size) but is unsure about the identity of the relevant variables. Traditionally in the empirical literature economists assumed that the model of interest is characterized by a small number of variables, and they limited the number of explanatory variables by hand, rather than choosing them in a data-driven manner. However, recent advances in high-dimensional modeling have prompted economists to take advantage of the recent big data revolution to deal with large dimensional data sets, in a way that maintains the interpretability of economic models by searching for the truely important variables (for example Belloni2011a,Fan2011). The sparsity assumption imposed there allows the effective use of a large set of covariates while at the same time maintaining the spirit of parsimonious models in the economic discipline. For this reason, we follow this strand of literature on the sparse modeling of economic processes.

Estimation

We use penalized techniques to solve the logistic regression and choose the $\ell_1$ penalty throughout our estimations. The $\ell_1$ penalty is suitable for sparsity because it forms a square constraint region for the parameter vector, and the least-squares contours are likely to intercept the constraint region at the extremes, such that certain coordinates of the parameter vector are set to zero. Moreover the $\ell_1$ penalty is a convex function, which makes it convenient to optimize over.

The goal is to estimate the support set $\mathcal{S} = \text{supp}(\ensuremath{\boldsymbol{\beta}^*})$ for the model in eq. (ref) with the family of regularized estimators

align[align omitted — 285 chars of source]

where $\mathcal{L}[\ensuremath{\boldsymbol{\beta}^*},X , \ensuremath{\boldsymbol{y}}] =\sum_{i=1}^{n} (\log(1+\exp(\ensuremath{\boldsymbol{x}}_i^{\top} \ \ensuremath{\boldsymbol{\beta}^*} ))- y_i \ensuremath{\boldsymbol{x}}_i^{\top} \ \ensuremath{\boldsymbol{\beta}^*}) /n $ is the negative log-likelihood function, $r \in [0,\infty)$ is a tuning parameter and $h:\mathbb{R}^p \to [0,\infty]$ is a prior function which we specify as sparsity inducing prior function

align*[align* omitted — 80 chars of source]

where $|\cdot|$ denotes the absolute value and $||\boldsymbol{\beta}^*||_1:= \sum_{j=1}^p|\beta_j^*|$ is the $\ell_1$-prior.

Next, to avoid an unwanted overall shrinkage of the estimates imposed by the $l_1$ penalty, we refit the regularized estimators $\hat{\boldsymbol{\beta}}$ with subsequent logistic regression estimation on the support $\hat{\mathcal{S}}:=\text{supp}(\hat{\boldsymbol{\beta}})$:

align[align omitted — 379 chars of source]

Inference

In this paper, we focus on controlling the aggregated false discovery rate (AFDR) Xie2019, which we can define as follows: letting $\mathcal{S}^* := \{j \in \{1, \dots,p\}: \beta_j^* \neq 0 \}$ be the true support set and $\hat{\mathcal{S}}_q[[X,\ensuremath{\boldsymbol{y}}], k]$ an estimate of $\mathcal{S}^*$ operating on data $[X,\ensuremath{\boldsymbol{y}}]$. Then, the AFDR at target level $q = \sum_{i=1}^k q_i \in [0,1]$ is

align[align omitted — 367 chars of source]

where $|\cdot|$ denotes the cardinality of a set. The AFDR scheme is nothing else as applying the FDR control method Benjamini1995 $k$ times with specific target levels $q_1, \dots, q_k$ and combining the results by taking the union. Consequently, for $k=1$, we consider the original FDR control scheme. Because the AFDR scheme is equipped with the same guarantees as the original FDR control method, the same procedures that are used to control the FDR can here be applied.\footnote{For the proof see Xie2019.}

Next, we give a quick introduction to the model-X knockoff filter which we use in our empirical application to control the AFDR. The model-X (MX) knockoff filter is a method for high-dimensional controlled variable selection in any class of generalized linear models (GLM) Candes2016. It extends the knockoff procedure that was originally designed for controlling the FDR in low-dimensional linear models Barber2015. The key ingredient of the knockoff filter are the generated knockoff copies $\tilde{\ensuremath{X}} \in \mathbb{R}^{n \times p}$ for the design matrix $X$, which mimic the correlation structures between the variables and therefore serve as a control group for them to ensure that not too many irrelevant variables are selected. The final goal is to perform FDR control on the specific statistics based on both $X$ and $\tilde{\ensuremath{X}}$.

Denote $X = (\ensuremath{\boldsymbol{x}}_1, \dots,\ensuremath{\boldsymbol{x}}_n)^{\top}$ and $\tilde{\ensuremath{X}}=(\tilde{\ensuremath{\boldsymbol{x}}}_1, \dots, \tilde{\ensuremath{\boldsymbol{x}}}_n)^{\top}$. In this paper, we generate MX knockoffs $\tilde{\ensuremath{X}}$ from a multivariate normal distribution obeying

align[align omitted — 84 chars of source]

where we assume $X \stackrel{d}{\sim} N_p(\boldsymbol{0}, \Sigma)$ with $\Sigma \in \mathbb{R}^{p \times p}$ being positive definite and $\boldsymbol{\mu}$ and $V$ satisfy

align[align omitted — 196 chars of source]

with $\boldsymbol{s} \in \mathbb{R}^p$ making $V$ positive definite. This way to construct knockoffs is most common in the literature (for example Barber2019,Candes2016,Barber2015,Xie2019).

We consider the following penalized estimator for logistic regression which is augmented by the knockoffs

align[align omitted — 338 chars of source]

where $\mathcal{L}[\ensuremath{\boldsymbol{\beta}^*},[X \ \tilde{\ensuremath{X}}], \ensuremath{\boldsymbol{y}}] =\sum_{i=1}^{n} (\log(1+\exp([\ensuremath{\boldsymbol{x}} \ \tilde{\ensuremath{\boldsymbol{x}}}]_i^{\top} \ \ensuremath{\boldsymbol{\beta}^*} ))- y_i [\ensuremath{\boldsymbol{x}} \ \tilde{\ensuremath{\boldsymbol{x}}}]_i^{\top} \ \ensuremath{\boldsymbol{\beta}^*}) /n $ is the negative log-likelihood function, $[X \ \tilde{\ensuremath{X}}] \in \mathbb{R}^{n \times 2p}$ is an augmented matrix, $r \in [0,\infty)$ is a tuning parameter, and $h:\mathbb{R}^{2p} \to [0,\infty]$ is a prior function which we specify as sparsity inducing prior function

align*[align* omitted — 82 chars of source]

We use Cross-Validation (CV) and a recently introduced novel calibration scheme Li2019 for $\ell_1$-penalized logistic regression for calibrating the tuning parameter in equation (ref). The latter calibration scheme is based on simple tests along the tuning parameter path and is equipped with finite sample guarantees for feature selection.

We consider here the Lasso Signed Max (LSM) Barber2015 and the Lasso coefficient-difference (LCD) Candes2016 as knockoff statistics. The LSM statistic denotes the maximum penalty coefficients of each variable entering in the model in (ref). The LCD statistic denotes the difference between the absolute values of the Lasso coefficients of the original variables and the knockoff copies. Denote the LSM and LCD statistics by $(Z_1,\dots,Z_p, \tilde{Z}_1,\dots, \tilde{Z}_p)$, that is,

align*[align* omitted — 583 chars of source]

for $j \in \{1,\dots, p\}$. Then the LSM and LCD statistic vectors $W:=(W_1,\dots,W_p)^{\top}$ can be defined by

align*[align* omitted — 267 chars of source]

for $j \in \{1, \dots, p \}$.

Then, the data-dependent thresholds of the knockoff and knockoff$^+$ filter methods (two types of knockoff procedures) for a given FDR target $q\in [0,1] $ are defined by

align*[align* omitted — 491 chars of source]

Thus, the corresponding estimated support recovery sets are defined as

align*[align* omitted — 306 chars of source]

The estimated support $\hat{\mathcal{S}}_q^+[X, \ensuremath{\boldsymbol{y}}]$ obtained by the knockoff$^+$ procedure satisfies the inquality in (ref), whereas $\hat{\mathcal{S}}_q[X, \ensuremath{\boldsymbol{y}}]$ satisfies this theoretical bound for an approximated FDR which is less or equal to the FDR.

We use the (A)FDR control methods for variable selection in our highdimensional empirical application. To apply (A)FDR control to our data, we use the knockoff filter for logistic regression. In line with the recommendations in Xie2019 we simulate $k=3$ independent knockoffs $\tilde{\boldsymbol{X}}^1,\tilde{\boldsymbol{X}}^2, \tilde{\boldsymbol{X}}^3$ from a Gaussian distribution for the AFDR control method. As target FDR, we choose $q=0.1$.

Data

Our analysis draws on employment records from German administrative data. In particular, we use the Sample of Integrated Employment Biographies (SIAB).\footnote{The data basis of this project is the Scientific Use File (SUF) of the SIAB (version SIAB-Regionalfile 1975 - 2014 Ganzer2017). The data was assessed via a guest stay at the Research Data Centre (FDZ) of the German Federal Employment Agency at the Institute for Employment Research (IAB) and then via controlled data processing at the FDZ.} It is provided by the Institut für Arbeitsmarkt- und Berufsforschung (IAB) in Nuremberg, and draws a 2 % random sample from the population of Integrated Employment Biographies (IEB). The source of the data set comes from the registration procedure for social insurance. The SIAB has been used in a number of studies on career and wage outcomes of women (for example Adda2017,Ejrnaes2013).

The sample of the IEB considered here covers all employees in Germany that contributed to the social security system between 1975 and 2014 or receiving transfer payments from the labor agency, or being registered job seekers. Civil servants, including teachers and self-employed, are not included in the sample. The sample provides detailed daily information on, for example, earnings, occupations, employment, tansitions in and out of work, periods of unemployment, unemployment and welfare benefits as well as basic demographics. An important advantage of the data set is that, because of the large sample size and the longitudinal nature of the administrative data, employment and wage histories are measured precisely.

Empirical data preparation

Our goal is to discover which variables from employment records and basic demographics are truely associated with professional careers of women. Due to methodological reasons, we concentrate on one particular year: 2010.\footnote{The proposed FDR pipeline requires the design matrix X to have independent and identically distributed rows.} As robustness check, we do the same analysis for two further years: 2009 and 2011. We focus on West-Germany and restrict our sample to birth cohorts of women between 1965 and 1975. The birth cohort restriction corresponds to the age groups between 35 and 45 years. Before age 35 many women still climb the career ladder, and have not yet reached the final top salary, while few women reach their career peak at an age older than 45. Because we do not predict when women will reach their professional career for the first time, we also leave women who already had a professional career before 2010 in the sample. Finally, we drop those women who do not have any employment, benefit, or job search record in the data either one or two years before the prediction year 2010. The idea behind this restriction is that we want to avoid having women in the sample who have emigrated from Germany or have died, and thus cannot have a career recorded in our data. Overall, the final sample includes middle-aged West German women with and without German citizenship, and with some recent attachement to the labor market. The sample consists of $N_1= 1\,567$ women with a career in 2010 and $N_2=78\,214$ women with no career in 2010.

To discover which variables from employment records and basic demographics affect future professional careers in 2010, we generate a plethora of separate variables from our SIAB data to capture the available information on the wage and (un-)employment histories in a very detailed and flexible way. Especially the longitudinal nature of the administrative employment data allows to capture complex information on individual employment and earning histories (for example types of employment, wages, skill level) for the entire elapsed working live possibly interacting in many different ways. To accurately record time-varying information of the elapsed (un-)employment history, we consider five lags for each of the baseline variables. For time-constant information (for example education) we only consider one lag. Consequently, we use information from data spells until 2009. Employment records from 2010 cannot be used for our analysis because of endogeneity issues. Consider a woman making a professional career in November 2010. She has most likely find out about her upcoming promotion in contract negotiations with her employer a few months earlier, and her knowledge of the career peak may induce her to work less.

We consider 5 sociodemographic and 107 (un-)employment and wage history baseline predictor variables. The reason for including also sociodemographic variables that do not explicitly account for the employment and wage histories is that they are strongly associated with professional career outcomes. For instance, highly educated women are more likely to achieve a professional career than low educated women. This means, the skill level might be a powerful predictor for a professional career.

Finally, to select the main effects that are truely associated with professional careers, it is important to allow for a wide range of plausible interactions between the baseline variables. Thus, we interact the lagged terms of the associated baseline predictors with each other in many different ways, so that we end up with 327 sociodemographic, and employment and wage history variables. A full list of the included predictor variables can be found in the appendix. We only consider those interactions that make sense from an economic point of view, and that do not lead to perfect multicollinearity, that is $X^{\top}X$ is invertible. Moreover, our interacted time-varying variables come from the same time period because we want to avoid modeling complex dynamics between different predictors for the ease of the interpretation later on. We also interact time-constant variables with time-varying variables. Further, we assume that some of the predictor variables can be expressed as a linear combination of a number of basis functions of the original measured variables. Indeed, our two wage profile predictor variables can be expressed through the group of the wage variables. We also include non-linear effects for women's age and working experience by using second-order polynomials.

In our empirical application, we consider a large-scale design matrix $X \in \mathbb{R}^{79\,782 \times 327}$. Thus, we consider a setting in which both the sample size and the feature size are large, but the feature size is much smaller than the sample size ($n,p \gg 1, p \ll n$). Although we do not consider the case where $n,p \gg 1, p \approx n$ or $p \gg n$, and high-dimensional statistics become indispensable, our case ($p$ is large) is the typical scope of high-dimensional statistics.

The outcome variable $y_i$ is an indicator that takes the value 1 if the woman makes a professional career in 2010, otherwise 0. Making a professional career means, according to our definition, to achieve a high wage level. A difficulty of the SIAB data is that the wage variable is censored from above to ensure that high earners cannot be identified. Therefore, we define the binary career outcome variable based on the social security contribution ceiling imposed on the SIAB data\footnote{Note that the social security contribution ceiling varies for each year and between West- and East-Germnay.}: whenever the wage variable coincides with the the social security contribution ceiling, the career dummy takes the value 1, and 0 otherwise.

Descriptives

Table (ref) in the appendix provides detailed descriptive statistics of all baseline variables included in the analysis. Among the 112 baseline predictors, 32 are continuous (like for example age in years) and 80 are binary (like for example part-time employed). The predictors collect information on age, education, non-German citizenship, employment, part-time employment, marginal employment, daily earnings, different wage profiles, change of establishment, wage growth, experience, unemployment benefits, welfare benefits, registered job search and occupations.

table*[table* omitted — 2,662 chars of source]

Table (ref) provides the descriptive statistics of some key baseline variables. The table shows the descriptives separately for the women with and without a professional career in 2010. The summary statistics show that the majority of women with no career in 2010 has vocational training (79%), around 10% are tertiary educated, and 10% have no postsecondary degree. Unlike these, most of the women with a career in 2010 are highly educated (61%), followed by women with middle education (37%) and low education (0.5%). Our sample of women has an average age of 40 years in 2009. The average employment rate is 87% for women with no career in 2010 and 98% for women with a career in 2010. Almost all women with a professional career in 2010 are full-time employed (92%), followed by part-time working women (6%). None of these women are marginally employed. On the other hand, only 39% of women with no career in 2010 are full-time employed, followed by part-time (30%) and marginally (18%) employment. In our sample, around 4% (5%) of the women have non-German citizenship. The average daily wage among working women with a professional career in 2010 is around 163 EUR. However, for women without a professional career in 2010 it is only 50 EUR. Making a career is highly persistent, since 74% of women with a professional career in 2010 also had a career in 2009. In the case of women without a career in 2010, however, it is only 0.4%.

Results

How important is a woman's employment history for her professional career? To approach this question, we establish an FDR pipeline based on Barber2015,Candes2016,Xie2019 to identify among a plethora of predictors derived from employment records those ones that are most predictive for a successful career.

In particular, we want to increase the estimation accuracy by effectively identifying the important predictor variables and improve the model interpretability. For this purpose, we take into account a large set of variables derived from employment records, and apply a new aggregation scheme Xie2019 for controlling the FDR to the data of the SIAB. We use this FDR approach beyond multiple testing for controlled variable selection as introduced in Barber2015 and Candes2016. To achieve correct inference, we apply the knockoff-filter to the aggregation scheme which solves the controlled variable selection problem in such a generality and for a wide variety of test-statistics that it is attractive to introduce it in the field of economics.

We assume that our data generating process can be modeled precisely by using only a small, but a priori unknown number of predictor variables. Therefore we include all predictors to search for the truely important ones with the FDR pipeline.

Main findings

Our main findings suggest that our data generating process can indeed be described by a sparse representation, since it improves the prediction performance and model interpretability. In particular, the (A)FDR methods show an improvement in terms of model interpretability and estimation accuracy compared to conventional variable selection methods such as the LASSO. In sum, we find that the employment status, working experience, daily wages, and wage increases in combination with a high level of education are truely associated with professional careers.

Prediction and variable selection performance

Our objective is to establish the FDR pipeline for variable selection in modern big data applications in economics. In doing so, we apply the proposed pipeline to labor market data and identify those variables from individual employment histories that are truely important for professional careers. Since there are no ground truths available for our application, that is, the true support $\mathcal{S}:=\text{supp}[\boldsymbol{\beta}]$ is unverifiable, we cannot measure variable selection accuracy directly. However, the ground truth is almost never available in empirical research. Therefore, we infer the methods' performance from the number of selected variables and the prediction accuracy. We apply the different methods as described in the methodology section. In addition, we estimate conventional logistic regression with the full set of variables and without any variable (only intercept). These two latter estimations mark the two extreme situations in which all variables or no variable is relevant for predicting professional careers.

For the ease of interpretation, we typically seek methods that yield a sparse model with a small number of variables. Moreover, we search for models with small prediction errors (good fit to the data). The model size is here crucial since our primary goal is variable selection. In particular, we are interested in the variable selection and prediction performances of the $l_1$-regularized logistic regression with 10-fold Cross-Validation refit. The prediction error of the 10-fold-CV with refit is defined as $\text{pred.~error}:= \frac{1}{10n^{\mathfrak{v}}}\sum_{j=1}^{10} \mathcal{L}[\boldsymbol{y}_{\mathcal{V}_j},X_{\mathcal{\hat{S}}_{\mathcal{V}_j}}\hat{\boldsymbol{\beta}}_{\text{refit}}[\boldsymbol{y}_{\mathcal{T}_j}, X_{\mathcal{\hat{S}}_{\mathcal{T}_j}}]]$, where $\mathcal{L}[\cdot]$ is the negative log-likelihood function, $X_{\mathcal{\hat{S}}}$ is the restricted design matrix after screening the non-zero coordinates with the proposed variable selection methods, $\hat{\boldsymbol{\beta}}_{\text{refit}}$ are the refitted estimators, and $\mathcal{T}_j:=\{1,\dots,n\} \backslash \mathcal{A}_j$ and $\mathcal{V}_j:= \mathcal{A}_j$, $j \in \{1, \dots, 10\}$ are our 10 training and validation sets, with $\mathcal{A}_1, \dots, \mathcal{A}_{10}$ having equal cardinality of $n/10$.

table*[table* omitted — 971 chars of source]

Table (ref) reports the model sizes and the prediction errors of the 10-fold-CV with refit. First, it is noticable that the empty model (only intercept) has a large prediction error. This means that the employment and wage history variables are indeed able to improve the prediction quality and are correlated with the professional careers. On the other hand, the large-scale model with the full set of variables does not lead to a better fit of the data than the models selected by our proposed FDR pipeline. Thus, it is profitable to search for the main variables which improve the model fit and model interpretability. Second, we observe that the LASSO method selects a considerably larger model than the (A)FDR control methods. This is expected, in view of LASSO-type estimators calibrated by cross-validation being designed for prediction and (A)FDR being designed for variable selection. What is added to the aggregated knockoff filter over the usual LASSO are the data-driven test statistics to control the FDR. This additional feature typically results in a larger increase in accuracy after refitting. This can actually be observed in Table (ref) that reports a slightly smaller refit error for the (A)FDR methods. Third, by comparing the FDR methods with each other it seems that the (A)FDR control with CV is slightly superior to the other methods in terms of model size and prediction accuracy. In contrast, (A)FDR control with the testing-based calibration is not usuable in our context since at least one variable should be selected.

Overall, no (A)FDR control method is significantly dominating in all measures, so that in empirical work the two aspects have to be weighted according to the primary goal. In our application, the model size is crucial since the primary objective is accurate variable selection, that is, the estimated support $\hat{\mathcal{S}}$ should be a good approximation of the true support $\mathcal{S}$. Below, we focus on the AFDR control methods for two main reasons. First, Xie2019 have shown in simulations that the aggregated knockoff filter can simultanously decrease the FDR and increase the power, while maintaining the original method's theoretical FDR guarantees. Second, we are interested in identifying a medium-sized model, because a model with only a few predictor variables leaves too little room for interpretation on the one hand, and on the other hand a model that is too large is difficult to interpret and harbors the risk of noise variables. Our results in Table (ref) show that the AFDR method selects a medium-sized model regardless of the knockoff statistics chosen.

Next, we provide the variables selected by the different applied methods. Table (ref) in the appendix contains the results for the six (A)FDR control estimations and Table (ref) in the appendix displays the results for the LASSO estimation. The upper panel of Table (ref) displays the results for the (A)FDR control method based on the LSM statistic, whereas the two lower panels display the results based on the LCD statistic.

We observe that variables that appear multiple times across methods are wage profile 1, employed${}_i$, marginal employed$_3$, daily wage${}_i$, strong positive wage growth$_2$, change of establishment${}_1$, \texttt{high educated}$_1$ $\times$ \texttt{change of establishment}${}_1$, \texttt{high educated}$_1$ $\times$ \texttt{part-time employed}$_j$, \texttt{high educated}$_1$ $\times$ \texttt{marginal employed}$_l$, \texttt{high educated}$_1$ $\times$ \texttt{daily wage}$_i$, \texttt{high educated}$_1$ $\times$ \texttt{strong positive wage growth}$_i$, with $i\in\{1,2,3,4,5\}$, $j\in\{1,5\}$, and $l\in\{2,3\}$. We can observe that the selected variables form two main groups: either predictors related to wages or predictors related to the employment status are selected. On the other hand, predictors related to unemployment benefits and welfare benefits are not selected by the methods at all and therefore are not associated with professional careers.

This finding is in line with our intuition: Daily wages and daily wages (or positive wage growths) combined with a high level of education are generally considered as important determinants for a professional career. Our proposed FDR pipeline confirms this assumption, since all delays of the above mentioned variables are selected. The fact that several delays of the employment variable are selected has to do with the career prediction signal inherent in this variable. A continuous employment is almost indispensable for making a successful career. Finally, we observe that in the majority of cases where two baseline variables interact with each other, the high level of education interacts with the variables from the employment and wage history. This is also not suprisingly, given that a high educational level is often associated with a successful career. A closer look at these interactions reveals that the high level of education interacts with the variables that have already been selected as meaningful baseline predictors.

Tables (ref) and (ref) in the appendix contain the selected variables of the robustness checks. To check the robustness, we carried out the same analysis for the LSM knockoff statistics for the years 2009 and 2011. The results show that, although fewer variables are selected overall, the same variables as in the 2010 sample are selected.

Overall, our results suggest that the proposed FDR pipeline is very useful in practice, especially if the researcher is faced with huge amounts of data. A distinctive feature of the FDR pipeline from conventional econometric tools is that it teases apart important from irrelevant variables while guaranteeing type I error control. Thus, the FDR approach enable economists to limit the number of covariates in a data-driven manner rather than by hand, in a way that the interpretability of economic models is preserved and correct inference is achieved. Therefore, we recommend economists who handle large amounts of data to use modern high-dimensional tools for variable selection. Due to the shift in the scale of economic data towards Big Data these high-dimensional models will become a new important toolbox in empirical economic reasearch alongside the conventional econometric toolbox. Having said this, a reasonable question is why should economists include so many covariates in the analysis? The answer is twofold. First, because we can due to the newly available large-scale administrative data sets. Second, and more importantly, even though economists may believe that a certain economic outcome depends on a small fraction of all available variables, we have a priori no idea which ones are the truely important ones.

Inference

In the previous paragraph, we showed that the (A)FDR control, which was originally designed for multiple testing, is extremely useful for variable selection in empirical applications where a plethora of predictor variables is available. The final goal of variable selection with (A)FDR control is to draw correct inferences about the estimates of the selected variables. Therefore, in this paragraph we focus on the inference and interpretation of the estimates. Since the $l_1$ penalty sets a certain fraction of parameters exactly to zero and favors small estimates, this can lead to an overall unwanted shrinkage of the estimates. To remove such biases, we refit the penalized estimators with subsequent logistic regression and least-squares estimation on the support. The advantage of this two-step procedure is that least-squares provides an accurate and unbiased estimator in low-dimensional and correctly specified models which can be interpreted straightforward. In addition, we calculate marginal effects of the logistic regression estimates to compare them with the respective least-squares estimates.

In Table (ref) we provide the estimation results for refitting AFDR $\text{LCD}_{\text{CV}}$. Tables (ref), (ref) and (ref) in the appendix report the corresponding estimation results for refitting AFDR LSM, FDR $\text{LCD}_{\text{CV}}$ and FDR LSM. The respective tables provide the OLS estimates (with refit), the logistic regression estimates (with refit) and the corresponding marginal effects.

A first look at the results shows that the estimates are robust against the respective specification, and that most estimates are statistically significant. The OLS and logistic regression results in columns 3 and 5 are very similar. For the sake of simplicity, we only interpret the statistically significant OLS estimates in the second column of Table (ref), since only the continuous variables are standardized in this column, so that a reasonable interpretation of the binary variables is possible. In addition, we only give an exemplary interpretation of one lagged term from the respective variable group.

table*[table* omitted — 2,044 chars of source]

Starting to interpret the OLS estimate on $\texttt{daily wage}_1$, we find that, by holding all other factors constant (h.a.o.f.c.), increasing the daily wage in 2009 by one standard deviation leads on average to a 4.0 standard deviation higher probability of making a professional career in 2010. It is striking that all lagged terms of the daily wage have a significant positive impact on the professional career. Continuing with the next group of selected variables that has statistically significant estimates, we find that being employed in 2008 decreases the probability of making a professional career in 2010 on average by 1.6 percentage points compared of being non-employed in 2008 (h.a.o.f.c.). Moreover, we find that having a strong negative wage growth in 2009 increases the probability of making a successful career in 2010 on average by 1.4 percentage points compared of having a normal or a positive wage growth in 2009. Finally, we give the interpretation on the interaction term $\texttt{experience}_1 \times \texttt{employed}_1$: h.a.o.f.c., women who are employed in 2009 have on average a 0.1 standard deviation lower probability of making a professional career in 2010 than women who are unemployed in 2009 with every additional standard deviation experience.

Table (ref) in the appendix gives further important insights regarding the interpretation of the OLS estimates of $\mathcal{\hat{S}}_{\text{AFDR}}$ with the $\text{LSM}$ statistic. We find that being employed in 2009 increases the probability of making a professional career in 2010 on average by 1.2 percentage points compared of being non-employed in 2009 (h.a.o.f.c). We further find that being marginal employed in 2007 decreases the probability of making a professional career in 2010 on average by 0.86 percentage point compared of being full-time, part-time or unemployed in 2007. On the other hand, a steady wage growth over the past five years increases the probability of making a professional career on average by 0.55 percentage point (h.a.o.f.c.).

Conclusion

Based on an empirical application towards the labor market, this paper introduced the potentials of high-dimensional tools aimed to controlled variable selection to economists. More specifically, we applied a new aggregation scheme for FDR control Xie2019 to a high-dimensional logistic regression model, which teases apart important from irrelevant variables while maximizing the (statistical) power and guaranteeing Type I error control. So far, the aggreagated FDR control method has only been used in the context of high-dimensional linear regression models Xie2019. However, we were easily able to extent the aggregation scheme for the non-linear case by exploiting the framework of MX knockoffs Candes2016. The MX knockoff-filter effectively controls the aggregated false discovery rate and performs valid inference in our high-dimensional logistic regression model by mimicking the correlation structure found within the potential covariates.

In particular, we applied the MX knockoff-filter to the data from the Sample of Integrated Employment Biographies to discover which variables from individual employment histories are truely associated with female professional careers. We assumed that the binary career outcome variable can be modeled accurately by using a sparse representation of the covariates, but unlike the conventional economic literature, we were unsure about the identity of the relevant covariates and selected them in a data-driven manner. To reach a parsimonious model, we used the $l_1$ penalized technique to solve the logistic regression throughout our estimations.

Our main results suggest that our high-resolution logistic regression model can indeed be described by a sparse representation, since it improves the prediction performance and model interpretability. Indeed, the (A)FDR methods, which are designed for variable selection, show an improvement in terms of estimation accuracy and model interpretability compared to conventional variable selection methods such as the LASSO, which are designed for prediction. Overall, the relatively small subset of the employment history variables that are genuinely related to female professional careers includes the working experience, employment status, daily wage and wage increases in combination with a high level of education.

Our results provide new insights for the empirical work in the economic discipline. First, the high-dimensional tools presented here enable economists to fit high-dimensional models. Due to the newly available large-scale data-sets these high-dimensional models will become a new important workhorse in empirical economic research along the conventional econometrics toolbox. Second, based on the empirical application towards the labor market, we have shown that data-driven (A)FDR control, which was originally designed for multiple testing, is also useful for variable selection and correct inference in high-dimensional settings. This is particularly an important insight given the fact that there is less knowledge about variable selection and certainly not about (A)FDR control in the economic literature. Third, the tools for controlled variable selection introduced here enable economists to limit the number of explanatory variables in a data-driven fashion, in a way that the interpretability of economic models is preserved by searching for a sparse representation of the data generating process.